{
  "id": 524111,
  "title": "7th place solution",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/524111",
  "author_name": "Ryota",
  "post_date": "2024-08-04T16:18:26.299000",
  "votes": 15,
  "comment_count": 0,
  "views": 0,
  "content": "<p>First of all, I would like to express my gratitude to the hosts <a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> and the Kaggle staff for organizing this interesting competition. It was a tough competition with issues of leaks, but the competition's task were very interesting and it was a great learning experience. I would also like to thank the community for sharing so much in the Discussions, including the discovery of leaks. And thank you to my team members <a href=\"https://www.kaggle.com/nomorevotch\" target=\"_blank\">@nomorevotch</a>, <a href=\"https://www.kaggle.com/masatomatsui\" target=\"_blank\">@masatomatsui</a>, <a href=\"https://www.kaggle.com/rheinmetall\" target=\"_blank\">@rheinmetall</a>, I learned a lot from all of you.</p>\n<h1>[Summary]</h1>\n<ul>\n<li>We used various models, including LSTM, Transformer, Conv1D, and Squeezeformer. LSTM and Squeezeformer were particularly strong performers.</li>\n<li>Additional features based on domain knowledge contributed to improved accuracy.</li>\n<li>Training with MAE or SmoothL1Loss, followed by additional training with MSE, led to increased accuracy.</li>\n<li>For the ensemble, we used a weighted average with weights optimized by the Nelder-Mead method. (Public: 0.78560 / Private: 0.79080)</li>\n<li>In the ensemble, it was crucial to include a few strong single models rather than many models.</li>\n<li>It was important to speed up experimentation by not using HF's full data until the final week.</li>\n</ul>\n<h1>[Ryota's Part]</h1>\n<h3>Data Preparation</h3>\n<ul>\n<li>Use full low-res dataset from HF</li>\n<li>We sampled data at a 1/7 ratio from the period [0008-02, 0009-01], similar to the competition data, and used only 625,000 samples for validation.</li>\n<li>Use StandardScaler for scaling both input and target.</li>\n<li>Additional Features<ul>\n<li>Diff features calculated by taking the differences along the vertical axis</li>\n<li>Diff features calculated by taking the differences of the aforementioned diff features</li>\n<li>Relative humidity ratio</li>\n<li>Pressure difference</li>\n<li>Water vapor pressure</li>\n<li>Ice rate</li>\n<li>(lat, lon)<ul>\n<li>Due to concerns that this could be considered leakage, I finally did not use it, but it gave a slight improvement (~0.0002)</li></ul></li></ul></li>\n<li>Tried the following additional features calculated along the vertical axis, but they were ineffective<ul>\n<li>Moving statistics (mean, std, max, min, median)</li>\n<li>Lag features</li></ul></li>\n</ul>\n<h3>Model</h3>\n<table>\n<thead>\n<tr>\n<th>model type</th>\n<th>CV</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Transformer + LSTM</td>\n<td>0.78734</td>\n<td>0.78567</td>\n<td>0.78058</td>\n</tr>\n<tr>\n<td>LSTM</td>\n<td>0.78794</td>\n<td>0.78682</td>\n<td>0.78120</td>\n</tr>\n<tr>\n<td>Conv1D</td>\n<td>0.78635</td>\n<td>0.78301</td>\n<td>0.77506</td>\n</tr>\n</tbody>\n</table>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3845682%2Fe55bcb587d0c489b2eaf46632acb07c1%2Farchitecture.png?generation=1722785104431299&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>Input / Output<ul>\n<li>Repeat the scaler features in the sequence direction, and the shape is (batch, 60, 25)</li>\n<li>The output shape is (batch, 60, 14)<ul>\n<li>The scaler features are averaged across the entire sequence</li></ul></li></ul></li>\n<li>Get diff features<ul>\n<li>Calculate the aforementioned diff features in the forward method</li>\n<li>The shape is (batch, 60, 86)</li></ul></li>\n<li>Convolution Feature Extractor<ul>\n<li>Using 2 layers of convolution with Linear layers before and after</li></ul></li>\n<li>Positional Embedding<ul>\n<li>Same as sinusoidal positional encoding but used as learnable parameters</li></ul></li>\n<li>Transformer Encoder<ul>\n<li>PyTorch's Transformer Encoder</li></ul></li>\n<li>Bi-LSTM Block<ul>\n<li>Each LSTM layer is followed by a Linear layer, with skip_connections applied to each layer, similar to a Transformer Block</li></ul></li>\n<li>ResNet Block<ul>\n<li>Similar to ResNet, each block contains two convolutional layers with a skip connection to the input</li>\n<li>In the latter 7 blocks of the Conv1D, an inception-like structure is used, applying a bottleneck structure and multiple parallel convolutional layers with different kernel sizes (1, 3, 5, 7).</li>\n<li>Use SE-Block</li></ul></li>\n<li>Head<ul>\n<li>2 layers of Linear</li></ul></li>\n<li>Activation<ul>\n<li>ELU for Conv1D</li>\n<li>GELU for Transformer and LSTM</li>\n<li>ReLU for Head</li></ul></li>\n<li>Normalization<ul>\n<li>Batch Normalization for Conv1D</li>\n<li>Layer Normalization for Transformer and LSTM</li></ul></li>\n<li>No Dropout</li>\n</ul>\n<h3>Loss</h3>\n<ul>\n<li>MAE<ul>\n<li>MAE performed better than HuberLoss or MSELoss.</li></ul></li>\n<li>Mask target columns where the weight is 0 or is included in ptend_q0002_[12, 26]</li>\n<li>Fine-Tuning by MSE<ul>\n<li>This trick consistently led to an improvement of about 0.002</li></ul></li>\n</ul>\n<h3>Training</h3>\n<ul>\n<li>epoch<ul>\n<li>MAE : 13 epochs</li>\n<li>MSE(Fine-Tuning) : 5 epochs</li></ul></li>\n<li>optimizer<ul>\n<li>AdamW<ul>\n<li>lr=[5e-4, 5e-6]</li>\n<li>weight_decay=0.01</li></ul></li></ul></li>\n<li>scheduler<ul>\n<li>Cosine schedule with warmup</li></ul></li>\n</ul>\n<h3>Post-processing</h3>\n<ul>\n<li>Applied the post-processing described <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/502484\" target=\"_blank\">here</a> to ptend_q0002_[12, 26]</li>\n<li>As an additional post-processing, after calculating next_state_[q0002, q0003]_[all_levels] from the predictions, apply the above post-processing if the values are below threshold.<ul>\n<li>This led to an improvement of about 0.001 when using only the competition data, but there was a negligible improvement after using all the low-res data.</li></ul></li>\n</ul>\n<h3>Source Code</h3>\n<ul>\n<li>All code is <a href=\"https://github.com/nocchi1/kaggle-leap-7th-place-solution\" target=\"_blank\">here</a></li>\n</ul>\n<h1>[sqrt4kaido's part]</h1>\n<h3>Overview</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3845682%2Fc5662ed50363fadc2d1b5e4d6facdc6d%2FLEAP.png?generation=1722784983889720&amp;alt=media\" alt=\"\"></p>\n<h3>Validation</h3>\n<p>From the low-res data, I extracted 625,000 rows from the future period (February 2008 to January 2009) relative to the kaggle data and used them for validation.</p>\n<h3>Feature Engineering</h3>\n<p>For features with sequences, we used the following:</p>\n<pre><code>,\n,\n,\n,\n,\n,\n</code></pre>\n<p>For non-sequence features, we used all of them.<br>\nIn addition to the data, the following are calculated:</p>\n<ul>\n<li>dp: Pressure difference</li>\n<li>RH: Relative humidity</li>\n<li>vp: Vapor pressure</li>\n<li>state_ice_rate: Ratio of ice in cloud water content (water + ice)</li>\n<li>ice_rate_diff: Difference between ice Ratio derived from temperature and state_ice_rate</li>\n</ul>\n<p>After adding the above features, standard scaler is applied. Using max(1e-6, std) for the std.<br>\nThen, the following process is applied:</p>\n<ul>\n<li><p>Sequence features<br>\nShaped into (60, num_feature) form. Diff and diff of diff features (both in negative and positive directions) are added.</p></li>\n<li><p>Non-sequence features<br>\nRepeated 60 times to match the sequence features.</p></li>\n</ul>\n<p>In the end, we used 11*5 sequence features and 16 non-sequence features.</p>\n<h3>Models</h3>\n<p>I used 1D sequence models.</p>\n<ul>\n<li>SqueezeFormer: Refer to <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/460316\" target=\"_blank\">RNA 2nd solution</a></li>\n<li>LSTM</li>\n</ul>\n<p>Using SmoothL1Loss as the loss function. Loss is calculated only for the columns targeted in the test set.</p>\n<h3>Post-processing</h3>\n<ul>\n<li><p><a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/499896\" target=\"_blank\">Replacement</a><br>\n\"ptend_q0002_12 ~ ptend_q0002_28\" columns was replaced with \"state_q0002_12 ~ state_q0002_28\".</p></li>\n<li><p>Adjustment to ensure that percentage values do not fall below 0<br>\nAs state_q0002 and state_q0003 must always be non-negative, I adjusted ptend_q0002 and ptend_q0003 to ensure that the next time step's state_q0002 and state_q0003 would not fall below 0.</p></li>\n</ul>\n<h3>Other Notes</h3>\n<ul>\n<li>As a second stage, additional training with MSELoss boosts performance by about 0.002.</li>\n<li>When using low-res dataset, training is possible without loading all data into memory by using hdf5py.</li>\n<li>I had not used the full low-res dataset until the final week, conducting experiments using only Kaggle data. (Result: rank jump-up on the final day)</li>\n</ul>\n<h3>Score</h3>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>public</th>\n<th>private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>squeezeformer</td>\n<td>0.78511</td>\n<td>0.78056</td>\n</tr>\n<tr>\n<td>LSTM</td>\n<td>0.78094</td>\n<td>0.77629</td>\n</tr>\n</tbody>\n</table>\n<h1>[e-toppo's part]</h1>\n<h3>Feature Engineering</h3>\n<p>In addition to the data, the following are calculated:</p>\n<ul>\n<li>dp: Pressure difference</li>\n<li>RH: Relative humidity (Reference: <a href=\"https://www.science.org/doi/10.1126/sciadv.adj7250\" target=\"_blank\"><strong>Climate-invariant machine learning</strong></a>)</li>\n<li>vp: Vapor pressure</li>\n<li>state_ice_rate: Ratio of ice in cloud water content (water + ice)</li>\n<li>ice_rate_diff: Difference between ice Ratio derived from temperature and state_ice_rate</li>\n<li>q0005: q0002 + q0003</li>\n</ul>\n<h3>Models &amp; Training</h3>\n<ul>\n<li>Model: LSTM</li>\n<li>Loss: Smooth L1</li>\n<li>Auiliary Loss: ptend_RH</li>\n</ul>\n<h3>Post-processing</h3>\n<ul>\n<li>Adjustment to ensure that percentage values do not fall below 0 As state_q0002 and state_q0003 must always be non-negative, I adjusted ptend_q0002 and ptend_q0003 to ensure that the next time step's state_q0002 and state_q0003 would not fall below 0.</li>\n<li>Temperature Adjustment As state_q0003 must be 0 for temperatures above 274 degrees, ptend_0003 was adjusted to ensure this condition is met.</li>\n</ul>\n<h1>[Rheinmetall's Part]</h1>\n<h3>Data</h3>\n<p>use all ClimSim_low-res in Hugging Face.</p>\n<h3>Preprocessing</h3>\n<p>Apply StandardScaler to both features and targets.</p>\n<h3>Input</h3>\n<ul>\n<li>There are two types of features and targets, one with height dimension and the other with scalar quantity, respectively. Therefore, for both 556 dimensional features and 368 dimensional targets, we split them into sequence features and scalar features. </li>\n<li>No additional input features are created.</li>\n</ul>\n<h3>Validation</h3>\n<ul>\n<li>After random shuffling, make the tail 625,000 as valid data.</li>\n</ul>\n<h3>Model</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3845682%2F9d88d5d990b46a838414280c30313846%2FLEAP_model.png?generation=1722783589992901&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>Sequence features are embedded in a linear layer and then input to LSTM, while scalar features are also embedded in a linear layer and then used as input to LSTM as hidden states. </li>\n<li>The LSTM block has six layers, which was determined by a trade-off between model training time and accuracy.</li>\n</ul>\n<h3>Outputs</h3>\n<ul>\n<li>Since my model has two outputs, a sequence head and a scalar head, I reconstruct this in competition format, 368 dimensions.</li>\n</ul>\n<h3>Loss function</h3>\n<ul>\n<li>calculated based on MSEloss at each head, but multiplied by 0.1 for the scalar head loss. (to prioritize training on sequence heads)</li>\n</ul>\n<h3>Post-processing</h3>\n<ul>\n<li>Check all data, and if a non-negative column is negative, fill it with 0.</li>\n</ul>\n<h1>[Ensemble]</h1>\n<ul>\n<li>Ensemble method is weighted average of top 6 single models with weights optimized by the Nelder-Mead method.</li>\n<li>The Public best and Private best were the same submission. (Public : 0.78560 / Private : 0.79080)</li>\n<li>Including derivative models with lower single scores in the ensemble did not lead to improved accuracy. The key was to generate a strong single model</li>\n</ul>",
  "messages": [
    {
      "id": 2946755,
      "postDate": "2024-08-04T16:18:26.300Z",
      "content": "<p>First of all, I would like to express my gratitude to the hosts <a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> and the Kaggle staff for organizing this interesting competition. It was a tough competition with issues of leaks, but the competition's task were very interesting and it was a great learning experience. I would also like to thank the community for sharing so much in the Discussions, including the discovery of leaks. And thank you to my team members <a href=\"https://www.kaggle.com/nomorevotch\" target=\"_blank\">@nomorevotch</a>, <a href=\"https://www.kaggle.com/masatomatsui\" target=\"_blank\">@masatomatsui</a>, <a href=\"https://www.kaggle.com/rheinmetall\" target=\"_blank\">@rheinmetall</a>, I learned a lot from all of you.</p>\n<h1>[Summary]</h1>\n<ul>\n<li>We used various models, including LSTM, Transformer, Conv1D, and Squeezeformer. LSTM and Squeezeformer were particularly strong performers.</li>\n<li>Additional features based on domain knowledge contributed to improved accuracy.</li>\n<li>Training with MAE or SmoothL1Loss, followed by additional training with MSE, led to increased accuracy.</li>\n<li>For the ensemble, we used a weighted average with weights optimized by the Nelder-Mead method. (Public: 0.78560 / Private: 0.79080)</li>\n<li>In the ensemble, it was crucial to include a few strong single models rather than many models.</li>\n<li>It was important to speed up experimentation by not using HF's full data until the final week.</li>\n</ul>\n<h1>[Ryota's Part]</h1>\n<h3>Data Preparation</h3>\n<ul>\n<li>Use full low-res dataset from HF</li>\n<li>We sampled data at a 1/7 ratio from the period [0008-02, 0009-01], similar to the competition data, and used only 625,000 samples for validation.</li>\n<li>Use StandardScaler for scaling both input and target.</li>\n<li>Additional Features<ul>\n<li>Diff features calculated by taking the differences along the vertical axis</li>\n<li>Diff features calculated by taking the differences of the aforementioned diff features</li>\n<li>Relative humidity ratio</li>\n<li>Pressure difference</li>\n<li>Water vapor pressure</li>\n<li>Ice rate</li>\n<li>(lat, lon)<ul>\n<li>Due to concerns that this could be considered leakage, I finally did not use it, but it gave a slight improvement (~0.0002)</li></ul></li></ul></li>\n<li>Tried the following additional features calculated along the vertical axis, but they were ineffective<ul>\n<li>Moving statistics (mean, std, max, min, median)</li>\n<li>Lag features</li></ul></li>\n</ul>\n<h3>Model</h3>\n<table>\n<thead>\n<tr>\n<th>model type</th>\n<th>CV</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Transformer + LSTM</td>\n<td>0.78734</td>\n<td>0.78567</td>\n<td>0.78058</td>\n</tr>\n<tr>\n<td>LSTM</td>\n<td>0.78794</td>\n<td>0.78682</td>\n<td>0.78120</td>\n</tr>\n<tr>\n<td>Conv1D</td>\n<td>0.78635</td>\n<td>0.78301</td>\n<td>0.77506</td>\n</tr>\n</tbody>\n</table>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3845682%2Fe55bcb587d0c489b2eaf46632acb07c1%2Farchitecture.png?generation=1722785104431299&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>Input / Output<ul>\n<li>Repeat the scaler features in the sequence direction, and the shape is (batch, 60, 25)</li>\n<li>The output shape is (batch, 60, 14)<ul>\n<li>The scaler features are averaged across the entire sequence</li></ul></li></ul></li>\n<li>Get diff features<ul>\n<li>Calculate the aforementioned diff features in the forward method</li>\n<li>The shape is (batch, 60, 86)</li></ul></li>\n<li>Convolution Feature Extractor<ul>\n<li>Using 2 layers of convolution with Linear layers before and after</li></ul></li>\n<li>Positional Embedding<ul>\n<li>Same as sinusoidal positional encoding but used as learnable parameters</li></ul></li>\n<li>Transformer Encoder<ul>\n<li>PyTorch's Transformer Encoder</li></ul></li>\n<li>Bi-LSTM Block<ul>\n<li>Each LSTM layer is followed by a Linear layer, with skip_connections applied to each layer, similar to a Transformer Block</li></ul></li>\n<li>ResNet Block<ul>\n<li>Similar to ResNet, each block contains two convolutional layers with a skip connection to the input</li>\n<li>In the latter 7 blocks of the Conv1D, an inception-like structure is used, applying a bottleneck structure and multiple parallel convolutional layers with different kernel sizes (1, 3, 5, 7).</li>\n<li>Use SE-Block</li></ul></li>\n<li>Head<ul>\n<li>2 layers of Linear</li></ul></li>\n<li>Activation<ul>\n<li>ELU for Conv1D</li>\n<li>GELU for Transformer and LSTM</li>\n<li>ReLU for Head</li></ul></li>\n<li>Normalization<ul>\n<li>Batch Normalization for Conv1D</li>\n<li>Layer Normalization for Transformer and LSTM</li></ul></li>\n<li>No Dropout</li>\n</ul>\n<h3>Loss</h3>\n<ul>\n<li>MAE<ul>\n<li>MAE performed better than HuberLoss or MSELoss.</li></ul></li>\n<li>Mask target columns where the weight is 0 or is included in ptend_q0002_[12, 26]</li>\n<li>Fine-Tuning by MSE<ul>\n<li>This trick consistently led to an improvement of about 0.002</li></ul></li>\n</ul>\n<h3>Training</h3>\n<ul>\n<li>epoch<ul>\n<li>MAE : 13 epochs</li>\n<li>MSE(Fine-Tuning) : 5 epochs</li></ul></li>\n<li>optimizer<ul>\n<li>AdamW<ul>\n<li>lr=[5e-4, 5e-6]</li>\n<li>weight_decay=0.01</li></ul></li></ul></li>\n<li>scheduler<ul>\n<li>Cosine schedule with warmup</li></ul></li>\n</ul>\n<h3>Post-processing</h3>\n<ul>\n<li>Applied the post-processing described <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/502484\" target=\"_blank\">here</a> to ptend_q0002_[12, 26]</li>\n<li>As an additional post-processing, after calculating next_state_[q0002, q0003]_[all_levels] from the predictions, apply the above post-processing if the values are below threshold.<ul>\n<li>This led to an improvement of about 0.001 when using only the competition data, but there was a negligible improvement after using all the low-res data.</li></ul></li>\n</ul>\n<h3>Source Code</h3>\n<ul>\n<li>All code is <a href=\"https://github.com/nocchi1/kaggle-leap-7th-place-solution\" target=\"_blank\">here</a></li>\n</ul>\n<h1>[sqrt4kaido's part]</h1>\n<h3>Overview</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3845682%2Fc5662ed50363fadc2d1b5e4d6facdc6d%2FLEAP.png?generation=1722784983889720&amp;alt=media\" alt=\"\"></p>\n<h3>Validation</h3>\n<p>From the low-res data, I extracted 625,000 rows from the future period (February 2008 to January 2009) relative to the kaggle data and used them for validation.</p>\n<h3>Feature Engineering</h3>\n<p>For features with sequences, we used the following:</p>\n<pre><code>,\n,\n,\n,\n,\n,\n</code></pre>\n<p>For non-sequence features, we used all of them.<br>\nIn addition to the data, the following are calculated:</p>\n<ul>\n<li>dp: Pressure difference</li>\n<li>RH: Relative humidity</li>\n<li>vp: Vapor pressure</li>\n<li>state_ice_rate: Ratio of ice in cloud water content (water + ice)</li>\n<li>ice_rate_diff: Difference between ice Ratio derived from temperature and state_ice_rate</li>\n</ul>\n<p>After adding the above features, standard scaler is applied. Using max(1e-6, std) for the std.<br>\nThen, the following process is applied:</p>\n<ul>\n<li><p>Sequence features<br>\nShaped into (60, num_feature) form. Diff and diff of diff features (both in negative and positive directions) are added.</p></li>\n<li><p>Non-sequence features<br>\nRepeated 60 times to match the sequence features.</p></li>\n</ul>\n<p>In the end, we used 11*5 sequence features and 16 non-sequence features.</p>\n<h3>Models</h3>\n<p>I used 1D sequence models.</p>\n<ul>\n<li>SqueezeFormer: Refer to <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/460316\" target=\"_blank\">RNA 2nd solution</a></li>\n<li>LSTM</li>\n</ul>\n<p>Using SmoothL1Loss as the loss function. Loss is calculated only for the columns targeted in the test set.</p>\n<h3>Post-processing</h3>\n<ul>\n<li><p><a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/499896\" target=\"_blank\">Replacement</a><br>\n\"ptend_q0002_12 ~ ptend_q0002_28\" columns was replaced with \"state_q0002_12 ~ state_q0002_28\".</p></li>\n<li><p>Adjustment to ensure that percentage values do not fall below 0<br>\nAs state_q0002 and state_q0003 must always be non-negative, I adjusted ptend_q0002 and ptend_q0003 to ensure that the next time step's state_q0002 and state_q0003 would not fall below 0.</p></li>\n</ul>\n<h3>Other Notes</h3>\n<ul>\n<li>As a second stage, additional training with MSELoss boosts performance by about 0.002.</li>\n<li>When using low-res dataset, training is possible without loading all data into memory by using hdf5py.</li>\n<li>I had not used the full low-res dataset until the final week, conducting experiments using only Kaggle data. (Result: rank jump-up on the final day)</li>\n</ul>\n<h3>Score</h3>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>public</th>\n<th>private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>squeezeformer</td>\n<td>0.78511</td>\n<td>0.78056</td>\n</tr>\n<tr>\n<td>LSTM</td>\n<td>0.78094</td>\n<td>0.77629</td>\n</tr>\n</tbody>\n</table>\n<h1>[e-toppo's part]</h1>\n<h3>Feature Engineering</h3>\n<p>In addition to the data, the following are calculated:</p>\n<ul>\n<li>dp: Pressure difference</li>\n<li>RH: Relative humidity (Reference: <a href=\"https://www.science.org/doi/10.1126/sciadv.adj7250\" target=\"_blank\"><strong>Climate-invariant machine learning</strong></a>)</li>\n<li>vp: Vapor pressure</li>\n<li>state_ice_rate: Ratio of ice in cloud water content (water + ice)</li>\n<li>ice_rate_diff: Difference between ice Ratio derived from temperature and state_ice_rate</li>\n<li>q0005: q0002 + q0003</li>\n</ul>\n<h3>Models &amp; Training</h3>\n<ul>\n<li>Model: LSTM</li>\n<li>Loss: Smooth L1</li>\n<li>Auiliary Loss: ptend_RH</li>\n</ul>\n<h3>Post-processing</h3>\n<ul>\n<li>Adjustment to ensure that percentage values do not fall below 0 As state_q0002 and state_q0003 must always be non-negative, I adjusted ptend_q0002 and ptend_q0003 to ensure that the next time step's state_q0002 and state_q0003 would not fall below 0.</li>\n<li>Temperature Adjustment As state_q0003 must be 0 for temperatures above 274 degrees, ptend_0003 was adjusted to ensure this condition is met.</li>\n</ul>\n<h1>[Rheinmetall's Part]</h1>\n<h3>Data</h3>\n<p>use all ClimSim_low-res in Hugging Face.</p>\n<h3>Preprocessing</h3>\n<p>Apply StandardScaler to both features and targets.</p>\n<h3>Input</h3>\n<ul>\n<li>There are two types of features and targets, one with height dimension and the other with scalar quantity, respectively. Therefore, for both 556 dimensional features and 368 dimensional targets, we split them into sequence features and scalar features. </li>\n<li>No additional input features are created.</li>\n</ul>\n<h3>Validation</h3>\n<ul>\n<li>After random shuffling, make the tail 625,000 as valid data.</li>\n</ul>\n<h3>Model</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3845682%2F9d88d5d990b46a838414280c30313846%2FLEAP_model.png?generation=1722783589992901&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>Sequence features are embedded in a linear layer and then input to LSTM, while scalar features are also embedded in a linear layer and then used as input to LSTM as hidden states. </li>\n<li>The LSTM block has six layers, which was determined by a trade-off between model training time and accuracy.</li>\n</ul>\n<h3>Outputs</h3>\n<ul>\n<li>Since my model has two outputs, a sequence head and a scalar head, I reconstruct this in competition format, 368 dimensions.</li>\n</ul>\n<h3>Loss function</h3>\n<ul>\n<li>calculated based on MSEloss at each head, but multiplied by 0.1 for the scalar head loss. (to prioritize training on sequence heads)</li>\n</ul>\n<h3>Post-processing</h3>\n<ul>\n<li>Check all data, and if a non-negative column is negative, fill it with 0.</li>\n</ul>\n<h1>[Ensemble]</h1>\n<ul>\n<li>Ensemble method is weighted average of top 6 single models with weights optimized by the Nelder-Mead method.</li>\n<li>The Public best and Private best were the same submission. (Public : 0.78560 / Private : 0.79080)</li>\n<li>Including derivative models with lower single scores in the ensemble did not lead to improved accuracy. The key was to generate a strong single model</li>\n</ul>",
      "rawMarkdown": "First of all, I would like to express my gratitude to the hosts @jerrylin96 and the Kaggle staff for organizing this interesting competition. It was a tough competition with issues of leaks, but the competition's task were very interesting and it was a great learning experience. I would also like to thank the community for sharing so much in the Discussions, including the discovery of leaks. And thank you to my team members @nomorevotch, @masatomatsui, @rheinmetall, I learned a lot from all of you.\n\n# [Summary]\n- We used various models, including LSTM, Transformer, Conv1D, and Squeezeformer. LSTM and Squeezeformer were particularly strong performers.\n- Additional features based on domain knowledge contributed to improved accuracy.\n- Training with MAE or SmoothL1Loss, followed by additional training with MSE, led to increased accuracy.\n- For the ensemble, we used a weighted average with weights optimized by the Nelder-Mead method. (Public: 0.78560 / Private: 0.79080)\n- In the ensemble, it was crucial to include a few strong single models rather than many models.\n- It was important to speed up experimentation by not using HF's full data until the final week.\n\n# [Ryota's Part]\n### Data Preparation\n- Use full low-res dataset from HF\n- We sampled data at a 1/7 ratio from the period [0008-02, 0009-01], similar to the competition data, and used only 625,000 samples for validation.\n- Use StandardScaler for scaling both input and target.\n- Additional Features\n    - Diff features calculated by taking the differences along the vertical axis\n    - Diff features calculated by taking the differences of the aforementioned diff features\n    - Relative humidity ratio\n    - Pressure difference\n    - Water vapor pressure\n    - Ice rate\n    - (lat, lon)\n        - Due to concerns that this could be considered leakage, I finally did not use it, but it gave a slight improvement (~0.0002)\n- Tried the following additional features calculated along the vertical axis, but they were ineffective\n    - Moving statistics (mean, std, max, min, median)\n    - Lag features\n### Model\n| model type | CV | Public | Private |\n| --- | --- | --- | --- |\n| Transformer + LSTM | 0.78734 | 0.78567 | 0.78058 |\n| LSTM | 0.78794 | 0.78682 | 0.78120 |\n| Conv1D | 0.78635 | 0.78301 | 0.77506 |\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3845682%2Fe55bcb587d0c489b2eaf46632acb07c1%2Farchitecture.png?generation=1722785104431299&alt=media)\n\n- Input / Output\n    - Repeat the scaler features in the sequence direction, and the shape is (batch, 60, 25)\n    - The output shape is (batch, 60, 14)\n        - The scaler features are averaged across the entire sequence\n- Get diff features\n    - Calculate the aforementioned diff features in the forward method\n    - The shape is (batch, 60, 86)\n- Convolution Feature Extractor\n    - Using 2 layers of convolution with Linear layers before and after\n- Positional Embedding\n    - Same as sinusoidal positional encoding but used as learnable parameters\n- Transformer Encoder\n    - PyTorch's Transformer Encoder\n- Bi-LSTM Block\n    - Each LSTM layer is followed by a Linear layer, with skip_connections applied to each layer, similar to a Transformer Block\n- ResNet Block\n    - Similar to ResNet, each block contains two convolutional layers with a skip connection to the input\n    - In the latter 7 blocks of the Conv1D, an inception-like structure is used, applying a bottleneck structure and multiple parallel convolutional layers with different kernel sizes (1, 3, 5, 7).\n    - Use SE-Block\n- Head\n    - 2 layers of Linear\n- Activation\n    - ELU for Conv1D\n    - GELU for Transformer and LSTM\n    - ReLU for Head\n- Normalization\n    - Batch Normalization for Conv1D\n    - Layer Normalization for Transformer and LSTM\n- No Dropout\n\n### Loss\n- MAE\n    - MAE performed better than HuberLoss or MSELoss.\n- Mask target columns where the weight is 0 or is included in ptend_q0002_[12, 26]\n- Fine-Tuning by MSE\n    - This trick consistently led to an improvement of about 0.002\n\n### Training\n- epoch\n    - MAE : 13 epochs\n    - MSE(Fine-Tuning) : 5 epochs\n- optimizer\n    - AdamW\n        - lr=[5e-4, 5e-6]\n        - weight_decay=0.01\n- scheduler\n    - Cosine schedule with warmup\n\n### Post-processing\n- Applied the post-processing described [here](https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/502484) to ptend_q0002_[12, 26]\n- As an additional post-processing, after calculating next_state_[q0002, q0003]_[all_levels] from the predictions, apply the above post-processing if the values are below threshold.\n    - This led to an improvement of about 0.001 when using only the competition data, but there was a negligible improvement after using all the low-res data.\n\n### Source Code\n- All code is [here](https://github.com/nocchi1/kaggle-leap-7th-place-solution)\n\n# [sqrt4kaido's part]\n### Overview\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3845682%2Fc5662ed50363fadc2d1b5e4d6facdc6d%2FLEAP.png?generation=1722784983889720&alt=media)\n\n### Validation\nFrom the low-res data, I extracted 625,000 rows from the future period (February 2008 to January 2009) relative to the kaggle data and used them for validation.\n\n### Feature Engineering\nFor features with sequences, we used the following:\n```\n\"state_t\",\n\"state_q0001\",\n\"state_q0002\",\n\"state_q0003\",\n\"state_u\",\n\"state_v\",\n```\nFor non-sequence features, we used all of them.\nIn addition to the data, the following are calculated:\n- dp: Pressure difference\n- RH: Relative humidity\n- vp: Vapor pressure\n- state_ice_rate: Ratio of ice in cloud water content (water + ice)\n- ice_rate_diff: Difference between ice Ratio derived from temperature and state_ice_rate\n\nAfter adding the above features, standard scaler is applied. Using max(1e-6, std) for the std.\nThen, the following process is applied:\n\n- Sequence features\nShaped into (60, num_feature) form. Diff and diff of diff features (both in negative and positive directions) are added.\n\n- Non-sequence features\nRepeated 60 times to match the sequence features.\n\nIn the end, we used 11*5 sequence features and 16 non-sequence features.\n\n\n### Models\n\nI used 1D sequence models.\n- SqueezeFormer: Refer to [RNA 2nd solution](https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/460316)\n- LSTM\n\nUsing SmoothL1Loss as the loss function. Loss is calculated only for the columns targeted in the test set.\n\n### Post-processing\n- [Replacement](https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/499896)\n\"ptend_q0002_12 ~ ptend_q0002_28\" columns was replaced with \"state_q0002_12 ~ state_q0002_28\".\n\n- Adjustment to ensure that percentage values do not fall below 0\nAs state_q0002 and state_q0003 must always be non-negative, I adjusted ptend_q0002 and ptend_q0003 to ensure that the next time step's state_q0002 and state_q0003 would not fall below 0.\n\n### Other Notes\n- As a second stage, additional training with MSELoss boosts performance by about 0.002.\n- When using low-res dataset, training is possible without loading all data into memory by using hdf5py.\n- I had not used the full low-res dataset until the final week, conducting experiments using only Kaggle data. (Result: rank jump-up on the final day)\n\n### Score\n|model|public|private|\n|---|---|---|\n|squeezeformer|0.78511|0.78056| \n|LSTM|0.78094|0.77629|\n\n# [e-toppo's part]\n### Feature Engineering\nIn addition to the data, the following are calculated:\n-   dp: Pressure difference\n-   RH: Relative humidity (Reference: [**Climate-invariant machine learning](https://www.science.org/doi/10.1126/sciadv.adj7250))**\n-   vp: Vapor pressure\n-   state_ice_rate: Ratio of ice in cloud water content (water + ice)\n-   ice_rate_diff: Difference between ice Ratio derived from temperature and state_ice_rate\n-   q0005: q0002 + q0003\n### Models & Training\n-   Model: LSTM\n-   Loss: Smooth L1\n-   Auiliary Loss: ptend_RH\n### Post-processing\n-   Adjustment to ensure that percentage values do not fall below 0 As state_q0002 and state_q0003 must always be non-negative, I adjusted ptend_q0002 and ptend_q0003 to ensure that the next time step's state_q0002 and state_q0003 would not fall below 0.\n-   Temperature Adjustment As state_q0003 must be 0 for temperatures above 274 degrees, ptend_0003 was adjusted to ensure this condition is met.\n\n# [Rheinmetall's Part]\n### Data\nuse all ClimSim_low-res in Hugging Face.\n### Preprocessing\nApply StandardScaler to both features and targets.\n### Input\n- There are two types of features and targets, one with height dimension and the other with scalar quantity, respectively. Therefore, for both 556 dimensional features and 368 dimensional targets, we split them into sequence features and scalar features. \n- No additional input features are created.\n### Validation\n- After random shuffling, make the tail 625,000 as valid data.\n### Model\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3845682%2F9d88d5d990b46a838414280c30313846%2FLEAP_model.png?generation=1722783589992901&alt=media)\n- Sequence features are embedded in a linear layer and then input to LSTM, while scalar features are also embedded in a linear layer and then used as input to LSTM as hidden states. \n- The LSTM block has six layers, which was determined by a trade-off between model training time and accuracy.\n### Outputs\n- Since my model has two outputs, a sequence head and a scalar head, I reconstruct this in competition format, 368 dimensions.\n### Loss function\n- calculated based on MSEloss at each head, but multiplied by 0.1 for the scalar head loss. (to prioritize training on sequence heads)\n### Post-processing\n- Check all data, and if a non-negative column is negative, fill it with 0.\n\n# [Ensemble]\n- Ensemble method is weighted average of top 6 single models with weights optimized by the Nelder-Mead method.\n- The Public best and Private best were the same submission. (Public : 0.78560 / Private : 0.79080)\n- Including derivative models with lower single scores in the ensemble did not lead to improved accuracy. The key was to generate a strong single model",
      "votes": 15
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2946755": "First of all, I would like to express my gratitude to the hosts @jerrylin96 and the Kaggle staff for organizing this interesting competition. It was a tough competition with issues of leaks, but the competition's task were very interesting and it was a great learning experience. I would also like to thank the community for sharing so much in the Discussions, including the discovery of leaks. And thank you to my team members @nomorevotch, @masatomatsui, @rheinmetall, I learned a lot from all of you.\n\n# [Summary]\n- We used various models, including LSTM, Transformer, Conv1D, and Squeezeformer. LSTM and Squeezeformer were particularly strong performers.\n- Additional features based on domain knowledge contributed to improved accuracy.\n- Training with MAE or SmoothL1Loss, followed by additional training with MSE, led to increased accuracy.\n- For the ensemble, we used a weighted average with weights optimized by the Nelder-Mead method. (Public: 0.78560 / Private: 0.79080)\n- In the ensemble, it was crucial to include a few strong single models rather than many models.\n- It was important to speed up experimentation by not using HF's full data until the final week.\n\n# [Ryota's Part]\n### Data Preparation\n- Use full low-res dataset from HF\n- We sampled data at a 1/7 ratio from the period [0008-02, 0009-01], similar to the competition data, and used only 625,000 samples for validation.\n- Use StandardScaler for scaling both input and target.\n- Additional Features\n    - Diff features calculated by taking the differences along the vertical axis\n    - Diff features calculated by taking the differences of the aforementioned diff features\n    - Relative humidity ratio\n    - Pressure difference\n    - Water vapor pressure\n    - Ice rate\n    - (lat, lon)\n        - Due to concerns that this could be considered leakage, I finally did not use it, but it gave a slight improvement (~0.0002)\n- Tried the following additional features calculated along the vertical axis, but they were ineffective\n    - Moving statistics (mean, std, max, min, median)\n    - Lag features\n### Model\n| model type | CV | Public | Private |\n| --- | --- | --- | --- |\n| Transformer + LSTM | 0.78734 | 0.78567 | 0.78058 |\n| LSTM | 0.78794 | 0.78682 | 0.78120 |\n| Conv1D | 0.78635 | 0.78301 | 0.77506 |\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3845682%2Fe55bcb587d0c489b2eaf46632acb07c1%2Farchitecture.png?generation=1722785104431299&alt=media)\n\n- Input / Output\n    - Repeat the scaler features in the sequence direction, and the shape is (batch, 60, 25)\n    - The output shape is (batch, 60, 14)\n        - The scaler features are averaged across the entire sequence\n- Get diff features\n    - Calculate the aforementioned diff features in the forward method\n    - The shape is (batch, 60, 86)\n- Convolution Feature Extractor\n    - Using 2 layers of convolution with Linear layers before and after\n- Positional Embedding\n    - Same as sinusoidal positional encoding but used as learnable parameters\n- Transformer Encoder\n    - PyTorch's Transformer Encoder\n- Bi-LSTM Block\n    - Each LSTM layer is followed by a Linear layer, with skip_connections applied to each layer, similar to a Transformer Block\n- ResNet Block\n    - Similar to ResNet, each block contains two convolutional layers with a skip connection to the input\n    - In the latter 7 blocks of the Conv1D, an inception-like structure is used, applying a bottleneck structure and multiple parallel convolutional layers with different kernel sizes (1, 3, 5, 7).\n    - Use SE-Block\n- Head\n    - 2 layers of Linear\n- Activation\n    - ELU for Conv1D\n    - GELU for Transformer and LSTM\n    - ReLU for Head\n- Normalization\n    - Batch Normalization for Conv1D\n    - Layer Normalization for Transformer and LSTM\n- No Dropout\n\n### Loss\n- MAE\n    - MAE performed better than HuberLoss or MSELoss.\n- Mask target columns where the weight is 0 or is included in ptend_q0002_[12, 26]\n- Fine-Tuning by MSE\n    - This trick consistently led to an improvement of about 0.002\n\n### Training\n- epoch\n    - MAE : 13 epochs\n    - MSE(Fine-Tuning) : 5 epochs\n- optimizer\n    - AdamW\n        - lr=[5e-4, 5e-6]\n        - weight_decay=0.01\n- scheduler\n    - Cosine schedule with warmup\n\n### Post-processing\n- Applied the post-processing described [here](https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/502484) to ptend_q0002_[12, 26]\n- As an additional post-processing, after calculating next_state_[q0002, q0003]_[all_levels] from the predictions, apply the above post-processing if the values are below threshold.\n    - This led to an improvement of about 0.001 when using only the competition data, but there was a negligible improvement after using all the low-res data.\n\n### Source Code\n- All code is [here](https://github.com/nocchi1/kaggle-leap-7th-place-solution)\n\n# [sqrt4kaido's part]\n### Overview\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3845682%2Fc5662ed50363fadc2d1b5e4d6facdc6d%2FLEAP.png?generation=1722784983889720&alt=media)\n\n### Validation\nFrom the low-res data, I extracted 625,000 rows from the future period (February 2008 to January 2009) relative to the kaggle data and used them for validation.\n\n### Feature Engineering\nFor features with sequences, we used the following:\n```\n\"state_t\",\n\"state_q0001\",\n\"state_q0002\",\n\"state_q0003\",\n\"state_u\",\n\"state_v\",\n```\nFor non-sequence features, we used all of them.\nIn addition to the data, the following are calculated:\n- dp: Pressure difference\n- RH: Relative humidity\n- vp: Vapor pressure\n- state_ice_rate: Ratio of ice in cloud water content (water + ice)\n- ice_rate_diff: Difference between ice Ratio derived from temperature and state_ice_rate\n\nAfter adding the above features, standard scaler is applied. Using max(1e-6, std) for the std.\nThen, the following process is applied:\n\n- Sequence features\nShaped into (60, num_feature) form. Diff and diff of diff features (both in negative and positive directions) are added.\n\n- Non-sequence features\nRepeated 60 times to match the sequence features.\n\nIn the end, we used 11*5 sequence features and 16 non-sequence features.\n\n\n### Models\n\nI used 1D sequence models.\n- SqueezeFormer: Refer to [RNA 2nd solution](https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/460316)\n- LSTM\n\nUsing SmoothL1Loss as the loss function. Loss is calculated only for the columns targeted in the test set.\n\n### Post-processing\n- [Replacement](https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/499896)\n\"ptend_q0002_12 ~ ptend_q0002_28\" columns was replaced with \"state_q0002_12 ~ state_q0002_28\".\n\n- Adjustment to ensure that percentage values do not fall below 0\nAs state_q0002 and state_q0003 must always be non-negative, I adjusted ptend_q0002 and ptend_q0003 to ensure that the next time step's state_q0002 and state_q0003 would not fall below 0.\n\n### Other Notes\n- As a second stage, additional training with MSELoss boosts performance by about 0.002.\n- When using low-res dataset, training is possible without loading all data into memory by using hdf5py.\n- I had not used the full low-res dataset until the final week, conducting experiments using only Kaggle data. (Result: rank jump-up on the final day)\n\n### Score\n|model|public|private|\n|---|---|---|\n|squeezeformer|0.78511|0.78056| \n|LSTM|0.78094|0.77629|\n\n# [e-toppo's part]\n### Feature Engineering\nIn addition to the data, the following are calculated:\n-   dp: Pressure difference\n-   RH: Relative humidity (Reference: [**Climate-invariant machine learning](https://www.science.org/doi/10.1126/sciadv.adj7250))**\n-   vp: Vapor pressure\n-   state_ice_rate: Ratio of ice in cloud water content (water + ice)\n-   ice_rate_diff: Difference between ice Ratio derived from temperature and state_ice_rate\n-   q0005: q0002 + q0003\n### Models & Training\n-   Model: LSTM\n-   Loss: Smooth L1\n-   Auiliary Loss: ptend_RH\n### Post-processing\n-   Adjustment to ensure that percentage values do not fall below 0 As state_q0002 and state_q0003 must always be non-negative, I adjusted ptend_q0002 and ptend_q0003 to ensure that the next time step's state_q0002 and state_q0003 would not fall below 0.\n-   Temperature Adjustment As state_q0003 must be 0 for temperatures above 274 degrees, ptend_0003 was adjusted to ensure this condition is met.\n\n# [Rheinmetall's Part]\n### Data\nuse all ClimSim_low-res in Hugging Face.\n### Preprocessing\nApply StandardScaler to both features and targets.\n### Input\n- There are two types of features and targets, one with height dimension and the other with scalar quantity, respectively. Therefore, for both 556 dimensional features and 368 dimensional targets, we split them into sequence features and scalar features. \n- No additional input features are created.\n### Validation\n- After random shuffling, make the tail 625,000 as valid data.\n### Model\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3845682%2F9d88d5d990b46a838414280c30313846%2FLEAP_model.png?generation=1722783589992901&alt=media)\n- Sequence features are embedded in a linear layer and then input to LSTM, while scalar features are also embedded in a linear layer and then used as input to LSTM as hidden states. \n- The LSTM block has six layers, which was determined by a trade-off between model training time and accuracy.\n### Outputs\n- Since my model has two outputs, a sequence head and a scalar head, I reconstruct this in competition format, 368 dimensions.\n### Loss function\n- calculated based on MSEloss at each head, but multiplied by 0.1 for the scalar head loss. (to prioritize training on sequence heads)\n### Post-processing\n- Check all data, and if a non-negative column is negative, fill it with 0.\n\n# [Ensemble]\n- Ensemble method is weighted average of top 6 single models with weights optimized by the Nelder-Mead method.\n- The Public best and Private best were the same submission. (Public : 0.78560 / Private : 0.79080)\n- Including derivative models with lower single scores in the ensemble did not lead to improved accuracy. The key was to generate a strong single model"
  }
}