{
  "id": 523055,
  "title": "2nd Place Solution",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/523055",
  "author_name": "Max2020",
  "post_date": "2024-07-30T01:26:46.566000",
  "votes": 47,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Thanks to the hosts <a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a>  and official kaggle team for organizing this interesting competition! Our team learned a lot from it!<br>\nI am very pleased to share our team's solution. The complete code is available on GitHub:  <a href=\"https://github.com/ChunhanLi/2nd-kaggle-LEAP\" target=\"_blank\">https://github.com/ChunhanLi/2nd-kaggle-LEAP</a></p>\n<h2>Summary</h2>\n<p>We do not have extensive domain knowledge in atmospheric forecasting; we simply framed the competition task as a 1D seq2seq multi-target regression task. We trained with various model structures and meticulously tuned them, ultimately using a hill-climbing algorithm for model ensembling based on validation set scores.</p>\n<p>The four highest-priority key points in the solution are summarized as follows:</p>\n<ul>\n<li><strong>Data size is super important in this competition.</strong> Using the entire <a href=\"https://huggingface.co/datasets/LEAP/ClimSim_low-res\" target=\"_blank\">low-resolution dataset</a> can significantly enhance performance.</li>\n<li><strong>The smoothl1 loss function, auxiliary diff loss and the cosine annealing learning rate schedule were adopted for model optimization.</strong></li>\n<li><strong>Group fine-tuning methods were introduced in the training pipeline.</strong></li>\n<li><strong>The team focused on developing diverse model structures and ensembling them.</strong></li>\n</ul>\n<h2>Solution</h2>\n<h3>Data and Cross validation</h3>\n<p><strong>Training data</strong>  <br>\nTraining data is crucial for this competition. According to this <a href=\"https://github.com/leap-stc/ClimSim\" target=\"_blank\">GitHub repository</a>, the Kaggle data is sampled from this <a href=\"https://huggingface.co/datasets/LEAP/ClimSim_low-res\" target=\"_blank\">low-resolution dataset</a>. Utilizing the entire low-resolution dataset can boost performance by almost 0.01 in this competition.</p>\n<p><strong>Cross validation</strong>  <br>\nOur models are trained using data from the first 7 years and the first half (January to June) of year 8, totaling approximately 75 million data points. We set the <code>stride_sample</code> parameter to 7 and used data from July to December of year 8 and January of year 9 as our hold-out validation dataset. This validation set has shown perfect correlation with the leaderboard scores.</p>\n<h3>Model optimization</h3>\n<p><strong>Smooth L1 loss</strong>  <br>\nThe Smooth L1 loss is a robust loss function used primarily in regression tasks, combining the advantages of L2 and L1 losses. It is less sensitive to outliers than L2 and avoids the non-differentiability at zero seen with L1. The function is defined as follows:</p>\n<div><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5724098%2F8c7335ec3f1f94f0f40e041832b28303%2Floss.png?generation=1722302507888184&amp;alt=media\" alt=\"jpg name\"></div>\n<p>This approach helps stabilize the training process and often improves performance in noisy data situations.</p>\n<div><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5724098%2F566356fb54e82e5218a83c0746d8de42%2Ff1.png?generation=1722302556372586&amp;alt=media\" alt=\"jpg name\"></div>\n<p><strong>Auxiliary diff loss</strong>  <br>\nWe use an auxiliary 'Diff Loss,' proposed by <a href=\"https://www.kaggle.com/zui0711\" target=\"_blank\">@zui0711</a> , to enhance learning. This loss calculates the difference between real and predicted values at adjacent levels of a target, using Smooth L1 loss to quantify the error. The approach is implemented as follows:</p>\n<pre><code> torch.no_grad():\n    outputs = model(inputs)\n    loss = criterion(outputs, labels)\n     i  ():\n        output_diff = outputs[:, +*i+:+*(i+)] - outputs[:, +*i:+*(i+)-]\n        label_diff = labels[:, +*i+:+*(i+)] - labels[:, +*i:+*(i+)-]\n        loss += criterion(output_diff, label_diff) / \n</code></pre>\n<p><strong>Cosine annealing scheduler</strong>  <br>\nThe cosine annealing schedule adjusts the learning rate in a cosine-shaped curve to help the model escape local minima and improve convergence. We implemented this with learning rate decays at the 3rd and 9th epochs.</p>\n<div><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5724098%2Fe965962b0d05d483bf9d1bf690014863%2Ff2.png?generation=1722302675640626&amp;alt=media\" alt=\"jpg name\"></div>\n<h3>Group fine-tune</h3>\n<p>During multi-objective learning tasks, different learning objectives can interact in complex ways, promoting or inhibiting each other. Our experiments found that while different target groups initially promoted each other, they began to interfere towards the end of training, potentially due to complex semantic constraints.</p>\n<p>Inspired by the <a href=\"https://www.kaggle.com/competitions/ventilator-pressure-prediction/discussion/285320\" target=\"_blank\">top solution from the 2021 VPP competition</a>, we divided 368 features into seven groups, six of which are series of measurements of different metrics along the atmospheric column, and one group consists of eight unique single targets. After the training process with 364 full outputs was completed, we fine-tuned these groups again. This allowed each model with different architectures to achieve an improvement ranging from 0.0005 to 0.0015. Due to time and resource constraints, we only fine-tuned each group for one epoch. <a href=\"https://www.kaggle.com/max2020\" target=\"_blank\">@max2020</a> </p>\n<h3>Post-processing</h3>\n<p>After obtaining raw model predictions, we denormalize them using the standard deviation and mean. Following community <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/499896#2791290\" target=\"_blank\">discussions</a>, we adjust certain variables of <em>ptend_q0002</em> based on their linear relationships. Finally, we apply the weight values from the official sample submission file to process the predictions.</p>\n<h2>Model structure</h2>\n<h3>Model by Forcewithme</h3>\n<p><strong>Table 1</strong></p>\n<table>\n<thead>\n<tr>\n<th>Model id</th>\n<th>pb</th>\n<th>fig</th>\n<th>Exp</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>forcewithme_reslstm_cv0.789_lb0.783</td>\n<td>0.78275</td>\n<td>Figure 3(a)</td>\n<td>18</td>\n</tr>\n<tr>\n<td>forcewithme_gf_reslstm_cv0.790_lb0.785</td>\n<td>0.7840</td>\n<td>Figure 3(a)</td>\n<td>32</td>\n</tr>\n<tr>\n<td>forcewithme_gf_cnnlstm_cv0.789_lb0.787</td>\n<td>0.7826</td>\n<td>Figure 3(b)</td>\n<td>39</td>\n</tr>\n<tr>\n<td>forcewithme_gf_lstmmamba_cv0.7885_lb0.7853</td>\n<td>0.7814</td>\n<td>Figure 3(c)</td>\n<td>40</td>\n</tr>\n<tr>\n<td>forcewithme_gf_LstmMambaMixed_cv0.7886_lb0.7858</td>\n<td>0.7821</td>\n<td>Figure 3(d)</td>\n<td>-</td>\n</tr>\n<tr>\n<td>forcewithme_gf037_1LSTM1mamba-5_state16_cv0.7896_LBunknown</td>\n<td>0.7830</td>\n<td>Figure 3(d)</td>\n<td>37</td>\n</tr>\n<tr>\n<td>forcewithme_gf038_2LSTM1mamba-3_state64_cv0.7897_LB0.787</td>\n<td>0.7836</td>\n<td>Figure 3(d)</td>\n<td>38</td>\n</tr>\n</tbody>\n</table>\n<div><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5724098%2F978e5b0a2e223d5024fb69cb1031e526%2Ff3.png?generation=1722301967906869&amp;alt=media\" alt=\"jpg name\"></div>\n<p>The final ensemble strategy contains 6 models of ForcewithMe. These 6 models, the private score, the corresponding architectures, and the corresponding \"exp id\" in table 2 are in the table above. The model architectures are easy to understand and implement with the corresponding figures and the code. Here are some details and insights:</p>\n<ul>\n<li><p>All models show local scores and public leaderboard (LB) scores within their model ids, which correspond directly to filenames, model names, and submission files (without extensions). The table supplements these with private LB scores and details on the model architectures.</p></li>\n<li><p>All models are primarily built with LSTM as main components. This is due to our finding that LSTMs significantly outperform other architectures such as CNNs, MHSA, NNs, and GRUs in this competition.</p></li>\n<li><p>Inspired by ResNet and Transformer, all my models incorporate residual connection. We believe residuals make models converge faster and perform better.</p></li>\n<li><p>Figure 3(a) depicts a simple ResLSTM structure, which surprisingly achieved the highest private LB score among all models. Moreover, It took over two weeks for my another submission to beat this residual LSTM model on both local score and on the public leaderboard. Additionally, ResLSTM is still better on the private LB. exp32 has the same architecture to exp 18. The former one resumes the weights of the latter one and applies group fine-tuning.</p></li>\n\n\n<li><p>Figure 3b illustrates a model combining small kernel convolutions, large kernel convolutions, and LSTMs as encoders, with the encoded results concatenated and fed into an LSTM backbone. This design leverages the presumed relationships between adjacent atmospheric layers, aiming to model local information and positional relations. Although it performs less well offline and on the private LB compared to the pure resLSTM model, it scores higher on the public LB and provides gains during the ensemble.</p></li>\n<li><p>The remaining four models combine LSTMs with MAMBA. While LSTMs typically have more layers, MAMBA occupies the same number of layers as the LSTMs in model <code>forcewithme_gf037_1LSTM1mamba-5_state16_cv0.7896_LBunknown</code>.</p></li>\n<li><p>The models mixing LSTM and Mamba (3d) in every block achieve very good scores offline and on the public LB.</p></li>\n<li><p>Despite individual MAMBA-based models not outperforming the pure LSTM structure, they play a significant role in the final ensemble.</p></li>\n<li><p>There are two versions of MAMBA available: MAMBA and MAMBA-2 in the \"mamba-ssm\" package. We exclusively used MAMBA.</p></li>\n<li><p>The default parameters for MAMBA are: <code>d_model=512</code>, <code>d_state=16</code>, <code>d_conv=4</code>, expand=2. Other than <code>d_model</code>, the parameter settings follow the official repository.</p></li>\n<li><p>The models <code>forcewithme_gf037_1LSTM1mamba-5_state16_cv0.7896_LBunknown</code> have 1 LSTM and 1 MAMBA in each block, and it has 5 blocks, the mamba uses default params setting. While the models <code>forcewithme_gf_LstmMambaMixed_cv0.7886_lb0.7858</code> and <code>forcewithme_gf038_2LSTM1mamba-3_state64</code> have 2 LSTMs and 1 mamba in each block, and they have 3 blocks.</p></li>\n<li><p>The models <code>forcewithme_gf_LstmMambaMixed_cv0.7886_lb0.7858</code> and <code>forcewithme_gf038_2LSTM1mamba-3_state64</code> differ only in that the latter has <code>d_state</code> set to <code>64</code>. They share figure 3(d) as they are almost the same.</p></li>\n<li><p>For installation and usage of MAMBA, please refer to the official repository: <a href=\"https://github.com/state-spaces/mamba?tab=readme-ov-file\" target=\"_blank\">https://github.com/state-spaces/mamba?tab=readme-ov-file</a>. It's licensed under Apache-2.0, allowing for free use, including commercial purposes, provided the license requirements are met.</p></li>\n</ul>\n<h3>Model by Max2020</h3>\n<p>In the Max2020 part, models 14, 15, 21, and 22 are all improvements based on the model by <a href=\"https://www.kaggle.com/forcewithme\" target=\"_blank\">@forcewithme</a>, integrating LSTM with skip connections. <strong>Model 22</strong> is our team’s highest-performing single model, providing the best results in local scoring, Leader Board scoring, and private scoring. Its architecture is shown in Figure 4. Regarding the learning rate schedule, a cosine decay learning rate was used, with decays occurring at 3 and 9 epochs. The loss function employed was the smooth L1 loss with a beta of 0.5. The Table 2 contains the detailed performance of my five models.</p>\n<div><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5724098%2F82ce64bb111824f2eee6c53954286c7a%2Ff4.png?generation=1722302228698175&amp;alt=media\" alt=\"jpg name\"></div>\n<h3>Model by Joseph Zhou</h3>\n<p>The structure of Joseph's models mainly consists of the use of multi-layers Res-ConvLSTM blocks and a TimeDistributed fully-connected layer. The input size is <code>[bs, 60, 25]</code> and output size is <code>[bs, 368]</code>, the last dimension of the input includes 16 sequence-features and 9 scalar-features. After passing the multi-layers Res-ConvLSTM blocks and fully-connected layer, it will result in a size of <code>[bs, 60, 14]</code>, the last dimension includes 6 sequence-labels and 8 scalar-labels. For scalar-labels part which is <code>[bs, 60, 8]</code>, we just average in the second axis to get the shape of <code>[bs, 8]</code>. For sequence-labels part which is <code>[bs, 60, 6]</code>, we reshape it to the shape of <code>[bs, 360]</code>. Finally, we concatenate the two parts to get the output. Besides, a reverse and shift type of augmentation is implemented in one of the models. The idea is we randomly reverse or shift the sequence-features and sequence-labels respectively in the data loader.</p>\n<h3>Model by Adam</h3>\n<p>The structure of Adam's models are like: Inputs -&gt; Position encoding -&gt; 1DCNN -&gt; 3-4 layers LSTM -&gt; 1 layer transformer -&gt; Outputs. Most models used all the low-resolution dataset and the others only use sampling data. The model trained by sampling data can contribute to the ensemble a bit. For the loss, Adam used SmoothL1 loss and auxiliary diff loss as mentioned before. For the scheduler, Adam used ReduceLROnPlateau scheduler with a factor of 0.2 and patience of 2.</p>\n<p>Adam's best single model has LB 0.78594 and PB 0.78141. The ensemble of Adam's models has LB 0.79050 and PB 0.78575.</p>\n<h3>Model by Zuiye</h3>\n<p>Zuiye's models are mainly based on two architectures. The first one consists of 2 LSTM layers followed by a MultiheadAttention layer. The other one consists of 3 parallel Convolutional layers with 3 different kernel sizes and next 2 LSTM layers followed by a MultiheadAttention layer just like the first architecture. Zuiye's best single model gets LB 0.78696 / PB 0.78205 and the ensemble of Zuiye's own models (with hill climb) gets LB 0.79050 / PB 0.78614.</p>\n<h3>Model ensembling</h3>\n<p>We use the hill climb method to search blend weights of each model. The steps of hill climb are the following: <a href=\"https://www.kaggle.com/hookman\" target=\"_blank\">@hookman</a> </p>\n<ol>\n<li>Take the best out-of-fold predictions as best_ensemble. This will be our baseline.</li>\n<li>Iteratively blend best_ensemble with different models with different weights, using the formula <code>new\\_ensemble = w * best\\_ensemble + (1-w) * new\\_oof</code></li>\n<li>Check the r2 score of the new ensemble. Choose the best new ensemble to replace best_ensemble as our new baseline.</li>\n<li>Repeat until the r2 score can't increase anymore or reaches the threshold.    </li>\n</ol>\n<p><strong>Weights of the best model are shown in the table below:</strong>    <br>\n<strong>Table 2</strong></p>\n<table>\n<thead>\n<tr>\n<th>Exp id</th>\n<th>weight</th>\n<th>cv</th>\n<th>lb</th>\n<th>pb</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>forcewithme_exp32</td>\n<td>0.166556</td>\n<td>0.790</td>\n<td>0.7865</td>\n<td>0.78398</td>\n</tr>\n<tr>\n<td>forcewithme_exp37</td>\n<td>0.158625</td>\n<td>0.7896</td>\n<td>0.78618</td>\n<td>0.78293</td>\n</tr>\n<tr>\n<td>forcewithme_exp38</td>\n<td>0.139194</td>\n<td>0.7897</td>\n<td>0.78719</td>\n<td>0.78362</td>\n</tr>\n<tr>\n<td>max_exp22</td>\n<td>0.120125</td>\n<td>0.7908</td>\n<td><strong>0.78793</strong></td>\n<td><strong>0.78434</strong></td>\n</tr>\n<tr>\n<td>Jo_exp912</td>\n<td>0.111971</td>\n<td>0.78935</td>\n<td>0.78528</td>\n<td>0.78150</td>\n</tr>\n<tr>\n<td>max_exp21</td>\n<td>0.104738</td>\n<td>0.7904</td>\n<td>0.78752</td>\n<td>0.78425</td>\n</tr>\n<tr>\n<td>forcewithme_exp39</td>\n<td>0.098977</td>\n<td>0.789</td>\n<td>0.78699</td>\n<td>0.78257</td>\n</tr>\n<tr>\n<td>max_exp14</td>\n<td>0.093088</td>\n<td>0.7905</td>\n<td>0.78641</td>\n<td>0.78214</td>\n</tr>\n<tr>\n<td>max_exp10</td>\n<td>0.092157</td>\n<td>0.7888</td>\n<td>0.78619</td>\n<td>0.78213</td>\n</tr>\n<tr>\n<td>forcewithme_exp40</td>\n<td>0.082941</td>\n<td>0.7885</td>\n<td>0.7853</td>\n<td>0.78261</td>\n</tr>\n<tr>\n<td>max_exp015</td>\n<td>0.052500</td>\n<td>0.7905</td>\n<td>0.78695</td>\n<td>0.78244</td>\n</tr>\n<tr>\n<td>adam_exp197</td>\n<td>0.048994</td>\n<td>0.7855</td>\n<td>0.78269</td>\n<td>0.777</td>\n</tr>\n<tr>\n<td>adam_exp200</td>\n<td>-0.047132</td>\n<td>0.7836</td>\n<td>0.78010</td>\n<td>0.77434</td>\n</tr>\n<tr>\n<td>adam_exp195</td>\n<td>-0.049875</td>\n<td>0.78569</td>\n<td>0.78334</td>\n<td>0.77753</td>\n</tr>\n<tr>\n<td>Jo_exp907</td>\n<td>-0.083779</td>\n<td>0.7855</td>\n<td>0.78289</td>\n<td>0.77873</td>\n</tr>\n<tr>\n<td>forcewithme_exp18</td>\n<td>-0.089079</td>\n<td>0.7890</td>\n<td>0.7863</td>\n<td>0.78272</td>\n</tr>\n<tr>\n<td><strong>Ensemble</strong></td>\n<td><strong>1.0</strong></td>\n<td><strong>0.7955</strong></td>\n<td><strong>0.79211</strong></td>\n<td><strong>0.78856</strong></td>\n</tr>\n</tbody>\n</table>\n<hr>\n<p>It's worth noting that we've provided the parameter weights files for all models on <a href=\"https://drive.google.com/drive/u/0/folders/1-1gavnxXqj2x6giAjPkPQcrTfpVE3fpC\" target=\"_blank\">Google Drive</a>. Using these for ensemble submissions results in slightly higher LB and PB scores, by a margin of 0.0002. The discrepancies stem from two main issues:</p>\n<p>Firstly, the <code>jo_exp907.pt</code> model file was missing, and the model had to be rerun post-competition, which led to some differences from the original. Secondly, during the competition, an incorrect model file was used under <code>forcewithme_exp18</code> (corresponding to the <code>forcewithme_reslstm_cv0.789_lb0.783</code> folder). This has now been corrected. These two points have caused a very minor difference in our final results. Although we believe this difference does not impact the reproducibility of our overall approach, we mention it here to avoid any confusion.</p>",
  "messages": [
    {
      "id": 2940309,
      "postDate": "2024-07-30T01:26:46.567Z",
      "content": "<p>Thanks to the hosts <a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a>  and official kaggle team for organizing this interesting competition! Our team learned a lot from it!<br>\nI am very pleased to share our team's solution. The complete code is available on GitHub:  <a href=\"https://github.com/ChunhanLi/2nd-kaggle-LEAP\" target=\"_blank\">https://github.com/ChunhanLi/2nd-kaggle-LEAP</a></p>\n<h2>Summary</h2>\n<p>We do not have extensive domain knowledge in atmospheric forecasting; we simply framed the competition task as a 1D seq2seq multi-target regression task. We trained with various model structures and meticulously tuned them, ultimately using a hill-climbing algorithm for model ensembling based on validation set scores.</p>\n<p>The four highest-priority key points in the solution are summarized as follows:</p>\n<ul>\n<li><strong>Data size is super important in this competition.</strong> Using the entire <a href=\"https://huggingface.co/datasets/LEAP/ClimSim_low-res\" target=\"_blank\">low-resolution dataset</a> can significantly enhance performance.</li>\n<li><strong>The smoothl1 loss function, auxiliary diff loss and the cosine annealing learning rate schedule were adopted for model optimization.</strong></li>\n<li><strong>Group fine-tuning methods were introduced in the training pipeline.</strong></li>\n<li><strong>The team focused on developing diverse model structures and ensembling them.</strong></li>\n</ul>\n<h2>Solution</h2>\n<h3>Data and Cross validation</h3>\n<p><strong>Training data</strong>  <br>\nTraining data is crucial for this competition. According to this <a href=\"https://github.com/leap-stc/ClimSim\" target=\"_blank\">GitHub repository</a>, the Kaggle data is sampled from this <a href=\"https://huggingface.co/datasets/LEAP/ClimSim_low-res\" target=\"_blank\">low-resolution dataset</a>. Utilizing the entire low-resolution dataset can boost performance by almost 0.01 in this competition.</p>\n<p><strong>Cross validation</strong>  <br>\nOur models are trained using data from the first 7 years and the first half (January to June) of year 8, totaling approximately 75 million data points. We set the <code>stride_sample</code> parameter to 7 and used data from July to December of year 8 and January of year 9 as our hold-out validation dataset. This validation set has shown perfect correlation with the leaderboard scores.</p>\n<h3>Model optimization</h3>\n<p><strong>Smooth L1 loss</strong>  <br>\nThe Smooth L1 loss is a robust loss function used primarily in regression tasks, combining the advantages of L2 and L1 losses. It is less sensitive to outliers than L2 and avoids the non-differentiability at zero seen with L1. The function is defined as follows:</p>\n<div><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5724098%2F8c7335ec3f1f94f0f40e041832b28303%2Floss.png?generation=1722302507888184&amp;alt=media\" alt=\"jpg name\"></div>\n<p>This approach helps stabilize the training process and often improves performance in noisy data situations.</p>\n<div><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5724098%2F566356fb54e82e5218a83c0746d8de42%2Ff1.png?generation=1722302556372586&amp;alt=media\" alt=\"jpg name\"></div>\n<p><strong>Auxiliary diff loss</strong>  <br>\nWe use an auxiliary 'Diff Loss,' proposed by <a href=\"https://www.kaggle.com/zui0711\" target=\"_blank\">@zui0711</a> , to enhance learning. This loss calculates the difference between real and predicted values at adjacent levels of a target, using Smooth L1 loss to quantify the error. The approach is implemented as follows:</p>\n<pre><code> torch.no_grad():\n    outputs = model(inputs)\n    loss = criterion(outputs, labels)\n     i  ():\n        output_diff = outputs[:, +*i+:+*(i+)] - outputs[:, +*i:+*(i+)-]\n        label_diff = labels[:, +*i+:+*(i+)] - labels[:, +*i:+*(i+)-]\n        loss += criterion(output_diff, label_diff) / \n</code></pre>\n<p><strong>Cosine annealing scheduler</strong>  <br>\nThe cosine annealing schedule adjusts the learning rate in a cosine-shaped curve to help the model escape local minima and improve convergence. We implemented this with learning rate decays at the 3rd and 9th epochs.</p>\n<div><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5724098%2Fe965962b0d05d483bf9d1bf690014863%2Ff2.png?generation=1722302675640626&amp;alt=media\" alt=\"jpg name\"></div>\n<h3>Group fine-tune</h3>\n<p>During multi-objective learning tasks, different learning objectives can interact in complex ways, promoting or inhibiting each other. Our experiments found that while different target groups initially promoted each other, they began to interfere towards the end of training, potentially due to complex semantic constraints.</p>\n<p>Inspired by the <a href=\"https://www.kaggle.com/competitions/ventilator-pressure-prediction/discussion/285320\" target=\"_blank\">top solution from the 2021 VPP competition</a>, we divided 368 features into seven groups, six of which are series of measurements of different metrics along the atmospheric column, and one group consists of eight unique single targets. After the training process with 364 full outputs was completed, we fine-tuned these groups again. This allowed each model with different architectures to achieve an improvement ranging from 0.0005 to 0.0015. Due to time and resource constraints, we only fine-tuned each group for one epoch. <a href=\"https://www.kaggle.com/max2020\" target=\"_blank\">@max2020</a> </p>\n<h3>Post-processing</h3>\n<p>After obtaining raw model predictions, we denormalize them using the standard deviation and mean. Following community <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/499896#2791290\" target=\"_blank\">discussions</a>, we adjust certain variables of <em>ptend_q0002</em> based on their linear relationships. Finally, we apply the weight values from the official sample submission file to process the predictions.</p>\n<h2>Model structure</h2>\n<h3>Model by Forcewithme</h3>\n<p><strong>Table 1</strong></p>\n<table>\n<thead>\n<tr>\n<th>Model id</th>\n<th>pb</th>\n<th>fig</th>\n<th>Exp</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>forcewithme_reslstm_cv0.789_lb0.783</td>\n<td>0.78275</td>\n<td>Figure 3(a)</td>\n<td>18</td>\n</tr>\n<tr>\n<td>forcewithme_gf_reslstm_cv0.790_lb0.785</td>\n<td>0.7840</td>\n<td>Figure 3(a)</td>\n<td>32</td>\n</tr>\n<tr>\n<td>forcewithme_gf_cnnlstm_cv0.789_lb0.787</td>\n<td>0.7826</td>\n<td>Figure 3(b)</td>\n<td>39</td>\n</tr>\n<tr>\n<td>forcewithme_gf_lstmmamba_cv0.7885_lb0.7853</td>\n<td>0.7814</td>\n<td>Figure 3(c)</td>\n<td>40</td>\n</tr>\n<tr>\n<td>forcewithme_gf_LstmMambaMixed_cv0.7886_lb0.7858</td>\n<td>0.7821</td>\n<td>Figure 3(d)</td>\n<td>-</td>\n</tr>\n<tr>\n<td>forcewithme_gf037_1LSTM1mamba-5_state16_cv0.7896_LBunknown</td>\n<td>0.7830</td>\n<td>Figure 3(d)</td>\n<td>37</td>\n</tr>\n<tr>\n<td>forcewithme_gf038_2LSTM1mamba-3_state64_cv0.7897_LB0.787</td>\n<td>0.7836</td>\n<td>Figure 3(d)</td>\n<td>38</td>\n</tr>\n</tbody>\n</table>\n<div><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5724098%2F978e5b0a2e223d5024fb69cb1031e526%2Ff3.png?generation=1722301967906869&amp;alt=media\" alt=\"jpg name\"></div>\n<p>The final ensemble strategy contains 6 models of ForcewithMe. These 6 models, the private score, the corresponding architectures, and the corresponding \"exp id\" in table 2 are in the table above. The model architectures are easy to understand and implement with the corresponding figures and the code. Here are some details and insights:</p>\n<ul>\n<li><p>All models show local scores and public leaderboard (LB) scores within their model ids, which correspond directly to filenames, model names, and submission files (without extensions). The table supplements these with private LB scores and details on the model architectures.</p></li>\n<li><p>All models are primarily built with LSTM as main components. This is due to our finding that LSTMs significantly outperform other architectures such as CNNs, MHSA, NNs, and GRUs in this competition.</p></li>\n<li><p>Inspired by ResNet and Transformer, all my models incorporate residual connection. We believe residuals make models converge faster and perform better.</p></li>\n<li><p>Figure 3(a) depicts a simple ResLSTM structure, which surprisingly achieved the highest private LB score among all models. Moreover, It took over two weeks for my another submission to beat this residual LSTM model on both local score and on the public leaderboard. Additionally, ResLSTM is still better on the private LB. exp32 has the same architecture to exp 18. The former one resumes the weights of the latter one and applies group fine-tuning.</p></li>\n\n\n<li><p>Figure 3b illustrates a model combining small kernel convolutions, large kernel convolutions, and LSTMs as encoders, with the encoded results concatenated and fed into an LSTM backbone. This design leverages the presumed relationships between adjacent atmospheric layers, aiming to model local information and positional relations. Although it performs less well offline and on the private LB compared to the pure resLSTM model, it scores higher on the public LB and provides gains during the ensemble.</p></li>\n<li><p>The remaining four models combine LSTMs with MAMBA. While LSTMs typically have more layers, MAMBA occupies the same number of layers as the LSTMs in model <code>forcewithme_gf037_1LSTM1mamba-5_state16_cv0.7896_LBunknown</code>.</p></li>\n<li><p>The models mixing LSTM and Mamba (3d) in every block achieve very good scores offline and on the public LB.</p></li>\n<li><p>Despite individual MAMBA-based models not outperforming the pure LSTM structure, they play a significant role in the final ensemble.</p></li>\n<li><p>There are two versions of MAMBA available: MAMBA and MAMBA-2 in the \"mamba-ssm\" package. We exclusively used MAMBA.</p></li>\n<li><p>The default parameters for MAMBA are: <code>d_model=512</code>, <code>d_state=16</code>, <code>d_conv=4</code>, expand=2. Other than <code>d_model</code>, the parameter settings follow the official repository.</p></li>\n<li><p>The models <code>forcewithme_gf037_1LSTM1mamba-5_state16_cv0.7896_LBunknown</code> have 1 LSTM and 1 MAMBA in each block, and it has 5 blocks, the mamba uses default params setting. While the models <code>forcewithme_gf_LstmMambaMixed_cv0.7886_lb0.7858</code> and <code>forcewithme_gf038_2LSTM1mamba-3_state64</code> have 2 LSTMs and 1 mamba in each block, and they have 3 blocks.</p></li>\n<li><p>The models <code>forcewithme_gf_LstmMambaMixed_cv0.7886_lb0.7858</code> and <code>forcewithme_gf038_2LSTM1mamba-3_state64</code> differ only in that the latter has <code>d_state</code> set to <code>64</code>. They share figure 3(d) as they are almost the same.</p></li>\n<li><p>For installation and usage of MAMBA, please refer to the official repository: <a href=\"https://github.com/state-spaces/mamba?tab=readme-ov-file\" target=\"_blank\">https://github.com/state-spaces/mamba?tab=readme-ov-file</a>. It's licensed under Apache-2.0, allowing for free use, including commercial purposes, provided the license requirements are met.</p></li>\n</ul>\n<h3>Model by Max2020</h3>\n<p>In the Max2020 part, models 14, 15, 21, and 22 are all improvements based on the model by <a href=\"https://www.kaggle.com/forcewithme\" target=\"_blank\">@forcewithme</a>, integrating LSTM with skip connections. <strong>Model 22</strong> is our team’s highest-performing single model, providing the best results in local scoring, Leader Board scoring, and private scoring. Its architecture is shown in Figure 4. Regarding the learning rate schedule, a cosine decay learning rate was used, with decays occurring at 3 and 9 epochs. The loss function employed was the smooth L1 loss with a beta of 0.5. The Table 2 contains the detailed performance of my five models.</p>\n<div><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5724098%2F82ce64bb111824f2eee6c53954286c7a%2Ff4.png?generation=1722302228698175&amp;alt=media\" alt=\"jpg name\"></div>\n<h3>Model by Joseph Zhou</h3>\n<p>The structure of Joseph's models mainly consists of the use of multi-layers Res-ConvLSTM blocks and a TimeDistributed fully-connected layer. The input size is <code>[bs, 60, 25]</code> and output size is <code>[bs, 368]</code>, the last dimension of the input includes 16 sequence-features and 9 scalar-features. After passing the multi-layers Res-ConvLSTM blocks and fully-connected layer, it will result in a size of <code>[bs, 60, 14]</code>, the last dimension includes 6 sequence-labels and 8 scalar-labels. For scalar-labels part which is <code>[bs, 60, 8]</code>, we just average in the second axis to get the shape of <code>[bs, 8]</code>. For sequence-labels part which is <code>[bs, 60, 6]</code>, we reshape it to the shape of <code>[bs, 360]</code>. Finally, we concatenate the two parts to get the output. Besides, a reverse and shift type of augmentation is implemented in one of the models. The idea is we randomly reverse or shift the sequence-features and sequence-labels respectively in the data loader.</p>\n<h3>Model by Adam</h3>\n<p>The structure of Adam's models are like: Inputs -&gt; Position encoding -&gt; 1DCNN -&gt; 3-4 layers LSTM -&gt; 1 layer transformer -&gt; Outputs. Most models used all the low-resolution dataset and the others only use sampling data. The model trained by sampling data can contribute to the ensemble a bit. For the loss, Adam used SmoothL1 loss and auxiliary diff loss as mentioned before. For the scheduler, Adam used ReduceLROnPlateau scheduler with a factor of 0.2 and patience of 2.</p>\n<p>Adam's best single model has LB 0.78594 and PB 0.78141. The ensemble of Adam's models has LB 0.79050 and PB 0.78575.</p>\n<h3>Model by Zuiye</h3>\n<p>Zuiye's models are mainly based on two architectures. The first one consists of 2 LSTM layers followed by a MultiheadAttention layer. The other one consists of 3 parallel Convolutional layers with 3 different kernel sizes and next 2 LSTM layers followed by a MultiheadAttention layer just like the first architecture. Zuiye's best single model gets LB 0.78696 / PB 0.78205 and the ensemble of Zuiye's own models (with hill climb) gets LB 0.79050 / PB 0.78614.</p>\n<h3>Model ensembling</h3>\n<p>We use the hill climb method to search blend weights of each model. The steps of hill climb are the following: <a href=\"https://www.kaggle.com/hookman\" target=\"_blank\">@hookman</a> </p>\n<ol>\n<li>Take the best out-of-fold predictions as best_ensemble. This will be our baseline.</li>\n<li>Iteratively blend best_ensemble with different models with different weights, using the formula <code>new\\_ensemble = w * best\\_ensemble + (1-w) * new\\_oof</code></li>\n<li>Check the r2 score of the new ensemble. Choose the best new ensemble to replace best_ensemble as our new baseline.</li>\n<li>Repeat until the r2 score can't increase anymore or reaches the threshold.    </li>\n</ol>\n<p><strong>Weights of the best model are shown in the table below:</strong>    <br>\n<strong>Table 2</strong></p>\n<table>\n<thead>\n<tr>\n<th>Exp id</th>\n<th>weight</th>\n<th>cv</th>\n<th>lb</th>\n<th>pb</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>forcewithme_exp32</td>\n<td>0.166556</td>\n<td>0.790</td>\n<td>0.7865</td>\n<td>0.78398</td>\n</tr>\n<tr>\n<td>forcewithme_exp37</td>\n<td>0.158625</td>\n<td>0.7896</td>\n<td>0.78618</td>\n<td>0.78293</td>\n</tr>\n<tr>\n<td>forcewithme_exp38</td>\n<td>0.139194</td>\n<td>0.7897</td>\n<td>0.78719</td>\n<td>0.78362</td>\n</tr>\n<tr>\n<td>max_exp22</td>\n<td>0.120125</td>\n<td>0.7908</td>\n<td><strong>0.78793</strong></td>\n<td><strong>0.78434</strong></td>\n</tr>\n<tr>\n<td>Jo_exp912</td>\n<td>0.111971</td>\n<td>0.78935</td>\n<td>0.78528</td>\n<td>0.78150</td>\n</tr>\n<tr>\n<td>max_exp21</td>\n<td>0.104738</td>\n<td>0.7904</td>\n<td>0.78752</td>\n<td>0.78425</td>\n</tr>\n<tr>\n<td>forcewithme_exp39</td>\n<td>0.098977</td>\n<td>0.789</td>\n<td>0.78699</td>\n<td>0.78257</td>\n</tr>\n<tr>\n<td>max_exp14</td>\n<td>0.093088</td>\n<td>0.7905</td>\n<td>0.78641</td>\n<td>0.78214</td>\n</tr>\n<tr>\n<td>max_exp10</td>\n<td>0.092157</td>\n<td>0.7888</td>\n<td>0.78619</td>\n<td>0.78213</td>\n</tr>\n<tr>\n<td>forcewithme_exp40</td>\n<td>0.082941</td>\n<td>0.7885</td>\n<td>0.7853</td>\n<td>0.78261</td>\n</tr>\n<tr>\n<td>max_exp015</td>\n<td>0.052500</td>\n<td>0.7905</td>\n<td>0.78695</td>\n<td>0.78244</td>\n</tr>\n<tr>\n<td>adam_exp197</td>\n<td>0.048994</td>\n<td>0.7855</td>\n<td>0.78269</td>\n<td>0.777</td>\n</tr>\n<tr>\n<td>adam_exp200</td>\n<td>-0.047132</td>\n<td>0.7836</td>\n<td>0.78010</td>\n<td>0.77434</td>\n</tr>\n<tr>\n<td>adam_exp195</td>\n<td>-0.049875</td>\n<td>0.78569</td>\n<td>0.78334</td>\n<td>0.77753</td>\n</tr>\n<tr>\n<td>Jo_exp907</td>\n<td>-0.083779</td>\n<td>0.7855</td>\n<td>0.78289</td>\n<td>0.77873</td>\n</tr>\n<tr>\n<td>forcewithme_exp18</td>\n<td>-0.089079</td>\n<td>0.7890</td>\n<td>0.7863</td>\n<td>0.78272</td>\n</tr>\n<tr>\n<td><strong>Ensemble</strong></td>\n<td><strong>1.0</strong></td>\n<td><strong>0.7955</strong></td>\n<td><strong>0.79211</strong></td>\n<td><strong>0.78856</strong></td>\n</tr>\n</tbody>\n</table>\n<hr>\n<p>It's worth noting that we've provided the parameter weights files for all models on <a href=\"https://drive.google.com/drive/u/0/folders/1-1gavnxXqj2x6giAjPkPQcrTfpVE3fpC\" target=\"_blank\">Google Drive</a>. Using these for ensemble submissions results in slightly higher LB and PB scores, by a margin of 0.0002. The discrepancies stem from two main issues:</p>\n<p>Firstly, the <code>jo_exp907.pt</code> model file was missing, and the model had to be rerun post-competition, which led to some differences from the original. Secondly, during the competition, an incorrect model file was used under <code>forcewithme_exp18</code> (corresponding to the <code>forcewithme_reslstm_cv0.789_lb0.783</code> folder). This has now been corrected. These two points have caused a very minor difference in our final results. Although we believe this difference does not impact the reproducibility of our overall approach, we mention it here to avoid any confusion.</p>",
      "rawMarkdown": "Thanks to the hosts @jerrylin96  and official kaggle team for organizing this interesting competition! Our team learned a lot from it!\nI am very pleased to share our team's solution. The complete code is available on GitHub:  [https://github.com/ChunhanLi/2nd-kaggle-LEAP](https://github.com/ChunhanLi/2nd-kaggle-LEAP)\n\n## Summary\nWe do not have extensive domain knowledge in atmospheric forecasting; we simply framed the competition task as a 1D seq2seq multi-target regression task. We trained with various model structures and meticulously tuned them, ultimately using a hill-climbing algorithm for model ensembling based on validation set scores.\n\nThe four highest-priority key points in the solution are summarized as follows:\n- **Data size is super important in this competition.** Using the entire [low-resolution dataset](https://huggingface.co/datasets/LEAP/ClimSim_low-res) can significantly enhance performance.\n- **The smoothl1 loss function, auxiliary diff loss and the cosine annealing learning rate schedule were adopted for model optimization.**\n- **Group fine-tuning methods were introduced in the training pipeline.**\n- **The team focused on developing diverse model structures and ensembling them.**\n\n\n## Solution\n\n### Data and Cross validation\n\n**Training data**  \nTraining data is crucial for this competition. According to this [GitHub repository](https://github.com/leap-stc/ClimSim), the Kaggle data is sampled from this [low-resolution dataset](https://huggingface.co/datasets/LEAP/ClimSim_low-res). Utilizing the entire low-resolution dataset can boost performance by almost 0.01 in this competition.\n\n**Cross validation**  \nOur models are trained using data from the first 7 years and the first half (January to June) of year 8, totaling approximately 75 million data points. We set the `stride_sample` parameter to 7 and used data from July to December of year 8 and January of year 9 as our hold-out validation dataset. This validation set has shown perfect correlation with the leaderboard scores.\n\n### Model optimization\n\n**Smooth L1 loss**  \nThe Smooth L1 loss is a robust loss function used primarily in regression tasks, combining the advantages of L2 and L1 losses. It is less sensitive to outliers than L2 and avoids the non-differentiability at zero seen with L1. The function is defined as follows:\n\n\n<div align=center><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5724098%2F8c7335ec3f1f94f0f40e041832b28303%2Floss.png?generation=1722302507888184&alt=media\" alt=\"jpg name\" width=\"60%\"/></div>\n\nThis approach helps stabilize the training process and often improves performance in noisy data situations.\n\n<div align=center><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5724098%2F566356fb54e82e5218a83c0746d8de42%2Ff1.png?generation=1722302556372586&alt=media\" alt=\"jpg name\" width=\"60%\"/></div>\n\n\n**Auxiliary diff loss**  \nWe use an auxiliary 'Diff Loss,' proposed by @zui0711 , to enhance learning. This loss calculates the difference between real and predicted values at adjacent levels of a target, using Smooth L1 loss to quantify the error. The approach is implemented as follows:\n```python\nwith torch.no_grad():\n    outputs = model(inputs)\n    loss = criterion(outputs, labels)\n    for i in range(6):\n        output_diff = outputs[:, 8+60*i+1:8+60*(i+1)] - outputs[:, 8+60*i:8+60*(i+1)-1]\n        label_diff = labels[:, 8+60*i+1:8+60*(i+1)] - labels[:, 8+60*i:8+60*(i+1)-1]\n        loss += criterion(output_diff, label_diff) / 6\n```\n\n**Cosine annealing scheduler**  \nThe cosine annealing schedule adjusts the learning rate in a cosine-shaped curve to help the model escape local minima and improve convergence. We implemented this with learning rate decays at the 3rd and 9th epochs.\n\n<div align=center><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5724098%2Fe965962b0d05d483bf9d1bf690014863%2Ff2.png?generation=1722302675640626&alt=media\" alt=\"jpg name\" width=\"60%\"/></div>\n\n\n### Group fine-tune\nDuring multi-objective learning tasks, different learning objectives can interact in complex ways, promoting or inhibiting each other. Our experiments found that while different target groups initially promoted each other, they began to interfere towards the end of training, potentially due to complex semantic constraints.\n\nInspired by the [top solution from the 2021 VPP competition](https://www.kaggle.com/competitions/ventilator-pressure-prediction/discussion/285320), we divided 368 features into seven groups, six of which are series of measurements of different metrics along the atmospheric column, and one group consists of eight unique single targets. After the training process with 364 full outputs was completed, we fine-tuned these groups again. This allowed each model with different architectures to achieve an improvement ranging from 0.0005 to 0.0015. Due to time and resource constraints, we only fine-tuned each group for one epoch. @max2020 \n\n### Post-processing\nAfter obtaining raw model predictions, we denormalize them using the standard deviation and mean. Following community [discussions](https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/499896#2791290), we adjust certain variables of *ptend_q0002* based on their linear relationships. Finally, we apply the weight values from the official sample submission file to process the predictions.\n\n## Model structure\n\n### Model by Forcewithme\n**Table 1**\n\n| Model id                                                     | pb      | fig                               | Exp  |\n| ------------------------------------------------------------ | ------- | --------------------------------- | ---- |\n| forcewithme\\_reslstm\\_cv0.789\\_lb0.783                       | 0.78275 | Figure 3(a)        | 18   |\n| forcewithme\\_gf\\_reslstm\\_cv0.790\\_lb0.785                   | 0.7840  | Figure 3(a)               | 32   |\n| forcewithme\\_gf\\_cnnlstm\\_cv0.789\\_lb0.787                   | 0.7826  | Figure 3(b)               | 39   |\n| forcewithme\\_gf\\_lstmmamba\\_cv0.7885\\_lb0.7853               | 0.7814  | Figure 3(c)           | 40   |\n| forcewithme\\_gf\\_LstmMambaMixed\\_cv0.7886\\_lb0.7858          | 0.7821  | Figure 3(d) | -    |\n| forcewithme\\_gf037\\_1LSTM1mamba-5\\_state16\\_cv0.7896\\_LBunknown | 0.7830  | Figure 3(d) | 37   |\n| forcewithme\\_gf038\\_2LSTM1mamba-3\\_state64\\_cv0.7897\\_LB0.787 | 0.7836  | Figure 3(d) | 38   |    \n    \n    \n    \n\n<div align=center><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5724098%2F978e5b0a2e223d5024fb69cb1031e526%2Ff3.png?generation=1722301967906869&alt=media\" alt=\"jpg name\" width=\"80%\"/></div>\n\nThe final ensemble strategy contains 6 models of ForcewithMe. These 6 models, the private score, the corresponding architectures, and the corresponding \"exp id\" in table 2 are in the table above. The model architectures are easy to understand and implement with the corresponding figures and the code. Here are some details and insights:\n\n- All models show local scores and public leaderboard (LB) scores within their model ids, which correspond directly to filenames, model names, and submission files (without extensions). The table supplements these with private LB scores and details on the model architectures.\n- All models are primarily built with LSTM as main components. This is due to our finding that LSTMs significantly outperform other architectures such as CNNs, MHSA, NNs, and GRUs in this competition.\n- Inspired by ResNet and Transformer, all my models incorporate residual connection. We believe residuals make models converge faster and perform better.\n- Figure 3(a) depicts a simple ResLSTM structure, which surprisingly achieved the highest private LB score among all models. Moreover, It took over two weeks for my another submission to beat this residual LSTM model on both local score and on the public leaderboard. Additionally, ResLSTM is still better on the private LB. exp32 has the same architecture to exp 18. The former one resumes the weights of the latter one and applies group fine-tuning.\n\n\n\n\n- Figure 3b illustrates a model combining small kernel convolutions, large kernel convolutions, and LSTMs as encoders, with the encoded results concatenated and fed into an LSTM backbone. This design leverages the presumed relationships between adjacent atmospheric layers, aiming to model local information and positional relations. Although it performs less well offline and on the private LB compared to the pure resLSTM model, it scores higher on the public LB and provides gains during the ensemble.\n- The remaining four models combine LSTMs with MAMBA. While LSTMs typically have more layers, MAMBA occupies the same number of layers as the LSTMs in model `forcewithme_gf037_1LSTM1mamba-5_state16_cv0.7896_LBunknown`.\n- The models mixing LSTM and Mamba (3d) in every block achieve very good scores offline and on the public LB.\n- Despite individual MAMBA-based models not outperforming the pure LSTM structure, they play a significant role in the final ensemble.\n- There are two versions of MAMBA available: MAMBA and MAMBA-2 in the \"mamba-ssm\" package. We exclusively used MAMBA.\n- The default parameters for MAMBA are: `d_model=512`, `d_state=16`, `d_conv=4`, expand=2. Other than `d_model`, the parameter settings follow the official repository.\n- The models `forcewithme_gf037_1LSTM1mamba-5_state16_cv0.7896_LBunknown` have 1 LSTM and 1 MAMBA in each block, and it has 5 blocks, the mamba uses default params setting. While the models `forcewithme_gf_LstmMambaMixed_cv0.7886_lb0.7858` and `forcewithme_gf038_2LSTM1mamba-3_state64` have 2 LSTMs and 1 mamba in each block, and they have 3 blocks.\n- The models `forcewithme_gf_LstmMambaMixed_cv0.7886_lb0.7858` and `forcewithme_gf038_2LSTM1mamba-3_state64` differ only in that the latter has `d_state` set to `64`. They share figure 3(d) as they are almost the same.\n- For installation and usage of MAMBA, please refer to the official repository: [https://github.com/state-spaces/mamba?tab=readme-ov-file](https://github.com/state-spaces/mamba?tab=readme-ov-file). It's licensed under Apache-2.0, allowing for free use, including commercial purposes, provided the license requirements are met.\n\n\n### Model by Max2020\n\nIn the Max2020 part, models 14, 15, 21, and 22 are all improvements based on the model by [@forcewithme](https://www.kaggle.com/forcewithme), integrating LSTM with skip connections. **Model 22** is our team’s highest-performing single model, providing the best results in local scoring, Leader Board scoring, and private scoring. Its architecture is shown in Figure 4. Regarding the learning rate schedule, a cosine decay learning rate was used, with decays occurring at 3 and 9 epochs. The loss function employed was the smooth L1 loss with a beta of 0.5. The Table 2 contains the detailed performance of my five models.\n\n\n<div align=center><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5724098%2F82ce64bb111824f2eee6c53954286c7a%2Ff4.png?generation=1722302228698175&alt=media\" alt=\"jpg name\" width=\"60%\"/></div>\n\n### Model by Joseph Zhou\n\nThe structure of Joseph's models mainly consists of the use of multi-layers Res-ConvLSTM blocks and a TimeDistributed fully-connected layer. The input size is `[bs, 60, 25]` and output size is `[bs, 368]`, the last dimension of the input includes 16 sequence-features and 9 scalar-features. After passing the multi-layers Res-ConvLSTM blocks and fully-connected layer, it will result in a size of `[bs, 60, 14]`, the last dimension includes 6 sequence-labels and 8 scalar-labels. For scalar-labels part which is `[bs, 60, 8]`, we just average in the second axis to get the shape of `[bs, 8]`. For sequence-labels part which is `[bs, 60, 6]`, we reshape it to the shape of `[bs, 360]`. Finally, we concatenate the two parts to get the output. Besides, a reverse and shift type of augmentation is implemented in one of the models. The idea is we randomly reverse or shift the sequence-features and sequence-labels respectively in the data loader.\n\n### Model by Adam\n\nThe structure of Adam's models are like: Inputs -> Position encoding -> 1DCNN -> 3-4 layers LSTM -> 1 layer transformer -> Outputs. Most models used all the low-resolution dataset and the others only use sampling data. The model trained by sampling data can contribute to the ensemble a bit. For the loss, Adam used SmoothL1 loss and auxiliary diff loss as mentioned before. For the scheduler, Adam used ReduceLROnPlateau scheduler with a factor of 0.2 and patience of 2.\n\nAdam's best single model has LB 0.78594 and PB 0.78141. The ensemble of Adam's models has LB 0.79050 and PB 0.78575.\n\n### Model by Zuiye\n\nZuiye's models are mainly based on two architectures. The first one consists of 2 LSTM layers followed by a MultiheadAttention layer. The other one consists of 3 parallel Convolutional layers with 3 different kernel sizes and next 2 LSTM layers followed by a MultiheadAttention layer just like the first architecture. Zuiye's best single model gets LB 0.78696 / PB 0.78205 and the ensemble of Zuiye's own models (with hill climb) gets LB 0.79050 / PB 0.78614.\n\n\n### Model ensembling\n\nWe use the hill climb method to search blend weights of each model. The steps of hill climb are the following: @hookman \n\n1. Take the best out-of-fold predictions as best\\_ensemble. This will be our baseline.\n2. Iteratively blend best\\_ensemble with different models with different weights, using the formula `new\\_ensemble = w * best\\_ensemble + (1-w) * new\\_oof`\n3. Check the r2 score of the new ensemble. Choose the best new ensemble to replace best\\_ensemble as our new baseline.\n4. Repeat until the r2 score can't increase anymore or reaches the threshold.    \n    \n    \n**Weights of the best model are shown in the table below:**    \n**Table 2**\n    \n| Exp id           | weight   | cv    | lb     | pb     |\n|------------------|----------|-------|--------|--------|\n| forcewithme\\_exp32| 0.166556 | 0.790 | 0.7865 | 0.78398|\n| forcewithme\\_exp37| 0.158625 | 0.7896| 0.78618| 0.78293|\n| forcewithme\\_exp38| 0.139194 | 0.7897| 0.78719| 0.78362|\n| max\\_exp22        | 0.120125 | 0.7908| **0.78793**| **0.78434**|\n| Jo\\_exp912        | 0.111971 | 0.78935| 0.78528| 0.78150|\n| max\\_exp21        | 0.104738 | 0.7904| 0.78752| 0.78425|\n| forcewithme\\_exp39| 0.098977 | 0.789 | 0.78699| 0.78257|\n| max\\_exp14        | 0.093088 | 0.7905| 0.78641| 0.78214|\n| max\\_exp10        | 0.092157 | 0.7888| 0.78619| 0.78213|\n| forcewithme\\_exp40| 0.082941 | 0.7885| 0.7853 | 0.78261|\n| max\\_exp015       | 0.052500 | 0.7905| 0.78695| 0.78244|\n| adam\\_exp197      | 0.048994 | 0.7855| 0.78269| 0.777  |\n| adam\\_exp200      | -0.047132| 0.7836| 0.78010| 0.77434|\n| adam\\_exp195      | -0.049875| 0.78569| 0.78334| 0.77753|\n| Jo\\_exp907        | -0.083779| 0.7855| 0.78289| 0.77873|\n| forcewithme\\_exp18| -0.089079| 0.7890| 0.7863 | 0.78272|\n| **Ensemble**      | **1.0**  | **0.7955**| **0.79211**| **0.78856**|\n    \n----\n    \nIt's worth noting that we've provided the parameter weights files for all models on [Google Drive](https://drive.google.com/drive/u/0/folders/1-1gavnxXqj2x6giAjPkPQcrTfpVE3fpC). Using these for ensemble submissions results in slightly higher LB and PB scores, by a margin of 0.0002. The discrepancies stem from two main issues:\n\nFirstly, the `jo_exp907.pt` model file was missing, and the model had to be rerun post-competition, which led to some differences from the original. Secondly, during the competition, an incorrect model file was used under `forcewithme_exp18` (corresponding to the `forcewithme_reslstm_cv0.789_lb0.783` folder). This has now been corrected. These two points have caused a very minor difference in our final results. Although we believe this difference does not impact the reproducibility of our overall approach, we mention it here to avoid any confusion.",
      "votes": 46
    },
    {
      "id": 2968112,
      "postDate": "2024-08-23T14:48:26.997Z",
      "content": "<p>How did you shuffle/load training data? I had trouble getting it into memory and if you don't adequately shuffle larger batches of data (if you load them separately) the learning becomes noisy. </p>",
      "rawMarkdown": "How did you shuffle/load training data? I had trouble getting it into memory and if you don't adequately shuffle larger batches of data (if you load them separately) the learning becomes noisy. ",
      "replies": [
        {
          "id": 2975095,
          "postDate": "2024-08-31T12:33:01.800Z",
          "content": "<p>We rented a high-memory server, and reading the entire dataset requires at least 300GB of memory space.</p>",
          "rawMarkdown": "We rented a high-memory server, and reading the entire dataset requires at least 300GB of memory space."
        }
      ]
    },
    {
      "id": 2940737,
      "postDate": "2024-07-30T12:05:46.200Z",
      "content": "<p>Congratulations on achieving second place in the LEAP - Atmospheric Physics using AI (ClimSim) competition! <a href=\"https://www.kaggle.com/max2020\" target=\"_blank\">@max2020</a><br>\nYour innovative approach, framing the task as a 1D seq2seq multi-target regression, and the use of a hill-climbing algorithm for model ensembling are truly commendable. The strategic use of the entire low-resolution dataset, Smooth L1 loss, auxiliary diff loss, and group fine-tuning methods significantly enhanced performance. Your detailed and transparent sharing of the solution on GitHub is a valuable resource for the community. Thank you for your contribution and for recognizing the support of the hosts and the Kaggle team.</p>",
      "rawMarkdown": "Congratulations on achieving second place in the LEAP - Atmospheric Physics using AI (ClimSim) competition! @max2020\nYour innovative approach, framing the task as a 1D seq2seq multi-target regression, and the use of a hill-climbing algorithm for model ensembling are truly commendable. The strategic use of the entire low-resolution dataset, Smooth L1 loss, auxiliary diff loss, and group fine-tuning methods significantly enhanced performance. Your detailed and transparent sharing of the solution on GitHub is a valuable resource for the community. Thank you for your contribution and for recognizing the support of the hosts and the Kaggle team."
    },
    {
      "id": 2940350,
      "postDate": "2024-07-30T02:33:45.503Z",
      "content": "<p>Thank you for the explanation. </p>",
      "rawMarkdown": "Thank you for the explanation. "
    }
  ],
  "comments": [
    {
      "id": 2968112,
      "author_name": "zpanj",
      "author_url": "",
      "post_date": "2024-08-23T14:48:26.997000",
      "content": "<p>How did you shuffle/load training data? I had trouble getting it into memory and if you don't adequately shuffle larger batches of data (if you load them separately) the learning becomes noisy. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2975095,
          "author_name": "Max2020",
          "author_url": "",
          "post_date": "2024-08-31T12:33:01.800000",
          "content": "<p>We rented a high-memory server, and reading the entire dataset requires at least 300GB of memory space.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2940737,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-07-30T12:05:46.200000",
      "content": "<p>Congratulations on achieving second place in the LEAP - Atmospheric Physics using AI (ClimSim) competition! <a href=\"https://www.kaggle.com/max2020\" target=\"_blank\">@max2020</a><br>\nYour innovative approach, framing the task as a 1D seq2seq multi-target regression, and the use of a hill-climbing algorithm for model ensembling are truly commendable. The strategic use of the entire low-resolution dataset, Smooth L1 loss, auxiliary diff loss, and group fine-tuning methods significantly enhanced performance. Your detailed and transparent sharing of the solution on GitHub is a valuable resource for the community. Thank you for your contribution and for recognizing the support of the hosts and the Kaggle team.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2940350,
      "author_name": "Sharan V",
      "author_url": "",
      "post_date": "2024-07-30T02:33:45.503000",
      "content": "<p>Thank you for the explanation. </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2940309": "Thanks to the hosts @jerrylin96  and official kaggle team for organizing this interesting competition! Our team learned a lot from it!\nI am very pleased to share our team's solution. The complete code is available on GitHub:  [https://github.com/ChunhanLi/2nd-kaggle-LEAP](https://github.com/ChunhanLi/2nd-kaggle-LEAP)\n\n## Summary\nWe do not have extensive domain knowledge in atmospheric forecasting; we simply framed the competition task as a 1D seq2seq multi-target regression task. We trained with various model structures and meticulously tuned them, ultimately using a hill-climbing algorithm for model ensembling based on validation set scores.\n\nThe four highest-priority key points in the solution are summarized as follows:\n- **Data size is super important in this competition.** Using the entire [low-resolution dataset](https://huggingface.co/datasets/LEAP/ClimSim_low-res) can significantly enhance performance.\n- **The smoothl1 loss function, auxiliary diff loss and the cosine annealing learning rate schedule were adopted for model optimization.**\n- **Group fine-tuning methods were introduced in the training pipeline.**\n- **The team focused on developing diverse model structures and ensembling them.**\n\n\n## Solution\n\n### Data and Cross validation\n\n**Training data**  \nTraining data is crucial for this competition. According to this [GitHub repository](https://github.com/leap-stc/ClimSim), the Kaggle data is sampled from this [low-resolution dataset](https://huggingface.co/datasets/LEAP/ClimSim_low-res). Utilizing the entire low-resolution dataset can boost performance by almost 0.01 in this competition.\n\n**Cross validation**  \nOur models are trained using data from the first 7 years and the first half (January to June) of year 8, totaling approximately 75 million data points. We set the `stride_sample` parameter to 7 and used data from July to December of year 8 and January of year 9 as our hold-out validation dataset. This validation set has shown perfect correlation with the leaderboard scores.\n\n### Model optimization\n\n**Smooth L1 loss**  \nThe Smooth L1 loss is a robust loss function used primarily in regression tasks, combining the advantages of L2 and L1 losses. It is less sensitive to outliers than L2 and avoids the non-differentiability at zero seen with L1. The function is defined as follows:\n\n\n<div align=center><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5724098%2F8c7335ec3f1f94f0f40e041832b28303%2Floss.png?generation=1722302507888184&alt=media\" alt=\"jpg name\" width=\"60%\"/></div>\n\nThis approach helps stabilize the training process and often improves performance in noisy data situations.\n\n<div align=center><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5724098%2F566356fb54e82e5218a83c0746d8de42%2Ff1.png?generation=1722302556372586&alt=media\" alt=\"jpg name\" width=\"60%\"/></div>\n\n\n**Auxiliary diff loss**  \nWe use an auxiliary 'Diff Loss,' proposed by @zui0711 , to enhance learning. This loss calculates the difference between real and predicted values at adjacent levels of a target, using Smooth L1 loss to quantify the error. The approach is implemented as follows:\n```python\nwith torch.no_grad():\n    outputs = model(inputs)\n    loss = criterion(outputs, labels)\n    for i in range(6):\n        output_diff = outputs[:, 8+60*i+1:8+60*(i+1)] - outputs[:, 8+60*i:8+60*(i+1)-1]\n        label_diff = labels[:, 8+60*i+1:8+60*(i+1)] - labels[:, 8+60*i:8+60*(i+1)-1]\n        loss += criterion(output_diff, label_diff) / 6\n```\n\n**Cosine annealing scheduler**  \nThe cosine annealing schedule adjusts the learning rate in a cosine-shaped curve to help the model escape local minima and improve convergence. We implemented this with learning rate decays at the 3rd and 9th epochs.\n\n<div align=center><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5724098%2Fe965962b0d05d483bf9d1bf690014863%2Ff2.png?generation=1722302675640626&alt=media\" alt=\"jpg name\" width=\"60%\"/></div>\n\n\n### Group fine-tune\nDuring multi-objective learning tasks, different learning objectives can interact in complex ways, promoting or inhibiting each other. Our experiments found that while different target groups initially promoted each other, they began to interfere towards the end of training, potentially due to complex semantic constraints.\n\nInspired by the [top solution from the 2021 VPP competition](https://www.kaggle.com/competitions/ventilator-pressure-prediction/discussion/285320), we divided 368 features into seven groups, six of which are series of measurements of different metrics along the atmospheric column, and one group consists of eight unique single targets. After the training process with 364 full outputs was completed, we fine-tuned these groups again. This allowed each model with different architectures to achieve an improvement ranging from 0.0005 to 0.0015. Due to time and resource constraints, we only fine-tuned each group for one epoch. @max2020 \n\n### Post-processing\nAfter obtaining raw model predictions, we denormalize them using the standard deviation and mean. Following community [discussions](https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/499896#2791290), we adjust certain variables of *ptend_q0002* based on their linear relationships. Finally, we apply the weight values from the official sample submission file to process the predictions.\n\n## Model structure\n\n### Model by Forcewithme\n**Table 1**\n\n| Model id                                                     | pb      | fig                               | Exp  |\n| ------------------------------------------------------------ | ------- | --------------------------------- | ---- |\n| forcewithme\\_reslstm\\_cv0.789\\_lb0.783                       | 0.78275 | Figure 3(a)        | 18   |\n| forcewithme\\_gf\\_reslstm\\_cv0.790\\_lb0.785                   | 0.7840  | Figure 3(a)               | 32   |\n| forcewithme\\_gf\\_cnnlstm\\_cv0.789\\_lb0.787                   | 0.7826  | Figure 3(b)               | 39   |\n| forcewithme\\_gf\\_lstmmamba\\_cv0.7885\\_lb0.7853               | 0.7814  | Figure 3(c)           | 40   |\n| forcewithme\\_gf\\_LstmMambaMixed\\_cv0.7886\\_lb0.7858          | 0.7821  | Figure 3(d) | -    |\n| forcewithme\\_gf037\\_1LSTM1mamba-5\\_state16\\_cv0.7896\\_LBunknown | 0.7830  | Figure 3(d) | 37   |\n| forcewithme\\_gf038\\_2LSTM1mamba-3\\_state64\\_cv0.7897\\_LB0.787 | 0.7836  | Figure 3(d) | 38   |    \n    \n    \n    \n\n<div align=center><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5724098%2F978e5b0a2e223d5024fb69cb1031e526%2Ff3.png?generation=1722301967906869&alt=media\" alt=\"jpg name\" width=\"80%\"/></div>\n\nThe final ensemble strategy contains 6 models of ForcewithMe. These 6 models, the private score, the corresponding architectures, and the corresponding \"exp id\" in table 2 are in the table above. The model architectures are easy to understand and implement with the corresponding figures and the code. Here are some details and insights:\n\n- All models show local scores and public leaderboard (LB) scores within their model ids, which correspond directly to filenames, model names, and submission files (without extensions). The table supplements these with private LB scores and details on the model architectures.\n- All models are primarily built with LSTM as main components. This is due to our finding that LSTMs significantly outperform other architectures such as CNNs, MHSA, NNs, and GRUs in this competition.\n- Inspired by ResNet and Transformer, all my models incorporate residual connection. We believe residuals make models converge faster and perform better.\n- Figure 3(a) depicts a simple ResLSTM structure, which surprisingly achieved the highest private LB score among all models. Moreover, It took over two weeks for my another submission to beat this residual LSTM model on both local score and on the public leaderboard. Additionally, ResLSTM is still better on the private LB. exp32 has the same architecture to exp 18. The former one resumes the weights of the latter one and applies group fine-tuning.\n\n\n\n\n- Figure 3b illustrates a model combining small kernel convolutions, large kernel convolutions, and LSTMs as encoders, with the encoded results concatenated and fed into an LSTM backbone. This design leverages the presumed relationships between adjacent atmospheric layers, aiming to model local information and positional relations. Although it performs less well offline and on the private LB compared to the pure resLSTM model, it scores higher on the public LB and provides gains during the ensemble.\n- The remaining four models combine LSTMs with MAMBA. While LSTMs typically have more layers, MAMBA occupies the same number of layers as the LSTMs in model `forcewithme_gf037_1LSTM1mamba-5_state16_cv0.7896_LBunknown`.\n- The models mixing LSTM and Mamba (3d) in every block achieve very good scores offline and on the public LB.\n- Despite individual MAMBA-based models not outperforming the pure LSTM structure, they play a significant role in the final ensemble.\n- There are two versions of MAMBA available: MAMBA and MAMBA-2 in the \"mamba-ssm\" package. We exclusively used MAMBA.\n- The default parameters for MAMBA are: `d_model=512`, `d_state=16`, `d_conv=4`, expand=2. Other than `d_model`, the parameter settings follow the official repository.\n- The models `forcewithme_gf037_1LSTM1mamba-5_state16_cv0.7896_LBunknown` have 1 LSTM and 1 MAMBA in each block, and it has 5 blocks, the mamba uses default params setting. While the models `forcewithme_gf_LstmMambaMixed_cv0.7886_lb0.7858` and `forcewithme_gf038_2LSTM1mamba-3_state64` have 2 LSTMs and 1 mamba in each block, and they have 3 blocks.\n- The models `forcewithme_gf_LstmMambaMixed_cv0.7886_lb0.7858` and `forcewithme_gf038_2LSTM1mamba-3_state64` differ only in that the latter has `d_state` set to `64`. They share figure 3(d) as they are almost the same.\n- For installation and usage of MAMBA, please refer to the official repository: [https://github.com/state-spaces/mamba?tab=readme-ov-file](https://github.com/state-spaces/mamba?tab=readme-ov-file). It's licensed under Apache-2.0, allowing for free use, including commercial purposes, provided the license requirements are met.\n\n\n### Model by Max2020\n\nIn the Max2020 part, models 14, 15, 21, and 22 are all improvements based on the model by [@forcewithme](https://www.kaggle.com/forcewithme), integrating LSTM with skip connections. **Model 22** is our team’s highest-performing single model, providing the best results in local scoring, Leader Board scoring, and private scoring. Its architecture is shown in Figure 4. Regarding the learning rate schedule, a cosine decay learning rate was used, with decays occurring at 3 and 9 epochs. The loss function employed was the smooth L1 loss with a beta of 0.5. The Table 2 contains the detailed performance of my five models.\n\n\n<div align=center><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5724098%2F82ce64bb111824f2eee6c53954286c7a%2Ff4.png?generation=1722302228698175&alt=media\" alt=\"jpg name\" width=\"60%\"/></div>\n\n### Model by Joseph Zhou\n\nThe structure of Joseph's models mainly consists of the use of multi-layers Res-ConvLSTM blocks and a TimeDistributed fully-connected layer. The input size is `[bs, 60, 25]` and output size is `[bs, 368]`, the last dimension of the input includes 16 sequence-features and 9 scalar-features. After passing the multi-layers Res-ConvLSTM blocks and fully-connected layer, it will result in a size of `[bs, 60, 14]`, the last dimension includes 6 sequence-labels and 8 scalar-labels. For scalar-labels part which is `[bs, 60, 8]`, we just average in the second axis to get the shape of `[bs, 8]`. For sequence-labels part which is `[bs, 60, 6]`, we reshape it to the shape of `[bs, 360]`. Finally, we concatenate the two parts to get the output. Besides, a reverse and shift type of augmentation is implemented in one of the models. The idea is we randomly reverse or shift the sequence-features and sequence-labels respectively in the data loader.\n\n### Model by Adam\n\nThe structure of Adam's models are like: Inputs -> Position encoding -> 1DCNN -> 3-4 layers LSTM -> 1 layer transformer -> Outputs. Most models used all the low-resolution dataset and the others only use sampling data. The model trained by sampling data can contribute to the ensemble a bit. For the loss, Adam used SmoothL1 loss and auxiliary diff loss as mentioned before. For the scheduler, Adam used ReduceLROnPlateau scheduler with a factor of 0.2 and patience of 2.\n\nAdam's best single model has LB 0.78594 and PB 0.78141. The ensemble of Adam's models has LB 0.79050 and PB 0.78575.\n\n### Model by Zuiye\n\nZuiye's models are mainly based on two architectures. The first one consists of 2 LSTM layers followed by a MultiheadAttention layer. The other one consists of 3 parallel Convolutional layers with 3 different kernel sizes and next 2 LSTM layers followed by a MultiheadAttention layer just like the first architecture. Zuiye's best single model gets LB 0.78696 / PB 0.78205 and the ensemble of Zuiye's own models (with hill climb) gets LB 0.79050 / PB 0.78614.\n\n\n### Model ensembling\n\nWe use the hill climb method to search blend weights of each model. The steps of hill climb are the following: @hookman \n\n1. Take the best out-of-fold predictions as best\\_ensemble. This will be our baseline.\n2. Iteratively blend best\\_ensemble with different models with different weights, using the formula `new\\_ensemble = w * best\\_ensemble + (1-w) * new\\_oof`\n3. Check the r2 score of the new ensemble. Choose the best new ensemble to replace best\\_ensemble as our new baseline.\n4. Repeat until the r2 score can't increase anymore or reaches the threshold.    \n    \n    \n**Weights of the best model are shown in the table below:**    \n**Table 2**\n    \n| Exp id           | weight   | cv    | lb     | pb     |\n|------------------|----------|-------|--------|--------|\n| forcewithme\\_exp32| 0.166556 | 0.790 | 0.7865 | 0.78398|\n| forcewithme\\_exp37| 0.158625 | 0.7896| 0.78618| 0.78293|\n| forcewithme\\_exp38| 0.139194 | 0.7897| 0.78719| 0.78362|\n| max\\_exp22        | 0.120125 | 0.7908| **0.78793**| **0.78434**|\n| Jo\\_exp912        | 0.111971 | 0.78935| 0.78528| 0.78150|\n| max\\_exp21        | 0.104738 | 0.7904| 0.78752| 0.78425|\n| forcewithme\\_exp39| 0.098977 | 0.789 | 0.78699| 0.78257|\n| max\\_exp14        | 0.093088 | 0.7905| 0.78641| 0.78214|\n| max\\_exp10        | 0.092157 | 0.7888| 0.78619| 0.78213|\n| forcewithme\\_exp40| 0.082941 | 0.7885| 0.7853 | 0.78261|\n| max\\_exp015       | 0.052500 | 0.7905| 0.78695| 0.78244|\n| adam\\_exp197      | 0.048994 | 0.7855| 0.78269| 0.777  |\n| adam\\_exp200      | -0.047132| 0.7836| 0.78010| 0.77434|\n| adam\\_exp195      | -0.049875| 0.78569| 0.78334| 0.77753|\n| Jo\\_exp907        | -0.083779| 0.7855| 0.78289| 0.77873|\n| forcewithme\\_exp18| -0.089079| 0.7890| 0.7863 | 0.78272|\n| **Ensemble**      | **1.0**  | **0.7955**| **0.79211**| **0.78856**|\n    \n----\n    \nIt's worth noting that we've provided the parameter weights files for all models on [Google Drive](https://drive.google.com/drive/u/0/folders/1-1gavnxXqj2x6giAjPkPQcrTfpVE3fpC). Using these for ensemble submissions results in slightly higher LB and PB scores, by a margin of 0.0002. The discrepancies stem from two main issues:\n\nFirstly, the `jo_exp907.pt` model file was missing, and the model had to be rerun post-competition, which led to some differences from the original. Secondly, during the competition, an incorrect model file was used under `forcewithme_exp18` (corresponding to the `forcewithme_reslstm_cv0.789_lb0.783` folder). This has now been corrected. These two points have caused a very minor difference in our final results. Although we believe this difference does not impact the reproducibility of our overall approach, we mention it here to avoid any confusion.",
    "2968112": "How did you shuffle/load training data? I had trouble getting it into memory and if you don't adequately shuffle larger batches of data (if you load them separately) the learning becomes noisy. ",
    "2940737": "Congratulations on achieving second place in the LEAP - Atmospheric Physics using AI (ClimSim) competition! @max2020\nYour innovative approach, framing the task as a 1D seq2seq multi-target regression, and the use of a hill-climbing algorithm for model ensembling are truly commendable. The strategic use of the entire low-resolution dataset, Smooth L1 loss, auxiliary diff loss, and group fine-tuning methods significantly enhanced performance. Your detailed and transparent sharing of the solution on GitHub is a valuable resource for the community. Thank you for your contribution and for recognizing the support of the hosts and the Kaggle team.",
    "2940350": "Thank you for the explanation. "
  }
}