{
  "id": 523129,
  "title": "27th Place Solution",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/523129",
  "author_name": "chumajin",
  "post_date": "2024-07-30T14:07:09.878000",
  "votes": 17,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Thank you for hosting such an interesting competition. We are deeply grateful to the hosts and the Kaggle staff. I would also like to thank , <a href=\"https://www.kaggle.com/sae20gorilla\" target=\"_blank\">@sae20gorilla</a> , <a href=\"https://www.kaggle.com/hidebu\" target=\"_blank\">@hidebu</a>, <a href=\"https://www.kaggle.com/takatoyoshikawa\" target=\"_blank\">@takatoyoshikawa</a> , and <a href=\"https://www.kaggle.com/khiroki\" target=\"_blank\">@khiroki</a> for teaming up with me. Even though we formed the team just a week before the close, we had many discussions that were very educational, and it was enjoyable to see our scores and public lb rank improve each time.</p>\n<h1>1. Summary</h1>\n<p>Our best private solution is an ensemble of 8 models.</p>\n<p>After forming the team, we started the ensemble using external data for validation that none of us had used for training. Since the relationship between CV and the public LB was favorable in this competition, we calculated the coefficients using the Nelder-Mead method for each model to significantly improve the CV (sub1). Additionally, anticipating the presence of outliers, we used the Nelder-Mead method to calculate coefficients for each column (sub2), but as expected, the private LB for this was not as good. Below, we will describe each part for sub1 in more detail.</p>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>part</th>\n<th>model</th>\n<th>loss</th>\n<th>use train volume</th>\n<th>Comment</th>\n<th>cv</th>\n<th>publicScore</th>\n<th>privateScore</th>\n<th>weights</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>tenten</td>\n<td>1dConv-Transformer-GRU</td>\n<td>Huberloss</td>\n<td>73.5M</td>\n<td></td>\n<td>0.77776</td>\n<td>0.77697</td>\n<td>0.77157</td>\n<td>0.49187</td>\n</tr>\n<tr>\n<td>2</td>\n<td>tenten</td>\n<td>1dConv-Transformer-GRU</td>\n<td>Huberloss</td>\n<td>35M</td>\n<td></td>\n<td>0.77257</td>\n<td>0.77292</td>\n<td>0.76602</td>\n<td>0.09789</td>\n</tr>\n<tr>\n<td>3</td>\n<td>chumajin</td>\n<td>squeeze former</td>\n<td>Huberloss</td>\n<td>16M</td>\n<td>dim 192 + retraining</td>\n<td>0.76961</td>\n<td>0.77050</td>\n<td>0.76239</td>\n<td>0.32106</td>\n</tr>\n<tr>\n<td>4</td>\n<td>sae</td>\n<td>1dConv - GRU</td>\n<td>Huberloss</td>\n<td>16M</td>\n<td></td>\n<td>0.76787</td>\n<td>0.76890</td>\n<td>0.76192</td>\n<td>0.01156</td>\n</tr>\n<tr>\n<td>5</td>\n<td>tenten</td>\n<td>1dUnet</td>\n<td>Huberloss</td>\n<td>73.5M</td>\n<td></td>\n<td>0.76723</td>\n<td>0.76611</td>\n<td>0.76139</td>\n<td>0.07264</td>\n</tr>\n<tr>\n<td>6</td>\n<td>sae</td>\n<td>1dConv - LSTM</td>\n<td>Huberloss</td>\n<td>16M</td>\n<td></td>\n<td>0.76618</td>\n<td>0.76739</td>\n<td>0.76001</td>\n<td>0.18994</td>\n</tr>\n<tr>\n<td>7</td>\n<td>chumajin</td>\n<td>squeeze former</td>\n<td>Huberloss</td>\n<td>16M</td>\n<td>dim 128</td>\n<td>0.76121</td>\n<td>0.76278</td>\n<td>0.75497</td>\n<td>-0.24416</td>\n</tr>\n<tr>\n<td>8</td>\n<td>tenten</td>\n<td>1dConv-Transformer-GRU</td>\n<td>MSEloss</td>\n<td>16M</td>\n<td></td>\n<td>0.76108</td>\n<td>0.75965</td>\n<td>0.75654</td>\n<td>0.07312</td>\n</tr>\n</tbody>\n</table>\n<p><strong>final ensemble cv : 0.78116375, public lb : 0.78052, private lb : 0.77477</strong></p>\n<h1>2. tenten's part</h1>\n<ul>\n<li>overview</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2F19847183370359028731b5040267ca88%2FKaggle_LEAP._tentenpng.png?generation=1722347182292758&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>Input<ul>\n<li>StandardScaler</li>\n<li>Scalars are expanded to match the series length (60).</li></ul></li>\n<li>Model<ul>\n<li>1dUnet<ul>\n<li>down_channels: [128, 256, 512]</li></ul></li>\n<li>1dConv-Transformer-GRU</li></ul></li>\n<li>Training<ul>\n<li>Loss function: HuberLoss(delta=1.0)</li>\n<li>Optimizer: AdamW</li>\n<li>Scheduler: Warmup Cosine Annealing</li>\n<li>Data: All low-resolution data</li></ul></li>\n<li>Postprocess<ul>\n<li>I used torch.min to limit all constant targets except cam_out_FLWDS to non-negative values (I apply this during training before calculating the loss)</li>\n<li>ptend_q0002_x =  -state_q0002_x * 1200 (12≤x≤27)</li></ul></li>\n<li>What didn’t work well<ul>\n<li>Auxiliary loss (latlon, grid id)</li>\n<li>Increase the sequence length by upsampling</li>\n<li>Robust scaler</li>\n<li>Tuning the loss weight </li></ul></li>\n</ul>\n<h1>3. saeNeko's part</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2F2ab9eabeabc302b73dffc382404d358f%2FKaggle_LEAP._sae.png?generation=1722347171155973&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li><p>Models:</p>\n<p>・ 1D CNN + GRU</p>\n<p>・ 1D CNN + LSTM</p></li>\n<li><p>Preprocessing:</p>\n<p>Applied StandardScaler to both features and targets.</p></li>\n<li><p>Training Data:<br>\n・Used Kaggle’s train.csv.<br>\n・6M additional data (Thanks to <a href=\"https://www.kaggle.com/khirokifor\" target=\"_blank\">@khirokifor</a> extracting data from Hugging Face!)</p></li>\n<li><p>Training:</p>\n<p>Trained separately for series targets and scalar targets.</p></li>\n<li><p>Loss Function:</p>\n<p>RMSE</p></li>\n<li><p>Postprocessing:</p>\n<p>For ptendq0002_x(1&lt;=x&lt;=27), used a linear model.<br>\nptend_q0002_x = state_q0002_x * coef + intercept_</p></li>\n</ul>\n<h1>4. chumajin's part</h1>\n<ul>\n<li><p>Preprocess</p>\n<p>・ The diff of the sequential part and StandardScaler<br>\n・ The scalar part was repeated after applying the StandardScaler.</p></li>\n<li><p>Training Data:</p>\n<p>・Used Kaggle’s train.csv.<br>\n・6M additional data</p></li>\n<li><p>Architecture</p>\n<p>SqueezeFormer from the 2nd place solution of the Stanford Ribonanza RNA Folding competition, (<a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/460316\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/460316</a>) </p></li>\n<li><p>Mask of loss</p>\n<p>I did not use the loss for the unpredictable target.</p></li>\n<li><p>Re-training</p>\n<p>Additional 10 epochs fine-tuning after 20 epochs.<br>\ncv 0.767055 → 0.768034112 (I call this chumajin model1.)</p></li>\n<li><p>Insights from my teammates<br>\n・ Huberloss(delta=1)<br>\ncv improved by about 0.01. </p>\n<p>・ Re-training(re-finetune) by separating the loss<br>\nBy training again with limiting the loss, and creating three additional models from the above chumajin model1, I averaged their predictions which was inferred from the limited loss parts of the model with chumajin model1</p>\n<ol>\n<li>model 2 : limited loss of ptend_q0001, q0002, and q0003</li>\n<li>model 3 : limited loss of ptend_t, ptend_u, and ptend_v</li>\n<li>model 4 : limited loss of scalar values </li></ol></li>\n</ul>\n<p>By doing this, I was able to improve my model's CV from 0.7680 to 0.7696. As a result, re-training led to a CV improvement of +0.0026.</p>\n<ul>\n<li><p>Postprocess</p>\n<p>・ Applying tenten's postprocess of torch.min for scaler part.<br>\n・ Applying sae's postprocess, the CV improved by about 0.002.</p></li>\n<li><p>Regrets and Acknowledgements</p>\n<p>With tenten's insights, the improved CV SqueezeFormer was created about a day and a half before the deadline. It took approximately 24 hours to train on 16M train data with an A100, so I couldn't train on all the 73.5M low-resolution data…But I learned a lot expecially from team up. Thank you for team mate!!</p></li>\n</ul>",
  "messages": [
    {
      "id": 2940856,
      "postDate": "2024-07-30T14:07:09.877Z",
      "content": "<p>Thank you for hosting such an interesting competition. We are deeply grateful to the hosts and the Kaggle staff. I would also like to thank , <a href=\"https://www.kaggle.com/sae20gorilla\" target=\"_blank\">@sae20gorilla</a> , <a href=\"https://www.kaggle.com/hidebu\" target=\"_blank\">@hidebu</a>, <a href=\"https://www.kaggle.com/takatoyoshikawa\" target=\"_blank\">@takatoyoshikawa</a> , and <a href=\"https://www.kaggle.com/khiroki\" target=\"_blank\">@khiroki</a> for teaming up with me. Even though we formed the team just a week before the close, we had many discussions that were very educational, and it was enjoyable to see our scores and public lb rank improve each time.</p>\n<h1>1. Summary</h1>\n<p>Our best private solution is an ensemble of 8 models.</p>\n<p>After forming the team, we started the ensemble using external data for validation that none of us had used for training. Since the relationship between CV and the public LB was favorable in this competition, we calculated the coefficients using the Nelder-Mead method for each model to significantly improve the CV (sub1). Additionally, anticipating the presence of outliers, we used the Nelder-Mead method to calculate coefficients for each column (sub2), but as expected, the private LB for this was not as good. Below, we will describe each part for sub1 in more detail.</p>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>part</th>\n<th>model</th>\n<th>loss</th>\n<th>use train volume</th>\n<th>Comment</th>\n<th>cv</th>\n<th>publicScore</th>\n<th>privateScore</th>\n<th>weights</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>tenten</td>\n<td>1dConv-Transformer-GRU</td>\n<td>Huberloss</td>\n<td>73.5M</td>\n<td></td>\n<td>0.77776</td>\n<td>0.77697</td>\n<td>0.77157</td>\n<td>0.49187</td>\n</tr>\n<tr>\n<td>2</td>\n<td>tenten</td>\n<td>1dConv-Transformer-GRU</td>\n<td>Huberloss</td>\n<td>35M</td>\n<td></td>\n<td>0.77257</td>\n<td>0.77292</td>\n<td>0.76602</td>\n<td>0.09789</td>\n</tr>\n<tr>\n<td>3</td>\n<td>chumajin</td>\n<td>squeeze former</td>\n<td>Huberloss</td>\n<td>16M</td>\n<td>dim 192 + retraining</td>\n<td>0.76961</td>\n<td>0.77050</td>\n<td>0.76239</td>\n<td>0.32106</td>\n</tr>\n<tr>\n<td>4</td>\n<td>sae</td>\n<td>1dConv - GRU</td>\n<td>Huberloss</td>\n<td>16M</td>\n<td></td>\n<td>0.76787</td>\n<td>0.76890</td>\n<td>0.76192</td>\n<td>0.01156</td>\n</tr>\n<tr>\n<td>5</td>\n<td>tenten</td>\n<td>1dUnet</td>\n<td>Huberloss</td>\n<td>73.5M</td>\n<td></td>\n<td>0.76723</td>\n<td>0.76611</td>\n<td>0.76139</td>\n<td>0.07264</td>\n</tr>\n<tr>\n<td>6</td>\n<td>sae</td>\n<td>1dConv - LSTM</td>\n<td>Huberloss</td>\n<td>16M</td>\n<td></td>\n<td>0.76618</td>\n<td>0.76739</td>\n<td>0.76001</td>\n<td>0.18994</td>\n</tr>\n<tr>\n<td>7</td>\n<td>chumajin</td>\n<td>squeeze former</td>\n<td>Huberloss</td>\n<td>16M</td>\n<td>dim 128</td>\n<td>0.76121</td>\n<td>0.76278</td>\n<td>0.75497</td>\n<td>-0.24416</td>\n</tr>\n<tr>\n<td>8</td>\n<td>tenten</td>\n<td>1dConv-Transformer-GRU</td>\n<td>MSEloss</td>\n<td>16M</td>\n<td></td>\n<td>0.76108</td>\n<td>0.75965</td>\n<td>0.75654</td>\n<td>0.07312</td>\n</tr>\n</tbody>\n</table>\n<p><strong>final ensemble cv : 0.78116375, public lb : 0.78052, private lb : 0.77477</strong></p>\n<h1>2. tenten's part</h1>\n<ul>\n<li>overview</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2F19847183370359028731b5040267ca88%2FKaggle_LEAP._tentenpng.png?generation=1722347182292758&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>Input<ul>\n<li>StandardScaler</li>\n<li>Scalars are expanded to match the series length (60).</li></ul></li>\n<li>Model<ul>\n<li>1dUnet<ul>\n<li>down_channels: [128, 256, 512]</li></ul></li>\n<li>1dConv-Transformer-GRU</li></ul></li>\n<li>Training<ul>\n<li>Loss function: HuberLoss(delta=1.0)</li>\n<li>Optimizer: AdamW</li>\n<li>Scheduler: Warmup Cosine Annealing</li>\n<li>Data: All low-resolution data</li></ul></li>\n<li>Postprocess<ul>\n<li>I used torch.min to limit all constant targets except cam_out_FLWDS to non-negative values (I apply this during training before calculating the loss)</li>\n<li>ptend_q0002_x =  -state_q0002_x * 1200 (12≤x≤27)</li></ul></li>\n<li>What didn’t work well<ul>\n<li>Auxiliary loss (latlon, grid id)</li>\n<li>Increase the sequence length by upsampling</li>\n<li>Robust scaler</li>\n<li>Tuning the loss weight </li></ul></li>\n</ul>\n<h1>3. saeNeko's part</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2F2ab9eabeabc302b73dffc382404d358f%2FKaggle_LEAP._sae.png?generation=1722347171155973&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li><p>Models:</p>\n<p>・ 1D CNN + GRU</p>\n<p>・ 1D CNN + LSTM</p></li>\n<li><p>Preprocessing:</p>\n<p>Applied StandardScaler to both features and targets.</p></li>\n<li><p>Training Data:<br>\n・Used Kaggle’s train.csv.<br>\n・6M additional data (Thanks to <a href=\"https://www.kaggle.com/khirokifor\" target=\"_blank\">@khirokifor</a> extracting data from Hugging Face!)</p></li>\n<li><p>Training:</p>\n<p>Trained separately for series targets and scalar targets.</p></li>\n<li><p>Loss Function:</p>\n<p>RMSE</p></li>\n<li><p>Postprocessing:</p>\n<p>For ptendq0002_x(1&lt;=x&lt;=27), used a linear model.<br>\nptend_q0002_x = state_q0002_x * coef + intercept_</p></li>\n</ul>\n<h1>4. chumajin's part</h1>\n<ul>\n<li><p>Preprocess</p>\n<p>・ The diff of the sequential part and StandardScaler<br>\n・ The scalar part was repeated after applying the StandardScaler.</p></li>\n<li><p>Training Data:</p>\n<p>・Used Kaggle’s train.csv.<br>\n・6M additional data</p></li>\n<li><p>Architecture</p>\n<p>SqueezeFormer from the 2nd place solution of the Stanford Ribonanza RNA Folding competition, (<a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/460316\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/460316</a>) </p></li>\n<li><p>Mask of loss</p>\n<p>I did not use the loss for the unpredictable target.</p></li>\n<li><p>Re-training</p>\n<p>Additional 10 epochs fine-tuning after 20 epochs.<br>\ncv 0.767055 → 0.768034112 (I call this chumajin model1.)</p></li>\n<li><p>Insights from my teammates<br>\n・ Huberloss(delta=1)<br>\ncv improved by about 0.01. </p>\n<p>・ Re-training(re-finetune) by separating the loss<br>\nBy training again with limiting the loss, and creating three additional models from the above chumajin model1, I averaged their predictions which was inferred from the limited loss parts of the model with chumajin model1</p>\n<ol>\n<li>model 2 : limited loss of ptend_q0001, q0002, and q0003</li>\n<li>model 3 : limited loss of ptend_t, ptend_u, and ptend_v</li>\n<li>model 4 : limited loss of scalar values </li></ol></li>\n</ul>\n<p>By doing this, I was able to improve my model's CV from 0.7680 to 0.7696. As a result, re-training led to a CV improvement of +0.0026.</p>\n<ul>\n<li><p>Postprocess</p>\n<p>・ Applying tenten's postprocess of torch.min for scaler part.<br>\n・ Applying sae's postprocess, the CV improved by about 0.002.</p></li>\n<li><p>Regrets and Acknowledgements</p>\n<p>With tenten's insights, the improved CV SqueezeFormer was created about a day and a half before the deadline. It took approximately 24 hours to train on 16M train data with an A100, so I couldn't train on all the 73.5M low-resolution data…But I learned a lot expecially from team up. Thank you for team mate!!</p></li>\n</ul>",
      "rawMarkdown": "Thank you for hosting such an interesting competition. We are deeply grateful to the hosts and the Kaggle staff. I would also like to thank , @sae20gorilla , @hidebu, @takatoyoshikawa , and @khiroki for teaming up with me. Even though we formed the team just a week before the close, we had many discussions that were very educational, and it was enjoyable to see our scores and public lb rank improve each time.\n\n# 1. Summary\n\nOur best private solution is an ensemble of 8 models.\n\nAfter forming the team, we started the ensemble using external data for validation that none of us had used for training. Since the relationship between CV and the public LB was favorable in this competition, we calculated the coefficients using the Nelder-Mead method for each model to significantly improve the CV (sub1). Additionally, anticipating the presence of outliers, we used the Nelder-Mead method to calculate coefficients for each column (sub2), but as expected, the private LB for this was not as good. Below, we will describe each part for sub1 in more detail.\n\n| model | part     | model                  | loss      | use train volume | Comment              | cv       | publicScore | privateScore | weights   |\n|-------|----------|------------------------|-----------|------------------|----------------------|----------|-------------|--------------|-----------|\n| 1     | tenten   | 1dConv-Transformer-GRU | Huberloss | 73.5M            |                      | 0.77776  | 0.77697     | 0.77157      | 0.49187   |\n| 2     | tenten   | 1dConv-Transformer-GRU | Huberloss | 35M              |                      | 0.77257  | 0.77292     | 0.76602      | 0.09789   |\n| 3     | chumajin | squeeze former         | Huberloss | 16M              | dim 192 + retraining | 0.76961  | 0.77050     | 0.76239      | 0.32106   |\n| 4     | sae      | 1dConv - GRU           | Huberloss | 16M              |                      | 0.76787  | 0.76890     | 0.76192      | 0.01156   |\n| 5     | tenten   | 1dUnet                 | Huberloss | 73.5M            |                      | 0.76723  | 0.76611     | 0.76139      | 0.07264   |\n| 6     | sae      | 1dConv - LSTM          | Huberloss | 16M              |                      | 0.76618  | 0.76739     | 0.76001      | 0.18994   |\n| 7     | chumajin | squeeze former         | Huberloss | 16M              | dim 128              | 0.76121  | 0.76278     | 0.75497      | -0.24416  |\n| 8     | tenten   | 1dConv-Transformer-GRU | MSEloss   | 16M              |                      | 0.76108  | 0.75965     | 0.75654      | 0.07312   |\n\n\n\n**final ensemble cv : 0.78116375, public lb : 0.78052, private lb : 0.77477**\n\n\n# 2. tenten's part\n\n- overview\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2F19847183370359028731b5040267ca88%2FKaggle_LEAP._tentenpng.png?generation=1722347182292758&alt=media)\n\n- Input\n    - StandardScaler\n    - Scalars are expanded to match the series length (60).\n- Model\n    - 1dUnet\n        - down_channels: [128, 256, 512]\n    - 1dConv-Transformer-GRU\n- Training\n    - Loss function: HuberLoss(delta=1.0)\n    - Optimizer: AdamW\n    - Scheduler: Warmup Cosine Annealing\n    - Data: All low-resolution data\n- Postprocess\n    - I used torch.min to limit all constant targets except cam_out_FLWDS to non-negative values (I apply this during training before calculating the loss)\n    - ptend_q0002_x =  -state_q0002_x * 1200 (12≤x≤27)\n- What didn’t work well\n    - Auxiliary loss (latlon, grid id)\n    - Increase the sequence length by upsampling\n    - Robust scaler\n    - Tuning the loss weight \n\n# 3. saeNeko's part\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2F2ab9eabeabc302b73dffc382404d358f%2FKaggle_LEAP._sae.png?generation=1722347171155973&alt=media)\n\n- Models:\n\n\n    ・ 1D CNN + GRU\n\n\n    ・ 1D CNN + LSTM\n\n- Preprocessing:\n \n    Applied StandardScaler to both features and targets.\n\n- Training Data:\n    ・Used Kaggle’s train.csv.\n    ・6M additional data (Thanks to @khirokifor extracting data from Hugging Face!)\n\n- Training:\n \n    Trained separately for series targets and scalar targets.\n\n- Loss Function:\n \n    RMSE\n\n- Postprocessing:\n \n    For ptendq0002_x(1<=x<=27), used a linear model.\n    ptend_q0002_x = state_q0002_x * coef + intercept_\n\n# 4. chumajin's part\n\n- Preprocess\n\n    ・ The diff of the sequential part and StandardScaler\n    ・ The scalar part was repeated after applying the StandardScaler.\n\n- Training Data:\n    \n    ・Used Kaggle’s train.csv.\n    ・6M additional data\n\n\n- Architecture\n\n    SqueezeFormer from the 2nd place solution of the Stanford Ribonanza RNA Folding competition, (https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/460316) \n\n- Mask of loss\n\n    I did not use the loss for the unpredictable target.\n\n- Re-training\n\n    Additional 10 epochs fine-tuning after 20 epochs.\n    cv 0.767055 → 0.768034112 (I call this chumajin model1.)\n\n- Insights from my teammates\n    ・ Huberloss(delta=1)\n    cv improved by about 0.01. \n\n    ・ Re-training(re-finetune) by separating the loss\n    By training again with limiting the loss, and creating three additional models from the above chumajin model1, I averaged their predictions which was inferred from the limited loss parts of the model with chumajin model1\n\n     1.   model 2 : limited loss of ptend_q0001, q0002, and q0003\n     2.  model 3 : limited loss of ptend_t, ptend_u, and ptend_v\n     3.  model 4 : limited loss of scalar values \n\nBy doing this, I was able to improve my model's CV from 0.7680 to 0.7696. As a result, re-training led to a CV improvement of +0.0026.\n\n- Postprocess\n\n    ・ Applying tenten's postprocess of torch.min for scaler part.\n    ・ Applying sae's postprocess, the CV improved by about 0.002.\n\n- Regrets and Acknowledgements\n\n    With tenten's insights, the improved CV SqueezeFormer was created about a day and a half before the deadline. It took approximately 24 hours to train on 16M train data with an A100, so I couldn't train on all the 73.5M low-resolution data...But I learned a lot expecially from team up. Thank you for team mate!!\n\n",
      "votes": 17
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2940856": "Thank you for hosting such an interesting competition. We are deeply grateful to the hosts and the Kaggle staff. I would also like to thank , @sae20gorilla , @hidebu, @takatoyoshikawa , and @khiroki for teaming up with me. Even though we formed the team just a week before the close, we had many discussions that were very educational, and it was enjoyable to see our scores and public lb rank improve each time.\n\n# 1. Summary\n\nOur best private solution is an ensemble of 8 models.\n\nAfter forming the team, we started the ensemble using external data for validation that none of us had used for training. Since the relationship between CV and the public LB was favorable in this competition, we calculated the coefficients using the Nelder-Mead method for each model to significantly improve the CV (sub1). Additionally, anticipating the presence of outliers, we used the Nelder-Mead method to calculate coefficients for each column (sub2), but as expected, the private LB for this was not as good. Below, we will describe each part for sub1 in more detail.\n\n| model | part     | model                  | loss      | use train volume | Comment              | cv       | publicScore | privateScore | weights   |\n|-------|----------|------------------------|-----------|------------------|----------------------|----------|-------------|--------------|-----------|\n| 1     | tenten   | 1dConv-Transformer-GRU | Huberloss | 73.5M            |                      | 0.77776  | 0.77697     | 0.77157      | 0.49187   |\n| 2     | tenten   | 1dConv-Transformer-GRU | Huberloss | 35M              |                      | 0.77257  | 0.77292     | 0.76602      | 0.09789   |\n| 3     | chumajin | squeeze former         | Huberloss | 16M              | dim 192 + retraining | 0.76961  | 0.77050     | 0.76239      | 0.32106   |\n| 4     | sae      | 1dConv - GRU           | Huberloss | 16M              |                      | 0.76787  | 0.76890     | 0.76192      | 0.01156   |\n| 5     | tenten   | 1dUnet                 | Huberloss | 73.5M            |                      | 0.76723  | 0.76611     | 0.76139      | 0.07264   |\n| 6     | sae      | 1dConv - LSTM          | Huberloss | 16M              |                      | 0.76618  | 0.76739     | 0.76001      | 0.18994   |\n| 7     | chumajin | squeeze former         | Huberloss | 16M              | dim 128              | 0.76121  | 0.76278     | 0.75497      | -0.24416  |\n| 8     | tenten   | 1dConv-Transformer-GRU | MSEloss   | 16M              |                      | 0.76108  | 0.75965     | 0.75654      | 0.07312   |\n\n\n\n**final ensemble cv : 0.78116375, public lb : 0.78052, private lb : 0.77477**\n\n\n# 2. tenten's part\n\n- overview\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2F19847183370359028731b5040267ca88%2FKaggle_LEAP._tentenpng.png?generation=1722347182292758&alt=media)\n\n- Input\n    - StandardScaler\n    - Scalars are expanded to match the series length (60).\n- Model\n    - 1dUnet\n        - down_channels: [128, 256, 512]\n    - 1dConv-Transformer-GRU\n- Training\n    - Loss function: HuberLoss(delta=1.0)\n    - Optimizer: AdamW\n    - Scheduler: Warmup Cosine Annealing\n    - Data: All low-resolution data\n- Postprocess\n    - I used torch.min to limit all constant targets except cam_out_FLWDS to non-negative values (I apply this during training before calculating the loss)\n    - ptend_q0002_x =  -state_q0002_x * 1200 (12≤x≤27)\n- What didn’t work well\n    - Auxiliary loss (latlon, grid id)\n    - Increase the sequence length by upsampling\n    - Robust scaler\n    - Tuning the loss weight \n\n# 3. saeNeko's part\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2F2ab9eabeabc302b73dffc382404d358f%2FKaggle_LEAP._sae.png?generation=1722347171155973&alt=media)\n\n- Models:\n\n\n    ・ 1D CNN + GRU\n\n\n    ・ 1D CNN + LSTM\n\n- Preprocessing:\n \n    Applied StandardScaler to both features and targets.\n\n- Training Data:\n    ・Used Kaggle’s train.csv.\n    ・6M additional data (Thanks to @khirokifor extracting data from Hugging Face!)\n\n- Training:\n \n    Trained separately for series targets and scalar targets.\n\n- Loss Function:\n \n    RMSE\n\n- Postprocessing:\n \n    For ptendq0002_x(1<=x<=27), used a linear model.\n    ptend_q0002_x = state_q0002_x * coef + intercept_\n\n# 4. chumajin's part\n\n- Preprocess\n\n    ・ The diff of the sequential part and StandardScaler\n    ・ The scalar part was repeated after applying the StandardScaler.\n\n- Training Data:\n    \n    ・Used Kaggle’s train.csv.\n    ・6M additional data\n\n\n- Architecture\n\n    SqueezeFormer from the 2nd place solution of the Stanford Ribonanza RNA Folding competition, (https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/460316) \n\n- Mask of loss\n\n    I did not use the loss for the unpredictable target.\n\n- Re-training\n\n    Additional 10 epochs fine-tuning after 20 epochs.\n    cv 0.767055 → 0.768034112 (I call this chumajin model1.)\n\n- Insights from my teammates\n    ・ Huberloss(delta=1)\n    cv improved by about 0.01. \n\n    ・ Re-training(re-finetune) by separating the loss\n    By training again with limiting the loss, and creating three additional models from the above chumajin model1, I averaged their predictions which was inferred from the limited loss parts of the model with chumajin model1\n\n     1.   model 2 : limited loss of ptend_q0001, q0002, and q0003\n     2.  model 3 : limited loss of ptend_t, ptend_u, and ptend_v\n     3.  model 4 : limited loss of scalar values \n\nBy doing this, I was able to improve my model's CV from 0.7680 to 0.7696. As a result, re-training led to a CV improvement of +0.0026.\n\n- Postprocess\n\n    ・ Applying tenten's postprocess of torch.min for scaler part.\n    ・ Applying sae's postprocess, the CV improved by about 0.002.\n\n- Regrets and Acknowledgements\n\n    With tenten's insights, the improved CV SqueezeFormer was created about a day and a half before the deadline. It took approximately 24 hours to train on 16M train data with an A100, so I couldn't train on all the 73.5M low-resolution data...But I learned a lot expecially from team up. Thank you for team mate!!\n\n"
  }
}