{
  "id": 523087,
  "title": "11th Place Solution for the LEAP - Atmospheric Physics using AI (ClimSim) Competition",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/523087",
  "author_name": "Uesugi Erii",
  "post_date": "2024-07-30T07:24:53.776000",
  "votes": 11,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Thanks to Learning the Earth with Artificial Intelligence and Physics (LEAP) NSF Science and Technology Center and Kaggle for hosting this fun competition, which was very stable in CV, LB and PB.</p>\n<p>Considering that my English is not particularly good, this article is being written with the assistance of chatgpt for translation. I will check and make corrections.</p>\n<p>infer code to prove that no leak: <a href=\"https://www.kaggle.com/code/uesugierii/leap-11th-pb0-78491/notebook\" target=\"_blank\">LEAP_11th_PB0.78491</a></p>\n<h2>Context</h2>\n<ul>\n<li><code>Business context</code>: <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/overview\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/overview</a></li>\n<li><code>Data context</code>: <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/data\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/data</a></li>\n</ul>\n<h2>Overview of the approach</h2>\n<p>My final submission consisted of 6 models (with 5 different architectures) integrated together(important). All models use residual structures. All models are trained on the full dataset(low-res) from 0001-02 to 0008-01, and validated on approximately 1.7 million rows of data sampled from 0008-02 to 0009-01. Using both standardization, log transformation and embedding for input process. The models have undergone meticulous training through multiple stages(important). The weights for the ensemble of models were determined using a hill-climbing algorithm base on data from 0008-02 to 0009-01.</p>\n<table>\n<thead>\n<tr>\n<th>Model Architecture</th>\n<th>Number of Block</th>\n<th>Embedding Dimension</th>\n<th>Number of Parameters</th>\n<th>CV</th>\n<th>file name</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>bi-LSTM + MLP</td>\n<td>7</td>\n<td>8</td>\n<td>5070534</td>\n<td>0.71969</td>\n<td><a href=\"https://www.kaggle.com/datasets/uesugierii/leap-final?select=model\" target=\"_blank\">v4_1</a></td>\n</tr>\n<tr>\n<td>bi-LSTM + MLP</td>\n<td>17</td>\n<td>16</td>\n<td>49089054</td>\n<td>0.72703</td>\n<td><a href=\"https://www.kaggle.com/datasets/uesugierii/leap-final?select=model\" target=\"_blank\">v4_1</a></td>\n</tr>\n<tr>\n<td>bi-LSTM + MLP + BN</td>\n<td>7</td>\n<td>8</td>\n<td>5238534</td>\n<td>0.72493</td>\n<td><a href=\"https://www.kaggle.com/datasets/uesugierii/leap-final?select=model\" target=\"_blank\">v4_8</a></td>\n</tr>\n<tr>\n<td>bi-GRU + MLP + BN</td>\n<td>9</td>\n<td>8</td>\n<td>5286134</td>\n<td>0.72536</td>\n<td><a href=\"https://www.kaggle.com/datasets/uesugierii/leap-final?select=model\" target=\"_blank\">v5_1</a></td>\n</tr>\n<tr>\n<td>bi-RNN + MLP + BN</td>\n<td>17</td>\n<td>8</td>\n<td>4511734</td>\n<td>0.71692</td>\n<td><a href=\"https://www.kaggle.com/datasets/uesugierii/leap-final?select=model\" target=\"_blank\">v8_1</a></td>\n</tr>\n<tr>\n<td>1D-CNN + BN</td>\n<td>11</td>\n<td>8</td>\n<td>16854334</td>\n<td>0.71459</td>\n<td><a href=\"https://www.kaggle.com/datasets/uesugierii/leap-final?select=model\" target=\"_blank\">v6</a></td>\n</tr>\n<tr>\n<td>Ensemble</td>\n<td></td>\n<td></td>\n<td></td>\n<td>0.73472 (PB: 78491)</td>\n<td></td>\n</tr>\n</tbody>\n</table>\n<h2>Solution Detail</h2>\n<h3>data process detail</h3>\n<p>Based on my adversarial verification experiments, there is a significant difference(AUC=1) in data between years. According to the organizer's information, there should be one billion pieces of data in theory (but there are only about 80 million on HuggingFace). Therefore, it can be speculated that the test set is from data in the later years. Considering the above information, I used data from 0001-02 to 0008-01 for training, and data from 0008-02 to 0009-01 for validation. Perhaps this is why I was able to climb up the PB ranking and enter the gold medal zone.</p>\n<p>The original input had 556 columns. I removed 66 columns that only had one value, leaving 490 columns. For each column of data, I perform two operations: one is to directly standardize it <code>(X - x_mean) / x_std</code>, and the other is to apply log after taking the absolute value and then standardize it (CV +0.003). <code>(log_X - log_x_mean) / log_x_std</code> So the shape of the input for my model is (980,)</p>\n<p>For each of the 980 inputs to the model, an embedding is assigned with an initialization range of [-0.05, 0.05]. This initialization range is <strong>crucial</strong> as other initialization methods may result in poorer performance.</p>\n<p>By padding and reshape, the input eventually becomes (bs, 60, 25*emb_dim).</p>\n<h3>model detail (important)</h3>\n<p>For <code>bi-LSTM + MLP + BN</code>, <code>bi-GRU + MLP + BN</code>, <code>bi-RNN + MLP + BN</code>, The model structures are all arranged in the following manner, with the only difference being the encoding layer.</p>\n<pre><code>pre_seq = seq  \n i  ((self.encoder_layers)):\n    encoder_layer = self.encoder_layers[i]\n    seq, _ = encoder_layer(seq)  \n    ffn = self.mlp[i]\n    seq = ffn(seq)  \n    seq = seq.contiguous().view(bs, -)  \n    seq = self.mlp_bn[i](seq)  \n    seq = seq.contiguous().view(bs, , -)  \n    seq = F.gelu(seq)\n    seq = seq + pre_seq  \n    pre_seq = seq\n</code></pre>\n<p>For <code>1D-CNN + BN</code>, model structure as below</p>\n<pre><code>pre_x = x\n i  ((self.cnns)):\n    x = self.cnns[i](x) + pre_x\n    pre_x = x\n</code></pre>\n<p>The complete code for the model can be found in the 'model' folder here. <a href=\"https://www.kaggle.com/datasets/uesugierii/leap-final?select=model\" target=\"_blank\">https://www.kaggle.com/datasets/uesugierii/leap-final?select=model</a></p>\n<h3>training detail (important)</h3>\n<p>During the training process, it was found that training directly on the full dataset (low-res) resulted in lower scores. After performing EDA, it was discovered that the main cause of the unstable training was outliers.</p>\n<p>To address this, I employed two independent methods for mitigation: </p>\n<ol>\n<li>First method involved clipping the value range of the input and target(<code>np.clip(x, -10, 10)</code>, <code>np.clip(y, -10, 10)</code>), first fitting on a smaller range (15 epochs, lr 1e-3), and then fitting on the original data (15 epochs, lr 1e-4). The rationale behind this is to first allow the model to learn and adapt to relatively normal weather conditions before generalizing to more extreme weather variations.</li>\n<li>Second method: first let the model train on one month's data(15 epochs, lr 1e-3), then on one year's data(15 epochs, lr 1e-3), and finally on all available data(15 epochs, lr 1e-4). The main idea behind this is to have the model initially learn the climate variations over short time spans, then gradually expand the time range to ensure a smoother learning process and reduce the difficulty of learning.</li>\n</ol>\n<p>To gain a slight improvement in performance towards the end, I would train the model for a few epochs on an extremely large batch size (10240) and a very low learning rate (1e-5). Some models were able to achieve further enhancements (CV +0.004).</p>\n<h3>What didn't work for me</h3>\n<ul>\n<li>Physics knowledge, neural network methods for solving partial differential equations<ul>\n<li>I'm not sure if it's my lack of understanding of these knowledge that's causing the lack of effect, or if there really don't work</li></ul></li>\n<li>U-Net, U-Net++</li>\n<li>Transformers</li>\n<li>Squeezeformer</li>\n<li>Assist in training using target columns with weight 0</li>\n<li>More complex embedding schemes, such as bucketing + embedding, AutoDis, and so on</li>\n</ul>\n<h2>Sources</h2>\n<p>ptend trick: <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/499896#2791290\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/499896#2791290</a></p>",
  "messages": [
    {
      "id": 2940546,
      "postDate": "2024-07-30T07:24:53.777Z",
      "content": "<p>Thanks to Learning the Earth with Artificial Intelligence and Physics (LEAP) NSF Science and Technology Center and Kaggle for hosting this fun competition, which was very stable in CV, LB and PB.</p>\n<p>Considering that my English is not particularly good, this article is being written with the assistance of chatgpt for translation. I will check and make corrections.</p>\n<p>infer code to prove that no leak: <a href=\"https://www.kaggle.com/code/uesugierii/leap-11th-pb0-78491/notebook\" target=\"_blank\">LEAP_11th_PB0.78491</a></p>\n<h2>Context</h2>\n<ul>\n<li><code>Business context</code>: <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/overview\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/overview</a></li>\n<li><code>Data context</code>: <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/data\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/data</a></li>\n</ul>\n<h2>Overview of the approach</h2>\n<p>My final submission consisted of 6 models (with 5 different architectures) integrated together(important). All models use residual structures. All models are trained on the full dataset(low-res) from 0001-02 to 0008-01, and validated on approximately 1.7 million rows of data sampled from 0008-02 to 0009-01. Using both standardization, log transformation and embedding for input process. The models have undergone meticulous training through multiple stages(important). The weights for the ensemble of models were determined using a hill-climbing algorithm base on data from 0008-02 to 0009-01.</p>\n<table>\n<thead>\n<tr>\n<th>Model Architecture</th>\n<th>Number of Block</th>\n<th>Embedding Dimension</th>\n<th>Number of Parameters</th>\n<th>CV</th>\n<th>file name</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>bi-LSTM + MLP</td>\n<td>7</td>\n<td>8</td>\n<td>5070534</td>\n<td>0.71969</td>\n<td><a href=\"https://www.kaggle.com/datasets/uesugierii/leap-final?select=model\" target=\"_blank\">v4_1</a></td>\n</tr>\n<tr>\n<td>bi-LSTM + MLP</td>\n<td>17</td>\n<td>16</td>\n<td>49089054</td>\n<td>0.72703</td>\n<td><a href=\"https://www.kaggle.com/datasets/uesugierii/leap-final?select=model\" target=\"_blank\">v4_1</a></td>\n</tr>\n<tr>\n<td>bi-LSTM + MLP + BN</td>\n<td>7</td>\n<td>8</td>\n<td>5238534</td>\n<td>0.72493</td>\n<td><a href=\"https://www.kaggle.com/datasets/uesugierii/leap-final?select=model\" target=\"_blank\">v4_8</a></td>\n</tr>\n<tr>\n<td>bi-GRU + MLP + BN</td>\n<td>9</td>\n<td>8</td>\n<td>5286134</td>\n<td>0.72536</td>\n<td><a href=\"https://www.kaggle.com/datasets/uesugierii/leap-final?select=model\" target=\"_blank\">v5_1</a></td>\n</tr>\n<tr>\n<td>bi-RNN + MLP + BN</td>\n<td>17</td>\n<td>8</td>\n<td>4511734</td>\n<td>0.71692</td>\n<td><a href=\"https://www.kaggle.com/datasets/uesugierii/leap-final?select=model\" target=\"_blank\">v8_1</a></td>\n</tr>\n<tr>\n<td>1D-CNN + BN</td>\n<td>11</td>\n<td>8</td>\n<td>16854334</td>\n<td>0.71459</td>\n<td><a href=\"https://www.kaggle.com/datasets/uesugierii/leap-final?select=model\" target=\"_blank\">v6</a></td>\n</tr>\n<tr>\n<td>Ensemble</td>\n<td></td>\n<td></td>\n<td></td>\n<td>0.73472 (PB: 78491)</td>\n<td></td>\n</tr>\n</tbody>\n</table>\n<h2>Solution Detail</h2>\n<h3>data process detail</h3>\n<p>Based on my adversarial verification experiments, there is a significant difference(AUC=1) in data between years. According to the organizer's information, there should be one billion pieces of data in theory (but there are only about 80 million on HuggingFace). Therefore, it can be speculated that the test set is from data in the later years. Considering the above information, I used data from 0001-02 to 0008-01 for training, and data from 0008-02 to 0009-01 for validation. Perhaps this is why I was able to climb up the PB ranking and enter the gold medal zone.</p>\n<p>The original input had 556 columns. I removed 66 columns that only had one value, leaving 490 columns. For each column of data, I perform two operations: one is to directly standardize it <code>(X - x_mean) / x_std</code>, and the other is to apply log after taking the absolute value and then standardize it (CV +0.003). <code>(log_X - log_x_mean) / log_x_std</code> So the shape of the input for my model is (980,)</p>\n<p>For each of the 980 inputs to the model, an embedding is assigned with an initialization range of [-0.05, 0.05]. This initialization range is <strong>crucial</strong> as other initialization methods may result in poorer performance.</p>\n<p>By padding and reshape, the input eventually becomes (bs, 60, 25*emb_dim).</p>\n<h3>model detail (important)</h3>\n<p>For <code>bi-LSTM + MLP + BN</code>, <code>bi-GRU + MLP + BN</code>, <code>bi-RNN + MLP + BN</code>, The model structures are all arranged in the following manner, with the only difference being the encoding layer.</p>\n<pre><code>pre_seq = seq  \n i  ((self.encoder_layers)):\n    encoder_layer = self.encoder_layers[i]\n    seq, _ = encoder_layer(seq)  \n    ffn = self.mlp[i]\n    seq = ffn(seq)  \n    seq = seq.contiguous().view(bs, -)  \n    seq = self.mlp_bn[i](seq)  \n    seq = seq.contiguous().view(bs, , -)  \n    seq = F.gelu(seq)\n    seq = seq + pre_seq  \n    pre_seq = seq\n</code></pre>\n<p>For <code>1D-CNN + BN</code>, model structure as below</p>\n<pre><code>pre_x = x\n i  ((self.cnns)):\n    x = self.cnns[i](x) + pre_x\n    pre_x = x\n</code></pre>\n<p>The complete code for the model can be found in the 'model' folder here. <a href=\"https://www.kaggle.com/datasets/uesugierii/leap-final?select=model\" target=\"_blank\">https://www.kaggle.com/datasets/uesugierii/leap-final?select=model</a></p>\n<h3>training detail (important)</h3>\n<p>During the training process, it was found that training directly on the full dataset (low-res) resulted in lower scores. After performing EDA, it was discovered that the main cause of the unstable training was outliers.</p>\n<p>To address this, I employed two independent methods for mitigation: </p>\n<ol>\n<li>First method involved clipping the value range of the input and target(<code>np.clip(x, -10, 10)</code>, <code>np.clip(y, -10, 10)</code>), first fitting on a smaller range (15 epochs, lr 1e-3), and then fitting on the original data (15 epochs, lr 1e-4). The rationale behind this is to first allow the model to learn and adapt to relatively normal weather conditions before generalizing to more extreme weather variations.</li>\n<li>Second method: first let the model train on one month's data(15 epochs, lr 1e-3), then on one year's data(15 epochs, lr 1e-3), and finally on all available data(15 epochs, lr 1e-4). The main idea behind this is to have the model initially learn the climate variations over short time spans, then gradually expand the time range to ensure a smoother learning process and reduce the difficulty of learning.</li>\n</ol>\n<p>To gain a slight improvement in performance towards the end, I would train the model for a few epochs on an extremely large batch size (10240) and a very low learning rate (1e-5). Some models were able to achieve further enhancements (CV +0.004).</p>\n<h3>What didn't work for me</h3>\n<ul>\n<li>Physics knowledge, neural network methods for solving partial differential equations<ul>\n<li>I'm not sure if it's my lack of understanding of these knowledge that's causing the lack of effect, or if there really don't work</li></ul></li>\n<li>U-Net, U-Net++</li>\n<li>Transformers</li>\n<li>Squeezeformer</li>\n<li>Assist in training using target columns with weight 0</li>\n<li>More complex embedding schemes, such as bucketing + embedding, AutoDis, and so on</li>\n</ul>\n<h2>Sources</h2>\n<p>ptend trick: <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/499896#2791290\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/499896#2791290</a></p>",
      "rawMarkdown": "Thanks to Learning the Earth with Artificial Intelligence and Physics (LEAP) NSF Science and Technology Center and Kaggle for hosting this fun competition, which was very stable in CV, LB and PB.\n\nConsidering that my English is not particularly good, this article is being written with the assistance of chatgpt for translation. I will check and make corrections.\n\ninfer code to prove that no leak: [LEAP_11th_PB0.78491](https://www.kaggle.com/code/uesugierii/leap-11th-pb0-78491/notebook)\n\n## Context\n\n+ `Business context`: [https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/overview](https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/overview)\n+ `Data context`: [https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/data](https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/data)\n\n## Overview of the approach\n\nMy final submission consisted of 6 models (with 5 different architectures) integrated together(important). All models use residual structures. All models are trained on the full dataset(low-res) from 0001-02 to 0008-01, and validated on approximately 1.7 million rows of data sampled from 0008-02 to 0009-01. Using both standardization, log transformation and embedding for input process. The models have undergone meticulous training through multiple stages(important). The weights for the ensemble of models were determined using a hill-climbing algorithm base on data from 0008-02 to 0009-01.\n\n| Model Architecture | Number of Block | Embedding Dimension | Number of Parameters |   CV    |                                 file name                                  |  \n|:------------------:|:---------------:|:-------------------:|:--------------------:|:-------:|:--------------------------------------------------------------------------:|  \n|   bi-LSTM + MLP    |        7        |          8          |       5070534        | 0.71969 | [v4_1](https://www.kaggle.com/datasets/uesugierii/leap-final?select=model) |  \n|   bi-LSTM + MLP    |       17        |         16          |       49089054       | 0.72703 | [v4_1](https://www.kaggle.com/datasets/uesugierii/leap-final?select=model) |  \n| bi-LSTM + MLP + BN |        7        |          8          |       5238534        | 0.72493 | [v4_8](https://www.kaggle.com/datasets/uesugierii/leap-final?select=model) |  \n| bi-GRU + MLP + BN  |        9        |          8          |       5286134        | 0.72536 | [v5_1](https://www.kaggle.com/datasets/uesugierii/leap-final?select=model) |  \n| bi-RNN + MLP + BN  |       17        |          8          |       4511734        | 0.71692 | [v8_1](https://www.kaggle.com/datasets/uesugierii/leap-final?select=model) |  \n|    1D-CNN + BN     |       11        |          8          |       16854334       | 0.71459 |  [v6](https://www.kaggle.com/datasets/uesugierii/leap-final?select=model)  |  \n|      Ensemble      |                 |                     |                      | 0.73472 (PB: 78491) |                                                                            | \n\n## Solution Detail\n\n### data process detail\n\nBased on my adversarial verification experiments, there is a significant difference(AUC=1) in data between years. According to the organizer's information, there should be one billion pieces of data in theory (but there are only about 80 million on HuggingFace). Therefore, it can be speculated that the test set is from data in the later years. Considering the above information, I used data from 0001-02 to 0008-01 for training, and data from 0008-02 to 0009-01 for validation. Perhaps this is why I was able to climb up the PB ranking and enter the gold medal zone.\n\nThe original input had 556 columns. I removed 66 columns that only had one value, leaving 490 columns. For each column of data, I perform two operations: one is to directly standardize it `(X - x_mean) / x_std`, and the other is to apply log after taking the absolute value and then standardize it (CV +0.003). `(log_X - log_x_mean) / log_x_std` So the shape of the input for my model is (980,)\n\nFor each of the 980 inputs to the model, an embedding is assigned with an initialization range of [-0.05, 0.05]. This initialization range is **crucial** as other initialization methods may result in poorer performance.\n\nBy padding and reshape, the input eventually becomes (bs, 60, 25*emb_dim).\n\n### model detail (important)\n\nFor `bi-LSTM + MLP + BN`, `bi-GRU + MLP + BN`, `bi-RNN + MLP + BN`, The model structures are all arranged in the following manner, with the only difference being the encoding layer.\n\n```python\npre_seq = seq  # (bs, 60, 25*emb_dim)\nfor i in range(len(self.encoder_layers)):\n    encoder_layer = self.encoder_layers[i]\n    seq, _ = encoder_layer(seq)  # (bs, 60, 25*emb_dim*2)\n    ffn = self.mlp[i]\n    seq = ffn(seq)  # (bs, 60, 25*emb_dim)\n    seq = seq.contiguous().view(bs, -1)  # (bs, 60*25*emb_dim)\n    seq = self.mlp_bn[i](seq)  # (bs, 60*25*emb_dim)\n    seq = seq.contiguous().view(bs, 60, -1)  # (bs, 60, 25*emb_dim)\n    seq = F.gelu(seq)\n    seq = seq + pre_seq  # (bs, 60*25*emb_dim)\n    pre_seq = seq\n```\n\nFor `1D-CNN + BN`, model structure as below\n\n```python\npre_x = x\nfor i in range(len(self.cnns)):\n    x = self.cnns[i](x) + pre_x\n    pre_x = x\n```\n\nThe complete code for the model can be found in the 'model' folder here. [https://www.kaggle.com/datasets/uesugierii/leap-final?select=model](https://www.kaggle.com/datasets/uesugierii/leap-final?select=model)\n\n### training detail (important)\n\nDuring the training process, it was found that training directly on the full dataset (low-res) resulted in lower scores. After performing EDA, it was discovered that the main cause of the unstable training was outliers.\n\nTo address this, I employed two independent methods for mitigation: \n\n1. First method involved clipping the value range of the input and target(`np.clip(x, -10, 10)`, `np.clip(y, -10, 10)`), first fitting on a smaller range (15 epochs, lr 1e-3), and then fitting on the original data (15 epochs, lr 1e-4). The rationale behind this is to first allow the model to learn and adapt to relatively normal weather conditions before generalizing to more extreme weather variations.\n2. Second method: first let the model train on one month's data(15 epochs, lr 1e-3), then on one year's data(15 epochs, lr 1e-3), and finally on all available data(15 epochs, lr 1e-4). The main idea behind this is to have the model initially learn the climate variations over short time spans, then gradually expand the time range to ensure a smoother learning process and reduce the difficulty of learning.\n\nTo gain a slight improvement in performance towards the end, I would train the model for a few epochs on an extremely large batch size (10240) and a very low learning rate (1e-5). Some models were able to achieve further enhancements (CV +0.004).\n\n### What didn't work for me\n\n+ Physics knowledge, neural network methods for solving partial differential equations\n    + I'm not sure if it's my lack of understanding of these knowledge that's causing the lack of effect, or if there really don't work\n+ U-Net, U-Net++\n+ Transformers\n+ Squeezeformer\n+ Assist in training using target columns with weight 0\n+ More complex embedding schemes, such as bucketing + embedding, AutoDis, and so on\n\n## Sources\n\nptend trick: https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/499896#2791290\n",
      "votes": 11
    },
    {
      "id": 2940710,
      "postDate": "2024-07-30T11:28:45.077Z",
      "content": "<p>Congrats on your 11th place finish, <a href=\"https://www.kaggle.com/uesugierii\" target=\"_blank\">@uesugierii</a>!<br>\nYour approach using multiple models and detailed training process is impressive. It's clear you put a lot of effort into optimizing your solution. Thanks for sharing your insights and methods.</p>",
      "rawMarkdown": "Congrats on your 11th place finish, @uesugierii!\nYour approach using multiple models and detailed training process is impressive. It's clear you put a lot of effort into optimizing your solution. Thanks for sharing your insights and methods."
    }
  ],
  "comments": [
    {
      "id": 2940710,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-07-30T11:28:45.077000",
      "content": "<p>Congrats on your 11th place finish, <a href=\"https://www.kaggle.com/uesugierii\" target=\"_blank\">@uesugierii</a>!<br>\nYour approach using multiple models and detailed training process is impressive. It's clear you put a lot of effort into optimizing your solution. Thanks for sharing your insights and methods.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2940546": "Thanks to Learning the Earth with Artificial Intelligence and Physics (LEAP) NSF Science and Technology Center and Kaggle for hosting this fun competition, which was very stable in CV, LB and PB.\n\nConsidering that my English is not particularly good, this article is being written with the assistance of chatgpt for translation. I will check and make corrections.\n\ninfer code to prove that no leak: [LEAP_11th_PB0.78491](https://www.kaggle.com/code/uesugierii/leap-11th-pb0-78491/notebook)\n\n## Context\n\n+ `Business context`: [https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/overview](https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/overview)\n+ `Data context`: [https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/data](https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/data)\n\n## Overview of the approach\n\nMy final submission consisted of 6 models (with 5 different architectures) integrated together(important). All models use residual structures. All models are trained on the full dataset(low-res) from 0001-02 to 0008-01, and validated on approximately 1.7 million rows of data sampled from 0008-02 to 0009-01. Using both standardization, log transformation and embedding for input process. The models have undergone meticulous training through multiple stages(important). The weights for the ensemble of models were determined using a hill-climbing algorithm base on data from 0008-02 to 0009-01.\n\n| Model Architecture | Number of Block | Embedding Dimension | Number of Parameters |   CV    |                                 file name                                  |  \n|:------------------:|:---------------:|:-------------------:|:--------------------:|:-------:|:--------------------------------------------------------------------------:|  \n|   bi-LSTM + MLP    |        7        |          8          |       5070534        | 0.71969 | [v4_1](https://www.kaggle.com/datasets/uesugierii/leap-final?select=model) |  \n|   bi-LSTM + MLP    |       17        |         16          |       49089054       | 0.72703 | [v4_1](https://www.kaggle.com/datasets/uesugierii/leap-final?select=model) |  \n| bi-LSTM + MLP + BN |        7        |          8          |       5238534        | 0.72493 | [v4_8](https://www.kaggle.com/datasets/uesugierii/leap-final?select=model) |  \n| bi-GRU + MLP + BN  |        9        |          8          |       5286134        | 0.72536 | [v5_1](https://www.kaggle.com/datasets/uesugierii/leap-final?select=model) |  \n| bi-RNN + MLP + BN  |       17        |          8          |       4511734        | 0.71692 | [v8_1](https://www.kaggle.com/datasets/uesugierii/leap-final?select=model) |  \n|    1D-CNN + BN     |       11        |          8          |       16854334       | 0.71459 |  [v6](https://www.kaggle.com/datasets/uesugierii/leap-final?select=model)  |  \n|      Ensemble      |                 |                     |                      | 0.73472 (PB: 78491) |                                                                            | \n\n## Solution Detail\n\n### data process detail\n\nBased on my adversarial verification experiments, there is a significant difference(AUC=1) in data between years. According to the organizer's information, there should be one billion pieces of data in theory (but there are only about 80 million on HuggingFace). Therefore, it can be speculated that the test set is from data in the later years. Considering the above information, I used data from 0001-02 to 0008-01 for training, and data from 0008-02 to 0009-01 for validation. Perhaps this is why I was able to climb up the PB ranking and enter the gold medal zone.\n\nThe original input had 556 columns. I removed 66 columns that only had one value, leaving 490 columns. For each column of data, I perform two operations: one is to directly standardize it `(X - x_mean) / x_std`, and the other is to apply log after taking the absolute value and then standardize it (CV +0.003). `(log_X - log_x_mean) / log_x_std` So the shape of the input for my model is (980,)\n\nFor each of the 980 inputs to the model, an embedding is assigned with an initialization range of [-0.05, 0.05]. This initialization range is **crucial** as other initialization methods may result in poorer performance.\n\nBy padding and reshape, the input eventually becomes (bs, 60, 25*emb_dim).\n\n### model detail (important)\n\nFor `bi-LSTM + MLP + BN`, `bi-GRU + MLP + BN`, `bi-RNN + MLP + BN`, The model structures are all arranged in the following manner, with the only difference being the encoding layer.\n\n```python\npre_seq = seq  # (bs, 60, 25*emb_dim)\nfor i in range(len(self.encoder_layers)):\n    encoder_layer = self.encoder_layers[i]\n    seq, _ = encoder_layer(seq)  # (bs, 60, 25*emb_dim*2)\n    ffn = self.mlp[i]\n    seq = ffn(seq)  # (bs, 60, 25*emb_dim)\n    seq = seq.contiguous().view(bs, -1)  # (bs, 60*25*emb_dim)\n    seq = self.mlp_bn[i](seq)  # (bs, 60*25*emb_dim)\n    seq = seq.contiguous().view(bs, 60, -1)  # (bs, 60, 25*emb_dim)\n    seq = F.gelu(seq)\n    seq = seq + pre_seq  # (bs, 60*25*emb_dim)\n    pre_seq = seq\n```\n\nFor `1D-CNN + BN`, model structure as below\n\n```python\npre_x = x\nfor i in range(len(self.cnns)):\n    x = self.cnns[i](x) + pre_x\n    pre_x = x\n```\n\nThe complete code for the model can be found in the 'model' folder here. [https://www.kaggle.com/datasets/uesugierii/leap-final?select=model](https://www.kaggle.com/datasets/uesugierii/leap-final?select=model)\n\n### training detail (important)\n\nDuring the training process, it was found that training directly on the full dataset (low-res) resulted in lower scores. After performing EDA, it was discovered that the main cause of the unstable training was outliers.\n\nTo address this, I employed two independent methods for mitigation: \n\n1. First method involved clipping the value range of the input and target(`np.clip(x, -10, 10)`, `np.clip(y, -10, 10)`), first fitting on a smaller range (15 epochs, lr 1e-3), and then fitting on the original data (15 epochs, lr 1e-4). The rationale behind this is to first allow the model to learn and adapt to relatively normal weather conditions before generalizing to more extreme weather variations.\n2. Second method: first let the model train on one month's data(15 epochs, lr 1e-3), then on one year's data(15 epochs, lr 1e-3), and finally on all available data(15 epochs, lr 1e-4). The main idea behind this is to have the model initially learn the climate variations over short time spans, then gradually expand the time range to ensure a smoother learning process and reduce the difficulty of learning.\n\nTo gain a slight improvement in performance towards the end, I would train the model for a few epochs on an extremely large batch size (10240) and a very low learning rate (1e-5). Some models were able to achieve further enhancements (CV +0.004).\n\n### What didn't work for me\n\n+ Physics knowledge, neural network methods for solving partial differential equations\n    + I'm not sure if it's my lack of understanding of these knowledge that's causing the lack of effect, or if there really don't work\n+ U-Net, U-Net++\n+ Transformers\n+ Squeezeformer\n+ Assist in training using target columns with weight 0\n+ More complex embedding schemes, such as bucketing + embedding, AutoDis, and so on\n\n## Sources\n\nptend trick: https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/499896#2791290\n",
    "2940710": "Congrats on your 11th place finish, @uesugierii!\nYour approach using multiple models and detailed training process is impressive. It's clear you put a lot of effort into optimizing your solution. Thanks for sharing your insights and methods."
  }
}