{
  "id": 523223,
  "title": "8th solution",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/523223",
  "author_name": "heng",
  "post_date": "2024-07-31T03:05:03.111000",
  "votes": 25,
  "comment_count": 5,
  "views": 0,
  "content": "<p>First of all, I would like to thank the organizers and Kaggle for hosting this competition. The quality of the competition data is great. Although there were some problems during the process, in any case, we have achieved results that satisfy most people.</p>\n<p>This is my first solo gold medal and I became the Competition GrandMaster. This seven-year journey has been quite long and exciting.</p>\n<h3>solution summary</h3>\n<p>I think my solution is very simple, and it is basically based on the seq2seq models derived from BiLSTM.</p>\n<p>(bs, 60, 25)  --&gt; seq2seq --&gt; (bs, 60, 14) --&gt; (bs, 368) </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1197382%2Fa02dd80af723ab676afa94c4f8f58b63%2FScreenshot%202024-07-31_10-19-57-734.png?generation=1722392449416839&amp;alt=media\" alt=\"\"></p>\n<h3>models</h3>\n<ul>\n<li>validate on last 6 months sample data</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>models</th>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>BiLSTM(layers=6)</td>\n<td>0.7844</td>\n<td>0.7812</td>\n</tr>\n<tr>\n<td>BiGRU(layers=8)</td>\n<td>0.7835</td>\n<td>0.7802</td>\n</tr>\n<tr>\n<td>BiLSTM+Transformer</td>\n<td>0.7858</td>\n<td>0.7821</td>\n</tr>\n<tr>\n<td>BiLSTM+Attention</td>\n<td>0.7865</td>\n<td>0.7834</td>\n</tr>\n<tr>\n<td>BiLSTM+TCN</td>\n<td>0.7855</td>\n<td>0.7832</td>\n</tr>\n<tr>\n<td>BiLSTM+CNN</td>\n<td>0.7842</td>\n<td>0.7821</td>\n</tr>\n<tr>\n<td>ensemble on models</td>\n<td>0.7923</td>\n<td>0.7890</td>\n</tr>\n<tr>\n<td>ensemble on targets</td>\n<td>0.7933</td>\n<td>0.7884</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>A base BiLSTM model:</li>\n</ul>\n<pre><code>class LeapModel(nn.Module):\n    def __init__(self,\n                 input_size,\n                 seq_len,\n                 hidden_size,\n                 output_size,\n                 =1,\n                 =,\n                 =.3,\n                 hidden_layers=[128, 256]):\n\n        super().__init__()\n        self.input_size = input_size\n        self.seq_len = seq_len\n        self.hidden_size = hidden_size\n        self.num_layers = num_layers\n        self.=bidirectional\n        self.=output_size\n\n        self.rnn = nn.LSTM(=input_size,\n                           =hidden_size,\n                           =num_layers,\n                           =bidirectional,\n                           =,\n                           =dropout)\n\n         hidden_layers  len(hidden_layers):\n            first_layer  = nn.Linear(hidden_size  bidirectional  hidden_size, hidden_layers[0])\n            self.hidden_layers = nn.ModuleList(\n                [first_layer] + \\\n                [nn.Linear(hidden_layers[i], hidden_layers[i+1])  i  range(len(hidden_layers) - 1)]\n            )\n             layer  self.hidden_layers:\n                nn.init.kaiming_normal_(layer.weight.data)\n            self.intermediate_layer = nn.Linear(hidden_layers[-1], self.input_size)\n            self.output_layer = nn.Linear(hidden_layers[-1], output_size)\n            nn.init.kaiming_normal_(self.output_layer.weight.data)\n        :\n            self.hidden_layers = []\n            self.intermediate_layer = nn.Linear(hidden_size  bidirectional  hidden_size, self.input_size)\n            self.output_layer = nn.Linear(hidden_size  bidirectional  hidden_size, output_size)\n            nn.init.kaiming_normal_(self.output_layer.weight.data)\n\n        self.activation_fn = torch.nn.GELU()\n        self.dropout = nn.Dropout(dropout)\n\n    def forward(self, x):\n        outputs, hidden = self.rnn(x)\n\n        x = self.dropout(self.activation_fn(outputs))\n         hidden_layer  self.hidden_layers:\n            x = self.activation_fn(hidden_layer(x))\n            x = self.dropout(x)\n        x = self.output_layer(x)\n\n        # (-1,60,14) -&gt; (-1,386)\n        o_s = x[:, :, :6]\n        o_s = o_s.permute(0,2,1).reshape(-1,360)\n        o_g = x[:, :, 6:]\n        o_g = o_g.mean(=1)\n        out = torch.cat([o_s, o_g], =1)\n\n        return out\n\ninput_size = 25\noutput_size = 14\nseq_len = 60\n\nhidden_size = 256\nhidden_layers = [256, 512]\nnum_layers = 6\ndropout = 0.1\n\nmodel = LeapModel(\n    =input_size,\n    =seq_len,\n    =hidden_size,\n    =output_size,\n    =num_layers,\n    =hidden_layers,\n    =dropout,\n    =,\n).(device)\n</code></pre>\n<p>Reference Links: <a href=\"https://www.kaggle.com/code/brandenkmurray/seq2seq-rnn-with-gru\" target=\"_blank\">https://www.kaggle.com/code/brandenkmurray/seq2seq-rnn-with-gru</a></p>\n<ul>\n<li>A BiLSTM with Transformer model</li>\n</ul>\n<pre><code>class LeapModel(nn.Module):\n    def __init__(self,\n                 input_size,\n                 seq_len,\n                 hidden_size,\n                 output_size,\n                 =1,\n                 =,\n                 =0.3,\n                 hidden_layers=[128, 256],\n                 =8,\n                 =2):\n\n        super().__init__()\n        self.input_size = input_size\n        self.seq_len = seq_len\n        self.hidden_size = hidden_size\n        self.num_layers = num_layers\n        self.bidirectional = bidirectional\n        self.output_size = output_size\n\n        # LSTM layer\n        self.rnn = nn.LSTM(=input_size,\n                           =hidden_size,\n                           =num_layers,\n                           =bidirectional,\n                           =,\n                           =dropout)\n\n        # Transformer layer\n        transformer_input_size = hidden_size * 2  bidirectional  hidden_size\n        self.transformer_layer = nn.TransformerEncoder(\n            nn.TransformerEncoderLayer(=transformer_input_size, =nhead, =dropout),\n            =num_transformer_layers\n        )\n\n        # Fully connected layers\n         hidden_layers  len(hidden_layers):\n            first_layer = nn.Linear(transformer_input_size, hidden_layers[0])\n            self.hidden_layers = nn.ModuleList(\n                [first_layer] + \\\n                [nn.Linear(hidden_layers[i], hidden_layers[i+1])  i  range(len(hidden_layers) - 1)]\n            )\n             layer  self.hidden_layers:\n                nn.init.kaiming_normal_(layer.weight.data)\n            self.intermediate_layer = nn.Linear(hidden_layers[-1], self.input_size)\n            self.output_layer = nn.Linear(hidden_layers[-1], output_size)\n            nn.init.kaiming_normal_(self.output_layer.weight.data)\n        :\n            self.hidden_layers = []\n            self.intermediate_layer = nn.Linear(transformer_input_size, self.input_size)\n            self.output_layer = nn.Linear(transformer_input_size, output_size)\n            nn.init.kaiming_normal_(self.output_layer.weight.data)\n\n        self.activation_fn = torch.nn.GELU()\n        self.dropout = nn.Dropout(dropout)\n\n    def forward(self, x):\n        # LSTM layer\n        lstm_output, _ = self.rnn(x)\n\n        # Transformer layer\n        transformer_output = self.transformer_layer(lstm_output)\n\n        # Apply dropout  activation\n        x = self.dropout(self.activation_fn(transformer_output))\n\n        # Fully connected layers\n         hidden_layer  self.hidden_layers:\n            x = self.activation_fn(hidden_layer(x))\n            x = self.dropout(x)\n        x = self.output_layer(x)\n\n        # Reshape output\n        o_s = x[:, :, :6]\n        o_s = o_s.permute(0, 2, 1).reshape(-1, 360)\n        o_g = x[:, :, 6:]\n        o_g = o_g.mean(=1)\n        out = torch.cat([o_s, o_g], =1)\n\n        return out\n\ninput_size = 25\noutput_size = 14\nseq_len = 60\n\nhidden_size = 256\nhidden_layers = [256, 512]\nnum_layers = 6\ndropout = 0.1\nnhead = 8\nnum_transformer_layers = 1\n\nmodel = LeapModel(\n    =input_size,\n    =seq_len,\n    =hidden_size,\n    =output_size,\n    =num_layers,\n    =hidden_layers,\n    =dropout,\n    =,\n    =nhead,\n    =num_transformer_layers\n).(device)\n</code></pre>\n<ul>\n<li>A BiLSTM with TCN model</li>\n</ul>\n<pre><code></code></pre>\n<ul>\n<li>A BiLSTM with Attention model</li>\n</ul>\n<pre><code>class LeapModel(nn.Module):\n    def __init__(self,\n                 input_size,\n                 seq_len,\n                 hidden_size,\n                 output_size,\n                 =1,\n                 =,\n                 =.3,\n                 hidden_layers=[128, 256]):\n\n        super().__init__()\n        self.input_size = input_size\n        self.seq_len = seq_len\n        self.hidden_size = hidden_size\n        self.num_layers = num_layers\n        self.=bidirectional\n        self.=output_size\n        self.rnn = nn.LSTM(=input_size,\n                           =hidden_size,\n                           =num_layers,\n                           =bidirectional,\n                           =,\n                           =0.1)\n\n        self.attention = nn.MultiheadAttention(=hidden_size*2  bidirectional  hidden_size,\n                                               =8,\n                                               =)\n\n         hidden_layers  len(hidden_layers):\n            first_layer  = nn.Linear(hidden_size  bidirectional  hidden_size, hidden_layers[0])\n            self.hidden_layers = nn.ModuleList(\n                [first_layer] + \\\n                [nn.Linear(hidden_layers[i], hidden_layers[i+1])  i  range(len(hidden_layers) - 1)]\n            )\n             layer  self.hidden_layers:\n                nn.init.kaiming_normal_(layer.weight.data)\n            self.intermediate_layer = nn.Linear(hidden_layers[-1], self.input_size)\n            self.output_layer = nn.Linear(hidden_layers[-1], output_size)\n            nn.init.kaiming_normal_(self.output_layer.weight.data)\n        :\n            self.hidden_layers = []\n            self.intermediate_layer = nn.Linear(hidden_size  bidirectional  hidden_siz, self.input_size)\n            self.output_layer = nn.Linear(hidden_size  bidirectional  hidden_size, output_size)\n            nn.init.kaiming_normal_(self.output_layer.weight.data)\n\n        self.activation_fn = torch.nn.GELU()\n        self.dropout = nn.Dropout(dropout)\n\n    def forward(self, x):\n        batch_size = x.size(0)\n        outputs, hidden = self.rnn(x)\n\n        outputs = outputs.permute(1, 0, 2)  # (seq_len, batch_size, hidden_size)\n        attn_output, _ = self.attention(outputs, outputs, outputs)\n        attn_output = attn_output.permute(1, 0, 2)  # (batch_size, seq_len, hidden_size)\n\n        x = self.dropout(self.activation_fn(attn_output))\n         hidden_layer  self.hidden_layers:\n            x = self.activation_fn(hidden_layer(x))\n            x = self.dropout(x)\n        x = self.output_layer(x)\n\n        # (-1,60,14) -&gt; (-1,386)\n        o_s = x[:, :, :6]\n        o_s = o_s.permute(0,2,1).reshape(-1,360)\n        o_g = x[:, :, 6:]\n        o_g = o_g.mean(=1)\n        out = torch.cat([o_s, o_g], =1)\n\n        return out\n\n\ninput_size = 25\noutput_size = 14\nseq_len = 60\n\nhidden_size = 256\nhidden_layers = [256, 512]\nnum_layers = 6\ndropout = 0.1\n\nmodel = LeapModel(\n    =input_size,\n    =seq_len,\n    =hidden_size,\n    =output_size,\n    =num_layers,\n    =hidden_layers,\n    =dropout,\n    =,\n).(device)\n</code></pre>\n<h3>Dataset</h3>\n<ol>\n<li>Download all 0001-02 to 0009-01  (low resolution) data from Huggingface.</li>\n<li>0001-02 to 0008-06 as training set, 0008-07 to 0009-01 as validate set (sampling to ~625000 rows).</li>\n</ol>\n<h3>Some training details</h3>\n<ol>\n<li>Loss: nn.SmoothL1Loss(reduction='mean') (0.005~0.008 better than mse)</li>\n<li>Scheduler: get_cosine_schedule_with_warmup</li>\n<li>Activation Function: GELU (0.002~0.004 better than relu)</li>\n<li>Trained on 4*RTX4090 with 360G RAM, 7.5 years training dataset, ~1 hour per epoch</li>\n</ol>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1197382%2Ff9d38b78021aa6aaa63bd77da4bea9c3%2Fscreencapture-a18044-8b4e-f62d61d2-westb-seetacloud-8443-monitor-2024-07-10-10_05_39.png?generation=1722405976252082&amp;alt=media\" alt=\"\"></p>\n<h3>Post-Processing</h3>\n<pre><code>targets_unpredictable = []\nfor target in weights:\n    if weights[target] == :\n        targets_unpredictable.append(target)\nfor target in targets_unpredictable:\n    df_pred[target] = \nfor target in [f for i in range(, )]:\n    df_pred[target] = -df_test[target.replace(, )] * weights[target] / \n</code></pre>\n<p>Reference Links: <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/502484\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/502484</a></p>\n<h3>Ensemble</h3>\n<ol>\n<li>on models: w0 * pred0 + w1 * pred1 + … + w5 * pred5</li>\n<li>on targets:</li>\n</ol>\n<pre><code>selects = \n idx_t, target  ((TARGETCOLS), total=(TARGETCOLS)):\n    di = {}\n     idx_p, prob  (probs):\n        di = (df_valid, probs)\n    selects((di, key=di, reverse=True))\n</code></pre>\n<p>`</p>\n<h3>What didn't work</h3>\n<ol>\n<li>All 8 years dataset with fix epochs, without validation.</li>\n<li>Data augment: mask 10% input and TTA.</li>\n</ol>",
  "messages": [
    {
      "id": 2941488,
      "postDate": "2024-07-31T03:05:03.110Z",
      "content": "<p>First of all, I would like to thank the organizers and Kaggle for hosting this competition. The quality of the competition data is great. Although there were some problems during the process, in any case, we have achieved results that satisfy most people.</p>\n<p>This is my first solo gold medal and I became the Competition GrandMaster. This seven-year journey has been quite long and exciting.</p>\n<h3>solution summary</h3>\n<p>I think my solution is very simple, and it is basically based on the seq2seq models derived from BiLSTM.</p>\n<p>(bs, 60, 25)  --&gt; seq2seq --&gt; (bs, 60, 14) --&gt; (bs, 368) </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1197382%2Fa02dd80af723ab676afa94c4f8f58b63%2FScreenshot%202024-07-31_10-19-57-734.png?generation=1722392449416839&amp;alt=media\" alt=\"\"></p>\n<h3>models</h3>\n<ul>\n<li>validate on last 6 months sample data</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>models</th>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>BiLSTM(layers=6)</td>\n<td>0.7844</td>\n<td>0.7812</td>\n</tr>\n<tr>\n<td>BiGRU(layers=8)</td>\n<td>0.7835</td>\n<td>0.7802</td>\n</tr>\n<tr>\n<td>BiLSTM+Transformer</td>\n<td>0.7858</td>\n<td>0.7821</td>\n</tr>\n<tr>\n<td>BiLSTM+Attention</td>\n<td>0.7865</td>\n<td>0.7834</td>\n</tr>\n<tr>\n<td>BiLSTM+TCN</td>\n<td>0.7855</td>\n<td>0.7832</td>\n</tr>\n<tr>\n<td>BiLSTM+CNN</td>\n<td>0.7842</td>\n<td>0.7821</td>\n</tr>\n<tr>\n<td>ensemble on models</td>\n<td>0.7923</td>\n<td>0.7890</td>\n</tr>\n<tr>\n<td>ensemble on targets</td>\n<td>0.7933</td>\n<td>0.7884</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>A base BiLSTM model:</li>\n</ul>\n<pre><code>class LeapModel(nn.Module):\n    def __init__(self,\n                 input_size,\n                 seq_len,\n                 hidden_size,\n                 output_size,\n                 =1,\n                 =,\n                 =.3,\n                 hidden_layers=[128, 256]):\n\n        super().__init__()\n        self.input_size = input_size\n        self.seq_len = seq_len\n        self.hidden_size = hidden_size\n        self.num_layers = num_layers\n        self.=bidirectional\n        self.=output_size\n\n        self.rnn = nn.LSTM(=input_size,\n                           =hidden_size,\n                           =num_layers,\n                           =bidirectional,\n                           =,\n                           =dropout)\n\n         hidden_layers  len(hidden_layers):\n            first_layer  = nn.Linear(hidden_size  bidirectional  hidden_size, hidden_layers[0])\n            self.hidden_layers = nn.ModuleList(\n                [first_layer] + \\\n                [nn.Linear(hidden_layers[i], hidden_layers[i+1])  i  range(len(hidden_layers) - 1)]\n            )\n             layer  self.hidden_layers:\n                nn.init.kaiming_normal_(layer.weight.data)\n            self.intermediate_layer = nn.Linear(hidden_layers[-1], self.input_size)\n            self.output_layer = nn.Linear(hidden_layers[-1], output_size)\n            nn.init.kaiming_normal_(self.output_layer.weight.data)\n        :\n            self.hidden_layers = []\n            self.intermediate_layer = nn.Linear(hidden_size  bidirectional  hidden_size, self.input_size)\n            self.output_layer = nn.Linear(hidden_size  bidirectional  hidden_size, output_size)\n            nn.init.kaiming_normal_(self.output_layer.weight.data)\n\n        self.activation_fn = torch.nn.GELU()\n        self.dropout = nn.Dropout(dropout)\n\n    def forward(self, x):\n        outputs, hidden = self.rnn(x)\n\n        x = self.dropout(self.activation_fn(outputs))\n         hidden_layer  self.hidden_layers:\n            x = self.activation_fn(hidden_layer(x))\n            x = self.dropout(x)\n        x = self.output_layer(x)\n\n        # (-1,60,14) -&gt; (-1,386)\n        o_s = x[:, :, :6]\n        o_s = o_s.permute(0,2,1).reshape(-1,360)\n        o_g = x[:, :, 6:]\n        o_g = o_g.mean(=1)\n        out = torch.cat([o_s, o_g], =1)\n\n        return out\n\ninput_size = 25\noutput_size = 14\nseq_len = 60\n\nhidden_size = 256\nhidden_layers = [256, 512]\nnum_layers = 6\ndropout = 0.1\n\nmodel = LeapModel(\n    =input_size,\n    =seq_len,\n    =hidden_size,\n    =output_size,\n    =num_layers,\n    =hidden_layers,\n    =dropout,\n    =,\n).(device)\n</code></pre>\n<p>Reference Links: <a href=\"https://www.kaggle.com/code/brandenkmurray/seq2seq-rnn-with-gru\" target=\"_blank\">https://www.kaggle.com/code/brandenkmurray/seq2seq-rnn-with-gru</a></p>\n<ul>\n<li>A BiLSTM with Transformer model</li>\n</ul>\n<pre><code>class LeapModel(nn.Module):\n    def __init__(self,\n                 input_size,\n                 seq_len,\n                 hidden_size,\n                 output_size,\n                 =1,\n                 =,\n                 =0.3,\n                 hidden_layers=[128, 256],\n                 =8,\n                 =2):\n\n        super().__init__()\n        self.input_size = input_size\n        self.seq_len = seq_len\n        self.hidden_size = hidden_size\n        self.num_layers = num_layers\n        self.bidirectional = bidirectional\n        self.output_size = output_size\n\n        # LSTM layer\n        self.rnn = nn.LSTM(=input_size,\n                           =hidden_size,\n                           =num_layers,\n                           =bidirectional,\n                           =,\n                           =dropout)\n\n        # Transformer layer\n        transformer_input_size = hidden_size * 2  bidirectional  hidden_size\n        self.transformer_layer = nn.TransformerEncoder(\n            nn.TransformerEncoderLayer(=transformer_input_size, =nhead, =dropout),\n            =num_transformer_layers\n        )\n\n        # Fully connected layers\n         hidden_layers  len(hidden_layers):\n            first_layer = nn.Linear(transformer_input_size, hidden_layers[0])\n            self.hidden_layers = nn.ModuleList(\n                [first_layer] + \\\n                [nn.Linear(hidden_layers[i], hidden_layers[i+1])  i  range(len(hidden_layers) - 1)]\n            )\n             layer  self.hidden_layers:\n                nn.init.kaiming_normal_(layer.weight.data)\n            self.intermediate_layer = nn.Linear(hidden_layers[-1], self.input_size)\n            self.output_layer = nn.Linear(hidden_layers[-1], output_size)\n            nn.init.kaiming_normal_(self.output_layer.weight.data)\n        :\n            self.hidden_layers = []\n            self.intermediate_layer = nn.Linear(transformer_input_size, self.input_size)\n            self.output_layer = nn.Linear(transformer_input_size, output_size)\n            nn.init.kaiming_normal_(self.output_layer.weight.data)\n\n        self.activation_fn = torch.nn.GELU()\n        self.dropout = nn.Dropout(dropout)\n\n    def forward(self, x):\n        # LSTM layer\n        lstm_output, _ = self.rnn(x)\n\n        # Transformer layer\n        transformer_output = self.transformer_layer(lstm_output)\n\n        # Apply dropout  activation\n        x = self.dropout(self.activation_fn(transformer_output))\n\n        # Fully connected layers\n         hidden_layer  self.hidden_layers:\n            x = self.activation_fn(hidden_layer(x))\n            x = self.dropout(x)\n        x = self.output_layer(x)\n\n        # Reshape output\n        o_s = x[:, :, :6]\n        o_s = o_s.permute(0, 2, 1).reshape(-1, 360)\n        o_g = x[:, :, 6:]\n        o_g = o_g.mean(=1)\n        out = torch.cat([o_s, o_g], =1)\n\n        return out\n\ninput_size = 25\noutput_size = 14\nseq_len = 60\n\nhidden_size = 256\nhidden_layers = [256, 512]\nnum_layers = 6\ndropout = 0.1\nnhead = 8\nnum_transformer_layers = 1\n\nmodel = LeapModel(\n    =input_size,\n    =seq_len,\n    =hidden_size,\n    =output_size,\n    =num_layers,\n    =hidden_layers,\n    =dropout,\n    =,\n    =nhead,\n    =num_transformer_layers\n).(device)\n</code></pre>\n<ul>\n<li>A BiLSTM with TCN model</li>\n</ul>\n<pre><code></code></pre>\n<ul>\n<li>A BiLSTM with Attention model</li>\n</ul>\n<pre><code>class LeapModel(nn.Module):\n    def __init__(self,\n                 input_size,\n                 seq_len,\n                 hidden_size,\n                 output_size,\n                 =1,\n                 =,\n                 =.3,\n                 hidden_layers=[128, 256]):\n\n        super().__init__()\n        self.input_size = input_size\n        self.seq_len = seq_len\n        self.hidden_size = hidden_size\n        self.num_layers = num_layers\n        self.=bidirectional\n        self.=output_size\n        self.rnn = nn.LSTM(=input_size,\n                           =hidden_size,\n                           =num_layers,\n                           =bidirectional,\n                           =,\n                           =0.1)\n\n        self.attention = nn.MultiheadAttention(=hidden_size*2  bidirectional  hidden_size,\n                                               =8,\n                                               =)\n\n         hidden_layers  len(hidden_layers):\n            first_layer  = nn.Linear(hidden_size  bidirectional  hidden_size, hidden_layers[0])\n            self.hidden_layers = nn.ModuleList(\n                [first_layer] + \\\n                [nn.Linear(hidden_layers[i], hidden_layers[i+1])  i  range(len(hidden_layers) - 1)]\n            )\n             layer  self.hidden_layers:\n                nn.init.kaiming_normal_(layer.weight.data)\n            self.intermediate_layer = nn.Linear(hidden_layers[-1], self.input_size)\n            self.output_layer = nn.Linear(hidden_layers[-1], output_size)\n            nn.init.kaiming_normal_(self.output_layer.weight.data)\n        :\n            self.hidden_layers = []\n            self.intermediate_layer = nn.Linear(hidden_size  bidirectional  hidden_siz, self.input_size)\n            self.output_layer = nn.Linear(hidden_size  bidirectional  hidden_size, output_size)\n            nn.init.kaiming_normal_(self.output_layer.weight.data)\n\n        self.activation_fn = torch.nn.GELU()\n        self.dropout = nn.Dropout(dropout)\n\n    def forward(self, x):\n        batch_size = x.size(0)\n        outputs, hidden = self.rnn(x)\n\n        outputs = outputs.permute(1, 0, 2)  # (seq_len, batch_size, hidden_size)\n        attn_output, _ = self.attention(outputs, outputs, outputs)\n        attn_output = attn_output.permute(1, 0, 2)  # (batch_size, seq_len, hidden_size)\n\n        x = self.dropout(self.activation_fn(attn_output))\n         hidden_layer  self.hidden_layers:\n            x = self.activation_fn(hidden_layer(x))\n            x = self.dropout(x)\n        x = self.output_layer(x)\n\n        # (-1,60,14) -&gt; (-1,386)\n        o_s = x[:, :, :6]\n        o_s = o_s.permute(0,2,1).reshape(-1,360)\n        o_g = x[:, :, 6:]\n        o_g = o_g.mean(=1)\n        out = torch.cat([o_s, o_g], =1)\n\n        return out\n\n\ninput_size = 25\noutput_size = 14\nseq_len = 60\n\nhidden_size = 256\nhidden_layers = [256, 512]\nnum_layers = 6\ndropout = 0.1\n\nmodel = LeapModel(\n    =input_size,\n    =seq_len,\n    =hidden_size,\n    =output_size,\n    =num_layers,\n    =hidden_layers,\n    =dropout,\n    =,\n).(device)\n</code></pre>\n<h3>Dataset</h3>\n<ol>\n<li>Download all 0001-02 to 0009-01  (low resolution) data from Huggingface.</li>\n<li>0001-02 to 0008-06 as training set, 0008-07 to 0009-01 as validate set (sampling to ~625000 rows).</li>\n</ol>\n<h3>Some training details</h3>\n<ol>\n<li>Loss: nn.SmoothL1Loss(reduction='mean') (0.005~0.008 better than mse)</li>\n<li>Scheduler: get_cosine_schedule_with_warmup</li>\n<li>Activation Function: GELU (0.002~0.004 better than relu)</li>\n<li>Trained on 4*RTX4090 with 360G RAM, 7.5 years training dataset, ~1 hour per epoch</li>\n</ol>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1197382%2Ff9d38b78021aa6aaa63bd77da4bea9c3%2Fscreencapture-a18044-8b4e-f62d61d2-westb-seetacloud-8443-monitor-2024-07-10-10_05_39.png?generation=1722405976252082&amp;alt=media\" alt=\"\"></p>\n<h3>Post-Processing</h3>\n<pre><code>targets_unpredictable = []\nfor target in weights:\n    if weights[target] == :\n        targets_unpredictable.append(target)\nfor target in targets_unpredictable:\n    df_pred[target] = \nfor target in [f for i in range(, )]:\n    df_pred[target] = -df_test[target.replace(, )] * weights[target] / \n</code></pre>\n<p>Reference Links: <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/502484\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/502484</a></p>\n<h3>Ensemble</h3>\n<ol>\n<li>on models: w0 * pred0 + w1 * pred1 + … + w5 * pred5</li>\n<li>on targets:</li>\n</ol>\n<pre><code>selects = \n idx_t, target  ((TARGETCOLS), total=(TARGETCOLS)):\n    di = {}\n     idx_p, prob  (probs):\n        di = (df_valid, probs)\n    selects((di, key=di, reverse=True))\n</code></pre>\n<p>`</p>\n<h3>What didn't work</h3>\n<ol>\n<li>All 8 years dataset with fix epochs, without validation.</li>\n<li>Data augment: mask 10% input and TTA.</li>\n</ol>",
      "rawMarkdown": "First of all, I would like to thank the organizers and Kaggle for hosting this competition. The quality of the competition data is great. Although there were some problems during the process, in any case, we have achieved results that satisfy most people.\n\nThis is my first solo gold medal and I became the Competition GrandMaster. This seven-year journey has been quite long and exciting.\n\n### solution summary \n\nI think my solution is very simple, and it is basically based on the seq2seq models derived from BiLSTM.\n\n(bs, 60, 25)  --> seq2seq --> (bs, 60, 14) --> (bs, 368) \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1197382%2Fa02dd80af723ab676afa94c4f8f58b63%2FScreenshot%202024-07-31_10-19-57-734.png?generation=1722392449416839&alt=media)\n\n### models\n\n* validate on last 6 months sample data\n\n| models | CV |  LB |\n| --- | --- | --- |\n| BiLSTM(layers=6) | 0.7844 | 0.7812 |\n| BiGRU(layers=8) | 0.7835 | 0.7802 |\n| BiLSTM+Transformer | 0.7858 | 0.7821 |\n| BiLSTM+Attention | 0.7865 | 0.7834 |\n| BiLSTM+TCN | 0.7855 | 0.7832 |\n| BiLSTM+CNN | 0.7842 | 0.7821 |\n| ensemble on models | 0.7923 | 0.7890 |\n| ensemble on targets | 0.7933 | 0.7884 |\n\n\n- A base BiLSTM model:\n\n```\nclass LeapModel(nn.Module):\n    def __init__(self,\n                 input_size,\n                 seq_len,\n                 hidden_size,\n                 output_size,\n                 num_layers=1,\n                 bidirectional=False,\n                 dropout=.3,\n                 hidden_layers=[128, 256]):\n\n        super().__init__()\n        self.input_size = input_size\n        self.seq_len = seq_len\n        self.hidden_size = hidden_size\n        self.num_layers = num_layers\n        self.bidirectional=bidirectional\n        self.output_size=output_size\n\n        self.rnn = nn.LSTM(input_size=input_size,\n                           hidden_size=hidden_size,\n                           num_layers=num_layers,\n                           bidirectional=bidirectional,\n                           batch_first=True,\n                           dropout=dropout)\n\n        if hidden_layers and len(hidden_layers):\n            first_layer  = nn.Linear(hidden_size*2 if bidirectional else hidden_size, hidden_layers[0])\n            self.hidden_layers = nn.ModuleList(\n                [first_layer] + \\\n                [nn.Linear(hidden_layers[i], hidden_layers[i+1]) for i in range(len(hidden_layers) - 1)]\n            )\n            for layer in self.hidden_layers:\n                nn.init.kaiming_normal_(layer.weight.data)\n            self.intermediate_layer = nn.Linear(hidden_layers[-1], self.input_size)\n            self.output_layer = nn.Linear(hidden_layers[-1], output_size)\n            nn.init.kaiming_normal_(self.output_layer.weight.data)\n        else:\n            self.hidden_layers = []\n            self.intermediate_layer = nn.Linear(hidden_size*2 if bidirectional else hidden_size, self.input_size)\n            self.output_layer = nn.Linear(hidden_size*2 if bidirectional else hidden_size, output_size)\n            nn.init.kaiming_normal_(self.output_layer.weight.data)\n\n        self.activation_fn = torch.nn.GELU()\n        self.dropout = nn.Dropout(dropout)\n\n    def forward(self, x):\n        outputs, hidden = self.rnn(x)\n\n        x = self.dropout(self.activation_fn(outputs))\n        for hidden_layer in self.hidden_layers:\n            x = self.activation_fn(hidden_layer(x))\n            x = self.dropout(x)\n        x = self.output_layer(x)\n\n        # (-1,60,14) -> (-1,386)\n        o_s = x[:, :, :6]\n        o_s = o_s.permute(0,2,1).reshape(-1,360)\n        o_g = x[:, :, 6:]\n        o_g = o_g.mean(dim=1)\n        out = torch.cat([o_s, o_g], dim=1)\n\n        return out\n\ninput_size = 25\noutput_size = 14\nseq_len = 60\n\nhidden_size = 256\nhidden_layers = [256, 512]\nnum_layers = 6\ndropout = 0.1\n\nmodel = LeapModel(\n    input_size=input_size,\n    seq_len=seq_len,\n    hidden_size=hidden_size,\n    output_size=output_size,\n    num_layers=num_layers,\n    hidden_layers=hidden_layers,\n    dropout=dropout,\n    bidirectional=True,\n).to(device)\n```\n\nReference Links: https://www.kaggle.com/code/brandenkmurray/seq2seq-rnn-with-gru\n\n- A BiLSTM with Transformer model\n\n```\nclass LeapModel(nn.Module):\n    def __init__(self,\n                 input_size,\n                 seq_len,\n                 hidden_size,\n                 output_size,\n                 num_layers=1,\n                 bidirectional=False,\n                 dropout=0.3,\n                 hidden_layers=[128, 256],\n                 nhead=8,\n                 num_transformer_layers=2):\n\n        super().__init__()\n        self.input_size = input_size\n        self.seq_len = seq_len\n        self.hidden_size = hidden_size\n        self.num_layers = num_layers\n        self.bidirectional = bidirectional\n        self.output_size = output_size\n\n        # LSTM layer\n        self.rnn = nn.LSTM(input_size=input_size,\n                           hidden_size=hidden_size,\n                           num_layers=num_layers,\n                           bidirectional=bidirectional,\n                           batch_first=True,\n                           dropout=dropout)\n\n        # Transformer layer\n        transformer_input_size = hidden_size * 2 if bidirectional else hidden_size\n        self.transformer_layer = nn.TransformerEncoder(\n            nn.TransformerEncoderLayer(d_model=transformer_input_size, nhead=nhead, dropout=dropout),\n            num_layers=num_transformer_layers\n        )\n\n        # Fully connected layers\n        if hidden_layers and len(hidden_layers):\n            first_layer = nn.Linear(transformer_input_size, hidden_layers[0])\n            self.hidden_layers = nn.ModuleList(\n                [first_layer] + \\\n                [nn.Linear(hidden_layers[i], hidden_layers[i+1]) for i in range(len(hidden_layers) - 1)]\n            )\n            for layer in self.hidden_layers:\n                nn.init.kaiming_normal_(layer.weight.data)\n            self.intermediate_layer = nn.Linear(hidden_layers[-1], self.input_size)\n            self.output_layer = nn.Linear(hidden_layers[-1], output_size)\n            nn.init.kaiming_normal_(self.output_layer.weight.data)\n        else:\n            self.hidden_layers = []\n            self.intermediate_layer = nn.Linear(transformer_input_size, self.input_size)\n            self.output_layer = nn.Linear(transformer_input_size, output_size)\n            nn.init.kaiming_normal_(self.output_layer.weight.data)\n\n        self.activation_fn = torch.nn.GELU()\n        self.dropout = nn.Dropout(dropout)\n\n    def forward(self, x):\n        # LSTM layer\n        lstm_output, _ = self.rnn(x)\n        \n        # Transformer layer\n        transformer_output = self.transformer_layer(lstm_output)\n        \n        # Apply dropout and activation\n        x = self.dropout(self.activation_fn(transformer_output))\n        \n        # Fully connected layers\n        for hidden_layer in self.hidden_layers:\n            x = self.activation_fn(hidden_layer(x))\n            x = self.dropout(x)\n        x = self.output_layer(x)\n\n        # Reshape output\n        o_s = x[:, :, :6]\n        o_s = o_s.permute(0, 2, 1).reshape(-1, 360)\n        o_g = x[:, :, 6:]\n        o_g = o_g.mean(dim=1)\n        out = torch.cat([o_s, o_g], dim=1)\n\n        return out\n\ninput_size = 25\noutput_size = 14\nseq_len = 60\n\nhidden_size = 256\nhidden_layers = [256, 512]\nnum_layers = 6\ndropout = 0.1\nnhead = 8\nnum_transformer_layers = 1\n\nmodel = LeapModel(\n    input_size=input_size,\n    seq_len=seq_len,\n    hidden_size=hidden_size,\n    output_size=output_size,\n    num_layers=num_layers,\n    hidden_layers=hidden_layers,\n    dropout=dropout,\n    bidirectional=True,\n    nhead=nhead,\n    num_transformer_layers=num_transformer_layers\n).to(device)\n```\n\n- A BiLSTM with TCN model\n\n```\nclass TCNBlock(nn.Module):\n    def __init__(self, in_channels, out_channels, kernel_size, dilation):\n        super(TCNBlock, self).__init__()\n        self.conv = nn.Conv1d(in_channels, out_channels, kernel_size, \n                              padding=(kernel_size-1) * dilation // 2, dilation=dilation)\n        self.bn = nn.BatchNorm1d(out_channels)\n        self.activation_fn = nn.GELU()\n\n    def forward(self, x):\n        return self.activation_fn(self.bn(self.conv(x)))\n\n\nclass LeapModel(nn.Module):\n    def __init__(self,\n                 input_size,\n                 seq_len,\n                 hidden_size,\n                 output_size,\n                 num_layers=1,\n                 bidirectional=False,\n                 dropout=0.3):\n\n        super().__init__()\n        self.input_size = input_size\n        self.seq_len = seq_len\n        self.hidden_size = hidden_size\n        self.num_layers = num_layers\n        self.bidirectional = bidirectional\n        self.output_size = output_size\n\n        # LSTM layer\n        self.rnn = nn.LSTM(input_size=input_size,\n                           hidden_size=hidden_size,\n                           num_layers=num_layers,\n                           bidirectional=bidirectional,\n                           batch_first=True,\n                           dropout=dropout)\n\n        self.se = nn.Sequential(\n            nn.Linear(hidden_size*2, hidden_size//2),\n            nn.GELU(),\n            nn.Linear(hidden_size//2, hidden_size*2),\n            nn.Sigmoid()\n        )\n        \n        self.tcn = nn.Sequential(\n            TCNBlock(hidden_size*2, hidden_size*2, kernel_size=3, dilation=1),\n            TCNBlock(hidden_size*2, hidden_size*2, kernel_size=3, dilation=2),\n            TCNBlock(hidden_size*2, hidden_size*2, kernel_size=3, dilation=4),\n            TCNBlock(hidden_size*2, hidden_size*2, kernel_size=3, dilation=8),\n        )\n        \n        self.fc = nn.Linear(hidden_size*2, output_size)\n        self.dropout = nn.Dropout(dropout)\n        \n\n    def forward(self, x):\n        # RNN layer\n        outputs, _ = self.rnn(x)\n        \n        se_weights = self.se(torch.mean(outputs, dim=1)).unsqueeze(1)\n        outputs = outputs * se_weights\n        \n        tcn_input = outputs.permute(0, 2, 1)\n        tcn_output = self.tcn(tcn_input)\n        tcn_output = tcn_output.permute(0, 2, 1)\n        \n        x = self.dropout(tcn_output)\n        x = self.fc(x)\n\n        # Reshape output\n        o_s = x[:, :, :6]\n        o_s = o_s.permute(0, 2, 1).reshape(-1, 360)\n        o_g = x[:, :, 6:]\n        o_g = o_g.mean(dim=1)\n        out = torch.cat([o_s, o_g], dim=1)  # (bs,368)\n\n        return out\n\n\ninput_size = 25\noutput_size = 14\nseq_len = 60\n\nhidden_size = 256\nnum_layers = 6\ndropout = 0.1\n\nmodel = LeapModel(\n    input_size=input_size,\n    seq_len=seq_len,\n    hidden_size=hidden_size,\n    output_size=output_size,\n    num_layers=num_layers,\n    dropout=dropout,\n    bidirectional=True,\n).to(device)\n```\n\n- A BiLSTM with Attention model\n\n```\nclass LeapModel(nn.Module):\n    def __init__(self,\n                 input_size,\n                 seq_len,\n                 hidden_size,\n                 output_size,\n                 num_layers=1,\n                 bidirectional=False,\n                 dropout=.3,\n                 hidden_layers=[128, 256]):\n\n        super().__init__()\n        self.input_size = input_size\n        self.seq_len = seq_len\n        self.hidden_size = hidden_size\n        self.num_layers = num_layers\n        self.bidirectional=bidirectional\n        self.output_size=output_size\n        self.rnn = nn.LSTM(input_size=input_size,\n                           hidden_size=hidden_size,\n                           num_layers=num_layers,\n                           bidirectional=bidirectional,\n                           batch_first=True,\n                           dropout=0.1)\n\n        self.attention = nn.MultiheadAttention(embed_dim=hidden_size*2 if bidirectional else hidden_size,\n                                               num_heads=8,\n                                               batch_first=True)\n\n        if hidden_layers and len(hidden_layers):\n            first_layer  = nn.Linear(hidden_size*2 if bidirectional else hidden_size, hidden_layers[0])\n            self.hidden_layers = nn.ModuleList(\n                [first_layer] + \\\n                [nn.Linear(hidden_layers[i], hidden_layers[i+1]) for i in range(len(hidden_layers) - 1)]\n            )\n            for layer in self.hidden_layers:\n                nn.init.kaiming_normal_(layer.weight.data)\n            self.intermediate_layer = nn.Linear(hidden_layers[-1], self.input_size)\n            self.output_layer = nn.Linear(hidden_layers[-1], output_size)\n            nn.init.kaiming_normal_(self.output_layer.weight.data)\n        else:\n            self.hidden_layers = []\n            self.intermediate_layer = nn.Linear(hidden_size*2 if bidirectional else hidden_siz, self.input_size)\n            self.output_layer = nn.Linear(hidden_size*2 if bidirectional else hidden_size, output_size)\n            nn.init.kaiming_normal_(self.output_layer.weight.data)\n\n        self.activation_fn = torch.nn.GELU()\n        self.dropout = nn.Dropout(dropout)\n\n    def forward(self, x):\n        batch_size = x.size(0)\n        outputs, hidden = self.rnn(x)\n\n        outputs = outputs.permute(1, 0, 2)  # (seq_len, batch_size, hidden_size)\n        attn_output, _ = self.attention(outputs, outputs, outputs)\n        attn_output = attn_output.permute(1, 0, 2)  # (batch_size, seq_len, hidden_size)\n\n        x = self.dropout(self.activation_fn(attn_output))\n        for hidden_layer in self.hidden_layers:\n            x = self.activation_fn(hidden_layer(x))\n            x = self.dropout(x)\n        x = self.output_layer(x)\n\n        # (-1,60,14) -> (-1,386)\n        o_s = x[:, :, :6]\n        o_s = o_s.permute(0,2,1).reshape(-1,360)\n        o_g = x[:, :, 6:]\n        o_g = o_g.mean(dim=1)\n        out = torch.cat([o_s, o_g], dim=1)\n\n        return out\n\n\ninput_size = 25\noutput_size = 14\nseq_len = 60\n\nhidden_size = 256\nhidden_layers = [256, 512]\nnum_layers = 6\ndropout = 0.1\n\nmodel = LeapModel(\n    input_size=input_size,\n    seq_len=seq_len,\n    hidden_size=hidden_size,\n    output_size=output_size,\n    num_layers=num_layers,\n    hidden_layers=hidden_layers,\n    dropout=dropout,\n    bidirectional=True,\n).to(device)\n```\n\n### Dataset\n\n1. Download all 0001-02 to 0009-01  (low resolution) data from Huggingface.\n2. 0001-02 to 0008-06 as training set, 0008-07 to 0009-01 as validate set (sampling to ~625000 rows).\n\n### Some training details\n\n1. Loss: nn.SmoothL1Loss(reduction='mean') (0.005~0.008 better than mse)\n2. Scheduler: get_cosine_schedule_with_warmup\n3. Activation Function: GELU (0.002~0.004 better than relu)\n4. Trained on 4*RTX4090 with 360G RAM, 7.5 years training dataset, ~1 hour per epoch\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1197382%2Ff9d38b78021aa6aaa63bd77da4bea9c3%2Fscreencapture-a18044-8b4e-f62d61d2-westb-seetacloud-8443-monitor-2024-07-10-10_05_39.png?generation=1722405976252082&alt=media)\n\n### Post-Processing\n\n```\ntargets_unpredictable = []\nfor target in weights:\n    if weights[target] == 0.:\n        targets_unpredictable.append(target)\nfor target in targets_unpredictable:\n    df_pred[target] = 0.\nfor target in [f'ptend_q0002_{i}' for i in range(12, 28)]:\n    df_pred[target] = -df_test[target.replace(\"ptend\", \"state\")] * weights[target] / 1200.\n```\n\nReference Links: https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/502484\n\n### Ensemble\n\n1. on models: w0 * pred0 + w1 * pred1 + ... + w5 * pred5\n2. on targets:\n\n```\nselects = []\nfor idx_t, target in tqdm(enumerate(TARGETCOLS), total=len(TARGETCOLS)):\n    di = {}\n    for idx_p, prob in enumerate(probs):\n        di[idx_p] = r2_score(df_valid[target], probs[idx_p][:, idx_t])\n    selects.append(sorted(di, key=di.get, reverse=True)[:4])\n````\n\n### What didn't work\n\n1. All 8 years dataset with fix epochs, without validation.\n2. Data augment: mask 10% input and TTA.\n",
      "votes": 25
    },
    {
      "id": 2941893,
      "postDate": "2024-07-31T12:17:47.613Z",
      "content": "<p>Congrats on solo gold and GM!</p>",
      "rawMarkdown": "Congrats on solo gold and GM!",
      "votes": 1,
      "replies": [
        {
          "id": 2942698,
          "postDate": "2024-08-01T02:44:38.760Z",
          "content": "<p>Congratulations on winning the championship, I learned a lot from your solution!</p>",
          "rawMarkdown": "Congratulations on winning the championship, I learned a lot from your solution!"
        }
      ]
    },
    {
      "id": 2941501,
      "postDate": "2024-07-31T03:23:44.230Z",
      "content": "<p>Congrats on achieving GM!</p>",
      "rawMarkdown": "Congrats on achieving GM!",
      "votes": 2
    },
    {
      "id": 2953728,
      "postDate": "2024-08-09T03:09:26.493Z",
      "content": "<p>Congratulations on becoming a GM. I wanted to ask if you've encountered any memory issues with dataloaders in your use of PyTorch?</p>\n<p>Specifically, I'd like to know how you train models on multiple GPUs? Do you use Data Parallel (DP), Distributed Data Parallel (DDP), or something else? For me, when I use DDP on multiple GPUs, the pseudocode below results in an inability to properly release the memory occupied by the dataloaders:</p>\n<pre><code>for epoch in (CFG.num_epochs):\n    for j in ():  #  months\n        train_loader = (CFG, j=j)\n        train_loss, grad_norm = (\n            model, criterion, optimizer, scheduler, train_loader, gpu_id, CFG\n        )\n        del train_loader  # failed to release memory\n        ()\n    ()\n</code></pre>",
      "rawMarkdown": "Congratulations on becoming a GM. I wanted to ask if you've encountered any memory issues with dataloaders in your use of PyTorch?\n\nSpecifically, I'd like to know how you train models on multiple GPUs? Do you use Data Parallel (DP), Distributed Data Parallel (DDP), or something else? For me, when I use DDP on multiple GPUs, the pseudocode below results in an inability to properly release the memory occupied by the dataloaders:\n\n```\nfor epoch in range(CFG.num_epochs):\n    for j in range(12):  # 12 months\n        train_loader = get_dataloaders(CFG, j=j)\n        train_loss, grad_norm = train(\n            model, criterion, optimizer, scheduler, train_loader, gpu_id, CFG\n        )\n        del train_loader  # failed to release memory\n        clean_memory()\n    valid()\n```",
      "replies": [
        {
          "id": 2953733,
          "postDate": "2024-08-09T03:27:28.533Z",
          "content": "<p>Congratulations on your solo gold.  </p>\n<p>I use DP, which is a very simple solution, although not as efficient as DDP. Also, I heard that DDP does have memory issues, so I chose DP and a large memory machine to circumvent this problem. Maybe I will study DDP next time. 😀</p>",
          "rawMarkdown": "Congratulations on your solo gold.  \n\nI use DP, which is a very simple solution, although not as efficient as DDP. Also, I heard that DDP does have memory issues, so I chose DP and a large memory machine to circumvent this problem. Maybe I will study DDP next time. 😀",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2941893,
      "author_name": "greySnow",
      "author_url": "",
      "post_date": "2024-07-31T12:17:47.613000",
      "content": "<p>Congrats on solo gold and GM!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2942698,
          "author_name": "heng",
          "author_url": "",
          "post_date": "2024-08-01T02:44:38.760000",
          "content": "<p>Congratulations on winning the championship, I learned a lot from your solution!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2941501,
      "author_name": "ADAM.",
      "author_url": "",
      "post_date": "2024-07-31T03:23:44.230000",
      "content": "<p>Congrats on achieving GM!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2953728,
      "author_name": "Uesugi Erii",
      "author_url": "",
      "post_date": "2024-08-09T03:09:26.493000",
      "content": "<p>Congratulations on becoming a GM. I wanted to ask if you've encountered any memory issues with dataloaders in your use of PyTorch?</p>\n<p>Specifically, I'd like to know how you train models on multiple GPUs? Do you use Data Parallel (DP), Distributed Data Parallel (DDP), or something else? For me, when I use DDP on multiple GPUs, the pseudocode below results in an inability to properly release the memory occupied by the dataloaders:</p>\n<pre><code>for epoch in (CFG.num_epochs):\n    for j in ():  #  months\n        train_loader = (CFG, j=j)\n        train_loss, grad_norm = (\n            model, criterion, optimizer, scheduler, train_loader, gpu_id, CFG\n        )\n        del train_loader  # failed to release memory\n        ()\n    ()\n</code></pre>",
      "votes": 0,
      "replies": [
        {
          "id": 2953733,
          "author_name": "heng",
          "author_url": "",
          "post_date": "2024-08-09T03:27:28.533000",
          "content": "<p>Congratulations on your solo gold.  </p>\n<p>I use DP, which is a very simple solution, although not as efficient as DDP. Also, I heard that DDP does have memory issues, so I chose DP and a large memory machine to circumvent this problem. Maybe I will study DDP next time. 😀</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2941488": "First of all, I would like to thank the organizers and Kaggle for hosting this competition. The quality of the competition data is great. Although there were some problems during the process, in any case, we have achieved results that satisfy most people.\n\nThis is my first solo gold medal and I became the Competition GrandMaster. This seven-year journey has been quite long and exciting.\n\n### solution summary \n\nI think my solution is very simple, and it is basically based on the seq2seq models derived from BiLSTM.\n\n(bs, 60, 25)  --> seq2seq --> (bs, 60, 14) --> (bs, 368) \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1197382%2Fa02dd80af723ab676afa94c4f8f58b63%2FScreenshot%202024-07-31_10-19-57-734.png?generation=1722392449416839&alt=media)\n\n### models\n\n* validate on last 6 months sample data\n\n| models | CV |  LB |\n| --- | --- | --- |\n| BiLSTM(layers=6) | 0.7844 | 0.7812 |\n| BiGRU(layers=8) | 0.7835 | 0.7802 |\n| BiLSTM+Transformer | 0.7858 | 0.7821 |\n| BiLSTM+Attention | 0.7865 | 0.7834 |\n| BiLSTM+TCN | 0.7855 | 0.7832 |\n| BiLSTM+CNN | 0.7842 | 0.7821 |\n| ensemble on models | 0.7923 | 0.7890 |\n| ensemble on targets | 0.7933 | 0.7884 |\n\n\n- A base BiLSTM model:\n\n```\nclass LeapModel(nn.Module):\n    def __init__(self,\n                 input_size,\n                 seq_len,\n                 hidden_size,\n                 output_size,\n                 num_layers=1,\n                 bidirectional=False,\n                 dropout=.3,\n                 hidden_layers=[128, 256]):\n\n        super().__init__()\n        self.input_size = input_size\n        self.seq_len = seq_len\n        self.hidden_size = hidden_size\n        self.num_layers = num_layers\n        self.bidirectional=bidirectional\n        self.output_size=output_size\n\n        self.rnn = nn.LSTM(input_size=input_size,\n                           hidden_size=hidden_size,\n                           num_layers=num_layers,\n                           bidirectional=bidirectional,\n                           batch_first=True,\n                           dropout=dropout)\n\n        if hidden_layers and len(hidden_layers):\n            first_layer  = nn.Linear(hidden_size*2 if bidirectional else hidden_size, hidden_layers[0])\n            self.hidden_layers = nn.ModuleList(\n                [first_layer] + \\\n                [nn.Linear(hidden_layers[i], hidden_layers[i+1]) for i in range(len(hidden_layers) - 1)]\n            )\n            for layer in self.hidden_layers:\n                nn.init.kaiming_normal_(layer.weight.data)\n            self.intermediate_layer = nn.Linear(hidden_layers[-1], self.input_size)\n            self.output_layer = nn.Linear(hidden_layers[-1], output_size)\n            nn.init.kaiming_normal_(self.output_layer.weight.data)\n        else:\n            self.hidden_layers = []\n            self.intermediate_layer = nn.Linear(hidden_size*2 if bidirectional else hidden_size, self.input_size)\n            self.output_layer = nn.Linear(hidden_size*2 if bidirectional else hidden_size, output_size)\n            nn.init.kaiming_normal_(self.output_layer.weight.data)\n\n        self.activation_fn = torch.nn.GELU()\n        self.dropout = nn.Dropout(dropout)\n\n    def forward(self, x):\n        outputs, hidden = self.rnn(x)\n\n        x = self.dropout(self.activation_fn(outputs))\n        for hidden_layer in self.hidden_layers:\n            x = self.activation_fn(hidden_layer(x))\n            x = self.dropout(x)\n        x = self.output_layer(x)\n\n        # (-1,60,14) -> (-1,386)\n        o_s = x[:, :, :6]\n        o_s = o_s.permute(0,2,1).reshape(-1,360)\n        o_g = x[:, :, 6:]\n        o_g = o_g.mean(dim=1)\n        out = torch.cat([o_s, o_g], dim=1)\n\n        return out\n\ninput_size = 25\noutput_size = 14\nseq_len = 60\n\nhidden_size = 256\nhidden_layers = [256, 512]\nnum_layers = 6\ndropout = 0.1\n\nmodel = LeapModel(\n    input_size=input_size,\n    seq_len=seq_len,\n    hidden_size=hidden_size,\n    output_size=output_size,\n    num_layers=num_layers,\n    hidden_layers=hidden_layers,\n    dropout=dropout,\n    bidirectional=True,\n).to(device)\n```\n\nReference Links: https://www.kaggle.com/code/brandenkmurray/seq2seq-rnn-with-gru\n\n- A BiLSTM with Transformer model\n\n```\nclass LeapModel(nn.Module):\n    def __init__(self,\n                 input_size,\n                 seq_len,\n                 hidden_size,\n                 output_size,\n                 num_layers=1,\n                 bidirectional=False,\n                 dropout=0.3,\n                 hidden_layers=[128, 256],\n                 nhead=8,\n                 num_transformer_layers=2):\n\n        super().__init__()\n        self.input_size = input_size\n        self.seq_len = seq_len\n        self.hidden_size = hidden_size\n        self.num_layers = num_layers\n        self.bidirectional = bidirectional\n        self.output_size = output_size\n\n        # LSTM layer\n        self.rnn = nn.LSTM(input_size=input_size,\n                           hidden_size=hidden_size,\n                           num_layers=num_layers,\n                           bidirectional=bidirectional,\n                           batch_first=True,\n                           dropout=dropout)\n\n        # Transformer layer\n        transformer_input_size = hidden_size * 2 if bidirectional else hidden_size\n        self.transformer_layer = nn.TransformerEncoder(\n            nn.TransformerEncoderLayer(d_model=transformer_input_size, nhead=nhead, dropout=dropout),\n            num_layers=num_transformer_layers\n        )\n\n        # Fully connected layers\n        if hidden_layers and len(hidden_layers):\n            first_layer = nn.Linear(transformer_input_size, hidden_layers[0])\n            self.hidden_layers = nn.ModuleList(\n                [first_layer] + \\\n                [nn.Linear(hidden_layers[i], hidden_layers[i+1]) for i in range(len(hidden_layers) - 1)]\n            )\n            for layer in self.hidden_layers:\n                nn.init.kaiming_normal_(layer.weight.data)\n            self.intermediate_layer = nn.Linear(hidden_layers[-1], self.input_size)\n            self.output_layer = nn.Linear(hidden_layers[-1], output_size)\n            nn.init.kaiming_normal_(self.output_layer.weight.data)\n        else:\n            self.hidden_layers = []\n            self.intermediate_layer = nn.Linear(transformer_input_size, self.input_size)\n            self.output_layer = nn.Linear(transformer_input_size, output_size)\n            nn.init.kaiming_normal_(self.output_layer.weight.data)\n\n        self.activation_fn = torch.nn.GELU()\n        self.dropout = nn.Dropout(dropout)\n\n    def forward(self, x):\n        # LSTM layer\n        lstm_output, _ = self.rnn(x)\n        \n        # Transformer layer\n        transformer_output = self.transformer_layer(lstm_output)\n        \n        # Apply dropout and activation\n        x = self.dropout(self.activation_fn(transformer_output))\n        \n        # Fully connected layers\n        for hidden_layer in self.hidden_layers:\n            x = self.activation_fn(hidden_layer(x))\n            x = self.dropout(x)\n        x = self.output_layer(x)\n\n        # Reshape output\n        o_s = x[:, :, :6]\n        o_s = o_s.permute(0, 2, 1).reshape(-1, 360)\n        o_g = x[:, :, 6:]\n        o_g = o_g.mean(dim=1)\n        out = torch.cat([o_s, o_g], dim=1)\n\n        return out\n\ninput_size = 25\noutput_size = 14\nseq_len = 60\n\nhidden_size = 256\nhidden_layers = [256, 512]\nnum_layers = 6\ndropout = 0.1\nnhead = 8\nnum_transformer_layers = 1\n\nmodel = LeapModel(\n    input_size=input_size,\n    seq_len=seq_len,\n    hidden_size=hidden_size,\n    output_size=output_size,\n    num_layers=num_layers,\n    hidden_layers=hidden_layers,\n    dropout=dropout,\n    bidirectional=True,\n    nhead=nhead,\n    num_transformer_layers=num_transformer_layers\n).to(device)\n```\n\n- A BiLSTM with TCN model\n\n```\nclass TCNBlock(nn.Module):\n    def __init__(self, in_channels, out_channels, kernel_size, dilation):\n        super(TCNBlock, self).__init__()\n        self.conv = nn.Conv1d(in_channels, out_channels, kernel_size, \n                              padding=(kernel_size-1) * dilation // 2, dilation=dilation)\n        self.bn = nn.BatchNorm1d(out_channels)\n        self.activation_fn = nn.GELU()\n\n    def forward(self, x):\n        return self.activation_fn(self.bn(self.conv(x)))\n\n\nclass LeapModel(nn.Module):\n    def __init__(self,\n                 input_size,\n                 seq_len,\n                 hidden_size,\n                 output_size,\n                 num_layers=1,\n                 bidirectional=False,\n                 dropout=0.3):\n\n        super().__init__()\n        self.input_size = input_size\n        self.seq_len = seq_len\n        self.hidden_size = hidden_size\n        self.num_layers = num_layers\n        self.bidirectional = bidirectional\n        self.output_size = output_size\n\n        # LSTM layer\n        self.rnn = nn.LSTM(input_size=input_size,\n                           hidden_size=hidden_size,\n                           num_layers=num_layers,\n                           bidirectional=bidirectional,\n                           batch_first=True,\n                           dropout=dropout)\n\n        self.se = nn.Sequential(\n            nn.Linear(hidden_size*2, hidden_size//2),\n            nn.GELU(),\n            nn.Linear(hidden_size//2, hidden_size*2),\n            nn.Sigmoid()\n        )\n        \n        self.tcn = nn.Sequential(\n            TCNBlock(hidden_size*2, hidden_size*2, kernel_size=3, dilation=1),\n            TCNBlock(hidden_size*2, hidden_size*2, kernel_size=3, dilation=2),\n            TCNBlock(hidden_size*2, hidden_size*2, kernel_size=3, dilation=4),\n            TCNBlock(hidden_size*2, hidden_size*2, kernel_size=3, dilation=8),\n        )\n        \n        self.fc = nn.Linear(hidden_size*2, output_size)\n        self.dropout = nn.Dropout(dropout)\n        \n\n    def forward(self, x):\n        # RNN layer\n        outputs, _ = self.rnn(x)\n        \n        se_weights = self.se(torch.mean(outputs, dim=1)).unsqueeze(1)\n        outputs = outputs * se_weights\n        \n        tcn_input = outputs.permute(0, 2, 1)\n        tcn_output = self.tcn(tcn_input)\n        tcn_output = tcn_output.permute(0, 2, 1)\n        \n        x = self.dropout(tcn_output)\n        x = self.fc(x)\n\n        # Reshape output\n        o_s = x[:, :, :6]\n        o_s = o_s.permute(0, 2, 1).reshape(-1, 360)\n        o_g = x[:, :, 6:]\n        o_g = o_g.mean(dim=1)\n        out = torch.cat([o_s, o_g], dim=1)  # (bs,368)\n\n        return out\n\n\ninput_size = 25\noutput_size = 14\nseq_len = 60\n\nhidden_size = 256\nnum_layers = 6\ndropout = 0.1\n\nmodel = LeapModel(\n    input_size=input_size,\n    seq_len=seq_len,\n    hidden_size=hidden_size,\n    output_size=output_size,\n    num_layers=num_layers,\n    dropout=dropout,\n    bidirectional=True,\n).to(device)\n```\n\n- A BiLSTM with Attention model\n\n```\nclass LeapModel(nn.Module):\n    def __init__(self,\n                 input_size,\n                 seq_len,\n                 hidden_size,\n                 output_size,\n                 num_layers=1,\n                 bidirectional=False,\n                 dropout=.3,\n                 hidden_layers=[128, 256]):\n\n        super().__init__()\n        self.input_size = input_size\n        self.seq_len = seq_len\n        self.hidden_size = hidden_size\n        self.num_layers = num_layers\n        self.bidirectional=bidirectional\n        self.output_size=output_size\n        self.rnn = nn.LSTM(input_size=input_size,\n                           hidden_size=hidden_size,\n                           num_layers=num_layers,\n                           bidirectional=bidirectional,\n                           batch_first=True,\n                           dropout=0.1)\n\n        self.attention = nn.MultiheadAttention(embed_dim=hidden_size*2 if bidirectional else hidden_size,\n                                               num_heads=8,\n                                               batch_first=True)\n\n        if hidden_layers and len(hidden_layers):\n            first_layer  = nn.Linear(hidden_size*2 if bidirectional else hidden_size, hidden_layers[0])\n            self.hidden_layers = nn.ModuleList(\n                [first_layer] + \\\n                [nn.Linear(hidden_layers[i], hidden_layers[i+1]) for i in range(len(hidden_layers) - 1)]\n            )\n            for layer in self.hidden_layers:\n                nn.init.kaiming_normal_(layer.weight.data)\n            self.intermediate_layer = nn.Linear(hidden_layers[-1], self.input_size)\n            self.output_layer = nn.Linear(hidden_layers[-1], output_size)\n            nn.init.kaiming_normal_(self.output_layer.weight.data)\n        else:\n            self.hidden_layers = []\n            self.intermediate_layer = nn.Linear(hidden_size*2 if bidirectional else hidden_siz, self.input_size)\n            self.output_layer = nn.Linear(hidden_size*2 if bidirectional else hidden_size, output_size)\n            nn.init.kaiming_normal_(self.output_layer.weight.data)\n\n        self.activation_fn = torch.nn.GELU()\n        self.dropout = nn.Dropout(dropout)\n\n    def forward(self, x):\n        batch_size = x.size(0)\n        outputs, hidden = self.rnn(x)\n\n        outputs = outputs.permute(1, 0, 2)  # (seq_len, batch_size, hidden_size)\n        attn_output, _ = self.attention(outputs, outputs, outputs)\n        attn_output = attn_output.permute(1, 0, 2)  # (batch_size, seq_len, hidden_size)\n\n        x = self.dropout(self.activation_fn(attn_output))\n        for hidden_layer in self.hidden_layers:\n            x = self.activation_fn(hidden_layer(x))\n            x = self.dropout(x)\n        x = self.output_layer(x)\n\n        # (-1,60,14) -> (-1,386)\n        o_s = x[:, :, :6]\n        o_s = o_s.permute(0,2,1).reshape(-1,360)\n        o_g = x[:, :, 6:]\n        o_g = o_g.mean(dim=1)\n        out = torch.cat([o_s, o_g], dim=1)\n\n        return out\n\n\ninput_size = 25\noutput_size = 14\nseq_len = 60\n\nhidden_size = 256\nhidden_layers = [256, 512]\nnum_layers = 6\ndropout = 0.1\n\nmodel = LeapModel(\n    input_size=input_size,\n    seq_len=seq_len,\n    hidden_size=hidden_size,\n    output_size=output_size,\n    num_layers=num_layers,\n    hidden_layers=hidden_layers,\n    dropout=dropout,\n    bidirectional=True,\n).to(device)\n```\n\n### Dataset\n\n1. Download all 0001-02 to 0009-01  (low resolution) data from Huggingface.\n2. 0001-02 to 0008-06 as training set, 0008-07 to 0009-01 as validate set (sampling to ~625000 rows).\n\n### Some training details\n\n1. Loss: nn.SmoothL1Loss(reduction='mean') (0.005~0.008 better than mse)\n2. Scheduler: get_cosine_schedule_with_warmup\n3. Activation Function: GELU (0.002~0.004 better than relu)\n4. Trained on 4*RTX4090 with 360G RAM, 7.5 years training dataset, ~1 hour per epoch\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1197382%2Ff9d38b78021aa6aaa63bd77da4bea9c3%2Fscreencapture-a18044-8b4e-f62d61d2-westb-seetacloud-8443-monitor-2024-07-10-10_05_39.png?generation=1722405976252082&alt=media)\n\n### Post-Processing\n\n```\ntargets_unpredictable = []\nfor target in weights:\n    if weights[target] == 0.:\n        targets_unpredictable.append(target)\nfor target in targets_unpredictable:\n    df_pred[target] = 0.\nfor target in [f'ptend_q0002_{i}' for i in range(12, 28)]:\n    df_pred[target] = -df_test[target.replace(\"ptend\", \"state\")] * weights[target] / 1200.\n```\n\nReference Links: https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/502484\n\n### Ensemble\n\n1. on models: w0 * pred0 + w1 * pred1 + ... + w5 * pred5\n2. on targets:\n\n```\nselects = []\nfor idx_t, target in tqdm(enumerate(TARGETCOLS), total=len(TARGETCOLS)):\n    di = {}\n    for idx_p, prob in enumerate(probs):\n        di[idx_p] = r2_score(df_valid[target], probs[idx_p][:, idx_t])\n    selects.append(sorted(di, key=di.get, reverse=True)[:4])\n````\n\n### What didn't work\n\n1. All 8 years dataset with fix epochs, without validation.\n2. Data augment: mask 10% input and TTA.\n",
    "2941893": "Congrats on solo gold and GM!",
    "2941501": "Congrats on achieving GM!",
    "2953728": "Congratulations on becoming a GM. I wanted to ask if you've encountered any memory issues with dataloaders in your use of PyTorch?\n\nSpecifically, I'd like to know how you train models on multiple GPUs? Do you use Data Parallel (DP), Distributed Data Parallel (DDP), or something else? For me, when I use DDP on multiple GPUs, the pseudocode below results in an inability to properly release the memory occupied by the dataloaders:\n\n```\nfor epoch in range(CFG.num_epochs):\n    for j in range(12):  # 12 months\n        train_loader = get_dataloaders(CFG, j=j)\n        train_loss, grad_norm = train(\n            model, criterion, optimizer, scheduler, train_loader, gpu_id, CFG\n        )\n        del train_loader  # failed to release memory\n        clean_memory()\n    valid()\n```"
  }
}