{
  "id": 523040,
  "title": "[5th solution]: Ensemble of Bidirectional LSTM based Models",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/523040",
  "author_name": "Marec Serlin",
  "post_date": "2024-07-29T22:39:22.661000",
  "votes": 32,
  "comment_count": 1,
  "views": 0,
  "content": "<h1>Github Code</h1>\n<p><a href=\"https://github.com/YusefAN/leap-climsim-kaggle-5th\" target=\"_blank\">https://github.com/YusefAN/leap-climsim-kaggle-5th</a></p>\n<h1>Explanation</h1>\n<p>We achieved our results by training an ensemble of models, each with its own unique variations. The core of each model architecture is a bidirectional LSTM, with different surrounding layers and engineered features. These variations, detailed below, enabled our models to ensemble effectively, leading to our competitive score and 5th-place finish.</p>\n<p>Overall, our model architectures remain very simple and manageable in size (each having 20 million parameters or so), enabling us to efficiently train on as much data as possible. Some models were trained on the <a href=\"https://huggingface.co/datasets/LEAP/ClimSim_low-res\" target=\"_blank\">ClimSim_low-res</a> dataset combined with the <a href=\"https://huggingface.co/datasets/LEAP/ClimSim_low-res_aqua-planet\" target=\"_blank\">ClimSim_low-res_aqua-planet</a> dataset, while others used the <a href=\"https://huggingface.co/datasets/LEAP/ClimSim_low-res\" target=\"_blank\">ClimSim_low-res</a> dataset combined with a subset of <a href=\"https://huggingface.co/datasets/LEAP/ClimSim_high-res\" target=\"_blank\">ClimSim_high-res</a> data. Our training time for even our hungriest models did not exceed 24 hours. </p>\n<h2>Preprocessing</h2>\n<h3>Low-Res-Aqua-Mixed</h3>\n<p>The 556 element inputs were separated into column inputs (60x9 = 540) and global inputs (1x16). The global inputs were then repeated 60 times and concatenated to the columns to form a (60x25) input to our models. Engineered features, described in the individual model sections, are added as additional columns. All inputs and outputs, including engineered features, were standardized using their mean and standard deviation without any additional transformations.</p>\n<h3>Low-Res-High-Res-Mixed</h3>\n<p>As for the previous model, the 556 element inputs are separated into column inputs (60x9 = 540) and global inputs (1x16). The global inputs are repeated 60 times and concatenated to the columns to form a (60x25) input. These models did not use any engineered features. The inputs 'state-_q0001', 'state_q0002', 'state_q0003', 'pbuf_ozone', 'pbuf_CH4', 'pbuf_N2O' were log-transformed. Then, all the inputs (both transformed and not) were standardized using their mean and standard deviation. The outputs are predicted directly without any normalization. The models are trained using all the low-res data and about 1/15th of the high-res data.</p>\n<h2>Models</h2>\n<h3>Low-Res-Aqua-Mixed</h3>\n<p>The models trained on the combination of the low-res and the aqua-planet datasets had the following simple architecture:</p>\n<ul>\n<li>MLP encoder-decoder on the input, outputting a 60x25 matrix which is concatenated to the input</li>\n<li>The concatenated input is fed into a wide (hidden dimension of 512) but shallow (3 layers deep) bidirectional LSTM </li>\n<li>The output is fed into a single bidirectional GRU layer</li>\n<li>Final MLP encoder to produce the 368-element output sequence</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20298348%2Fc07d755c212f9de991b58c53a37c2a1a%2FLowResAquaArchitecture.jpg?generation=1722292377019977&amp;alt=media\" alt=\"\"></p>\n<p>Several models in our final ensemble followed this procedure without additional feature engineering. Others included the following features:</p>\n<ol>\n<li>liq_partition STATE_Q0002 / (STATE_Q0002 + STATE_Q0003)</li>\n<li>imbalance (STATE_Q0002 - STATE_Q0003) / (STATE_Q0002_IDX + STATE_Q0003)</li>\n<li>moisture = (STATE_Q0001) * (STATE_U)**2 + (STATE_V)**2)</li>\n<li>air_total = PBUF_OZONE + PBUF_CH4 + PBUF_N2O</li>\n<li>temp_humid = STATE_T / STATE_Q0001</li>\n<li>temp_diff1 = STATE_T_{i} - STATE_T_IDX_{i+1}</li>\n<li>temp_diff2 = STATE_Q0001_{i} - STATE_Q0001_IDX_{i+1}</li>\n<li>wind_diff1 = STATE_U_{i} - STATE_U_{i+1}</li>\n<li>wind_diff2 = STATE_V_{i} - STATE_V_{i+1}</li>\n</ol>\n<p>Model code: </p>\n<pre><code> (nn.Module):\n     ():\n        ().__init__()\n        \n        self.encoder = MLP([input_dim * , input_dim * , input_dim * , input_dim * ])\n        self.decoder = MLP([input_dim * , input_dim * , input_dim * , input_dim * ])\n\n        self.lstm = nn.LSTM(input_dim * , hidden_dim, num_layers, batch_first=, dropout=, bidirectional=)\n\n        self.lstm2 = nn.GRU(hidden_dim * , , batch_first=, dropout=, bidirectional=)\n\n        \n        self.fc_lstm = MLP([ * ,  * , output_dim])\n\n        self.criterion = nn.HuberLoss(delta=) \n\n     ():\n        x_0 = self.encoder(x.view(x.size(), -))\n        x_1 = self.decoder(x_0)\n        x_1 = x_1.view(x.size(), , -)\n\n        lstm_out, _ = self.lstm(torch.cat([x, x_1], dim = -))\n        \n        lstm_out, _ = self.lstm2(lstm_out)\n        lstm_out = lstm_out.contiguous().view(x.size(), -)  \n        output = self.fc_lstm(lstm_out)\n         output\n</code></pre>\n<h3>Low-Res-High-Res-Mixed</h3>\n<p>The 3 best-performing models trained on the combination of the low-res and the high-res datasets were bidirectional LSTMs with outfitted with various FFNNs. They all shared the following features:</p>\n<ul>\n<li>Linear layer expanding the input dimension</li>\n<li>Bidirectional LSTM (6 layers deep) with hidden dimension ranging from 256 to 320</li>\n<li>The LSTM output is put into a 1D average pooling layer and a mean layer (along the dimension with length 60, which is repeated 60 times)</li>\n<li>The original input, the LSTM output, the 1D average pooling output, and the mean are all concatenated</li>\n<li>A series of linear layers then produce the final 368-element output sequence</li>\n</ul>\n<p>For 2 out of 3 of the models, an additional linear layer (not pictured in the diagram) with a softmax activation is used to compute an \"attention\" for computing the 8x1 flattened output.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20298348%2F5e6dc8f287158a7820bb3aee3b02fe5f%2FLowResHighResArchitecture.jpg?generation=1722292441625060&amp;alt=media\" alt=\"\"></p>\n<p>Model code: </p>\n<pre><code> (nn.Module):\n     ():\n        (FFNN_LSTM_6_AVG, self).__init__()\n\n        self.encode_dim = \n        self.hidden_dim = \n        self.iter_dim = \n\n        self.LSTM_1 = LSTM(self.encode_dim,self.hidden_dim,,batch_first=,dropout=,bidirectional=)\n        self.input_size = input_size\n        self.Linear_1 = nn.Linear((seq_fea_list)+(num_fea_list), self.encode_dim)\n        self.Linear_2 = nn.Linear(*self.hidden_dim+self.encode_dim, self.iter_dim)\n        self.Linear_3 = nn.Linear(self.iter_dim, (seq_y_list))\n        self.Linear_4_0 = nn.Linear(self.iter_dim, self.iter_dim*)\n\n        self.Linear_4 = nn.Linear(self.iter_dim*, (num_y_list))\n        self.bias = nn.Linear((seq_y_list)*+(num_y_list),)\n        self.weight = nn.Linear((seq_y_list)*+(num_y_list),)\n        self.avg_pool_1 = AvgPool1d(kernel_size=,stride=,padding=)\n\n     ():\n        x_seq = x[:,:*(seq_fea_list)]\n        x_seq = x_seq.reshape((-,(seq_fea_list),))\n        x_seq = torch.transpose(x_seq, , )\n\n        x_num = x[:,*(seq_fea_list):x.shape[]]\n        x_num_repeat = x_num.reshape((-,,(num_fea_list)))\n        x_num_repeat = x_num_repeat.repeat((,,))\n\n        x_seq = F.elu(self.Linear_1(torch.concat((x_seq,x_num_repeat),dim=-)/))\n\n        x_seq_1,_ = self.LSTM_1(x_seq/)\n\n        x_seq_1_mean = torch.mean(x_seq_1,dim=,keepdim=)\n        x_seq_1_mean = x_seq_1_mean.repeat((,,))\n\n        x_seq_1_avg_pool = self.avg_pool_1(torch.transpose(x_seq_1, , ))\n        x_seq_1_avg_pool = torch.transpose(x_seq_1_avg_pool,, )\n\n        x_seq_1 = F.elu(self.Linear_2(torch.cat((x_seq_1,x_seq_1_mean,x_seq,x_seq_1_avg_pool),dim=-)/))\n\n        x_seq_out = self.Linear_3(x_seq_1)\n        x_seq_out = torch.transpose(x_seq_out, , )\n        x_seq_out = x_seq_out.reshape((-,*(seq_y_list)))\n\n        x_num_out = F.elu(self.Linear_4_0(torch.mean(x_seq_1,dim=)))\n        x_num_out = self.Linear_4(x_num_out)\n\n        output = self.weight.weight*(torch.concat((x_seq_out,x_num_out),dim=-))/+self.bias.weight/\n\n        output[:,zeroout_index] =  output[:,zeroout_index]*\n\n         output\n</code></pre>\n<h2>Training</h2>\n<h3>Low-Res-Aqua-Mixed</h3>\n<p>Typically, our batch sizes were 4-5 files or 4000-5000 data points. We found that training using a Huber loss function with delta = 1 noticeably improves the model's performance. A single model is trained on the mixed low-res and aqua-planet data to completion. Then, the same model is fine-tuned using slightly different procedures (varying SWA and checkpoint averaging parameters) to create distinct models that ensemble effectively. Training to completion takes approximately 24 hours on a single 4090 GPU.</p>\n<h3>Low-Res-High-Res-Mixed</h3>\n<p>For each step during training, a batch of low-res data and a batch of high-res data are evaluated using the model. The loss function for the step is a weighted combination of the loss on the low-res data and the high-res data. The loss functions used are a combination of MSE and MAE loss functions. About 72% of the total loss weight is placed on the MAE of the low-res data, 6% on the MSE of the low-res data, 22% on the MAE of the high-res data, and &lt;1% on the MSE of the high-res data. The best-performing models were equipped with EMA and trained with a batch size of about 7000 data points. Training takes about 12 hours when trained locally on a computer with 6 RTX 4090 GPUs.</p>\n<h2>Ensemble Inferencing</h2>\n<p>The best individual model was a Low-Res-High-Res-Mixed model that achieved a public LB score of 0.78553. The best Low-Res-Aqua-Mixed model was not too different, with an LB public score of 0.78457.</p>\n<p>The optimal weighting for the final ensemble was determined experimentally building on the intuition that the Low-Res-High-Res-Mixed model performed slightly better, and therefore received slightly larger weights. Predictions from models and ensembles from various stages of development were included in the final ensemble. Prior to submitting the final ensemble, an additional postprocessing step was applied to variables q0001, q0002, and q0003 whereby the maximum between the predicted ptend value and -state/1200 is chosen. The final ensemble had a public LB score of 0.79071.</p>",
  "messages": [
    {
      "id": 2940257,
      "postDate": "2024-07-29T22:39:22.660Z",
      "content": "<h1>Github Code</h1>\n<p><a href=\"https://github.com/YusefAN/leap-climsim-kaggle-5th\" target=\"_blank\">https://github.com/YusefAN/leap-climsim-kaggle-5th</a></p>\n<h1>Explanation</h1>\n<p>We achieved our results by training an ensemble of models, each with its own unique variations. The core of each model architecture is a bidirectional LSTM, with different surrounding layers and engineered features. These variations, detailed below, enabled our models to ensemble effectively, leading to our competitive score and 5th-place finish.</p>\n<p>Overall, our model architectures remain very simple and manageable in size (each having 20 million parameters or so), enabling us to efficiently train on as much data as possible. Some models were trained on the <a href=\"https://huggingface.co/datasets/LEAP/ClimSim_low-res\" target=\"_blank\">ClimSim_low-res</a> dataset combined with the <a href=\"https://huggingface.co/datasets/LEAP/ClimSim_low-res_aqua-planet\" target=\"_blank\">ClimSim_low-res_aqua-planet</a> dataset, while others used the <a href=\"https://huggingface.co/datasets/LEAP/ClimSim_low-res\" target=\"_blank\">ClimSim_low-res</a> dataset combined with a subset of <a href=\"https://huggingface.co/datasets/LEAP/ClimSim_high-res\" target=\"_blank\">ClimSim_high-res</a> data. Our training time for even our hungriest models did not exceed 24 hours. </p>\n<h2>Preprocessing</h2>\n<h3>Low-Res-Aqua-Mixed</h3>\n<p>The 556 element inputs were separated into column inputs (60x9 = 540) and global inputs (1x16). The global inputs were then repeated 60 times and concatenated to the columns to form a (60x25) input to our models. Engineered features, described in the individual model sections, are added as additional columns. All inputs and outputs, including engineered features, were standardized using their mean and standard deviation without any additional transformations.</p>\n<h3>Low-Res-High-Res-Mixed</h3>\n<p>As for the previous model, the 556 element inputs are separated into column inputs (60x9 = 540) and global inputs (1x16). The global inputs are repeated 60 times and concatenated to the columns to form a (60x25) input. These models did not use any engineered features. The inputs 'state-_q0001', 'state_q0002', 'state_q0003', 'pbuf_ozone', 'pbuf_CH4', 'pbuf_N2O' were log-transformed. Then, all the inputs (both transformed and not) were standardized using their mean and standard deviation. The outputs are predicted directly without any normalization. The models are trained using all the low-res data and about 1/15th of the high-res data.</p>\n<h2>Models</h2>\n<h3>Low-Res-Aqua-Mixed</h3>\n<p>The models trained on the combination of the low-res and the aqua-planet datasets had the following simple architecture:</p>\n<ul>\n<li>MLP encoder-decoder on the input, outputting a 60x25 matrix which is concatenated to the input</li>\n<li>The concatenated input is fed into a wide (hidden dimension of 512) but shallow (3 layers deep) bidirectional LSTM </li>\n<li>The output is fed into a single bidirectional GRU layer</li>\n<li>Final MLP encoder to produce the 368-element output sequence</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20298348%2Fc07d755c212f9de991b58c53a37c2a1a%2FLowResAquaArchitecture.jpg?generation=1722292377019977&amp;alt=media\" alt=\"\"></p>\n<p>Several models in our final ensemble followed this procedure without additional feature engineering. Others included the following features:</p>\n<ol>\n<li>liq_partition STATE_Q0002 / (STATE_Q0002 + STATE_Q0003)</li>\n<li>imbalance (STATE_Q0002 - STATE_Q0003) / (STATE_Q0002_IDX + STATE_Q0003)</li>\n<li>moisture = (STATE_Q0001) * (STATE_U)**2 + (STATE_V)**2)</li>\n<li>air_total = PBUF_OZONE + PBUF_CH4 + PBUF_N2O</li>\n<li>temp_humid = STATE_T / STATE_Q0001</li>\n<li>temp_diff1 = STATE_T_{i} - STATE_T_IDX_{i+1}</li>\n<li>temp_diff2 = STATE_Q0001_{i} - STATE_Q0001_IDX_{i+1}</li>\n<li>wind_diff1 = STATE_U_{i} - STATE_U_{i+1}</li>\n<li>wind_diff2 = STATE_V_{i} - STATE_V_{i+1}</li>\n</ol>\n<p>Model code: </p>\n<pre><code> (nn.Module):\n     ():\n        ().__init__()\n        \n        self.encoder = MLP([input_dim * , input_dim * , input_dim * , input_dim * ])\n        self.decoder = MLP([input_dim * , input_dim * , input_dim * , input_dim * ])\n\n        self.lstm = nn.LSTM(input_dim * , hidden_dim, num_layers, batch_first=, dropout=, bidirectional=)\n\n        self.lstm2 = nn.GRU(hidden_dim * , , batch_first=, dropout=, bidirectional=)\n\n        \n        self.fc_lstm = MLP([ * ,  * , output_dim])\n\n        self.criterion = nn.HuberLoss(delta=) \n\n     ():\n        x_0 = self.encoder(x.view(x.size(), -))\n        x_1 = self.decoder(x_0)\n        x_1 = x_1.view(x.size(), , -)\n\n        lstm_out, _ = self.lstm(torch.cat([x, x_1], dim = -))\n        \n        lstm_out, _ = self.lstm2(lstm_out)\n        lstm_out = lstm_out.contiguous().view(x.size(), -)  \n        output = self.fc_lstm(lstm_out)\n         output\n</code></pre>\n<h3>Low-Res-High-Res-Mixed</h3>\n<p>The 3 best-performing models trained on the combination of the low-res and the high-res datasets were bidirectional LSTMs with outfitted with various FFNNs. They all shared the following features:</p>\n<ul>\n<li>Linear layer expanding the input dimension</li>\n<li>Bidirectional LSTM (6 layers deep) with hidden dimension ranging from 256 to 320</li>\n<li>The LSTM output is put into a 1D average pooling layer and a mean layer (along the dimension with length 60, which is repeated 60 times)</li>\n<li>The original input, the LSTM output, the 1D average pooling output, and the mean are all concatenated</li>\n<li>A series of linear layers then produce the final 368-element output sequence</li>\n</ul>\n<p>For 2 out of 3 of the models, an additional linear layer (not pictured in the diagram) with a softmax activation is used to compute an \"attention\" for computing the 8x1 flattened output.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20298348%2F5e6dc8f287158a7820bb3aee3b02fe5f%2FLowResHighResArchitecture.jpg?generation=1722292441625060&amp;alt=media\" alt=\"\"></p>\n<p>Model code: </p>\n<pre><code> (nn.Module):\n     ():\n        (FFNN_LSTM_6_AVG, self).__init__()\n\n        self.encode_dim = \n        self.hidden_dim = \n        self.iter_dim = \n\n        self.LSTM_1 = LSTM(self.encode_dim,self.hidden_dim,,batch_first=,dropout=,bidirectional=)\n        self.input_size = input_size\n        self.Linear_1 = nn.Linear((seq_fea_list)+(num_fea_list), self.encode_dim)\n        self.Linear_2 = nn.Linear(*self.hidden_dim+self.encode_dim, self.iter_dim)\n        self.Linear_3 = nn.Linear(self.iter_dim, (seq_y_list))\n        self.Linear_4_0 = nn.Linear(self.iter_dim, self.iter_dim*)\n\n        self.Linear_4 = nn.Linear(self.iter_dim*, (num_y_list))\n        self.bias = nn.Linear((seq_y_list)*+(num_y_list),)\n        self.weight = nn.Linear((seq_y_list)*+(num_y_list),)\n        self.avg_pool_1 = AvgPool1d(kernel_size=,stride=,padding=)\n\n     ():\n        x_seq = x[:,:*(seq_fea_list)]\n        x_seq = x_seq.reshape((-,(seq_fea_list),))\n        x_seq = torch.transpose(x_seq, , )\n\n        x_num = x[:,*(seq_fea_list):x.shape[]]\n        x_num_repeat = x_num.reshape((-,,(num_fea_list)))\n        x_num_repeat = x_num_repeat.repeat((,,))\n\n        x_seq = F.elu(self.Linear_1(torch.concat((x_seq,x_num_repeat),dim=-)/))\n\n        x_seq_1,_ = self.LSTM_1(x_seq/)\n\n        x_seq_1_mean = torch.mean(x_seq_1,dim=,keepdim=)\n        x_seq_1_mean = x_seq_1_mean.repeat((,,))\n\n        x_seq_1_avg_pool = self.avg_pool_1(torch.transpose(x_seq_1, , ))\n        x_seq_1_avg_pool = torch.transpose(x_seq_1_avg_pool,, )\n\n        x_seq_1 = F.elu(self.Linear_2(torch.cat((x_seq_1,x_seq_1_mean,x_seq,x_seq_1_avg_pool),dim=-)/))\n\n        x_seq_out = self.Linear_3(x_seq_1)\n        x_seq_out = torch.transpose(x_seq_out, , )\n        x_seq_out = x_seq_out.reshape((-,*(seq_y_list)))\n\n        x_num_out = F.elu(self.Linear_4_0(torch.mean(x_seq_1,dim=)))\n        x_num_out = self.Linear_4(x_num_out)\n\n        output = self.weight.weight*(torch.concat((x_seq_out,x_num_out),dim=-))/+self.bias.weight/\n\n        output[:,zeroout_index] =  output[:,zeroout_index]*\n\n         output\n</code></pre>\n<h2>Training</h2>\n<h3>Low-Res-Aqua-Mixed</h3>\n<p>Typically, our batch sizes were 4-5 files or 4000-5000 data points. We found that training using a Huber loss function with delta = 1 noticeably improves the model's performance. A single model is trained on the mixed low-res and aqua-planet data to completion. Then, the same model is fine-tuned using slightly different procedures (varying SWA and checkpoint averaging parameters) to create distinct models that ensemble effectively. Training to completion takes approximately 24 hours on a single 4090 GPU.</p>\n<h3>Low-Res-High-Res-Mixed</h3>\n<p>For each step during training, a batch of low-res data and a batch of high-res data are evaluated using the model. The loss function for the step is a weighted combination of the loss on the low-res data and the high-res data. The loss functions used are a combination of MSE and MAE loss functions. About 72% of the total loss weight is placed on the MAE of the low-res data, 6% on the MSE of the low-res data, 22% on the MAE of the high-res data, and &lt;1% on the MSE of the high-res data. The best-performing models were equipped with EMA and trained with a batch size of about 7000 data points. Training takes about 12 hours when trained locally on a computer with 6 RTX 4090 GPUs.</p>\n<h2>Ensemble Inferencing</h2>\n<p>The best individual model was a Low-Res-High-Res-Mixed model that achieved a public LB score of 0.78553. The best Low-Res-Aqua-Mixed model was not too different, with an LB public score of 0.78457.</p>\n<p>The optimal weighting for the final ensemble was determined experimentally building on the intuition that the Low-Res-High-Res-Mixed model performed slightly better, and therefore received slightly larger weights. Predictions from models and ensembles from various stages of development were included in the final ensemble. Prior to submitting the final ensemble, an additional postprocessing step was applied to variables q0001, q0002, and q0003 whereby the maximum between the predicted ptend value and -state/1200 is chosen. The final ensemble had a public LB score of 0.79071.</p>",
      "rawMarkdown": "# Github Code\n\nhttps://github.com/YusefAN/leap-climsim-kaggle-5th\n\n# Explanation\n\nWe achieved our results by training an ensemble of models, each with its own unique variations. The core of each model architecture is a bidirectional LSTM, with different surrounding layers and engineered features. These variations, detailed below, enabled our models to ensemble effectively, leading to our competitive score and 5th-place finish.\n\nOverall, our model architectures remain very simple and manageable in size (each having 20 million parameters or so), enabling us to efficiently train on as much data as possible. Some models were trained on the [ClimSim_low-res](https://huggingface.co/datasets/LEAP/ClimSim_low-res) dataset combined with the [ClimSim_low-res_aqua-planet](https://huggingface.co/datasets/LEAP/ClimSim_low-res_aqua-planet) dataset, while others used the [ClimSim_low-res](https://huggingface.co/datasets/LEAP/ClimSim_low-res) dataset combined with a subset of [ClimSim_high-res](https://huggingface.co/datasets/LEAP/ClimSim_high-res) data. Our training time for even our hungriest models did not exceed 24 hours. \n\n## Preprocessing\n\n### Low-Res-Aqua-Mixed\n\nThe 556 element inputs were separated into column inputs (60x9 = 540) and global inputs (1x16). The global inputs were then repeated 60 times and concatenated to the columns to form a (60x25) input to our models. Engineered features, described in the individual model sections, are added as additional columns. All inputs and outputs, including engineered features, were standardized using their mean and standard deviation without any additional transformations.\n\n### Low-Res-High-Res-Mixed \n\nAs for the previous model, the 556 element inputs are separated into column inputs (60x9 = 540) and global inputs (1x16). The global inputs are repeated 60 times and concatenated to the columns to form a (60x25) input. These models did not use any engineered features. The inputs 'state-\\_q0001', 'state\\_q0002', 'state\\_q0003', 'pbuf\\_ozone', 'pbuf\\_CH4', 'pbuf\\_N2O' were log-transformed. Then, all the inputs (both transformed and not) were standardized using their mean and standard deviation. The outputs are predicted directly without any normalization. The models are trained using all the low-res data and about 1/15th of the high-res data.\n\n## Models\n\n### Low-Res-Aqua-Mixed\n\nThe models trained on the combination of the low-res and the aqua-planet datasets had the following simple architecture:\n- MLP encoder-decoder on the input, outputting a 60x25 matrix which is concatenated to the input\n- The concatenated input is fed into a wide (hidden dimension of 512) but shallow (3 layers deep) bidirectional LSTM \n- The output is fed into a single bidirectional GRU layer\n- Final MLP encoder to produce the 368-element output sequence\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20298348%2Fc07d755c212f9de991b58c53a37c2a1a%2FLowResAquaArchitecture.jpg?generation=1722292377019977&alt=media)\n\nSeveral models in our final ensemble followed this procedure without additional feature engineering. Others included the following features:\n1. liq_partition STATE_Q0002 / (STATE_Q0002 + STATE_Q0003)\n2. imbalance (STATE_Q0002 - STATE_Q0003) / (STATE_Q0002_IDX + STATE_Q0003)\n3. moisture = (STATE_Q0001) * (STATE_U)\\*\\*2 + (STATE_V)\\*\\*2)\n4. air_total = PBUF_OZONE + PBUF_CH4 + PBUF_N2O\n5. temp_humid = STATE_T / STATE_Q0001\n6. temp_diff1 = STATE_T_{i} - STATE_T_IDX_{i+1}\n7. temp_diff2 = STATE_Q0001_{i} - STATE_Q0001_IDX_{i+1}\n8. wind_diff1 = STATE_U_{i} - STATE_U_{i+1}\n9. wind_diff2 = STATE_V_{i} - STATE_V_{i+1}\n\nModel code: \n\n```python\nclass LeapModel(nn.Module):\n    def __init__(self, input_dim, hidden_dim, output_dim, num_layers=2):\n        super().__init__()\n        # self.N_SEQ_OUTPUTS = N_SEQ_OUTPUTS\n        self.encoder = MLP([input_dim * 60, input_dim * 15, input_dim * 8, input_dim * 4])\n        self.decoder = MLP([input_dim * 4, input_dim * 8, input_dim * 15, input_dim * 60])\n\n        self.lstm = nn.LSTM(input_dim * 2, hidden_dim, num_layers, batch_first=True, dropout=0.09, bidirectional=True)\n        \n        self.lstm2 = nn.GRU(hidden_dim * 2, 16, batch_first=True, dropout=0.00, bidirectional=True)\n\n        # self.fc_lstm = nn.Linear((32) * 60 , output_dim)\n        self.fc_lstm = MLP([32 * 60, 16 * 60, output_dim])#nn.Linear((32) * 60 , output_dim)\n        \n        self.criterion = nn.HuberLoss(delta=1) ########################################\n\n    def forward(self, x):\n        x_0 = self.encoder(x.view(x.size(0), -1))\n        x_1 = self.decoder(x_0)\n        x_1 = x_1.view(x.size(0), 60, -1)\n        \n        lstm_out, _ = self.lstm(torch.cat([x, x_1], dim = -1))\n        # lstm_out, _ = self.gru(lstm_out)\n        lstm_out, _ = self.lstm2(lstm_out)\n        lstm_out = lstm_out.contiguous().view(x.size(0), -1)  # Flatten the LSTM output\n        output = self.fc_lstm(lstm_out)\n        return output\n```\n\n### Low-Res-High-Res-Mixed\n\nThe 3 best-performing models trained on the combination of the low-res and the high-res datasets were bidirectional LSTMs with outfitted with various FFNNs. They all shared the following features:\n\n* Linear layer expanding the input dimension\n* Bidirectional LSTM (6 layers deep) with hidden dimension ranging from 256 to 320\n* The LSTM output is put into a 1D average pooling layer and a mean layer (along the dimension with length 60, which is repeated 60 times)\n* The original input, the LSTM output, the 1D average pooling output, and the mean are all concatenated\n* A series of linear layers then produce the final 368-element output sequence\n\nFor 2 out of 3 of the models, an additional linear layer (not pictured in the diagram) with a softmax activation is used to compute an \"attention\" for computing the 8x1 flattened output.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20298348%2F5e6dc8f287158a7820bb3aee3b02fe5f%2FLowResHighResArchitecture.jpg?generation=1722292441625060&alt=media)\n\nModel code: \n\n```python\nclass FFNN_LSTM_6_AVG(nn.Module):\n    def __init__(self, input_size, output_size):\n        super(FFNN_LSTM_6_AVG, self).__init__()\n        \n        self.encode_dim = 300\n        self.hidden_dim = 280\n        self.iter_dim = 800\n\n        self.LSTM_1 = LSTM(self.encode_dim,self.hidden_dim,6,batch_first=True,dropout=0.01,bidirectional=True)\n        self.input_size = input_size\n        self.Linear_1 = nn.Linear(len(seq_fea_list)+len(num_fea_list), self.encode_dim)\n        self.Linear_2 = nn.Linear(6*self.hidden_dim+self.encode_dim, self.iter_dim)\n        self.Linear_3 = nn.Linear(self.iter_dim, len(seq_y_list))\n        self.Linear_4_0 = nn.Linear(self.iter_dim, self.iter_dim*2)\n\n        self.Linear_4 = nn.Linear(self.iter_dim*2, len(num_y_list))\n        self.bias = nn.Linear(len(seq_y_list)*60+len(num_y_list),1)\n        self.weight = nn.Linear(len(seq_y_list)*60+len(num_y_list),1)\n        self.avg_pool_1 = AvgPool1d(kernel_size=3,stride=1,padding=1)\n        \n    def forward(self, x):\n        x_seq = x[:,0:60*len(seq_fea_list)]\n        x_seq = x_seq.reshape((-1,len(seq_fea_list),60))\n        x_seq = torch.transpose(x_seq, 1, 2)\n        \n        x_num = x[:,60*len(seq_fea_list):x.shape[1]]\n        x_num_repeat = x_num.reshape((-1,1,len(num_fea_list)))\n        x_num_repeat = x_num_repeat.repeat((1,60,1))\n        \n        x_seq = F.elu(self.Linear_1(torch.concat((x_seq,x_num_repeat),dim=-1)/5))\n        \n        x_seq_1,_ = self.LSTM_1(x_seq/5)\n        \n        x_seq_1_mean = torch.mean(x_seq_1,dim=1,keepdim=True)\n        x_seq_1_mean = x_seq_1_mean.repeat((1,60,1))\n\n        x_seq_1_avg_pool = self.avg_pool_1(torch.transpose(x_seq_1, 1, 2))\n        x_seq_1_avg_pool = torch.transpose(x_seq_1_avg_pool,1, 2)\n        \n        x_seq_1 = F.elu(self.Linear_2(torch.cat((x_seq_1,x_seq_1_mean,x_seq,x_seq_1_avg_pool),dim=-1)/5))\n        \n        x_seq_out = self.Linear_3(x_seq_1)\n        x_seq_out = torch.transpose(x_seq_out, 1, 2)\n        x_seq_out = x_seq_out.reshape((-1,60*len(seq_y_list)))\n        \n        x_num_out = F.elu(self.Linear_4_0(torch.mean(x_seq_1,dim=1)))\n        x_num_out = self.Linear_4(x_num_out)\n\n        output = self.weight.weight*(torch.concat((x_seq_out,x_num_out),dim=-1))/3+self.bias.weight/3\n        \n        output[:,zeroout_index] =  output[:,zeroout_index]*0.0\n        \n        return output\n```\n\n## Training \n\n### Low-Res-Aqua-Mixed\n\nTypically, our batch sizes were 4-5 files or 4000-5000 data points. We found that training using a Huber loss function with delta = 1 noticeably improves the model's performance. A single model is trained on the mixed low-res and aqua-planet data to completion. Then, the same model is fine-tuned using slightly different procedures (varying SWA and checkpoint averaging parameters) to create distinct models that ensemble effectively. Training to completion takes approximately 24 hours on a single 4090 GPU.\n\n### Low-Res-High-Res-Mixed\n\nFor each step during training, a batch of low-res data and a batch of high-res data are evaluated using the model. The loss function for the step is a weighted combination of the loss on the low-res data and the high-res data. The loss functions used are a combination of MSE and MAE loss functions. About 72% of the total loss weight is placed on the MAE of the low-res data, 6% on the MSE of the low-res data, 22% on the MAE of the high-res data, and <1% on the MSE of the high-res data. The best-performing models were equipped with EMA and trained with a batch size of about 7000 data points. Training takes about 12 hours when trained locally on a computer with 6 RTX 4090 GPUs.\n\n## Ensemble Inferencing\n\nThe best individual model was a Low-Res-High-Res-Mixed model that achieved a public LB score of 0.78553. The best Low-Res-Aqua-Mixed model was not too different, with an LB public score of 0.78457.\n\nThe optimal weighting for the final ensemble was determined experimentally building on the intuition that the Low-Res-High-Res-Mixed model performed slightly better, and therefore received slightly larger weights. Predictions from models and ensembles from various stages of development were included in the final ensemble. Prior to submitting the final ensemble, an additional postprocessing step was applied to variables q0001, q0002, and q0003 whereby the maximum between the predicted ptend value and -state/1200 is chosen. The final ensemble had a public LB score of 0.79071.",
      "votes": 32
    },
    {
      "id": 2940744,
      "postDate": "2024-07-30T12:11:28.830Z",
      "content": "<p>Congratulations, <a href=\"https://www.kaggle.com/marecserlin\" target=\"_blank\">@marecserlin</a>, on the impressive 5th place finish with your ensemble of Bidirectional LSTM models in the LEAP ClimSim competition! Your approach of leveraging a combination of low-resolution and high-resolution datasets, and engineering features, is highly commendable. The use of Bidirectional LSTMs and GRUs, along with effective preprocessing and model ensembling, clearly showcases the potential of deep learning in atmospheric simulations.</p>\n<p>Your detailed methodology, from preprocessing strategies to model architecture and training techniques, offers valuable insights for others in the field. The results highlight the effectiveness of combining various LSTM configurations and engineered features to enhance predictive performance. It’s also notable how you fine-tuned the ensemble and applied postprocessing to achieve a public LB score of 0.79071.</p>\n<p>Thank you for sharing this detailed explanation and code. It provides a clear understanding of how to tackle complex climate simulation problems with AI.</p>",
      "rawMarkdown": "Congratulations, @marecserlin, on the impressive 5th place finish with your ensemble of Bidirectional LSTM models in the LEAP ClimSim competition! Your approach of leveraging a combination of low-resolution and high-resolution datasets, and engineering features, is highly commendable. The use of Bidirectional LSTMs and GRUs, along with effective preprocessing and model ensembling, clearly showcases the potential of deep learning in atmospheric simulations.\n\nYour detailed methodology, from preprocessing strategies to model architecture and training techniques, offers valuable insights for others in the field. The results highlight the effectiveness of combining various LSTM configurations and engineered features to enhance predictive performance. It’s also notable how you fine-tuned the ensemble and applied postprocessing to achieve a public LB score of 0.79071.\n\nThank you for sharing this detailed explanation and code. It provides a clear understanding of how to tackle complex climate simulation problems with AI."
    }
  ],
  "comments": [
    {
      "id": 2940744,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-07-30T12:11:28.830000",
      "content": "<p>Congratulations, <a href=\"https://www.kaggle.com/marecserlin\" target=\"_blank\">@marecserlin</a>, on the impressive 5th place finish with your ensemble of Bidirectional LSTM models in the LEAP ClimSim competition! Your approach of leveraging a combination of low-resolution and high-resolution datasets, and engineering features, is highly commendable. The use of Bidirectional LSTMs and GRUs, along with effective preprocessing and model ensembling, clearly showcases the potential of deep learning in atmospheric simulations.</p>\n<p>Your detailed methodology, from preprocessing strategies to model architecture and training techniques, offers valuable insights for others in the field. The results highlight the effectiveness of combining various LSTM configurations and engineered features to enhance predictive performance. It’s also notable how you fine-tuned the ensemble and applied postprocessing to achieve a public LB score of 0.79071.</p>\n<p>Thank you for sharing this detailed explanation and code. It provides a clear understanding of how to tackle complex climate simulation problems with AI.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2940257": "# Github Code\n\nhttps://github.com/YusefAN/leap-climsim-kaggle-5th\n\n# Explanation\n\nWe achieved our results by training an ensemble of models, each with its own unique variations. The core of each model architecture is a bidirectional LSTM, with different surrounding layers and engineered features. These variations, detailed below, enabled our models to ensemble effectively, leading to our competitive score and 5th-place finish.\n\nOverall, our model architectures remain very simple and manageable in size (each having 20 million parameters or so), enabling us to efficiently train on as much data as possible. Some models were trained on the [ClimSim_low-res](https://huggingface.co/datasets/LEAP/ClimSim_low-res) dataset combined with the [ClimSim_low-res_aqua-planet](https://huggingface.co/datasets/LEAP/ClimSim_low-res_aqua-planet) dataset, while others used the [ClimSim_low-res](https://huggingface.co/datasets/LEAP/ClimSim_low-res) dataset combined with a subset of [ClimSim_high-res](https://huggingface.co/datasets/LEAP/ClimSim_high-res) data. Our training time for even our hungriest models did not exceed 24 hours. \n\n## Preprocessing\n\n### Low-Res-Aqua-Mixed\n\nThe 556 element inputs were separated into column inputs (60x9 = 540) and global inputs (1x16). The global inputs were then repeated 60 times and concatenated to the columns to form a (60x25) input to our models. Engineered features, described in the individual model sections, are added as additional columns. All inputs and outputs, including engineered features, were standardized using their mean and standard deviation without any additional transformations.\n\n### Low-Res-High-Res-Mixed \n\nAs for the previous model, the 556 element inputs are separated into column inputs (60x9 = 540) and global inputs (1x16). The global inputs are repeated 60 times and concatenated to the columns to form a (60x25) input. These models did not use any engineered features. The inputs 'state-\\_q0001', 'state\\_q0002', 'state\\_q0003', 'pbuf\\_ozone', 'pbuf\\_CH4', 'pbuf\\_N2O' were log-transformed. Then, all the inputs (both transformed and not) were standardized using their mean and standard deviation. The outputs are predicted directly without any normalization. The models are trained using all the low-res data and about 1/15th of the high-res data.\n\n## Models\n\n### Low-Res-Aqua-Mixed\n\nThe models trained on the combination of the low-res and the aqua-planet datasets had the following simple architecture:\n- MLP encoder-decoder on the input, outputting a 60x25 matrix which is concatenated to the input\n- The concatenated input is fed into a wide (hidden dimension of 512) but shallow (3 layers deep) bidirectional LSTM \n- The output is fed into a single bidirectional GRU layer\n- Final MLP encoder to produce the 368-element output sequence\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20298348%2Fc07d755c212f9de991b58c53a37c2a1a%2FLowResAquaArchitecture.jpg?generation=1722292377019977&alt=media)\n\nSeveral models in our final ensemble followed this procedure without additional feature engineering. Others included the following features:\n1. liq_partition STATE_Q0002 / (STATE_Q0002 + STATE_Q0003)\n2. imbalance (STATE_Q0002 - STATE_Q0003) / (STATE_Q0002_IDX + STATE_Q0003)\n3. moisture = (STATE_Q0001) * (STATE_U)\\*\\*2 + (STATE_V)\\*\\*2)\n4. air_total = PBUF_OZONE + PBUF_CH4 + PBUF_N2O\n5. temp_humid = STATE_T / STATE_Q0001\n6. temp_diff1 = STATE_T_{i} - STATE_T_IDX_{i+1}\n7. temp_diff2 = STATE_Q0001_{i} - STATE_Q0001_IDX_{i+1}\n8. wind_diff1 = STATE_U_{i} - STATE_U_{i+1}\n9. wind_diff2 = STATE_V_{i} - STATE_V_{i+1}\n\nModel code: \n\n```python\nclass LeapModel(nn.Module):\n    def __init__(self, input_dim, hidden_dim, output_dim, num_layers=2):\n        super().__init__()\n        # self.N_SEQ_OUTPUTS = N_SEQ_OUTPUTS\n        self.encoder = MLP([input_dim * 60, input_dim * 15, input_dim * 8, input_dim * 4])\n        self.decoder = MLP([input_dim * 4, input_dim * 8, input_dim * 15, input_dim * 60])\n\n        self.lstm = nn.LSTM(input_dim * 2, hidden_dim, num_layers, batch_first=True, dropout=0.09, bidirectional=True)\n        \n        self.lstm2 = nn.GRU(hidden_dim * 2, 16, batch_first=True, dropout=0.00, bidirectional=True)\n\n        # self.fc_lstm = nn.Linear((32) * 60 , output_dim)\n        self.fc_lstm = MLP([32 * 60, 16 * 60, output_dim])#nn.Linear((32) * 60 , output_dim)\n        \n        self.criterion = nn.HuberLoss(delta=1) ########################################\n\n    def forward(self, x):\n        x_0 = self.encoder(x.view(x.size(0), -1))\n        x_1 = self.decoder(x_0)\n        x_1 = x_1.view(x.size(0), 60, -1)\n        \n        lstm_out, _ = self.lstm(torch.cat([x, x_1], dim = -1))\n        # lstm_out, _ = self.gru(lstm_out)\n        lstm_out, _ = self.lstm2(lstm_out)\n        lstm_out = lstm_out.contiguous().view(x.size(0), -1)  # Flatten the LSTM output\n        output = self.fc_lstm(lstm_out)\n        return output\n```\n\n### Low-Res-High-Res-Mixed\n\nThe 3 best-performing models trained on the combination of the low-res and the high-res datasets were bidirectional LSTMs with outfitted with various FFNNs. They all shared the following features:\n\n* Linear layer expanding the input dimension\n* Bidirectional LSTM (6 layers deep) with hidden dimension ranging from 256 to 320\n* The LSTM output is put into a 1D average pooling layer and a mean layer (along the dimension with length 60, which is repeated 60 times)\n* The original input, the LSTM output, the 1D average pooling output, and the mean are all concatenated\n* A series of linear layers then produce the final 368-element output sequence\n\nFor 2 out of 3 of the models, an additional linear layer (not pictured in the diagram) with a softmax activation is used to compute an \"attention\" for computing the 8x1 flattened output.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20298348%2F5e6dc8f287158a7820bb3aee3b02fe5f%2FLowResHighResArchitecture.jpg?generation=1722292441625060&alt=media)\n\nModel code: \n\n```python\nclass FFNN_LSTM_6_AVG(nn.Module):\n    def __init__(self, input_size, output_size):\n        super(FFNN_LSTM_6_AVG, self).__init__()\n        \n        self.encode_dim = 300\n        self.hidden_dim = 280\n        self.iter_dim = 800\n\n        self.LSTM_1 = LSTM(self.encode_dim,self.hidden_dim,6,batch_first=True,dropout=0.01,bidirectional=True)\n        self.input_size = input_size\n        self.Linear_1 = nn.Linear(len(seq_fea_list)+len(num_fea_list), self.encode_dim)\n        self.Linear_2 = nn.Linear(6*self.hidden_dim+self.encode_dim, self.iter_dim)\n        self.Linear_3 = nn.Linear(self.iter_dim, len(seq_y_list))\n        self.Linear_4_0 = nn.Linear(self.iter_dim, self.iter_dim*2)\n\n        self.Linear_4 = nn.Linear(self.iter_dim*2, len(num_y_list))\n        self.bias = nn.Linear(len(seq_y_list)*60+len(num_y_list),1)\n        self.weight = nn.Linear(len(seq_y_list)*60+len(num_y_list),1)\n        self.avg_pool_1 = AvgPool1d(kernel_size=3,stride=1,padding=1)\n        \n    def forward(self, x):\n        x_seq = x[:,0:60*len(seq_fea_list)]\n        x_seq = x_seq.reshape((-1,len(seq_fea_list),60))\n        x_seq = torch.transpose(x_seq, 1, 2)\n        \n        x_num = x[:,60*len(seq_fea_list):x.shape[1]]\n        x_num_repeat = x_num.reshape((-1,1,len(num_fea_list)))\n        x_num_repeat = x_num_repeat.repeat((1,60,1))\n        \n        x_seq = F.elu(self.Linear_1(torch.concat((x_seq,x_num_repeat),dim=-1)/5))\n        \n        x_seq_1,_ = self.LSTM_1(x_seq/5)\n        \n        x_seq_1_mean = torch.mean(x_seq_1,dim=1,keepdim=True)\n        x_seq_1_mean = x_seq_1_mean.repeat((1,60,1))\n\n        x_seq_1_avg_pool = self.avg_pool_1(torch.transpose(x_seq_1, 1, 2))\n        x_seq_1_avg_pool = torch.transpose(x_seq_1_avg_pool,1, 2)\n        \n        x_seq_1 = F.elu(self.Linear_2(torch.cat((x_seq_1,x_seq_1_mean,x_seq,x_seq_1_avg_pool),dim=-1)/5))\n        \n        x_seq_out = self.Linear_3(x_seq_1)\n        x_seq_out = torch.transpose(x_seq_out, 1, 2)\n        x_seq_out = x_seq_out.reshape((-1,60*len(seq_y_list)))\n        \n        x_num_out = F.elu(self.Linear_4_0(torch.mean(x_seq_1,dim=1)))\n        x_num_out = self.Linear_4(x_num_out)\n\n        output = self.weight.weight*(torch.concat((x_seq_out,x_num_out),dim=-1))/3+self.bias.weight/3\n        \n        output[:,zeroout_index] =  output[:,zeroout_index]*0.0\n        \n        return output\n```\n\n## Training \n\n### Low-Res-Aqua-Mixed\n\nTypically, our batch sizes were 4-5 files or 4000-5000 data points. We found that training using a Huber loss function with delta = 1 noticeably improves the model's performance. A single model is trained on the mixed low-res and aqua-planet data to completion. Then, the same model is fine-tuned using slightly different procedures (varying SWA and checkpoint averaging parameters) to create distinct models that ensemble effectively. Training to completion takes approximately 24 hours on a single 4090 GPU.\n\n### Low-Res-High-Res-Mixed\n\nFor each step during training, a batch of low-res data and a batch of high-res data are evaluated using the model. The loss function for the step is a weighted combination of the loss on the low-res data and the high-res data. The loss functions used are a combination of MSE and MAE loss functions. About 72% of the total loss weight is placed on the MAE of the low-res data, 6% on the MSE of the low-res data, 22% on the MAE of the high-res data, and <1% on the MSE of the high-res data. The best-performing models were equipped with EMA and trained with a batch size of about 7000 data points. Training takes about 12 hours when trained locally on a computer with 6 RTX 4090 GPUs.\n\n## Ensemble Inferencing\n\nThe best individual model was a Low-Res-High-Res-Mixed model that achieved a public LB score of 0.78553. The best Low-Res-Aqua-Mixed model was not too different, with an LB public score of 0.78457.\n\nThe optimal weighting for the final ensemble was determined experimentally building on the intuition that the Low-Res-High-Res-Mixed model performed slightly better, and therefore received slightly larger weights. Predictions from models and ensembles from various stages of development were included in the final ensemble. Prior to submitting the final ensemble, an additional postprocessing step was applied to variables q0001, q0002, and q0003 whereby the maximum between the predicted ptend value and -state/1200 is chosen. The final ensemble had a public LB score of 0.79071.",
    "2940744": "Congratulations, @marecserlin, on the impressive 5th place finish with your ensemble of Bidirectional LSTM models in the LEAP ClimSim competition! Your approach of leveraging a combination of low-resolution and high-resolution datasets, and engineering features, is highly commendable. The use of Bidirectional LSTMs and GRUs, along with effective preprocessing and model ensembling, clearly showcases the potential of deep learning in atmospheric simulations.\n\nYour detailed methodology, from preprocessing strategies to model architecture and training techniques, offers valuable insights for others in the field. The results highlight the effectiveness of combining various LSTM configurations and engineered features to enhance predictive performance. It’s also notable how you fine-tuned the ensemble and applied postprocessing to achieve a public LB score of 0.79071.\n\nThank you for sharing this detailed explanation and code. It provides a clear understanding of how to tackle complex climate simulation problems with AI."
  }
}