{
  "id": 523077,
  "title": "3rd place solution",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/523077",
  "author_name": "pao",
  "post_date": "2024-07-30T05:31:40.672000",
  "votes": 29,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Everyone, thank you for your hard work on the competition.</p>\n<p><a href=\"https://www.kaggle.com/jerrylin\" target=\"_blank\">@jerrylin</a>  Thank you for organizing the competition. I know it must have been tough, but I believe we were able to get this far thanks to your sincere dedication up to the final check.<br>\nCongratulations to everyone who ranked high and get good results. It was enjoyable to compete alongside you, and you all served as great motivation.<br>\nTo my teammates <a href=\"https://www.kaggle.com/bamps53\" target=\"_blank\">@bamps53</a>  <a href=\"https://www.kaggle.com/kmat2019\" target=\"_blank\">@kmat2019</a> , it was fun and I learned a lot. Thank you very much.<br>\nAlso I am happy to become Grandmaster at this competition. Also <a href=\"https://www.kaggle.com/kmat2019\" target=\"_blank\">@kmat2019</a> became grandmaster. Congrats!!</p>\n<h2>Model Summary</h2>\n<h3>Overall Summary</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1305588%2Fa0b89d5c5ca6aca62baa4cfb090ac7bb%2Fimage%20(2).png?generation=1722316979148053&amp;alt=media\" alt=\"Overall pipeline\"></p>\n<p>Each team member built their own neural network models. </p>\n<p>After obtaining the Camaro model's predictions, we created several features to input into GBDT regressors. These regressors refined the model's predictions. Although this second stage could be applied to other models, we only applied it to the Camaro model because the Pao and Kmat models lacked a validation dataset.</p>\n<p>The final prediction was calculated as a weighted average of these predictions.</p>\n<h3>Pao Part</h3>\n<h4>Overview</h4>\n<ul>\n<li><strong>Model:</strong> 1d CNN + Transformer + LSTM</li>\n<li><strong>Feature:</strong> Original and relative humidity with those sequence 1st derivative and 2nd derivative at heights (diff and diff-diff)</li>\n<li><strong>Auxiliary Loss:</strong> Predicting the difference between adjacent vertical levels</li>\n<li><strong>Row-less full training</strong></li>\n</ul>\n<h4>Model Architecture</h4>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1305588%2F3142f36d09f2e57ddcdcb023222eaba6%2F2024-07-30%2014.24.46.png?generation=1722317122204086&amp;alt=media\" alt=\"Pao model\"></p>\n<ul>\n<li><strong>Input layer:</strong> Linear transformation with bias at each feature<ul>\n<li>Example: <code>feature1 = feature1_original * a1 + b1</code> (a1 and b1 are trainable parameters)</li></ul></li>\n<li><strong>Feature Concatenation:</strong> Concatenate scalar features to each height sequence feature after each dense layer<ul>\n<li>Example: <code>features_level0 = concat([seq_features_level0, dense_level0(scalar_features)])</code></li></ul></li>\n<li><strong>Positional Encoding:</strong> Adding embedding per height to the hidden layer</li>\n<li><strong>Model Blocks:</strong><ul>\n<li><strong>Residual block:</strong> 1dCNN (Conv1d + BN + GELU) * 2 + Transformer</li>\n<li>Conv1d: kernel size = 5</li>\n<li>Transformer: n_head = 8, n_layers = 1</li>\n<li><strong>LSTM:</strong> Bidirectional 2 layers</li>\n<li><strong>MLP:</strong> Simple MLP with (Linear and GELU with no dropout)</li></ul></li>\n</ul>\n<h4>Training</h4>\n<ul>\n<li><strong>Auxiliary Loss:</strong> Predicting the difference between adjacent vertical levels (similar to Camaro part)</li>\n<li><strong>Loss:</strong> Huber Loss using EMA</li>\n<li><strong>Learning rate:</strong> Cosine annealing 1e-3 to 1e-5, 10 epochs</li>\n<li><strong>Optimizer:</strong> AdamW</li>\n</ul>\n<h4>Others</h4>\n<ul>\n<li><strong>Dataset:</strong> WebDataset for LowRes full-training</li>\n<li><strong>Normalization:</strong> <ul>\n<li><strong>Feature:</strong> <code>(input - mean(input)) / std(input)</code></li>\n<li><strong>Target:</strong> Multiply old_sample_submission weight</li></ul></li>\n<li><strong>Postprocess:</strong> Replace predictions for <code>ptend_q0002_0-27</code> with <code>-1 * input / 1200</code></li>\n<li><strong>Ensemble:</strong> 3 models with variations in dropout, hidden size, and Huber Loss delta</li>\n</ul>\n<h3>Camaro Part</h3>\n<h4>Overview</h4>\n<ul>\n<li><strong>Full Dataset and Long Training:</strong> Training on the full dataset</li>\n<li><strong>Model Architecture:</strong> Combination of CNN and Transformer or Transformer-only using the CLIP Encoder</li>\n<li><strong>Auxiliary Loss:</strong> Predicting the difference between adjacent vertical levels</li>\n</ul>\n<h4>Dataset and Preprocessing</h4>\n<ul>\n<li><strong>HuggingFace Dataset:</strong> Full dataset for training</li>\n<li><strong>Feature Engineering:</strong> Added saturation vapor pressure as a feature</li>\n<li><strong>Normalization:</strong> Pre-computed using Kaggle train and test datasets</li>\n<li><strong>WebDataset:</strong> Efficient loading for faster training</li>\n</ul>\n<h4>Model</h4>\n<ul>\n<li><strong>Architecture:</strong> Combination of CNN and Transformer or Transformer-only with heavy embedding and head layers</li>\n</ul>\n<h4>Training</h4>\n<ul>\n<li><strong>Target Transformation:</strong> Dividing by old sample submission weights and subtracting the mean</li>\n<li><strong>Auxiliary Loss:</strong> Incorporating information about derivatives and second derivatives</li>\n<li><strong>Loss Function:</strong> Huber loss with a delta of 2.0</li>\n</ul>\n<h4>Post-Processing</h4>\n<ul>\n<li>Replace predictions for <code>ptend_q0002_0-26</code> with <code>-1 * input / 1200</code></li>\n</ul>\n<h4>Results</h4>\n<table>\n<thead>\n<tr>\n<th>exp</th>\n<th>bs</th>\n<th>arch</th>\n<th>dim</th>\n<th>loss</th>\n<th>Public LB</th>\n<th>Private LB</th>\n<th>Public LB (2nd stage)</th>\n<th>Private LB (2nd stage)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>256</td>\n<td>ConvTransformer</td>\n<td>256</td>\n<td>Huber2</td>\n<td>0.78468</td>\n<td>0.78176</td>\n<td>0.78504</td>\n<td>0.78199</td>\n</tr>\n<tr>\n<td>2</td>\n<td>1024</td>\n<td>ConvTransformer</td>\n<td>384</td>\n<td>Huber2</td>\n<td>0.78418</td>\n<td>0.78131</td>\n<td>0.78451</td>\n<td>0.78159</td>\n</tr>\n<tr>\n<td>3</td>\n<td>512</td>\n<td>Transformer n_layer=8</td>\n<td>256</td>\n<td>Huber2</td>\n<td>0.78545</td>\n<td>0.78154</td>\n<td>0.78572</td>\n<td>0.78154</td>\n</tr>\n<tr>\n<td>4</td>\n<td>768</td>\n<td>ConvTransformer x 2</td>\n<td>256</td>\n<td>Huber4</td>\n<td>0.78395</td>\n<td>0.78122</td>\n<td>0.78427</td>\n<td>0.78141</td>\n</tr>\n<tr>\n<td>5</td>\n<td>768</td>\n<td>Transformer n_layer=6</td>\n<td>256</td>\n<td>Huber8</td>\n<td>0.78245</td>\n<td>0.77951</td>\n<td>0.78255</td>\n<td>0.77957</td>\n</tr>\n<tr>\n<td><strong>Ensemble (1+2+3+4)</strong></td>\n<td></td>\n<td></td>\n<td></td>\n<td></td>\n<td></td>\n<td></td>\n<td>0.79025</td>\n<td>0.78694</td>\n</tr>\n<tr>\n<td><strong>Ensemble (1+2+3+4+5)</strong></td>\n<td></td>\n<td></td>\n<td></td>\n<td></td>\n<td></td>\n<td></td>\n<td>0.78998</td>\n<td>0.78685</td>\n</tr>\n</tbody>\n</table>\n<h3>Kmat Part</h3>\n<h4>Overview</h4>\n<p>As shown in Fig.6-1, Kmat part consists of:</p>\n<ul>\n<li>Add Features</li>\n<li>Normalize by averages and standards</li>\n<li>1D CNN model to predict climate</li>\n<li>Postprocess (some predictions are replaced by <code>-input/1200</code>)</li>\n</ul>\n<h4>Feature Engineering</h4>\n<ul>\n<li>Diff features from 1D data: <code>x[z] - x[z-1]</code></li>\n<li>Relative humidity-related features such as dew_point and vapor_pressure / saturation_pressure</li>\n</ul>\n<h4>Normalization</h4>\n<ul>\n<li><strong>Inputs:</strong><ol>\n<li><code>(x - x_mean(axis=0)) / x_std(axis=0)</code></li>\n<li><code>(x - x_mean(axis=(0,1))) / x_std(axis=(0,1))</code></li>\n<li><code>(log_x - log_x_mean(axis=(0,1))) / log_x_std(axis=(0,1))</code></li></ol></li>\n<li><strong>Targets:</strong><ul>\n<li><code>(x - x_mean(axis=0)) / x_std(axis=0)</code></li></ul></li>\n</ul>\n<h4>Model Architecture</h4>\n<ul>\n<li><strong>Core Architecture:</strong> FiLM 1D UNet<ul>\n<li>Scalar features processed by fully connected layers</li>\n<li>1D features processed by 1D FiLM Convolution layers</li>\n<li>Initial and final convolutions divided into multiple branches</li>\n<li>Three head branches for temperature, q000X, and wind vector prediction</li>\n<li>Classification branch for state_q drops</li></ul></li>\n<li><strong>Loss Function:</strong> Huber loss (beta=2)</li>\n</ul>\n<h4>Training</h4>\n<ul>\n<li><strong>Optimizer:</strong> Adam with clip norm</li>\n<li><strong>Scheduler:</strong> Cosine scheduler for 7 epochs</li>\n<li><strong>Batch Size:</strong> 384 with lr 0.0012</li>\n<li><strong>Training Time:</strong> 3 days on RTX 3090 for the entire LowRes dataset</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1305588%2F5930445971412b47277c800212ccef86%2F2024-07-30%2014.25.58.png?generation=1722317176416622&amp;alt=media\" alt=\"Kmat pipeline\"></p>\n<h3>2nd Stage Modeling</h3>\n<p>The final submission from our team capaomat (3rd place) is an ensemble of predictions from three members. Some Camaro predictions (ptend_q0001, q0002, q0003) are refined by the 2nd stage. The score improvement is less than 0.0004. Neural network modeling is much more dominant.</p>\n<h4>Features</h4>\n<p>We used a few features from raw inputs and predictions of 1st stage to prevent overfitting. State_t, state_q, ptend_q, future_state_q features and ratio of future_state_q2 to q3 are provided to the model.</p>\n<p><img src=\"https://arxiv.org/abs/2407.00124\" alt=\"Fraction of liquid cloud over total cloud as a function of temperature\"></p>\n<h4>Model / Training</h4>\n<p>We employed lightGBM to predict ptend_q at each level. Specifically, we trained the models and updated predictions for various levels. It took less than 20 minutes on CPU to train all 91 models.</p>\n<h3>Ensemble</h3>\n<p>As a final submission, we blended the following 9 models. Our solution achieved the following results:</p>\n<table>\n<thead>\n<tr>\n<th>exp</th>\n<th>Public LB</th>\n<th>Private LB</th>\n<th>Weight1</th>\n<th>Weight2</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Camaro1_v2</td>\n<td>0.78504</td>\n<td>0.78199</td>\n<td>3.0</td>\n<td>3.0</td>\n</tr>\n<tr>\n<td>Camaro2_v2</td>\n<td>0.78451</td>\n<td>0.78159</td>\n<td>3.0</td>\n<td>3.0</td>\n</tr>\n<tr>\n<td>Camaro3_v2</td>\n<td>0.78572</td>\n<td>0.78154</td>\n<td>3.0</td>\n<td>3.5</td>\n</tr>\n<tr>\n<td>Camaro4_v2</td>\n<td>0.78427</td>\n<td>0.78141</td>\n<td>3.0</td>\n<td>2.0</td>\n</tr>\n<tr>\n<td>Camaro5_v2</td>\n<td>0.78255</td>\n<td>0.77957</td>\n<td>3.0</td>\n<td>1.0</td>\n</tr>\n<tr>\n<td>Pao1</td>\n<td>0.78139</td>\n<td>0.77770</td>\n<td>1.0</td>\n<td>0.5</td>\n</tr>\n<tr>\n<td>Pao2</td>\n<td>0.78252</td>\n<td>0.77864</td>\n<td>3.0</td>\n<td>2.0</td>\n</tr>\n<tr>\n<td>Pao3</td>\n<td>0.77985</td>\n<td>0.77801</td>\n<td>1.0</td>\n<td>0.5</td>\n</tr>\n<tr>\n<td>Kmat1</td>\n<td>0.78120</td>\n<td>0.77647</td>\n<td>3.0</td>\n<td>1.5</td>\n</tr>\n<tr>\n<td>Ensemble with weight1 (private best)</td>\n<td>0.79026</td>\n<td>0.78810</td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>Ensemble with weight2 (final submission)</td>\n<td>0.79048</td>\n<td>0.78792</td>\n<td></td>\n<td></td>\n</tr>\n</tbody>\n</table>",
  "messages": [
    {
      "id": 2940455,
      "postDate": "2024-07-30T05:31:40.673Z",
      "content": "<p>Everyone, thank you for your hard work on the competition.</p>\n<p><a href=\"https://www.kaggle.com/jerrylin\" target=\"_blank\">@jerrylin</a>  Thank you for organizing the competition. I know it must have been tough, but I believe we were able to get this far thanks to your sincere dedication up to the final check.<br>\nCongratulations to everyone who ranked high and get good results. It was enjoyable to compete alongside you, and you all served as great motivation.<br>\nTo my teammates <a href=\"https://www.kaggle.com/bamps53\" target=\"_blank\">@bamps53</a>  <a href=\"https://www.kaggle.com/kmat2019\" target=\"_blank\">@kmat2019</a> , it was fun and I learned a lot. Thank you very much.<br>\nAlso I am happy to become Grandmaster at this competition. Also <a href=\"https://www.kaggle.com/kmat2019\" target=\"_blank\">@kmat2019</a> became grandmaster. Congrats!!</p>\n<h2>Model Summary</h2>\n<h3>Overall Summary</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1305588%2Fa0b89d5c5ca6aca62baa4cfb090ac7bb%2Fimage%20(2).png?generation=1722316979148053&amp;alt=media\" alt=\"Overall pipeline\"></p>\n<p>Each team member built their own neural network models. </p>\n<p>After obtaining the Camaro model's predictions, we created several features to input into GBDT regressors. These regressors refined the model's predictions. Although this second stage could be applied to other models, we only applied it to the Camaro model because the Pao and Kmat models lacked a validation dataset.</p>\n<p>The final prediction was calculated as a weighted average of these predictions.</p>\n<h3>Pao Part</h3>\n<h4>Overview</h4>\n<ul>\n<li><strong>Model:</strong> 1d CNN + Transformer + LSTM</li>\n<li><strong>Feature:</strong> Original and relative humidity with those sequence 1st derivative and 2nd derivative at heights (diff and diff-diff)</li>\n<li><strong>Auxiliary Loss:</strong> Predicting the difference between adjacent vertical levels</li>\n<li><strong>Row-less full training</strong></li>\n</ul>\n<h4>Model Architecture</h4>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1305588%2F3142f36d09f2e57ddcdcb023222eaba6%2F2024-07-30%2014.24.46.png?generation=1722317122204086&amp;alt=media\" alt=\"Pao model\"></p>\n<ul>\n<li><strong>Input layer:</strong> Linear transformation with bias at each feature<ul>\n<li>Example: <code>feature1 = feature1_original * a1 + b1</code> (a1 and b1 are trainable parameters)</li></ul></li>\n<li><strong>Feature Concatenation:</strong> Concatenate scalar features to each height sequence feature after each dense layer<ul>\n<li>Example: <code>features_level0 = concat([seq_features_level0, dense_level0(scalar_features)])</code></li></ul></li>\n<li><strong>Positional Encoding:</strong> Adding embedding per height to the hidden layer</li>\n<li><strong>Model Blocks:</strong><ul>\n<li><strong>Residual block:</strong> 1dCNN (Conv1d + BN + GELU) * 2 + Transformer</li>\n<li>Conv1d: kernel size = 5</li>\n<li>Transformer: n_head = 8, n_layers = 1</li>\n<li><strong>LSTM:</strong> Bidirectional 2 layers</li>\n<li><strong>MLP:</strong> Simple MLP with (Linear and GELU with no dropout)</li></ul></li>\n</ul>\n<h4>Training</h4>\n<ul>\n<li><strong>Auxiliary Loss:</strong> Predicting the difference between adjacent vertical levels (similar to Camaro part)</li>\n<li><strong>Loss:</strong> Huber Loss using EMA</li>\n<li><strong>Learning rate:</strong> Cosine annealing 1e-3 to 1e-5, 10 epochs</li>\n<li><strong>Optimizer:</strong> AdamW</li>\n</ul>\n<h4>Others</h4>\n<ul>\n<li><strong>Dataset:</strong> WebDataset for LowRes full-training</li>\n<li><strong>Normalization:</strong> <ul>\n<li><strong>Feature:</strong> <code>(input - mean(input)) / std(input)</code></li>\n<li><strong>Target:</strong> Multiply old_sample_submission weight</li></ul></li>\n<li><strong>Postprocess:</strong> Replace predictions for <code>ptend_q0002_0-27</code> with <code>-1 * input / 1200</code></li>\n<li><strong>Ensemble:</strong> 3 models with variations in dropout, hidden size, and Huber Loss delta</li>\n</ul>\n<h3>Camaro Part</h3>\n<h4>Overview</h4>\n<ul>\n<li><strong>Full Dataset and Long Training:</strong> Training on the full dataset</li>\n<li><strong>Model Architecture:</strong> Combination of CNN and Transformer or Transformer-only using the CLIP Encoder</li>\n<li><strong>Auxiliary Loss:</strong> Predicting the difference between adjacent vertical levels</li>\n</ul>\n<h4>Dataset and Preprocessing</h4>\n<ul>\n<li><strong>HuggingFace Dataset:</strong> Full dataset for training</li>\n<li><strong>Feature Engineering:</strong> Added saturation vapor pressure as a feature</li>\n<li><strong>Normalization:</strong> Pre-computed using Kaggle train and test datasets</li>\n<li><strong>WebDataset:</strong> Efficient loading for faster training</li>\n</ul>\n<h4>Model</h4>\n<ul>\n<li><strong>Architecture:</strong> Combination of CNN and Transformer or Transformer-only with heavy embedding and head layers</li>\n</ul>\n<h4>Training</h4>\n<ul>\n<li><strong>Target Transformation:</strong> Dividing by old sample submission weights and subtracting the mean</li>\n<li><strong>Auxiliary Loss:</strong> Incorporating information about derivatives and second derivatives</li>\n<li><strong>Loss Function:</strong> Huber loss with a delta of 2.0</li>\n</ul>\n<h4>Post-Processing</h4>\n<ul>\n<li>Replace predictions for <code>ptend_q0002_0-26</code> with <code>-1 * input / 1200</code></li>\n</ul>\n<h4>Results</h4>\n<table>\n<thead>\n<tr>\n<th>exp</th>\n<th>bs</th>\n<th>arch</th>\n<th>dim</th>\n<th>loss</th>\n<th>Public LB</th>\n<th>Private LB</th>\n<th>Public LB (2nd stage)</th>\n<th>Private LB (2nd stage)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>256</td>\n<td>ConvTransformer</td>\n<td>256</td>\n<td>Huber2</td>\n<td>0.78468</td>\n<td>0.78176</td>\n<td>0.78504</td>\n<td>0.78199</td>\n</tr>\n<tr>\n<td>2</td>\n<td>1024</td>\n<td>ConvTransformer</td>\n<td>384</td>\n<td>Huber2</td>\n<td>0.78418</td>\n<td>0.78131</td>\n<td>0.78451</td>\n<td>0.78159</td>\n</tr>\n<tr>\n<td>3</td>\n<td>512</td>\n<td>Transformer n_layer=8</td>\n<td>256</td>\n<td>Huber2</td>\n<td>0.78545</td>\n<td>0.78154</td>\n<td>0.78572</td>\n<td>0.78154</td>\n</tr>\n<tr>\n<td>4</td>\n<td>768</td>\n<td>ConvTransformer x 2</td>\n<td>256</td>\n<td>Huber4</td>\n<td>0.78395</td>\n<td>0.78122</td>\n<td>0.78427</td>\n<td>0.78141</td>\n</tr>\n<tr>\n<td>5</td>\n<td>768</td>\n<td>Transformer n_layer=6</td>\n<td>256</td>\n<td>Huber8</td>\n<td>0.78245</td>\n<td>0.77951</td>\n<td>0.78255</td>\n<td>0.77957</td>\n</tr>\n<tr>\n<td><strong>Ensemble (1+2+3+4)</strong></td>\n<td></td>\n<td></td>\n<td></td>\n<td></td>\n<td></td>\n<td></td>\n<td>0.79025</td>\n<td>0.78694</td>\n</tr>\n<tr>\n<td><strong>Ensemble (1+2+3+4+5)</strong></td>\n<td></td>\n<td></td>\n<td></td>\n<td></td>\n<td></td>\n<td></td>\n<td>0.78998</td>\n<td>0.78685</td>\n</tr>\n</tbody>\n</table>\n<h3>Kmat Part</h3>\n<h4>Overview</h4>\n<p>As shown in Fig.6-1, Kmat part consists of:</p>\n<ul>\n<li>Add Features</li>\n<li>Normalize by averages and standards</li>\n<li>1D CNN model to predict climate</li>\n<li>Postprocess (some predictions are replaced by <code>-input/1200</code>)</li>\n</ul>\n<h4>Feature Engineering</h4>\n<ul>\n<li>Diff features from 1D data: <code>x[z] - x[z-1]</code></li>\n<li>Relative humidity-related features such as dew_point and vapor_pressure / saturation_pressure</li>\n</ul>\n<h4>Normalization</h4>\n<ul>\n<li><strong>Inputs:</strong><ol>\n<li><code>(x - x_mean(axis=0)) / x_std(axis=0)</code></li>\n<li><code>(x - x_mean(axis=(0,1))) / x_std(axis=(0,1))</code></li>\n<li><code>(log_x - log_x_mean(axis=(0,1))) / log_x_std(axis=(0,1))</code></li></ol></li>\n<li><strong>Targets:</strong><ul>\n<li><code>(x - x_mean(axis=0)) / x_std(axis=0)</code></li></ul></li>\n</ul>\n<h4>Model Architecture</h4>\n<ul>\n<li><strong>Core Architecture:</strong> FiLM 1D UNet<ul>\n<li>Scalar features processed by fully connected layers</li>\n<li>1D features processed by 1D FiLM Convolution layers</li>\n<li>Initial and final convolutions divided into multiple branches</li>\n<li>Three head branches for temperature, q000X, and wind vector prediction</li>\n<li>Classification branch for state_q drops</li></ul></li>\n<li><strong>Loss Function:</strong> Huber loss (beta=2)</li>\n</ul>\n<h4>Training</h4>\n<ul>\n<li><strong>Optimizer:</strong> Adam with clip norm</li>\n<li><strong>Scheduler:</strong> Cosine scheduler for 7 epochs</li>\n<li><strong>Batch Size:</strong> 384 with lr 0.0012</li>\n<li><strong>Training Time:</strong> 3 days on RTX 3090 for the entire LowRes dataset</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1305588%2F5930445971412b47277c800212ccef86%2F2024-07-30%2014.25.58.png?generation=1722317176416622&amp;alt=media\" alt=\"Kmat pipeline\"></p>\n<h3>2nd Stage Modeling</h3>\n<p>The final submission from our team capaomat (3rd place) is an ensemble of predictions from three members. Some Camaro predictions (ptend_q0001, q0002, q0003) are refined by the 2nd stage. The score improvement is less than 0.0004. Neural network modeling is much more dominant.</p>\n<h4>Features</h4>\n<p>We used a few features from raw inputs and predictions of 1st stage to prevent overfitting. State_t, state_q, ptend_q, future_state_q features and ratio of future_state_q2 to q3 are provided to the model.</p>\n<p><img src=\"https://arxiv.org/abs/2407.00124\" alt=\"Fraction of liquid cloud over total cloud as a function of temperature\"></p>\n<h4>Model / Training</h4>\n<p>We employed lightGBM to predict ptend_q at each level. Specifically, we trained the models and updated predictions for various levels. It took less than 20 minutes on CPU to train all 91 models.</p>\n<h3>Ensemble</h3>\n<p>As a final submission, we blended the following 9 models. Our solution achieved the following results:</p>\n<table>\n<thead>\n<tr>\n<th>exp</th>\n<th>Public LB</th>\n<th>Private LB</th>\n<th>Weight1</th>\n<th>Weight2</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Camaro1_v2</td>\n<td>0.78504</td>\n<td>0.78199</td>\n<td>3.0</td>\n<td>3.0</td>\n</tr>\n<tr>\n<td>Camaro2_v2</td>\n<td>0.78451</td>\n<td>0.78159</td>\n<td>3.0</td>\n<td>3.0</td>\n</tr>\n<tr>\n<td>Camaro3_v2</td>\n<td>0.78572</td>\n<td>0.78154</td>\n<td>3.0</td>\n<td>3.5</td>\n</tr>\n<tr>\n<td>Camaro4_v2</td>\n<td>0.78427</td>\n<td>0.78141</td>\n<td>3.0</td>\n<td>2.0</td>\n</tr>\n<tr>\n<td>Camaro5_v2</td>\n<td>0.78255</td>\n<td>0.77957</td>\n<td>3.0</td>\n<td>1.0</td>\n</tr>\n<tr>\n<td>Pao1</td>\n<td>0.78139</td>\n<td>0.77770</td>\n<td>1.0</td>\n<td>0.5</td>\n</tr>\n<tr>\n<td>Pao2</td>\n<td>0.78252</td>\n<td>0.77864</td>\n<td>3.0</td>\n<td>2.0</td>\n</tr>\n<tr>\n<td>Pao3</td>\n<td>0.77985</td>\n<td>0.77801</td>\n<td>1.0</td>\n<td>0.5</td>\n</tr>\n<tr>\n<td>Kmat1</td>\n<td>0.78120</td>\n<td>0.77647</td>\n<td>3.0</td>\n<td>1.5</td>\n</tr>\n<tr>\n<td>Ensemble with weight1 (private best)</td>\n<td>0.79026</td>\n<td>0.78810</td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>Ensemble with weight2 (final submission)</td>\n<td>0.79048</td>\n<td>0.78792</td>\n<td></td>\n<td></td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "Everyone, thank you for your hard work on the competition.\n\n\n@jerrylin  Thank you for organizing the competition. I know it must have been tough, but I believe we were able to get this far thanks to your sincere dedication up to the final check.\nCongratulations to everyone who ranked high and get good results. It was enjoyable to compete alongside you, and you all served as great motivation.\nTo my teammates @bamps53  @kmat2019 , it was fun and I learned a lot. Thank you very much.\nAlso I am happy to become Grandmaster at this competition. Also @kmat2019 became grandmaster. Congrats!!\n\n\n## Model Summary\n### Overall Summary\n\n![Overall pipeline](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1305588%2Fa0b89d5c5ca6aca62baa4cfb090ac7bb%2Fimage%20(2).png?generation=1722316979148053&alt=media)\n\nEach team member built their own neural network models. \n\nAfter obtaining the Camaro model's predictions, we created several features to input into GBDT regressors. These regressors refined the model's predictions. Although this second stage could be applied to other models, we only applied it to the Camaro model because the Pao and Kmat models lacked a validation dataset.\n\nThe final prediction was calculated as a weighted average of these predictions.\n\n### Pao Part\n\n#### Overview\n- **Model:** 1d CNN + Transformer + LSTM\n- **Feature:** Original and relative humidity with those sequence 1st derivative and 2nd derivative at heights (diff and diff-diff)\n- **Auxiliary Loss:** Predicting the difference between adjacent vertical levels\n- **Row-less full training**\n\n#### Model Architecture\n\n![Pao model](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1305588%2F3142f36d09f2e57ddcdcb023222eaba6%2F2024-07-30%2014.24.46.png?generation=1722317122204086&alt=media)\n\n- **Input layer:** Linear transformation with bias at each feature\n  - Example: `feature1 = feature1_original * a1 + b1` (a1 and b1 are trainable parameters)\n- **Feature Concatenation:** Concatenate scalar features to each height sequence feature after each dense layer\n  - Example: `features_level0 = concat([seq_features_level0, dense_level0(scalar_features)])`\n- **Positional Encoding:** Adding embedding per height to the hidden layer\n- **Model Blocks:**\n  - **Residual block:** 1dCNN (Conv1d + BN + GELU) * 2 + Transformer\n    - Conv1d: kernel size = 5\n    - Transformer: n_head = 8, n_layers = 1\n  - **LSTM:** Bidirectional 2 layers\n  - **MLP:** Simple MLP with (Linear and GELU with no dropout)\n\n\n\n#### Training\n- **Auxiliary Loss:** Predicting the difference between adjacent vertical levels (similar to Camaro part)\n- **Loss:** Huber Loss using EMA\n- **Learning rate:** Cosine annealing 1e-3 to 1e-5, 10 epochs\n- **Optimizer:** AdamW\n\n#### Others\n- **Dataset:** WebDataset for LowRes full-training\n- **Normalization:** \n  - **Feature:** `(input - mean(input)) / std(input)`\n  - **Target:** Multiply old_sample_submission weight\n- **Postprocess:** Replace predictions for `ptend_q0002_0-27` with `-1 * input / 1200`\n- **Ensemble:** 3 models with variations in dropout, hidden size, and Huber Loss delta\n\n### Camaro Part\n\n#### Overview\n- **Full Dataset and Long Training:** Training on the full dataset\n- **Model Architecture:** Combination of CNN and Transformer or Transformer-only using the CLIP Encoder\n- **Auxiliary Loss:** Predicting the difference between adjacent vertical levels\n\n#### Dataset and Preprocessing\n- **HuggingFace Dataset:** Full dataset for training\n- **Feature Engineering:** Added saturation vapor pressure as a feature\n- **Normalization:** Pre-computed using Kaggle train and test datasets\n- **WebDataset:** Efficient loading for faster training\n\n#### Model\n- **Architecture:** Combination of CNN and Transformer or Transformer-only with heavy embedding and head layers\n\n#### Training\n- **Target Transformation:** Dividing by old sample submission weights and subtracting the mean\n- **Auxiliary Loss:** Incorporating information about derivatives and second derivatives\n- **Loss Function:** Huber loss with a delta of 2.0\n\n#### Post-Processing\n- Replace predictions for `ptend_q0002_0-26` with `-1 * input / 1200`\n\n#### Results\n| exp | bs  | arch               | dim | loss   | Public LB | Private LB | Public LB (2nd stage) | Private LB (2nd stage) |\n|-----|-----|--------------------|-----|--------|-----------|------------|-----------------------|------------------------|\n| 1   | 256 | ConvTransformer    | 256 | Huber2 | 0.78468   | 0.78176    | 0.78504               | 0.78199                |\n| 2   | 1024| ConvTransformer    | 384 | Huber2 | 0.78418   | 0.78131    | 0.78451               | 0.78159                |\n| 3   | 512 | Transformer n_layer=8 | 256 | Huber2 | 0.78545   | 0.78154    | 0.78572               | 0.78154                |\n| 4   | 768 | ConvTransformer x 2| 256 | Huber4 | 0.78395   | 0.78122    | 0.78427               | 0.78141                |\n| 5   | 768 | Transformer n_layer=6 | 256 | Huber8 | 0.78245   | 0.77951    | 0.78255               | 0.77957                |\n| **Ensemble (1+2+3+4)** |     |                    |     |        |           |            | 0.79025               | 0.78694                |\n| **Ensemble (1+2+3+4+5)** |     |                    |     |        |           |            | 0.78998               | 0.78685                | 0.79023                | 0.78703                |\n\n### Kmat Part\n\n#### Overview\nAs shown in Fig.6-1, Kmat part consists of:\n- Add Features\n- Normalize by averages and standards\n- 1D CNN model to predict climate\n- Postprocess (some predictions are replaced by `-input/1200`)\n\n#### Feature Engineering\n- Diff features from 1D data: `x[z] - x[z-1]`\n- Relative humidity-related features such as dew_point and vapor_pressure / saturation_pressure\n\n#### Normalization\n- **Inputs:**\n  1. `(x - x_mean(axis=0)) / x_std(axis=0)`\n  2. `(x - x_mean(axis=(0,1))) / x_std(axis=(0,1))`\n  3. `(log_x - log_x_mean(axis=(0,1))) / log_x_std(axis=(0,1))`\n- **Targets:**\n  - `(x - x_mean(axis=0)) / x_std(axis=0)`\n\n#### Model Architecture\n- **Core Architecture:** FiLM 1D UNet\n  - Scalar features processed by fully connected layers\n  - 1D features processed by 1D FiLM Convolution layers\n  - Initial and final convolutions divided into multiple branches\n  - Three head branches for temperature, q000X, and wind vector prediction\n  - Classification branch for state_q drops\n- **Loss Function:** Huber loss (beta=2)\n\n#### Training\n- **Optimizer:** Adam with clip norm\n- **Scheduler:** Cosine scheduler for 7 epochs\n- **Batch Size:** 384 with lr 0.0012\n- **Training Time:** 3 days on RTX 3090 for the entire LowRes dataset\n\n![Kmat pipeline](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1305588%2F5930445971412b47277c800212ccef86%2F2024-07-30%2014.25.58.png?generation=1722317176416622&alt=media)\n\n### 2nd Stage Modeling\nThe final submission from our team capaomat (3rd place) is an ensemble of predictions from three members. Some Camaro predictions (ptend_q0001, q0002, q0003) are refined by the 2nd stage. The score improvement is less than 0.0004. Neural network modeling is much more dominant.\n\n#### Features\nWe used a few features from raw inputs and predictions of 1st stage to prevent overfitting. State_t, state_q, ptend_q, future_state_q features and ratio of future_state_q2 to q3 are provided to the model.\n\n![Fraction of liquid cloud over total cloud as a function of temperature](https://arxiv.org/abs/2407.00124)\n\n#### Model / Training\nWe employed lightGBM to predict ptend_q at each level. Specifically, we trained the models and updated predictions for various levels. It took less than 20 minutes on CPU to train all 91 models.\n\n### Ensemble\nAs a final submission, we blended the following 9 models. Our solution achieved the following results:\n\n| exp                                      | Public LB | Private LB | Weight1 | Weight2 |\n| ---------------------------------------- | --------- | ---------- | ------- | ------- |\n| Camaro1_v2                               | 0.78504   | 0.78199    | 3.0     | 3.0     |\n| Camaro2_v2                               | 0.78451   | 0.78159    | 3.0     | 3.0     |\n| Camaro3_v2                               | 0.78572   | 0.78154    | 3.0     | 3.5     |\n| Camaro4_v2                               | 0.78427   | 0.78141    | 3.0     | 2.0     |\n| Camaro5_v2                               | 0.78255   | 0.77957    | 3.0     | 1.0     |\n| Pao1                                     | 0.78139   | 0.77770    | 1.0     | 0.5     |\n| Pao2                                     | 0.78252   | 0.77864    | 3.0     | 2.0     |\n| Pao3                                     | 0.77985   | 0.77801    | 1.0     | 0.5     |\n| Kmat1                                    | 0.78120   | 0.77647    | 3.0     | 1.5     |\n| Ensemble with weight1 (private best)     | 0.79026   | 0.78810    |         |         |\n| Ensemble with weight2 (final submission) | 0.79048   | 0.78792    |         |         |",
      "votes": 29
    },
    {
      "id": 2959182,
      "postDate": "2024-08-14T17:01:16.427Z",
      "content": "<p>Hi!, Congratulation for the 3rd place and thank you for your solution.</p>\n<p>i confused on one thing in <a href=\"https://www.kaggle.com/go5kuramubon\" target=\"_blank\">@go5kuramubon</a> part, that is the loss function. I don't get it how you are using EMA with huberloss, also i wasn't able to find any literature for the same. Can you help me with this, please ? It would be great if you can suggest some literature for the same as it seems quite interesting to me. Thank you in advance.</p>",
      "rawMarkdown": "Hi!, Congratulation for the 3rd place and thank you for your solution.\n\ni confused on one thing in @go5kuramubon part, that is the loss function. I don't get it how you are using EMA with huberloss, also i wasn't able to find any literature for the same. Can you help me with this, please ? It would be great if you can suggest some literature for the same as it seems quite interesting to me. Thank you in advance."
    },
    {
      "id": 2950022,
      "postDate": "2024-08-07T06:50:31.983Z",
      "content": "<p>Congratulation for the 3rd place. </p>\n<p>I have a question for the \"feature1 = feature1_original * a1 + b1 (a1 and b1 are trainable parameters)\" part. </p>\n<p>I'm still new so sorry if this is a dumb question. What do you mean by \"trainable parameters\"?<br>\nDo you mean that you trained so that the final R2 score become minimum, or something else? <br>\nI'm confused as to why this was introduced.</p>",
      "rawMarkdown": "Congratulation for the 3rd place. \n\nI have a question for the \"feature1 = feature1_original * a1 + b1 (a1 and b1 are trainable parameters)\" part. \n\nI'm still new so sorry if this is a dumb question. What do you mean by \"trainable parameters\"?\nDo you mean that you trained so that the final R2 score become minimum, or something else? \nI'm confused as to why this was introduced."
    },
    {
      "id": 2940579,
      "postDate": "2024-07-30T08:17:24.533Z",
      "content": "<p>Congratulation for the 3rd place.</p>\n<p><a href=\"https://www.kaggle.com/bamps53\" target=\"_blank\">@bamps53</a> </p>\n<ol>\n<li>Could you elaborate on what CLIP encoder is? Did you use OpenAI's multi-modal model architecture and pre-trained weight? <a href=\"https://arxiv.org/abs/2103.00020\" target=\"_blank\">https://arxiv.org/abs/2103.00020</a></li>\n<li>If Q1 is yes, how did you transfer the pre-trained weight to this competition task? And could you share the motivation to adapt this model to this competition task?</li>\n<li>How your validation data is split?</li>\n</ol>\n<p><a href=\"https://www.kaggle.com/kmat2019\" target=\"_blank\">@kmat2019</a> </p>\n<ol>\n<li>What FILM means? This paper? <a href=\"https://arxiv.org/abs/1709.07871\" target=\"_blank\">https://arxiv.org/abs/1709.07871</a></li>\n<li>Could you explain the motivation of introducing this technique to this competition task?</li>\n</ol>",
      "rawMarkdown": "Congratulation for the 3rd place.\n\n@bamps53 \n\n1. Could you elaborate on what CLIP encoder is? Did you use OpenAI's multi-modal model architecture and pre-trained weight? https://arxiv.org/abs/2103.00020\n1. If Q1 is yes, how did you transfer the pre-trained weight to this competition task? And could you share the motivation to adapt this model to this competition task?\n1. How your validation data is split?\n\n@kmat2019 \n\n1. What FILM means? This paper? https://arxiv.org/abs/1709.07871\n1. Could you explain the motivation of introducing this technique to this competition task?",
      "replies": [
        {
          "id": 2940782,
          "postDate": "2024-07-30T12:45:55.743Z",
          "content": "<p><a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> </p>\n<p>In this solution, it just means that I borrowed the transformer implementation from here:<br>\n<a href=\"https://github.com/huggingface/transformers/blob/main/src/transformers/models/clip/modeling_clip.py\" target=\"_blank\">https://github.com/huggingface/transformers/blob/main/src/transformers/models/clip/modeling_clip.py</a></p>\n<p>BUt sometimes I use pretrained weights from the CLIP text encoder to initialize the transformer like this:</p>\n<pre><code> transformers  CLIPModel\nself.transformer = CLIPModel.from_pretrained().text_model.encoder\nself.transformer.layers = self.transformer.layers[model_cfg.start_layer : model_cfg.end_layer]\n</code></pre>\n<p>In some sense, it knows how to process sequential data, so it is better than random initialization if the dataset size is small.</p>\n<p>However, the problem is the encoder has 512 dimensions, so it was too heavy for this challenge. I changed the encoder dimensions to 256 or 384 to iterate faster.</p>\n<p>My validation split was a random split with train:val = 99:1 for all low-resolution datasets.</p>",
          "rawMarkdown": "@tatamikenn \n\nIn this solution, it just means that I borrowed the transformer implementation from here:\nhttps://github.com/huggingface/transformers/blob/main/src/transformers/models/clip/modeling_clip.py\n\nBUt sometimes I use pretrained weights from the CLIP text encoder to initialize the transformer like this:\n\n```python\nfrom transformers import CLIPModel\nself.transformer = CLIPModel.from_pretrained(\"openai/clip-vit-base-patch32\").text_model.encoder\nself.transformer.layers = self.transformer.layers[model_cfg.start_layer : model_cfg.end_layer]\n```\n\nIn some sense, it knows how to process sequential data, so it is better than random initialization if the dataset size is small.\n\nHowever, the problem is the encoder has 512 dimensions, so it was too heavy for this challenge. I changed the encoder dimensions to 256 or 384 to iterate faster.\n\nMy validation split was a random split with train:val = 99:1 for all low-resolution datasets.\n",
          "votes": 1,
          "replies": [
            {
              "id": 2940799,
              "postDate": "2024-07-30T13:05:00.793Z",
              "content": "<p><a href=\"https://www.kaggle.com/bamps53\" target=\"_blank\">@bamps53</a> I got it. Thank you for detailed explanation.</p>",
              "rawMarkdown": "@bamps53 I got it. Thank you for detailed explanation."
            }
          ]
        },
        {
          "id": 2941040,
          "postDate": "2024-07-30T15:53:14.763Z",
          "content": "<p><a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> </p>\n<ol>\n<li>You're right.</li>\n<li>FiLM is one of the common conditioning technique for CNN architecture. It's also used in 1D CNN. I prefer FiLM to concatenation because I feel it conditions features more naturally. (However, there was almost no difference in this competition.)</li>\n</ol>",
          "rawMarkdown": "@tatamikenn \n1. You're right.\n2. FiLM is one of the common conditioning technique for CNN architecture. It's also used in 1D CNN. I prefer FiLM to concatenation because I feel it conditions features more naturally. (However, there was almost no difference in this competition.)",
          "votes": 1
        }
      ]
    },
    {
      "id": 2940747,
      "postDate": "2024-07-30T12:14:47.297Z",
      "rawMarkdown": "",
      "votes": -1,
      "isDeleted": true,
      "replies": [
        {
          "id": 2940772,
          "postDate": "2024-07-30T12:36:09.243Z",
          "content": "<p>Sorry if I'm wrong, but the discriminator in my head says this is LLM-generated text with 99% confidence.<br>\nWhy did you shoutout to Jerry here?</p>",
          "rawMarkdown": "Sorry if I'm wrong, but the discriminator in my head says this is LLM-generated text with 99% confidence.\nWhy did you shoutout to Jerry here?",
          "votes": 9
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2959182,
      "author_name": "Icees8",
      "author_url": "",
      "post_date": "2024-08-14T17:01:16.427000",
      "content": "<p>Hi!, Congratulation for the 3rd place and thank you for your solution.</p>\n<p>i confused on one thing in <a href=\"https://www.kaggle.com/go5kuramubon\" target=\"_blank\">@go5kuramubon</a> part, that is the loss function. I don't get it how you are using EMA with huberloss, also i wasn't able to find any literature for the same. Can you help me with this, please ? It would be great if you can suggest some literature for the same as it seems quite interesting to me. Thank you in advance.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2950022,
      "author_name": "professional food critic",
      "author_url": "",
      "post_date": "2024-08-07T06:50:31.983000",
      "content": "<p>Congratulation for the 3rd place. </p>\n<p>I have a question for the \"feature1 = feature1_original * a1 + b1 (a1 and b1 are trainable parameters)\" part. </p>\n<p>I'm still new so sorry if this is a dumb question. What do you mean by \"trainable parameters\"?<br>\nDo you mean that you trained so that the final R2 score become minimum, or something else? <br>\nI'm confused as to why this was introduced.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2940579,
      "author_name": "Bilzard",
      "author_url": "",
      "post_date": "2024-07-30T08:17:24.533000",
      "content": "<p>Congratulation for the 3rd place.</p>\n<p><a href=\"https://www.kaggle.com/bamps53\" target=\"_blank\">@bamps53</a> </p>\n<ol>\n<li>Could you elaborate on what CLIP encoder is? Did you use OpenAI's multi-modal model architecture and pre-trained weight? <a href=\"https://arxiv.org/abs/2103.00020\" target=\"_blank\">https://arxiv.org/abs/2103.00020</a></li>\n<li>If Q1 is yes, how did you transfer the pre-trained weight to this competition task? And could you share the motivation to adapt this model to this competition task?</li>\n<li>How your validation data is split?</li>\n</ol>\n<p><a href=\"https://www.kaggle.com/kmat2019\" target=\"_blank\">@kmat2019</a> </p>\n<ol>\n<li>What FILM means? This paper? <a href=\"https://arxiv.org/abs/1709.07871\" target=\"_blank\">https://arxiv.org/abs/1709.07871</a></li>\n<li>Could you explain the motivation of introducing this technique to this competition task?</li>\n</ol>",
      "votes": 0,
      "replies": [
        {
          "id": 2940782,
          "author_name": "Camaro",
          "author_url": "",
          "post_date": "2024-07-30T12:45:55.743000",
          "content": "<p><a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> </p>\n<p>In this solution, it just means that I borrowed the transformer implementation from here:<br>\n<a href=\"https://github.com/huggingface/transformers/blob/main/src/transformers/models/clip/modeling_clip.py\" target=\"_blank\">https://github.com/huggingface/transformers/blob/main/src/transformers/models/clip/modeling_clip.py</a></p>\n<p>BUt sometimes I use pretrained weights from the CLIP text encoder to initialize the transformer like this:</p>\n<pre><code> transformers  CLIPModel\nself.transformer = CLIPModel.from_pretrained().text_model.encoder\nself.transformer.layers = self.transformer.layers[model_cfg.start_layer : model_cfg.end_layer]\n</code></pre>\n<p>In some sense, it knows how to process sequential data, so it is better than random initialization if the dataset size is small.</p>\n<p>However, the problem is the encoder has 512 dimensions, so it was too heavy for this challenge. I changed the encoder dimensions to 256 or 384 to iterate faster.</p>\n<p>My validation split was a random split with train:val = 99:1 for all low-resolution datasets.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2940799,
              "author_name": "Bilzard",
              "author_url": "",
              "post_date": "2024-07-30T13:05:00.793000",
              "content": "<p><a href=\"https://www.kaggle.com/bamps53\" target=\"_blank\">@bamps53</a> I got it. Thank you for detailed explanation.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2941040,
          "author_name": "K_mat",
          "author_url": "",
          "post_date": "2024-07-30T15:53:14.763000",
          "content": "<p><a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> </p>\n<ol>\n<li>You're right.</li>\n<li>FiLM is one of the common conditioning technique for CNN architecture. It's also used in 1D CNN. I prefer FiLM to concatenation because I feel it conditions features more naturally. (However, there was almost no difference in this competition.)</li>\n</ol>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2940747,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-07-30T12:14:47.297000",
      "content": "",
      "votes": -1,
      "replies": [
        {
          "id": 2940772,
          "author_name": "Camaro",
          "author_url": "",
          "post_date": "2024-07-30T12:36:09.243000",
          "content": "<p>Sorry if I'm wrong, but the discriminator in my head says this is LLM-generated text with 99% confidence.<br>\nWhy did you shoutout to Jerry here?</p>",
          "votes": 9,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2940455": "Everyone, thank you for your hard work on the competition.\n\n\n@jerrylin  Thank you for organizing the competition. I know it must have been tough, but I believe we were able to get this far thanks to your sincere dedication up to the final check.\nCongratulations to everyone who ranked high and get good results. It was enjoyable to compete alongside you, and you all served as great motivation.\nTo my teammates @bamps53  @kmat2019 , it was fun and I learned a lot. Thank you very much.\nAlso I am happy to become Grandmaster at this competition. Also @kmat2019 became grandmaster. Congrats!!\n\n\n## Model Summary\n### Overall Summary\n\n![Overall pipeline](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1305588%2Fa0b89d5c5ca6aca62baa4cfb090ac7bb%2Fimage%20(2).png?generation=1722316979148053&alt=media)\n\nEach team member built their own neural network models. \n\nAfter obtaining the Camaro model's predictions, we created several features to input into GBDT regressors. These regressors refined the model's predictions. Although this second stage could be applied to other models, we only applied it to the Camaro model because the Pao and Kmat models lacked a validation dataset.\n\nThe final prediction was calculated as a weighted average of these predictions.\n\n### Pao Part\n\n#### Overview\n- **Model:** 1d CNN + Transformer + LSTM\n- **Feature:** Original and relative humidity with those sequence 1st derivative and 2nd derivative at heights (diff and diff-diff)\n- **Auxiliary Loss:** Predicting the difference between adjacent vertical levels\n- **Row-less full training**\n\n#### Model Architecture\n\n![Pao model](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1305588%2F3142f36d09f2e57ddcdcb023222eaba6%2F2024-07-30%2014.24.46.png?generation=1722317122204086&alt=media)\n\n- **Input layer:** Linear transformation with bias at each feature\n  - Example: `feature1 = feature1_original * a1 + b1` (a1 and b1 are trainable parameters)\n- **Feature Concatenation:** Concatenate scalar features to each height sequence feature after each dense layer\n  - Example: `features_level0 = concat([seq_features_level0, dense_level0(scalar_features)])`\n- **Positional Encoding:** Adding embedding per height to the hidden layer\n- **Model Blocks:**\n  - **Residual block:** 1dCNN (Conv1d + BN + GELU) * 2 + Transformer\n    - Conv1d: kernel size = 5\n    - Transformer: n_head = 8, n_layers = 1\n  - **LSTM:** Bidirectional 2 layers\n  - **MLP:** Simple MLP with (Linear and GELU with no dropout)\n\n\n\n#### Training\n- **Auxiliary Loss:** Predicting the difference between adjacent vertical levels (similar to Camaro part)\n- **Loss:** Huber Loss using EMA\n- **Learning rate:** Cosine annealing 1e-3 to 1e-5, 10 epochs\n- **Optimizer:** AdamW\n\n#### Others\n- **Dataset:** WebDataset for LowRes full-training\n- **Normalization:** \n  - **Feature:** `(input - mean(input)) / std(input)`\n  - **Target:** Multiply old_sample_submission weight\n- **Postprocess:** Replace predictions for `ptend_q0002_0-27` with `-1 * input / 1200`\n- **Ensemble:** 3 models with variations in dropout, hidden size, and Huber Loss delta\n\n### Camaro Part\n\n#### Overview\n- **Full Dataset and Long Training:** Training on the full dataset\n- **Model Architecture:** Combination of CNN and Transformer or Transformer-only using the CLIP Encoder\n- **Auxiliary Loss:** Predicting the difference between adjacent vertical levels\n\n#### Dataset and Preprocessing\n- **HuggingFace Dataset:** Full dataset for training\n- **Feature Engineering:** Added saturation vapor pressure as a feature\n- **Normalization:** Pre-computed using Kaggle train and test datasets\n- **WebDataset:** Efficient loading for faster training\n\n#### Model\n- **Architecture:** Combination of CNN and Transformer or Transformer-only with heavy embedding and head layers\n\n#### Training\n- **Target Transformation:** Dividing by old sample submission weights and subtracting the mean\n- **Auxiliary Loss:** Incorporating information about derivatives and second derivatives\n- **Loss Function:** Huber loss with a delta of 2.0\n\n#### Post-Processing\n- Replace predictions for `ptend_q0002_0-26` with `-1 * input / 1200`\n\n#### Results\n| exp | bs  | arch               | dim | loss   | Public LB | Private LB | Public LB (2nd stage) | Private LB (2nd stage) |\n|-----|-----|--------------------|-----|--------|-----------|------------|-----------------------|------------------------|\n| 1   | 256 | ConvTransformer    | 256 | Huber2 | 0.78468   | 0.78176    | 0.78504               | 0.78199                |\n| 2   | 1024| ConvTransformer    | 384 | Huber2 | 0.78418   | 0.78131    | 0.78451               | 0.78159                |\n| 3   | 512 | Transformer n_layer=8 | 256 | Huber2 | 0.78545   | 0.78154    | 0.78572               | 0.78154                |\n| 4   | 768 | ConvTransformer x 2| 256 | Huber4 | 0.78395   | 0.78122    | 0.78427               | 0.78141                |\n| 5   | 768 | Transformer n_layer=6 | 256 | Huber8 | 0.78245   | 0.77951    | 0.78255               | 0.77957                |\n| **Ensemble (1+2+3+4)** |     |                    |     |        |           |            | 0.79025               | 0.78694                |\n| **Ensemble (1+2+3+4+5)** |     |                    |     |        |           |            | 0.78998               | 0.78685                | 0.79023                | 0.78703                |\n\n### Kmat Part\n\n#### Overview\nAs shown in Fig.6-1, Kmat part consists of:\n- Add Features\n- Normalize by averages and standards\n- 1D CNN model to predict climate\n- Postprocess (some predictions are replaced by `-input/1200`)\n\n#### Feature Engineering\n- Diff features from 1D data: `x[z] - x[z-1]`\n- Relative humidity-related features such as dew_point and vapor_pressure / saturation_pressure\n\n#### Normalization\n- **Inputs:**\n  1. `(x - x_mean(axis=0)) / x_std(axis=0)`\n  2. `(x - x_mean(axis=(0,1))) / x_std(axis=(0,1))`\n  3. `(log_x - log_x_mean(axis=(0,1))) / log_x_std(axis=(0,1))`\n- **Targets:**\n  - `(x - x_mean(axis=0)) / x_std(axis=0)`\n\n#### Model Architecture\n- **Core Architecture:** FiLM 1D UNet\n  - Scalar features processed by fully connected layers\n  - 1D features processed by 1D FiLM Convolution layers\n  - Initial and final convolutions divided into multiple branches\n  - Three head branches for temperature, q000X, and wind vector prediction\n  - Classification branch for state_q drops\n- **Loss Function:** Huber loss (beta=2)\n\n#### Training\n- **Optimizer:** Adam with clip norm\n- **Scheduler:** Cosine scheduler for 7 epochs\n- **Batch Size:** 384 with lr 0.0012\n- **Training Time:** 3 days on RTX 3090 for the entire LowRes dataset\n\n![Kmat pipeline](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1305588%2F5930445971412b47277c800212ccef86%2F2024-07-30%2014.25.58.png?generation=1722317176416622&alt=media)\n\n### 2nd Stage Modeling\nThe final submission from our team capaomat (3rd place) is an ensemble of predictions from three members. Some Camaro predictions (ptend_q0001, q0002, q0003) are refined by the 2nd stage. The score improvement is less than 0.0004. Neural network modeling is much more dominant.\n\n#### Features\nWe used a few features from raw inputs and predictions of 1st stage to prevent overfitting. State_t, state_q, ptend_q, future_state_q features and ratio of future_state_q2 to q3 are provided to the model.\n\n![Fraction of liquid cloud over total cloud as a function of temperature](https://arxiv.org/abs/2407.00124)\n\n#### Model / Training\nWe employed lightGBM to predict ptend_q at each level. Specifically, we trained the models and updated predictions for various levels. It took less than 20 minutes on CPU to train all 91 models.\n\n### Ensemble\nAs a final submission, we blended the following 9 models. Our solution achieved the following results:\n\n| exp                                      | Public LB | Private LB | Weight1 | Weight2 |\n| ---------------------------------------- | --------- | ---------- | ------- | ------- |\n| Camaro1_v2                               | 0.78504   | 0.78199    | 3.0     | 3.0     |\n| Camaro2_v2                               | 0.78451   | 0.78159    | 3.0     | 3.0     |\n| Camaro3_v2                               | 0.78572   | 0.78154    | 3.0     | 3.5     |\n| Camaro4_v2                               | 0.78427   | 0.78141    | 3.0     | 2.0     |\n| Camaro5_v2                               | 0.78255   | 0.77957    | 3.0     | 1.0     |\n| Pao1                                     | 0.78139   | 0.77770    | 1.0     | 0.5     |\n| Pao2                                     | 0.78252   | 0.77864    | 3.0     | 2.0     |\n| Pao3                                     | 0.77985   | 0.77801    | 1.0     | 0.5     |\n| Kmat1                                    | 0.78120   | 0.77647    | 3.0     | 1.5     |\n| Ensemble with weight1 (private best)     | 0.79026   | 0.78810    |         |         |\n| Ensemble with weight2 (final submission) | 0.79048   | 0.78792    |         |         |",
    "2959182": "Hi!, Congratulation for the 3rd place and thank you for your solution.\n\ni confused on one thing in @go5kuramubon part, that is the loss function. I don't get it how you are using EMA with huberloss, also i wasn't able to find any literature for the same. Can you help me with this, please ? It would be great if you can suggest some literature for the same as it seems quite interesting to me. Thank you in advance.",
    "2950022": "Congratulation for the 3rd place. \n\nI have a question for the \"feature1 = feature1_original * a1 + b1 (a1 and b1 are trainable parameters)\" part. \n\nI'm still new so sorry if this is a dumb question. What do you mean by \"trainable parameters\"?\nDo you mean that you trained so that the final R2 score become minimum, or something else? \nI'm confused as to why this was introduced.",
    "2940579": "Congratulation for the 3rd place.\n\n@bamps53 \n\n1. Could you elaborate on what CLIP encoder is? Did you use OpenAI's multi-modal model architecture and pre-trained weight? https://arxiv.org/abs/2103.00020\n1. If Q1 is yes, how did you transfer the pre-trained weight to this competition task? And could you share the motivation to adapt this model to this competition task?\n1. How your validation data is split?\n\n@kmat2019 \n\n1. What FILM means? This paper? https://arxiv.org/abs/1709.07871\n1. Could you explain the motivation of introducing this technique to this competition task?",
    "2940747": ""
  }
}