{
  "id": 523041,
  "title": "10th Place Solution",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/523041",
  "author_name": "Bilzard",
  "post_date": "2024-07-29T22:41:24.873000",
  "votes": 33,
  "comment_count": 2,
  "views": 0,
  "content": "<h1>10th Place Solution</h1>\n<p>Thank you for holding this competition. We enjoyed a lot, as well as a lot of leaning experience. This was also really tough competition we ever participated. We literary training the model and making ensemble few hours before the competition closed, and finally we managed to win a gold medal.</p>\n<p>The following is the summary of 10th place solution by team Updraft. More detailed writeup might be shared by each members.</p>\n<h2>Team Members</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> (team leader)</li>\n<li><a href=\"https://www.kaggle.com/tereka\" target=\"_blank\">@tereka</a></li>\n<li><a href=\"https://www.kaggle.com/phalanx\" target=\"_blank\">@phalanx</a></li>\n<li><a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a></li>\n</ul>\n<h2>Glossary</h2>\n<ul>\n<li>low resolution data = LR</li>\n<li>high resolution data = HR</li>\n</ul>\n<h2>Summary</h2>\n<p>Our solution is ensemble of each 4 member's regression models.<br>\nThe models in the final submission contains 23 models with different architectures and learning conditions.<br>\nWe developed the following architectures.</p>\n<ul>\n<li>Conv + Transformer (+ LSTM)</li>\n<li>CNN / U-Net</li>\n<li>LSTM</li>\n</ul>\n<h2>Train/Validation Data</h2>\n<p>We basically trained our models using 1-7y LR data, and evaluate with 1/6 subsample of 8y LR data.<br>\nSome model also use:</p>\n<ul>\n<li>8y LR data</li>\n<li>1-8y HR data</li>\n<li>pseudo labeled old/new test data.</li>\n</ul>\n<h2>Evaluation Metrics</h2>\n<p>We used 1-MSE instead of R2 after std normalized target.<br>\nWe use this metrics because some member reported R2 loses correlation after host replaced the test dataset.<br>\n(However, later we found both are correlated when using sufficiently small eps for normalization.)<br>\nIn normalization, we used the following formula:</p>\n<pre><code>eps =  ** (-)\ntarget = target / maximum(std(target), eps)\n</code></pre>\n<h2>Model Design</h2>\n<h3>Bilzard's Pipeline</h3>\n<p>Basically 1/7 ~ 1/6 subsample of LR data is used because of relatively <em>poor</em> machine resource (RTX4090 x 1).<br>\nWhat worked most was:</p>\n<ul>\n<li>feature extraction of Climate-invariant features (RH, plume buoyancy, normalized Heat Flux etc.)[1]</li>\n<li>confidence-aware MSE loss</li>\n<li>doping hard sample of 1/1 full LR data</li>\n</ul>\n<h4>Climate-invariant feature extraction</h4>\n<p>We extracted climate-invariant features which is suggested by Tom Beucler et. al. [1].</p>\n<ul>\n<li>Relative Humidity (RH)</li>\n<li>Plume buoyancy</li>\n<li>normalized Heat Flux</li>\n</ul>\n<p>Conceptually, these features are invariant under both cool and warm temperature.<br>\nSo these features are expected to cancel out the effect of yearly warm-up in ClimSim dataset (which I observed by EDA).</p>\n<p>During the competition, we couldn't train models with this features on the full LR dataset, but it showed a significant gain of ~1% on a 1/7 to 1/6 subsample of LR dataset.</p>\n<p>Note that the late submission results shows this features still have a significant gain on the full LR dataset (see Appendix A1).</p>\n<h4>Confidence-aware MSE loss</h4>\n<p>We used the following loss which we called <em>confidence-aware MSE loss</em> instead of MSE.<br>\nIntuitively it has the effect of loosening the contributions of the large error by hard samples.<br>\nThis idea was variant of the <a href=\"https://www.kaggle.com/ryomak\" target=\"_blank\">@ryomak</a> 's previously shared solution[2].<br>\nNote that variance of target was regressed by model in addition to mean value. So our model's output variable is 2x larger than basic models.</p>\n<pre><code>loss =  * (target - pred) **  / var +  * torch.log(var)\n</code></pre>\n<h4>Doping Hard Examples</h4>\n<p>As we shared on a public discussion[3], most of error came from extreme outliers. So we came to think that hard samples are good representatives of all dataset: if we doped small amount of hard samples from 1/1 full LR data, it would help to reduces regression error even if we used 1/n sub-sampled LR data.</p>\n<p>We extracted ~14% of hard samples which have the squared error greater than threshold using one of our best models, and doped to train other models. We found this has ~0.x% gain using this method.<br>\nWe also found adding tiny amount of neighbor time-frame sample from the same location also helps gaining accuracy.</p>\n<h4>Model Architecture</h4>\n<p>Basically we used the following two architecture</p>\n<ol>\n<li>Conv-Transformer (variant of <a href=\"https://www.kaggle.com/hoyso48\" target=\"_blank\">@hoyso48</a> 's architecture[4])</li>\n<li>Pixel-Shuffle stacked UNet</li>\n</ol>\n<p>The second architecture was inspired by stacked hourglass architecture[5] which stacks U-Net and gradually improves the accuracy of output variable by regressing error of previous stack (like boosting). In addition, replacing ConvTranspose decoder in the usual U-Net to PixelShuffle gained ~0.x%.</p>\n<h4>Other modeling Tips</h4>\n<h5>Tanh normalization</h5>\n<p>We often encounter gradient explosion when training deeper transformers (~8-12 layers).<br>\nWe mitigated this issue by utilizing trick we called <em>tanh normalization</em>.<br>\nIt is just inserting tanh activation before calculating softmax attention:</p>\n<pre><code>     ():\n        x = rearrange(x, )\n\n        qkv = self.qkv(x)\n        qkv = rearrange(qkv, , h=self.num_heads)\n        q, k, v = qkv.chunk(, dim=-)\n\n        attn = (q @ k.transpose(-, -)) * self.scale\n        attn = attn.tanh() * self.attn_clip_val  \n        attn = F.softmax(attn, dim=-)\n\n        x = attn @ v\n        x = rearrange(x, )\n\n        x = self.proj(x)\n        x = rearrange(x, )\n         x\n</code></pre>\n<h5>Removing Dropout</h5>\n<p>In the previous competition, some Kagglers reported removing Dropout layer has significant improvement on regression tasks[6]. We observed similar phenomena on this competition.</p>\n<h3>Tereka's Pipeline</h3>\n<pre><code>Public/Private LB: 0.77784/0.77023\n\nModel\n Transformer + Convolution + LSTM\n Architecture Points\n Learnable Positional Embedding(nn.Embedding 60)\n several Transformer Layers\n Final Layer LSTM\n\nFeatures\n Original Dataset(standard scaler)\n\nDataset\n 70M(full low resolution dataset)\n\nLoss\n Huber Loss(&gt; MSE)\n\nNot Worked\n Mixup\n EMA\n Deeper model, I tried some experiment, but it’s not stable.\n pretrained High resolution model, but low resolution and some high resolution data is worked in public LB.\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F0d32daa11aecc54e4573a294b244ec79%2F21_architecture_tereka.png?generation=1722302269909845&amp;alt=media\" alt=\"\"></p>\n<h3>Phalanx's Pipeline</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F194b6d4d267ae9c693f3fba5ab444cad%2F21_architecture_phalanx.png?generation=1722302291403618&amp;alt=media\" alt=\"\"></p>\n<h3>Ryches' Pipeline</h3>\n<p>TBD</p>\n<h2>Ensemble Method</h2>\n<p>We used tree type of ensemble method and final submission is weighted ensemble of these:</p>\n<ul>\n<li>type1) Stacking using models trained with 1-7y low resolution data</li>\n<li>type2) Stacking using models trained with 1-8y low resolution data</li>\n<li>type3) Mean ensemble of models trained with 1-8y models</li>\n</ul>\n<p>For type2, basically we can't use the data which trained the seed model to fit parameter in the stacking.<br>\nSo we take the following scheme:</p>\n<ol>\n<li>calculate the ensemble weight using model trained with 1-7y data</li>\n<li>re-train the model of the same architecture, same training condition except swapping train data to 1-8y data</li>\n<li>using ensemble weight calculated on step2, but replacing the model from 1-7y to 1-8y data to make the test prediction</li>\n</ol>\n<p>For the final submission, we ensemble the type-1 ~ type-3 model with manually fitted weight:</p>\n<pre><code>final_prediction = type1 * w1 + type2 * w2 + type3 * w3\n</code></pre>\n<p>The late submission result shows assigning 100% weight for type1 model performs the best on the private LB.</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>LB(Public)</th>\n<th>LB(Private)</th>\n<th>final submission</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>type1 * 0.5 + type2 * 0.3 + type3 * 0.2</td>\n<td>0.78803</td>\n<td>0.78503</td>\n<td>✔️</td>\n</tr>\n<tr>\n<td>type1 * 0.75 + type3 * 0.25</td>\n<td>0.78801</td>\n<td>0.78487</td>\n<td>✔️</td>\n</tr>\n<tr>\n<td>type1 * 0.6 * type2 * 0.4</td>\n<td>0.78777</td>\n<td>0.78535</td>\n<td></td>\n</tr>\n<tr>\n<td>type1</td>\n<td>0.78771</td>\n<td><strong>0.78536</strong></td>\n<td></td>\n</tr>\n</tbody>\n</table>\n<h2>Detail of Stacking</h2>\n<p>We trained ensemble weight per (target, model). This is based on the study of EDA where we found each model has their expertize: some models performs good at regressing mixing ratio and some are good at regressing wind velocity etc.</p>\n<p>We tested the following variations and found ensemble without normalization and with bias performs the best.</p>\n<table>\n<thead>\n<tr>\n<th>normalize weight?</th>\n<th>with bias?</th>\n<th>LB(Public)</th>\n<th>LB(Private)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>✔️</td>\n<td>-</td>\n<td>0.78727</td>\n<td>0.78532</td>\n</tr>\n<tr>\n<td>-</td>\n<td>-</td>\n<td>0.78739</td>\n<td>0.78563</td>\n</tr>\n<tr>\n<td>-</td>\n<td>✔️</td>\n<td>0.78740</td>\n<td><strong>0.78569</strong></td>\n</tr>\n</tbody>\n</table>\n<h2>Appendix</h2>\n<h3>A1. LR full data + Climate Invariant Feature</h3>\n<p>During the competition, we only trained our models using climate-invariant features on a subsample of the LR data.However, after the competition ended, we successfully trained Phalanx's lightweight CNN model train model with climate-invariant features on the full sample of the LR dataset in feasible training time (~1 day) with  a relatively \"low-throughput\" GPU (RTX 4090).</p>\n<p>The result of late submission shows it has significant gain of 0.05% on private LB. So climate-invariant feature has actually significant impact on this task.</p>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>CV(1-MSE)</th>\n<th>LB(Public)</th>\n<th>LB(Private)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>single w/o LR-full-clim-invariant</td>\n<td>0.78231</td>\n<td>0.78025</td>\n<td>0.77721</td>\n</tr>\n<tr>\n<td>single w LR-full-clim-invariant</td>\n<td><strong>0.78434</strong></td>\n<td><strong>0.7814</strong></td>\n<td><strong>0.78017</strong></td>\n</tr>\n<tr>\n<td>ensemble w/o LR-full-clim-invariant</td>\n<td>0.79192</td>\n<td>0.78740</td>\n<td>0.78569</td>\n</tr>\n<tr>\n<td>ensemble w LR-full-clim-invariant</td>\n<td><strong>0.79218</strong></td>\n<td><strong>0.78771</strong></td>\n<td><strong>0.78618</strong></td>\n</tr>\n</tbody>\n</table>\n<h2>Reference</h2>\n<ul>\n<li>[1] <a href=\"https://arxiv.org/abs/2112.08440\" target=\"_blank\">Climate-Invariant Machine Learning</a></li>\n<li>[2] <a href=\"https://www.kaggle.com/competitions/ventilator-pressure-prediction/discussion/285353\" target=\"_blank\">9th solution(with github code): Laplace distribution noise likelihood optimization</a></li>\n<li>[3] <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/511274\" target=\"_blank\">Is Tropical Cylones Key to Win this Competition?</a></li>\n<li>[4] <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/406684\" target=\"_blank\">1st place solution - 1DCNN combined with Transformer</a></li>\n<li>[5] <a href=\"https://arxiv.org/abs/1603.06937\" target=\"_blank\">Stacked Hourglass Networks for Human Pose Estimation</a></li>\n<li>[6] <a href=\"https://www.kaggle.com/competitions/commonlitreadabilityprize/discussion/260729\" target=\"_blank\">The Magic of No Dropout</a></li>\n</ul>",
  "messages": [
    {
      "id": 2940258,
      "postDate": "2024-07-29T22:41:24.873Z",
      "content": "<h1>10th Place Solution</h1>\n<p>Thank you for holding this competition. We enjoyed a lot, as well as a lot of leaning experience. This was also really tough competition we ever participated. We literary training the model and making ensemble few hours before the competition closed, and finally we managed to win a gold medal.</p>\n<p>The following is the summary of 10th place solution by team Updraft. More detailed writeup might be shared by each members.</p>\n<h2>Team Members</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> (team leader)</li>\n<li><a href=\"https://www.kaggle.com/tereka\" target=\"_blank\">@tereka</a></li>\n<li><a href=\"https://www.kaggle.com/phalanx\" target=\"_blank\">@phalanx</a></li>\n<li><a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a></li>\n</ul>\n<h2>Glossary</h2>\n<ul>\n<li>low resolution data = LR</li>\n<li>high resolution data = HR</li>\n</ul>\n<h2>Summary</h2>\n<p>Our solution is ensemble of each 4 member's regression models.<br>\nThe models in the final submission contains 23 models with different architectures and learning conditions.<br>\nWe developed the following architectures.</p>\n<ul>\n<li>Conv + Transformer (+ LSTM)</li>\n<li>CNN / U-Net</li>\n<li>LSTM</li>\n</ul>\n<h2>Train/Validation Data</h2>\n<p>We basically trained our models using 1-7y LR data, and evaluate with 1/6 subsample of 8y LR data.<br>\nSome model also use:</p>\n<ul>\n<li>8y LR data</li>\n<li>1-8y HR data</li>\n<li>pseudo labeled old/new test data.</li>\n</ul>\n<h2>Evaluation Metrics</h2>\n<p>We used 1-MSE instead of R2 after std normalized target.<br>\nWe use this metrics because some member reported R2 loses correlation after host replaced the test dataset.<br>\n(However, later we found both are correlated when using sufficiently small eps for normalization.)<br>\nIn normalization, we used the following formula:</p>\n<pre><code>eps =  ** (-)\ntarget = target / maximum(std(target), eps)\n</code></pre>\n<h2>Model Design</h2>\n<h3>Bilzard's Pipeline</h3>\n<p>Basically 1/7 ~ 1/6 subsample of LR data is used because of relatively <em>poor</em> machine resource (RTX4090 x 1).<br>\nWhat worked most was:</p>\n<ul>\n<li>feature extraction of Climate-invariant features (RH, plume buoyancy, normalized Heat Flux etc.)[1]</li>\n<li>confidence-aware MSE loss</li>\n<li>doping hard sample of 1/1 full LR data</li>\n</ul>\n<h4>Climate-invariant feature extraction</h4>\n<p>We extracted climate-invariant features which is suggested by Tom Beucler et. al. [1].</p>\n<ul>\n<li>Relative Humidity (RH)</li>\n<li>Plume buoyancy</li>\n<li>normalized Heat Flux</li>\n</ul>\n<p>Conceptually, these features are invariant under both cool and warm temperature.<br>\nSo these features are expected to cancel out the effect of yearly warm-up in ClimSim dataset (which I observed by EDA).</p>\n<p>During the competition, we couldn't train models with this features on the full LR dataset, but it showed a significant gain of ~1% on a 1/7 to 1/6 subsample of LR dataset.</p>\n<p>Note that the late submission results shows this features still have a significant gain on the full LR dataset (see Appendix A1).</p>\n<h4>Confidence-aware MSE loss</h4>\n<p>We used the following loss which we called <em>confidence-aware MSE loss</em> instead of MSE.<br>\nIntuitively it has the effect of loosening the contributions of the large error by hard samples.<br>\nThis idea was variant of the <a href=\"https://www.kaggle.com/ryomak\" target=\"_blank\">@ryomak</a> 's previously shared solution[2].<br>\nNote that variance of target was regressed by model in addition to mean value. So our model's output variable is 2x larger than basic models.</p>\n<pre><code>loss =  * (target - pred) **  / var +  * torch.log(var)\n</code></pre>\n<h4>Doping Hard Examples</h4>\n<p>As we shared on a public discussion[3], most of error came from extreme outliers. So we came to think that hard samples are good representatives of all dataset: if we doped small amount of hard samples from 1/1 full LR data, it would help to reduces regression error even if we used 1/n sub-sampled LR data.</p>\n<p>We extracted ~14% of hard samples which have the squared error greater than threshold using one of our best models, and doped to train other models. We found this has ~0.x% gain using this method.<br>\nWe also found adding tiny amount of neighbor time-frame sample from the same location also helps gaining accuracy.</p>\n<h4>Model Architecture</h4>\n<p>Basically we used the following two architecture</p>\n<ol>\n<li>Conv-Transformer (variant of <a href=\"https://www.kaggle.com/hoyso48\" target=\"_blank\">@hoyso48</a> 's architecture[4])</li>\n<li>Pixel-Shuffle stacked UNet</li>\n</ol>\n<p>The second architecture was inspired by stacked hourglass architecture[5] which stacks U-Net and gradually improves the accuracy of output variable by regressing error of previous stack (like boosting). In addition, replacing ConvTranspose decoder in the usual U-Net to PixelShuffle gained ~0.x%.</p>\n<h4>Other modeling Tips</h4>\n<h5>Tanh normalization</h5>\n<p>We often encounter gradient explosion when training deeper transformers (~8-12 layers).<br>\nWe mitigated this issue by utilizing trick we called <em>tanh normalization</em>.<br>\nIt is just inserting tanh activation before calculating softmax attention:</p>\n<pre><code>     ():\n        x = rearrange(x, )\n\n        qkv = self.qkv(x)\n        qkv = rearrange(qkv, , h=self.num_heads)\n        q, k, v = qkv.chunk(, dim=-)\n\n        attn = (q @ k.transpose(-, -)) * self.scale\n        attn = attn.tanh() * self.attn_clip_val  \n        attn = F.softmax(attn, dim=-)\n\n        x = attn @ v\n        x = rearrange(x, )\n\n        x = self.proj(x)\n        x = rearrange(x, )\n         x\n</code></pre>\n<h5>Removing Dropout</h5>\n<p>In the previous competition, some Kagglers reported removing Dropout layer has significant improvement on regression tasks[6]. We observed similar phenomena on this competition.</p>\n<h3>Tereka's Pipeline</h3>\n<pre><code>Public/Private LB: 0.77784/0.77023\n\nModel\n Transformer + Convolution + LSTM\n Architecture Points\n Learnable Positional Embedding(nn.Embedding 60)\n several Transformer Layers\n Final Layer LSTM\n\nFeatures\n Original Dataset(standard scaler)\n\nDataset\n 70M(full low resolution dataset)\n\nLoss\n Huber Loss(&gt; MSE)\n\nNot Worked\n Mixup\n EMA\n Deeper model, I tried some experiment, but it’s not stable.\n pretrained High resolution model, but low resolution and some high resolution data is worked in public LB.\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F0d32daa11aecc54e4573a294b244ec79%2F21_architecture_tereka.png?generation=1722302269909845&amp;alt=media\" alt=\"\"></p>\n<h3>Phalanx's Pipeline</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F194b6d4d267ae9c693f3fba5ab444cad%2F21_architecture_phalanx.png?generation=1722302291403618&amp;alt=media\" alt=\"\"></p>\n<h3>Ryches' Pipeline</h3>\n<p>TBD</p>\n<h2>Ensemble Method</h2>\n<p>We used tree type of ensemble method and final submission is weighted ensemble of these:</p>\n<ul>\n<li>type1) Stacking using models trained with 1-7y low resolution data</li>\n<li>type2) Stacking using models trained with 1-8y low resolution data</li>\n<li>type3) Mean ensemble of models trained with 1-8y models</li>\n</ul>\n<p>For type2, basically we can't use the data which trained the seed model to fit parameter in the stacking.<br>\nSo we take the following scheme:</p>\n<ol>\n<li>calculate the ensemble weight using model trained with 1-7y data</li>\n<li>re-train the model of the same architecture, same training condition except swapping train data to 1-8y data</li>\n<li>using ensemble weight calculated on step2, but replacing the model from 1-7y to 1-8y data to make the test prediction</li>\n</ol>\n<p>For the final submission, we ensemble the type-1 ~ type-3 model with manually fitted weight:</p>\n<pre><code>final_prediction = type1 * w1 + type2 * w2 + type3 * w3\n</code></pre>\n<p>The late submission result shows assigning 100% weight for type1 model performs the best on the private LB.</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>LB(Public)</th>\n<th>LB(Private)</th>\n<th>final submission</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>type1 * 0.5 + type2 * 0.3 + type3 * 0.2</td>\n<td>0.78803</td>\n<td>0.78503</td>\n<td>✔️</td>\n</tr>\n<tr>\n<td>type1 * 0.75 + type3 * 0.25</td>\n<td>0.78801</td>\n<td>0.78487</td>\n<td>✔️</td>\n</tr>\n<tr>\n<td>type1 * 0.6 * type2 * 0.4</td>\n<td>0.78777</td>\n<td>0.78535</td>\n<td></td>\n</tr>\n<tr>\n<td>type1</td>\n<td>0.78771</td>\n<td><strong>0.78536</strong></td>\n<td></td>\n</tr>\n</tbody>\n</table>\n<h2>Detail of Stacking</h2>\n<p>We trained ensemble weight per (target, model). This is based on the study of EDA where we found each model has their expertize: some models performs good at regressing mixing ratio and some are good at regressing wind velocity etc.</p>\n<p>We tested the following variations and found ensemble without normalization and with bias performs the best.</p>\n<table>\n<thead>\n<tr>\n<th>normalize weight?</th>\n<th>with bias?</th>\n<th>LB(Public)</th>\n<th>LB(Private)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>✔️</td>\n<td>-</td>\n<td>0.78727</td>\n<td>0.78532</td>\n</tr>\n<tr>\n<td>-</td>\n<td>-</td>\n<td>0.78739</td>\n<td>0.78563</td>\n</tr>\n<tr>\n<td>-</td>\n<td>✔️</td>\n<td>0.78740</td>\n<td><strong>0.78569</strong></td>\n</tr>\n</tbody>\n</table>\n<h2>Appendix</h2>\n<h3>A1. LR full data + Climate Invariant Feature</h3>\n<p>During the competition, we only trained our models using climate-invariant features on a subsample of the LR data.However, after the competition ended, we successfully trained Phalanx's lightweight CNN model train model with climate-invariant features on the full sample of the LR dataset in feasible training time (~1 day) with  a relatively \"low-throughput\" GPU (RTX 4090).</p>\n<p>The result of late submission shows it has significant gain of 0.05% on private LB. So climate-invariant feature has actually significant impact on this task.</p>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>CV(1-MSE)</th>\n<th>LB(Public)</th>\n<th>LB(Private)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>single w/o LR-full-clim-invariant</td>\n<td>0.78231</td>\n<td>0.78025</td>\n<td>0.77721</td>\n</tr>\n<tr>\n<td>single w LR-full-clim-invariant</td>\n<td><strong>0.78434</strong></td>\n<td><strong>0.7814</strong></td>\n<td><strong>0.78017</strong></td>\n</tr>\n<tr>\n<td>ensemble w/o LR-full-clim-invariant</td>\n<td>0.79192</td>\n<td>0.78740</td>\n<td>0.78569</td>\n</tr>\n<tr>\n<td>ensemble w LR-full-clim-invariant</td>\n<td><strong>0.79218</strong></td>\n<td><strong>0.78771</strong></td>\n<td><strong>0.78618</strong></td>\n</tr>\n</tbody>\n</table>\n<h2>Reference</h2>\n<ul>\n<li>[1] <a href=\"https://arxiv.org/abs/2112.08440\" target=\"_blank\">Climate-Invariant Machine Learning</a></li>\n<li>[2] <a href=\"https://www.kaggle.com/competitions/ventilator-pressure-prediction/discussion/285353\" target=\"_blank\">9th solution(with github code): Laplace distribution noise likelihood optimization</a></li>\n<li>[3] <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/511274\" target=\"_blank\">Is Tropical Cylones Key to Win this Competition?</a></li>\n<li>[4] <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/406684\" target=\"_blank\">1st place solution - 1DCNN combined with Transformer</a></li>\n<li>[5] <a href=\"https://arxiv.org/abs/1603.06937\" target=\"_blank\">Stacked Hourglass Networks for Human Pose Estimation</a></li>\n<li>[6] <a href=\"https://www.kaggle.com/competitions/commonlitreadabilityprize/discussion/260729\" target=\"_blank\">The Magic of No Dropout</a></li>\n</ul>",
      "rawMarkdown": "# 10th Place Solution\n\nThank you for holding this competition. We enjoyed a lot, as well as a lot of leaning experience. This was also really tough competition we ever participated. We literary training the model and making ensemble few hours before the competition closed, and finally we managed to win a gold medal.\n\nThe following is the summary of 10th place solution by team Updraft. More detailed writeup might be shared by each members.\n\n## Team Members\n\n* @tatamikenn (team leader)\n* @tereka\n* @phalanx\n* @ryches\n\n## Glossary\n\n* low resolution data = LR\n* high resolution data = HR\n\n## Summary\n\nOur solution is ensemble of each 4 member's regression models.\nThe models in the final submission contains 23 models with different architectures and learning conditions.\nWe developed the following architectures.\n\n* Conv + Transformer (+ LSTM)\n* CNN / U-Net\n* LSTM\n\n## Train/Validation Data\n\nWe basically trained our models using 1-7y LR data, and evaluate with 1/6 subsample of 8y LR data.\nSome model also use:\n\n* 8y LR data\n* 1-8y HR data\n* pseudo labeled old/new test data.\n\n## Evaluation Metrics\n\nWe used 1-MSE instead of R2 after std normalized target.\nWe use this metrics because some member reported R2 loses correlation after host replaced the test dataset.\n(However, later we found both are correlated when using sufficiently small eps for normalization.)\nIn normalization, we used the following formula:\n\n```python\neps = 2.0 ** (-95)\ntarget = target / maximum(std(target), eps)\n```\n\n## Model Design\n\n### Bilzard's Pipeline\n\nBasically 1/7 ~ 1/6 subsample of LR data is used because of relatively *poor* machine resource (RTX4090 x 1).\nWhat worked most was:\n\n* feature extraction of Climate-invariant features (RH, plume buoyancy, normalized Heat Flux etc.)[1]\n* confidence-aware MSE loss\n* doping hard sample of 1/1 full LR data\n\n#### Climate-invariant feature extraction\n\nWe extracted climate-invariant features which is suggested by Tom Beucler et. al. [1].\n\n- Relative Humidity (RH)\n- Plume buoyancy\n- normalized Heat Flux\n\nConceptually, these features are invariant under both cool and warm temperature.\nSo these features are expected to cancel out the effect of yearly warm-up in ClimSim dataset (which I observed by EDA).\n\nDuring the competition, we couldn't train models with this features on the full LR dataset, but it showed a significant gain of ~1% on a 1/7 to 1/6 subsample of LR dataset.\n\nNote that the late submission results shows this features still have a significant gain on the full LR dataset (see Appendix A1).\n\n#### Confidence-aware MSE loss\n\nWe used the following loss which we called *confidence-aware MSE loss* instead of MSE.\nIntuitively it has the effect of loosening the contributions of the large error by hard samples.\nThis idea was variant of the @ryomak 's previously shared solution[2].\nNote that variance of target was regressed by model in addition to mean value. So our model's output variable is 2x larger than basic models.\n\n```python\nloss = 0.5 * (target - pred) ** 2 / var + 0.5 * torch.log(var)\n```\n\n#### Doping Hard Examples\n\nAs we shared on a public discussion[3], most of error came from extreme outliers. So we came to think that hard samples are good representatives of all dataset: if we doped small amount of hard samples from 1/1 full LR data, it would help to reduces regression error even if we used 1/n sub-sampled LR data.\n\nWe extracted ~14% of hard samples which have the squared error greater than threshold using one of our best models, and doped to train other models. We found this has ~0.x% gain using this method.\nWe also found adding tiny amount of neighbor time-frame sample from the same location also helps gaining accuracy.\n\n#### Model Architecture\n\nBasically we used the following two architecture\n\n1. Conv-Transformer (variant of @hoyso48 's architecture[4])\n2. Pixel-Shuffle stacked UNet\n\nThe second architecture was inspired by stacked hourglass architecture[5] which stacks U-Net and gradually improves the accuracy of output variable by regressing error of previous stack (like boosting). In addition, replacing ConvTranspose decoder in the usual U-Net to PixelShuffle gained ~0.x%.\n\n#### Other modeling Tips\n\n##### Tanh normalization\n\nWe often encounter gradient explosion when training deeper transformers (~8-12 layers).\nWe mitigated this issue by utilizing trick we called *tanh normalization*.\nIt is just inserting tanh activation before calculating softmax attention:\n\n```python\n    def forward(self, x):\n        x = rearrange(x, \"b c n -> b n c\")\n\n        qkv = self.qkv(x)\n        qkv = rearrange(qkv, \"b n (h d) -> b h n d\", h=self.num_heads)\n        q, k, v = qkv.chunk(3, dim=-1)\n\n        attn = (q @ k.transpose(-2, -1)) * self.scale\n        attn = attn.tanh() * self.attn_clip_val  # stabilize attention weight\n        attn = F.softmax(attn, dim=-1)\n\n        x = attn @ v\n        x = rearrange(x, \"b h n d -> b n (h d)\")\n\n        x = self.proj(x)\n        x = rearrange(x, \"b n c -> b c n\")\n        return x\n```\n\n##### Removing Dropout\n\nIn the previous competition, some Kagglers reported removing Dropout layer has significant improvement on regression tasks[6]. We observed similar phenomena on this competition.\n\n### Tereka's Pipeline\n\n```markdown\nPublic/Private LB: 0.77784/0.77023\n\nModel\n- Transformer + Convolution + LSTM\n- Architecture Points\n - Learnable Positional Embedding(nn.Embedding 60)\n - several Transformer Layers\n - Final Layer LSTM\n\nFeatures\n- Original Dataset(standard scaler)\n\nDataset\n- 70M(full low resolution dataset)\n\nLoss\n- Huber Loss(> MSE)\n\nNot Worked\n- Mixup\n- EMA\n- Deeper model, I tried some experiment, but it’s not stable.\n- pretrained High resolution model, but low resolution and some high resolution data is worked in public LB.\n```\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F0d32daa11aecc54e4573a294b244ec79%2F21_architecture_tereka.png?generation=1722302269909845&alt=media)\n\n### Phalanx's Pipeline\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F194b6d4d267ae9c693f3fba5ab444cad%2F21_architecture_phalanx.png?generation=1722302291403618&alt=media)\n\n### Ryches' Pipeline\n\nTBD\n\n## Ensemble Method\n\nWe used tree type of ensemble method and final submission is weighted ensemble of these:\n\n- type1) Stacking using models trained with 1-7y low resolution data\n- type2) Stacking using models trained with 1-8y low resolution data\n- type3) Mean ensemble of models trained with 1-8y models\n\nFor type2, basically we can't use the data which trained the seed model to fit parameter in the stacking.\nSo we take the following scheme:\n\n1. calculate the ensemble weight using model trained with 1-7y data\n2. re-train the model of the same architecture, same training condition except swapping train data to 1-8y data\n3. using ensemble weight calculated on step2, but replacing the model from 1-7y to 1-8y data to make the test prediction\n\nFor the final submission, we ensemble the type-1 ~ type-3 model with manually fitted weight:\n\n```python\nfinal_prediction = type1 * w1 + type2 * w2 + type3 * w3\n```\n\nThe late submission result shows assigning 100% weight for type1 model performs the best on the private LB.\n\nModel|LB(Public)|LB(Private)|final submission\n--|--|--|--\ntype1 * 0.5 + type2 * 0.3 + type3 * 0.2 | 0.78803 | 0.78503 | ✔️\ntype1 * 0.75 + type3 * 0.25 | 0.78801 | 0.78487 | ✔️\ntype1 * 0.6 * type2 * 0.4 | 0.78777 | 0.78535 |\ntype1 | 0.78771 | **0.78536** |\n\n## Detail of Stacking\n\nWe trained ensemble weight per (target, model). This is based on the study of EDA where we found each model has their expertize: some models performs good at regressing mixing ratio and some are good at regressing wind velocity etc.\n\nWe tested the following variations and found ensemble without normalization and with bias performs the best.\n\nnormalize weight?|with bias?|LB(Public)|LB(Private)\n--|--|--|--\n✔️|-|0.78727|0.78532\n-|-|0.78739|0.78563\n-|✔️|0.78740|**0.78569**\n\n## Appendix\n\n### A1. LR full data + Climate Invariant Feature\n\nDuring the competition, we only trained our models using climate-invariant features on a subsample of the LR data.However, after the competition ended, we successfully trained Phalanx's lightweight CNN model train model with climate-invariant features on the full sample of the LR dataset in feasible training time (~1 day) with  a relatively \"low-throughput\" GPU (RTX 4090).\n\nThe result of late submission shows it has significant gain of 0.05% on private LB. So climate-invariant feature has actually significant impact on this task.\n\nmodel|CV(1-MSE)|LB(Public)|LB(Private)\n--|--|--|--\nsingle w/o LR-full-clim-invariant|0.78231|0.78025|0.77721\nsingle w LR-full-clim-invariant|**0.78434**|**0.7814**|**0.78017**\nensemble w/o LR-full-clim-invariant|0.79192|0.78740|0.78569\nensemble w LR-full-clim-invariant|**0.79218**|**0.78771**|**0.78618**\n\n## Reference\n\n- [1] [Climate-Invariant Machine Learning](https://arxiv.org/abs/2112.08440)\n- [2] [9th solution(with github code): Laplace distribution noise likelihood optimization](https://www.kaggle.com/competitions/ventilator-pressure-prediction/discussion/285353)\n- [3] [Is Tropical Cylones Key to Win this Competition?](https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/511274)\n- [4] [1st place solution - 1DCNN combined with Transformer](https://www.kaggle.com/competitions/asl-signs/discussion/406684)\n- [5] [Stacked Hourglass Networks for Human Pose Estimation](https://arxiv.org/abs/1603.06937)\n- [6] [The Magic of No Dropout](https://www.kaggle.com/competitions/commonlitreadabilityprize/discussion/260729)",
      "votes": 33
    },
    {
      "id": 2940265,
      "postDate": "2024-07-29T23:02:56.543Z",
      "content": "<p>Good job and congrats on on gold! Can you please elaborate on 'confidence aware' and 'variance prediction'? (I actually also used 'confidence loss' but I'm not sure if you mean the same thing)</p>",
      "rawMarkdown": "Good job and congrats on on gold! Can you please elaborate on 'confidence aware' and 'variance prediction'? (I actually also used 'confidence loss' but I'm not sure if you mean the same thing)",
      "votes": 1,
      "replies": [
        {
          "id": 2940269,
          "postDate": "2024-07-29T23:07:39.163Z",
          "content": "<p>Thanks. I introduced this loss so that the model predicts difficulty of targets. I think it’s close to Huber loss, which other team mates used.</p>\n<ul>\n<li>confidence aware: the loss which aware model's estimation of confidence(=variance) of targets</li>\n<li>variance prediction: the estimate of variance of each targets (just adding another head to predict variance)</li>\n</ul>",
          "rawMarkdown": "Thanks. I introduced this loss so that the model predicts difficulty of targets. I think it’s close to Huber loss, which other team mates used.\n\n* confidence aware: the loss which aware model's estimation of confidence(=variance) of targets\n* variance prediction: the estimate of variance of each targets (just adding another head to predict variance)"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2940265,
      "author_name": "greySnow",
      "author_url": "",
      "post_date": "2024-07-29T23:02:56.543000",
      "content": "<p>Good job and congrats on on gold! Can you please elaborate on 'confidence aware' and 'variance prediction'? (I actually also used 'confidence loss' but I'm not sure if you mean the same thing)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2940269,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2024-07-29T23:07:39.163000",
          "content": "<p>Thanks. I introduced this loss so that the model predicts difficulty of targets. I think it’s close to Huber loss, which other team mates used.</p>\n<ul>\n<li>confidence aware: the loss which aware model's estimation of confidence(=variance) of targets</li>\n<li>variance prediction: the estimate of variance of each targets (just adding another head to predict variance)</li>\n</ul>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2940258": "# 10th Place Solution\n\nThank you for holding this competition. We enjoyed a lot, as well as a lot of leaning experience. This was also really tough competition we ever participated. We literary training the model and making ensemble few hours before the competition closed, and finally we managed to win a gold medal.\n\nThe following is the summary of 10th place solution by team Updraft. More detailed writeup might be shared by each members.\n\n## Team Members\n\n* @tatamikenn (team leader)\n* @tereka\n* @phalanx\n* @ryches\n\n## Glossary\n\n* low resolution data = LR\n* high resolution data = HR\n\n## Summary\n\nOur solution is ensemble of each 4 member's regression models.\nThe models in the final submission contains 23 models with different architectures and learning conditions.\nWe developed the following architectures.\n\n* Conv + Transformer (+ LSTM)\n* CNN / U-Net\n* LSTM\n\n## Train/Validation Data\n\nWe basically trained our models using 1-7y LR data, and evaluate with 1/6 subsample of 8y LR data.\nSome model also use:\n\n* 8y LR data\n* 1-8y HR data\n* pseudo labeled old/new test data.\n\n## Evaluation Metrics\n\nWe used 1-MSE instead of R2 after std normalized target.\nWe use this metrics because some member reported R2 loses correlation after host replaced the test dataset.\n(However, later we found both are correlated when using sufficiently small eps for normalization.)\nIn normalization, we used the following formula:\n\n```python\neps = 2.0 ** (-95)\ntarget = target / maximum(std(target), eps)\n```\n\n## Model Design\n\n### Bilzard's Pipeline\n\nBasically 1/7 ~ 1/6 subsample of LR data is used because of relatively *poor* machine resource (RTX4090 x 1).\nWhat worked most was:\n\n* feature extraction of Climate-invariant features (RH, plume buoyancy, normalized Heat Flux etc.)[1]\n* confidence-aware MSE loss\n* doping hard sample of 1/1 full LR data\n\n#### Climate-invariant feature extraction\n\nWe extracted climate-invariant features which is suggested by Tom Beucler et. al. [1].\n\n- Relative Humidity (RH)\n- Plume buoyancy\n- normalized Heat Flux\n\nConceptually, these features are invariant under both cool and warm temperature.\nSo these features are expected to cancel out the effect of yearly warm-up in ClimSim dataset (which I observed by EDA).\n\nDuring the competition, we couldn't train models with this features on the full LR dataset, but it showed a significant gain of ~1% on a 1/7 to 1/6 subsample of LR dataset.\n\nNote that the late submission results shows this features still have a significant gain on the full LR dataset (see Appendix A1).\n\n#### Confidence-aware MSE loss\n\nWe used the following loss which we called *confidence-aware MSE loss* instead of MSE.\nIntuitively it has the effect of loosening the contributions of the large error by hard samples.\nThis idea was variant of the @ryomak 's previously shared solution[2].\nNote that variance of target was regressed by model in addition to mean value. So our model's output variable is 2x larger than basic models.\n\n```python\nloss = 0.5 * (target - pred) ** 2 / var + 0.5 * torch.log(var)\n```\n\n#### Doping Hard Examples\n\nAs we shared on a public discussion[3], most of error came from extreme outliers. So we came to think that hard samples are good representatives of all dataset: if we doped small amount of hard samples from 1/1 full LR data, it would help to reduces regression error even if we used 1/n sub-sampled LR data.\n\nWe extracted ~14% of hard samples which have the squared error greater than threshold using one of our best models, and doped to train other models. We found this has ~0.x% gain using this method.\nWe also found adding tiny amount of neighbor time-frame sample from the same location also helps gaining accuracy.\n\n#### Model Architecture\n\nBasically we used the following two architecture\n\n1. Conv-Transformer (variant of @hoyso48 's architecture[4])\n2. Pixel-Shuffle stacked UNet\n\nThe second architecture was inspired by stacked hourglass architecture[5] which stacks U-Net and gradually improves the accuracy of output variable by regressing error of previous stack (like boosting). In addition, replacing ConvTranspose decoder in the usual U-Net to PixelShuffle gained ~0.x%.\n\n#### Other modeling Tips\n\n##### Tanh normalization\n\nWe often encounter gradient explosion when training deeper transformers (~8-12 layers).\nWe mitigated this issue by utilizing trick we called *tanh normalization*.\nIt is just inserting tanh activation before calculating softmax attention:\n\n```python\n    def forward(self, x):\n        x = rearrange(x, \"b c n -> b n c\")\n\n        qkv = self.qkv(x)\n        qkv = rearrange(qkv, \"b n (h d) -> b h n d\", h=self.num_heads)\n        q, k, v = qkv.chunk(3, dim=-1)\n\n        attn = (q @ k.transpose(-2, -1)) * self.scale\n        attn = attn.tanh() * self.attn_clip_val  # stabilize attention weight\n        attn = F.softmax(attn, dim=-1)\n\n        x = attn @ v\n        x = rearrange(x, \"b h n d -> b n (h d)\")\n\n        x = self.proj(x)\n        x = rearrange(x, \"b n c -> b c n\")\n        return x\n```\n\n##### Removing Dropout\n\nIn the previous competition, some Kagglers reported removing Dropout layer has significant improvement on regression tasks[6]. We observed similar phenomena on this competition.\n\n### Tereka's Pipeline\n\n```markdown\nPublic/Private LB: 0.77784/0.77023\n\nModel\n- Transformer + Convolution + LSTM\n- Architecture Points\n - Learnable Positional Embedding(nn.Embedding 60)\n - several Transformer Layers\n - Final Layer LSTM\n\nFeatures\n- Original Dataset(standard scaler)\n\nDataset\n- 70M(full low resolution dataset)\n\nLoss\n- Huber Loss(> MSE)\n\nNot Worked\n- Mixup\n- EMA\n- Deeper model, I tried some experiment, but it’s not stable.\n- pretrained High resolution model, but low resolution and some high resolution data is worked in public LB.\n```\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F0d32daa11aecc54e4573a294b244ec79%2F21_architecture_tereka.png?generation=1722302269909845&alt=media)\n\n### Phalanx's Pipeline\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F194b6d4d267ae9c693f3fba5ab444cad%2F21_architecture_phalanx.png?generation=1722302291403618&alt=media)\n\n### Ryches' Pipeline\n\nTBD\n\n## Ensemble Method\n\nWe used tree type of ensemble method and final submission is weighted ensemble of these:\n\n- type1) Stacking using models trained with 1-7y low resolution data\n- type2) Stacking using models trained with 1-8y low resolution data\n- type3) Mean ensemble of models trained with 1-8y models\n\nFor type2, basically we can't use the data which trained the seed model to fit parameter in the stacking.\nSo we take the following scheme:\n\n1. calculate the ensemble weight using model trained with 1-7y data\n2. re-train the model of the same architecture, same training condition except swapping train data to 1-8y data\n3. using ensemble weight calculated on step2, but replacing the model from 1-7y to 1-8y data to make the test prediction\n\nFor the final submission, we ensemble the type-1 ~ type-3 model with manually fitted weight:\n\n```python\nfinal_prediction = type1 * w1 + type2 * w2 + type3 * w3\n```\n\nThe late submission result shows assigning 100% weight for type1 model performs the best on the private LB.\n\nModel|LB(Public)|LB(Private)|final submission\n--|--|--|--\ntype1 * 0.5 + type2 * 0.3 + type3 * 0.2 | 0.78803 | 0.78503 | ✔️\ntype1 * 0.75 + type3 * 0.25 | 0.78801 | 0.78487 | ✔️\ntype1 * 0.6 * type2 * 0.4 | 0.78777 | 0.78535 |\ntype1 | 0.78771 | **0.78536** |\n\n## Detail of Stacking\n\nWe trained ensemble weight per (target, model). This is based on the study of EDA where we found each model has their expertize: some models performs good at regressing mixing ratio and some are good at regressing wind velocity etc.\n\nWe tested the following variations and found ensemble without normalization and with bias performs the best.\n\nnormalize weight?|with bias?|LB(Public)|LB(Private)\n--|--|--|--\n✔️|-|0.78727|0.78532\n-|-|0.78739|0.78563\n-|✔️|0.78740|**0.78569**\n\n## Appendix\n\n### A1. LR full data + Climate Invariant Feature\n\nDuring the competition, we only trained our models using climate-invariant features on a subsample of the LR data.However, after the competition ended, we successfully trained Phalanx's lightweight CNN model train model with climate-invariant features on the full sample of the LR dataset in feasible training time (~1 day) with  a relatively \"low-throughput\" GPU (RTX 4090).\n\nThe result of late submission shows it has significant gain of 0.05% on private LB. So climate-invariant feature has actually significant impact on this task.\n\nmodel|CV(1-MSE)|LB(Public)|LB(Private)\n--|--|--|--\nsingle w/o LR-full-clim-invariant|0.78231|0.78025|0.77721\nsingle w LR-full-clim-invariant|**0.78434**|**0.7814**|**0.78017**\nensemble w/o LR-full-clim-invariant|0.79192|0.78740|0.78569\nensemble w LR-full-clim-invariant|**0.79218**|**0.78771**|**0.78618**\n\n## Reference\n\n- [1] [Climate-Invariant Machine Learning](https://arxiv.org/abs/2112.08440)\n- [2] [9th solution(with github code): Laplace distribution noise likelihood optimization](https://www.kaggle.com/competitions/ventilator-pressure-prediction/discussion/285353)\n- [3] [Is Tropical Cylones Key to Win this Competition?](https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/511274)\n- [4] [1st place solution - 1DCNN combined with Transformer](https://www.kaggle.com/competitions/asl-signs/discussion/406684)\n- [5] [Stacked Hourglass Networks for Human Pose Estimation](https://arxiv.org/abs/1603.06937)\n- [6] [The Magic of No Dropout](https://www.kaggle.com/competitions/commonlitreadabilityprize/discussion/260729)",
    "2940265": "Good job and congrats on on gold! Can you please elaborate on 'confidence aware' and 'variance prediction'? (I actually also used 'confidence loss' but I'm not sure if you mean the same thing)"
  }
}