{
  "id": 523271,
  "title": "[9th place] brief solution",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/523271",
  "author_name": "slime",
  "post_date": "2024-07-31T09:09:59.857000",
  "votes": 20,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Thanks Kaggle and host for organizing this competition, it was a good practice for training sequence models. <br>\nAlso thanks to my teammates <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> and <a href=\"https://www.kaggle.com/nikhilmishradev\" target=\"_blank\">@nikhilmishradev</a>, I learned a lot from you! </p>\n<h3>Model summary</h3>\n<p>We used simple transformer model (vanilla), but with couple of changes, - </p>\n<h4>projection:</h4>\n<p>We first broadcast global features to match the shape of vertical features tensors, also we create some handcrafted features e.g. differences between adjacent timestamps for vertical features.</p>\n<p>We separate embedding projection between features from different so called spaces (we learn projection weights separately for vertical features, global features and hand-crafted features), after we just concat the embedding, apply abs. positional encoding and pass to transformer</p>\n<h4>encoder:</h4>\n<p>Simple vanilla transformer w/ beit-v2 blocks, please refer to the <a href=\"https://github.com/DrHB/icecube-2nd-place\" target=\"_blank\">2nd place icecube solution</a>, thanks <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> and <a href=\"https://www.kaggle.com/drhabib\" target=\"_blank\">@drhabib</a> for publishing the code</p>\n<h4>decoder:</h4>\n<p>Instead of taking the mean of transformer hidden states and applying MLP to get predicitons, we use fno1d layers to predict per-token targets, please refer to the <a href=\"https://github.com/neuraloperator/neuraloperator\" target=\"_blank\">neuraloperator</a> lib, conv decoder should've worked as well</p>\n<p>Num params of the model ~16M, n=8 transformer layers d=320 and m=4 fno1d layers.</p>\n<h3>Training</h3>\n<ul>\n<li>use multiple losses applied to the single head (weighted), MSE and 0.1 * L1, we haven’t tried huberloss, should perform as well</li>\n<li>AdamW with 4e-3, batch_size=4096, train for 5-7 epochs on the low-res dataset</li>\n</ul>\n<h3>Dataset</h3>\n<p>One of the main challenges of this competition was optimizing the code in way that ensures fast random reading from 80M samples dataset that takes 700Gb~ space. </p>\n<p>We used .hdf5 files for that, - we first create .hdf5 file for each month and during training, in each worker we open a separate instance of .h5py file for each month. </p>\n<p>It was improtant to have a fast reading drive (e.g. M2 NVME, simple SATA SSD didn’t work) in order to achieve good GPU utilization, with this setup whole training on 8x4090 took around 100Gb~ of RAM in DDP setting, which is a reasonable amount for such machine.</p>\n<h3>Normalization</h3>\n<p>We collected statistics from original kaggle data (mean/std) for both features and targets multipled by sample_sub weights and applied mean/std normalization whenever we trained our models (that be kaggle data or full low-res dataset)</p>\n<h3>Notes</h3>\n<ul>\n<li>It was important to use bf16/fp32 instead of fp16 training due to wide range of features even after normalization</li>\n<li><code>transformer-engine</code> library is a very good thing if you want to train vanilla transformer (like here, when you don't need custom things such as learnable attention bias ref. recent Ribonanza competition)<br>\nIt takes way less memory than eager pytorch implementation allowing for larger batch sizes and generally we observed 30% speedup compared to vanilla transformer with flashattention implemented using eager torch, to run it we used latest <code>nvcr.io/nvidia/pytorch:24.05-py3</code> docker image</li>\n</ul>\n<h3>Tried, but didn't work</h3>\n<p>Target was <code>(x(t+1) - x(t))/1200</code>, </p>\n<ul>\n<li>predict <code>x(t+1)</code> directly, reconstruct original target later</li>\n<li>autoencoder pretrain (given t0, predict x(t0) when doing continious MLM)</li>\n<li>train diffusion transformer, the loss was good, but the reconstruction quality suffered</li>\n</ul>\n<h3>code:</h3>\n<p>TBA</p>",
  "messages": [
    {
      "id": 2941753,
      "postDate": "2024-07-31T09:09:59.857Z",
      "content": "<p>Thanks Kaggle and host for organizing this competition, it was a good practice for training sequence models. <br>\nAlso thanks to my teammates <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> and <a href=\"https://www.kaggle.com/nikhilmishradev\" target=\"_blank\">@nikhilmishradev</a>, I learned a lot from you! </p>\n<h3>Model summary</h3>\n<p>We used simple transformer model (vanilla), but with couple of changes, - </p>\n<h4>projection:</h4>\n<p>We first broadcast global features to match the shape of vertical features tensors, also we create some handcrafted features e.g. differences between adjacent timestamps for vertical features.</p>\n<p>We separate embedding projection between features from different so called spaces (we learn projection weights separately for vertical features, global features and hand-crafted features), after we just concat the embedding, apply abs. positional encoding and pass to transformer</p>\n<h4>encoder:</h4>\n<p>Simple vanilla transformer w/ beit-v2 blocks, please refer to the <a href=\"https://github.com/DrHB/icecube-2nd-place\" target=\"_blank\">2nd place icecube solution</a>, thanks <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> and <a href=\"https://www.kaggle.com/drhabib\" target=\"_blank\">@drhabib</a> for publishing the code</p>\n<h4>decoder:</h4>\n<p>Instead of taking the mean of transformer hidden states and applying MLP to get predicitons, we use fno1d layers to predict per-token targets, please refer to the <a href=\"https://github.com/neuraloperator/neuraloperator\" target=\"_blank\">neuraloperator</a> lib, conv decoder should've worked as well</p>\n<p>Num params of the model ~16M, n=8 transformer layers d=320 and m=4 fno1d layers.</p>\n<h3>Training</h3>\n<ul>\n<li>use multiple losses applied to the single head (weighted), MSE and 0.1 * L1, we haven’t tried huberloss, should perform as well</li>\n<li>AdamW with 4e-3, batch_size=4096, train for 5-7 epochs on the low-res dataset</li>\n</ul>\n<h3>Dataset</h3>\n<p>One of the main challenges of this competition was optimizing the code in way that ensures fast random reading from 80M samples dataset that takes 700Gb~ space. </p>\n<p>We used .hdf5 files for that, - we first create .hdf5 file for each month and during training, in each worker we open a separate instance of .h5py file for each month. </p>\n<p>It was improtant to have a fast reading drive (e.g. M2 NVME, simple SATA SSD didn’t work) in order to achieve good GPU utilization, with this setup whole training on 8x4090 took around 100Gb~ of RAM in DDP setting, which is a reasonable amount for such machine.</p>\n<h3>Normalization</h3>\n<p>We collected statistics from original kaggle data (mean/std) for both features and targets multipled by sample_sub weights and applied mean/std normalization whenever we trained our models (that be kaggle data or full low-res dataset)</p>\n<h3>Notes</h3>\n<ul>\n<li>It was important to use bf16/fp32 instead of fp16 training due to wide range of features even after normalization</li>\n<li><code>transformer-engine</code> library is a very good thing if you want to train vanilla transformer (like here, when you don't need custom things such as learnable attention bias ref. recent Ribonanza competition)<br>\nIt takes way less memory than eager pytorch implementation allowing for larger batch sizes and generally we observed 30% speedup compared to vanilla transformer with flashattention implemented using eager torch, to run it we used latest <code>nvcr.io/nvidia/pytorch:24.05-py3</code> docker image</li>\n</ul>\n<h3>Tried, but didn't work</h3>\n<p>Target was <code>(x(t+1) - x(t))/1200</code>, </p>\n<ul>\n<li>predict <code>x(t+1)</code> directly, reconstruct original target later</li>\n<li>autoencoder pretrain (given t0, predict x(t0) when doing continious MLM)</li>\n<li>train diffusion transformer, the loss was good, but the reconstruction quality suffered</li>\n</ul>\n<h3>code:</h3>\n<p>TBA</p>",
      "rawMarkdown": "Thanks Kaggle and host for organizing this competition, it was a good practice for training sequence models. \nAlso thanks to my teammates @nischaydnk and @nikhilmishradev, I learned a lot from you! \n\n### Model summary\n\nWe used simple transformer model (vanilla), but with couple of changes, - \n\n#### projection:\nWe first broadcast global features to match the shape of vertical features tensors, also we create some handcrafted features e.g. differences between adjacent timestamps for vertical features.\n\nWe separate embedding projection between features from different so called spaces (we learn projection weights separately for vertical features, global features and hand-crafted features), after we just concat the embedding, apply abs. positional encoding and pass to transformer\n\n#### encoder:\n\nSimple vanilla transformer w/ beit-v2 blocks, please refer to the [2nd place icecube solution](https://github.com/DrHB/icecube-2nd-place), thanks @iafoss and @drhabib for publishing the code\n\n#### decoder:\n\nInstead of taking the mean of transformer hidden states and applying MLP to get predicitons, we use fno1d layers to predict per-token targets, please refer to the [neuraloperator](https://github.com/neuraloperator/neuraloperator) lib, conv decoder should've worked as well\n\nNum params of the model ~16M, n=8 transformer layers d=320 and m=4 fno1d layers.\n\n### Training\n\n- use multiple losses applied to the single head (weighted), MSE and 0.1 * L1, we haven’t tried huberloss, should perform as well\n- AdamW with 4e-3, batch_size=4096, train for 5-7 epochs on the low-res dataset\n\n### Dataset\n\nOne of the main challenges of this competition was optimizing the code in way that ensures fast random reading from 80M samples dataset that takes 700Gb~ space. \n\nWe used .hdf5 files for that, - we first create .hdf5 file for each month and during training, in each worker we open a separate instance of .h5py file for each month. \n\nIt was improtant to have a fast reading drive (e.g. M2 NVME, simple SATA SSD didn’t work) in order to achieve good GPU utilization, with this setup whole training on 8x4090 took around 100Gb~ of RAM in DDP setting, which is a reasonable amount for such machine.\n\n### Normalization\nWe collected statistics from original kaggle data (mean/std) for both features and targets multipled by sample_sub weights and applied mean/std normalization whenever we trained our models (that be kaggle data or full low-res dataset)\n\n### Notes\n\n- It was important to use bf16/fp32 instead of fp16 training due to wide range of features even after normalization\n- `transformer-engine` library is a very good thing if you want to train vanilla transformer (like here, when you don't need custom things such as learnable attention bias ref. recent Ribonanza competition)\nIt takes way less memory than eager pytorch implementation allowing for larger batch sizes and generally we observed 30% speedup compared to vanilla transformer with flashattention implemented using eager torch, to run it we used latest `nvcr.io/nvidia/pytorch:24.05-py3` docker image\n\n### Tried, but didn't work\nTarget was `(x(t+1) - x(t))/1200`, \n\n- predict `x(t+1)` directly, reconstruct original target later\n- autoencoder pretrain (given t0, predict x(t0) when doing continious MLM)\n- train diffusion transformer, the loss was good, but the reconstruction quality suffered\n\n### code:\nTBA",
      "votes": 19
    },
    {
      "id": 2945014,
      "postDate": "2024-08-03T01:00:14.483Z",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/martynoveduard\" target=\"_blank\">@martynoveduard</a> , thank you for publishing the solution and Congratulation to the Team for the Gold Medal !!</p>\n<p>I had a question, as you have mentioned that it was important to use bf16/fp32 instead of fp16, even after normalizing the data first… How big of an impact did you observe on the score because of it?</p>",
      "rawMarkdown": "Hey @martynoveduard , thank you for publishing the solution and Congratulation to the Team for the Gold Medal !!\n\nI had a question, as you have mentioned that it was important to use bf16/fp32 instead of fp16, even after normalizing the data first... How big of an impact did you observe on the score because of it?",
      "replies": [
        {
          "id": 2945082,
          "postDate": "2024-08-03T03:58:44.490Z",
          "content": "<p>Thank you! For some old single model I observed quite a difference, - </p>\n<table>\n<thead>\n<tr>\n<th>training precision</th>\n<th>public score</th>\n<th>private score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>bf16</td>\n<td>0.7796</td>\n<td>0.7766</td>\n</tr>\n<tr>\n<td>fp16</td>\n<td>0.7703</td>\n<td>0.7664</td>\n</tr>\n</tbody>\n</table>",
          "rawMarkdown": "Thank you! For some old single model I observed quite a difference, - \n|training precision | public score | private score |\n| --- | --- | --- |\n| bf16 | 0.7796 | 0.7766 |\n| fp16 | 0.7703| 0.7664 |  \n",
          "votes": 2
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2945014,
      "author_name": "Icees8",
      "author_url": "",
      "post_date": "2024-08-03T01:00:14.483000",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/martynoveduard\" target=\"_blank\">@martynoveduard</a> , thank you for publishing the solution and Congratulation to the Team for the Gold Medal !!</p>\n<p>I had a question, as you have mentioned that it was important to use bf16/fp32 instead of fp16, even after normalizing the data first… How big of an impact did you observe on the score because of it?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2945082,
          "author_name": "slime",
          "author_url": "",
          "post_date": "2024-08-03T03:58:44.490000",
          "content": "<p>Thank you! For some old single model I observed quite a difference, - </p>\n<table>\n<thead>\n<tr>\n<th>training precision</th>\n<th>public score</th>\n<th>private score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>bf16</td>\n<td>0.7796</td>\n<td>0.7766</td>\n</tr>\n<tr>\n<td>fp16</td>\n<td>0.7703</td>\n<td>0.7664</td>\n</tr>\n</tbody>\n</table>",
          "votes": 2,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2941753": "Thanks Kaggle and host for organizing this competition, it was a good practice for training sequence models. \nAlso thanks to my teammates @nischaydnk and @nikhilmishradev, I learned a lot from you! \n\n### Model summary\n\nWe used simple transformer model (vanilla), but with couple of changes, - \n\n#### projection:\nWe first broadcast global features to match the shape of vertical features tensors, also we create some handcrafted features e.g. differences between adjacent timestamps for vertical features.\n\nWe separate embedding projection between features from different so called spaces (we learn projection weights separately for vertical features, global features and hand-crafted features), after we just concat the embedding, apply abs. positional encoding and pass to transformer\n\n#### encoder:\n\nSimple vanilla transformer w/ beit-v2 blocks, please refer to the [2nd place icecube solution](https://github.com/DrHB/icecube-2nd-place), thanks @iafoss and @drhabib for publishing the code\n\n#### decoder:\n\nInstead of taking the mean of transformer hidden states and applying MLP to get predicitons, we use fno1d layers to predict per-token targets, please refer to the [neuraloperator](https://github.com/neuraloperator/neuraloperator) lib, conv decoder should've worked as well\n\nNum params of the model ~16M, n=8 transformer layers d=320 and m=4 fno1d layers.\n\n### Training\n\n- use multiple losses applied to the single head (weighted), MSE and 0.1 * L1, we haven’t tried huberloss, should perform as well\n- AdamW with 4e-3, batch_size=4096, train for 5-7 epochs on the low-res dataset\n\n### Dataset\n\nOne of the main challenges of this competition was optimizing the code in way that ensures fast random reading from 80M samples dataset that takes 700Gb~ space. \n\nWe used .hdf5 files for that, - we first create .hdf5 file for each month and during training, in each worker we open a separate instance of .h5py file for each month. \n\nIt was improtant to have a fast reading drive (e.g. M2 NVME, simple SATA SSD didn’t work) in order to achieve good GPU utilization, with this setup whole training on 8x4090 took around 100Gb~ of RAM in DDP setting, which is a reasonable amount for such machine.\n\n### Normalization\nWe collected statistics from original kaggle data (mean/std) for both features and targets multipled by sample_sub weights and applied mean/std normalization whenever we trained our models (that be kaggle data or full low-res dataset)\n\n### Notes\n\n- It was important to use bf16/fp32 instead of fp16 training due to wide range of features even after normalization\n- `transformer-engine` library is a very good thing if you want to train vanilla transformer (like here, when you don't need custom things such as learnable attention bias ref. recent Ribonanza competition)\nIt takes way less memory than eager pytorch implementation allowing for larger batch sizes and generally we observed 30% speedup compared to vanilla transformer with flashattention implemented using eager torch, to run it we used latest `nvcr.io/nvidia/pytorch:24.05-py3` docker image\n\n### Tried, but didn't work\nTarget was `(x(t+1) - x(t))/1200`, \n\n- predict `x(t+1)` directly, reconstruct original target later\n- autoencoder pretrain (given t0, predict x(t0) when doing continious MLM)\n- train diffusion transformer, the loss was good, but the reconstruction quality suffered\n\n### code:\nTBA",
    "2945014": "Hey @martynoveduard , thank you for publishing the solution and Congratulation to the Team for the Gold Medal !!\n\nI had a question, as you have mentioned that it was important to use bf16/fp32 instead of fp16, even after normalizing the data first... How big of an impact did you observe on the score because of it?"
  }
}