{
  "id": 506984,
  "title": "Two tips to get LB0.7+",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/506984",
  "author_name": "phalanx",
  "post_date": "2024-05-24T02:46:33.428000",
  "votes": 109,
  "comment_count": 28,
  "views": 0,
  "content": "<p>It seems that there are few participants with an LB score above 0.7, and many participants are struggling in this competition. Considering lowering the entry barrier, I thought about providing starter code. However, since there is limited knowledge about this competition and fairness as a competitive event is maintained,  I'd like to limit our support to sharing simple tips.<br>\nBy considering just the following two points, you should be able to easily achieve a score above 0.7.</p>\n<h2>1. Pay attention to the data type of floating-point numbers</h2>\n<p>The provided data is in float64. As shared in <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/500816\" target=\"_blank\">this discussion</a>, converting to float32 poses a risk of underflow. You must not convert to float32.</p>\n<h2>2. Model architecture</h2>\n<p>Sharing detailed information while maintaining fairness is difficult, so I will limit it to a few tips. EDA and domain knowledge are extremely important. Methods from past competitions may not be useful (at least they are not useful for me). By refining my model architecture, I achieved CV 0.72+ from the first epoch.</p>\n<p>By focusing solely on improving the model architecture, I was able to achieve LB 0.766. I did not use pseudo labeling or add any additional training data. Since the model is extremely important, please review your data again and gather domain knowledge. <br>\nHave fun experimenting! Good luck!</p>",
  "messages": [
    {
      "id": 2832988,
      "postDate": "2024-05-24T02:46:33.427Z",
      "content": "<p>It seems that there are few participants with an LB score above 0.7, and many participants are struggling in this competition. Considering lowering the entry barrier, I thought about providing starter code. However, since there is limited knowledge about this competition and fairness as a competitive event is maintained,  I'd like to limit our support to sharing simple tips.<br>\nBy considering just the following two points, you should be able to easily achieve a score above 0.7.</p>\n<h2>1. Pay attention to the data type of floating-point numbers</h2>\n<p>The provided data is in float64. As shared in <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/500816\" target=\"_blank\">this discussion</a>, converting to float32 poses a risk of underflow. You must not convert to float32.</p>\n<h2>2. Model architecture</h2>\n<p>Sharing detailed information while maintaining fairness is difficult, so I will limit it to a few tips. EDA and domain knowledge are extremely important. Methods from past competitions may not be useful (at least they are not useful for me). By refining my model architecture, I achieved CV 0.72+ from the first epoch.</p>\n<p>By focusing solely on improving the model architecture, I was able to achieve LB 0.766. I did not use pseudo labeling or add any additional training data. Since the model is extremely important, please review your data again and gather domain knowledge. <br>\nHave fun experimenting! Good luck!</p>",
      "rawMarkdown": "It seems that there are few participants with an LB score above 0.7, and many participants are struggling in this competition. Considering lowering the entry barrier, I thought about providing starter code. However, since there is limited knowledge about this competition and fairness as a competitive event is maintained,  I'd like to limit our support to sharing simple tips.\nBy considering just the following two points, you should be able to easily achieve a score above 0.7.\n\n## 1. Pay attention to the data type of floating-point numbers\nThe provided data is in float64. As shared in [this discussion](https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/500816), converting to float32 poses a risk of underflow. You must not convert to float32.\n\n## 2. Model architecture\nSharing detailed information while maintaining fairness is difficult, so I will limit it to a few tips. EDA and domain knowledge are extremely important. Methods from past competitions may not be useful (at least they are not useful for me). By refining my model architecture, I achieved CV 0.72+ from the first epoch.\n\n\n\nBy focusing solely on improving the model architecture, I was able to achieve LB 0.766. I did not use pseudo labeling or add any additional training data. Since the model is extremely important, please review your data again and gather domain knowledge. \nHave fun experimenting! Good luck!",
      "votes": 108
    },
    {
      "id": 2833069,
      "postDate": "2024-05-24T04:14:09.160Z",
      "content": "<p>Thanks for sharing your strategy. Have you tried training with both float 32 and 64 and check the difference?</p>",
      "rawMarkdown": "Thanks for sharing your strategy. Have you tried training with both float 32 and 64 and check the difference?",
      "votes": 6,
      "replies": [
        {
          "id": 2833454,
          "postDate": "2024-05-24T07:53:51.353Z",
          "content": "<p>Interested in this as well. By extension, is AMP (mixed precision) useless in this comp?</p>\n<p>Let me share my current fumbles:</p>\n<ul>\n<li>Model: transformer encoder (tried 1m-6m params)</li>\n<li>Data: F64 -&gt; rescale -&gt; downcast -&gt; B/F16</li>\n<li>Sub: Pred -&gt; unscale -&gt; multiply by sample weight</li>\n<li>CV: ~0.5 LB: trash</li>\n<li>I think I need to do some post-processing (if mixed precision is not bunk)</li>\n</ul>\n<p>Edit: For anyone looking for more tips, post-processing is important (to get right).</p>",
          "rawMarkdown": "Interested in this as well. By extension, is AMP (mixed precision) useless in this comp?\n\nLet me share my current fumbles:\n- Model: transformer encoder (tried 1m-6m params)\n- Data: F64 -> rescale -> downcast -> B/F16\n- Sub: Pred -> unscale -> multiply by sample weight\n- CV: ~0.5 LB: trash\n- I think I need to do some post-processing (if mixed precision is not bunk)\n\nEdit: For anyone looking for more tips, post-processing is important (to get right).",
          "votes": 4,
          "replies": [
            {
              "id": 2837487,
              "postDate": "2024-05-26T14:28:52.907Z",
              "content": "<p>Sorry for late reply. <br>\nI believe that float64 is sufficient for preprocessing, so I have only tried training with float32. <br>\nAdditionally, since I only have a 3090 GPU, I don't have the capacity to try anything else.</p>",
              "rawMarkdown": "Sorry for late reply. \nI believe that float64 is sufficient for preprocessing, so I have only tried training with float32. \nAdditionally, since I only have a 3090 GPU, I don't have the capacity to try anything else.",
              "votes": 6
            },
            {
              "id": 2838649,
              "postDate": "2024-05-27T06:49:23.427Z",
              "content": "<p>Gotcha, thanks for the tips.<br>\nNice to have confirmation that downcasting after rescaling works (well)!</p>",
              "rawMarkdown": "Gotcha, thanks for the tips.\nNice to have confirmation that downcasting after rescaling works (well)!",
              "votes": 1
            },
            {
              "id": 2916937,
              "postDate": "2024-07-11T11:16:22.343Z",
              "content": "<p>How much memory does your GPU have to run out of code?</p>",
              "rawMarkdown": "How much memory does your GPU have to run out of code?",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2835512,
      "postDate": "2024-05-25T10:56:49.283Z",
      "content": "<p>Any suggested material to get some domain knowledge?</p>",
      "rawMarkdown": "Any suggested material to get some domain knowledge?",
      "votes": 3,
      "replies": [
        {
          "id": 2836186,
          "postDate": "2024-05-25T18:04:04.910Z",
          "content": "<p>I tried to find the papers where the implemented equations are shown. Unfortunately it is not so easy. Atmospheric codes are solving partial differential equations in the form d phi / dt = F(phi,x,t). d phi / dt are the tendencies and F(phi,x,t) is a non linear function of the variables (or features) like u,v,t which are given. So I try to do some feature engineering to reproduce the funktional relations F(phi). Unfortunately the time t and location x are not given. But i guess this is intentional since the derived machine learning model should be independent of the location and time in order to generalize well. </p>",
          "rawMarkdown": "I tried to find the papers where the implemented equations are shown. Unfortunately it is not so easy. Atmospheric codes are solving partial differential equations in the form d phi / dt = F(phi,x,t). d phi / dt are the tendencies and F(phi,x,t) is a non linear function of the variables (or features) like u,v,t which are given. So I try to do some feature engineering to reproduce the funktional relations F(phi). Unfortunately the time t and location x are not given. But i guess this is intentional since the derived machine learning model should be independent of the location and time in order to generalize well. ",
          "votes": 5,
          "replies": [
            {
              "id": 2836199,
              "postDate": "2024-05-25T18:09:16.250Z",
              "content": "<p>What it one has also to mention is that variables at pressure levels which are far away (e.g. label 59 and label 40) are not very correlated. It is also very unlikely the velocity at label 40 will influence the tendency of the velocity at label 59</p>",
              "rawMarkdown": "What it one has also to mention is that variables at pressure levels which are far away (e.g. label 59 and label 40) are not very correlated. It is also very unlikely the velocity at label 40 will influence the tendency of the velocity at label 59"
            }
          ]
        }
      ]
    },
    {
      "id": 2865824,
      "postDate": "2024-06-11T02:50:59.613Z",
      "content": "<p>Thank you for your great knowledge!<br>\nI wonder why methods from past competitions might not be useful.<br>\nIs there any fundamental difference compared to other competitions?</p>",
      "rawMarkdown": "Thank you for your great knowledge!\nI wonder why methods from past competitions might not be useful.\nIs there any fundamental difference compared to other competitions?",
      "votes": 1
    },
    {
      "id": 2856969,
      "postDate": "2024-06-05T16:29:20.040Z",
      "content": "<p>Hey phalanx, when you achieved 0.72 from the first epoch with what bach size? or how many training iterations ?</p>",
      "rawMarkdown": "Hey phalanx, when you achieved 0.72 from the first epoch with what bach size? or how many training iterations ?",
      "votes": 2
    },
    {
      "id": 2909809,
      "postDate": "2024-07-07T10:39:25.527Z",
      "content": "<p>Try Unet with attention block before the bottle neck</p>",
      "rawMarkdown": "Try Unet with attention block before the bottle neck"
    },
    {
      "id": 2853321,
      "postDate": "2024-06-03T18:06:21.337Z",
      "content": "<p>how the heck are people getting 0.72!? I guess my models are pretty basic, I am getting 0.57 with a very simple MLP.</p>",
      "rawMarkdown": "how the heck are people getting 0.72!? I guess my models are pretty basic, I am getting 0.57 with a very simple MLP."
    },
    {
      "id": 2839420,
      "postDate": "2024-05-27T14:13:55.240Z",
      "content": "<p>Thanks for sharing.</p>\n<p>I have one more question, in terms of NN architecture, which path do you think is best to start with? Seq2Seq as discussed in another thread, or is it possible to achieve 0.7+ with FFNN?</p>",
      "rawMarkdown": "Thanks for sharing.\n\nI have one more question, in terms of NN architecture, which path do you think is best to start with? Seq2Seq as discussed in another thread, or is it possible to achieve 0.7+ with FFNN?",
      "replies": [
        {
          "id": 2840302,
          "postDate": "2024-05-28T04:30:46.220Z",
          "content": "<p>I am confident of getting to 0.7+ using MLP, but I think mayde need to train with some techniques based on EDA</p>",
          "rawMarkdown": "I am confident of getting to 0.7+ using MLP, but I think mayde need to train with some techniques based on EDA",
          "votes": 3,
          "replies": [
            {
              "id": 2856457,
              "postDate": "2024-06-05T09:55:01.927Z",
              "content": "<p>which are these techniques?</p>",
              "rawMarkdown": "which are these techniques?"
            },
            {
              "id": 2856980,
              "postDate": "2024-06-05T16:45:04.350Z",
              "content": "<p>\"Did you ever hear the training of Darth Perceptron the Wide?\"</p>",
              "rawMarkdown": "\"Did you ever hear the training of Darth Perceptron the Wide?\"",
              "votes": 4
            }
          ]
        }
      ]
    },
    {
      "id": 2839068,
      "postDate": "2024-05-27T11:23:08.620Z",
      "content": "<p><a href=\"https://www.kaggle.com/muhammadawaistayyab\" target=\"_blank\">@muhammadawaistayyab</a> </p>",
      "rawMarkdown": "@muhammadawaistayyab "
    },
    {
      "id": 2837131,
      "postDate": "2024-05-26T09:39:56.940Z",
      "content": "<p>i have changed from float32 to float64 and i have reduced my batch_size from 2048 to 1024(because when i use float64 i was left out of memory) and now 100 steps of transformer from ~700 seconds takes 1560 … does this make sense?</p>",
      "rawMarkdown": "i have changed from float32 to float64 and i have reduced my batch_size from 2048 to 1024(because when i use float64 i was left out of memory) and now 100 steps of transformer from ~700 seconds takes 1560 ... does this make sense?",
      "replies": [
        {
          "id": 2837182,
          "postDate": "2024-05-26T10:43:21.533Z",
          "content": "<p>so i asked chatGPT and it seems very obvious that the train time went so up :(</p>",
          "rawMarkdown": "so i asked chatGPT and it seems very obvious that the train time went so up :(",
          "votes": 1
        },
        {
          "id": 2837243,
          "postDate": "2024-05-26T11:39:55.077Z",
          "content": "<p>Suggestion: float64 -&gt; normalize -&gt; float32 -&gt; train</p>",
          "rawMarkdown": "Suggestion: float64 -> normalize -> float32 -> train",
          "votes": 9,
          "replies": [
            {
              "id": 2838950,
              "postDate": "2024-05-27T09:41:08.877Z",
              "content": "<p>How did you do that, when i try that i get underflows, the loss functions sky rockets randomly so i need to train at float64. This happens only when i multiply y with the weights before training. If i dont multiply with the weights then i dont have underflows but i have red in the discussions that if you multiply with the weights you have an improvement of ~0.1</p>\n<p>Actually i noticed that even with float64 i have the same phenomenon, any idea why?<br>\nepoch: 0 , chunk: 0, Step 100, Training Loss: 0.5917 iterations: 100 time: 2783.370905160904<br>\nepoch: 0 , chunk: 0, Step 200, Training Loss: 0.5177 iterations: 200 time: 3947.4694044589996<br>\nepoch: 0 , chunk: 0, Step 300, Training Loss: 0.5088 iterations: 300 time: 5116.634320020676<br>\nepoch: 0 , chunk: 0, Step 400, Training Loss: 0.5138 iterations: 400 time: 6274.054301977158<br>\nepoch: 0 model saved chunk: 0 iterations: 486 val loss: 0.5341707844085616 time: 3782.210193157196</p>\n<p>tensor_data input shape: torch.Size([497123, 60, 25])<br>\nepoch: 0 chunk: 1 prep chunk time: 15.340328693389893<br>\nepoch: 0 , chunk: 1, Step 100, Training Loss: 0.4475 iterations: 586 time: 2813.811069250107<br>\nepoch: 0 , chunk: 1, Step 200, Training Loss: 0.4961 iterations: 686 time: 6220.386344671249<br>\nepoch: 0 , chunk: 1, Step 300, Training Loss: 0.4782 iterations: 786 time: 7684.73356628418<br>\nepoch: 0 , chunk: 1, Step 400, Training Loss: 1.3982 iterations: 886 time: 9230.30635881424</p>\n<p>The train loss spikes, this happens only when i multiply the y with the weights before training.</p>",
              "rawMarkdown": "How did you do that, when i try that i get underflows, the loss functions sky rockets randomly so i need to train at float64. This happens only when i multiply y with the weights before training. If i dont multiply with the weights then i dont have underflows but i have red in the discussions that if you multiply with the weights you have an improvement of ~0.1\n\nActually i noticed that even with float64 i have the same phenomenon, any idea why?\nepoch: 0 , chunk: 0, Step 100, Training Loss: 0.5917 iterations: 100 time: 2783.370905160904\nepoch: 0 , chunk: 0, Step 200, Training Loss: 0.5177 iterations: 200 time: 3947.4694044589996\nepoch: 0 , chunk: 0, Step 300, Training Loss: 0.5088 iterations: 300 time: 5116.634320020676\nepoch: 0 , chunk: 0, Step 400, Training Loss: 0.5138 iterations: 400 time: 6274.054301977158\nepoch: 0 model saved chunk: 0 iterations: 486 val loss: 0.5341707844085616 time: 3782.210193157196\n\ntensor_data input shape: torch.Size([497123, 60, 25])\nepoch: 0 chunk: 1 prep chunk time: 15.340328693389893\nepoch: 0 , chunk: 1, Step 100, Training Loss: 0.4475 iterations: 586 time: 2813.811069250107\nepoch: 0 , chunk: 1, Step 200, Training Loss: 0.4961 iterations: 686 time: 6220.386344671249\nepoch: 0 , chunk: 1, Step 300, Training Loss: 0.4782 iterations: 786 time: 7684.73356628418\nepoch: 0 , chunk: 1, Step 400, Training Loss: 1.3982 iterations: 886 time: 9230.30635881424\n\nThe train loss spikes, this happens only when i multiply the y with the weights before training."
            },
            {
              "id": 2839138,
              "postDate": "2024-05-27T11:56:52.490Z",
              "content": "<p>I have a baseline notebook where I directly train in float32 (also the normalization is in float32). My public score is also using the exact same logic, so it is possibile to get to at least 0.75 training in float32 without pre-normalizing. I didn't try to multiply by the weights</p>",
              "rawMarkdown": "I have a baseline notebook where I directly train in float32 (also the normalization is in float32). My public score is also using the exact same logic, so it is possibile to get to at least 0.75 training in float32 without pre-normalizing. I didn't try to multiply by the weights",
              "votes": 5
            },
            {
              "id": 2839495,
              "postDate": "2024-05-27T14:57:27.603Z",
              "content": "<p>Probably it generate some big gradient, I would try to clip the grad norm</p>",
              "rawMarkdown": "Probably it generate some big gradient, I would try to clip the grad norm",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2834706,
      "postDate": "2024-05-25T00:14:06.407Z",
      "content": "<p>Great sharing</p>",
      "rawMarkdown": "Great sharing"
    },
    {
      "id": 2834127,
      "postDate": "2024-05-24T15:03:12.090Z",
      "content": "<p>\" By refining my model architecture, I achieved CV 0.72+ from the first epoch.\"<br>\nIncredible! Congrats :)</p>",
      "rawMarkdown": "\" By refining my model architecture, I achieved CV 0.72+ from the first epoch.\"\nIncredible! Congrats :)",
      "replies": [
        {
          "id": 2874995,
          "postDate": "2024-06-16T18:35:36.520Z",
          "content": "<p>it seemed so impressive at the time, but now even i managed to get to CV of 0.726 after the first epoch . Only took me a month!</p>",
          "rawMarkdown": "it seemed so impressive at the time, but now even i managed to get to CV of 0.726 after the first epoch . Only took me a month!",
          "votes": 1
        }
      ]
    },
    {
      "id": 2837393,
      "postDate": "2024-05-26T13:04:36.287Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2837869,
      "postDate": "2024-05-26T17:50:45.400Z",
      "content": "<p>Thank you very much </p>",
      "rawMarkdown": "Thank you very much "
    }
  ],
  "comments": [
    {
      "id": 2833069,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2024-05-24T04:14:09.160000",
      "content": "<p>Thanks for sharing your strategy. Have you tried training with both float 32 and 64 and check the difference?</p>",
      "votes": 6,
      "replies": [
        {
          "id": 2833454,
          "author_name": "sroger",
          "author_url": "",
          "post_date": "2024-05-24T07:53:51.353000",
          "content": "<p>Interested in this as well. By extension, is AMP (mixed precision) useless in this comp?</p>\n<p>Let me share my current fumbles:</p>\n<ul>\n<li>Model: transformer encoder (tried 1m-6m params)</li>\n<li>Data: F64 -&gt; rescale -&gt; downcast -&gt; B/F16</li>\n<li>Sub: Pred -&gt; unscale -&gt; multiply by sample weight</li>\n<li>CV: ~0.5 LB: trash</li>\n<li>I think I need to do some post-processing (if mixed precision is not bunk)</li>\n</ul>\n<p>Edit: For anyone looking for more tips, post-processing is important (to get right).</p>",
          "votes": 4,
          "replies": [
            {
              "id": 2837487,
              "author_name": "phalanx",
              "author_url": "",
              "post_date": "2024-05-26T14:28:52.907000",
              "content": "<p>Sorry for late reply. <br>\nI believe that float64 is sufficient for preprocessing, so I have only tried training with float32. <br>\nAdditionally, since I only have a 3090 GPU, I don't have the capacity to try anything else.</p>",
              "votes": 6,
              "replies": []
            },
            {
              "id": 2838649,
              "author_name": "sroger",
              "author_url": "",
              "post_date": "2024-05-27T06:49:23.427000",
              "content": "<p>Gotcha, thanks for the tips.<br>\nNice to have confirmation that downcasting after rescaling works (well)!</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2916937,
              "author_name": "xj",
              "author_url": "",
              "post_date": "2024-07-11T11:16:22.343000",
              "content": "<p>How much memory does your GPU have to run out of code?</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2835512,
      "author_name": "Amedeo Biolatti",
      "author_url": "",
      "post_date": "2024-05-25T10:56:49.283000",
      "content": "<p>Any suggested material to get some domain knowledge?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2836186,
          "author_name": "HappyKiter",
          "author_url": "",
          "post_date": "2024-05-25T18:04:04.910000",
          "content": "<p>I tried to find the papers where the implemented equations are shown. Unfortunately it is not so easy. Atmospheric codes are solving partial differential equations in the form d phi / dt = F(phi,x,t). d phi / dt are the tendencies and F(phi,x,t) is a non linear function of the variables (or features) like u,v,t which are given. So I try to do some feature engineering to reproduce the funktional relations F(phi). Unfortunately the time t and location x are not given. But i guess this is intentional since the derived machine learning model should be independent of the location and time in order to generalize well. </p>",
          "votes": 5,
          "replies": [
            {
              "id": 2836199,
              "author_name": "HappyKiter",
              "author_url": "",
              "post_date": "2024-05-25T18:09:16.250000",
              "content": "<p>What it one has also to mention is that variables at pressure levels which are far away (e.g. label 59 and label 40) are not very correlated. It is also very unlikely the velocity at label 40 will influence the tendency of the velocity at label 59</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2865824,
      "author_name": "Rio",
      "author_url": "",
      "post_date": "2024-06-11T02:50:59.613000",
      "content": "<p>Thank you for your great knowledge!<br>\nI wonder why methods from past competitions might not be useful.<br>\nIs there any fundamental difference compared to other competitions?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2856969,
      "author_name": "Vasilis",
      "author_url": "",
      "post_date": "2024-06-05T16:29:20.040000",
      "content": "<p>Hey phalanx, when you achieved 0.72 from the first epoch with what bach size? or how many training iterations ?</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2909809,
      "author_name": "FarisML",
      "author_url": "",
      "post_date": "2024-07-07T10:39:25.527000",
      "content": "<p>Try Unet with attention block before the bottle neck</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2853321,
      "author_name": "Juan D C F",
      "author_url": "",
      "post_date": "2024-06-03T18:06:21.337000",
      "content": "<p>how the heck are people getting 0.72!? I guess my models are pretty basic, I am getting 0.57 with a very simple MLP.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2839420,
      "author_name": "Fernando Melo",
      "author_url": "",
      "post_date": "2024-05-27T14:13:55.240000",
      "content": "<p>Thanks for sharing.</p>\n<p>I have one more question, in terms of NN architecture, which path do you think is best to start with? Seq2Seq as discussed in another thread, or is it possible to achieve 0.7+ with FFNN?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2840302,
          "author_name": "Zhuoqun Li",
          "author_url": "",
          "post_date": "2024-05-28T04:30:46.220000",
          "content": "<p>I am confident of getting to 0.7+ using MLP, but I think mayde need to train with some techniques based on EDA</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2856457,
              "author_name": "Vasilis",
              "author_url": "",
              "post_date": "2024-06-05T09:55:01.927000",
              "content": "<p>which are these techniques?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2856980,
              "author_name": "Amedeo Biolatti",
              "author_url": "",
              "post_date": "2024-06-05T16:45:04.350000",
              "content": "<p>\"Did you ever hear the training of Darth Perceptron the Wide?\"</p>",
              "votes": 4,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2839068,
      "author_name": "Talha Barkaat Ahmad ☑️",
      "author_url": "",
      "post_date": "2024-05-27T11:23:08.620000",
      "content": "<p><a href=\"https://www.kaggle.com/muhammadawaistayyab\" target=\"_blank\">@muhammadawaistayyab</a> </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2837131,
      "author_name": "Vasilis",
      "author_url": "",
      "post_date": "2024-05-26T09:39:56.940000",
      "content": "<p>i have changed from float32 to float64 and i have reduced my batch_size from 2048 to 1024(because when i use float64 i was left out of memory) and now 100 steps of transformer from ~700 seconds takes 1560 … does this make sense?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2837182,
          "author_name": "Vasilis",
          "author_url": "",
          "post_date": "2024-05-26T10:43:21.533000",
          "content": "<p>so i asked chatGPT and it seems very obvious that the train time went so up :(</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2837243,
          "author_name": "Amedeo Biolatti",
          "author_url": "",
          "post_date": "2024-05-26T11:39:55.077000",
          "content": "<p>Suggestion: float64 -&gt; normalize -&gt; float32 -&gt; train</p>",
          "votes": 9,
          "replies": [
            {
              "id": 2838950,
              "author_name": "Vasilis",
              "author_url": "",
              "post_date": "2024-05-27T09:41:08.877000",
              "content": "<p>How did you do that, when i try that i get underflows, the loss functions sky rockets randomly so i need to train at float64. This happens only when i multiply y with the weights before training. If i dont multiply with the weights then i dont have underflows but i have red in the discussions that if you multiply with the weights you have an improvement of ~0.1</p>\n<p>Actually i noticed that even with float64 i have the same phenomenon, any idea why?<br>\nepoch: 0 , chunk: 0, Step 100, Training Loss: 0.5917 iterations: 100 time: 2783.370905160904<br>\nepoch: 0 , chunk: 0, Step 200, Training Loss: 0.5177 iterations: 200 time: 3947.4694044589996<br>\nepoch: 0 , chunk: 0, Step 300, Training Loss: 0.5088 iterations: 300 time: 5116.634320020676<br>\nepoch: 0 , chunk: 0, Step 400, Training Loss: 0.5138 iterations: 400 time: 6274.054301977158<br>\nepoch: 0 model saved chunk: 0 iterations: 486 val loss: 0.5341707844085616 time: 3782.210193157196</p>\n<p>tensor_data input shape: torch.Size([497123, 60, 25])<br>\nepoch: 0 chunk: 1 prep chunk time: 15.340328693389893<br>\nepoch: 0 , chunk: 1, Step 100, Training Loss: 0.4475 iterations: 586 time: 2813.811069250107<br>\nepoch: 0 , chunk: 1, Step 200, Training Loss: 0.4961 iterations: 686 time: 6220.386344671249<br>\nepoch: 0 , chunk: 1, Step 300, Training Loss: 0.4782 iterations: 786 time: 7684.73356628418<br>\nepoch: 0 , chunk: 1, Step 400, Training Loss: 1.3982 iterations: 886 time: 9230.30635881424</p>\n<p>The train loss spikes, this happens only when i multiply the y with the weights before training.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2839138,
              "author_name": "Amedeo Biolatti",
              "author_url": "",
              "post_date": "2024-05-27T11:56:52.490000",
              "content": "<p>I have a baseline notebook where I directly train in float32 (also the normalization is in float32). My public score is also using the exact same logic, so it is possibile to get to at least 0.75 training in float32 without pre-normalizing. I didn't try to multiply by the weights</p>",
              "votes": 5,
              "replies": []
            },
            {
              "id": 2839495,
              "author_name": "Amedeo Biolatti",
              "author_url": "",
              "post_date": "2024-05-27T14:57:27.603000",
              "content": "<p>Probably it generate some big gradient, I would try to clip the grad norm</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2834706,
      "author_name": "Benjamin Atiemo",
      "author_url": "",
      "post_date": "2024-05-25T00:14:06.407000",
      "content": "<p>Great sharing</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2834127,
      "author_name": "DennisSakva",
      "author_url": "",
      "post_date": "2024-05-24T15:03:12.090000",
      "content": "<p>\" By refining my model architecture, I achieved CV 0.72+ from the first epoch.\"<br>\nIncredible! Congrats :)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2874995,
          "author_name": "Youri Matiounine",
          "author_url": "",
          "post_date": "2024-06-16T18:35:36.520000",
          "content": "<p>it seemed so impressive at the time, but now even i managed to get to CV of 0.726 after the first epoch . Only took me a month!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2837393,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-05-26T13:04:36.287000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2837869,
      "author_name": "Benjamin Atiemo",
      "author_url": "",
      "post_date": "2024-05-26T17:50:45.400000",
      "content": "<p>Thank you very much </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2832988": "It seems that there are few participants with an LB score above 0.7, and many participants are struggling in this competition. Considering lowering the entry barrier, I thought about providing starter code. However, since there is limited knowledge about this competition and fairness as a competitive event is maintained,  I'd like to limit our support to sharing simple tips.\nBy considering just the following two points, you should be able to easily achieve a score above 0.7.\n\n## 1. Pay attention to the data type of floating-point numbers\nThe provided data is in float64. As shared in [this discussion](https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/500816), converting to float32 poses a risk of underflow. You must not convert to float32.\n\n## 2. Model architecture\nSharing detailed information while maintaining fairness is difficult, so I will limit it to a few tips. EDA and domain knowledge are extremely important. Methods from past competitions may not be useful (at least they are not useful for me). By refining my model architecture, I achieved CV 0.72+ from the first epoch.\n\n\n\nBy focusing solely on improving the model architecture, I was able to achieve LB 0.766. I did not use pseudo labeling or add any additional training data. Since the model is extremely important, please review your data again and gather domain knowledge. \nHave fun experimenting! Good luck!",
    "2833069": "Thanks for sharing your strategy. Have you tried training with both float 32 and 64 and check the difference?",
    "2835512": "Any suggested material to get some domain knowledge?",
    "2865824": "Thank you for your great knowledge!\nI wonder why methods from past competitions might not be useful.\nIs there any fundamental difference compared to other competitions?",
    "2856969": "Hey phalanx, when you achieved 0.72 from the first epoch with what bach size? or how many training iterations ?",
    "2909809": "Try Unet with attention block before the bottle neck",
    "2853321": "how the heck are people getting 0.72!? I guess my models are pretty basic, I am getting 0.57 with a very simple MLP.",
    "2839420": "Thanks for sharing.\n\nI have one more question, in terms of NN architecture, which path do you think is best to start with? Seq2Seq as discussed in another thread, or is it possible to achieve 0.7+ with FFNN?",
    "2839068": "@muhammadawaistayyab ",
    "2837131": "i have changed from float32 to float64 and i have reduced my batch_size from 2048 to 1024(because when i use float64 i was left out of memory) and now 100 steps of transformer from ~700 seconds takes 1560 ... does this make sense?",
    "2834706": "Great sharing",
    "2834127": "\" By refining my model architecture, I achieved CV 0.72+ from the first epoch.\"\nIncredible! Congrats :)",
    "2837393": "",
    "2837869": "Thank you very much "
  }
}