{
  "id": 513514,
  "title": "Large negative values when predicting with new test data&weights",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/513514",
  "author_name": "Julian",
  "post_date": "2024-06-20T12:08:45.459000",
  "votes": 6,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Hey all,</p>\n<p>Since the update to the weights and test file my score on the leaderboard has been deep in the negative, whereas before they were around ~0.4. I'm not sure exactly what caused this but I noticed that 2 of my submissions were very close,  (specifically: -52053.30714 and -16326.32076, within +-1.0 score difference) to some other submissions, which leads me to suspect that other participants are having similar issues.</p>\n<p>I'm following the same pipeline to <a href=\"https://www.kaggle.com/code/airazusta014/pytorch-nn\" target=\"_blank\">this notebook</a>. I follow the standard normalization procedures as followed in the existing notebooks for the old server data i.e. <br>\nload data -&gt; multiply weights to training targets \"y\" -&gt; calculate mean and std. for x, y and x_test and apply normalization on each set -&gt; train model -&gt; predict normalized y_hat -&gt; denormalize y_hat via the mean/std. calculated on y -&gt; submit.</p>\n<p>I additionally tried an experiment using float64 for all data and model parameters, whereas before I was using float32 for everything, however this got the -52053.30714 score. The -16326.32076 score, on the other hand, was achieved by running the <strong>exact same procedure</strong> from <a href=\"https://www.kaggle.com/code/airazusta014/pytorch-nn\" target=\"_blank\">this notebook</a> on the new test data &amp; sample submission weights (if there's a hotfix that can be made in that implementation that would be awesome btw)</p>\n<p>I reckon something is going wrong here with the de/normalization, but I can't quite figure it out. If anyone has any suggestions or is able to upload a new submission that scored &gt;0 to diagnose where the differences lie in the predictions, myself and other participants having these issues with large negative scores would be very grateful! 😊</p>\n<p>Thanks in advance.</p>\n<p>Julian</p>",
  "messages": [
    {
      "id": 2883677,
      "postDate": "2024-06-22T04:52:56.423Z",
      "content": "<p>For all those struggling with negative scores, - it's best to view per-class R2 scores first and see which columns ruin your score. Choose a small validation set (manageable for your machine), save it during training and analyse it after.</p>\n<p>It's pretty hard to debug it if you look only at the average, since your model and data can be fine, but still something like <code>ptend_q0002_15</code> must be replaced as a post-processing to get positive score.</p>",
      "rawMarkdown": "For all those struggling with negative scores, - it's best to view per-class R2 scores first and see which columns ruin your score. Choose a small validation set (manageable for your machine), save it during training and analyse it after.\n\nIt's pretty hard to debug it if you look only at the average, since your model and data can be fine, but still something like ```ptend_q0002_15``` must be replaced as a post-processing to get positive score.",
      "votes": 5,
      "replies": [
        {
          "id": 2885226,
          "postDate": "2024-06-23T01:09:20.117Z",
          "content": "<p>Yeah that's a good idea, up until now I've only been looking at the average MSELoss for batch during training &amp; validation loss for each epoch 😅</p>",
          "rawMarkdown": "Yeah that's a good idea, up until now I've only been looking at the average MSELoss for batch during training & validation loss for each epoch 😅"
        }
      ]
    },
    {
      "id": 2880839,
      "postDate": "2024-06-20T12:08:45.460Z",
      "content": "<p>Hey all,</p>\n<p>Since the update to the weights and test file my score on the leaderboard has been deep in the negative, whereas before they were around ~0.4. I'm not sure exactly what caused this but I noticed that 2 of my submissions were very close,  (specifically: -52053.30714 and -16326.32076, within +-1.0 score difference) to some other submissions, which leads me to suspect that other participants are having similar issues.</p>\n<p>I'm following the same pipeline to <a href=\"https://www.kaggle.com/code/airazusta014/pytorch-nn\" target=\"_blank\">this notebook</a>. I follow the standard normalization procedures as followed in the existing notebooks for the old server data i.e. <br>\nload data -&gt; multiply weights to training targets \"y\" -&gt; calculate mean and std. for x, y and x_test and apply normalization on each set -&gt; train model -&gt; predict normalized y_hat -&gt; denormalize y_hat via the mean/std. calculated on y -&gt; submit.</p>\n<p>I additionally tried an experiment using float64 for all data and model parameters, whereas before I was using float32 for everything, however this got the -52053.30714 score. The -16326.32076 score, on the other hand, was achieved by running the <strong>exact same procedure</strong> from <a href=\"https://www.kaggle.com/code/airazusta014/pytorch-nn\" target=\"_blank\">this notebook</a> on the new test data &amp; sample submission weights (if there's a hotfix that can be made in that implementation that would be awesome btw)</p>\n<p>I reckon something is going wrong here with the de/normalization, but I can't quite figure it out. If anyone has any suggestions or is able to upload a new submission that scored &gt;0 to diagnose where the differences lie in the predictions, myself and other participants having these issues with large negative scores would be very grateful! 😊</p>\n<p>Thanks in advance.</p>\n<p>Julian</p>",
      "rawMarkdown": "Hey all,\n\nSince the update to the weights and test file my score on the leaderboard has been deep in the negative, whereas before they were around ~0.4. I'm not sure exactly what caused this but I noticed that 2 of my submissions were very close,  (specifically: -52053.30714 and -16326.32076, within +-1.0 score difference) to some other submissions, which leads me to suspect that other participants are having similar issues.\n\nI'm following the same pipeline to [this notebook](https://www.kaggle.com/code/airazusta014/pytorch-nn). I follow the standard normalization procedures as followed in the existing notebooks for the old server data i.e. \nload data -> multiply weights to training targets \"y\" -> calculate mean and std. for x, y and x_test and apply normalization on each set -> train model -> predict normalized y_hat -> denormalize y_hat via the mean/std. calculated on y -> submit.\n\nI additionally tried an experiment using float64 for all data and model parameters, whereas before I was using float32 for everything, however this got the -52053.30714 score. The -16326.32076 score, on the other hand, was achieved by running the **exact same procedure** from [this notebook](https://www.kaggle.com/code/airazusta014/pytorch-nn) on the new test data & sample submission weights (if there's a hotfix that can be made in that implementation that would be awesome btw)\n\nI reckon something is going wrong here with the de/normalization, but I can't quite figure it out. If anyone has any suggestions or is able to upload a new submission that scored >0 to diagnose where the differences lie in the predictions, myself and other participants having these issues with large negative scores would be very grateful! 😊\n\nThanks in advance.\n\nJulian",
      "votes": 5
    },
    {
      "id": 2883378,
      "postDate": "2024-06-21T22:59:37.523Z",
      "content": "<p>Ok quick update, in an attempt to achieve a score &gt;0 on the LB I ran <a href=\"https://www.kaggle.com/code/asarvazyan/leap-predict-the-mean\" target=\"_blank\">this code</a> which naïvely calculates the mean along each column in the test data and fills each column in the submission with this value, and finally multiplies this by the weights in the sample_submission. Naturally, I was expecting a score similar to the notebook and in fact I obtained the exact same mean values as shown in this notebook. By simply applying the code that <a href=\"https://www.kaggle.com/vasileioscharatsidis\" target=\"_blank\">@vasileioscharatsidis</a> mentioned <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/513193#2882807\" target=\"_blank\">here</a> (and was also mentioned in a few other previous threads somewhere) i.e.:</p>\n<pre><code> idx  (, ):\n     df_p_test = -df_test() / \n</code></pre>\n<p>I was finally able to submit a positive result on the LB. I hope this also helps others who are submitting predictions that are around -16,000 or -52,000, turns out there was hotfix after all… and thank you for ending my 3-day headache 😂. Now time to get back to the fun machine learning stuff!</p>",
      "rawMarkdown": "Ok quick update, in an attempt to achieve a score >0 on the LB I ran [this code](https://www.kaggle.com/code/asarvazyan/leap-predict-the-mean) which naïvely calculates the mean along each column in the test data and fills each column in the submission with this value, and finally multiplies this by the weights in the sample_submission. Naturally, I was expecting a score similar to the notebook and in fact I obtained the exact same mean values as shown in this notebook. By simply applying the code that @vasileioscharatsidis mentioned [here](https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/513193#2882807) (and was also mentioned in a few other previous threads somewhere) i.e.:\n\n```\nfor idx in range(0, 27):\n     df_p_test[f\"ptend_q0002_{idx}\"] = -df_test[f\"state_q0002_{idx}\"].to_numpy() / 1200\n```\n\nI was finally able to submit a positive result on the LB. I hope this also helps others who are submitting predictions that are around -16,000 or -52,000, turns out there was hotfix after all... and thank you for ending my 3-day headache 😂. Now time to get back to the fun machine learning stuff!",
      "votes": 1
    },
    {
      "id": 2880895,
      "postDate": "2024-06-20T12:48:19.657Z",
      "content": "<p>yes i have the exact same problem, haven't solved it yet :(. Indeed when i de-normalise i get very strange values for the validation set but also when i train and use R2 as criterion i get nan values sometimes, non of these happened with the previous weights</p>",
      "rawMarkdown": "yes i have the exact same problem, haven't solved it yet :(. Indeed when i de-normalise i get very strange values for the validation set but also when i train and use R2 as criterion i get nan values sometimes, non of these happened with the previous weights",
      "replies": [
        {
          "id": 2880930,
          "postDate": "2024-06-20T13:13:00.870Z",
          "content": "<p>I have tried to train without normalise target variables and i get the exact same result. Actually even worse, now the train loss seems to go down too fast.</p>",
          "rawMarkdown": "I have tried to train without normalise target variables and i get the exact same result. Actually even worse, now the train loss seems to go down too fast."
        },
        {
          "id": 2880965,
          "postDate": "2024-06-20T13:27:13.300Z",
          "content": "<p>Ok i think i have solved it. </p>\n<pre><code> # Preprocess the features  target columns\n     col  FEAT_COLS:\n        X = df.select(FEAT_COLS)..cast(pl.Float32))\n     col  TARGET_COLS:\n        y = df.select(TARGET_COLS)..cast(pl.Float32))\n</code></pre>\n<p>I have changed this code from pl.Float64 to pl.Float32 and it seems i get much more sensible results for now (i hope it last it is still first iterations). I nornalise x and y using mean and std   calculated with float64 though.</p>",
          "rawMarkdown": "Ok i think i have solved it. \n```\n # Preprocess the features and target columns\n    for col in FEAT_COLS:\n        X = df.select(FEAT_COLS).with_columns(pl.col(col).cast(pl.Float32))\n    for col in TARGET_COLS:\n        y = df.select(TARGET_COLS).with_columns(pl.col(col).cast(pl.Float32))\n```\nI have changed this code from pl.Float64 to pl.Float32 and it seems i get much more sensible results for now (i hope it last it is still first iterations). I nornalise x and y using mean and std   calculated with float64 though.",
          "votes": 2,
          "replies": [
            {
              "id": 2881119,
              "postDate": "2024-06-20T14:46:01.797Z",
              "content": "<p>yes it seems fixed, when i de-normalise and validate i get sensible results, only thing left is that when i optimize the R2 instead of MSE i sometimes get nan values, that was not happening with the old weights. So now if i optimize MSE everything seems fine and the evaluation problems are gone.</p>",
              "rawMarkdown": "yes it seems fixed, when i de-normalise and validate i get sensible results, only thing left is that when i optimize the R2 instead of MSE i sometimes get nan values, that was not happening with the old weights. So now if i optimize MSE everything seems fine and the evaluation problems are gone."
            },
            {
              "id": 2881306,
              "postDate": "2024-06-20T16:09:29.900Z",
              "content": "<p>Ok yeah, I had the same idea to normalize via float64 and then to train with float32 parameters in the model :)</p>",
              "rawMarkdown": "Ok yeah, I had the same idea to normalize via float64 and then to train with float32 parameters in the model :)"
            }
          ]
        }
      ]
    },
    {
      "id": 2881014,
      "postDate": "2024-06-20T13:45:32.580Z",
      "content": "<p>Previous weights were 1/standard_deviation, now they are just 1.<br>\nI think what you could do (and works for me): load data -&gt; normalize with mean and standard deviation -&gt; train -&gt; predict -&gt; de-normalize predicted values with the same mean and standard deviation -&gt; submit. </p>\n<p>I also think that by normalizing like this and then training, the R2 (that is calculated during submission) should be: 1 - MSE (mean squared error used as loss during training).</p>",
      "rawMarkdown": "Previous weights were 1/standard_deviation, now they are just 1.\nI think what you could do (and works for me): load data -> normalize with mean and standard deviation -> train -> predict -> de-normalize predicted values with the same mean and standard deviation -> submit. \n\nI also think that by normalizing like this and then training, the R2 (that is calculated during submission) should be: 1 - MSE (mean squared error used as loss during training).",
      "votes": 2,
      "isDeleted": true,
      "replies": [
        {
          "id": 2881322,
          "postDate": "2024-06-20T16:15:41.140Z",
          "content": "<p>Hm interesting, do you still multiply by the weights from the sample submission at the beginning and/or before submission, or do you not multiply with weights at all then? 🤔 I also had a look at the weights and there are 0s as well as 1s, so if the targets (y) were weighted with these weights then the model would learn to always predict 0 for the indices where the weights are 0 in the sample submission.</p>",
          "rawMarkdown": "Hm interesting, do you still multiply by the weights from the sample submission at the beginning and/or before submission, or do you not multiply with weights at all then? 🤔 I also had a look at the weights and there are 0s as well as 1s, so if the targets (y) were weighted with these weights then the model would learn to always predict 0 for the indices where the weights are 0 in the sample submission.",
          "replies": [
            {
              "id": 2881399,
              "postDate": "2024-06-20T17:15:26.790Z",
              "content": "<p>No need to multiply with the weights before or after the training. <br>\nColumns that are 0 (that have the weight of 0) are completely excluded from training. During the submission I additionally enter these columns with zero values (I don't know if it is necessary, if during the submission process they do it automatically). </p>\n<p>(the columns with the weight 0 is meant to signify that we should ignore these columns. this is the case for both after and before the update)</p>",
              "rawMarkdown": "No need to multiply with the weights before or after the training. \nColumns that are 0 (that have the weight of 0) are completely excluded from training. During the submission I additionally enter these columns with zero values (I don't know if it is necessary, if during the submission process they do it automatically). \n\n(the columns with the weight 0 is meant to signify that we should ignore these columns. this is the case for both after and before the update)",
              "isDeleted": true
            },
            {
              "id": 2881546,
              "postDate": "2024-06-20T18:54:35.410Z",
              "content": "<p>Aha good to know, thanks for the insight 😄! I've had no luck yet submitting &gt;0 R2 LB predictions with a simple FFN, but I'm retraining now without pre-multiplying the weights and using a lower min_std=1e-12 as mentioned <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/513220#2879496\" target=\"_blank\">here</a>. Maybe this could have been causing underflow somewhere, but let's see 😅. This bug seems really tricky to diagnose without really digging deep into the data and how it's being manipulated in python..</p>",
              "rawMarkdown": "Aha good to know, thanks for the insight 😄! I've had no luck yet submitting >0 R2 LB predictions with a simple FFN, but I'm retraining now without pre-multiplying the weights and using a lower min_std=1e-12 as mentioned [here](https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/513220#2879496). Maybe this could have been causing underflow somewhere, but let's see 😅. This bug seems really tricky to diagnose without really digging deep into the data and how it's being manipulated in python.."
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2883677,
      "author_name": "slime",
      "author_url": "",
      "post_date": "2024-06-22T04:52:56.423000",
      "content": "<p>For all those struggling with negative scores, - it's best to view per-class R2 scores first and see which columns ruin your score. Choose a small validation set (manageable for your machine), save it during training and analyse it after.</p>\n<p>It's pretty hard to debug it if you look only at the average, since your model and data can be fine, but still something like <code>ptend_q0002_15</code> must be replaced as a post-processing to get positive score.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 2885226,
          "author_name": "Julian",
          "author_url": "",
          "post_date": "2024-06-23T01:09:20.117000",
          "content": "<p>Yeah that's a good idea, up until now I've only been looking at the average MSELoss for batch during training &amp; validation loss for each epoch 😅</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2883378,
      "author_name": "Julian",
      "author_url": "",
      "post_date": "2024-06-21T22:59:37.523000",
      "content": "<p>Ok quick update, in an attempt to achieve a score &gt;0 on the LB I ran <a href=\"https://www.kaggle.com/code/asarvazyan/leap-predict-the-mean\" target=\"_blank\">this code</a> which naïvely calculates the mean along each column in the test data and fills each column in the submission with this value, and finally multiplies this by the weights in the sample_submission. Naturally, I was expecting a score similar to the notebook and in fact I obtained the exact same mean values as shown in this notebook. By simply applying the code that <a href=\"https://www.kaggle.com/vasileioscharatsidis\" target=\"_blank\">@vasileioscharatsidis</a> mentioned <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/513193#2882807\" target=\"_blank\">here</a> (and was also mentioned in a few other previous threads somewhere) i.e.:</p>\n<pre><code> idx  (, ):\n     df_p_test = -df_test() / \n</code></pre>\n<p>I was finally able to submit a positive result on the LB. I hope this also helps others who are submitting predictions that are around -16,000 or -52,000, turns out there was hotfix after all… and thank you for ending my 3-day headache 😂. Now time to get back to the fun machine learning stuff!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2880895,
      "author_name": "Vasilis",
      "author_url": "",
      "post_date": "2024-06-20T12:48:19.657000",
      "content": "<p>yes i have the exact same problem, haven't solved it yet :(. Indeed when i de-normalise i get very strange values for the validation set but also when i train and use R2 as criterion i get nan values sometimes, non of these happened with the previous weights</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2880930,
          "author_name": "Vasilis",
          "author_url": "",
          "post_date": "2024-06-20T13:13:00.870000",
          "content": "<p>I have tried to train without normalise target variables and i get the exact same result. Actually even worse, now the train loss seems to go down too fast.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2880965,
          "author_name": "Vasilis",
          "author_url": "",
          "post_date": "2024-06-20T13:27:13.300000",
          "content": "<p>Ok i think i have solved it. </p>\n<pre><code> # Preprocess the features  target columns\n     col  FEAT_COLS:\n        X = df.select(FEAT_COLS)..cast(pl.Float32))\n     col  TARGET_COLS:\n        y = df.select(TARGET_COLS)..cast(pl.Float32))\n</code></pre>\n<p>I have changed this code from pl.Float64 to pl.Float32 and it seems i get much more sensible results for now (i hope it last it is still first iterations). I nornalise x and y using mean and std   calculated with float64 though.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2881119,
              "author_name": "Vasilis",
              "author_url": "",
              "post_date": "2024-06-20T14:46:01.797000",
              "content": "<p>yes it seems fixed, when i de-normalise and validate i get sensible results, only thing left is that when i optimize the R2 instead of MSE i sometimes get nan values, that was not happening with the old weights. So now if i optimize MSE everything seems fine and the evaluation problems are gone.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2881306,
              "author_name": "Julian",
              "author_url": "",
              "post_date": "2024-06-20T16:09:29.900000",
              "content": "<p>Ok yeah, I had the same idea to normalize via float64 and then to train with float32 parameters in the model :)</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2881014,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-06-20T13:45:32.580000",
      "content": "<p>Previous weights were 1/standard_deviation, now they are just 1.<br>\nI think what you could do (and works for me): load data -&gt; normalize with mean and standard deviation -&gt; train -&gt; predict -&gt; de-normalize predicted values with the same mean and standard deviation -&gt; submit. </p>\n<p>I also think that by normalizing like this and then training, the R2 (that is calculated during submission) should be: 1 - MSE (mean squared error used as loss during training).</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2881322,
          "author_name": "Julian",
          "author_url": "",
          "post_date": "2024-06-20T16:15:41.140000",
          "content": "<p>Hm interesting, do you still multiply by the weights from the sample submission at the beginning and/or before submission, or do you not multiply with weights at all then? 🤔 I also had a look at the weights and there are 0s as well as 1s, so if the targets (y) were weighted with these weights then the model would learn to always predict 0 for the indices where the weights are 0 in the sample submission.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2881399,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-06-20T17:15:26.790000",
              "content": "<p>No need to multiply with the weights before or after the training. <br>\nColumns that are 0 (that have the weight of 0) are completely excluded from training. During the submission I additionally enter these columns with zero values (I don't know if it is necessary, if during the submission process they do it automatically). </p>\n<p>(the columns with the weight 0 is meant to signify that we should ignore these columns. this is the case for both after and before the update)</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2881546,
              "author_name": "Julian",
              "author_url": "",
              "post_date": "2024-06-20T18:54:35.410000",
              "content": "<p>Aha good to know, thanks for the insight 😄! I've had no luck yet submitting &gt;0 R2 LB predictions with a simple FFN, but I'm retraining now without pre-multiplying the weights and using a lower min_std=1e-12 as mentioned <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/513220#2879496\" target=\"_blank\">here</a>. Maybe this could have been causing underflow somewhere, but let's see 😅. This bug seems really tricky to diagnose without really digging deep into the data and how it's being manipulated in python..</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2883677": "For all those struggling with negative scores, - it's best to view per-class R2 scores first and see which columns ruin your score. Choose a small validation set (manageable for your machine), save it during training and analyse it after.\n\nIt's pretty hard to debug it if you look only at the average, since your model and data can be fine, but still something like ```ptend_q0002_15``` must be replaced as a post-processing to get positive score.",
    "2880839": "Hey all,\n\nSince the update to the weights and test file my score on the leaderboard has been deep in the negative, whereas before they were around ~0.4. I'm not sure exactly what caused this but I noticed that 2 of my submissions were very close,  (specifically: -52053.30714 and -16326.32076, within +-1.0 score difference) to some other submissions, which leads me to suspect that other participants are having similar issues.\n\nI'm following the same pipeline to [this notebook](https://www.kaggle.com/code/airazusta014/pytorch-nn). I follow the standard normalization procedures as followed in the existing notebooks for the old server data i.e. \nload data -> multiply weights to training targets \"y\" -> calculate mean and std. for x, y and x_test and apply normalization on each set -> train model -> predict normalized y_hat -> denormalize y_hat via the mean/std. calculated on y -> submit.\n\nI additionally tried an experiment using float64 for all data and model parameters, whereas before I was using float32 for everything, however this got the -52053.30714 score. The -16326.32076 score, on the other hand, was achieved by running the **exact same procedure** from [this notebook](https://www.kaggle.com/code/airazusta014/pytorch-nn) on the new test data & sample submission weights (if there's a hotfix that can be made in that implementation that would be awesome btw)\n\nI reckon something is going wrong here with the de/normalization, but I can't quite figure it out. If anyone has any suggestions or is able to upload a new submission that scored >0 to diagnose where the differences lie in the predictions, myself and other participants having these issues with large negative scores would be very grateful! 😊\n\nThanks in advance.\n\nJulian",
    "2883378": "Ok quick update, in an attempt to achieve a score >0 on the LB I ran [this code](https://www.kaggle.com/code/asarvazyan/leap-predict-the-mean) which naïvely calculates the mean along each column in the test data and fills each column in the submission with this value, and finally multiplies this by the weights in the sample_submission. Naturally, I was expecting a score similar to the notebook and in fact I obtained the exact same mean values as shown in this notebook. By simply applying the code that @vasileioscharatsidis mentioned [here](https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/513193#2882807) (and was also mentioned in a few other previous threads somewhere) i.e.:\n\n```\nfor idx in range(0, 27):\n     df_p_test[f\"ptend_q0002_{idx}\"] = -df_test[f\"state_q0002_{idx}\"].to_numpy() / 1200\n```\n\nI was finally able to submit a positive result on the LB. I hope this also helps others who are submitting predictions that are around -16,000 or -52,000, turns out there was hotfix after all... and thank you for ending my 3-day headache 😂. Now time to get back to the fun machine learning stuff!",
    "2880895": "yes i have the exact same problem, haven't solved it yet :(. Indeed when i de-normalise i get very strange values for the validation set but also when i train and use R2 as criterion i get nan values sometimes, non of these happened with the previous weights",
    "2881014": "Previous weights were 1/standard_deviation, now they are just 1.\nI think what you could do (and works for me): load data -> normalize with mean and standard deviation -> train -> predict -> de-normalize predicted values with the same mean and standard deviation -> submit. \n\nI also think that by normalizing like this and then training, the R2 (that is calculated during submission) should be: 1 - MSE (mean squared error used as loss during training)."
  }
}