{
  "id": 513220,
  "title": "Lots of negative r2 scores after multiplying with new weights",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/513220",
  "author_name": "Gunes Evitan",
  "post_date": "2024-06-19T05:38:39.338000",
  "votes": 13,
  "comment_count": 17,
  "views": 0,
  "content": "<p>I'm investigating this right now. Does that happen to anyone else?</p>\n<p>Edit: Could it be related to q0001, q0002 and q0003 being too small and r2 score drops due to fp precision error? It happens when I use mixed precision.</p>",
  "messages": [
    {
      "id": 2878655,
      "postDate": "2024-06-19T05:38:39.337Z",
      "content": "<p>I'm investigating this right now. Does that happen to anyone else?</p>\n<p>Edit: Could it be related to q0001, q0002 and q0003 being too small and r2 score drops due to fp precision error? It happens when I use mixed precision.</p>",
      "rawMarkdown": "I'm investigating this right now. Does that happen to anyone else?\n\nEdit: Could it be related to q0001, q0002 and q0003 being too small and r2 score drops due to fp precision error? It happens when I use mixed precision.",
      "votes": 13
    },
    {
      "id": 2879385,
      "postDate": "2024-06-19T14:20:17.353Z",
      "content": "<p>The type float32 in python have a lower bound of 7.007e-46. Any value lower is clipped to zero (0) making r2 explode. When computing the metric with original target values, make sure to use float64.</p>",
      "rawMarkdown": "The type float32 in python have a lower bound of 7.007e-46. Any value lower is clipped to zero (0) making r2 explode. When computing the metric with original target values, make sure to use float64.",
      "votes": 8,
      "replies": [
        {
          "id": 2880520,
          "postDate": "2024-06-20T08:16:23.400Z",
          "content": "<p>Ok but underflow isn't happening on my local machine when I calculate r2 score on 32-bit targets and predictions. It was also like that before weights are changed. I tried 32-bit inference and casting to 64-bit later and 64-bit inference. Both of them are scoring negative on LB for some reason.</p>",
          "rawMarkdown": "Ok but underflow isn't happening on my local machine when I calculate r2 score on 32-bit targets and predictions. It was also like that before weights are changed. I tried 32-bit inference and casting to 64-bit later and 64-bit inference. Both of them are scoring negative on LB for some reason."
        }
      ]
    },
    {
      "id": 2878984,
      "postDate": "2024-06-19T09:51:00.783Z",
      "content": "<p>When i calculate  mean ((preds - targets) ** 2) the result makes sense (around 0.4 very early in training) but when i calculate the R2 i get values very close to 1 like 0.95 and when i submit its negative of course</p>",
      "rawMarkdown": "When i calculate  mean ((preds - targets) ** 2) the result makes sense (around 0.4 very early in training) but when i calculate the R2 i get values very close to 1 like 0.95 and when i submit its negative of course",
      "votes": 1,
      "replies": [
        {
          "id": 2878995,
          "postDate": "2024-06-19T10:04:14.677Z",
          "content": "<p>There is nothing changed on the training side unless you are training on targets * weight. Even if you do that, r2 score shouldn't change since scale doesn't affect it. MSE would definitely change though.</p>\n<p>I was calculating my validation score as r2(targets * weight, predictions * weight) and it didn't change either. The only thing that is changed is LB score which is very unusual. That's why I'm suspecting of fp precision error on Kaggle side.</p>",
          "rawMarkdown": "There is nothing changed on the training side unless you are training on targets * weight. Even if you do that, r2 score shouldn't change since scale doesn't affect it. MSE would definitely change though.\n\nI was calculating my validation score as r2(targets * weight, predictions * weight) and it didn't change either. The only thing that is changed is LB score which is very unusual. That's why I'm suspecting of fp precision error on Kaggle side.",
          "votes": 3,
          "replies": [
            {
              "id": 2878999,
              "postDate": "2024-06-19T10:10:28.863Z",
              "content": "<p>yes i see, i had a bug in my R2, i was taking the whole mean of targets and not in axis=0. Very strange that with this bug my R2 score calculation was very very similar with the submission, while now the bug makes a big difference</p>",
              "rawMarkdown": "yes i see, i had a bug in my R2, i was taking the whole mean of targets and not in axis=0. Very strange that with this bug my R2 score calculation was very very similar with the submission, while now the bug makes a big difference"
            },
            {
              "id": 2880002,
              "postDate": "2024-06-20T01:16:32.043Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2880003,
              "postDate": "2024-06-20T01:18:41.363Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        },
        {
          "id": 2880066,
          "postDate": "2024-06-20T02:52:30.757Z",
          "content": "<p>I've also encountered this problem. How did you solve it? </p>",
          "rawMarkdown": "I've also encountered this problem. How did you solve it? ",
          "replies": [
            {
              "id": 2881211,
              "postDate": "2024-06-20T15:26:09.327Z",
              "content": "<p>i use pytorch for me the problem was when i was going from polars float64 to torch float32, when i start read the X and y as float32 from polars it was okay</p>",
              "rawMarkdown": "i use pytorch for me the problem was when i was going from polars float64 to torch float32, when i start read the X and y as float32 from polars it was okay"
            },
            {
              "id": 2881241,
              "postDate": "2024-06-20T15:39:41.903Z",
              "content": "<p>this helped me get back into positive scores: <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/513220#2880523\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/513220#2880523</a></p>",
              "rawMarkdown": "this helped me get back into positive scores: https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/513220#2880523"
            }
          ]
        }
      ]
    },
    {
      "id": 2880482,
      "postDate": "2024-06-20T07:57:20.917Z",
      "content": "<p>Hey guys, apparently the old weights (sample_submission) were 0 for ptend_q0002_{12-14}, but are now 1.0 for these columns! I was training using the 0 weights, so I am guessing that is why my model is so bad… good luck!</p>",
      "rawMarkdown": "Hey guys, apparently the old weights (sample_submission) were 0 for ptend_q0002_{12-14}, but are now 1.0 for these columns! I was training using the 0 weights, so I am guessing that is why my model is so bad... good luck!",
      "votes": 2,
      "replies": [
        {
          "id": 2880523,
          "postDate": "2024-06-20T08:18:05.843Z",
          "content": "<p>We can get perfect r2 score by overwriting -state_q0002 / 1200 for those targets. They don't hurt the overall performance during training either.</p>",
          "rawMarkdown": "We can get perfect r2 score by overwriting -state_q0002 / 1200 for those targets. They don't hurt the overall performance during training either."
        }
      ]
    },
    {
      "id": 2878955,
      "postDate": "2024-06-19T09:18:10.417Z",
      "content": "<p>yes happened to me too. I have changed the weights but while my previous score was around 0.7 now i get negatives. i used to do something with masks like mask = 1.1 * std &lt; min_std and then preds[:, mask] *= 0, i wonder if this  has to do with it. </p>",
      "rawMarkdown": "yes happened to me too. I have changed the weights but while my previous score was around 0.7 now i get negatives. i used to do something with masks like mask = 1.1 * std < min_std and then preds[:, mask] *= 0, i wonder if this  has to do with it. ",
      "votes": 2,
      "replies": [
        {
          "id": 2879492,
          "postDate": "2024-06-19T15:49:28.873Z",
          "content": "<p><code>mask = 1.1 * std &lt; min_std</code><br>\nI don't understand the meaning of this code. What does this code mean?</p>",
          "rawMarkdown": "`mask = 1.1 * std < min_std`\nI don't understand the meaning of this code. What does this code mean?",
          "replies": [
            {
              "id": 2879496,
              "postDate": "2024-06-19T16:02:08.770Z",
              "content": "<p>i had some instabilities with some variables so i used this mask that i saw in a notebook to zero out variables that had very very small standard deviation. i used min_std = 1e-12.</p>",
              "rawMarkdown": "i had some instabilities with some variables so i used this mask that i saw in a notebook to zero out variables that had very very small standard deviation. i used min_std = 1e-12."
            }
          ]
        }
      ]
    },
    {
      "id": 2879440,
      "postDate": "2024-06-19T15:14:45.130Z",
      "content": "<p>I didn't change anything.<br>\nI am using Amadeo's notebook as baseline.</p>",
      "rawMarkdown": "I didn't change anything.\nI am using Amadeo's notebook as baseline."
    },
    {
      "id": 2879371,
      "postDate": "2024-06-19T14:10:41.587Z",
      "content": "<p>Lots of 0.0 scores on LGB model - polars used to load train and test.  The old weights let me get away with float32 as I was multiplying before model fit and got those very small feature values above the float32 minimum.</p>",
      "rawMarkdown": "Lots of 0.0 scores on LGB model - polars used to load train and test.  The old weights let me get away with float32 as I was multiplying before model fit and got those very small feature values above the float32 minimum."
    }
  ],
  "comments": [
    {
      "id": 2879385,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2024-06-19T14:20:17.353000",
      "content": "<p>The type float32 in python have a lower bound of 7.007e-46. Any value lower is clipped to zero (0) making r2 explode. When computing the metric with original target values, make sure to use float64.</p>",
      "votes": 8,
      "replies": [
        {
          "id": 2880520,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2024-06-20T08:16:23.400000",
          "content": "<p>Ok but underflow isn't happening on my local machine when I calculate r2 score on 32-bit targets and predictions. It was also like that before weights are changed. I tried 32-bit inference and casting to 64-bit later and 64-bit inference. Both of them are scoring negative on LB for some reason.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2878984,
      "author_name": "Vasilis",
      "author_url": "",
      "post_date": "2024-06-19T09:51:00.783000",
      "content": "<p>When i calculate  mean ((preds - targets) ** 2) the result makes sense (around 0.4 very early in training) but when i calculate the R2 i get values very close to 1 like 0.95 and when i submit its negative of course</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2878995,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2024-06-19T10:04:14.677000",
          "content": "<p>There is nothing changed on the training side unless you are training on targets * weight. Even if you do that, r2 score shouldn't change since scale doesn't affect it. MSE would definitely change though.</p>\n<p>I was calculating my validation score as r2(targets * weight, predictions * weight) and it didn't change either. The only thing that is changed is LB score which is very unusual. That's why I'm suspecting of fp precision error on Kaggle side.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2878999,
              "author_name": "Vasilis",
              "author_url": "",
              "post_date": "2024-06-19T10:10:28.863000",
              "content": "<p>yes i see, i had a bug in my R2, i was taking the whole mean of targets and not in axis=0. Very strange that with this bug my R2 score calculation was very very similar with the submission, while now the bug makes a big difference</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2880002,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-06-20T01:16:32.043000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2880003,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-06-20T01:18:41.363000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2880066,
          "author_name": "Chengwei Yan",
          "author_url": "",
          "post_date": "2024-06-20T02:52:30.757000",
          "content": "<p>I've also encountered this problem. How did you solve it? </p>",
          "votes": 0,
          "replies": [
            {
              "id": 2881211,
              "author_name": "Vasilis",
              "author_url": "",
              "post_date": "2024-06-20T15:26:09.327000",
              "content": "<p>i use pytorch for me the problem was when i was going from polars float64 to torch float32, when i start read the X and y as float32 from polars it was okay</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2881241,
              "author_name": "Juan D C F",
              "author_url": "",
              "post_date": "2024-06-20T15:39:41.903000",
              "content": "<p>this helped me get back into positive scores: <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/513220#2880523\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/513220#2880523</a></p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2880482,
      "author_name": "Juan D C F",
      "author_url": "",
      "post_date": "2024-06-20T07:57:20.917000",
      "content": "<p>Hey guys, apparently the old weights (sample_submission) were 0 for ptend_q0002_{12-14}, but are now 1.0 for these columns! I was training using the 0 weights, so I am guessing that is why my model is so bad… good luck!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2880523,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2024-06-20T08:18:05.843000",
          "content": "<p>We can get perfect r2 score by overwriting -state_q0002 / 1200 for those targets. They don't hurt the overall performance during training either.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2878955,
      "author_name": "Vasilis",
      "author_url": "",
      "post_date": "2024-06-19T09:18:10.417000",
      "content": "<p>yes happened to me too. I have changed the weights but while my previous score was around 0.7 now i get negatives. i used to do something with masks like mask = 1.1 * std &lt; min_std and then preds[:, mask] *= 0, i wonder if this  has to do with it. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 2879492,
          "author_name": "Zhuoqun Li",
          "author_url": "",
          "post_date": "2024-06-19T15:49:28.873000",
          "content": "<p><code>mask = 1.1 * std &lt; min_std</code><br>\nI don't understand the meaning of this code. What does this code mean?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2879496,
              "author_name": "Vasilis",
              "author_url": "",
              "post_date": "2024-06-19T16:02:08.770000",
              "content": "<p>i had some instabilities with some variables so i used this mask that i saw in a notebook to zero out variables that had very very small standard deviation. i used min_std = 1e-12.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2879440,
      "author_name": "Fernando Melo",
      "author_url": "",
      "post_date": "2024-06-19T15:14:45.130000",
      "content": "<p>I didn't change anything.<br>\nI am using Amadeo's notebook as baseline.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2879371,
      "author_name": "PC Jimmmy",
      "author_url": "",
      "post_date": "2024-06-19T14:10:41.587000",
      "content": "<p>Lots of 0.0 scores on LGB model - polars used to load train and test.  The old weights let me get away with float32 as I was multiplying before model fit and got those very small feature values above the float32 minimum.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2878655": "I'm investigating this right now. Does that happen to anyone else?\n\nEdit: Could it be related to q0001, q0002 and q0003 being too small and r2 score drops due to fp precision error? It happens when I use mixed precision.",
    "2879385": "The type float32 in python have a lower bound of 7.007e-46. Any value lower is clipped to zero (0) making r2 explode. When computing the metric with original target values, make sure to use float64.",
    "2878984": "When i calculate  mean ((preds - targets) ** 2) the result makes sense (around 0.4 very early in training) but when i calculate the R2 i get values very close to 1 like 0.95 and when i submit its negative of course",
    "2880482": "Hey guys, apparently the old weights (sample_submission) were 0 for ptend_q0002_{12-14}, but are now 1.0 for these columns! I was training using the 0 weights, so I am guessing that is why my model is so bad... good luck!",
    "2878955": "yes happened to me too. I have changed the weights but while my previous score was around 0.7 now i get negatives. i used to do something with masks like mask = 1.1 * std < min_std and then preds[:, mask] *= 0, i wonder if this  has to do with it. ",
    "2879440": "I didn't change anything.\nI am using Amadeo's notebook as baseline.",
    "2879371": "Lots of 0.0 scores on LGB model - polars used to load train and test.  The old weights let me get away with float32 as I was multiplying before model fit and got those very small feature values above the float32 minimum."
  }
}