{
  "id": 495255,
  "title": "CONFUSION and POTENTIAL issue about evaluation metric",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/495255",
  "author_name": "Ulrich G.",
  "post_date": "2024-04-20T11:09:46.681000",
  "votes": 26,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I think there is an issue with the evaluation metric. </p>\n<p>I'd like to know if we</p>\n<ul>\n<li>a) take the average of 368 individual R2s :<br>\n$$  \\text{SCORE1} = \\frac{1}{368}  \\sum_{k=1}^{368} 1- \\frac{  \\sum_{i=1}^n (y_{ik} - \\hat{y}_{ik} )^2  }{   SSE_k    }$$</li>\n<li>b) take the R2 of all data<br>\n$$ \\text{SCORE2} =   1 -  \\frac{ \\sum_{k=1}^{368} \\sum_{i=1}^n (y_{ik} - \\hat{y}_{ik} )^2  }{   SSE    }$$ </li>\n</ul>\n<p>Because I noticed that in the sample submission <code>sample_submission.csv</code> all the 368 colums have the same value per column. So weighting columns with <strong>SCORE1</strong> will be useless. Again if we are good with 367 variables and performed poorly on a 368-th with a low variance, then we will perform poorly overall. So are we using <strong>SCORE2</strong> with the weights of  <code>sample_submission.csv</code> ?</p>\n<p><a href=\"https://www.kaggle.com/mylesoneill\" target=\"_blank\">@mylesoneill</a>, <a href=\"https://www.kaggle.com/ashleychow\" target=\"_blank\">@ashleychow</a></p>",
  "messages": [
    {
      "id": 2763190,
      "postDate": "2024-04-20T11:09:46.680Z",
      "content": "<p>I think there is an issue with the evaluation metric. </p>\n<p>I'd like to know if we</p>\n<ul>\n<li>a) take the average of 368 individual R2s :<br>\n$$  \\text{SCORE1} = \\frac{1}{368}  \\sum_{k=1}^{368} 1- \\frac{  \\sum_{i=1}^n (y_{ik} - \\hat{y}_{ik} )^2  }{   SSE_k    }$$</li>\n<li>b) take the R2 of all data<br>\n$$ \\text{SCORE2} =   1 -  \\frac{ \\sum_{k=1}^{368} \\sum_{i=1}^n (y_{ik} - \\hat{y}_{ik} )^2  }{   SSE    }$$ </li>\n</ul>\n<p>Because I noticed that in the sample submission <code>sample_submission.csv</code> all the 368 colums have the same value per column. So weighting columns with <strong>SCORE1</strong> will be useless. Again if we are good with 367 variables and performed poorly on a 368-th with a low variance, then we will perform poorly overall. So are we using <strong>SCORE2</strong> with the weights of  <code>sample_submission.csv</code> ?</p>\n<p><a href=\"https://www.kaggle.com/mylesoneill\" target=\"_blank\">@mylesoneill</a>, <a href=\"https://www.kaggle.com/ashleychow\" target=\"_blank\">@ashleychow</a></p>",
      "rawMarkdown": "I think there is an issue with the evaluation metric. \n\nI'd like to know if we\n* a) take the average of 368 individual R2s :\n$$  \\text{SCORE1} = \\frac{1}{368}  \\sum_{k=1}^{368} 1- \\frac{  \\sum_{i=1}^n (y_{ik} - \\hat{y}_{ik} )^2  }{   SSE_k    }$$\n* b) take the R2 of all data\n$$ \\text{SCORE2} =   1 -  \\frac{ \\sum_{k=1}^{368} \\sum_{i=1}^n (y_{ik} - \\hat{y}_{ik} )^2  }{   SSE    }$$ \n\nBecause I noticed that in the sample submission `sample_submission.csv` all the 368 colums have the same value per column. So weighting columns with **SCORE1** will be useless. Again if we are good with 367 variables and performed poorly on a 368-th with a low variance, then we will perform poorly overall. So are we using **SCORE2** with the weights of  `sample_submission.csv` ?\n\n@mylesoneill, @ashleychow",
      "votes": 25
    },
    {
      "id": 2769052,
      "postDate": "2024-04-23T06:31:20.647Z",
      "content": "<p>According to the source code <a href=\"https://www.kaggle.com/code/jerrylin96/r2-score-default\" target=\"_blank\">here</a> and <a href=\"https://github.com/scikit-learn/scikit-learn/blob/8721245511de2f225ff5f9aa5f5fadce663cd4a3/sklearn/metrics/_regression.py#L857\" target=\"_blank\">here</a>, it appears that Score 1 is what's being used. </p>\n<p>You raise a very fair point that this would negate the effect of much of the prediction weightings. The important exception would be the variables that get zero'd out.</p>",
      "rawMarkdown": "According to the source code [here](https://www.kaggle.com/code/jerrylin96/r2-score-default) and [here](https://github.com/scikit-learn/scikit-learn/blob/8721245511de2f225ff5f9aa5f5fadce663cd4a3/sklearn/metrics/_regression.py#L857), it appears that Score 1 is what's being used. \n\nYou raise a very fair point that this would negate the effect of much of the prediction weightings. The important exception would be the variables that get zero'd out.",
      "votes": 9,
      "replies": [
        {
          "id": 2770926,
          "postDate": "2024-04-24T04:52:13.330Z",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a>. I noticed that as well. So can you confirm that we use R1 ?</p>",
          "rawMarkdown": "Thank you @jerrylin96. I noticed that as well. So can you confirm that we use R1 ?",
          "votes": 2,
          "replies": [
            {
              "id": 2789407,
              "postDate": "2024-05-02T16:55:49.407Z",
              "content": "<p>Should be SCORE1</p>",
              "rawMarkdown": "Should be SCORE1",
              "votes": 1
            }
          ]
        },
        {
          "id": 2800412,
          "postDate": "2024-05-08T07:15:58.787Z",
          "content": "<p>I think the submission weights do help, because they stop columns with small absolute values from having tiny variance compared to other columns. And having tiny variance is a problem because the R2 formula explodes to infinity for zero variance, so a relatively small (absolute) prediction error in a column with very low variance will make the overall average-across-columns R2 go very large. If we multiply a column of tiny numbers by 1e15 say (the largest submission weightings I think), the variance is increased to something reasonable compared with other columns, unless the data is truly invariant (in which case R2 is no good as a metric).</p>\n<p>PS <a href=\"https://www.kaggle.com/ulrich07\" target=\"_blank\">@ulrich07</a> thanks very much for the question, I'm a newbie but you prompted me to look into this much harder. Aren't your (y-y^)/SSE sums upside-down though? But we know what you mean, don't worry if so!</p>",
          "rawMarkdown": "I think the submission weights do help, because they stop columns with small absolute values from having tiny variance compared to other columns. And having tiny variance is a problem because the R2 formula explodes to infinity for zero variance, so a relatively small (absolute) prediction error in a column with very low variance will make the overall average-across-columns R2 go very large. If we multiply a column of tiny numbers by 1e15 say (the largest submission weightings I think), the variance is increased to something reasonable compared with other columns, unless the data is truly invariant (in which case R2 is no good as a metric).\n\nPS @ulrich07 thanks very much for the question, I'm a newbie but you prompted me to look into this much harder. Aren't your (y-y^)/SSE sums upside-down though? But we know what you mean, don't worry if so!"
        }
      ]
    },
    {
      "id": 2763197,
      "postDate": "2024-04-20T11:20:47.817Z",
      "content": "<p>I don't think it's Score 1 which would be mean column-wise R2 score and it's not described like that on evaluation section. Besides, I calculated my validation score with Score 2 and it was pretty close to LB score. </p>",
      "rawMarkdown": "I don't think it's Score 1 which would be mean column-wise R2 score and it's not described like that on evaluation section. Besides, I calculated my validation score with Score 2 and it was pretty close to LB score. ",
      "votes": 3,
      "replies": [
        {
          "id": 2763249,
          "postDate": "2024-04-20T12:26:46.290Z",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a>. For your validation score, do you use the weights from <code>submission.csv</code> ? For me it does not seem to work. But it <strong>may</strong> come from the fact that I use only 300k samples to train and 100k samples to evaluate ( 4% of the dataset).</p>",
          "rawMarkdown": "Thank you @gunesevitan. For your validation score, do you use the weights from `submission.csv` ? For me it does not seem to work. But it **may** come from the fact that I use only 300k samples to train and 100k samples to evaluate ( 4% of the dataset).",
          "votes": 3,
          "replies": [
            {
              "id": 2763257,
              "postDate": "2024-04-20T12:34:41.947Z",
              "content": "<p>I'm not using the weights for validation. They are only for submissions.</p>",
              "rawMarkdown": "I'm not using the weights for validation. They are only for submissions.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2823410,
      "postDate": "2024-05-19T07:22:45.033Z",
      "content": "<p>Hi, I just joined this game. I wanna ask: is there any update on this topic?</p>\n<p>So the competition use score1 as evaluation metric and the weights don't make any difference.</p>\n<p>do I understand right?</p>",
      "rawMarkdown": "Hi, I just joined this game. I wanna ask: is there any update on this topic?\n\nSo the competition use score1 as evaluation metric and the weights don't make any difference.\n\ndo I understand right?",
      "replies": [
        {
          "id": 2823519,
          "postDate": "2024-05-19T08:57:35.077Z",
          "content": "<p>some weights are zeros, you probably should just make a constant zero prediction for these columns<br>\nafter compute per-class R1-score and then average them (this method worked for me in terms of CV/LB correlation)</p>",
          "rawMarkdown": "some weights are zeros, you probably should just make a constant zero prediction for these columns\nafter compute per-class R1-score and then average them (this method worked for me in terms of CV/LB correlation)",
          "votes": 3,
          "replies": [
            {
              "id": 2823526,
              "postDate": "2024-05-19T09:09:26.053Z",
              "content": "<p>Thanks for answering!</p>",
              "rawMarkdown": "Thanks for answering!"
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2769052,
      "author_name": "Jerry Lin",
      "author_url": "",
      "post_date": "2024-04-23T06:31:20.647000",
      "content": "<p>According to the source code <a href=\"https://www.kaggle.com/code/jerrylin96/r2-score-default\" target=\"_blank\">here</a> and <a href=\"https://github.com/scikit-learn/scikit-learn/blob/8721245511de2f225ff5f9aa5f5fadce663cd4a3/sklearn/metrics/_regression.py#L857\" target=\"_blank\">here</a>, it appears that Score 1 is what's being used. </p>\n<p>You raise a very fair point that this would negate the effect of much of the prediction weightings. The important exception would be the variables that get zero'd out.</p>",
      "votes": 9,
      "replies": [
        {
          "id": 2770926,
          "author_name": "Ulrich G.",
          "author_url": "",
          "post_date": "2024-04-24T04:52:13.330000",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a>. I noticed that as well. So can you confirm that we use R1 ?</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2789407,
              "author_name": "olougern111",
              "author_url": "",
              "post_date": "2024-05-02T16:55:49.407000",
              "content": "<p>Should be SCORE1</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 2800412,
          "author_name": "Charlie Wartnaby",
          "author_url": "",
          "post_date": "2024-05-08T07:15:58.787000",
          "content": "<p>I think the submission weights do help, because they stop columns with small absolute values from having tiny variance compared to other columns. And having tiny variance is a problem because the R2 formula explodes to infinity for zero variance, so a relatively small (absolute) prediction error in a column with very low variance will make the overall average-across-columns R2 go very large. If we multiply a column of tiny numbers by 1e15 say (the largest submission weightings I think), the variance is increased to something reasonable compared with other columns, unless the data is truly invariant (in which case R2 is no good as a metric).</p>\n<p>PS <a href=\"https://www.kaggle.com/ulrich07\" target=\"_blank\">@ulrich07</a> thanks very much for the question, I'm a newbie but you prompted me to look into this much harder. Aren't your (y-y^)/SSE sums upside-down though? But we know what you mean, don't worry if so!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2763197,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2024-04-20T11:20:47.817000",
      "content": "<p>I don't think it's Score 1 which would be mean column-wise R2 score and it's not described like that on evaluation section. Besides, I calculated my validation score with Score 2 and it was pretty close to LB score. </p>",
      "votes": 3,
      "replies": [
        {
          "id": 2763249,
          "author_name": "Ulrich G.",
          "author_url": "",
          "post_date": "2024-04-20T12:26:46.290000",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a>. For your validation score, do you use the weights from <code>submission.csv</code> ? For me it does not seem to work. But it <strong>may</strong> come from the fact that I use only 300k samples to train and 100k samples to evaluate ( 4% of the dataset).</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2763257,
              "author_name": "Gunes Evitan",
              "author_url": "",
              "post_date": "2024-04-20T12:34:41.947000",
              "content": "<p>I'm not using the weights for validation. They are only for submissions.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2823410,
      "author_name": "ADAM.",
      "author_url": "",
      "post_date": "2024-05-19T07:22:45.033000",
      "content": "<p>Hi, I just joined this game. I wanna ask: is there any update on this topic?</p>\n<p>So the competition use score1 as evaluation metric and the weights don't make any difference.</p>\n<p>do I understand right?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2823519,
          "author_name": "slime",
          "author_url": "",
          "post_date": "2024-05-19T08:57:35.077000",
          "content": "<p>some weights are zeros, you probably should just make a constant zero prediction for these columns<br>\nafter compute per-class R1-score and then average them (this method worked for me in terms of CV/LB correlation)</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2823526,
              "author_name": "ADAM.",
              "author_url": "",
              "post_date": "2024-05-19T09:09:26.053000",
              "content": "<p>Thanks for answering!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2763190": "I think there is an issue with the evaluation metric. \n\nI'd like to know if we\n* a) take the average of 368 individual R2s :\n$$  \\text{SCORE1} = \\frac{1}{368}  \\sum_{k=1}^{368} 1- \\frac{  \\sum_{i=1}^n (y_{ik} - \\hat{y}_{ik} )^2  }{   SSE_k    }$$\n* b) take the R2 of all data\n$$ \\text{SCORE2} =   1 -  \\frac{ \\sum_{k=1}^{368} \\sum_{i=1}^n (y_{ik} - \\hat{y}_{ik} )^2  }{   SSE    }$$ \n\nBecause I noticed that in the sample submission `sample_submission.csv` all the 368 colums have the same value per column. So weighting columns with **SCORE1** will be useless. Again if we are good with 367 variables and performed poorly on a 368-th with a low variance, then we will perform poorly overall. So are we using **SCORE2** with the weights of  `sample_submission.csv` ?\n\n@mylesoneill, @ashleychow",
    "2769052": "According to the source code [here](https://www.kaggle.com/code/jerrylin96/r2-score-default) and [here](https://github.com/scikit-learn/scikit-learn/blob/8721245511de2f225ff5f9aa5f5fadce663cd4a3/sklearn/metrics/_regression.py#L857), it appears that Score 1 is what's being used. \n\nYou raise a very fair point that this would negate the effect of much of the prediction weightings. The important exception would be the variables that get zero'd out.",
    "2763197": "I don't think it's Score 1 which would be mean column-wise R2 score and it's not described like that on evaluation section. Besides, I calculated my validation score with Score 2 and it was pretty close to LB score. ",
    "2823410": "Hi, I just joined this game. I wanna ask: is there any update on this topic?\n\nSo the competition use score1 as evaluation metric and the weights don't make any difference.\n\ndo I understand right?"
  }
}