{
  "id": 513193,
  "title": "Kaggle Competition Update [IMPORTANT]",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/513193",
  "author_name": "Jerry Lin",
  "post_date": "2024-06-19T00:30:57.174000",
  "votes": 31,
  "comment_count": 36,
  "views": 0,
  "content": "<p>Hi Kagglers,</p>\n<p>We have some good news and bad news. The good news is that we are extending the competition by 2 weeks! The new deadline is <strong>July 15, 2024 11:59PM UTC</strong>. We are incredibly impressed with your work so far and the engaging discussion that has been going on between competing teams. It’s truly awe-inspiring.</p>\n<p>Now for the bad news: we are releasing a new test set and sample submission (weighting). As Kaggle user <a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> pointed out, it is possible to reverse-engineer location information from the original test set. Many users (very fairly) entered this competition with the assumption that it would be a column to column regression problem, not a multi-column to column regression problem. Because we don’t want to give an unfair advantage to users who discovered this exploit early, we are making the tough decision to release a fresh test set + weighting + solution. For those that have invested a great deal of time in the multi-column approach, you must have realized that the sample IDs were pre-scrambled and of no use. Knowing this, it is fairly obvious that the intention of the competition designers was to not grant you access to this information. As Kaggle user <a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a> puts it, using this information indeed goes against the spirit of the competition.</p>\n<p>This unfortunately means all previous submissions will be invalidated, and you will need to submit new predictions using the new test set and weighting. We sincerely apologize for the changes and we are very grateful to <a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> for pointing this out. Thank you so much for your continued engagement and happy Kaggling!</p>\n<p>EDIT 1: Thanks to <a href=\"https://www.kaggle.com/ymatioun\" target=\"_blank\">@ymatioun</a> for originally pointing the location leak out.<br>\nEDIT 2: As <a href=\"https://www.kaggle.com/churkinnikita\" target=\"_blank\">@churkinnikita</a> points out, ptend_q0002 12-14 are no longer zeroed out.</p>",
  "messages": [
    {
      "id": 2878482,
      "postDate": "2024-06-19T01:19:25.940Z",
      "content": "<p>For those who don't want to download everything, you can use -f option</p>\n<pre><code>kaggle competitions download -c  leap-atmospheric-physics-ai-climsim -f sample_submission.csv\nkaggle competitions download -c  leap-atmospheric-physics-ai-climsim -f test.csv\n</code></pre>",
      "rawMarkdown": "For those who don't want to download everything, you can use -f option\n```shell\nkaggle competitions download -c  leap-atmospheric-physics-ai-climsim -f sample_submission.csv\nkaggle competitions download -c  leap-atmospheric-physics-ai-climsim -f test.csv\n```\n",
      "votes": 25,
      "replies": [
        {
          "id": 2882982,
          "postDate": "2024-06-21T17:18:57.650Z",
          "content": "<p>Thank you very much!</p>",
          "rawMarkdown": "Thank you very much!",
          "votes": 1
        }
      ]
    },
    {
      "id": 2878453,
      "postDate": "2024-06-19T00:30:57.173Z",
      "content": "<p>Hi Kagglers,</p>\n<p>We have some good news and bad news. The good news is that we are extending the competition by 2 weeks! The new deadline is <strong>July 15, 2024 11:59PM UTC</strong>. We are incredibly impressed with your work so far and the engaging discussion that has been going on between competing teams. It’s truly awe-inspiring.</p>\n<p>Now for the bad news: we are releasing a new test set and sample submission (weighting). As Kaggle user <a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> pointed out, it is possible to reverse-engineer location information from the original test set. Many users (very fairly) entered this competition with the assumption that it would be a column to column regression problem, not a multi-column to column regression problem. Because we don’t want to give an unfair advantage to users who discovered this exploit early, we are making the tough decision to release a fresh test set + weighting + solution. For those that have invested a great deal of time in the multi-column approach, you must have realized that the sample IDs were pre-scrambled and of no use. Knowing this, it is fairly obvious that the intention of the competition designers was to not grant you access to this information. As Kaggle user <a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a> puts it, using this information indeed goes against the spirit of the competition.</p>\n<p>This unfortunately means all previous submissions will be invalidated, and you will need to submit new predictions using the new test set and weighting. We sincerely apologize for the changes and we are very grateful to <a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> for pointing this out. Thank you so much for your continued engagement and happy Kaggling!</p>\n<p>EDIT 1: Thanks to <a href=\"https://www.kaggle.com/ymatioun\" target=\"_blank\">@ymatioun</a> for originally pointing the location leak out.<br>\nEDIT 2: As <a href=\"https://www.kaggle.com/churkinnikita\" target=\"_blank\">@churkinnikita</a> points out, ptend_q0002 12-14 are no longer zeroed out.</p>",
      "rawMarkdown": "Hi Kagglers,\n\nWe have some good news and bad news. The good news is that we are extending the competition by 2 weeks! The new deadline is **July 15, 2024 11:59PM UTC**. We are incredibly impressed with your work so far and the engaging discussion that has been going on between competing teams. It’s truly awe-inspiring.\n\nNow for the bad news: we are releasing a new test set and sample submission (weighting). As Kaggle user @tatamikenn pointed out, it is possible to reverse-engineer location information from the original test set. Many users (very fairly) entered this competition with the assumption that it would be a column to column regression problem, not a multi-column to column regression problem. Because we don’t want to give an unfair advantage to users who discovered this exploit early, we are making the tough decision to release a fresh test set + weighting + solution. For those that have invested a great deal of time in the multi-column approach, you must have realized that the sample IDs were pre-scrambled and of no use. Knowing this, it is fairly obvious that the intention of the competition designers was to not grant you access to this information. As Kaggle user @ryches puts it, using this information indeed goes against the spirit of the competition.\n\nThis unfortunately means all previous submissions will be invalidated, and you will need to submit new predictions using the new test set and weighting. We sincerely apologize for the changes and we are very grateful to @tatamikenn for pointing this out. Thank you so much for your continued engagement and happy Kaggling!\n\nEDIT 1: Thanks to @ymatioun for originally pointing the location leak out.\nEDIT 2: As @churkinnikita points out, ptend_q0002 12-14 are no longer zeroed out.",
      "votes": 30
    },
    {
      "id": 2878528,
      "postDate": "2024-06-19T03:23:51.943Z",
      "content": "<p>When I check the new sample_submission.csv, the weight of 12-14 in ptend_q0002 is 1 instead of 0. <br>\nIs this correct?</p>",
      "rawMarkdown": "When I check the new sample_submission.csv, the weight of 12-14 in ptend_q0002 is 1 instead of 0. \nIs this correct?",
      "votes": 21
    },
    {
      "id": 2878461,
      "postDate": "2024-06-19T00:57:38.087Z",
      "content": "<p>I want to mention that the original person to reveal that <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/511274#2868420\" target=\"_blank\">there was a leak</a> in test was <a href=\"https://www.kaggle.com/ymatioun\" target=\"_blank\">@ymatioun</a> so kudos to him too.</p>",
      "rawMarkdown": "I want to mention that the original person to reveal that [there was a leak](https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/511274#2868420) in test was @ymatioun so kudos to him too.",
      "votes": 20
    },
    {
      "id": 2878746,
      "postDate": "2024-06-19T06:16:07.517Z",
      "content": "<blockquote>\n  <p>We have some good news and bad news. The good news is that we are extending the competition by 2 weeks! </p>\n</blockquote>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5487737%2Fd89f52ecf20de68b91f838ebe7d18563%2Fshrek-shrek-rizz.gif?generation=1718777749835349&amp;alt=media\"></p>",
      "rawMarkdown": "> We have some good news and bad news. The good news is that we are extending the competition by 2 weeks! \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5487737%2Fd89f52ecf20de68b91f838ebe7d18563%2Fshrek-shrek-rizz.gif?generation=1718777749835349&alt=media)",
      "votes": 7
    },
    {
      "id": 2878951,
      "postDate": "2024-06-19T09:14:51.503Z",
      "content": "<p><a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> The weights of ptend_q0002 12–14 are now ones instead of zeros. Is this intended? This introduces a discrepancy between pre-update scoring and current scoring. Now we have three more \"extreme\" targets to predict.</p>",
      "rawMarkdown": "@jerrylin96 The weights of ptend_q0002 12–14 are now ones instead of zeros. Is this intended? This introduces a discrepancy between pre-update scoring and current scoring. Now we have three more \"extreme\" targets to predict.",
      "votes": 8,
      "replies": [
        {
          "id": 2878959,
          "postDate": "2024-06-19T09:22:12.443Z",
          "content": "<p>They are overwritten from -state_q0002 i / 1200 so it doesn't change anything. </p>",
          "rawMarkdown": "They are overwritten from -state_q0002 i / 1200 so it doesn't change anything. ",
          "votes": 3
        },
        {
          "id": 2879388,
          "postDate": "2024-06-19T14:28:41.747Z",
          "content": "<p>This is correct. Apologies for not mentioning this earlier. I will update the post accordingly.</p>",
          "rawMarkdown": "This is correct. Apologies for not mentioning this earlier. I will update the post accordingly.",
          "votes": 3,
          "replies": [
            {
              "id": 2879400,
              "postDate": "2024-06-19T14:46:33.460Z",
              "content": "<p><a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> In the context of Gunes' comment above, can you confirm whether you now want us to predict ptend_q0002_12-14? I'm a little confused, as I'm not sure what this update has to do with preventing the multi-column approach. Can you explain why this change has been made?</p>",
              "rawMarkdown": "@jerrylin96 In the context of Gunes' comment above, can you confirm whether you now want us to predict ptend_q0002_12-14? I'm a little confused, as I'm not sure what this update has to do with preventing the multi-column approach. Can you explain why this change has been made?",
              "votes": 2
            },
            {
              "id": 2879564,
              "postDate": "2024-06-19T17:02:55.933Z",
              "content": "<p>You're correct, that part of the update does not have anything to do with preventing the multi-column approach. This was an unintentional update spurred by addressing the other issue quickly.</p>\n<p>While it is unlikely we will reverse course on this, we can revisit this issue if the Kaggle community feels sufficiently strongly about zero'ing out those values again.</p>",
              "rawMarkdown": "You're correct, that part of the update does not have anything to do with preventing the multi-column approach. This was an unintentional update spurred by addressing the other issue quickly.\n\nWhile it is unlikely we will reverse course on this, we can revisit this issue if the Kaggle community feels sufficiently strongly about zero'ing out those values again."
            }
          ]
        }
      ]
    },
    {
      "id": 2878523,
      "postDate": "2024-06-19T03:15:23.287Z",
      "content": "<p><a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> Thanks for the update!</p>\n<p>Just to be clear.<br>\nWhat if someone predicts location, reconstructs map and aggregates location specific Information?(like lat, lon and more sophisticated)<br>\nIt's still possible even after this shuffling as top team says they can predict location with very high accuracy.</p>",
      "rawMarkdown": "@jerrylin96 Thanks for the update!\n\nJust to be clear.\nWhat if someone predicts location, reconstructs map and aggregates location specific Information?(like lat, lon and more sophisticated)\nIt's still possible even after this shuffling as top team says they can predict location with very high accuracy.",
      "votes": 5,
      "replies": [
        {
          "id": 2878590,
          "postDate": "2024-06-19T04:45:35.813Z",
          "content": "<p>I think models are already capturing that information from the current features. Explicitly feeding the location via an embedding might be giving the edge.</p>",
          "rawMarkdown": "I think models are already capturing that information from the current features. Explicitly feeding the location via an embedding might be giving the edge.",
          "votes": 2
        },
        {
          "id": 2878779,
          "postDate": "2024-06-19T06:42:31.990Z",
          "content": "<p>Even if you can predict with 100% accuracy latlon you will not be able to use it for multiple-columns-to-one model since the test set is also shuffled in time. And good luck predicting also time stamp with sufficient accuracy.</p>",
          "rawMarkdown": "Even if you can predict with 100% accuracy latlon you will not be able to use it for multiple-columns-to-one model since the test set is also shuffled in time. And good luck predicting also time stamp with sufficient accuracy.",
          "votes": 2
        },
        {
          "id": 2879029,
          "postDate": "2024-06-19T10:27:08.933Z",
          "content": "<p>I think explicitly aggregating location information will be more advantageous, rather than model feeding it, but will it be allowed ? I think its a very important question to be addressed</p>",
          "rawMarkdown": "I think explicitly aggregating location information will be more advantageous, rather than model feeding it, but will it be allowed ? I think its a very important question to be addressed",
          "votes": 3,
          "replies": [
            {
              "id": 2890460,
              "postDate": "2024-06-26T06:04:53.707Z",
              "content": "<p>I have the same question, is it allowed to restore the lat/lon information of the grid? <a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> </p>",
              "rawMarkdown": "I have the same question, is it allowed to restore the lat/lon information of the grid? @jerrylin96 "
            }
          ]
        }
      ]
    },
    {
      "id": 2878487,
      "postDate": "2024-06-19T01:26:37.927Z",
      "content": "<p><a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> Can you confirm that train.csv has not changed? Because I remember that the size of this file was about 180G before, no more than 200G, but now it is 359.31G.</p>",
      "rawMarkdown": "@jerrylin96 Can you confirm that train.csv has not changed? Because I remember that the size of this file was about 180G before, no more than 200G, but now it is 359.31G.",
      "votes": 4,
      "replies": [
        {
          "id": 2878511,
          "postDate": "2024-06-19T02:34:17.967Z",
          "content": "<p>i just downloaded the train data, and it's about 169G</p>",
          "rawMarkdown": "i just downloaded the train data, and it's about 169G",
          "votes": 1,
          "replies": [
            {
              "id": 2878520,
              "postDate": "2024-06-19T03:08:36.407Z",
              "content": "<p>Thanks, maybe it's just a display issue.</p>",
              "rawMarkdown": "Thanks, maybe it's just a display issue."
            }
          ]
        }
      ]
    },
    {
      "id": 2886907,
      "postDate": "2024-06-23T22:17:34.320Z",
      "content": "<p><a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> Can I ask the score difference of the benchmark model between new Public test and new Private test?<br>\nAre they close?</p>\n<p>c.f.) In your previous comment, it had 0.005 score difference between old public test and old private test.<br>\n<a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/506490#2834423\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/506490#2834423</a></p>",
      "rawMarkdown": "@jerrylin96 Can I ask the score difference of the benchmark model between new Public test and new Private test?\nAre they close?\n\nc.f.) In your previous comment, it had 0.005 score difference between old public test and old private test.\nhttps://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/506490#2834423",
      "votes": 1
    },
    {
      "id": 2878485,
      "postDate": "2024-06-19T01:23:33.163Z",
      "content": "<p>sample_submission.csv   weight file is   0 or 1,    The difference between the new weights and the previous ones is quite large, with a CV  of over 0.994. This needs to be confirmed - should we indeed adopt the new weights?</p>",
      "rawMarkdown": "sample_submission.csv   weight file is   0 or 1,    The difference between the new weights and the previous ones is quite large, with a CV  of over 0.994. This needs to be confirmed - should we indeed adopt the new weights?",
      "votes": 2,
      "replies": [
        {
          "id": 2878490,
          "postDate": "2024-06-19T01:35:14.710Z",
          "content": "<p>Yes, use the new weights.</p>",
          "rawMarkdown": "Yes, use the new weights."
        },
        {
          "id": 2879001,
          "postDate": "2024-06-19T10:13:01.857Z",
          "content": "<p>You're probably using the wrong calculation formula.<br>\nThe discussion below may be helpful.</p>\n<p><a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/495255\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/495255</a></p>",
          "rawMarkdown": "You're probably using the wrong calculation formula.\nThe discussion below may be helpful.\n\nhttps://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/495255",
          "votes": 3
        },
        {
          "id": 2880062,
          "postDate": "2024-06-20T02:45:55.670Z",
          "content": "<p>I've also encountered this problem. How did you solve it? I changed the calculation method, but it still seems to be 0.994. Maybe there's something wrong with my calculation?</p>",
          "rawMarkdown": "I've also encountered this problem. How did you solve it? I changed the calculation method, but it still seems to be 0.994. Maybe there's something wrong with my calculation?",
          "votes": 2,
          "replies": [
            {
              "id": 2880927,
              "postDate": "2024-06-20T13:12:04.917Z",
              "content": "<p>Use sklearn not pytorch implements </p>",
              "rawMarkdown": "Use sklearn not pytorch implements "
            },
            {
              "id": 2880949,
              "postDate": "2024-06-20T13:23:43.607Z",
              "content": "<p>I also use pytorch, what is wrong with pytorch? 😅</p>",
              "rawMarkdown": "I also use pytorch, what is wrong with pytorch? 😅"
            },
            {
              "id": 2881016,
              "postDate": "2024-06-20T13:45:35.403Z",
              "content": "<p>Thank you, I'll double-check</p>",
              "rawMarkdown": "Thank you, I'll double-check"
            },
            {
              "id": 2881784,
              "postDate": "2024-06-21T01:02:23.243Z",
              "content": "<p>I have checked. When torchmetrics calculates very small numbers, r2_score is 1.</p>",
              "rawMarkdown": "I have checked. When torchmetrics calculates very small numbers, r2_score is 1.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2884450,
      "postDate": "2024-06-22T14:08:25.930Z",
      "content": "<p>Project System is fast but lagging</p>",
      "rawMarkdown": "Project System is fast but lagging\n"
    },
    {
      "id": 2880096,
      "postDate": "2024-06-20T03:56:47.560Z",
      "content": "<p><a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> </p>\n<p>I'm new in this competition, and I have one question about r2 calculation.</p>\n<p>In the description page, the following statement can be seen:<br>\n<code>This weighting can also be calculated without downloading this file. The top 12 levels (0-11 inclusive) for 'ptend_q0001', 'ptend_q0002', 'ptend_q0003', 'ptend_u', and 'ptend_v' can be ignored.</code></p>\n<p>Is the r2 calculated with skipping these columns in submission stage?<br>\nIn other words, columns with weight=0 are skipped in r2 calculation?</p>",
      "rawMarkdown": "@jerrylin96 \n\nI'm new in this competition, and I have one question about r2 calculation.\n\nIn the description page, the following statement can be seen:\n`This weighting can also be calculated without downloading this file. The top 12 levels (0-11 inclusive) for 'ptend_q0001', 'ptend_q0002', 'ptend_q0003', 'ptend_u', and 'ptend_v' can be ignored.`\n\nIs the r2 calculated with skipping these columns in submission stage?\nIn other words, columns with weight=0 are skipped in r2 calculation?",
      "replies": [
        {
          "id": 2880390,
          "postDate": "2024-06-20T06:52:24.407Z",
          "content": "<p>yes, pretty much…</p>",
          "rawMarkdown": "yes, pretty much..."
        },
        {
          "id": 2880432,
          "postDate": "2024-06-20T07:32:37.540Z",
          "content": "<p>No. It is 1 for these cols.</p>",
          "rawMarkdown": "No. It is 1 for these cols.",
          "votes": 3,
          "replies": [
            {
              "id": 2881786,
              "postDate": "2024-06-21T01:08:29.960Z",
              "content": "<p>I don't understand what you mean, I checked these columns and they are 0.</p>",
              "rawMarkdown": "I don't understand what you mean, I checked these columns and they are 0."
            },
            {
              "id": 2881844,
              "postDate": "2024-06-21T02:52:32.243Z",
              "content": "<p>R2 value should be 1 for those columns with weight 0</p>",
              "rawMarkdown": "R2 value should be 1 for those columns with weight 0",
              "votes": 1
            },
            {
              "id": 2882139,
              "postDate": "2024-06-21T07:19:46.847Z",
              "content": "<p>Thanks everyone.<br>\nI have another question. For convenience, I split the case 1 and 2 depending on target columns.</p>\n<p><strong>1). About the top 12 levels (0-11 inclusive) for 'ptend_q0001', 'ptend_q0002', 'ptend_q0003', 'ptend_u', and 'ptend_v'</strong></p>\n<p>The reason why I asked the above question is that if y_true is approximately equal to y_pred, then ss_tot will become tiny in r2 calculation, leading to large negative values in each R2_i where i denotes each class.<br>\nIf the host set R2_i=1 for these columns, then there is no afraid in these targets.</p>\n<p><strong>2). About the other target columns having tiny values</strong><br>\nBut the situation is not resolved for other targets such as ptend_q0002_18, 19 etc which has tiny values.<br>\nSince R2 calculation is performed on class-wise and then averaged over all classes, very tiny values of ss_tot would be again appeared in some R2_i calculation. Though the problem is alleviated in <strong>our local calculation</strong> by using some tricks such as adding epsilon to ss_tot or standardization, these approach is not applicable to submission stage because of the difference of r2 calculation. </p>\n<p>from these points, my another question is:<br>\nHow should we avoid this negative explosion in some targets? </p>\n<p>p.s.</p>\n<ul>\n<li>since r2 is sensitive to ss_tot, adding-epsilon seems not good</li>\n<li>in my case, standardization before putting y into r2 func works on local calculation (I haven't submitted yet).</li>\n</ul>",
              "rawMarkdown": "Thanks everyone.\nI have another question. For convenience, I split the case 1 and 2 depending on target columns.\n\n**1). About the top 12 levels (0-11 inclusive) for 'ptend_q0001', 'ptend_q0002', 'ptend_q0003', 'ptend_u', and 'ptend_v'**\n\nThe reason why I asked the above question is that if y_true is approximately equal to y_pred, then ss_tot will become tiny in r2 calculation, leading to large negative values in each R2_i where i denotes each class.\nIf the host set R2_i=1 for these columns, then there is no afraid in these targets.\n\n**2). About the other target columns having tiny values**\nBut the situation is not resolved for other targets such as ptend_q0002_18, 19 etc which has tiny values.\nSince R2 calculation is performed on class-wise and then averaged over all classes, very tiny values of ss_tot would be again appeared in some R2_i calculation. Though the problem is alleviated in **our local calculation** by using some tricks such as adding epsilon to ss_tot or standardization, these approach is not applicable to submission stage because of the difference of r2 calculation. \n\nfrom these points, my another question is:\nHow should we avoid this negative explosion in some targets? \n\np.s.\n- since r2 is sensitive to ss_tot, adding-epsilon seems not good\n- in my case, standardization before putting y into r2 func works on local calculation (I haven't submitted yet)."
            },
            {
              "id": 2882358,
              "postDate": "2024-06-21T10:07:29.023Z",
              "content": "<p>There is a thread somewhere about different scoring methods of R2, i was always using the one that you first sum of ss_res then sum of ss_tot and then divide res/tot and whatever score i had in my validation sets i had also in my submissions (with the old weights, with the new weights i am currently still training but i am confident its going to be the same).</p>",
              "rawMarkdown": "There is a thread somewhere about different scoring methods of R2, i was always using the one that you first sum of ss_res then sum of ss_tot and then divide res/tot and whatever score i had in my validation sets i had also in my submissions (with the old weights, with the new weights i am currently still training but i am confident its going to be the same)."
            },
            {
              "id": 2882807,
              "postDate": "2024-06-21T15:13:26.243Z",
              "content": "<p>i think you were right. When i omit this code in inference i get very negative values, with this code i am at 0.71 ```<br>\nREPLACE_TO = ['ptend_q0002_0', 'ptend_q0002_1', 'ptend_q0002_2', 'ptend_q0002_3', 'ptend_q0002_4', 'ptend_q0002_5', 'ptend_q0002_6', 'ptend_q0002_7', 'ptend_q0002_8', 'ptend_q0002_9', 'ptend_q0002_10', 'ptend_q0002_11', 'ptend_q0002_12', 'ptend_q0002_13', 'ptend_q0002_14', 'ptend_q0002_15', 'ptend_q0002_16', 'ptend_q0002_17', 'ptend_q0002_18', 'ptend_q0002_19', 'ptend_q0002_20', 'ptend_q0002_21', 'ptend_q0002_22', 'ptend_q0002_23', 'ptend_q0002_24', 'ptend_q0002_25', 'ptend_q0002_26']<br>\nREPLACE_FROM = ['state_q0002_0', 'state_q0002_1', 'state_q0002_2', 'state_q0002_3', 'state_q0002_4', 'state_q0002_5', 'state_q0002_6', 'state_q0002_7', 'state_q0002_8', 'state_q0002_9', 'state_q0002_10', 'state_q0002_11', 'state_q0002_12', 'state_q0002_13', 'state_q0002_14', 'state_q0002_15', 'state_q0002_16', 'state_q0002_17', 'state_q0002_18', 'state_q0002_19', 'state_q0002_20', 'state_q0002_21', 'state_q0002_22', 'state_q0002_23', 'state_q0002_24', 'state_q0002_25', 'state_q0002_26']</p>\n<p>df_test = pd.read_csv(test_file)<br>\nfor idx in range(0, 27):<br>\n    df_p_test[f\"ptend_q0002_{idx}\"] = -df_test[f\"state_q0002_{idx}\"].to_numpy() / 1200<br>\n```</p>",
              "rawMarkdown": "i think you were right. When i omit this code in inference i get very negative values, with this code i am at 0.71 ```\nREPLACE_TO = ['ptend_q0002_0', 'ptend_q0002_1', 'ptend_q0002_2', 'ptend_q0002_3', 'ptend_q0002_4', 'ptend_q0002_5', 'ptend_q0002_6', 'ptend_q0002_7', 'ptend_q0002_8', 'ptend_q0002_9', 'ptend_q0002_10', 'ptend_q0002_11', 'ptend_q0002_12', 'ptend_q0002_13', 'ptend_q0002_14', 'ptend_q0002_15', 'ptend_q0002_16', 'ptend_q0002_17', 'ptend_q0002_18', 'ptend_q0002_19', 'ptend_q0002_20', 'ptend_q0002_21', 'ptend_q0002_22', 'ptend_q0002_23', 'ptend_q0002_24', 'ptend_q0002_25', 'ptend_q0002_26']\nREPLACE_FROM = ['state_q0002_0', 'state_q0002_1', 'state_q0002_2', 'state_q0002_3', 'state_q0002_4', 'state_q0002_5', 'state_q0002_6', 'state_q0002_7', 'state_q0002_8', 'state_q0002_9', 'state_q0002_10', 'state_q0002_11', 'state_q0002_12', 'state_q0002_13', 'state_q0002_14', 'state_q0002_15', 'state_q0002_16', 'state_q0002_17', 'state_q0002_18', 'state_q0002_19', 'state_q0002_20', 'state_q0002_21', 'state_q0002_22', 'state_q0002_23', 'state_q0002_24', 'state_q0002_25', 'state_q0002_26']\n\ndf_test = pd.read_csv(test_file)\nfor idx in range(0, 27):\n    df_p_test[f\"ptend_q0002_{idx}\"] = -df_test[f\"state_q0002_{idx}\"].to_numpy() / 1200\n```",
              "votes": 2
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2878482,
      "author_name": "ONODERA",
      "author_url": "",
      "post_date": "2024-06-19T01:19:25.940000",
      "content": "<p>For those who don't want to download everything, you can use -f option</p>\n<pre><code>kaggle competitions download -c  leap-atmospheric-physics-ai-climsim -f sample_submission.csv\nkaggle competitions download -c  leap-atmospheric-physics-ai-climsim -f test.csv\n</code></pre>",
      "votes": 25,
      "replies": [
        {
          "id": 2882982,
          "author_name": "James Kimura",
          "author_url": "",
          "post_date": "2024-06-21T17:18:57.650000",
          "content": "<p>Thank you very much!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2878528,
      "author_name": "kuto",
      "author_url": "",
      "post_date": "2024-06-19T03:23:51.943000",
      "content": "<p>When I check the new sample_submission.csv, the weight of 12-14 in ptend_q0002 is 1 instead of 0. <br>\nIs this correct?</p>",
      "votes": 21,
      "replies": []
    },
    {
      "id": 2878461,
      "author_name": "greySnow",
      "author_url": "",
      "post_date": "2024-06-19T00:57:38.087000",
      "content": "<p>I want to mention that the original person to reveal that <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/511274#2868420\" target=\"_blank\">there was a leak</a> in test was <a href=\"https://www.kaggle.com/ymatioun\" target=\"_blank\">@ymatioun</a> so kudos to him too.</p>",
      "votes": 20,
      "replies": []
    },
    {
      "id": 2878746,
      "author_name": "slime",
      "author_url": "",
      "post_date": "2024-06-19T06:16:07.517000",
      "content": "<blockquote>\n  <p>We have some good news and bad news. The good news is that we are extending the competition by 2 weeks! </p>\n</blockquote>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5487737%2Fd89f52ecf20de68b91f838ebe7d18563%2Fshrek-shrek-rizz.gif?generation=1718777749835349&amp;alt=media\"></p>",
      "votes": 7,
      "replies": []
    },
    {
      "id": 2878951,
      "author_name": "Nikita Churkin",
      "author_url": "",
      "post_date": "2024-06-19T09:14:51.503000",
      "content": "<p><a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> The weights of ptend_q0002 12–14 are now ones instead of zeros. Is this intended? This introduces a discrepancy between pre-update scoring and current scoring. Now we have three more \"extreme\" targets to predict.</p>",
      "votes": 8,
      "replies": [
        {
          "id": 2878959,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2024-06-19T09:22:12.443000",
          "content": "<p>They are overwritten from -state_q0002 i / 1200 so it doesn't change anything. </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 2879388,
          "author_name": "Jerry Lin",
          "author_url": "",
          "post_date": "2024-06-19T14:28:41.747000",
          "content": "<p>This is correct. Apologies for not mentioning this earlier. I will update the post accordingly.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2879400,
              "author_name": "Sarah Jeffreson",
              "author_url": "",
              "post_date": "2024-06-19T14:46:33.460000",
              "content": "<p><a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> In the context of Gunes' comment above, can you confirm whether you now want us to predict ptend_q0002_12-14? I'm a little confused, as I'm not sure what this update has to do with preventing the multi-column approach. Can you explain why this change has been made?</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2879564,
              "author_name": "Jerry Lin",
              "author_url": "",
              "post_date": "2024-06-19T17:02:55.933000",
              "content": "<p>You're correct, that part of the update does not have anything to do with preventing the multi-column approach. This was an unintentional update spurred by addressing the other issue quickly.</p>\n<p>While it is unlikely we will reverse course on this, we can revisit this issue if the Kaggle community feels sufficiently strongly about zero'ing out those values again.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2878523,
      "author_name": "Camaro",
      "author_url": "",
      "post_date": "2024-06-19T03:15:23.287000",
      "content": "<p><a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> Thanks for the update!</p>\n<p>Just to be clear.<br>\nWhat if someone predicts location, reconstructs map and aggregates location specific Information?(like lat, lon and more sophisticated)<br>\nIt's still possible even after this shuffling as top team says they can predict location with very high accuracy.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 2878590,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2024-06-19T04:45:35.813000",
          "content": "<p>I think models are already capturing that information from the current features. Explicitly feeding the location via an embedding might be giving the edge.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2878779,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2024-06-19T06:42:31.990000",
          "content": "<p>Even if you can predict with 100% accuracy latlon you will not be able to use it for multiple-columns-to-one model since the test set is also shuffled in time. And good luck predicting also time stamp with sufficient accuracy.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2879029,
          "author_name": "NikhilMishra",
          "author_url": "",
          "post_date": "2024-06-19T10:27:08.933000",
          "content": "<p>I think explicitly aggregating location information will be more advantageous, rather than model feeding it, but will it be allowed ? I think its a very important question to be addressed</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2890460,
              "author_name": "Xintong Zhao",
              "author_url": "",
              "post_date": "2024-06-26T06:04:53.707000",
              "content": "<p>I have the same question, is it allowed to restore the lat/lon information of the grid? <a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2878487,
      "author_name": "heng",
      "author_url": "",
      "post_date": "2024-06-19T01:26:37.927000",
      "content": "<p><a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> Can you confirm that train.csv has not changed? Because I remember that the size of this file was about 180G before, no more than 200G, but now it is 359.31G.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2878511,
          "author_name": "yelim421",
          "author_url": "",
          "post_date": "2024-06-19T02:34:17.967000",
          "content": "<p>i just downloaded the train data, and it's about 169G</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2878520,
              "author_name": "heng",
              "author_url": "",
              "post_date": "2024-06-19T03:08:36.407000",
              "content": "<p>Thanks, maybe it's just a display issue.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2886907,
      "author_name": "Bilzard",
      "author_url": "",
      "post_date": "2024-06-23T22:17:34.320000",
      "content": "<p><a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> Can I ask the score difference of the benchmark model between new Public test and new Private test?<br>\nAre they close?</p>\n<p>c.f.) In your previous comment, it had 0.005 score difference between old public test and old private test.<br>\n<a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/506490#2834423\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/506490#2834423</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2878485,
      "author_name": "lhwcv",
      "author_url": "",
      "post_date": "2024-06-19T01:23:33.163000",
      "content": "<p>sample_submission.csv   weight file is   0 or 1,    The difference between the new weights and the previous ones is quite large, with a CV  of over 0.994. This needs to be confirmed - should we indeed adopt the new weights?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2878490,
          "author_name": "Jerry Lin",
          "author_url": "",
          "post_date": "2024-06-19T01:35:14.710000",
          "content": "<p>Yes, use the new weights.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2879001,
          "author_name": "ohkawa3",
          "author_url": "",
          "post_date": "2024-06-19T10:13:01.857000",
          "content": "<p>You're probably using the wrong calculation formula.<br>\nThe discussion below may be helpful.</p>\n<p><a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/495255\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/495255</a></p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 2880062,
          "author_name": "Chengwei Yan",
          "author_url": "",
          "post_date": "2024-06-20T02:45:55.670000",
          "content": "<p>I've also encountered this problem. How did you solve it? I changed the calculation method, but it still seems to be 0.994. Maybe there's something wrong with my calculation?</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2880927,
              "author_name": "lhwcv",
              "author_url": "",
              "post_date": "2024-06-20T13:12:04.917000",
              "content": "<p>Use sklearn not pytorch implements </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2880949,
              "author_name": "Vasilis",
              "author_url": "",
              "post_date": "2024-06-20T13:23:43.607000",
              "content": "<p>I also use pytorch, what is wrong with pytorch? 😅</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2881016,
              "author_name": "Chengwei Yan",
              "author_url": "",
              "post_date": "2024-06-20T13:45:35.403000",
              "content": "<p>Thank you, I'll double-check</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2881784,
              "author_name": "ynhuhu",
              "author_url": "",
              "post_date": "2024-06-21T01:02:23.243000",
              "content": "<p>I have checked. When torchmetrics calculates very small numbers, r2_score is 1.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2884450,
      "author_name": "Devesh Kumar Gola",
      "author_url": "",
      "post_date": "2024-06-22T14:08:25.930000",
      "content": "<p>Project System is fast but lagging</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2880096,
      "author_name": "MBOOK",
      "author_url": "",
      "post_date": "2024-06-20T03:56:47.560000",
      "content": "<p><a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> </p>\n<p>I'm new in this competition, and I have one question about r2 calculation.</p>\n<p>In the description page, the following statement can be seen:<br>\n<code>This weighting can also be calculated without downloading this file. The top 12 levels (0-11 inclusive) for 'ptend_q0001', 'ptend_q0002', 'ptend_q0003', 'ptend_u', and 'ptend_v' can be ignored.</code></p>\n<p>Is the r2 calculated with skipping these columns in submission stage?<br>\nIn other words, columns with weight=0 are skipped in r2 calculation?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2880390,
          "author_name": "Juan D C F",
          "author_url": "",
          "post_date": "2024-06-20T06:52:24.407000",
          "content": "<p>yes, pretty much…</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2880432,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2024-06-20T07:32:37.540000",
          "content": "<p>No. It is 1 for these cols.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2881786,
              "author_name": "ynhuhu",
              "author_url": "",
              "post_date": "2024-06-21T01:08:29.960000",
              "content": "<p>I don't understand what you mean, I checked these columns and they are 0.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2881844,
              "author_name": "PC Jimmmy",
              "author_url": "",
              "post_date": "2024-06-21T02:52:32.243000",
              "content": "<p>R2 value should be 1 for those columns with weight 0</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2882139,
              "author_name": "MBOOK",
              "author_url": "",
              "post_date": "2024-06-21T07:19:46.847000",
              "content": "<p>Thanks everyone.<br>\nI have another question. For convenience, I split the case 1 and 2 depending on target columns.</p>\n<p><strong>1). About the top 12 levels (0-11 inclusive) for 'ptend_q0001', 'ptend_q0002', 'ptend_q0003', 'ptend_u', and 'ptend_v'</strong></p>\n<p>The reason why I asked the above question is that if y_true is approximately equal to y_pred, then ss_tot will become tiny in r2 calculation, leading to large negative values in each R2_i where i denotes each class.<br>\nIf the host set R2_i=1 for these columns, then there is no afraid in these targets.</p>\n<p><strong>2). About the other target columns having tiny values</strong><br>\nBut the situation is not resolved for other targets such as ptend_q0002_18, 19 etc which has tiny values.<br>\nSince R2 calculation is performed on class-wise and then averaged over all classes, very tiny values of ss_tot would be again appeared in some R2_i calculation. Though the problem is alleviated in <strong>our local calculation</strong> by using some tricks such as adding epsilon to ss_tot or standardization, these approach is not applicable to submission stage because of the difference of r2 calculation. </p>\n<p>from these points, my another question is:<br>\nHow should we avoid this negative explosion in some targets? </p>\n<p>p.s.</p>\n<ul>\n<li>since r2 is sensitive to ss_tot, adding-epsilon seems not good</li>\n<li>in my case, standardization before putting y into r2 func works on local calculation (I haven't submitted yet).</li>\n</ul>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2882358,
              "author_name": "Vasilis",
              "author_url": "",
              "post_date": "2024-06-21T10:07:29.023000",
              "content": "<p>There is a thread somewhere about different scoring methods of R2, i was always using the one that you first sum of ss_res then sum of ss_tot and then divide res/tot and whatever score i had in my validation sets i had also in my submissions (with the old weights, with the new weights i am currently still training but i am confident its going to be the same).</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2882807,
              "author_name": "Vasilis",
              "author_url": "",
              "post_date": "2024-06-21T15:13:26.243000",
              "content": "<p>i think you were right. When i omit this code in inference i get very negative values, with this code i am at 0.71 ```<br>\nREPLACE_TO = ['ptend_q0002_0', 'ptend_q0002_1', 'ptend_q0002_2', 'ptend_q0002_3', 'ptend_q0002_4', 'ptend_q0002_5', 'ptend_q0002_6', 'ptend_q0002_7', 'ptend_q0002_8', 'ptend_q0002_9', 'ptend_q0002_10', 'ptend_q0002_11', 'ptend_q0002_12', 'ptend_q0002_13', 'ptend_q0002_14', 'ptend_q0002_15', 'ptend_q0002_16', 'ptend_q0002_17', 'ptend_q0002_18', 'ptend_q0002_19', 'ptend_q0002_20', 'ptend_q0002_21', 'ptend_q0002_22', 'ptend_q0002_23', 'ptend_q0002_24', 'ptend_q0002_25', 'ptend_q0002_26']<br>\nREPLACE_FROM = ['state_q0002_0', 'state_q0002_1', 'state_q0002_2', 'state_q0002_3', 'state_q0002_4', 'state_q0002_5', 'state_q0002_6', 'state_q0002_7', 'state_q0002_8', 'state_q0002_9', 'state_q0002_10', 'state_q0002_11', 'state_q0002_12', 'state_q0002_13', 'state_q0002_14', 'state_q0002_15', 'state_q0002_16', 'state_q0002_17', 'state_q0002_18', 'state_q0002_19', 'state_q0002_20', 'state_q0002_21', 'state_q0002_22', 'state_q0002_23', 'state_q0002_24', 'state_q0002_25', 'state_q0002_26']</p>\n<p>df_test = pd.read_csv(test_file)<br>\nfor idx in range(0, 27):<br>\n    df_p_test[f\"ptend_q0002_{idx}\"] = -df_test[f\"state_q0002_{idx}\"].to_numpy() / 1200<br>\n```</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2878482": "For those who don't want to download everything, you can use -f option\n```shell\nkaggle competitions download -c  leap-atmospheric-physics-ai-climsim -f sample_submission.csv\nkaggle competitions download -c  leap-atmospheric-physics-ai-climsim -f test.csv\n```\n",
    "2878453": "Hi Kagglers,\n\nWe have some good news and bad news. The good news is that we are extending the competition by 2 weeks! The new deadline is **July 15, 2024 11:59PM UTC**. We are incredibly impressed with your work so far and the engaging discussion that has been going on between competing teams. It’s truly awe-inspiring.\n\nNow for the bad news: we are releasing a new test set and sample submission (weighting). As Kaggle user @tatamikenn pointed out, it is possible to reverse-engineer location information from the original test set. Many users (very fairly) entered this competition with the assumption that it would be a column to column regression problem, not a multi-column to column regression problem. Because we don’t want to give an unfair advantage to users who discovered this exploit early, we are making the tough decision to release a fresh test set + weighting + solution. For those that have invested a great deal of time in the multi-column approach, you must have realized that the sample IDs were pre-scrambled and of no use. Knowing this, it is fairly obvious that the intention of the competition designers was to not grant you access to this information. As Kaggle user @ryches puts it, using this information indeed goes against the spirit of the competition.\n\nThis unfortunately means all previous submissions will be invalidated, and you will need to submit new predictions using the new test set and weighting. We sincerely apologize for the changes and we are very grateful to @tatamikenn for pointing this out. Thank you so much for your continued engagement and happy Kaggling!\n\nEDIT 1: Thanks to @ymatioun for originally pointing the location leak out.\nEDIT 2: As @churkinnikita points out, ptend_q0002 12-14 are no longer zeroed out.",
    "2878528": "When I check the new sample_submission.csv, the weight of 12-14 in ptend_q0002 is 1 instead of 0. \nIs this correct?",
    "2878461": "I want to mention that the original person to reveal that [there was a leak](https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/511274#2868420) in test was @ymatioun so kudos to him too.",
    "2878746": "> We have some good news and bad news. The good news is that we are extending the competition by 2 weeks! \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5487737%2Fd89f52ecf20de68b91f838ebe7d18563%2Fshrek-shrek-rizz.gif?generation=1718777749835349&alt=media)",
    "2878951": "@jerrylin96 The weights of ptend_q0002 12–14 are now ones instead of zeros. Is this intended? This introduces a discrepancy between pre-update scoring and current scoring. Now we have three more \"extreme\" targets to predict.",
    "2878523": "@jerrylin96 Thanks for the update!\n\nJust to be clear.\nWhat if someone predicts location, reconstructs map and aggregates location specific Information?(like lat, lon and more sophisticated)\nIt's still possible even after this shuffling as top team says they can predict location with very high accuracy.",
    "2878487": "@jerrylin96 Can you confirm that train.csv has not changed? Because I remember that the size of this file was about 180G before, no more than 200G, but now it is 359.31G.",
    "2886907": "@jerrylin96 Can I ask the score difference of the benchmark model between new Public test and new Private test?\nAre they close?\n\nc.f.) In your previous comment, it had 0.005 score difference between old public test and old private test.\nhttps://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/506490#2834423",
    "2878485": "sample_submission.csv   weight file is   0 or 1,    The difference between the new weights and the previous ones is quite large, with a CV  of over 0.994. This needs to be confirmed - should we indeed adopt the new weights?",
    "2884450": "Project System is fast but lagging\n",
    "2880096": "@jerrylin96 \n\nI'm new in this competition, and I have one question about r2 calculation.\n\nIn the description page, the following statement can be seen:\n`This weighting can also be calculated without downloading this file. The top 12 levels (0-11 inclusive) for 'ptend_q0001', 'ptend_q0002', 'ptend_q0003', 'ptend_u', and 'ptend_v' can be ignored.`\n\nIs the r2 calculated with skipping these columns in submission stage?\nIn other words, columns with weight=0 are skipped in r2 calculation?"
  }
}