{
  "id": 511911,
  "title": "[Confirmation Request]: Availability of Reverse-engineered Timestamp & Location Data on the Test Set",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/511911",
  "author_name": "Bilzard",
  "post_date": "2024-06-12T15:39:14.040000",
  "votes": 76,
  "comment_count": 118,
  "views": 0,
  "content": "<p>Hi. Kaggle Team &amp; Host members,</p>\n<p>Related to previous discussion[1], we found location and timestamp data on the test set is possible to be reverse-engineered.<br>\nHowever, in the early days of this competition, one of host member commented on the availability of these data, and answered they don't expect to utilize them[2]. So we want to make sure using these information is valid on this competition.</p>\n<p>Could you confirm we can utilize these information?</p>\n<p><a href=\"https://www.kaggle.com/ashleychow\" target=\"_blank\">@ashleychow</a> <a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a></p>\n<h2>How to reproduce</h2>\n<ol>\n<li>drop <code>sample_id</code></li>\n<li>add new column <code>sample_id</code> based on the current order of test data</li>\n</ol>\n<p>Then, we found the processed test data is ordered by each 384 locations and time steps.</p>\n<pre><code> polars  pl\n\nprefix = \ntest = (\n    pl.read_csv(Path())\n    .drop()\n    .with_row_index()\n    .with_columns(pl.col().mod().alias())\n)\n</code></pre>\n<h2>Evidence</h2>\n<p>When taking statistics per reverse-engineered location IDs, we can clearly see specific patterns which found on ordered train/validation data.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F140f93019fc2b9fc55c2ea06979620b3%2FScreenshot%202024-06-13%20at%200.42.31.png?generation=1718206981274040&amp;alt=media\"></p>\n<h2>Reference</h2>\n<ol>\n<li><a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/511274#286771\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/511274#286771</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/495258#2768949\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/495258#2768949</a></li>\n</ol>",
  "messages": [
    {
      "id": 2868744,
      "postDate": "2024-06-12T15:39:14.040Z",
      "content": "<p>Hi. Kaggle Team &amp; Host members,</p>\n<p>Related to previous discussion[1], we found location and timestamp data on the test set is possible to be reverse-engineered.<br>\nHowever, in the early days of this competition, one of host member commented on the availability of these data, and answered they don't expect to utilize them[2]. So we want to make sure using these information is valid on this competition.</p>\n<p>Could you confirm we can utilize these information?</p>\n<p><a href=\"https://www.kaggle.com/ashleychow\" target=\"_blank\">@ashleychow</a> <a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a></p>\n<h2>How to reproduce</h2>\n<ol>\n<li>drop <code>sample_id</code></li>\n<li>add new column <code>sample_id</code> based on the current order of test data</li>\n</ol>\n<p>Then, we found the processed test data is ordered by each 384 locations and time steps.</p>\n<pre><code> polars  pl\n\nprefix = \ntest = (\n    pl.read_csv(Path())\n    .drop()\n    .with_row_index()\n    .with_columns(pl.col().mod().alias())\n)\n</code></pre>\n<h2>Evidence</h2>\n<p>When taking statistics per reverse-engineered location IDs, we can clearly see specific patterns which found on ordered train/validation data.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F140f93019fc2b9fc55c2ea06979620b3%2FScreenshot%202024-06-13%20at%200.42.31.png?generation=1718206981274040&amp;alt=media\"></p>\n<h2>Reference</h2>\n<ol>\n<li><a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/511274#286771\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/511274#286771</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/495258#2768949\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/495258#2768949</a></li>\n</ol>",
      "rawMarkdown": "Hi. Kaggle Team & Host members,\n\nRelated to previous discussion[1], we found location and timestamp data on the test set is possible to be reverse-engineered.\nHowever, in the early days of this competition, one of host member commented on the availability of these data, and answered they don't expect to utilize them[2]. So we want to make sure using these information is valid on this competition.\n\nCould you confirm we can utilize these information?\n\n@ashleychow @jerrylin96\n\n\n## How to reproduce\n\n1. drop `sample_id`\n2. add new column `sample_id` based on the current order of test data\n\nThen, we found the processed test data is ordered by each 384 locations and time steps.\n\n```python\nimport polars as pl\n\nprefix = \"test\"\ntest = (\n    pl.read_csv(Path(\"../../../input/leap-atmospheric-physics-ai-climsim/test.csv\"))\n    .drop(\"sample_id\")\n    .with_row_index(\"sample_id\")\n    .with_columns(pl.col(\"sample_id\").mod(384).alias(\"grid_id\"))\n)\n```\n\n## Evidence\n\nWhen taking statistics per reverse-engineered location IDs, we can clearly see specific patterns which found on ordered train/validation data.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F140f93019fc2b9fc55c2ea06979620b3%2FScreenshot%202024-06-13%20at%200.42.31.png?generation=1718206981274040&alt=media)\n\n## Reference\n\n1. https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/511274#286771\n2. https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/495258#2768949",
      "votes": 76
    },
    {
      "id": 2869607,
      "postDate": "2024-06-13T06:49:38.677Z",
      "content": "<p>We are all likely already using geographical information whether the data are shuffled or not.</p>\n<p>The atmosphere has a consistent structure that can be related to geographical location. For example, tropopause height alone is a fairly good predictor of distance from the equator (|latitude|). Some of the per-column predictors also obviously convey information about location (e.g. SZA, land fraction, and sea ice fraction).</p>\n<p>This evening, I tried training a 4-layer MLP classifier to predict the E3SM grid-point that a sample originated from. It reaches ~96% accuracy after one pass through the “low-res” dataset and could surely be improved with minimal effort and additional training.</p>\n<p>My point is that lat/lon information is already encoded in the individual rows of the test set and shuffling the data won't remove it. A relatively simple MLP can extract locations so, to the extent that this information is relevant to the problem (and I think that it is), those of us using more sophisticated models are likely leveraging this information already without providing it explicitly as an input.</p>",
      "rawMarkdown": "We are all likely already using geographical information whether the data are shuffled or not.\n\nThe atmosphere has a consistent structure that can be related to geographical location. For example, tropopause height alone is a fairly good predictor of distance from the equator (|latitude|). Some of the per-column predictors also obviously convey information about location (e.g. SZA, land fraction, and sea ice fraction).\n\nThis evening, I tried training a 4-layer MLP classifier to predict the E3SM grid-point that a sample originated from. It reaches ~96% accuracy after one pass through the “low-res” dataset and could surely be improved with minimal effort and additional training.\n\nMy point is that lat/lon information is already encoded in the individual rows of the test set and shuffling the data won't remove it. A relatively simple MLP can extract locations so, to the extent that this information is relevant to the problem (and I think that it is), those of us using more sophisticated models are likely leveraging this information already without providing it explicitly as an input.",
      "votes": 35,
      "replies": [
        {
          "id": 2869631,
          "postDate": "2024-06-13T06:56:35.890Z",
          "content": "<p>I see. Thank you for providing us the domain knowledge.</p>",
          "rawMarkdown": "I see. Thank you for providing us the domain knowledge.",
          "replies": [
            {
              "id": 2869637,
              "postDate": "2024-06-13T06:58:59.690Z",
              "content": "<blockquote>\n  <p>This evening, I tried training a 4-layer MLP classifier to predict the E3SM grid-point that a sample originated from. It reaches ~96% accuracy after one pass through the “low-res” dataset and could surely be improved with minimal effort and additional training.</p>\n</blockquote>\n<p>It corresponds with the fact many participant already utilizing lat/lon feature report that they fail to improves their model.</p>",
              "rawMarkdown": "> This evening, I tried training a 4-layer MLP classifier to predict the E3SM grid-point that a sample originated from. It reaches ~96% accuracy after one pass through the “low-res” dataset and could surely be improved with minimal effort and additional training.\n\nIt corresponds with the fact many participant already utilizing lat/lon feature report that they fail to improves their model."
            }
          ]
        },
        {
          "id": 2870262,
          "postDate": "2024-06-13T13:41:25.500Z",
          "content": "<p>Just curious, are you using latlon or timestamp? Can we get 0.80 on LB without those?</p>",
          "rawMarkdown": "Just curious, are you using latlon or timestamp? Can we get 0.80 on LB without those?",
          "votes": 3
        },
        {
          "id": 2874007,
          "postDate": "2024-06-16T00:12:10.927Z",
          "content": "<p>I agree with this… if geographical location is implicit in the data, a good, complex ML would capture that implicitly. Also, we know the test set gets swapped at the end of the comp, so overfitting to the current test set might actually be detrimental in the long run.</p>",
          "rawMarkdown": "I agree with this… if geographical location is implicit in the data, a good, complex ML would capture that implicitly. Also, we know the test set gets swapped at the end of the comp, so overfitting to the current test set might actually be detrimental in the long run.",
          "replies": [
            {
              "id": 2874557,
              "postDate": "2024-06-16T11:42:09.383Z",
              "content": "<p>Who said that test set get swapped at the end of comp ? 🤔</p>",
              "rawMarkdown": "Who said that test set get swapped at the end of comp ? 🤔",
              "votes": 1
            },
            {
              "id": 2875179,
              "postDate": "2024-06-16T23:45:49.413Z",
              "content": "<p>Hmmm. I could have sworn I read something about it, but maybe I was confused with some other competition. This is my first Kaggle comp, but I was under the impression that for the final score they either swap or extend the test set with new data, so you should avoid overfitting to the current test set. But I couldn't find anything about that for this competition, so maybe this isn't the case.</p>",
              "rawMarkdown": "Hmmm. I could have sworn I read something about it, but maybe I was confused with some other competition. This is my first Kaggle comp, but I was under the impression that for the final score they either swap or extend the test set with new data, so you should avoid overfitting to the current test set. But I couldn't find anything about that for this competition, so maybe this isn't the case.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2876590,
      "postDate": "2024-06-17T21:31:18.187Z",
      "content": "<p><strong>UPDATE</strong>: Some big changes to the competition are happening tomorrow.</p>\n<p>1.) There's going to be a <strong>brand new test-set, sample submission (weighting file), and benchmark submission</strong>.<br>\n2.) The competition <strong>deadline will be extended two weeks</strong> for everyone.<br>\n3.) All current submissions will be invalidated. <strong>You must resubmit using the new data</strong> to appear on the leaderboard.<br>\n4.) If you have been using the single-atmospheric-column to single-atmospheric-column regression approach, your score shouldn't change much.<br>\n5.) <em>A more detailed post will appear tomorrow</em>.</p>\n<p>For those who have been using the multi-column to single-column approach, we apologize. However, making full use of this approach would have required knowing the sample IDs were scrambled, implying you knew this was \"against the spirit of the competition\" to quote <a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a> . </p>\n<p>TLDR: we cannot have multi-atmospheric-column to single-atmospheric-column and single-atmospheric-column to single-atmospheric-column approaches coexisting on the leaderboard, especially if the former was done knowing it went against the intentions of the competition.</p>\n<p>EDIT: By column I mean \"atmospheric column\" not csv column. Sorry if I scared you!</p>",
      "rawMarkdown": "**UPDATE**: Some big changes to the competition are happening tomorrow.\n\n1.) There's going to be a **brand new test-set, sample submission (weighting file), and benchmark submission**.\n2.) The competition **deadline will be extended two weeks** for everyone.\n3.) All current submissions will be invalidated. **You must resubmit using the new data** to appear on the leaderboard.\n4.) If you have been using the single-atmospheric-column to single-atmospheric-column regression approach, your score shouldn't change much.\n5.) *A more detailed post will appear tomorrow*.\n\nFor those who have been using the multi-column to single-column approach, we apologize. However, making full use of this approach would have required knowing the sample IDs were scrambled, implying you knew this was \"against the spirit of the competition\" to quote @ryches . \n\nTLDR: we cannot have multi-atmospheric-column to single-atmospheric-column and single-atmospheric-column to single-atmospheric-column approaches coexisting on the leaderboard, especially if the former was done knowing it went against the intentions of the competition.\n\nEDIT: By column I mean \"atmospheric column\" not csv column. Sorry if I scared you!",
      "votes": 20,
      "replies": [
        {
          "id": 2876609,
          "postDate": "2024-06-17T22:00:26.023Z",
          "content": "<p>I think you meant single-row and multi-row? I got scared for a second :D</p>",
          "rawMarkdown": "I think you meant single-row and multi-row? I got scared for a second :D",
          "votes": 2,
          "replies": [
            {
              "id": 2876633,
              "postDate": "2024-06-17T22:23:42.683Z",
              "content": "<p>Sorry by column, I mean \"atmospheric column\"</p>",
              "rawMarkdown": "Sorry by column, I mean \"atmospheric column\"",
              "votes": 4
            },
            {
              "id": 2876636,
              "postDate": "2024-06-17T22:25:59.917Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        },
        {
          "id": 2876661,
          "postDate": "2024-06-17T23:29:30.027Z",
          "content": "<p>Thank you for the update. I hope you double-check this time that the data is scrambled for good and there is no leak…😅 Also I would like to know if the new test is from the same years as the old test (only different downsampled split) of from a totally new years. (And if so- what years?)</p>",
          "rawMarkdown": "Thank you for the update. I hope you double-check this time that the data is scrambled for good and there is no leak...😅 Also I would like to know if the new test is from the same years as the old test (only different downsampled split) of from a totally new years. (And if so- what years?)",
          "votes": 4
        },
        {
          "id": 2876673,
          "postDate": "2024-06-17T23:36:57.653Z",
          "content": "<blockquote>\n  <p>For those who have been using the multi-column to single-column approach, we apologize. However, making full use of this approach would have required knowing the sample IDs were scrambled, implying you knew this was \"against the spirit of the competition\" to quote <a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a></p>\n</blockquote>\n<p>Well I guess it's a lesson, if you see a leak report instead of exploit (though the Kaggle way seems to be more on the exploit side lol) 😅</p>",
          "rawMarkdown": "> For those who have been using the multi-column to single-column approach, we apologize. However, making full use of this approach would have required knowing the sample IDs were scrambled, implying you knew this was \"against the spirit of the competition\" to quote @ryches\n\nWell I guess it's a lesson, if you see a leak report instead of exploit (though the Kaggle way seems to be more on the exploit side lol) 😅",
          "votes": 2
        },
        {
          "id": 2876675,
          "postDate": "2024-06-17T23:37:31.733Z",
          "content": "<p><a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> Thank you for the announcement.</p>\n<p>Could you also confirm the validity of solution which <em>estimates</em> location and timestamp using below data?</p>\n<ol>\n<li>currently released test.csv</li>\n<li>newly releasing test.csv</li>\n</ol>\n<p>c.f. (my previous comment)<br>\n<a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/511911#2870998\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/511911#2870998</a></p>",
          "rawMarkdown": "@jerrylin96 Thank you for the announcement.\n\nCould you also confirm the validity of solution which *estimates* location and timestamp using below data?\n\n1. currently released test.csv\n2. newly releasing test.csv\n\nc.f. (my previous comment)\nhttps://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/511911#2870998",
          "votes": 7
        },
        {
          "id": 2876709,
          "postDate": "2024-06-18T00:31:51.363Z",
          "content": "<p>r2_score is not sensitive to the scale, so why change weighting file?</p>",
          "rawMarkdown": "r2_score is not sensitive to the scale, so why change weighting file?",
          "votes": 2,
          "replies": [
            {
              "id": 2876717,
              "postDate": "2024-06-18T01:11:10.783Z",
              "content": "<p>For convenience purposes (on my end)</p>",
              "rawMarkdown": "For convenience purposes (on my end)",
              "votes": 1
            },
            {
              "id": 2876742,
              "postDate": "2024-06-18T01:37:09.253Z",
              "content": "<p>Ok, got it.</p>",
              "rawMarkdown": "Ok, got it."
            },
            {
              "id": 2877463,
              "postDate": "2024-06-18T11:36:52.513Z",
              "content": "<p>hey, since you are changing the weights anyway, how about you exclude all columns where weights are very high - like &gt;10^11? Then we would not have to fill them with -x/1200 or 0. Just make those weights 0. I think that would be in the spirit of the objective of this competition - predicting results only where they make sense.<br>\nOtherwise somebody may find a way to reverse engineer those weights and hack the results again.</p>",
              "rawMarkdown": "hey, since you are changing the weights anyway, how about you exclude all columns where weights are very high - like >10^11? Then we would not have to fill them with -x/1200 or 0. Just make those weights 0. I think that would be in the spirit of the objective of this competition - predicting results only where they make sense.\nOtherwise somebody may find a way to reverse engineer those weights and hack the results again.",
              "votes": 3
            }
          ]
        },
        {
          "id": 2876757,
          "postDate": "2024-06-18T01:58:55.953Z",
          "content": "<p>I think two weeks is too long, I'm curious what others think. The training data won't change, and this competition has a high correlation between CV and LB, so I don't think it will change what we need to do. Wouldn't a one-week extension be enough?</p>",
          "rawMarkdown": "I think two weeks is too long, I'm curious what others think. The training data won't change, and this competition has a high correlation between CV and LB, so I don't think it will change what we need to do. Wouldn't a one-week extension be enough?",
          "votes": 2,
          "replies": [
            {
              "id": 2876830,
              "postDate": "2024-06-18T03:41:30.790Z",
              "content": "<p>I think they made it 2 weeks for the \"someone\" who is using the location/timestamp information, but two or three days extension is enough for me. I was joined from the beginning of the competition and I am exhausted😭.  </p>",
              "rawMarkdown": "I think they made it 2 weeks for the \"someone\" who is using the location/timestamp information, but two or three days extension is enough for me. I was joined from the beginning of the competition and I am exhausted😭.  ",
              "votes": 7
            },
            {
              "id": 2876844,
              "postDate": "2024-06-18T03:56:51.957Z",
              "content": "<p>I agree with the one week timeline.<br>\nTwo weeks would veer the competition unnecessarily towards \"who has the most compute\".</p>",
              "rawMarkdown": "I agree with the one week timeline.\nTwo weeks would veer the competition unnecessarily towards \"who has the most compute\".",
              "votes": 2
            },
            {
              "id": 2876854,
              "postDate": "2024-06-18T04:08:21.520Z",
              "content": "<p>Personally, I have no counter opinion to the length of the extension, however, I think deciding by hearing with few people in this discussion is not fair for most of participants not here. So I think it’s the best we obey the decision what host made (i.e. 2 weeks extension)</p>",
              "rawMarkdown": "Personally, I have no counter opinion to the length of the extension, however, I think deciding by hearing with few people in this discussion is not fair for most of participants not here. So I think it’s the best we obey the decision what host made (i.e. 2 weeks extension)",
              "votes": 7
            }
          ]
        },
        {
          "id": 2876762,
          "postDate": "2024-06-18T02:06:26.697Z",
          "content": "<p>Would you disqualify the winner if he/she used the location or timestamp feature?</p>",
          "rawMarkdown": "Would you disqualify the winner if he/she used the location or timestamp feature?",
          "votes": 6
        },
        {
          "id": 2876816,
          "postDate": "2024-06-18T03:23:44.543Z",
          "content": "<p><a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> </p>\n<p>Thank you for the update! I'm impressed by your quick decision-making despite the scale of these changes.</p>\n<p>I have a proposal regarding the update method:</p>\n<ul>\n<li>Release the current test set with the correct labels.</li>\n<li>Create a new test set from the next year's data, with the sample_ids completely shuffled.</li>\n<li>Allow the use of location and timestamps for the training set, including the data published on Huggingface.</li>\n</ul>\n<p>I believe this is the best way to frame the problem you want to solve without significantly changing its nature.  <br>\nIt's especially important to shift the timestamps between the current and new test sets to prevent unproductive solutions like matching between the two sets.</p>",
          "rawMarkdown": "@jerrylin96 \n\nThank you for the update! I'm impressed by your quick decision-making despite the scale of these changes.\n\nI have a proposal regarding the update method:\n\n- Release the current test set with the correct labels.\n- Create a new test set from the next year's data, with the sample_ids completely shuffled.\n- Allow the use of location and timestamps for the training set, including the data published on Huggingface.\n\nI believe this is the best way to frame the problem you want to solve without significantly changing its nature.  \nIt's especially important to shift the timestamps between the current and new test sets to prevent unproductive solutions like matching between the two sets.",
          "votes": -4,
          "replies": [
            {
              "id": 2877629,
              "postDate": "2024-06-18T13:34:45.803Z",
              "content": "<blockquote>\n  <p>Allow the use of location and timestamps for the training set, including the data published on Huggingface.</p>\n</blockquote>\n<p>It is already allowed by virtue of not being forbidden…</p>",
              "rawMarkdown": ">Allow the use of location and timestamps for the training set, including the data published on Huggingface.\n\nIt is already allowed by virtue of not being forbidden...",
              "votes": 1
            },
            {
              "id": 2877633,
              "postDate": "2024-06-18T13:37:03.713Z",
              "content": "<p>Wait, why does this comment have so many downvotes?😢<br>\nWhat do you have against this proposal? I seriously have no idea…</p>",
              "rawMarkdown": "Wait, why does this comment have so many downvotes?😢\nWhat do you have against this proposal? I seriously have no idea...",
              "votes": 1
            },
            {
              "id": 2877646,
              "postDate": "2024-06-18T13:43:37.397Z",
              "content": "<p>lol I wondered the same. I am not one of them just to be clear…</p>",
              "rawMarkdown": "lol I wondered the same. I am not one of them just to be clear..."
            },
            {
              "id": 2877652,
              "postDate": "2024-06-18T13:45:18.963Z",
              "content": "<p><a href=\"https://www.kaggle.com/shlomoron\" target=\"_blank\">@shlomoron</a> <br>\nWhat I wanted to propose was to refrain from banning location and timestamps. Completely banning them is impractical as they cannot be detected, and we would have to rely on the integrity of the participants.  <br>\nInstead, I suggested continuing to allow their use for the train set, but reconstructing the test set so that the information cannot be identified from the IDs at the very least.</p>",
              "rawMarkdown": "@shlomoron \nWhat I wanted to propose was to refrain from banning location and timestamps. Completely banning them is impractical as they cannot be detected, and we would have to rely on the integrity of the participants.  \nInstead, I suggested continuing to allow their use for the train set, but reconstructing the test set so that the information cannot be identified from the IDs at the very least."
            },
            {
              "id": 2877653,
              "postDate": "2024-06-18T13:47:43.927Z",
              "content": "<p><a href=\"https://www.kaggle.com/shlomoron\" target=\"_blank\">@shlomoron</a> <br>\nHaha, thanks!<br>\nI'm not angry with them, just curious.. 🤔</p>",
              "rawMarkdown": "@shlomoron \nHaha, thanks!\nI'm not angry with them, just curious.. 🤔"
            },
            {
              "id": 2877660,
              "postDate": "2024-06-18T13:57:40.300Z",
              "content": "<p>New labels means re-training, I would rather avoid that ;D</p>",
              "rawMarkdown": "New labels means re-training, I would rather avoid that ;D",
              "votes": 1
            },
            {
              "id": 2877667,
              "postDate": "2024-06-18T14:04:06.017Z",
              "content": "<p>Well, they weren't banned yet.<br>\nIt's not like hosts don't know we have access to train spacetime. It's not some secret. only test data was shuffled and obviously meant to not have spacetime data.</p>",
              "rawMarkdown": "Well, they weren't banned yet.\nIt's not like hosts don't know we have access to train spacetime. It's not some secret. only test data was shuffled and obviously meant to not have spacetime data.",
              "votes": 1
            },
            {
              "id": 2877673,
              "postDate": "2024-06-18T14:08:52.450Z",
              "content": "<p><a href=\"https://www.kaggle.com/martynoveduard\" target=\"_blank\">@martynoveduard</a> <br>\nAh, that part? But even if there's no label, we might use the old test set with pseudo-labeling or something like that. Especially if the new test set is sampled from the same timestamp, that kind of approach would shine.<br>\nI feel it's a more meaningless effort. 🤔<br>\nBut thanks for the feedback;)</p>",
              "rawMarkdown": "@martynoveduard \nAh, that part? But even if there's no label, we might use the old test set with pseudo-labeling or something like that. Especially if the new test set is sampled from the same timestamp, that kind of approach would shine.\nI feel it's a more meaningless effort. 🤔\nBut thanks for the feedback;)"
            },
            {
              "id": 2877720,
              "postDate": "2024-06-18T14:34:39.877Z",
              "content": "<p>Releasing the ground truth of a test set before the competition ends, even if it's an old test set with an undesirable side effect, is not a good thing to do.  The probability of <a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> doing that is at most zero.</p>",
              "rawMarkdown": "Releasing the ground truth of a test set before the competition ends, even if it's an old test set with an undesirable side effect, is not a good thing to do.  The probability of @jerrylin96 doing that is at most zero.",
              "votes": 1
            },
            {
              "id": 2877744,
              "postDate": "2024-06-18T14:47:26.520Z",
              "content": "<p>Well, it's almost a synthetic dataset, so there is no inherent difference between train and test by nature. But if the host creates a new test set for the next year to avoid overlap between the old and new test sets, there will be a larger gap between train and test, which (maybe) makes prediction more difficult.</p>\n<p>But I'm now understanding everyone's concerns. Thanks!</p>",
              "rawMarkdown": "Well, it's almost a synthetic dataset, so there is no inherent difference between train and test by nature. But if the host creates a new test set for the next year to avoid overlap between the old and new test sets, there will be a larger gap between train and test, which (maybe) makes prediction more difficult.\n\nBut I'm now understanding everyone's concerns. Thanks!"
            }
          ]
        },
        {
          "id": 2876930,
          "postDate": "2024-06-18T04:51:18.290Z",
          "content": "<p><a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> <br>\nThank you for update.<br>\nAnd I have a question.</p>\n<ul>\n<li>What does <code>multi-atmospheric-column to single-atmospheric-column and single-atmospheric-column to single-atmospheric-column approaches</code> mean?</li>\n</ul>",
          "rawMarkdown": "@jerrylin96 \nThank you for update.\nAnd I have a question.\n- What does `multi-atmospheric-column to single-atmospheric-column and single-atmospheric-column to single-atmospheric-column approaches` mean?\n",
          "votes": 2,
          "replies": [
            {
              "id": 2877576,
              "postDate": "2024-06-18T13:05:46.927Z",
              "content": "<p>essentially it means that methods that explicitly use location information (lat/lon) are against the rules of the competition</p>",
              "rawMarkdown": "essentially it means that methods that explicitly use location information (lat/lon) are against the rules of the competition",
              "votes": 1
            },
            {
              "id": 2877613,
              "postDate": "2024-06-18T13:19:58.810Z",
              "content": "<p>No. The rules are the rules (under the tab 'Rules') and nothing in the rules says anything like that. We have train spacetime labels and we definitely can use them. (for example for pseudo-labeling etc.)<br>\nIf host want to forbid it, he needs to do it clearly and formally by changing the rules. Which he did not.</p>",
              "rawMarkdown": "No. The rules are the rules (under the tab 'Rules') and nothing in the rules says anything like that. We have train spacetime labels and we definitely can use them. (for example for pseudo-labeling etc.)\nIf host want to forbid it, he needs to do it clearly and formally by changing the rules. Which he did not.",
              "votes": 2
            },
            {
              "id": 2878313,
              "postDate": "2024-06-18T21:21:21.447Z",
              "content": "<p>The entire purpose of this update is to remove multi-atmospheric-column approaches (e.g. 2D to 2D or 2D to 1D regression models) from the leaderboard.</p>",
              "rawMarkdown": "The entire purpose of this update is to remove multi-atmospheric-column approaches (e.g. 2D to 2D or 2D to 1D regression models) from the leaderboard.",
              "votes": 3
            }
          ]
        },
        {
          "id": 2876994,
          "postDate": "2024-06-18T05:29:23.733Z",
          "content": "<p><a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> , thank you for update.</p>\n<p>Could you elaborate further on the new sample submission (weighting file)? In my opinion, only the first column (sample_id) should change, while the weights for each target should remain unchanged. If this is not the case, assumptions might be made about the new test set.</p>",
          "rawMarkdown": "@jerrylin96 , thank you for update.\n\nCould you elaborate further on the new sample submission (weighting file)? In my opinion, only the first column (sample_id) should change, while the weights for each target should remain unchanged. If this is not the case, assumptions might be made about the new test set.",
          "replies": [
            {
              "id": 2877025,
              "postDate": "2024-06-18T06:16:42.213Z",
              "content": "<p>The only purpose of the sample submission is to zero out stratospheric moistening variables that are extremely difficult to predict and cause abnormally low r squared values. I made the rest of the values 1 for convenience sake (they weren't 1 before).</p>",
              "rawMarkdown": "The only purpose of the sample submission is to zero out stratospheric moistening variables that are extremely difficult to predict and cause abnormally low r squared values. I made the rest of the values 1 for convenience sake (they weren't 1 before).",
              "votes": 8
            }
          ]
        },
        {
          "id": 2877747,
          "postDate": "2024-06-18T14:49:10.040Z",
          "content": "<p>Suggest you pin this to the top </p>",
          "rawMarkdown": "Suggest you pin this to the top ",
          "votes": 1
        }
      ]
    },
    {
      "id": 2868875,
      "postDate": "2024-06-12T17:26:02.667Z",
      "content": "<p>Looking into this now. Thank you for the heads up---this is alarming.</p>",
      "rawMarkdown": "Looking into this now. Thank you for the heads up---this is alarming.",
      "votes": 16,
      "replies": [
        {
          "id": 2875441,
          "postDate": "2024-06-17T06:32:47.913Z",
          "rawMarkdown": "",
          "votes": -1,
          "isDeleted": true
        }
      ]
    },
    {
      "id": 2870966,
      "postDate": "2024-06-14T00:13:06.780Z",
      "content": "<p>Update: We are addressing this issue and will provide an update by next week.</p>",
      "rawMarkdown": "Update: We are addressing this issue and will provide an update by next week.",
      "votes": 12,
      "replies": [
        {
          "id": 2870973,
          "postDate": "2024-06-14T00:29:55.803Z",
          "content": "<p>This is very close to comp' end already. Depend on the update and how much impact it will have on existing solutions, it might be worth to consider an extension to the competition. <a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> </p>",
          "rawMarkdown": "This is very close to comp' end already. Depend on the update and how much impact it will have on existing solutions, it might be worth to consider an extension to the competition. @jerrylin96 ",
          "votes": 5,
          "replies": [
            {
              "id": 2870979,
              "postDate": "2024-06-14T01:00:24.887Z",
              "content": "<p>Yes we are thinking about this.</p>",
              "rawMarkdown": "Yes we are thinking about this.",
              "votes": 7
            },
            {
              "id": 2872169,
              "postDate": "2024-06-14T16:02:47.180Z",
              "content": "<p>Don't forget to extend the deadline if you replace the test dataset!:)</p>",
              "rawMarkdown": "Don't forget to extend the deadline if you replace the test dataset!:)",
              "votes": -1
            }
          ]
        },
        {
          "id": 2870991,
          "postDate": "2024-06-14T01:25:41.670Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 2870998,
          "postDate": "2024-06-14T01:50:23.077Z",
          "content": "<p><a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> Thank you for the reply.</p>\n<p>If \"update\" means replacement of test set to different sub-sampling of year 9-10, I guess we can infer timestamp and location data of new dataset from currently released test set (by comparing similarity of two feature vectors etc.). And according to  Andrew Gaiss's experiment[1], location data is almost perfectly predictable from the atmospheric features (~96% according to his report). So, could you also confirm the indirect/implicit usage of these information? I believe this must be inevitable according to the fact that current atmospheric feature already contains good amount of location-specific information.</p>\n<ul>\n<li>[1] <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/511911#2869607\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/511911#2869607</a></li>\n</ul>",
          "rawMarkdown": "@jerrylin96 Thank you for the reply.\n\nIf \"update\" means replacement of test set to different sub-sampling of year 9-10, I guess we can infer timestamp and location data of new dataset from currently released test set (by comparing similarity of two feature vectors etc.). And according to  Andrew Gaiss's experiment[1], location data is almost perfectly predictable from the atmospheric features (~96% according to his report). So, could you also confirm the indirect/implicit usage of these information? I believe this must be inevitable according to the fact that current atmospheric feature already contains good amount of location-specific information.\n\n- [1] https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/511911#2869607",
          "votes": 7
        },
        {
          "id": 2871326,
          "postDate": "2024-06-14T06:57:22.137Z",
          "content": "<p><a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> </p>\n<p>I urge caution when considering updating the test data.</p>\n<ul>\n<li><p>Changing the data in the middle of the competition would alter the competition itself. Participants who have invested time and money following the existing rules may find their efforts wasted, which is unfair. Many participants who are not violating any rules may be disadvantaged.</p></li>\n<li><p>Using the new test data could increase the risk of it being hacked, as the already provided test data is out there [1]. The update may cause more harm than good, with little to no benefit.</p></li>\n<li><p>We are using location data. We attempted to use future data with timestamps but were unsuccessful. Our solution, which does not use future data, remains valuable. Additionally, <strong>comparing solutions that use location data with those that do not can provide insights into the usefulness of this data, which could be beneficial to the field.</strong></p></li>\n</ul>\n<p>[1] <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/511911#2869193\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/511911#2869193</a></p>",
          "rawMarkdown": "@jerrylin96 \n\nI urge caution when considering updating the test data.\n\n- Changing the data in the middle of the competition would alter the competition itself. Participants who have invested time and money following the existing rules may find their efforts wasted, which is unfair. Many participants who are not violating any rules may be disadvantaged.\n\n- Using the new test data could increase the risk of it being hacked, as the already provided test data is out there [1]. The update may cause more harm than good, with little to no benefit.\n\n- We are using location data. We attempted to use future data with timestamps but were unsuccessful. Our solution, which does not use future data, remains valuable. Additionally, **comparing solutions that use location data with those that do not can provide insights into the usefulness of this data, which could be beneficial to the field.**\n\n[1] https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/511911#2869193",
          "votes": 11,
          "replies": [
            {
              "id": 2871374,
              "postDate": "2024-06-14T07:12:15.820Z",
              "content": "<p>So the gap between 0.788 and 0.798 is because of the location leak?</p>",
              "rawMarkdown": "So the gap between 0.788 and 0.798 is because of the location leak?",
              "votes": 2
            },
            {
              "id": 2871384,
              "postDate": "2024-06-14T07:18:21.810Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2871402,
              "postDate": "2024-06-14T07:25:57.037Z",
              "content": "<blockquote>\n  <p>We are using location data. We attempted to use future data with timestamps but were unsuccessful.</p>\n</blockquote>\n<p>Good to know. Your team could exploit something from latlon, but lead data.<br>\nMaybe NearestNeighbor or something like that.</p>\n<p>I think if the host provides new shuffled test data, we can estimate latlon and timestamp. This might be the problem.</p>",
              "rawMarkdown": "> We are using location data. We attempted to use future data with timestamps but were unsuccessful.\n\nGood to know. Your team could exploit something from latlon, but lead data.\nMaybe NearestNeighbor or something like that.\n\nI think if the host provides new shuffled test data, we can estimate latlon and timestamp. This might be the problem.",
              "votes": 3
            },
            {
              "id": 2871431,
              "postDate": "2024-06-14T07:41:14.913Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2871433,
              "postDate": "2024-06-14T07:42:06.587Z",
              "content": "<blockquote>\n  <p>So the gap between 0.788 and 0.798 is because of the location leak?</p>\n</blockquote>\n<p>We use it, but I don't know if other teams use it.</p>",
              "rawMarkdown": "> So the gap between 0.788 and 0.798 is because of the location leak?\n\nWe use it, but I don't know if other teams use it."
            },
            {
              "id": 2871528,
              "postDate": "2024-06-14T08:58:08.193Z",
              "content": "<p>Well there is a very big difference between 'estimate' based on train data and between 'know' from leak in test…</p>",
              "rawMarkdown": "Well there is a very big difference between 'estimate' based on train data and between 'know' from leak in test...",
              "votes": 1
            },
            {
              "id": 2871679,
              "postDate": "2024-06-14T10:38:18.263Z",
              "content": "<p>I can confirm that LB 0.788 can be achieved without using Location and Timestamp leak, I also wonder to what extent can location leak improve predictive performance</p>",
              "rawMarkdown": "I can confirm that LB 0.788 can be achieved without using Location and Timestamp leak, I also wonder to what extent can location leak improve predictive performance",
              "votes": 6
            },
            {
              "id": 2871744,
              "postDate": "2024-06-14T11:29:20.973Z",
              "content": "<p>In one of my experiments I fed the model location and timestamp and got an improvement of ~0.005. Only on train data of course, I could not apply this model to test (well it was before I knew about leak) it was just out of curiosity hahaha.<br>\nI would say it can definitely help, but it's not a game changer.<br>\nIt's not some secret to propel people from ~0.75 to ~0.8. But if 1st and 2nd are very close, it can make the difference.</p>",
              "rawMarkdown": "In one of my experiments I fed the model location and timestamp and got an improvement of ~0.005. Only on train data of course, I could not apply this model to test (well it was before I knew about leak) it was just out of curiosity hahaha.\nI would say it can definitely help, but it's not a game changer.\nIt's not some secret to propel people from ~0.75 to ~0.8. But if 1st and 2nd are very close, it can make the difference.",
              "votes": 7
            },
            {
              "id": 2871872,
              "postDate": "2024-06-14T12:41:22.900Z",
              "content": "<p>Yep, second confirmation I am not using latlon or any doing anything to exploit the time series</p>",
              "rawMarkdown": "Yep, second confirmation I am not using latlon or any doing anything to exploit the time series",
              "votes": 7
            },
            {
              "id": 2871900,
              "postDate": "2024-06-14T12:58:44.403Z",
              "content": "<p>There’s already a lot of research data showing that location is important for reproducing model output, but a non-location dependent weather parameterization would be more ideal.</p>",
              "rawMarkdown": "There’s already a lot of research data showing that location is important for reproducing model output, but a non-location dependent weather parameterization would be more ideal.",
              "votes": 2
            },
            {
              "id": 2871932,
              "postDate": "2024-06-14T13:23:04.683Z",
              "content": "<p>My guess is that those using lat/lon are using it to reorganize the data into spatial maps of atmospheric variables, and then applying some sort of CNN to those spatial maps. </p>\n<p>A way to guard against this (since these solutions are not what the organizers want) would be to sprinkle a new NaNs into some rows of the full test data set. Full NaN rows would break any CNNs based on global maps, but not those using a column approach. </p>",
              "rawMarkdown": "My guess is that those using lat/lon are using it to reorganize the data into spatial maps of atmospheric variables, and then applying some sort of CNN to those spatial maps. \n\nA way to guard against this (since these solutions are not what the organizers want) would be to sprinkle a new NaNs into some rows of the full test data set. Full NaN rows would break any CNNs based on global maps, but not those using a column approach. "
            },
            {
              "id": 2871968,
              "postDate": "2024-06-14T13:52:18.407Z",
              "content": "<p>It's not likely. Using location most probably means just feeding the model the location of each sample. I don't believe a 2D CNN has any chance in this competition. Good luck to anyone who tries that.</p>",
              "rawMarkdown": "It's not likely. Using location most probably means just feeding the model the location of each sample. I don't believe a 2D CNN has any chance in this competition. Good luck to anyone who tries that."
            },
            {
              "id": 2874089,
              "postDate": "2024-06-16T03:31:09.143Z",
              "content": "<p><a href=\"https://www.kaggle.com/ekffar\" target=\"_blank\">@ekffar</a> Also without original(HugginFace) data?</p>",
              "rawMarkdown": "@ekffar Also without original(HugginFace) data?"
            },
            {
              "id": 2874437,
              "postDate": "2024-06-16T09:05:52.453Z",
              "content": "<p><a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a></p>\n<p>It is best for any dataset update (should the host decide its necessity) to not depend on any assumption whatsoever about the learning/inference algorithms being used.</p>",
              "rawMarkdown": "@jerrylin96\n\nIt is best for any dataset update (should the host decide its necessity) to not depend on any assumption whatsoever about the learning/inference algorithms being used."
            }
          ]
        },
        {
          "id": 2871790,
          "postDate": "2024-06-14T11:55:41.807Z",
          "rawMarkdown": "",
          "votes": -1,
          "isDeleted": true
        },
        {
          "id": 2871930,
          "postDate": "2024-06-14T13:21:52.327Z",
          "content": "<p><a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> </p>\n<p>If the competition host releases another new test set for us to work on instead, this should <em>not</em> be a problem at all as long as the new test set is generated/produced/collected/whatever-verb-is-appropriate from the same way as the training set.  In machine learning jargon, we say that \"The training and test sets come from the same/similar distribution.\"</p>\n<p>In fact, if a model should work well on the old test set but not on a new test set, that probably means the model does not generalize well.  But the whole point of machine learning is to find a model that generalizes well.  If there were a God of Machine Learning, I'm sure He would care most about learning algorithms that generalize well, and could care less about any leaderboard score shakeup.</p>\n<p>If the competition host decides it's necessary to also release another new training set for us to work on, that is also <em>not</em> a problem because of the following reasons:</p>\n<p>(1) If the new training set is provided with any undesirable leak information somehow removed, then those models that have not been taking advantage of the leak should have minimal impact, if at all.</p>\n<p>(2) Even if there exists any nonzero impact to any models from the new training set, orthogonal to the removal of the undesirable leak, it would apply approximately equally to every model anyway.  So, everyone would be impacted roughly the same way, anyway.</p>\n<p>(3) Worst comes worst, we just need to retrain our model on the new training set.  Big deal (sarcasm intended).  The only actual problem I see with this is the time it would take to train again.  But that is easily fixed if we get some extension to the deadline.</p>\n<p>In any case, this is a great competition.  A super large multiple-target regression problem to work on is a heaven Carl Friedrich Gauss could only dream of.  Competition host, thank you for providing this wonderful competition, and I will take it to the end whatever the host decides to do.</p>",
          "rawMarkdown": "@jerrylin96 \n\nIf the competition host releases another new test set for us to work on instead, this should *not* be a problem at all as long as the new test set is generated/produced/collected/whatever-verb-is-appropriate from the same way as the training set.  In machine learning jargon, we say that \"The training and test sets come from the same/similar distribution.\"\n\nIn fact, if a model should work well on the old test set but not on a new test set, that probably means the model does not generalize well.  But the whole point of machine learning is to find a model that generalizes well.  If there were a God of Machine Learning, I'm sure He would care most about learning algorithms that generalize well, and could care less about any leaderboard score shakeup.\n\nIf the competition host decides it's necessary to also release another new training set for us to work on, that is also *not* a problem because of the following reasons:\n\n(1) If the new training set is provided with any undesirable leak information somehow removed, then those models that have not been taking advantage of the leak should have minimal impact, if at all.\n\n(2) Even if there exists any nonzero impact to any models from the new training set, orthogonal to the removal of the undesirable leak, it would apply approximately equally to every model anyway.  So, everyone would be impacted roughly the same way, anyway.\n\n(3) Worst comes worst, we just need to retrain our model on the new training set.  Big deal (sarcasm intended).  The only actual problem I see with this is the time it would take to train again.  But that is easily fixed if we get some extension to the deadline.\n\nIn any case, this is a great competition.  A super large multiple-target regression problem to work on is a heaven Carl Friedrich Gauss could only dream of.  Competition host, thank you for providing this wonderful competition, and I will take it to the end whatever the host decides to do.",
          "replies": [
            {
              "id": 2872003,
              "postDate": "2024-06-14T14:11:07.530Z",
              "content": "<blockquote>\n  <p>The only actual problem I see with this is the time it would take to train again. But that is easily fixed if we get some extension to the deadline.</p>\n</blockquote>\n<p>you have no idea how much compute this comp requires :D </p>",
              "rawMarkdown": "> The only actual problem I see with this is the time it would take to train again. But that is easily fixed if we get some extension to the deadline.\n\nyou have no idea how much compute this comp requires :D \n",
              "votes": 5
            },
            {
              "id": 2872024,
              "postDate": "2024-06-14T14:21:03.990Z",
              "content": "<p>Hahahahaha true that. I won't be surprised if it's the most demanding in the history of Kaggle, outside of genAi comp' where there is no limit to compute requirements if you want to generate more samples.</p>",
              "rawMarkdown": "Hahahahaha true that. I won't be surprised if it's the most demanding in the history of Kaggle, outside of genAi comp' where there is no limit to compute requirements if you want to generate more samples."
            },
            {
              "id": 2872030,
              "postDate": "2024-06-14T14:26:37.627Z",
              "content": "<blockquote>\n  <p>you have no idea how much compute this comp requires :D </p>\n</blockquote>\n<p>I'm very aware of it :D</p>\n<p>Efficiency, and not just test set accuracy, is part of the competition.  That's just how things work in real life, too.</p>",
              "rawMarkdown": ">you have no idea how much compute this comp requires :D \n\nI'm very aware of it :D\n\nEfficiency, and not just test set accuracy, is part of the competition.  That's just how things work in real life, too."
            },
            {
              "id": 2872039,
              "postDate": "2024-06-14T14:33:53.657Z",
              "content": "<blockquote>\n  <blockquote>\n    <p>you have no idea how much compute this comp requires :D </p>\n  </blockquote>\n  <p>I'm very aware of it :D</p>\n  <p>Efficiency, and not just test set accuracy, is part of the competition.  That's just how things work in real life, too.</p>\n</blockquote>\n<p>Nah efficiency is maybe good for silver range.<br>\nFor gold range, I'm very doubtful anyone manged to get there 'efficiently' without throwing a LOT of compute on the data.</p>",
              "rawMarkdown": "> >you have no idea how much compute this comp requires :D \n> \n> I'm very aware of it :D\n> \n> Efficiency, and not just test set accuracy, is part of the competition.  That's just how things work in real life, too.\n\nNah efficiency is maybe good for silver range.\nFor gold range, I'm very doubtful anyone manged to get there 'efficiently' without throwing a LOT of compute on the data."
            },
            {
              "id": 2872133,
              "postDate": "2024-06-14T15:45:13.720Z",
              "content": "<p><a href=\"https://www.kaggle.com/shlomoron\" target=\"_blank\">@shlomoron</a> , sure, I hear you.  But an extension of sufficient length should eliminate such a concern.  My point is, we should all strive for sound machine learning practices, even if it might mean my leaderboard score might drop.  To that regard, I think we're on the same side 😀</p>",
              "rawMarkdown": "@shlomoron , sure, I hear you.  But an extension of sufficient length should eliminate such a concern.  My point is, we should all strive for sound machine learning practices, even if it might mean my leaderboard score might drop.  To that regard, I think we're on the same side 😀"
            },
            {
              "id": 2872183,
              "postDate": "2024-06-14T16:15:05.347Z",
              "content": "<p>I don't think I've thrown that much compute at the problem but I guess everyone has different thresholds for that. I can retrain my model on two 4090's in a couple days.</p>\n<p>Maybe people are doing way more than me for this but I don't feel like it's the most compute intensive that I've done. Deep fake, Lyft, some of the other large scale image ones</p>",
              "rawMarkdown": "I don't think I've thrown that much compute at the problem but I guess everyone has different thresholds for that. I can retrain my model on two 4090's in a couple days.\n\nMaybe people are doing way more than me for this but I don't feel like it's the most compute intensive that I've done. Deep fake, Lyft, some of the other large scale image ones",
              "votes": 3
            },
            {
              "id": 2872708,
              "postDate": "2024-06-15T02:12:34.893Z",
              "content": "<p>I’ve only trained my model for one hour maximum, I guess that’s what’s been holding me back 😂</p>",
              "rawMarkdown": "I’ve only trained my model for one hour maximum, I guess that’s what’s been holding me back 😂",
              "votes": 3
            },
            {
              "id": 2874513,
              "postDate": "2024-06-16T10:36:19.297Z",
              "content": "<p>Unless you mean 1 hour on a H100 cluster</p>",
              "rawMarkdown": "Unless you mean 1 hour on a H100 cluster",
              "votes": 2
            }
          ]
        },
        {
          "id": 2872031,
          "postDate": "2024-06-14T14:27:12.580Z",
          "content": "<p><a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a><br>\nFirst and foremost, we should consider the following two points separately, prioritizing the former:</p>\n<ul>\n<li>Not changing the competition rules</li>\n<li>Adhering to the host's objectives</li>\n</ul>\n<p>If there is a \"critical leak,\" such as the inclusion of training data in the test data, as happened in the past with the Foursquare - Location Matching competition, then it may be unavoidable to change the rules in such cases.</p>\n<p>However, in this particular instance, I don't believe it is critical enough to warrant changing the rules near the end of the competition. Although the current data is ordered by timestamps, the time gaps are too large to be effectively used. I don't believe this is a critical leak significant enough to change the data and compromise the \"competitiveness.\"</p>\n<p>Additionally, can we really say that being able to determine the latitude and longitude is a \"critical leak\"?  The E3SM simulator appears to use information from all grids. I believe that solutions using simulations that incorporate information from all grids have value. (at least, I am confident that our team can give you a solution that is much more valuable than the \"stop lightgbm early stopping\" and \"make the neural network as long as possible\" asked as foursquare's solution. )</p>",
          "rawMarkdown": "@jerrylin96\nFirst and foremost, we should consider the following two points separately, prioritizing the former:\n\n- Not changing the competition rules\n- Adhering to the host's objectives\n\nIf there is a \"critical leak,\" such as the inclusion of training data in the test data, as happened in the past with the Foursquare - Location Matching competition, then it may be unavoidable to change the rules in such cases.\n\nHowever, in this particular instance, I don't believe it is critical enough to warrant changing the rules near the end of the competition. Although the current data is ordered by timestamps, the time gaps are too large to be effectively used. I don't believe this is a critical leak significant enough to change the data and compromise the \"competitiveness.\"\n\nAdditionally, can we really say that being able to determine the latitude and longitude is a \"critical leak\"?  The E3SM simulator appears to use information from all grids. I believe that solutions using simulations that incorporate information from all grids have value. (at least, I am confident that our team can give you a solution that is much more valuable than the \"stop lightgbm early stopping\" and \"make the neural network as long as possible\" asked as foursquare's solution. )",
          "votes": 1,
          "replies": [
            {
              "id": 2872065,
              "postDate": "2024-06-14T14:58:01.680Z",
              "content": "<p>The problem with spacetime, imho is that it distracts from building the best model that depends only on the inputs/physical data.<br>\nSure, spacetime data can help, but we already knew that. Now, we won't know how far we can push without it. People will go crazy in the next two weeks trying to build the smartest spacetime models instead of the best model that depends only on the same sample data. It's a shame, really.<br>\nWell, it's also a shame for teams that used leak like yours. No good solutions here, really, only less harmful ones.</p>",
              "rawMarkdown": "The problem with spacetime, imho is that it distracts from building the best model that depends only on the inputs/physical data.\nSure, spacetime data can help, but we already knew that. Now, we won't know how far we can push without it. People will go crazy in the next two weeks trying to build the smartest spacetime models instead of the best model that depends only on the same sample data. It's a shame, really.\nWell, it's also a shame for teams that used leak like yours. No good solutions here, really, only less harmful ones.",
              "votes": 3
            },
            {
              "id": 2872148,
              "postDate": "2024-06-14T15:49:40.303Z",
              "content": "<blockquote>\n  <p>I believe that solutions using simulations that incorporate information from all grids have value. </p>\n</blockquote>\n<p>We are not host. Who knows</p>",
              "rawMarkdown": "> I believe that solutions using simulations that incorporate information from all grids have value. \n\nWe are not host. Who knows"
            },
            {
              "id": 2872181,
              "postDate": "2024-06-14T16:14:42.307Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2872184,
              "postDate": "2024-06-14T16:15:18.077Z",
              "content": "<p>Now I'm a bit surprised to see them so persistent for this leak and started to think that it must be a key to break 0.79 and more…🤔</p>",
              "rawMarkdown": "Now I'm a bit surprised to see them so persistent for this leak and started to think that it must be a key to break 0.79 and more...🤔",
              "votes": 3
            },
            {
              "id": 2872225,
              "postDate": "2024-06-14T17:01:40.450Z",
              "content": "<blockquote>\n  <p>Now I'm a bit surprised to see them so persistent for this leak and started to think that it must be a key to break 0.79 and more…🤔</p>\n</blockquote>\n<p>Had the same thought lol. Probably there is some smart way to use adjacent /mean grid data or something similar</p>",
              "rawMarkdown": "> Now I'm a bit surprised to see them so persistent for this leak and started to think that it must be a key to break 0.79 and more...🤔\n\nHad the same thought lol. Probably there is some smart way to use adjacent /mean grid data or something similar\n",
              "votes": 1
            },
            {
              "id": 2872602,
              "postDate": "2024-06-14T23:15:16.433Z",
              "content": "<blockquote>\n  <p>We are not host. Who knows</p>\n</blockquote>\n<p>As you said, only the host can know whether it will be useful or not. So, it’s just that \"I believe\"😀.</p>",
              "rawMarkdown": "> We are not host. Who knows\n\nAs you said, only the host can know whether it will be useful or not. So, it’s just that \"I believe\"😀.",
              "votes": 5
            }
          ]
        }
      ]
    },
    {
      "id": 2869411,
      "postDate": "2024-06-13T05:03:06.087Z",
      "content": "<p>Nice shot <a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> !</p>\n<p>I made a visualization of air temperature at level 60 in Celsius for train and test (using your trick) at time t=0  </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F214989%2F13f0df1992afedd1e0d56f78cc854fd1%2Fleap.png?generation=1718254862105057&amp;alt=media\"></p>\n<p>Here’s my two cents :</p>\n<ul>\n<li>I don't really like competitions with synthetic data because, in the end, your goal becomes to reverse engineer the data generator. But this is a reverse engineering competition. So, it's hard to say: your goal is to reverse engineer this emulator, but you can't use this or that.</li>\n<li>Never underestimate kagglers 😀.</li>\n</ul>",
      "rawMarkdown": "Nice shot @tatamikenn !\n\nI made a visualization of air temperature at level 60 in Celsius for train and test (using your trick) at time t=0  \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F214989%2F13f0df1992afedd1e0d56f78cc854fd1%2Fleap.png?generation=1718254862105057&alt=media)\n\nHere’s my two cents :\n\n- I don't really like competitions with synthetic data because, in the end, your goal becomes to reverse engineer the data generator. But this is a reverse engineering competition. So, it's hard to say: your goal is to reverse engineer this emulator, but you can't use this or that.\n- Never underestimate kagglers 😀.\n",
      "votes": 9,
      "replies": [
        {
          "id": 2869423,
          "postDate": "2024-06-13T05:14:17.517Z",
          "content": "<p>Year, definitely this competition is kind of reverse-engineering competition for climate simulator LOL.</p>",
          "rawMarkdown": "Year, definitely this competition is kind of reverse-engineering competition for climate simulator LOL.",
          "votes": 3
        },
        {
          "id": 2869436,
          "postDate": "2024-06-13T05:26:59.047Z",
          "content": "<p>Never underestimate kagglers should by this site motto.<br>\nIn Latin, as per tradition.<br>\nNUMQUAM MINORIS AESTIMO KAGGLERS</p>",
          "rawMarkdown": "Never underestimate kagglers should by this site motto.\nIn Latin, as per tradition.\nNUMQUAM MINORIS AESTIMO KAGGLERS",
          "votes": 3
        }
      ]
    },
    {
      "id": 2868822,
      "postDate": "2024-06-12T16:43:57.387Z",
      "content": "<p>I guess the easiest fix would be to release a new version of the test set with all rows properly shuffled and all the sample IDs scrambled. This would break all current submissions, but we could simply restore the previous leaderboard scores by rerunning our models on the new test set, since the data rows are still the same, just reordered. Otherwise, everyone would be forced to retrain their models using this leak to remain competitive, which is probably not the purpose of this competition?</p>",
      "rawMarkdown": "I guess the easiest fix would be to release a new version of the test set with all rows properly shuffled and all the sample IDs scrambled. This would break all current submissions, but we could simply restore the previous leaderboard scores by rerunning our models on the new test set, since the data rows are still the same, just reordered. Otherwise, everyone would be forced to retrain their models using this leak to remain competitive, which is probably not the purpose of this competition?",
      "votes": 6,
      "replies": [
        {
          "id": 2869193,
          "postDate": "2024-06-12T22:54:09.317Z",
          "content": "<blockquote>\n  <p>I guess the easiest fix would be to release a new version of the test set with all rows properly shuffled and all the sample IDs scrambled.</p>\n</blockquote>\n<p>In my feeling, this would not solve the issue, because knowing subsampling of the dataset, Kaggle’s would successfully reverse engineering the new subsampling using information of already released subsampling (checking similarity of feature vector etc.). The only way of releasing non-leaked dataset is to newly simulate another two year data (namely year 11-12), but I doubt host could make this considering not cheap amount of computing resource doing this.</p>",
          "rawMarkdown": "> I guess the easiest fix would be to release a new version of the test set with all rows properly shuffled and all the sample IDs scrambled.\n\nIn my feeling, this would not solve the issue, because knowing subsampling of the dataset, Kaggle’s would successfully reverse engineering the new subsampling using information of already released subsampling (checking similarity of feature vector etc.). The only way of releasing non-leaked dataset is to newly simulate another two year data (namely year 11-12), but I doubt host could make this considering not cheap amount of computing resource doing this.",
          "votes": 1,
          "replies": [
            {
              "id": 2869195,
              "postDate": "2024-06-12T22:59:33.203Z",
              "content": "<blockquote>\n  <p>Otherwise, everyone would be forced to retrain their models using this leak to remain competitive, which is probably not the purpose of this competition?</p>\n</blockquote>\n<p>I am also not sure this is possible, considering the case the model implicitly exploiting this leak.</p>",
              "rawMarkdown": "> Otherwise, everyone would be forced to retrain their models using this leak to remain competitive, which is probably not the purpose of this competition?\n\nI am also not sure this is possible, considering the case the model implicitly exploiting this leak."
            },
            {
              "id": 2869304,
              "postDate": "2024-06-13T02:37:40.763Z",
              "content": "<p>Considering the point I wrote above, I must say the most probable action they can take is <em>leave everything as it is, and just grant usage of timestamp and location data</em>. However, final decision is leave for host and Kaggle team.</p>",
              "rawMarkdown": "Considering the point I wrote above, I must say the most probable action they can take is *leave everything as it is, and just grant usage of timestamp and location data*. However, final decision is leave for host and Kaggle team."
            },
            {
              "id": 2869601,
              "postDate": "2024-06-13T06:47:44.797Z",
              "content": "<blockquote>\n  <p>because knowing subsampling of the dataset, Kaggle’s would successfully reverse engineering the new subsampling using information of already released subsampling</p>\n</blockquote>\n<p>You're right! I didn't think of that.</p>",
              "rawMarkdown": ">because knowing subsampling of the dataset, Kaggle’s would successfully reverse engineering the new subsampling using information of already released subsampling\n\nYou're right! I didn't think of that.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2869008,
      "postDate": "2024-06-12T19:24:02.537Z",
      "content": "<p>Wow, respect for sharing this when you could possibly have found ways to exploit this.</p>",
      "rawMarkdown": "Wow, respect for sharing this when you could possibly have found ways to exploit this.",
      "votes": 2,
      "replies": [
        {
          "id": 2869176,
          "postDate": "2024-06-12T22:34:17.837Z",
          "content": "<p>I was surprised to hear not few Kaggles already found this in the early stage of competition.<br>\nI don’t know Kaggle’s convention, but I have to say I am as novice as reporting this as an anomaly when someone is walking in the street upside down.</p>",
          "rawMarkdown": "I was surprised to hear not few Kaggles already found this in the early stage of competition.\nI don’t know Kaggle’s convention, but I have to say I am as novice as reporting this as an anomaly when someone is walking in the street upside down.",
          "votes": 2,
          "replies": [
            {
              "id": 2869334,
              "postDate": "2024-06-13T03:15:06.707Z",
              "content": "<p>Apparently many have found this, but they did not share it as you did. And those who found it and disclose it now say it is not useful.</p>",
              "rawMarkdown": "Apparently many have found this, but they did not share it as you did. And those who found it and disclose it now say it is not useful.",
              "votes": 6
            }
          ]
        }
      ]
    },
    {
      "id": 2868971,
      "postDate": "2024-06-12T18:55:35.943Z",
      "content": "<p>I made some plots of this early on but haven't done anything to exploit it. Assumed it was out of the spirit of the competition or simply wouldn't be of any value on the private set</p>",
      "rawMarkdown": "I made some plots of this early on but haven't done anything to exploit it. Assumed it was out of the spirit of the competition or simply wouldn't be of any value on the private set",
      "votes": 2,
      "replies": [
        {
          "id": 2869007,
          "postDate": "2024-06-12T19:23:47.550Z",
          "content": "<p>This is Kaggle, I know nothing about this 'spirit of the competition' lol. Whatever we can exploit, as long as it is within the rules- we exploit haha</p>",
          "rawMarkdown": "This is Kaggle, I know nothing about this 'spirit of the competition' lol. Whatever we can exploit, as long as it is within the rules- we exploit haha"
        }
      ]
    },
    {
      "id": 2878456,
      "postDate": "2024-06-19T00:36:33.447Z",
      "content": "<p>The competition should be good to go now. Thank you for your patience.</p>",
      "rawMarkdown": "The competition should be good to go now. Thank you for your patience.",
      "votes": 1
    },
    {
      "id": 2868774,
      "postDate": "2024-06-12T16:08:04.037Z",
      "content": "<p>What kind of action are you requesting?</p>",
      "rawMarkdown": "What kind of action are you requesting?",
      "votes": 1,
      "replies": [
        {
          "id": 2868824,
          "postDate": "2024-06-12T16:47:28.373Z",
          "content": "<p>Well, in the early days of this competition, one of host member commented on usage of location &amp; timestamp data, and from this comment, I thought the hosts are expecting models without these information.</p>\n<p>So at least, I want to make sure using these information is valid on this competition.</p>\n<p><a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/511911\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/511911</a></p>\n<p>Since some specific action is not on my mind, I rewrote this discussion's title to \"request to confirmation\".</p>",
          "rawMarkdown": "Well, in the early days of this competition, one of host member commented on usage of location & timestamp data, and from this comment, I thought the hosts are expecting models without these information.\n\nSo at least, I want to make sure using these information is valid on this competition.\n\nhttps://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/511911\n\nSince some specific action is not on my mind, I rewrote this discussion's title to \"request to confirmation\".",
          "replies": [
            {
              "id": 2868830,
              "postDate": "2024-06-12T16:55:19.927Z",
              "content": "<p>I'm sure a lot of kaggle grandmasters were aware of this property. But not sure if top teams could utilize them.</p>",
              "rawMarkdown": "I'm sure a lot of kaggle grandmasters were aware of this property. But not sure if top teams could utilize them."
            }
          ]
        }
      ]
    },
    {
      "id": 2868820,
      "postDate": "2024-06-12T16:42:34.753Z",
      "content": "<p>good catch!</p>",
      "rawMarkdown": "good catch!",
      "votes": 1
    },
    {
      "id": 2877488,
      "postDate": "2024-06-18T12:01:41.133Z",
      "content": "<p>Just for clarity, train data and train saved weights will be the same \"if we haven't utilize or reverse the timestamp and location\"<br>\nso that we don't repeat all from the beginning ? </p>",
      "rawMarkdown": "Just for clarity, train data and train saved weights will be the same \"if we haven't utilize or reverse the timestamp and location\"\nso that we don't repeat all from the beginning ? ",
      "replies": [
        {
          "id": 2878104,
          "postDate": "2024-06-18T18:25:38.183Z",
          "content": "<p>Assuming you are not using a multi-atmospheric-column approach (which requires exploiting the structure of the data in a way that we did not intend), you shouldn't have to do any retraining.</p>",
          "rawMarkdown": "Assuming you are not using a multi-atmospheric-column approach (which requires exploiting the structure of the data in a way that we did not intend), you shouldn't have to do any retraining.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2870247,
      "postDate": "2024-06-13T13:33:39.540Z",
      "content": "<p>Can the authors clarify whether it is legal to use latitude and longitude? <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/495258#2768949\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/495258#2768949</a> from this looks like not.</p>",
      "rawMarkdown": "Can the authors clarify whether it is legal to use latitude and longitude? https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/495258#2768949 from this looks like not.",
      "replies": [
        {
          "id": 2870961,
          "postDate": "2024-06-13T23:50:33.647Z",
          "content": "<p>For sure it can be used to root cause solutions if r2 values have a localized location.  If you identify that the north pole causes some labels to be bad, feature engineering (not including the lat/long) could be added based on research into the world of weather prediction.</p>",
          "rawMarkdown": "For sure it can be used to root cause solutions if r2 values have a localized location.  If you identify that the north pole causes some labels to be bad, feature engineering (not including the lat/long) could be added based on research into the world of weather prediction."
        }
      ]
    },
    {
      "id": 2868983,
      "postDate": "2024-06-12T19:09:53.307Z",
      "content": "<p>Are models trained with location/time id any better? Did someone try?</p>",
      "rawMarkdown": "Are models trained with location/time id any better? Did someone try?",
      "replies": [
        {
          "id": 2869172,
          "postDate": "2024-06-12T22:15:17.900Z",
          "content": "<p>My guess would be yes, most climate emulators take spatial information and time into account. Would be interesting to see if more people have tried this</p>",
          "rawMarkdown": "My guess would be yes, most climate emulators take spatial information and time into account. Would be interesting to see if more people have tried this"
        },
        {
          "id": 2870333,
          "postDate": "2024-06-13T14:17:36.193Z",
          "content": "<p>Purely intuitive there will be an improvement</p>",
          "rawMarkdown": "Purely intuitive there will be an improvement"
        }
      ]
    },
    {
      "id": 2875318,
      "postDate": "2024-06-17T03:30:07.680Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 2875321,
          "postDate": "2024-06-17T03:37:15.527Z",
          "content": "<p>Thanks, that is what I had read. </p>",
          "rawMarkdown": "Thanks, that is what I had read. "
        },
        {
          "id": 2875325,
          "postDate": "2024-06-17T03:39:50.113Z",
          "content": "<p>“This leaderboard is calculated with approximately 20% of the test data. The final results will be based on the other 80%, so the final standings may be different.”</p>",
          "rawMarkdown": "“This leaderboard is calculated with approximately 20% of the test data. The final results will be based on the other 80%, so the final standings may be different.”"
        },
        {
          "id": 2875364,
          "postDate": "2024-06-17T04:55:22.813Z",
          "content": "<p><a href=\"https://www.kaggle.com/juandcf\" target=\"_blank\">@juandcf</a> </p>\n<blockquote>\n  <p>This leaderboard is calculated with approximately 20% of the test data. The final results will be based on the other 80%, so the final standings may be different.</p>\n</blockquote>\n<p>This means public LB <em>score</em> is calculated by 20% of test set. Private LB is calculated by rest of <em>80%</em> of data. So in usual competition setup, data swap never happens.</p>",
          "rawMarkdown": "@juandcf \n\n> This leaderboard is calculated with approximately 20% of the test data. The final results will be based on the other 80%, so the final standings may be different.\n\nThis means public LB *score* is calculated by 20% of test set. Private LB is calculated by rest of *80%* of data. So in usual competition setup, data swap never happens."
        },
        {
          "id": 2875417,
          "postDate": "2024-06-17T06:11:51.570Z",
          "content": "<p>Got it, thanks!</p>",
          "rawMarkdown": "Got it, thanks!"
        }
      ]
    },
    {
      "id": 2873023,
      "postDate": "2024-06-15T09:09:43.313Z",
      "rawMarkdown": "",
      "votes": -3,
      "isDeleted": true
    },
    {
      "id": 2870934,
      "postDate": "2024-06-13T22:30:31.367Z",
      "rawMarkdown": "",
      "votes": -4,
      "isDeleted": true
    },
    {
      "id": 2870336,
      "postDate": "2024-06-13T14:18:52.613Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2869376,
      "postDate": "2024-06-13T03:56:57.313Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 2868812,
      "postDate": "2024-06-12T16:29:58.920Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2877664,
      "postDate": "2024-06-18T14:02:37.310Z",
      "content": "<p>Thank you for the update</p>",
      "rawMarkdown": "Thank you for the update"
    },
    {
      "id": 2871076,
      "postDate": "2024-06-14T04:08:08.447Z",
      "content": "<p>Thank you </p>",
      "rawMarkdown": "Thank you "
    }
  ],
  "comments": [
    {
      "id": 2869607,
      "author_name": "Andrew Geiss",
      "author_url": "",
      "post_date": "2024-06-13T06:49:38.677000",
      "content": "<p>We are all likely already using geographical information whether the data are shuffled or not.</p>\n<p>The atmosphere has a consistent structure that can be related to geographical location. For example, tropopause height alone is a fairly good predictor of distance from the equator (|latitude|). Some of the per-column predictors also obviously convey information about location (e.g. SZA, land fraction, and sea ice fraction).</p>\n<p>This evening, I tried training a 4-layer MLP classifier to predict the E3SM grid-point that a sample originated from. It reaches ~96% accuracy after one pass through the “low-res” dataset and could surely be improved with minimal effort and additional training.</p>\n<p>My point is that lat/lon information is already encoded in the individual rows of the test set and shuffling the data won't remove it. A relatively simple MLP can extract locations so, to the extent that this information is relevant to the problem (and I think that it is), those of us using more sophisticated models are likely leveraging this information already without providing it explicitly as an input.</p>",
      "votes": 35,
      "replies": [
        {
          "id": 2869631,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2024-06-13T06:56:35.890000",
          "content": "<p>I see. Thank you for providing us the domain knowledge.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2869637,
              "author_name": "Bilzard",
              "author_url": "",
              "post_date": "2024-06-13T06:58:59.690000",
              "content": "<blockquote>\n  <p>This evening, I tried training a 4-layer MLP classifier to predict the E3SM grid-point that a sample originated from. It reaches ~96% accuracy after one pass through the “low-res” dataset and could surely be improved with minimal effort and additional training.</p>\n</blockquote>\n<p>It corresponds with the fact many participant already utilizing lat/lon feature report that they fail to improves their model.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2870262,
          "author_name": "ONODERA",
          "author_url": "",
          "post_date": "2024-06-13T13:41:25.500000",
          "content": "<p>Just curious, are you using latlon or timestamp? Can we get 0.80 on LB without those?</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 2874007,
          "author_name": "Juan D C F",
          "author_url": "",
          "post_date": "2024-06-16T00:12:10.927000",
          "content": "<p>I agree with this… if geographical location is implicit in the data, a good, complex ML would capture that implicitly. Also, we know the test set gets swapped at the end of the comp, so overfitting to the current test set might actually be detrimental in the long run.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2874557,
              "author_name": "steubk",
              "author_url": "",
              "post_date": "2024-06-16T11:42:09.383000",
              "content": "<p>Who said that test set get swapped at the end of comp ? 🤔</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2875179,
              "author_name": "Juan D C F",
              "author_url": "",
              "post_date": "2024-06-16T23:45:49.413000",
              "content": "<p>Hmmm. I could have sworn I read something about it, but maybe I was confused with some other competition. This is my first Kaggle comp, but I was under the impression that for the final score they either swap or extend the test set with new data, so you should avoid overfitting to the current test set. But I couldn't find anything about that for this competition, so maybe this isn't the case.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2876590,
      "author_name": "Jerry Lin",
      "author_url": "",
      "post_date": "2024-06-17T21:31:18.187000",
      "content": "<p><strong>UPDATE</strong>: Some big changes to the competition are happening tomorrow.</p>\n<p>1.) There's going to be a <strong>brand new test-set, sample submission (weighting file), and benchmark submission</strong>.<br>\n2.) The competition <strong>deadline will be extended two weeks</strong> for everyone.<br>\n3.) All current submissions will be invalidated. <strong>You must resubmit using the new data</strong> to appear on the leaderboard.<br>\n4.) If you have been using the single-atmospheric-column to single-atmospheric-column regression approach, your score shouldn't change much.<br>\n5.) <em>A more detailed post will appear tomorrow</em>.</p>\n<p>For those who have been using the multi-column to single-column approach, we apologize. However, making full use of this approach would have required knowing the sample IDs were scrambled, implying you knew this was \"against the spirit of the competition\" to quote <a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a> . </p>\n<p>TLDR: we cannot have multi-atmospheric-column to single-atmospheric-column and single-atmospheric-column to single-atmospheric-column approaches coexisting on the leaderboard, especially if the former was done knowing it went against the intentions of the competition.</p>\n<p>EDIT: By column I mean \"atmospheric column\" not csv column. Sorry if I scared you!</p>",
      "votes": 20,
      "replies": [
        {
          "id": 2876609,
          "author_name": "slime",
          "author_url": "",
          "post_date": "2024-06-17T22:00:26.023000",
          "content": "<p>I think you meant single-row and multi-row? I got scared for a second :D</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2876633,
              "author_name": "Jerry Lin",
              "author_url": "",
              "post_date": "2024-06-17T22:23:42.683000",
              "content": "<p>Sorry by column, I mean \"atmospheric column\"</p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 2876636,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-06-17T22:25:59.917000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2876661,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2024-06-17T23:29:30.027000",
          "content": "<p>Thank you for the update. I hope you double-check this time that the data is scrambled for good and there is no leak…😅 Also I would like to know if the new test is from the same years as the old test (only different downsampled split) of from a totally new years. (And if so- what years?)</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 2876673,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2024-06-17T23:36:57.653000",
          "content": "<blockquote>\n  <p>For those who have been using the multi-column to single-column approach, we apologize. However, making full use of this approach would have required knowing the sample IDs were scrambled, implying you knew this was \"against the spirit of the competition\" to quote <a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a></p>\n</blockquote>\n<p>Well I guess it's a lesson, if you see a leak report instead of exploit (though the Kaggle way seems to be more on the exploit side lol) 😅</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2876675,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2024-06-17T23:37:31.733000",
          "content": "<p><a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> Thank you for the announcement.</p>\n<p>Could you also confirm the validity of solution which <em>estimates</em> location and timestamp using below data?</p>\n<ol>\n<li>currently released test.csv</li>\n<li>newly releasing test.csv</li>\n</ol>\n<p>c.f. (my previous comment)<br>\n<a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/511911#2870998\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/511911#2870998</a></p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 2876709,
          "author_name": "ADAM.",
          "author_url": "",
          "post_date": "2024-06-18T00:31:51.363000",
          "content": "<p>r2_score is not sensitive to the scale, so why change weighting file?</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2876717,
              "author_name": "Jerry Lin",
              "author_url": "",
              "post_date": "2024-06-18T01:11:10.783000",
              "content": "<p>For convenience purposes (on my end)</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2876742,
              "author_name": "ADAM.",
              "author_url": "",
              "post_date": "2024-06-18T01:37:09.253000",
              "content": "<p>Ok, got it.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2877463,
              "author_name": "Youri Matiounine",
              "author_url": "",
              "post_date": "2024-06-18T11:36:52.513000",
              "content": "<p>hey, since you are changing the weights anyway, how about you exclude all columns where weights are very high - like &gt;10^11? Then we would not have to fill them with -x/1200 or 0. Just make those weights 0. I think that would be in the spirit of the objective of this competition - predicting results only where they make sense.<br>\nOtherwise somebody may find a way to reverse engineer those weights and hack the results again.</p>",
              "votes": 3,
              "replies": []
            }
          ]
        },
        {
          "id": 2876757,
          "author_name": "HB",
          "author_url": "",
          "post_date": "2024-06-18T01:58:55.953000",
          "content": "<p>I think two weeks is too long, I'm curious what others think. The training data won't change, and this competition has a high correlation between CV and LB, so I don't think it will change what we need to do. Wouldn't a one-week extension be enough?</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2876830,
              "author_name": "kurupical",
              "author_url": "",
              "post_date": "2024-06-18T03:41:30.790000",
              "content": "<p>I think they made it 2 weeks for the \"someone\" who is using the location/timestamp information, but two or three days extension is enough for me. I was joined from the beginning of the competition and I am exhausted😭.  </p>",
              "votes": 7,
              "replies": []
            },
            {
              "id": 2876844,
              "author_name": "sroger",
              "author_url": "",
              "post_date": "2024-06-18T03:56:51.957000",
              "content": "<p>I agree with the one week timeline.<br>\nTwo weeks would veer the competition unnecessarily towards \"who has the most compute\".</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2876854,
              "author_name": "Bilzard",
              "author_url": "",
              "post_date": "2024-06-18T04:08:21.520000",
              "content": "<p>Personally, I have no counter opinion to the length of the extension, however, I think deciding by hearing with few people in this discussion is not fair for most of participants not here. So I think it’s the best we obey the decision what host made (i.e. 2 weeks extension)</p>",
              "votes": 7,
              "replies": []
            }
          ]
        },
        {
          "id": 2876762,
          "author_name": "ONODERA",
          "author_url": "",
          "post_date": "2024-06-18T02:06:26.697000",
          "content": "<p>Would you disqualify the winner if he/she used the location or timestamp feature?</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 2876816,
          "author_name": "Camaro",
          "author_url": "",
          "post_date": "2024-06-18T03:23:44.543000",
          "content": "<p><a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> </p>\n<p>Thank you for the update! I'm impressed by your quick decision-making despite the scale of these changes.</p>\n<p>I have a proposal regarding the update method:</p>\n<ul>\n<li>Release the current test set with the correct labels.</li>\n<li>Create a new test set from the next year's data, with the sample_ids completely shuffled.</li>\n<li>Allow the use of location and timestamps for the training set, including the data published on Huggingface.</li>\n</ul>\n<p>I believe this is the best way to frame the problem you want to solve without significantly changing its nature.  <br>\nIt's especially important to shift the timestamps between the current and new test sets to prevent unproductive solutions like matching between the two sets.</p>",
          "votes": -4,
          "replies": [
            {
              "id": 2877629,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-06-18T13:34:45.803000",
              "content": "<blockquote>\n  <p>Allow the use of location and timestamps for the training set, including the data published on Huggingface.</p>\n</blockquote>\n<p>It is already allowed by virtue of not being forbidden…</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2877633,
              "author_name": "Camaro",
              "author_url": "",
              "post_date": "2024-06-18T13:37:03.713000",
              "content": "<p>Wait, why does this comment have so many downvotes?😢<br>\nWhat do you have against this proposal? I seriously have no idea…</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2877646,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-06-18T13:43:37.397000",
              "content": "<p>lol I wondered the same. I am not one of them just to be clear…</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2877652,
              "author_name": "Camaro",
              "author_url": "",
              "post_date": "2024-06-18T13:45:18.963000",
              "content": "<p><a href=\"https://www.kaggle.com/shlomoron\" target=\"_blank\">@shlomoron</a> <br>\nWhat I wanted to propose was to refrain from banning location and timestamps. Completely banning them is impractical as they cannot be detected, and we would have to rely on the integrity of the participants.  <br>\nInstead, I suggested continuing to allow their use for the train set, but reconstructing the test set so that the information cannot be identified from the IDs at the very least.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2877653,
              "author_name": "Camaro",
              "author_url": "",
              "post_date": "2024-06-18T13:47:43.927000",
              "content": "<p><a href=\"https://www.kaggle.com/shlomoron\" target=\"_blank\">@shlomoron</a> <br>\nHaha, thanks!<br>\nI'm not angry with them, just curious.. 🤔</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2877660,
              "author_name": "slime",
              "author_url": "",
              "post_date": "2024-06-18T13:57:40.300000",
              "content": "<p>New labels means re-training, I would rather avoid that ;D</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2877667,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-06-18T14:04:06.017000",
              "content": "<p>Well, they weren't banned yet.<br>\nIt's not like hosts don't know we have access to train spacetime. It's not some secret. only test data was shuffled and obviously meant to not have spacetime data.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2877673,
              "author_name": "Camaro",
              "author_url": "",
              "post_date": "2024-06-18T14:08:52.450000",
              "content": "<p><a href=\"https://www.kaggle.com/martynoveduard\" target=\"_blank\">@martynoveduard</a> <br>\nAh, that part? But even if there's no label, we might use the old test set with pseudo-labeling or something like that. Especially if the new test set is sampled from the same timestamp, that kind of approach would shine.<br>\nI feel it's a more meaningless effort. 🤔<br>\nBut thanks for the feedback;)</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2877720,
              "author_name": "Truth Seeker",
              "author_url": "",
              "post_date": "2024-06-18T14:34:39.877000",
              "content": "<p>Releasing the ground truth of a test set before the competition ends, even if it's an old test set with an undesirable side effect, is not a good thing to do.  The probability of <a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> doing that is at most zero.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2877744,
              "author_name": "Camaro",
              "author_url": "",
              "post_date": "2024-06-18T14:47:26.520000",
              "content": "<p>Well, it's almost a synthetic dataset, so there is no inherent difference between train and test by nature. But if the host creates a new test set for the next year to avoid overlap between the old and new test sets, there will be a larger gap between train and test, which (maybe) makes prediction more difficult.</p>\n<p>But I'm now understanding everyone's concerns. Thanks!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2876930,
          "author_name": "HideBu",
          "author_url": "",
          "post_date": "2024-06-18T04:51:18.290000",
          "content": "<p><a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> <br>\nThank you for update.<br>\nAnd I have a question.</p>\n<ul>\n<li>What does <code>multi-atmospheric-column to single-atmospheric-column and single-atmospheric-column to single-atmospheric-column approaches</code> mean?</li>\n</ul>",
          "votes": 2,
          "replies": [
            {
              "id": 2877576,
              "author_name": "Anthony Meza",
              "author_url": "",
              "post_date": "2024-06-18T13:05:46.927000",
              "content": "<p>essentially it means that methods that explicitly use location information (lat/lon) are against the rules of the competition</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2877613,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-06-18T13:19:58.810000",
              "content": "<p>No. The rules are the rules (under the tab 'Rules') and nothing in the rules says anything like that. We have train spacetime labels and we definitely can use them. (for example for pseudo-labeling etc.)<br>\nIf host want to forbid it, he needs to do it clearly and formally by changing the rules. Which he did not.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2878313,
              "author_name": "Jerry Lin",
              "author_url": "",
              "post_date": "2024-06-18T21:21:21.447000",
              "content": "<p>The entire purpose of this update is to remove multi-atmospheric-column approaches (e.g. 2D to 2D or 2D to 1D regression models) from the leaderboard.</p>",
              "votes": 3,
              "replies": []
            }
          ]
        },
        {
          "id": 2876994,
          "author_name": "steubk",
          "author_url": "",
          "post_date": "2024-06-18T05:29:23.733000",
          "content": "<p><a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> , thank you for update.</p>\n<p>Could you elaborate further on the new sample submission (weighting file)? In my opinion, only the first column (sample_id) should change, while the weights for each target should remain unchanged. If this is not the case, assumptions might be made about the new test set.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2877025,
              "author_name": "Jerry Lin",
              "author_url": "",
              "post_date": "2024-06-18T06:16:42.213000",
              "content": "<p>The only purpose of the sample submission is to zero out stratospheric moistening variables that are extremely difficult to predict and cause abnormally low r squared values. I made the rest of the values 1 for convenience sake (they weren't 1 before).</p>",
              "votes": 8,
              "replies": []
            }
          ]
        },
        {
          "id": 2877747,
          "author_name": "PC Jimmmy",
          "author_url": "",
          "post_date": "2024-06-18T14:49:10.040000",
          "content": "<p>Suggest you pin this to the top </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2868875,
      "author_name": "Jerry Lin",
      "author_url": "",
      "post_date": "2024-06-12T17:26:02.667000",
      "content": "<p>Looking into this now. Thank you for the heads up---this is alarming.</p>",
      "votes": 16,
      "replies": [
        {
          "id": 2875441,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-06-17T06:32:47.913000",
          "content": "",
          "votes": -1,
          "replies": []
        }
      ]
    },
    {
      "id": 2870966,
      "author_name": "Jerry Lin",
      "author_url": "",
      "post_date": "2024-06-14T00:13:06.780000",
      "content": "<p>Update: We are addressing this issue and will provide an update by next week.</p>",
      "votes": 12,
      "replies": [
        {
          "id": 2870973,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2024-06-14T00:29:55.803000",
          "content": "<p>This is very close to comp' end already. Depend on the update and how much impact it will have on existing solutions, it might be worth to consider an extension to the competition. <a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> </p>",
          "votes": 5,
          "replies": [
            {
              "id": 2870979,
              "author_name": "Jerry Lin",
              "author_url": "",
              "post_date": "2024-06-14T01:00:24.887000",
              "content": "<p>Yes we are thinking about this.</p>",
              "votes": 7,
              "replies": []
            },
            {
              "id": 2872169,
              "author_name": "Camaro",
              "author_url": "",
              "post_date": "2024-06-14T16:02:47.180000",
              "content": "<p>Don't forget to extend the deadline if you replace the test dataset!:)</p>",
              "votes": -1,
              "replies": []
            }
          ]
        },
        {
          "id": 2870991,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-06-14T01:25:41.670000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2870998,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2024-06-14T01:50:23.077000",
          "content": "<p><a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> Thank you for the reply.</p>\n<p>If \"update\" means replacement of test set to different sub-sampling of year 9-10, I guess we can infer timestamp and location data of new dataset from currently released test set (by comparing similarity of two feature vectors etc.). And according to  Andrew Gaiss's experiment[1], location data is almost perfectly predictable from the atmospheric features (~96% according to his report). So, could you also confirm the indirect/implicit usage of these information? I believe this must be inevitable according to the fact that current atmospheric feature already contains good amount of location-specific information.</p>\n<ul>\n<li>[1] <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/511911#2869607\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/511911#2869607</a></li>\n</ul>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 2871326,
          "author_name": "kami",
          "author_url": "",
          "post_date": "2024-06-14T06:57:22.137000",
          "content": "<p><a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> </p>\n<p>I urge caution when considering updating the test data.</p>\n<ul>\n<li><p>Changing the data in the middle of the competition would alter the competition itself. Participants who have invested time and money following the existing rules may find their efforts wasted, which is unfair. Many participants who are not violating any rules may be disadvantaged.</p></li>\n<li><p>Using the new test data could increase the risk of it being hacked, as the already provided test data is out there [1]. The update may cause more harm than good, with little to no benefit.</p></li>\n<li><p>We are using location data. We attempted to use future data with timestamps but were unsuccessful. Our solution, which does not use future data, remains valuable. Additionally, <strong>comparing solutions that use location data with those that do not can provide insights into the usefulness of this data, which could be beneficial to the field.</strong></p></li>\n</ul>\n<p>[1] <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/511911#2869193\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/511911#2869193</a></p>",
          "votes": 11,
          "replies": [
            {
              "id": 2871374,
              "author_name": "Gunes Evitan",
              "author_url": "",
              "post_date": "2024-06-14T07:12:15.820000",
              "content": "<p>So the gap between 0.788 and 0.798 is because of the location leak?</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2871384,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-06-14T07:18:21.810000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2871402,
              "author_name": "ONODERA",
              "author_url": "",
              "post_date": "2024-06-14T07:25:57.037000",
              "content": "<blockquote>\n  <p>We are using location data. We attempted to use future data with timestamps but were unsuccessful.</p>\n</blockquote>\n<p>Good to know. Your team could exploit something from latlon, but lead data.<br>\nMaybe NearestNeighbor or something like that.</p>\n<p>I think if the host provides new shuffled test data, we can estimate latlon and timestamp. This might be the problem.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2871431,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-06-14T07:41:14.913000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2871433,
              "author_name": "kami",
              "author_url": "",
              "post_date": "2024-06-14T07:42:06.587000",
              "content": "<blockquote>\n  <p>So the gap between 0.788 and 0.798 is because of the location leak?</p>\n</blockquote>\n<p>We use it, but I don't know if other teams use it.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2871528,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-06-14T08:58:08.193000",
              "content": "<p>Well there is a very big difference between 'estimate' based on train data and between 'know' from leak in test…</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2871679,
              "author_name": "ekffar",
              "author_url": "",
              "post_date": "2024-06-14T10:38:18.263000",
              "content": "<p>I can confirm that LB 0.788 can be achieved without using Location and Timestamp leak, I also wonder to what extent can location leak improve predictive performance</p>",
              "votes": 6,
              "replies": []
            },
            {
              "id": 2871744,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-06-14T11:29:20.973000",
              "content": "<p>In one of my experiments I fed the model location and timestamp and got an improvement of ~0.005. Only on train data of course, I could not apply this model to test (well it was before I knew about leak) it was just out of curiosity hahaha.<br>\nI would say it can definitely help, but it's not a game changer.<br>\nIt's not some secret to propel people from ~0.75 to ~0.8. But if 1st and 2nd are very close, it can make the difference.</p>",
              "votes": 7,
              "replies": []
            },
            {
              "id": 2871872,
              "author_name": "ryches",
              "author_url": "",
              "post_date": "2024-06-14T12:41:22.900000",
              "content": "<p>Yep, second confirmation I am not using latlon or any doing anything to exploit the time series</p>",
              "votes": 7,
              "replies": []
            },
            {
              "id": 2871900,
              "author_name": "Anthony Meza",
              "author_url": "",
              "post_date": "2024-06-14T12:58:44.403000",
              "content": "<p>There’s already a lot of research data showing that location is important for reproducing model output, but a non-location dependent weather parameterization would be more ideal.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2871932,
              "author_name": "Anthony Meza",
              "author_url": "",
              "post_date": "2024-06-14T13:23:04.683000",
              "content": "<p>My guess is that those using lat/lon are using it to reorganize the data into spatial maps of atmospheric variables, and then applying some sort of CNN to those spatial maps. </p>\n<p>A way to guard against this (since these solutions are not what the organizers want) would be to sprinkle a new NaNs into some rows of the full test data set. Full NaN rows would break any CNNs based on global maps, but not those using a column approach. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2871968,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-06-14T13:52:18.407000",
              "content": "<p>It's not likely. Using location most probably means just feeding the model the location of each sample. I don't believe a 2D CNN has any chance in this competition. Good luck to anyone who tries that.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2874089,
              "author_name": "Joseph Zhou",
              "author_url": "",
              "post_date": "2024-06-16T03:31:09.143000",
              "content": "<p><a href=\"https://www.kaggle.com/ekffar\" target=\"_blank\">@ekffar</a> Also without original(HugginFace) data?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2874437,
              "author_name": "Truth Seeker",
              "author_url": "",
              "post_date": "2024-06-16T09:05:52.453000",
              "content": "<p><a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a></p>\n<p>It is best for any dataset update (should the host decide its necessity) to not depend on any assumption whatsoever about the learning/inference algorithms being used.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2871790,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-06-14T11:55:41.807000",
          "content": "",
          "votes": -1,
          "replies": []
        },
        {
          "id": 2871930,
          "author_name": "Truth Seeker",
          "author_url": "",
          "post_date": "2024-06-14T13:21:52.327000",
          "content": "<p><a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> </p>\n<p>If the competition host releases another new test set for us to work on instead, this should <em>not</em> be a problem at all as long as the new test set is generated/produced/collected/whatever-verb-is-appropriate from the same way as the training set.  In machine learning jargon, we say that \"The training and test sets come from the same/similar distribution.\"</p>\n<p>In fact, if a model should work well on the old test set but not on a new test set, that probably means the model does not generalize well.  But the whole point of machine learning is to find a model that generalizes well.  If there were a God of Machine Learning, I'm sure He would care most about learning algorithms that generalize well, and could care less about any leaderboard score shakeup.</p>\n<p>If the competition host decides it's necessary to also release another new training set for us to work on, that is also <em>not</em> a problem because of the following reasons:</p>\n<p>(1) If the new training set is provided with any undesirable leak information somehow removed, then those models that have not been taking advantage of the leak should have minimal impact, if at all.</p>\n<p>(2) Even if there exists any nonzero impact to any models from the new training set, orthogonal to the removal of the undesirable leak, it would apply approximately equally to every model anyway.  So, everyone would be impacted roughly the same way, anyway.</p>\n<p>(3) Worst comes worst, we just need to retrain our model on the new training set.  Big deal (sarcasm intended).  The only actual problem I see with this is the time it would take to train again.  But that is easily fixed if we get some extension to the deadline.</p>\n<p>In any case, this is a great competition.  A super large multiple-target regression problem to work on is a heaven Carl Friedrich Gauss could only dream of.  Competition host, thank you for providing this wonderful competition, and I will take it to the end whatever the host decides to do.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2872003,
              "author_name": "slime",
              "author_url": "",
              "post_date": "2024-06-14T14:11:07.530000",
              "content": "<blockquote>\n  <p>The only actual problem I see with this is the time it would take to train again. But that is easily fixed if we get some extension to the deadline.</p>\n</blockquote>\n<p>you have no idea how much compute this comp requires :D </p>",
              "votes": 5,
              "replies": []
            },
            {
              "id": 2872024,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-06-14T14:21:03.990000",
              "content": "<p>Hahahahaha true that. I won't be surprised if it's the most demanding in the history of Kaggle, outside of genAi comp' where there is no limit to compute requirements if you want to generate more samples.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2872030,
              "author_name": "Truth Seeker",
              "author_url": "",
              "post_date": "2024-06-14T14:26:37.627000",
              "content": "<blockquote>\n  <p>you have no idea how much compute this comp requires :D </p>\n</blockquote>\n<p>I'm very aware of it :D</p>\n<p>Efficiency, and not just test set accuracy, is part of the competition.  That's just how things work in real life, too.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2872039,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-06-14T14:33:53.657000",
              "content": "<blockquote>\n  <blockquote>\n    <p>you have no idea how much compute this comp requires :D </p>\n  </blockquote>\n  <p>I'm very aware of it :D</p>\n  <p>Efficiency, and not just test set accuracy, is part of the competition.  That's just how things work in real life, too.</p>\n</blockquote>\n<p>Nah efficiency is maybe good for silver range.<br>\nFor gold range, I'm very doubtful anyone manged to get there 'efficiently' without throwing a LOT of compute on the data.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2872133,
              "author_name": "Truth Seeker",
              "author_url": "",
              "post_date": "2024-06-14T15:45:13.720000",
              "content": "<p><a href=\"https://www.kaggle.com/shlomoron\" target=\"_blank\">@shlomoron</a> , sure, I hear you.  But an extension of sufficient length should eliminate such a concern.  My point is, we should all strive for sound machine learning practices, even if it might mean my leaderboard score might drop.  To that regard, I think we're on the same side 😀</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2872183,
              "author_name": "ryches",
              "author_url": "",
              "post_date": "2024-06-14T16:15:05.347000",
              "content": "<p>I don't think I've thrown that much compute at the problem but I guess everyone has different thresholds for that. I can retrain my model on two 4090's in a couple days.</p>\n<p>Maybe people are doing way more than me for this but I don't feel like it's the most compute intensive that I've done. Deep fake, Lyft, some of the other large scale image ones</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2872708,
              "author_name": "Anthony Meza",
              "author_url": "",
              "post_date": "2024-06-15T02:12:34.893000",
              "content": "<p>I’ve only trained my model for one hour maximum, I guess that’s what’s been holding me back 😂</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2874513,
              "author_name": "Amedeo Biolatti",
              "author_url": "",
              "post_date": "2024-06-16T10:36:19.297000",
              "content": "<p>Unless you mean 1 hour on a H100 cluster</p>",
              "votes": 2,
              "replies": []
            }
          ]
        },
        {
          "id": 2872031,
          "author_name": "kurupical",
          "author_url": "",
          "post_date": "2024-06-14T14:27:12.580000",
          "content": "<p><a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a><br>\nFirst and foremost, we should consider the following two points separately, prioritizing the former:</p>\n<ul>\n<li>Not changing the competition rules</li>\n<li>Adhering to the host's objectives</li>\n</ul>\n<p>If there is a \"critical leak,\" such as the inclusion of training data in the test data, as happened in the past with the Foursquare - Location Matching competition, then it may be unavoidable to change the rules in such cases.</p>\n<p>However, in this particular instance, I don't believe it is critical enough to warrant changing the rules near the end of the competition. Although the current data is ordered by timestamps, the time gaps are too large to be effectively used. I don't believe this is a critical leak significant enough to change the data and compromise the \"competitiveness.\"</p>\n<p>Additionally, can we really say that being able to determine the latitude and longitude is a \"critical leak\"?  The E3SM simulator appears to use information from all grids. I believe that solutions using simulations that incorporate information from all grids have value. (at least, I am confident that our team can give you a solution that is much more valuable than the \"stop lightgbm early stopping\" and \"make the neural network as long as possible\" asked as foursquare's solution. )</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2872065,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-06-14T14:58:01.680000",
              "content": "<p>The problem with spacetime, imho is that it distracts from building the best model that depends only on the inputs/physical data.<br>\nSure, spacetime data can help, but we already knew that. Now, we won't know how far we can push without it. People will go crazy in the next two weeks trying to build the smartest spacetime models instead of the best model that depends only on the same sample data. It's a shame, really.<br>\nWell, it's also a shame for teams that used leak like yours. No good solutions here, really, only less harmful ones.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2872148,
              "author_name": "ONODERA",
              "author_url": "",
              "post_date": "2024-06-14T15:49:40.303000",
              "content": "<blockquote>\n  <p>I believe that solutions using simulations that incorporate information from all grids have value. </p>\n</blockquote>\n<p>We are not host. Who knows</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2872181,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-06-14T16:14:42.307000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2872184,
              "author_name": "Camaro",
              "author_url": "",
              "post_date": "2024-06-14T16:15:18.077000",
              "content": "<p>Now I'm a bit surprised to see them so persistent for this leak and started to think that it must be a key to break 0.79 and more…🤔</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2872225,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-06-14T17:01:40.450000",
              "content": "<blockquote>\n  <p>Now I'm a bit surprised to see them so persistent for this leak and started to think that it must be a key to break 0.79 and more…🤔</p>\n</blockquote>\n<p>Had the same thought lol. Probably there is some smart way to use adjacent /mean grid data or something similar</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2872602,
              "author_name": "kurupical",
              "author_url": "",
              "post_date": "2024-06-14T23:15:16.433000",
              "content": "<blockquote>\n  <p>We are not host. Who knows</p>\n</blockquote>\n<p>As you said, only the host can know whether it will be useful or not. So, it’s just that \"I believe\"😀.</p>",
              "votes": 5,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2869411,
      "author_name": "steubk",
      "author_url": "",
      "post_date": "2024-06-13T05:03:06.087000",
      "content": "<p>Nice shot <a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> !</p>\n<p>I made a visualization of air temperature at level 60 in Celsius for train and test (using your trick) at time t=0  </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F214989%2F13f0df1992afedd1e0d56f78cc854fd1%2Fleap.png?generation=1718254862105057&amp;alt=media\"></p>\n<p>Here’s my two cents :</p>\n<ul>\n<li>I don't really like competitions with synthetic data because, in the end, your goal becomes to reverse engineer the data generator. But this is a reverse engineering competition. So, it's hard to say: your goal is to reverse engineer this emulator, but you can't use this or that.</li>\n<li>Never underestimate kagglers 😀.</li>\n</ul>",
      "votes": 9,
      "replies": [
        {
          "id": 2869423,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2024-06-13T05:14:17.517000",
          "content": "<p>Year, definitely this competition is kind of reverse-engineering competition for climate simulator LOL.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 2869436,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2024-06-13T05:26:59.047000",
          "content": "<p>Never underestimate kagglers should by this site motto.<br>\nIn Latin, as per tradition.<br>\nNUMQUAM MINORIS AESTIMO KAGGLERS</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2868822,
      "author_name": "Felix Yang",
      "author_url": "",
      "post_date": "2024-06-12T16:43:57.387000",
      "content": "<p>I guess the easiest fix would be to release a new version of the test set with all rows properly shuffled and all the sample IDs scrambled. This would break all current submissions, but we could simply restore the previous leaderboard scores by rerunning our models on the new test set, since the data rows are still the same, just reordered. Otherwise, everyone would be forced to retrain their models using this leak to remain competitive, which is probably not the purpose of this competition?</p>",
      "votes": 6,
      "replies": [
        {
          "id": 2869193,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2024-06-12T22:54:09.317000",
          "content": "<blockquote>\n  <p>I guess the easiest fix would be to release a new version of the test set with all rows properly shuffled and all the sample IDs scrambled.</p>\n</blockquote>\n<p>In my feeling, this would not solve the issue, because knowing subsampling of the dataset, Kaggle’s would successfully reverse engineering the new subsampling using information of already released subsampling (checking similarity of feature vector etc.). The only way of releasing non-leaked dataset is to newly simulate another two year data (namely year 11-12), but I doubt host could make this considering not cheap amount of computing resource doing this.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2869195,
              "author_name": "Bilzard",
              "author_url": "",
              "post_date": "2024-06-12T22:59:33.203000",
              "content": "<blockquote>\n  <p>Otherwise, everyone would be forced to retrain their models using this leak to remain competitive, which is probably not the purpose of this competition?</p>\n</blockquote>\n<p>I am also not sure this is possible, considering the case the model implicitly exploiting this leak.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2869304,
              "author_name": "Bilzard",
              "author_url": "",
              "post_date": "2024-06-13T02:37:40.763000",
              "content": "<p>Considering the point I wrote above, I must say the most probable action they can take is <em>leave everything as it is, and just grant usage of timestamp and location data</em>. However, final decision is leave for host and Kaggle team.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2869601,
              "author_name": "Felix Yang",
              "author_url": "",
              "post_date": "2024-06-13T06:47:44.797000",
              "content": "<blockquote>\n  <p>because knowing subsampling of the dataset, Kaggle’s would successfully reverse engineering the new subsampling using information of already released subsampling</p>\n</blockquote>\n<p>You're right! I didn't think of that.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2869008,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2024-06-12T19:24:02.537000",
      "content": "<p>Wow, respect for sharing this when you could possibly have found ways to exploit this.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2869176,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2024-06-12T22:34:17.837000",
          "content": "<p>I was surprised to hear not few Kaggles already found this in the early stage of competition.<br>\nI don’t know Kaggle’s convention, but I have to say I am as novice as reporting this as an anomaly when someone is walking in the street upside down.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2869334,
              "author_name": "CPMP",
              "author_url": "",
              "post_date": "2024-06-13T03:15:06.707000",
              "content": "<p>Apparently many have found this, but they did not share it as you did. And those who found it and disclose it now say it is not useful.</p>",
              "votes": 6,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2868971,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "2024-06-12T18:55:35.943000",
      "content": "<p>I made some plots of this early on but haven't done anything to exploit it. Assumed it was out of the spirit of the competition or simply wouldn't be of any value on the private set</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2869007,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2024-06-12T19:23:47.550000",
          "content": "<p>This is Kaggle, I know nothing about this 'spirit of the competition' lol. Whatever we can exploit, as long as it is within the rules- we exploit haha</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2878456,
      "author_name": "Jerry Lin",
      "author_url": "",
      "post_date": "2024-06-19T00:36:33.447000",
      "content": "<p>The competition should be good to go now. Thank you for your patience.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2868774,
      "author_name": "ONODERA",
      "author_url": "",
      "post_date": "2024-06-12T16:08:04.037000",
      "content": "<p>What kind of action are you requesting?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2868824,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2024-06-12T16:47:28.373000",
          "content": "<p>Well, in the early days of this competition, one of host member commented on usage of location &amp; timestamp data, and from this comment, I thought the hosts are expecting models without these information.</p>\n<p>So at least, I want to make sure using these information is valid on this competition.</p>\n<p><a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/511911\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/511911</a></p>\n<p>Since some specific action is not on my mind, I rewrote this discussion's title to \"request to confirmation\".</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2868830,
              "author_name": "ONODERA",
              "author_url": "",
              "post_date": "2024-06-12T16:55:19.927000",
              "content": "<p>I'm sure a lot of kaggle grandmasters were aware of this property. But not sure if top teams could utilize them.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2868820,
      "author_name": "ADAM.",
      "author_url": "",
      "post_date": "2024-06-12T16:42:34.753000",
      "content": "<p>good catch!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2877488,
      "author_name": "Mahmoud Elshahed",
      "author_url": "",
      "post_date": "2024-06-18T12:01:41.133000",
      "content": "<p>Just for clarity, train data and train saved weights will be the same \"if we haven't utilize or reverse the timestamp and location\"<br>\nso that we don't repeat all from the beginning ? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2878104,
          "author_name": "Jerry Lin",
          "author_url": "",
          "post_date": "2024-06-18T18:25:38.183000",
          "content": "<p>Assuming you are not using a multi-atmospheric-column approach (which requires exploiting the structure of the data in a way that we did not intend), you shouldn't have to do any retraining.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2870247,
      "author_name": "VITALIY",
      "author_url": "",
      "post_date": "2024-06-13T13:33:39.540000",
      "content": "<p>Can the authors clarify whether it is legal to use latitude and longitude? <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/495258#2768949\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/495258#2768949</a> from this looks like not.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2870961,
          "author_name": "PC Jimmmy",
          "author_url": "",
          "post_date": "2024-06-13T23:50:33.647000",
          "content": "<p>For sure it can be used to root cause solutions if r2 values have a localized location.  If you identify that the north pole causes some labels to be bad, feature engineering (not including the lat/long) could be added based on research into the world of weather prediction.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2868983,
      "author_name": "Amedeo Biolatti",
      "author_url": "",
      "post_date": "2024-06-12T19:09:53.307000",
      "content": "<p>Are models trained with location/time id any better? Did someone try?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2869172,
          "author_name": "Anthony Meza",
          "author_url": "",
          "post_date": "2024-06-12T22:15:17.900000",
          "content": "<p>My guess would be yes, most climate emulators take spatial information and time into account. Would be interesting to see if more people have tried this</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2870333,
          "author_name": "Zhuoqun Li",
          "author_url": "",
          "post_date": "2024-06-13T14:17:36.193000",
          "content": "<p>Purely intuitive there will be an improvement</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2875318,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-06-17T03:30:07.680000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 2875321,
          "author_name": "Juan D C F",
          "author_url": "",
          "post_date": "2024-06-17T03:37:15.527000",
          "content": "<p>Thanks, that is what I had read. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2875325,
          "author_name": "Juan D C F",
          "author_url": "",
          "post_date": "2024-06-17T03:39:50.113000",
          "content": "<p>“This leaderboard is calculated with approximately 20% of the test data. The final results will be based on the other 80%, so the final standings may be different.”</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2875364,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2024-06-17T04:55:22.813000",
          "content": "<p><a href=\"https://www.kaggle.com/juandcf\" target=\"_blank\">@juandcf</a> </p>\n<blockquote>\n  <p>This leaderboard is calculated with approximately 20% of the test data. The final results will be based on the other 80%, so the final standings may be different.</p>\n</blockquote>\n<p>This means public LB <em>score</em> is calculated by 20% of test set. Private LB is calculated by rest of <em>80%</em> of data. So in usual competition setup, data swap never happens.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2875417,
          "author_name": "Juan D C F",
          "author_url": "",
          "post_date": "2024-06-17T06:11:51.570000",
          "content": "<p>Got it, thanks!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2873023,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-06-15T09:09:43.313000",
      "content": "",
      "votes": -3,
      "replies": []
    },
    {
      "id": 2870934,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-06-13T22:30:31.367000",
      "content": "",
      "votes": -4,
      "replies": []
    },
    {
      "id": 2870336,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-06-13T14:18:52.613000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2869376,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-06-13T03:56:57.313000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2868812,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-06-12T16:29:58.920000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2877664,
      "author_name": "German Arley",
      "author_url": "",
      "post_date": "2024-06-18T14:02:37.310000",
      "content": "<p>Thank you for the update</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2871076,
      "author_name": "Harshita Sapra",
      "author_url": "",
      "post_date": "2024-06-14T04:08:08.447000",
      "content": "<p>Thank you </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2868744": "Hi. Kaggle Team & Host members,\n\nRelated to previous discussion[1], we found location and timestamp data on the test set is possible to be reverse-engineered.\nHowever, in the early days of this competition, one of host member commented on the availability of these data, and answered they don't expect to utilize them[2]. So we want to make sure using these information is valid on this competition.\n\nCould you confirm we can utilize these information?\n\n@ashleychow @jerrylin96\n\n\n## How to reproduce\n\n1. drop `sample_id`\n2. add new column `sample_id` based on the current order of test data\n\nThen, we found the processed test data is ordered by each 384 locations and time steps.\n\n```python\nimport polars as pl\n\nprefix = \"test\"\ntest = (\n    pl.read_csv(Path(\"../../../input/leap-atmospheric-physics-ai-climsim/test.csv\"))\n    .drop(\"sample_id\")\n    .with_row_index(\"sample_id\")\n    .with_columns(pl.col(\"sample_id\").mod(384).alias(\"grid_id\"))\n)\n```\n\n## Evidence\n\nWhen taking statistics per reverse-engineered location IDs, we can clearly see specific patterns which found on ordered train/validation data.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F140f93019fc2b9fc55c2ea06979620b3%2FScreenshot%202024-06-13%20at%200.42.31.png?generation=1718206981274040&alt=media)\n\n## Reference\n\n1. https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/511274#286771\n2. https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/495258#2768949",
    "2869607": "We are all likely already using geographical information whether the data are shuffled or not.\n\nThe atmosphere has a consistent structure that can be related to geographical location. For example, tropopause height alone is a fairly good predictor of distance from the equator (|latitude|). Some of the per-column predictors also obviously convey information about location (e.g. SZA, land fraction, and sea ice fraction).\n\nThis evening, I tried training a 4-layer MLP classifier to predict the E3SM grid-point that a sample originated from. It reaches ~96% accuracy after one pass through the “low-res” dataset and could surely be improved with minimal effort and additional training.\n\nMy point is that lat/lon information is already encoded in the individual rows of the test set and shuffling the data won't remove it. A relatively simple MLP can extract locations so, to the extent that this information is relevant to the problem (and I think that it is), those of us using more sophisticated models are likely leveraging this information already without providing it explicitly as an input.",
    "2876590": "**UPDATE**: Some big changes to the competition are happening tomorrow.\n\n1.) There's going to be a **brand new test-set, sample submission (weighting file), and benchmark submission**.\n2.) The competition **deadline will be extended two weeks** for everyone.\n3.) All current submissions will be invalidated. **You must resubmit using the new data** to appear on the leaderboard.\n4.) If you have been using the single-atmospheric-column to single-atmospheric-column regression approach, your score shouldn't change much.\n5.) *A more detailed post will appear tomorrow*.\n\nFor those who have been using the multi-column to single-column approach, we apologize. However, making full use of this approach would have required knowing the sample IDs were scrambled, implying you knew this was \"against the spirit of the competition\" to quote @ryches . \n\nTLDR: we cannot have multi-atmospheric-column to single-atmospheric-column and single-atmospheric-column to single-atmospheric-column approaches coexisting on the leaderboard, especially if the former was done knowing it went against the intentions of the competition.\n\nEDIT: By column I mean \"atmospheric column\" not csv column. Sorry if I scared you!",
    "2868875": "Looking into this now. Thank you for the heads up---this is alarming.",
    "2870966": "Update: We are addressing this issue and will provide an update by next week.",
    "2869411": "Nice shot @tatamikenn !\n\nI made a visualization of air temperature at level 60 in Celsius for train and test (using your trick) at time t=0  \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F214989%2F13f0df1992afedd1e0d56f78cc854fd1%2Fleap.png?generation=1718254862105057&alt=media)\n\nHere’s my two cents :\n\n- I don't really like competitions with synthetic data because, in the end, your goal becomes to reverse engineer the data generator. But this is a reverse engineering competition. So, it's hard to say: your goal is to reverse engineer this emulator, but you can't use this or that.\n- Never underestimate kagglers 😀.\n",
    "2868822": "I guess the easiest fix would be to release a new version of the test set with all rows properly shuffled and all the sample IDs scrambled. This would break all current submissions, but we could simply restore the previous leaderboard scores by rerunning our models on the new test set, since the data rows are still the same, just reordered. Otherwise, everyone would be forced to retrain their models using this leak to remain competitive, which is probably not the purpose of this competition?",
    "2869008": "Wow, respect for sharing this when you could possibly have found ways to exploit this.",
    "2868971": "I made some plots of this early on but haven't done anything to exploit it. Assumed it was out of the spirit of the competition or simply wouldn't be of any value on the private set",
    "2878456": "The competition should be good to go now. Thank you for your patience.",
    "2868774": "What kind of action are you requesting?",
    "2868820": "good catch!",
    "2877488": "Just for clarity, train data and train saved weights will be the same \"if we haven't utilize or reverse the timestamp and location\"\nso that we don't repeat all from the beginning ? ",
    "2870247": "Can the authors clarify whether it is legal to use latitude and longitude? https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/495258#2768949 from this looks like not.",
    "2868983": "Are models trained with location/time id any better? Did someone try?",
    "2875318": "",
    "2873023": "",
    "2870934": "",
    "2870336": "",
    "2869376": "",
    "2868812": "",
    "2877664": "Thank you for the update",
    "2871076": "Thank you "
  }
}