{
  "id": 502484,
  "title": "A severely underrated comment which raises a lot of questions",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/502484",
  "author_name": "DennisSakva",
  "post_date": "2024-05-13T16:56:50.448000",
  "votes": 37,
  "comment_count": 14,
  "views": 0,
  "content": "<p>If you haven't seen it already there's a simple heuristic that yields +0.03 to your LB score.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F349155%2F503540951a5e148f54901fd643b8fd21%2Fptent_q2.PNG?generation=1715619199918582&amp;alt=media\"><br>\nCourtesy of <a href=\"https://www.kaggle.com/jano123\" target=\"_blank\">@jano123</a> (show him some love here <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/499896\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/499896</a>)<br>\nBut my question is: why is my X million parameter model not picking this simple linear relationship up and just outputs garbage for these columns? What's going on here?</p>",
  "messages": [
    {
      "id": 2811230,
      "postDate": "2024-05-13T16:56:50.447Z",
      "content": "<p>If you haven't seen it already there's a simple heuristic that yields +0.03 to your LB score.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F349155%2F503540951a5e148f54901fd643b8fd21%2Fptent_q2.PNG?generation=1715619199918582&amp;alt=media\"><br>\nCourtesy of <a href=\"https://www.kaggle.com/jano123\" target=\"_blank\">@jano123</a> (show him some love here <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/499896\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/499896</a>)<br>\nBut my question is: why is my X million parameter model not picking this simple linear relationship up and just outputs garbage for these columns? What's going on here?</p>",
      "rawMarkdown": "If you haven't seen it already there's a simple heuristic that yields +0.03 to your LB score.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F349155%2F503540951a5e148f54901fd643b8fd21%2Fptent_q2.PNG?generation=1715619199918582&alt=media)\nCourtesy of @jano123 (show him some love here https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/499896)\nBut my question is: why is my X million parameter model not picking this simple linear relationship up and just outputs garbage for these columns? What's going on here?",
      "votes": 36
    },
    {
      "id": 2811881,
      "postDate": "2024-05-14T01:26:29.360Z",
      "content": "<p>if you look at the code preparing the training data, you will see that all the vector targets are calculated as y[t] = (x[t+1]-x[t]) / 1200 (samples are 20 minutes=1200 seconds apart).</p>\n<p>x[t] is given to us, x[t+1] is not. For columns where large proportion of values are zeros, assuming that x[i+1]=0 will be correct most of the time. Then y[t]= (0-x[t]) / 1200 = -x[t] / 1200 is the correct answer.</p>\n<p>This only works for the few columns where most values are 0.</p>",
      "rawMarkdown": "if you look at the code preparing the training data, you will see that all the vector targets are calculated as y[t] = (x[t+1]-x[t]) / 1200 (samples are 20 minutes=1200 seconds apart).\n\nx[t] is given to us, x[t+1] is not. For columns where large proportion of values are zeros, assuming that x[i+1]=0 will be correct most of the time. Then y[t]= (0-x[t]) / 1200 = -x[t] / 1200 is the correct answer.\n\nThis only works for the few columns where most values are 0.",
      "votes": 11,
      "replies": [
        {
          "id": 2812239,
          "postDate": "2024-05-14T06:49:21.553Z",
          "content": "<p>but take for example q0002_25 it seems to be mostly non zero, yet this relation still holds and you get a score of 1 with the relation over the entire train set. it seems more like they made an approximation here, maybe because these columns might be very negligible to the overall simulation or something, and for some reason they decide it starts becoming appreciable at around q0002_27 again</p>",
          "rawMarkdown": "but take for example q0002_25 it seems to be mostly non zero, yet this relation still holds and you get a score of 1 with the relation over the entire train set. it seems more like they made an approximation here, maybe because these columns might be very negligible to the overall simulation or something, and for some reason they decide it starts becoming appreciable at around q0002_27 again",
          "votes": 1,
          "replies": [
            {
              "id": 2813590,
              "postDate": "2024-05-14T21:14:33.450Z",
              "content": "<p>On further reading i see that targets are defined as (Tafter − Tbefore )/∆t, where variables are recorded before versus after the convection and radiation calculations.</p>\n<p>Looks like for situations with low prevalence of clouds they just set Tafter to 0, instead of setting it to Tbefore. </p>\n<p>So not very clean data. As a result a large portion of score improvement may be coming from hacking the data - similar to what you discovered in \"Home Credit\" competition.</p>\n<p>Perhaps organizers should drop the questionable columns from objective - then the results would better serve their purpose.</p>",
              "rawMarkdown": "On further reading i see that targets are defined as (Tafter − Tbefore )/∆t, where variables are recorded before versus after the convection and radiation calculations.\n\nLooks like for situations with low prevalence of clouds they just set Tafter to 0, instead of setting it to Tbefore. \n\nSo not very clean data. As a result a large portion of score improvement may be coming from hacking the data - similar to what you discovered in \"Home Credit\" competition.\n\nPerhaps organizers should drop the questionable columns from objective - then the results would better serve their purpose.",
              "votes": 1
            },
            {
              "id": 2814181,
              "postDate": "2024-05-15T07:49:37.583Z",
              "content": "<p>could be a million reasons, they're basically saying \"if cloud then cloud decay exponentially\", maybe there is a very small variable dependent on many things, which just gets rounded to 0 high up in the atmosphere in the low resolution data (of course it's probably very reasonable to assume that any cloud there is gone in 1200s)…</p>\n<p>who knows</p>",
              "rawMarkdown": "could be a million reasons, they're basically saying \"if cloud then cloud decay exponentially\", maybe there is a very small variable dependent on many things, which just gets rounded to 0 high up in the atmosphere in the low resolution data (of course it's probably very reasonable to assume that any cloud there is gone in 1200s)...\n\nwho knows"
            }
          ]
        }
      ]
    },
    {
      "id": 2811598,
      "postDate": "2024-05-13T19:14:22.027Z",
      "content": "<p>First of all, thank you <a href=\"https://www.kaggle.com/jano123\" target=\"_blank\">@jano123</a> for sharing this insight!</p>\n<p>The relationship between these q0002 variables is indeed very simple:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14891093%2Fc136ac66f62af4750204b4d2a746377e%2F01.png?generation=1715626817753826&amp;alt=media\"></p>\n<p>However, if you look at the distribution of these input variables, you can see that almost all samples have values very close to 0. Even though the relationship to the corresponding output variables is very simple, most deep learning-based models will likely fail to extrapolate when given a test sample with an unseen value much larger than the majority of the train samples, and predict something that is completely off. I think this is most likely the reason why many get very bad R2 scores for these target variables.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14891093%2Fc094b6dce557dc4874dd58c61866801a%2F02.png?generation=1715627035383583&amp;alt=media\"></p>\n<p>(I created these plots using the last 1 million train samples)</p>",
      "rawMarkdown": "First of all, thank you @jano123 for sharing this insight!\n\nThe relationship between these q0002 variables is indeed very simple:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14891093%2Fc136ac66f62af4750204b4d2a746377e%2F01.png?generation=1715626817753826&alt=media)\n\nHowever, if you look at the distribution of these input variables, you can see that almost all samples have values very close to 0. Even though the relationship to the corresponding output variables is very simple, most deep learning-based models will likely fail to extrapolate when given a test sample with an unseen value much larger than the majority of the train samples, and predict something that is completely off. I think this is most likely the reason why many get very bad R2 scores for these target variables.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14891093%2Fc094b6dce557dc4874dd58c61866801a%2F02.png?generation=1715627035383583&alt=media)\n\n(I created these plots using the last 1 million train samples)",
      "votes": 7
    },
    {
      "id": 2840480,
      "postDate": "2024-05-28T05:44:11.370Z",
      "content": "<p>This is such an amazing discussion! It's been a pleasure reading your analyses. 😃</p>",
      "rawMarkdown": "This is such an amazing discussion! It's been a pleasure reading your analyses. 😃",
      "votes": 1
    },
    {
      "id": 2814989,
      "postDate": "2024-05-15T16:26:43.047Z",
      "content": "<p>This is interesting.!</p>",
      "rawMarkdown": "This is interesting.!"
    },
    {
      "id": 2814457,
      "postDate": "2024-05-15T10:18:41.687Z",
      "content": "<p>This trick might not be useful on the test set.</p>",
      "rawMarkdown": "This trick might not be useful on the test set.",
      "replies": [
        {
          "id": 2814691,
          "postDate": "2024-05-15T13:57:57.373Z",
          "content": "<p>We know it is useful on the public part of the test dataset</p>",
          "rawMarkdown": "We know it is useful on the public part of the test dataset",
          "replies": [
            {
              "id": 2837120,
              "postDate": "2024-05-26T09:31:24.307Z",
              "content": "<p>Thank you, but I noticed that some people only used this trick for the first 26 indexes, you used this trick for all 60 indexes？<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11486426%2Fb81e53073d9e1577e7b9f80b8440702f%2F2024-05-26%20171833.png?generation=1716715861816528&amp;alt=media\"></p>",
              "rawMarkdown": "Thank you, but I noticed that some people only used this trick for the first 26 indexes, you used this trick for all 60 indexes？\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11486426%2Fb81e53073d9e1577e7b9f80b8440702f%2F2024-05-26%20171833.png?generation=1716715861816528&alt=media)"
            },
            {
              "id": 2842739,
              "postDate": "2024-05-29T08:37:04.190Z",
              "content": "<p>That formula works only for up to 26, but the first 14 are just zeros.</p>",
              "rawMarkdown": "That formula works only for up to 26, but the first 14 are just zeros."
            }
          ]
        }
      ]
    },
    {
      "id": 2814907,
      "postDate": "2024-05-15T15:54:30.570Z",
      "content": "<p>This is interesting.</p>\n<p>What I noticed is that, when calculating R2 for each of the 60 columns (after assuming y_pred = - x/1200), I get \"divide by zero\" for the first 12 (as expected since all y_true values are 0), then 1 for then next 16 columns, and then even negative values for the rest of the columns.</p>\n<p>This is done using data of about 1 million rows (by extracting every 10th row from train.csv).</p>\n<p>(edit: by applying this to the submission file for those columns where R2 is 1, the LB indeed increased on the order of ~0.03)</p>\n<p>edit2: As to the why doesn't NN peak this up. I just checked  \"ptend_q0002_26\" for example. Its minimum values is about -1e-10, max value 0, mean value -1e-16, median value -1e-35, standard deviation 1e-13, median absolute deviation 1e-35. <br>\nSo, if you normalize these target data by dividing with standard deviation, most of the values will still be very closely concentrated towards single value, and I guess NN doesn't peak up this slight difference. If you normalize these values by dividing with median absolute deviation, then it might.<br>\nI also think by default, for example, Keras NN uses float32 which has max exponent of 38. In some cases even this might make NN not able to differentiate target values.</p>",
      "rawMarkdown": "This is interesting.\n\nWhat I noticed is that, when calculating R2 for each of the 60 columns (after assuming y_pred = - x/1200), I get \"divide by zero\" for the first 12 (as expected since all y_true values are 0), then 1 for then next 16 columns, and then even negative values for the rest of the columns.\n\nThis is done using data of about 1 million rows (by extracting every 10th row from train.csv).\n\n(edit: by applying this to the submission file for those columns where R2 is 1, the LB indeed increased on the order of ~0.03)\n\nedit2: As to the why doesn't NN peak this up. I just checked  \"ptend_q0002_26\" for example. Its minimum values is about -1e-10, max value 0, mean value -1e-16, median value -1e-35, standard deviation 1e-13, median absolute deviation 1e-35. \nSo, if you normalize these target data by dividing with standard deviation, most of the values will still be very closely concentrated towards single value, and I guess NN doesn't peak up this slight difference. If you normalize these values by dividing with median absolute deviation, then it might.\nI also think by default, for example, Keras NN uses float32 which has max exponent of 38. In some cases even this might make NN not able to differentiate target values.",
      "isDeleted": true
    },
    {
      "id": 2814837,
      "postDate": "2024-05-15T15:23:23Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2811761,
      "postDate": "2024-05-13T21:58:44.210Z",
      "content": "<p>Wow I totally missed it! Thanks</p>",
      "rawMarkdown": "Wow I totally missed it! Thanks",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 2811881,
      "author_name": "Youri Matiounine",
      "author_url": "",
      "post_date": "2024-05-14T01:26:29.360000",
      "content": "<p>if you look at the code preparing the training data, you will see that all the vector targets are calculated as y[t] = (x[t+1]-x[t]) / 1200 (samples are 20 minutes=1200 seconds apart).</p>\n<p>x[t] is given to us, x[t+1] is not. For columns where large proportion of values are zeros, assuming that x[i+1]=0 will be correct most of the time. Then y[t]= (0-x[t]) / 1200 = -x[t] / 1200 is the correct answer.</p>\n<p>This only works for the few columns where most values are 0.</p>",
      "votes": 11,
      "replies": [
        {
          "id": 2812239,
          "author_name": "at7459",
          "author_url": "",
          "post_date": "2024-05-14T06:49:21.553000",
          "content": "<p>but take for example q0002_25 it seems to be mostly non zero, yet this relation still holds and you get a score of 1 with the relation over the entire train set. it seems more like they made an approximation here, maybe because these columns might be very negligible to the overall simulation or something, and for some reason they decide it starts becoming appreciable at around q0002_27 again</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2813590,
              "author_name": "Youri Matiounine",
              "author_url": "",
              "post_date": "2024-05-14T21:14:33.450000",
              "content": "<p>On further reading i see that targets are defined as (Tafter − Tbefore )/∆t, where variables are recorded before versus after the convection and radiation calculations.</p>\n<p>Looks like for situations with low prevalence of clouds they just set Tafter to 0, instead of setting it to Tbefore. </p>\n<p>So not very clean data. As a result a large portion of score improvement may be coming from hacking the data - similar to what you discovered in \"Home Credit\" competition.</p>\n<p>Perhaps organizers should drop the questionable columns from objective - then the results would better serve their purpose.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2814181,
              "author_name": "at7459",
              "author_url": "",
              "post_date": "2024-05-15T07:49:37.583000",
              "content": "<p>could be a million reasons, they're basically saying \"if cloud then cloud decay exponentially\", maybe there is a very small variable dependent on many things, which just gets rounded to 0 high up in the atmosphere in the low resolution data (of course it's probably very reasonable to assume that any cloud there is gone in 1200s)…</p>\n<p>who knows</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2811598,
      "author_name": "Felix Yang",
      "author_url": "",
      "post_date": "2024-05-13T19:14:22.027000",
      "content": "<p>First of all, thank you <a href=\"https://www.kaggle.com/jano123\" target=\"_blank\">@jano123</a> for sharing this insight!</p>\n<p>The relationship between these q0002 variables is indeed very simple:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14891093%2Fc136ac66f62af4750204b4d2a746377e%2F01.png?generation=1715626817753826&amp;alt=media\"></p>\n<p>However, if you look at the distribution of these input variables, you can see that almost all samples have values very close to 0. Even though the relationship to the corresponding output variables is very simple, most deep learning-based models will likely fail to extrapolate when given a test sample with an unseen value much larger than the majority of the train samples, and predict something that is completely off. I think this is most likely the reason why many get very bad R2 scores for these target variables.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14891093%2Fc094b6dce557dc4874dd58c61866801a%2F02.png?generation=1715627035383583&amp;alt=media\"></p>\n<p>(I created these plots using the last 1 million train samples)</p>",
      "votes": 7,
      "replies": []
    },
    {
      "id": 2840480,
      "author_name": "Jerry Lin",
      "author_url": "",
      "post_date": "2024-05-28T05:44:11.370000",
      "content": "<p>This is such an amazing discussion! It's been a pleasure reading your analyses. 😃</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2814989,
      "author_name": "Sheema Zain",
      "author_url": "",
      "post_date": "2024-05-15T16:26:43.047000",
      "content": "<p>This is interesting.!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2814457,
      "author_name": "Zhuoqun Li",
      "author_url": "",
      "post_date": "2024-05-15T10:18:41.687000",
      "content": "<p>This trick might not be useful on the test set.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2814691,
          "author_name": "DennisSakva",
          "author_url": "",
          "post_date": "2024-05-15T13:57:57.373000",
          "content": "<p>We know it is useful on the public part of the test dataset</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2837120,
              "author_name": "Zhuoqun Li",
              "author_url": "",
              "post_date": "2024-05-26T09:31:24.307000",
              "content": "<p>Thank you, but I noticed that some people only used this trick for the first 26 indexes, you used this trick for all 60 indexes？<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11486426%2Fb81e53073d9e1577e7b9f80b8440702f%2F2024-05-26%20171833.png?generation=1716715861816528&amp;alt=media\"></p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2842739,
              "author_name": "DennisSakva",
              "author_url": "",
              "post_date": "2024-05-29T08:37:04.190000",
              "content": "<p>That formula works only for up to 26, but the first 14 are just zeros.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2814907,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-05-15T15:54:30.570000",
      "content": "<p>This is interesting.</p>\n<p>What I noticed is that, when calculating R2 for each of the 60 columns (after assuming y_pred = - x/1200), I get \"divide by zero\" for the first 12 (as expected since all y_true values are 0), then 1 for then next 16 columns, and then even negative values for the rest of the columns.</p>\n<p>This is done using data of about 1 million rows (by extracting every 10th row from train.csv).</p>\n<p>(edit: by applying this to the submission file for those columns where R2 is 1, the LB indeed increased on the order of ~0.03)</p>\n<p>edit2: As to the why doesn't NN peak this up. I just checked  \"ptend_q0002_26\" for example. Its minimum values is about -1e-10, max value 0, mean value -1e-16, median value -1e-35, standard deviation 1e-13, median absolute deviation 1e-35. <br>\nSo, if you normalize these target data by dividing with standard deviation, most of the values will still be very closely concentrated towards single value, and I guess NN doesn't peak up this slight difference. If you normalize these values by dividing with median absolute deviation, then it might.<br>\nI also think by default, for example, Keras NN uses float32 which has max exponent of 38. In some cases even this might make NN not able to differentiate target values.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2814837,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-05-15T15:23:23",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2811761,
      "author_name": "greySnow",
      "author_url": "",
      "post_date": "2024-05-13T21:58:44.210000",
      "content": "<p>Wow I totally missed it! Thanks</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2811230": "If you haven't seen it already there's a simple heuristic that yields +0.03 to your LB score.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F349155%2F503540951a5e148f54901fd643b8fd21%2Fptent_q2.PNG?generation=1715619199918582&alt=media)\nCourtesy of @jano123 (show him some love here https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/499896)\nBut my question is: why is my X million parameter model not picking this simple linear relationship up and just outputs garbage for these columns? What's going on here?",
    "2811881": "if you look at the code preparing the training data, you will see that all the vector targets are calculated as y[t] = (x[t+1]-x[t]) / 1200 (samples are 20 minutes=1200 seconds apart).\n\nx[t] is given to us, x[t+1] is not. For columns where large proportion of values are zeros, assuming that x[i+1]=0 will be correct most of the time. Then y[t]= (0-x[t]) / 1200 = -x[t] / 1200 is the correct answer.\n\nThis only works for the few columns where most values are 0.",
    "2811598": "First of all, thank you @jano123 for sharing this insight!\n\nThe relationship between these q0002 variables is indeed very simple:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14891093%2Fc136ac66f62af4750204b4d2a746377e%2F01.png?generation=1715626817753826&alt=media)\n\nHowever, if you look at the distribution of these input variables, you can see that almost all samples have values very close to 0. Even though the relationship to the corresponding output variables is very simple, most deep learning-based models will likely fail to extrapolate when given a test sample with an unseen value much larger than the majority of the train samples, and predict something that is completely off. I think this is most likely the reason why many get very bad R2 scores for these target variables.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14891093%2Fc094b6dce557dc4874dd58c61866801a%2F02.png?generation=1715627035383583&alt=media)\n\n(I created these plots using the last 1 million train samples)",
    "2840480": "This is such an amazing discussion! It's been a pleasure reading your analyses. 😃",
    "2814989": "This is interesting.!",
    "2814457": "This trick might not be useful on the test set.",
    "2814907": "This is interesting.\n\nWhat I noticed is that, when calculating R2 for each of the 60 columns (after assuming y_pred = - x/1200), I get \"divide by zero\" for the first 12 (as expected since all y_true values are 0), then 1 for then next 16 columns, and then even negative values for the rest of the columns.\n\nThis is done using data of about 1 million rows (by extracting every 10th row from train.csv).\n\n(edit: by applying this to the submission file for those columns where R2 is 1, the LB indeed increased on the order of ~0.03)\n\nedit2: As to the why doesn't NN peak this up. I just checked  \"ptend_q0002_26\" for example. Its minimum values is about -1e-10, max value 0, mean value -1e-16, median value -1e-35, standard deviation 1e-13, median absolute deviation 1e-35. \nSo, if you normalize these target data by dividing with standard deviation, most of the values will still be very closely concentrated towards single value, and I guess NN doesn't peak up this slight difference. If you normalize these values by dividing with median absolute deviation, then it might.\nI also think by default, for example, Keras NN uses float32 which has max exponent of 38. In some cases even this might make NN not able to differentiate target values.",
    "2814837": "",
    "2811761": "Wow I totally missed it! Thanks"
  }
}