{
  "id": 511274,
  "title": "Is Tropical Cylones Key to Win this Competition?",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/511274",
  "author_name": "Bilzard",
  "post_date": "2024-06-10T05:55:36.932000",
  "votes": 26,
  "comment_count": 25,
  "views": 0,
  "content": "<p>Currently, I am searching the path to LB&gt;0.79~0.80, but not found one(s) yet.</p>\n<p>However, I noticed some clue: <em>tropical cyclones</em>.</p>\n<p>The picture below is MSE of one of my model mapped to location.<br>\nThe red area is where MSE is high, which means difficult to predict.<br>\nThe place where my model struggles is biased on Pacific ocean, and April to September.<br>\nSo as you know, the place where tropical cyclones tend to resolve.</p>\n<p>I couldn't find any clever method for utilize this information, so I share this hoping someone will and making this competition more competitive.</p>\n<p>Good Luck.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F60f105df991d392a8546fc3b496dfb10%2Fmse_location.jpeg?generation=1717998482104168&amp;alt=media\"></p>",
  "messages": [
    {
      "id": 2864354,
      "postDate": "2024-06-10T05:55:36.933Z",
      "content": "<p>Currently, I am searching the path to LB&gt;0.79~0.80, but not found one(s) yet.</p>\n<p>However, I noticed some clue: <em>tropical cyclones</em>.</p>\n<p>The picture below is MSE of one of my model mapped to location.<br>\nThe red area is where MSE is high, which means difficult to predict.<br>\nThe place where my model struggles is biased on Pacific ocean, and April to September.<br>\nSo as you know, the place where tropical cyclones tend to resolve.</p>\n<p>I couldn't find any clever method for utilize this information, so I share this hoping someone will and making this competition more competitive.</p>\n<p>Good Luck.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F60f105df991d392a8546fc3b496dfb10%2Fmse_location.jpeg?generation=1717998482104168&amp;alt=media\"></p>",
      "rawMarkdown": "Currently, I am searching the path to LB>0.79~0.80, but not found one(s) yet.\n\nHowever, I noticed some clue: *tropical cyclones*.\n\nThe picture below is MSE of one of my model mapped to location.\nThe red area is where MSE is high, which means difficult to predict.\nThe place where my model struggles is biased on Pacific ocean, and April to September.\nSo as you know, the place where tropical cyclones tend to resolve.\n\nI couldn't find any clever method for utilize this information, so I share this hoping someone will and making this competition more competitive.\n\nGood Luck.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F60f105df991d392a8546fc3b496dfb10%2Fmse_location.jpeg?generation=1717998482104168&alt=media)",
      "votes": 26
    },
    {
      "id": 2869171,
      "postDate": "2024-06-12T22:09:14.923Z",
      "content": "<p>The phenomenon that you are likely encoutering  is something called the Madden-Julian Oscillation. As the name suggests, it has semi-periodic behavior. Hope this helps!</p>",
      "rawMarkdown": "The phenomenon that you are likely encoutering  is something called the Madden-Julian Oscillation. As the name suggests, it has semi-periodic behavior. Hope this helps!",
      "votes": 3,
      "replies": [
        {
          "id": 2869174,
          "postDate": "2024-06-12T22:28:38.960Z",
          "content": "<p>Thank you for sharing domain knowledge. It surely will help.</p>",
          "rawMarkdown": "Thank you for sharing domain knowledge. It surely will help.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2872749,
      "postDate": "2024-06-15T03:56:25.340Z",
      "content": "<blockquote>\n  <p>The place where my model struggles is biased on Pacific ocean, and April to September.</p>\n</blockquote>\n<p>It might be a silly question: What does \"April to September\" mean? Can we get the time information for each sample?</p>",
      "rawMarkdown": ">The place where my model struggles is biased on Pacific ocean, and April to September.\n\nIt might be a silly question: What does \"April to September\" mean? Can we get the time information for each sample?",
      "votes": 1,
      "replies": [
        {
          "id": 2872950,
          "postDate": "2024-06-15T07:32:50.823Z",
          "content": "<p>Yes we can, by using the full data set from Hugging Face</p>",
          "rawMarkdown": "Yes we can, by using the full data set from Hugging Face",
          "votes": 1,
          "replies": [
            {
              "id": 2873080,
              "postDate": "2024-06-15T10:05:04.313Z",
              "content": "<p>Yes, but then you can only get the timestamps for the training dataset, right? So this information would be useful for analysis of your model, gaining insights into what would be missing in the local validation set… but in the end, you are not using timestamp information for the test set. Or am I wrong?</p>",
              "rawMarkdown": "Yes, but then you can only get the timestamps for the training dataset, right? So this information would be useful for analysis of your model, gaining insights into what would be missing in the local validation set... but in the end, you are not using timestamp information for the test set. Or am I wrong?",
              "votes": 1
            },
            {
              "id": 2873124,
              "postDate": "2024-06-15T10:30:47.720Z",
              "content": "<p>Well…yes and no. We can get exact timestamp only for train set, but by looking on moving averages of test columns we can get approximate time stamp for test too. Since there is a temporal behaviour of the moving average and the test is not shuffled, only downsampled. </p>",
              "rawMarkdown": "Well...yes and no. We can get exact timestamp only for train set, but by looking on moving averages of test columns we can get approximate time stamp for test too. Since there is a temporal behaviour of the moving average and the test is not shuffled, only downsampled. ",
              "votes": 1
            },
            {
              "id": 2875522,
              "postDate": "2024-06-17T08:02:07.813Z",
              "content": "<p>Did you use all of huggingface's data? How much does using all huggingface data improve your cv score?</p>",
              "rawMarkdown": "Did you use all of huggingface's data? How much does using all huggingface data improve your cv score?"
            }
          ]
        }
      ]
    },
    {
      "id": 2864552,
      "postDate": "2024-06-10T07:51:25.657Z",
      "content": "<p>Are you also inserting the location information as input to the model somehow?</p>",
      "rawMarkdown": "Are you also inserting the location information as input to the model somehow?",
      "votes": 1,
      "replies": [
        {
          "id": 2864561,
          "postDate": "2024-06-10T07:54:59.723Z",
          "content": "<p>No. However, it’s worth trying (pseudo labeling etc.)</p>",
          "rawMarkdown": "No. However, it’s worth trying (pseudo labeling etc.)",
          "votes": 1
        }
      ]
    },
    {
      "id": 2864510,
      "postDate": "2024-06-10T07:13:37.580Z",
      "content": "<p>Nice find, thanks for sharing. I suppose this is where domain expertise comes in.</p>",
      "rawMarkdown": "Nice find, thanks for sharing. I suppose this is where domain expertise comes in.",
      "votes": 1,
      "replies": [
        {
          "id": 2864547,
          "postDate": "2024-06-10T07:47:50.120Z",
          "content": "<p>Maybe. And 1152km x 1152km grid size of low resolution seems pretty challenging to precisely describe physical phenomena with ML models.</p>",
          "rawMarkdown": "Maybe. And 1152km x 1152km grid size of low resolution seems pretty challenging to precisely describe physical phenomena with ML models.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2864367,
      "postDate": "2024-06-10T06:07:40.767Z",
      "content": "<p>It's a interesting finding. Could you share the code for plotting this image ?</p>",
      "rawMarkdown": "It's a interesting finding. Could you share the code for plotting this image ?",
      "votes": 2,
      "replies": [
        {
          "id": 2864448,
          "postDate": "2024-06-10T06:40:23.290Z",
          "content": "<p>I omit the technical detail, but we can obtain location information from sample_id, and latitude and longitude are available from original ClimSim dataset.</p>\n<ul>\n<li>location_id = sample_id % 384</li>\n</ul>\n<p>Maybe their demo notebook is helpful <br>\n<a href=\"https://github.com/leap-stc/ClimSim/blob/main/demo_notebooks/water_conservation.ipynb\" target=\"_blank\">https://github.com/leap-stc/ClimSim/blob/main/demo_notebooks/water_conservation.ipynb</a></p>",
          "rawMarkdown": "I omit the technical detail, but we can obtain location information from sample_id, and latitude and longitude are available from original ClimSim dataset.\n\n- location_id = sample_id % 384\n\nMaybe their demo notebook is helpful \nhttps://github.com/leap-stc/ClimSim/blob/main/demo_notebooks/water_conservation.ipynb",
          "votes": 10,
          "replies": [
            {
              "id": 2866136,
              "postDate": "2024-06-11T06:47:55.693Z",
              "content": "<p>Hi, I am wondering whether <code>location_id = sample_id % 384</code> has the same meaning for test.csv.</p>",
              "rawMarkdown": "Hi, I am wondering whether `location_id = sample_id % 384` has the same meaning for test.csv."
            },
            {
              "id": 2866179,
              "postDate": "2024-06-11T07:15:19.323Z",
              "content": "<p><a href=\"https://www.kaggle.com/sijunxu\" target=\"_blank\">@sijunxu</a> I guess no, but you can examine with comparing statistic feature from train and test.</p>",
              "rawMarkdown": "@sijunxu I guess no, but you can examine with comparing statistic feature from train and test."
            },
            {
              "id": 2866319,
              "postDate": "2024-06-11T08:53:48.923Z",
              "content": "<p>Maybe, but probably not since 625000 % 384 != 0. While for train set 10091520 % 384 = 0</p>",
              "rawMarkdown": "Maybe, but probably not since 625000 % 384 != 0. While for train set 10091520 % 384 = 0",
              "votes": 3
            },
            {
              "id": 2867718,
              "postDate": "2024-06-12T03:52:34.050Z",
              "content": "<p>I scrambled the sample IDs for the test set. 🙃</p>",
              "rawMarkdown": "I scrambled the sample IDs for the test set. 🙃",
              "votes": 6
            },
            {
              "id": 2867762,
              "postDate": "2024-06-12T04:55:46.483Z",
              "content": "<p><a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> Yeah, it should be. Thank you for confirmation.</p>",
              "rawMarkdown": "@jerrylin96 Yeah, it should be. Thank you for confirmation."
            },
            {
              "id": 2868420,
              "postDate": "2024-06-12T11:16:19.167Z",
              "content": "<p>but you left \"sample_id\" column unchanged, so test data can be sorted back to its original order. From that label it is easy to deduce location and timestamp of each test sample.</p>",
              "rawMarkdown": "but you left \"sample_id\" column unchanged, so test data can be sorted back to its original order. From that label it is easy to deduce location and timestamp of each test sample.",
              "votes": 10
            },
            {
              "id": 2868459,
              "postDate": "2024-06-12T11:39:37.183Z",
              "content": "<p>Hahaha good catch <a href=\"https://www.kaggle.com/ymatioun\" target=\"_blank\">@ymatioun</a> I totally missed it. I guess timestamp and location are back on the table then. Unless the labels themselves are from a scrambled set too.</p>",
              "rawMarkdown": "Hahaha good catch @ymatioun I totally missed it. I guess timestamp and location are back on the table then. Unless the labels themselves are from a scrambled set too."
            },
            {
              "id": 2868466,
              "postDate": "2024-06-12T11:45:25.830Z",
              "content": "<p>I found in testset <code>max(sample_id) + 1 % 384=0</code>, but im not sure how the testset are scrambled.</p>",
              "rawMarkdown": "I found in testset `max(sample_id) + 1 % 384=0`, but im not sure how the testset are scrambled.",
              "votes": 1
            },
            {
              "id": 2868486,
              "postDate": "2024-06-12T12:08:52.060Z",
              "content": "<blockquote>\n  <p>but you left \"sample_id\" column unchanged, so test data can be sorted back to its original order. From that label it is easy to deduce location and timestamp of each test sample.</p>\n</blockquote>\n<p>I was hoping I won't have to dive into this, but it seems that there is no choice. Thank you for sharing, big mistake from the host side if true.</p>",
              "rawMarkdown": "> but you left \"sample_id\" column unchanged, so test data can be sorted back to its original order. From that label it is easy to deduce location and timestamp of each test sample.\n\nI was hoping I won't have to dive into this, but it seems that there is no choice. Thank you for sharing, big mistake from the host side if true.",
              "votes": 2
            },
            {
              "id": 2868487,
              "postDate": "2024-06-12T12:09:16.597Z",
              "content": "<p>Hmm.., in my understanding, \"scrambled\" means not just shuffle sample ids, but after re-ordering, they are totally random. However, it should be checked…</p>",
              "rawMarkdown": "Hmm.., in my understanding, \"scrambled\" means not just shuffle sample ids, but after re-ordering, they are totally random. However, it should be checked...",
              "votes": 1
            },
            {
              "id": 2868521,
              "postDate": "2024-06-12T12:40:30.953Z",
              "content": "<p><a href=\"https://www.kaggle.com/martynoveduard\" target=\"_blank\">@martynoveduard</a> same here…</p>",
              "rawMarkdown": "@martynoveduard same here..."
            }
          ]
        }
      ]
    },
    {
      "id": 2870726,
      "postDate": "2024-06-13T18:34:59.423Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2869171,
      "author_name": "Anthony Meza",
      "author_url": "",
      "post_date": "2024-06-12T22:09:14.923000",
      "content": "<p>The phenomenon that you are likely encoutering  is something called the Madden-Julian Oscillation. As the name suggests, it has semi-periodic behavior. Hope this helps!</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2869174,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2024-06-12T22:28:38.960000",
          "content": "<p>Thank you for sharing domain knowledge. It surely will help.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2872749,
      "author_name": "Joseph Zhou",
      "author_url": "",
      "post_date": "2024-06-15T03:56:25.340000",
      "content": "<blockquote>\n  <p>The place where my model struggles is biased on Pacific ocean, and April to September.</p>\n</blockquote>\n<p>It might be a silly question: What does \"April to September\" mean? Can we get the time information for each sample?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2872950,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2024-06-15T07:32:50.823000",
          "content": "<p>Yes we can, by using the full data set from Hugging Face</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2873080,
              "author_name": "Federico Peccia",
              "author_url": "",
              "post_date": "2024-06-15T10:05:04.313000",
              "content": "<p>Yes, but then you can only get the timestamps for the training dataset, right? So this information would be useful for analysis of your model, gaining insights into what would be missing in the local validation set… but in the end, you are not using timestamp information for the test set. Or am I wrong?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2873124,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-06-15T10:30:47.720000",
              "content": "<p>Well…yes and no. We can get exact timestamp only for train set, but by looking on moving averages of test columns we can get approximate time stamp for test too. Since there is a temporal behaviour of the moving average and the test is not shuffled, only downsampled. </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2875522,
              "author_name": "Zhuoqun Li",
              "author_url": "",
              "post_date": "2024-06-17T08:02:07.813000",
              "content": "<p>Did you use all of huggingface's data? How much does using all huggingface data improve your cv score?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2864552,
      "author_name": "Federico Peccia",
      "author_url": "",
      "post_date": "2024-06-10T07:51:25.657000",
      "content": "<p>Are you also inserting the location information as input to the model somehow?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2864561,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2024-06-10T07:54:59.723000",
          "content": "<p>No. However, it’s worth trying (pseudo labeling etc.)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2864510,
      "author_name": "sroger",
      "author_url": "",
      "post_date": "2024-06-10T07:13:37.580000",
      "content": "<p>Nice find, thanks for sharing. I suppose this is where domain expertise comes in.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2864547,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2024-06-10T07:47:50.120000",
          "content": "<p>Maybe. And 1152km x 1152km grid size of low resolution seems pretty challenging to precisely describe physical phenomena with ML models.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2864367,
      "author_name": "yuanzhe zhou",
      "author_url": "",
      "post_date": "2024-06-10T06:07:40.767000",
      "content": "<p>It's a interesting finding. Could you share the code for plotting this image ?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2864448,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2024-06-10T06:40:23.290000",
          "content": "<p>I omit the technical detail, but we can obtain location information from sample_id, and latitude and longitude are available from original ClimSim dataset.</p>\n<ul>\n<li>location_id = sample_id % 384</li>\n</ul>\n<p>Maybe their demo notebook is helpful <br>\n<a href=\"https://github.com/leap-stc/ClimSim/blob/main/demo_notebooks/water_conservation.ipynb\" target=\"_blank\">https://github.com/leap-stc/ClimSim/blob/main/demo_notebooks/water_conservation.ipynb</a></p>",
          "votes": 10,
          "replies": [
            {
              "id": 2866136,
              "author_name": "Sijun Xu",
              "author_url": "",
              "post_date": "2024-06-11T06:47:55.693000",
              "content": "<p>Hi, I am wondering whether <code>location_id = sample_id % 384</code> has the same meaning for test.csv.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2866179,
              "author_name": "Bilzard",
              "author_url": "",
              "post_date": "2024-06-11T07:15:19.323000",
              "content": "<p><a href=\"https://www.kaggle.com/sijunxu\" target=\"_blank\">@sijunxu</a> I guess no, but you can examine with comparing statistic feature from train and test.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2866319,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-06-11T08:53:48.923000",
              "content": "<p>Maybe, but probably not since 625000 % 384 != 0. While for train set 10091520 % 384 = 0</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2867718,
              "author_name": "Jerry Lin",
              "author_url": "",
              "post_date": "2024-06-12T03:52:34.050000",
              "content": "<p>I scrambled the sample IDs for the test set. 🙃</p>",
              "votes": 6,
              "replies": []
            },
            {
              "id": 2867762,
              "author_name": "Bilzard",
              "author_url": "",
              "post_date": "2024-06-12T04:55:46.483000",
              "content": "<p><a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> Yeah, it should be. Thank you for confirmation.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2868420,
              "author_name": "Youri Matiounine",
              "author_url": "",
              "post_date": "2024-06-12T11:16:19.167000",
              "content": "<p>but you left \"sample_id\" column unchanged, so test data can be sorted back to its original order. From that label it is easy to deduce location and timestamp of each test sample.</p>",
              "votes": 10,
              "replies": []
            },
            {
              "id": 2868459,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-06-12T11:39:37.183000",
              "content": "<p>Hahaha good catch <a href=\"https://www.kaggle.com/ymatioun\" target=\"_blank\">@ymatioun</a> I totally missed it. I guess timestamp and location are back on the table then. Unless the labels themselves are from a scrambled set too.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2868466,
              "author_name": "Sijun Xu",
              "author_url": "",
              "post_date": "2024-06-12T11:45:25.830000",
              "content": "<p>I found in testset <code>max(sample_id) + 1 % 384=0</code>, but im not sure how the testset are scrambled.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2868486,
              "author_name": "slime",
              "author_url": "",
              "post_date": "2024-06-12T12:08:52.060000",
              "content": "<blockquote>\n  <p>but you left \"sample_id\" column unchanged, so test data can be sorted back to its original order. From that label it is easy to deduce location and timestamp of each test sample.</p>\n</blockquote>\n<p>I was hoping I won't have to dive into this, but it seems that there is no choice. Thank you for sharing, big mistake from the host side if true.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2868487,
              "author_name": "Bilzard",
              "author_url": "",
              "post_date": "2024-06-12T12:09:16.597000",
              "content": "<p>Hmm.., in my understanding, \"scrambled\" means not just shuffle sample ids, but after re-ordering, they are totally random. However, it should be checked…</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2868521,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-06-12T12:40:30.953000",
              "content": "<p><a href=\"https://www.kaggle.com/martynoveduard\" target=\"_blank\">@martynoveduard</a> same here…</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2870726,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-06-13T18:34:59.423000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2864354": "Currently, I am searching the path to LB>0.79~0.80, but not found one(s) yet.\n\nHowever, I noticed some clue: *tropical cyclones*.\n\nThe picture below is MSE of one of my model mapped to location.\nThe red area is where MSE is high, which means difficult to predict.\nThe place where my model struggles is biased on Pacific ocean, and April to September.\nSo as you know, the place where tropical cyclones tend to resolve.\n\nI couldn't find any clever method for utilize this information, so I share this hoping someone will and making this competition more competitive.\n\nGood Luck.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F60f105df991d392a8546fc3b496dfb10%2Fmse_location.jpeg?generation=1717998482104168&alt=media)",
    "2869171": "The phenomenon that you are likely encoutering  is something called the Madden-Julian Oscillation. As the name suggests, it has semi-periodic behavior. Hope this helps!",
    "2872749": ">The place where my model struggles is biased on Pacific ocean, and April to September.\n\nIt might be a silly question: What does \"April to September\" mean? Can we get the time information for each sample?",
    "2864552": "Are you also inserting the location information as input to the model somehow?",
    "2864510": "Nice find, thanks for sharing. I suppose this is where domain expertise comes in.",
    "2864367": "It's a interesting finding. Could you share the code for plotting this image ?",
    "2870726": ""
  }
}