{
  "id": 520424,
  "title": "Leaky Solutions",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/520424",
  "author_name": "yu4u",
  "post_date": "2024-07-16T00:38:13.675000",
  "votes": 8,
  "comment_count": 2,
  "views": 0,
  "content": "<p>We cannot release solutions that do not use leaks until the checking of the solutions from the prize-winning teams is completed. Although it was prohibited in this competition, I am also very interested in the technical aspects of the 2D to 2D or 2D to 1D approaches. Thus, in the meantime, let's share or discuss solutions that use leaks to pass the time.</p>\n<p>I understand that the basic idea of a solution using leaks is to use multiple columns of the test data as inputs during training and inference. The host focused on the use of multiple columns with the same timestamp, but I think that the use of data with timestamps before and after the same point is also effective.</p>\n<p>I have not submitted this leaky solutions, but I have trained a model that takes as input the data from 8 hours before and after the same point, as well as the average and std of the data from 4 surrounding points at the same time, in addition to the normal input. This gave me about 0.01 improvement in CV score.</p>\n<p>I would like to know if there is a more elegant approach to utilizing spatio-temporal neighborhood information.</p>",
  "messages": [
    {
      "id": 2923511,
      "postDate": "2024-07-16T00:38:13.677Z",
      "content": "<p>We cannot release solutions that do not use leaks until the checking of the solutions from the prize-winning teams is completed. Although it was prohibited in this competition, I am also very interested in the technical aspects of the 2D to 2D or 2D to 1D approaches. Thus, in the meantime, let's share or discuss solutions that use leaks to pass the time.</p>\n<p>I understand that the basic idea of a solution using leaks is to use multiple columns of the test data as inputs during training and inference. The host focused on the use of multiple columns with the same timestamp, but I think that the use of data with timestamps before and after the same point is also effective.</p>\n<p>I have not submitted this leaky solutions, but I have trained a model that takes as input the data from 8 hours before and after the same point, as well as the average and std of the data from 4 surrounding points at the same time, in addition to the normal input. This gave me about 0.01 improvement in CV score.</p>\n<p>I would like to know if there is a more elegant approach to utilizing spatio-temporal neighborhood information.</p>",
      "rawMarkdown": "We cannot release solutions that do not use leaks until the checking of the solutions from the prize-winning teams is completed. Although it was prohibited in this competition, I am also very interested in the technical aspects of the 2D to 2D or 2D to 1D approaches. Thus, in the meantime, let's share or discuss solutions that use leaks to pass the time.\n\nI understand that the basic idea of a solution using leaks is to use multiple columns of the test data as inputs during training and inference. The host focused on the use of multiple columns with the same timestamp, but I think that the use of data with timestamps before and after the same point is also effective.\n\nI have not submitted this leaky solutions, but I have trained a model that takes as input the data from 8 hours before and after the same point, as well as the average and std of the data from 4 surrounding points at the same time, in addition to the normal input. This gave me about 0.01 improvement in CV score.\n\nI would like to know if there is a more elegant approach to utilizing spatio-temporal neighborhood information.",
      "votes": 8
    },
    {
      "id": 2923581,
      "postDate": "2024-07-16T02:27:38.450Z",
      "content": "<p>I also tried adding “timestamps before and after” 2 points in the very beginning before host changed dataset. At that time, it give me 0.004 in both lb and cv.</p>",
      "rawMarkdown": "I also tried adding “timestamps before and after” 2 points in the very beginning before host changed dataset. At that time, it give me 0.004 in both lb and cv.",
      "votes": 1
    },
    {
      "id": 2923853,
      "postDate": "2024-07-16T07:18:32.147Z",
      "content": "<p>I have thought about using a stacking approach, inputting predictions from near time-steps and locations to improve the original prediction. I haven't implemented it but has anyone tried?</p>",
      "rawMarkdown": "I have thought about using a stacking approach, inputting predictions from near time-steps and locations to improve the original prediction. I haven't implemented it but has anyone tried?"
    }
  ],
  "comments": [
    {
      "id": 2923581,
      "author_name": "ADAM.",
      "author_url": "",
      "post_date": "2024-07-16T02:27:38.450000",
      "content": "<p>I also tried adding “timestamps before and after” 2 points in the very beginning before host changed dataset. At that time, it give me 0.004 in both lb and cv.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2923853,
      "author_name": "Fritz Cremer",
      "author_url": "",
      "post_date": "2024-07-16T07:18:32.147000",
      "content": "<p>I have thought about using a stacking approach, inputting predictions from near time-steps and locations to improve the original prediction. I haven't implemented it but has anyone tried?</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2923511": "We cannot release solutions that do not use leaks until the checking of the solutions from the prize-winning teams is completed. Although it was prohibited in this competition, I am also very interested in the technical aspects of the 2D to 2D or 2D to 1D approaches. Thus, in the meantime, let's share or discuss solutions that use leaks to pass the time.\n\nI understand that the basic idea of a solution using leaks is to use multiple columns of the test data as inputs during training and inference. The host focused on the use of multiple columns with the same timestamp, but I think that the use of data with timestamps before and after the same point is also effective.\n\nI have not submitted this leaky solutions, but I have trained a model that takes as input the data from 8 hours before and after the same point, as well as the average and std of the data from 4 surrounding points at the same time, in addition to the normal input. This gave me about 0.01 improvement in CV score.\n\nI would like to know if there is a more elegant approach to utilizing spatio-temporal neighborhood information.",
    "2923581": "I also tried adding “timestamps before and after” 2 points in the very beginning before host changed dataset. At that time, it give me 0.004 in both lb and cv.",
    "2923853": "I have thought about using a stacking approach, inputting predictions from near time-steps and locations to improve the original prediction. I haven't implemented it but has anyone tried?"
  }
}