{
  "id": 501938,
  "title": "Drift phenomenon in the distribution of training and test dataset",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/501938",
  "author_name": "Zhuoqun Li",
  "post_date": "2024-05-11T12:10:17.218000",
  "votes": 4,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I noticed significant differences in the probability density distribution of certain features between the training and test sets, such as the 'state_q0001_0' feature below. How to handle these features? Should they be deleted, retained, or subjected to some transformation?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11486426%2F0bff91430df41cf60a809249b8911bd0%2Fstate_q0001_0.png?generation=1715429399685794&amp;alt=media\"></p>",
  "messages": [
    {
      "id": 2807515,
      "postDate": "2024-05-11T17:46:16.477Z",
      "content": "<p>'should they'- no one can answer this before experimenting and finding out, which is what I suggest you do (and I will do myself probably at some point)<br>\nHowever, as long as the test values are included in the train distribution (which seems to be the case here), I expect the models to generalize well. The problems usually start to arise when the models need to generalize to never-seen-before values.</p>",
      "rawMarkdown": "'should they'- no one can answer this before experimenting and finding out, which is what I suggest you do (and I will do myself probably at some point)\nHowever, as long as the test values are included in the train distribution (which seems to be the case here), I expect the models to generalize well. The problems usually start to arise when the models need to generalize to never-seen-before values.",
      "votes": 3
    },
    {
      "id": 2806942,
      "postDate": "2024-05-11T12:10:17.220Z",
      "content": "<p>I noticed significant differences in the probability density distribution of certain features between the training and test sets, such as the 'state_q0001_0' feature below. How to handle these features? Should they be deleted, retained, or subjected to some transformation?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11486426%2F0bff91430df41cf60a809249b8911bd0%2Fstate_q0001_0.png?generation=1715429399685794&amp;alt=media\"></p>",
      "rawMarkdown": "I noticed significant differences in the probability density distribution of certain features between the training and test sets, such as the 'state_q0001_0' feature below. How to handle these features? Should they be deleted, retained, or subjected to some transformation?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11486426%2F0bff91430df41cf60a809249b8911bd0%2Fstate_q0001_0.png?generation=1715429399685794&alt=media)",
      "votes": 3
    }
  ],
  "comments": [
    {
      "id": 2807515,
      "author_name": "greySnow",
      "author_url": "",
      "post_date": "2024-05-11T17:46:16.477000",
      "content": "<p>'should they'- no one can answer this before experimenting and finding out, which is what I suggest you do (and I will do myself probably at some point)<br>\nHowever, as long as the test values are included in the train distribution (which seems to be the case here), I expect the models to generalize well. The problems usually start to arise when the models need to generalize to never-seen-before values.</p>",
      "votes": 3,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2807515": "'should they'- no one can answer this before experimenting and finding out, which is what I suggest you do (and I will do myself probably at some point)\nHowever, as long as the test values are included in the train distribution (which seems to be the case here), I expect the models to generalize well. The problems usually start to arise when the models need to generalize to never-seen-before values.",
    "2806942": "I noticed significant differences in the probability density distribution of certain features between the training and test sets, such as the 'state_q0001_0' feature below. How to handle these features? Should they be deleted, retained, or subjected to some transformation?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11486426%2F0bff91430df41cf60a809249b8911bd0%2Fstate_q0001_0.png?generation=1715429399685794&alt=media)"
  }
}