{
  "id": 506427,
  "title": "Big distribution differences between features in train and test set",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/506427",
  "author_name": "Federico Peccia",
  "post_date": "2024-05-21T20:41:42.655000",
  "votes": 11,
  "comment_count": 2,
  "views": 0,
  "content": "<p>As shown in <a href=\"https://www.kaggle.com/code/docxian/leap-climsim-visual-eda#Targets-and-Features\" target=\"_blank\">this notebook</a>, there are features with extremely different distributions when comparing the train and test sets. For example, features state_q0001_[0-12] are completely different, and the same happens for features state_q0002_[0-29] and state_q0003_[0-15]. Additionally, some pbuf_ozone features have pretty different distributions.</p>\n<p>This is an example taken from the notebook, for features of the state_q0001 family:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1583716%2F93c8aeef41ac1205ac48acfe7c8dbef5%2FCaptura%20de%20pantalla%202024-05-21%20224049.png?generation=1716324083867018&amp;alt=media\"></p>\n<p>How are you managing these features in your models? Are you excluding them completely? Are the records with the missing values hidden in the original low-res dataset? Are these features scaled improperly in the test set, and is this why they are so different?</p>",
  "messages": [
    {
      "id": 2828054,
      "postDate": "2024-05-21T20:41:42.657Z",
      "content": "<p>As shown in <a href=\"https://www.kaggle.com/code/docxian/leap-climsim-visual-eda#Targets-and-Features\" target=\"_blank\">this notebook</a>, there are features with extremely different distributions when comparing the train and test sets. For example, features state_q0001_[0-12] are completely different, and the same happens for features state_q0002_[0-29] and state_q0003_[0-15]. Additionally, some pbuf_ozone features have pretty different distributions.</p>\n<p>This is an example taken from the notebook, for features of the state_q0001 family:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1583716%2F93c8aeef41ac1205ac48acfe7c8dbef5%2FCaptura%20de%20pantalla%202024-05-21%20224049.png?generation=1716324083867018&amp;alt=media\"></p>\n<p>How are you managing these features in your models? Are you excluding them completely? Are the records with the missing values hidden in the original low-res dataset? Are these features scaled improperly in the test set, and is this why they are so different?</p>",
      "rawMarkdown": "As shown in [this notebook](https://www.kaggle.com/code/docxian/leap-climsim-visual-eda#Targets-and-Features), there are features with extremely different distributions when comparing the train and test sets. For example, features state_q0001_[0-12] are completely different, and the same happens for features state_q0002_[0-29] and state_q0003_[0-15]. Additionally, some pbuf_ozone features have pretty different distributions.\n\nThis is an example taken from the notebook, for features of the state_q0001 family:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1583716%2F93c8aeef41ac1205ac48acfe7c8dbef5%2FCaptura%20de%20pantalla%202024-05-21%20224049.png?generation=1716324083867018&alt=media)\n\nHow are you managing these features in your models? Are you excluding them completely? Are the records with the missing values hidden in the original low-res dataset? Are these features scaled improperly in the test set, and is this why they are so different?",
      "votes": 11
    },
    {
      "id": 2837074,
      "postDate": "2024-05-26T09:02:05.720Z",
      "content": "<p>How to solve this problem?</p>",
      "rawMarkdown": "How to solve this problem?",
      "votes": 2
    },
    {
      "id": 2839880,
      "postDate": "2024-05-27T18:42:52.767Z",
      "content": "<p>Actually the tendency from 0-12 are not considered except temperature. I would not consider the state variables from the pressure levels 0-12 except temperature. I'll try this in the next days </p>",
      "rawMarkdown": "Actually the tendency from 0-12 are not considered except temperature. I would not consider the state variables from the pressure levels 0-12 except temperature. I'll try this in the next days "
    }
  ],
  "comments": [
    {
      "id": 2837074,
      "author_name": "Zhuoqun Li",
      "author_url": "",
      "post_date": "2024-05-26T09:02:05.720000",
      "content": "<p>How to solve this problem?</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2839880,
      "author_name": "HappyKiter",
      "author_url": "",
      "post_date": "2024-05-27T18:42:52.767000",
      "content": "<p>Actually the tendency from 0-12 are not considered except temperature. I would not consider the state variables from the pressure levels 0-12 except temperature. I'll try this in the next days </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2828054": "As shown in [this notebook](https://www.kaggle.com/code/docxian/leap-climsim-visual-eda#Targets-and-Features), there are features with extremely different distributions when comparing the train and test sets. For example, features state_q0001_[0-12] are completely different, and the same happens for features state_q0002_[0-29] and state_q0003_[0-15]. Additionally, some pbuf_ozone features have pretty different distributions.\n\nThis is an example taken from the notebook, for features of the state_q0001 family:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1583716%2F93c8aeef41ac1205ac48acfe7c8dbef5%2FCaptura%20de%20pantalla%202024-05-21%20224049.png?generation=1716324083867018&alt=media)\n\nHow are you managing these features in your models? Are you excluding them completely? Are the records with the missing values hidden in the original low-res dataset? Are these features scaled improperly in the test set, and is this why they are so different?",
    "2837074": "How to solve this problem?",
    "2839880": "Actually the tendency from 0-12 are not considered except temperature. I would not consider the state variables from the pressure levels 0-12 except temperature. I'll try this in the next days "
  }
}