{
  "id": 501430,
  "title": "Does the size of train dataset have a big impact on the test score?",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/501430",
  "author_name": "嘴爷",
  "post_date": "2024-05-09T07:06:53.067000",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I currently just use 3,125,000 rows (5 files of Giba's dataset) as train dataset and 125,000 rows as val dataset due to the limitation of my computer. I wonder how much improvement can I get if I expand the train dataset to 6,000,000 rows for example or event more.</p>",
  "messages": [
    {
      "id": 2802780,
      "postDate": "2024-05-09T07:06:53.067Z",
      "content": "<p>I currently just use 3,125,000 rows (5 files of Giba's dataset) as train dataset and 125,000 rows as val dataset due to the limitation of my computer. I wonder how much improvement can I get if I expand the train dataset to 6,000,000 rows for example or event more.</p>",
      "rawMarkdown": "I currently just use 3,125,000 rows (5 files of Giba's dataset) as train dataset and 125,000 rows as val dataset due to the limitation of my computer. I wonder how much improvement can I get if I expand the train dataset to 6,000,000 rows for example or event more.",
      "votes": 1
    },
    {
      "id": 2805618,
      "postDate": "2024-05-10T16:36:02.820Z",
      "content": "<p>I have only been working with very small subsets. (1k, 10k, 100k), and I've found quite the difference in cv score with each one.  about a ~.1 cv score increase with each</p>",
      "rawMarkdown": "I have only been working with very small subsets. (1k, 10k, 100k), and I've found quite the difference in cv score with each one.  about a ~.1 cv score increase with each",
      "replies": [
        {
          "id": 2806206,
          "postDate": "2024-05-11T00:48:47.933Z",
          "content": "<p>Maybe it's because 1k/10k is too small?</p>",
          "rawMarkdown": "Maybe it's because 1k/10k is too small?",
          "votes": 1
        }
      ]
    },
    {
      "id": 2803630,
      "postDate": "2024-05-09T15:27:12.950Z",
      "content": "<p>I'm also not sure, because I checked the probability density plots of the features in the training and test sets, and the distributions of about 20 features are completely different, which may have a significant impact on the scores.</p>",
      "rawMarkdown": "I'm also not sure, because I checked the probability density plots of the features in the training and test sets, and the distributions of about 20 features are completely different, which may have a significant impact on the scores.",
      "replies": [
        {
          "id": 2804339,
          "postDate": "2024-05-10T01:53:05.510Z",
          "content": "<p>It sounds strange because my CV works very well, just about ~0.01 gap with PB. And many of us get the same conclusion, see <a href=\"url\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/498194</a></p>",
          "rawMarkdown": "It sounds strange because my CV works very well, just about ~0.01 gap with PB. And many of us get the same conclusion, see [https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/498194](url)",
          "replies": [
            {
              "id": 2804345,
              "postDate": "2024-05-10T02:16:30.013Z",
              "content": "<p>Thank you for the reminder, I will give it another try.</p>",
              "rawMarkdown": "Thank you for the reminder, I will give it another try."
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2805618,
      "author_name": "Jekasm19",
      "author_url": "",
      "post_date": "2024-05-10T16:36:02.820000",
      "content": "<p>I have only been working with very small subsets. (1k, 10k, 100k), and I've found quite the difference in cv score with each one.  about a ~.1 cv score increase with each</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2806206,
          "author_name": "嘴爷",
          "author_url": "",
          "post_date": "2024-05-11T00:48:47.933000",
          "content": "<p>Maybe it's because 1k/10k is too small?</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2803630,
      "author_name": "Zhuoqun Li",
      "author_url": "",
      "post_date": "2024-05-09T15:27:12.950000",
      "content": "<p>I'm also not sure, because I checked the probability density plots of the features in the training and test sets, and the distributions of about 20 features are completely different, which may have a significant impact on the scores.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2804339,
          "author_name": "嘴爷",
          "author_url": "",
          "post_date": "2024-05-10T01:53:05.510000",
          "content": "<p>It sounds strange because my CV works very well, just about ~0.01 gap with PB. And many of us get the same conclusion, see <a href=\"url\" target=\"_blank\">https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/498194</a></p>",
          "votes": 0,
          "replies": [
            {
              "id": 2804345,
              "author_name": "Zhuoqun Li",
              "author_url": "",
              "post_date": "2024-05-10T02:16:30.013000",
              "content": "<p>Thank you for the reminder, I will give it another try.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2802780": "I currently just use 3,125,000 rows (5 files of Giba's dataset) as train dataset and 125,000 rows as val dataset due to the limitation of my computer. I wonder how much improvement can I get if I expand the train dataset to 6,000,000 rows for example or event more.",
    "2805618": "I have only been working with very small subsets. (1k, 10k, 100k), and I've found quite the difference in cv score with each one.  about a ~.1 cv score increase with each",
    "2803630": "I'm also not sure, because I checked the probability density plots of the features in the training and test sets, and the distributions of about 20 features are completely different, which may have a significant impact on the scores."
  }
}