{
  "id": 497597,
  "title": "Using full low-res dataset?",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/497597",
  "author_name": "slime",
  "post_date": "2024-04-25T05:49:13.137000",
  "votes": 4,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Do the orgs have the intuition if more (I assume temporal) data would bring the model performance closer to the achievable optimum? </p>\n<p>The Kaggle dataset is a 7x-downsampled version of the original dataset, that can be seen from this <a href=\"https://github.com/leap-stc/ClimSim/blob/58cb6a2e4fbc27ffd6acb9050a2f1df44b933a05/for_kaggle_users.py#L138\" target=\"_blank\">code</a></p>\n<p>Just wondering if we should go through the trouble of downloading raw dataset and converting it to the competition format ourselves</p>",
  "messages": [
    {
      "id": 2774270,
      "postDate": "2024-04-25T05:49:13.137Z",
      "content": "<p>Do the orgs have the intuition if more (I assume temporal) data would bring the model performance closer to the achievable optimum? </p>\n<p>The Kaggle dataset is a 7x-downsampled version of the original dataset, that can be seen from this <a href=\"https://github.com/leap-stc/ClimSim/blob/58cb6a2e4fbc27ffd6acb9050a2f1df44b933a05/for_kaggle_users.py#L138\" target=\"_blank\">code</a></p>\n<p>Just wondering if we should go through the trouble of downloading raw dataset and converting it to the competition format ourselves</p>",
      "rawMarkdown": "Do the orgs have the intuition if more (I assume temporal) data would bring the model performance closer to the achievable optimum? \n\nThe Kaggle dataset is a 7x-downsampled version of the original dataset, that can be seen from this [code](https://github.com/leap-stc/ClimSim/blob/58cb6a2e4fbc27ffd6acb9050a2f1df44b933a05/for_kaggle_users.py#L138)\n\nJust wondering if we should go through the trouble of downloading raw dataset and converting it to the competition format ourselves",
      "votes": 4
    },
    {
      "id": 2800893,
      "postDate": "2024-05-08T11:55:12.653Z",
      "content": "<p>Using more data helps, i can confirm that.</p>",
      "rawMarkdown": "Using more data helps, i can confirm that.",
      "votes": 1,
      "replies": [
        {
          "id": 2800923,
          "postDate": "2024-05-08T12:15:19.813Z",
          "content": "<p>Thank you, I failed to re-create hosts dataset, - for some reasons, even for the first month ('01_02') samples start to diverge from the hosts dataset from 384th row, I couldn't figure out the reason yet</p>\n<p>upd:<br>\nnvm, it worked</p>",
          "rawMarkdown": "Thank you, I failed to re-create hosts dataset, - for some reasons, even for the first month ('01_02') samples start to diverge from the hosts dataset from 384th row, I couldn't figure out the reason yet\n\nupd:\nnvm, it worked",
          "replies": [
            {
              "id": 2846470,
              "postDate": "2024-05-31T05:57:46.713Z",
              "content": "<p><a href=\"https://www.kaggle.com/martynoveduard\" target=\"_blank\">@martynoveduard</a> May I ask, how much have you improved in terms of LB by using all the data?</p>",
              "rawMarkdown": "@martynoveduard May I ask, how much have you improved in terms of LB by using all the data?",
              "votes": 2
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2800893,
      "author_name": "Youri Matiounine",
      "author_url": "",
      "post_date": "2024-05-08T11:55:12.653000",
      "content": "<p>Using more data helps, i can confirm that.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2800923,
          "author_name": "slime",
          "author_url": "",
          "post_date": "2024-05-08T12:15:19.813000",
          "content": "<p>Thank you, I failed to re-create hosts dataset, - for some reasons, even for the first month ('01_02') samples start to diverge from the hosts dataset from 384th row, I couldn't figure out the reason yet</p>\n<p>upd:<br>\nnvm, it worked</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2846470,
              "author_name": "lhwcv",
              "author_url": "",
              "post_date": "2024-05-31T05:57:46.713000",
              "content": "<p><a href=\"https://www.kaggle.com/martynoveduard\" target=\"_blank\">@martynoveduard</a> May I ask, how much have you improved in terms of LB by using all the data?</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2774270": "Do the orgs have the intuition if more (I assume temporal) data would bring the model performance closer to the achievable optimum? \n\nThe Kaggle dataset is a 7x-downsampled version of the original dataset, that can be seen from this [code](https://github.com/leap-stc/ClimSim/blob/58cb6a2e4fbc27ffd6acb9050a2f1df44b933a05/for_kaggle_users.py#L138)\n\nJust wondering if we should go through the trouble of downloading raw dataset and converting it to the competition format ourselves",
    "2800893": "Using more data helps, i can confirm that."
  }
}