{
  "id": 591540,
  "title": "Workarounds for relatively bigger data(300GB)?",
  "url": "/competitions/rsna-intracranial-aneurysm-detection/discussion/591540",
  "author_name": "Ryo Nakamura",
  "post_date": "2025-07-29T00:20:05.995000",
  "votes": 3,
  "comment_count": 1,
  "views": 0,
  "content": "<p>It takes time to load notebooks, <br>\nGCP instance stopped while installing datasets,</p>\n<p>Any ideas to go about this?</p>",
  "messages": [
    {
      "id": 3255548,
      "postDate": "2025-07-29T00:20:05.997Z",
      "content": "<p>It takes time to load notebooks, <br>\nGCP instance stopped while installing datasets,</p>\n<p>Any ideas to go about this?</p>",
      "rawMarkdown": "It takes time to load notebooks, \nGCP instance stopped while installing datasets,\n\nAny ideas to go about this?",
      "votes": 3
    },
    {
      "id": 3262411,
      "postDate": "2025-08-03T15:15:19.713Z",
      "content": "<p>I'm running into the same issue on GCP. My instance is set up and ready to train, but downloading all the series is really slow due to Kaggle's rate limits (getting 429 errors frequently). It’s also hard to clean or verify the data without pulling everything first.</p>\n<p>If anyone has a good pipeline for chunked downloading, zipping, or uploading subsets to Kaggle Datasets (or GCS mirrors), I’d love to coordinate.</p>\n<p>Also curious if Kaggle staff might allow direct access via GCS or batch archives for this competition?</p>\n<p>Appreciate any suggestions or collab ideas!</p>",
      "rawMarkdown": "I'm running into the same issue on GCP. My instance is set up and ready to train, but downloading all the series is really slow due to Kaggle's rate limits (getting 429 errors frequently). It’s also hard to clean or verify the data without pulling everything first.\n\nIf anyone has a good pipeline for chunked downloading, zipping, or uploading subsets to Kaggle Datasets (or GCS mirrors), I’d love to coordinate.\n\nAlso curious if Kaggle staff might allow direct access via GCS or batch archives for this competition?\n\nAppreciate any suggestions or collab ideas!\n"
    }
  ],
  "comments": [
    {
      "id": 3262411,
      "author_name": "Carmen Montero",
      "author_url": "",
      "post_date": "2025-08-03T15:15:19.713000",
      "content": "<p>I'm running into the same issue on GCP. My instance is set up and ready to train, but downloading all the series is really slow due to Kaggle's rate limits (getting 429 errors frequently). It’s also hard to clean or verify the data without pulling everything first.</p>\n<p>If anyone has a good pipeline for chunked downloading, zipping, or uploading subsets to Kaggle Datasets (or GCS mirrors), I’d love to coordinate.</p>\n<p>Also curious if Kaggle staff might allow direct access via GCS or batch archives for this competition?</p>\n<p>Appreciate any suggestions or collab ideas!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3255548": "It takes time to load notebooks, \nGCP instance stopped while installing datasets,\n\nAny ideas to go about this?",
    "3262411": "I'm running into the same issue on GCP. My instance is set up and ready to train, but downloading all the series is really slow due to Kaggle's rate limits (getting 429 errors frequently). It’s also hard to clean or verify the data without pulling everything first.\n\nIf anyone has a good pipeline for chunked downloading, zipping, or uploading subsets to Kaggle Datasets (or GCS mirrors), I’d love to coordinate.\n\nAlso curious if Kaggle staff might allow direct access via GCS or batch archives for this competition?\n\nAppreciate any suggestions or collab ideas!\n"
  }
}