{
  "id": 593830,
  "title": "Dataset loading",
  "url": "/competitions/rsna-intracranial-aneurysm-detection/discussion/593830",
  "author_name": "beaustroms",
  "post_date": "2025-07-30T20:15:35.033000",
  "votes": -1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hi all! Curious to give this challenge a go, but I'm struggling to figure out how to load such a massive dataset. Does Kaggle offer some sort of streaming? Mainly struggling with disk space (I don't have 300+ GB to spare).</p>\n<p>What have you all been doing to overcome this?</p>",
  "messages": [
    {
      "id": 3289974,
      "postDate": "2025-09-16T22:09:54.727Z",
      "content": "<p>I made a streamable webdataset version here: <a href=\"https://huggingface.co/datasets/NilanE/rsna-intracranial-aneurysm-detection-2025-WDS\" target=\"_blank\">https://huggingface.co/datasets/NilanE/rsna-intracranial-aneurysm-detection-2025-WDS</a></p>\n<p>It's still lossless though, so ~130gb when using FFV1 to encode the volumes.</p>\n<p>Edit:<br>\n…whoops. Had to make it private after re-reading the rules:<br>\n'You agree to use reasonable and suitable measures to prevent persons who have not formally agreed to these Rules from gaining access to the Competition Data. You agree not to transmit, duplicate, publish, redistribute or otherwise provide or make available the Competition Data to any party not participating in the Competition. You agree to notify Kaggle immediately upon learning of any possible unauthorized transmission of or unauthorized access to the Competition Data and agree to work with Kaggle to rectify any unauthorized transmission or access.'</p>",
      "rawMarkdown": "I made a streamable webdataset version here: https://huggingface.co/datasets/NilanE/rsna-intracranial-aneurysm-detection-2025-WDS\n\nIt's still lossless though, so ~130gb when using FFV1 to encode the volumes.\n\nEdit:\n...whoops. Had to make it private after re-reading the rules:\n'You agree to use reasonable and suitable measures to prevent persons who have not formally agreed to these Rules from gaining access to the Competition Data. You agree not to transmit, duplicate, publish, redistribute or otherwise provide or make available the Competition Data to any party not participating in the Competition. You agree to notify Kaggle immediately upon learning of any possible unauthorized transmission of or unauthorized access to the Competition Data and agree to work with Kaggle to rectify any unauthorized transmission or access.'"
    },
    {
      "id": 3265332,
      "postDate": "2025-08-07T10:51:04.700Z",
      "content": "<p>For data loading, you can check out previous RSNA competitions. And for storage, just get a 1TB SSD.</p>",
      "rawMarkdown": "For data loading, you can check out previous RSNA competitions. And for storage, just get a 1TB SSD."
    },
    {
      "id": 3262412,
      "postDate": "2025-08-03T15:17:47.837Z",
      "content": "<p>I’m running into the same thing. I don’t have 300GB of local space either, and downloading everything to GCP has been slow and kind of frustrating with API limits.</p>\n<p>One thing I’m trying is just working with a small subset of the data directly inside a Kaggle Notebook for now like picking 10–20 series and using that to build and test my cleaning + preprocessing pipeline. Once that’s working, I can scale it up later if I manage to get more data downloaded.</p>\n<p>Would love to know what others are doing too. Still trying to figure out the best approach myself!</p>",
      "rawMarkdown": "I’m running into the same thing. I don’t have 300GB of local space either, and downloading everything to GCP has been slow and kind of frustrating with API limits.\n\nOne thing I’m trying is just working with a small subset of the data directly inside a Kaggle Notebook for now like picking 10–20 series and using that to build and test my cleaning + preprocessing pipeline. Once that’s working, I can scale it up later if I manage to get more data downloaded.\n\nWould love to know what others are doing too. Still trying to figure out the best approach myself!\n"
    },
    {
      "id": 3258531,
      "postDate": "2025-07-30T20:15:35.033Z",
      "content": "<p>Hi all! Curious to give this challenge a go, but I'm struggling to figure out how to load such a massive dataset. Does Kaggle offer some sort of streaming? Mainly struggling with disk space (I don't have 300+ GB to spare).</p>\n<p>What have you all been doing to overcome this?</p>",
      "rawMarkdown": "Hi all! Curious to give this challenge a go, but I'm struggling to figure out how to load such a massive dataset. Does Kaggle offer some sort of streaming? Mainly struggling with disk space (I don't have 300+ GB to spare).\n\nWhat have you all been doing to overcome this?",
      "votes": -1
    }
  ],
  "comments": [
    {
      "id": 3289974,
      "author_name": "NilanE",
      "author_url": "",
      "post_date": "2025-09-16T22:09:54.727000",
      "content": "<p>I made a streamable webdataset version here: <a href=\"https://huggingface.co/datasets/NilanE/rsna-intracranial-aneurysm-detection-2025-WDS\" target=\"_blank\">https://huggingface.co/datasets/NilanE/rsna-intracranial-aneurysm-detection-2025-WDS</a></p>\n<p>It's still lossless though, so ~130gb when using FFV1 to encode the volumes.</p>\n<p>Edit:<br>\n…whoops. Had to make it private after re-reading the rules:<br>\n'You agree to use reasonable and suitable measures to prevent persons who have not formally agreed to these Rules from gaining access to the Competition Data. You agree not to transmit, duplicate, publish, redistribute or otherwise provide or make available the Competition Data to any party not participating in the Competition. You agree to notify Kaggle immediately upon learning of any possible unauthorized transmission of or unauthorized access to the Competition Data and agree to work with Kaggle to rectify any unauthorized transmission or access.'</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3265332,
      "author_name": "Seeing Times",
      "author_url": "",
      "post_date": "2025-08-07T10:51:04.700000",
      "content": "<p>For data loading, you can check out previous RSNA competitions. And for storage, just get a 1TB SSD.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3262412,
      "author_name": "Carmen Montero",
      "author_url": "",
      "post_date": "2025-08-03T15:17:47.837000",
      "content": "<p>I’m running into the same thing. I don’t have 300GB of local space either, and downloading everything to GCP has been slow and kind of frustrating with API limits.</p>\n<p>One thing I’m trying is just working with a small subset of the data directly inside a Kaggle Notebook for now like picking 10–20 series and using that to build and test my cleaning + preprocessing pipeline. Once that’s working, I can scale it up later if I manage to get more data downloaded.</p>\n<p>Would love to know what others are doing too. Still trying to figure out the best approach myself!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3289974": "I made a streamable webdataset version here: https://huggingface.co/datasets/NilanE/rsna-intracranial-aneurysm-detection-2025-WDS\n\nIt's still lossless though, so ~130gb when using FFV1 to encode the volumes.\n\nEdit:\n...whoops. Had to make it private after re-reading the rules:\n'You agree to use reasonable and suitable measures to prevent persons who have not formally agreed to these Rules from gaining access to the Competition Data. You agree not to transmit, duplicate, publish, redistribute or otherwise provide or make available the Competition Data to any party not participating in the Competition. You agree to notify Kaggle immediately upon learning of any possible unauthorized transmission of or unauthorized access to the Competition Data and agree to work with Kaggle to rectify any unauthorized transmission or access.'",
    "3265332": "For data loading, you can check out previous RSNA competitions. And for storage, just get a 1TB SSD.",
    "3262412": "I’m running into the same thing. I don’t have 300GB of local space either, and downloading everything to GCP has been slow and kind of frustrating with API limits.\n\nOne thing I’m trying is just working with a small subset of the data directly inside a Kaggle Notebook for now like picking 10–20 series and using that to build and test my cleaning + preprocessing pipeline. Once that’s working, I can scale it up later if I manage to get more data downloaded.\n\nWould love to know what others are doing too. Still trying to figure out the best approach myself!\n",
    "3258531": "Hi all! Curious to give this challenge a go, but I'm struggling to figure out how to load such a massive dataset. Does Kaggle offer some sort of streaming? Mainly struggling with disk space (I don't have 300+ GB to spare).\n\nWhat have you all been doing to overcome this?"
  }
}