{
  "id": 600151,
  "title": "Handling and Training with Large (370 GB) RSNA Intracranial Aneurysm Dataset in Kaggle Kernels",
  "url": "/competitions/rsna-intracranial-aneurysm-detection/discussion/600151",
  "author_name": "TarunSingh931",
  "post_date": "2025-08-21T09:14:29.824000",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p>1) I’m participating in the RSNA Intracranial Aneurysm Detection challenge, and the dataset size is about 370 GB. I want to understand the best approach to manage this large dataset within Kaggle kernels:</p>\n<p>2) Is it feasible to load and preprocess the entire 370 GB dataset directly inside a Kaggle kernel for training, or will this be too slow or impractical?</p>\n<p>3) What strategies have you used to handle such large medical imaging datasets on Kaggle? </p>\n<p>Any recommendations on balancing computational resource limits and training efficiency for this challenge?</p>\n<p>I’d appreciate your insights and experiences—thank you!</p>",
  "messages": [
    {
      "id": 3273022,
      "postDate": "2025-08-22T01:03:45.137Z",
      "content": "<ol>\n<li>Load and convert the data into some efficient form like .npy files first. </li>\n<li>I would not do this, but some people have done it. They load the .npy files into kaggle and run the training here. Though I don't speak from experience, I read it in some discussion post. </li>\n<li>Convert the 3d data into some 3 channel data. Some people have posted their inference code, you could take some suggestions from them. Using 2d 3 channel data for training would make it more efficient. </li>\n</ol>",
      "rawMarkdown": "1. Load and convert the data into some efficient form like .npy files first. \n2. I would not do this, but some people have done it. They load the .npy files into kaggle and run the training here. Though I don't speak from experience, I read it in some discussion post. \n3. Convert the 3d data into some 3 channel data. Some people have posted their inference code, you could take some suggestions from them. Using 2d 3 channel data for training would make it more efficient. \n",
      "votes": 2,
      "replies": [
        {
          "id": 3273186,
          "postDate": "2025-08-22T10:08:53.530Z",
          "content": "<p>Thanks  🫡</p>",
          "rawMarkdown": "Thanks  🫡"
        }
      ]
    },
    {
      "id": 3272586,
      "postDate": "2025-08-21T09:14:29.823Z",
      "content": "<p>Hi everyone,</p>\n<p>1) I’m participating in the RSNA Intracranial Aneurysm Detection challenge, and the dataset size is about 370 GB. I want to understand the best approach to manage this large dataset within Kaggle kernels:</p>\n<p>2) Is it feasible to load and preprocess the entire 370 GB dataset directly inside a Kaggle kernel for training, or will this be too slow or impractical?</p>\n<p>3) What strategies have you used to handle such large medical imaging datasets on Kaggle? </p>\n<p>Any recommendations on balancing computational resource limits and training efficiency for this challenge?</p>\n<p>I’d appreciate your insights and experiences—thank you!</p>",
      "rawMarkdown": "Hi everyone,\n\n1) I’m participating in the RSNA Intracranial Aneurysm Detection challenge, and the dataset size is about 370 GB. I want to understand the best approach to manage this large dataset within Kaggle kernels:\n\n2) Is it feasible to load and preprocess the entire 370 GB dataset directly inside a Kaggle kernel for training, or will this be too slow or impractical?\n\n3) What strategies have you used to handle such large medical imaging datasets on Kaggle? \n\nAny recommendations on balancing computational resource limits and training efficiency for this challenge?\n\nI’d appreciate your insights and experiences—thank you!",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 3273022,
      "author_name": "jucious",
      "author_url": "",
      "post_date": "2025-08-22T01:03:45.137000",
      "content": "<ol>\n<li>Load and convert the data into some efficient form like .npy files first. </li>\n<li>I would not do this, but some people have done it. They load the .npy files into kaggle and run the training here. Though I don't speak from experience, I read it in some discussion post. </li>\n<li>Convert the 3d data into some 3 channel data. Some people have posted their inference code, you could take some suggestions from them. Using 2d 3 channel data for training would make it more efficient. </li>\n</ol>",
      "votes": 2,
      "replies": [
        {
          "id": 3273186,
          "author_name": "TarunSingh931",
          "author_url": "",
          "post_date": "2025-08-22T10:08:53.530000",
          "content": "<p>Thanks  🫡</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3273022": "1. Load and convert the data into some efficient form like .npy files first. \n2. I would not do this, but some people have done it. They load the .npy files into kaggle and run the training here. Though I don't speak from experience, I read it in some discussion post. \n3. Convert the 3d data into some 3 channel data. Some people have posted their inference code, you could take some suggestions from them. Using 2d 3 channel data for training would make it more efficient. \n",
    "3272586": "Hi everyone,\n\n1) I’m participating in the RSNA Intracranial Aneurysm Detection challenge, and the dataset size is about 370 GB. I want to understand the best approach to manage this large dataset within Kaggle kernels:\n\n2) Is it feasible to load and preprocess the entire 370 GB dataset directly inside a Kaggle kernel for training, or will this be too slow or impractical?\n\n3) What strategies have you used to handle such large medical imaging datasets on Kaggle? \n\nAny recommendations on balancing computational resource limits and training efficiency for this challenge?\n\nI’d appreciate your insights and experiences—thank you!"
  }
}