{
  "id": 356025,
  "title": "Data too huge for colab",
  "url": "/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/356025",
  "author_name": "th3y3llowbird",
  "post_date": "2022-09-29T04:29:50.019000",
  "votes": 8,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi kagglers,<br>\nI had a very generic and maybe somewhat stupid question , but I have been trying to use this dataset on colab but its too huge for colab or drive.<br>\nNeed some suggestion if  there is any way to use these kind of huge competition datasets  on colab as it is,<br>\nThanks</p>",
  "messages": [
    {
      "id": 1961193,
      "postDate": "2022-09-29T04:29:50.020Z",
      "content": "<p>Hi kagglers,<br>\nI had a very generic and maybe somewhat stupid question , but I have been trying to use this dataset on colab but its too huge for colab or drive.<br>\nNeed some suggestion if  there is any way to use these kind of huge competition datasets  on colab as it is,<br>\nThanks</p>",
      "rawMarkdown": "Hi kagglers,\nI had a very generic and maybe somewhat stupid question , but I have been trying to use this dataset on colab but its too huge for colab or drive.\nNeed some suggestion if  there is any way to use these kind of huge competition datasets  on colab as it is,\nThanks",
      "votes": 8
    },
    {
      "id": 1976374,
      "postDate": "2022-10-07T10:05:49.940Z",
      "content": "<p>I saw a similar discussion in another competition, where @ min fuka described a solution how to access such a large dataset in Colab (how to mount it). Please see below the outline and the link to the discussion if you find it helpful:</p>\n<p><a href=\"https://www.kaggle.com/competitions/mayo-clinic-strip-ai/discussion/343470\" target=\"_blank\">https://www.kaggle.com/competitions/mayo-clinic-strip-ai/discussion/343470</a></p>\n<pre><code>1) Open new notebok in kaggle_notebook and display GCS path in kaggle_notebook.\n\nfrom kaggle_datasets import KaggleDatasets\nKaggleDatasets().get_gcs_path()\n\ngs://{path-name}\n\n2) Start colab and allow authentication of Google Cloud SDK via web browser.\n\nfrom google.colab import auth\nauth.authenticate_user()\n\nPaste the authentication code copied above into the colab blank and enter to authenticate.\n\n3) GPG key acquisition / Google Cloud Strage FUSE installation in colab environment.\n\n!echo \"deb http://packages.cloud.google.com/apt gcsfuse-lsb_release -c -s main\" | sudo tee /etc/apt/sources.list.d/gcsfuse.list\n!curl https://packages.cloud.google.com/apt/doc/apt-key.gpg | sudo apt-key add -\n!apt-get -y -q update\n!apt-get -y -q install gcsfuse\n\n4) Create a directory for mount with any name you like / mount the GCS path with the directory you just created using the gscfuse command.\n\n!mkdir -p tmp\n!gcsfuse --implicit-dirs --limit-bytes-per-sec -1 --limit-ops-per-sec -1 \"｛①path-name｝\" tmp\n</code></pre>",
      "rawMarkdown": "I saw a similar discussion in another competition, where @ min fuka described a solution how to access such a large dataset in Colab (how to mount it). Please see below the outline and the link to the discussion if you find it helpful:\n\nhttps://www.kaggle.com/competitions/mayo-clinic-strip-ai/discussion/343470\n\n```\n1) Open new notebok in kaggle_notebook and display GCS path in kaggle_notebook.\n\nfrom kaggle_datasets import KaggleDatasets\nKaggleDatasets().get_gcs_path()\n\ngs://{path-name}\n\n2) Start colab and allow authentication of Google Cloud SDK via web browser.\n\nfrom google.colab import auth\nauth.authenticate_user()\n\nPaste the authentication code copied above into the colab blank and enter to authenticate.\n\n3) GPG key acquisition / Google Cloud Strage FUSE installation in colab environment.\n\n!echo \"deb http://packages.cloud.google.com/apt gcsfuse-lsb_release -c -s main\" | sudo tee /etc/apt/sources.list.d/gcsfuse.list\n!curl https://packages.cloud.google.com/apt/doc/apt-key.gpg | sudo apt-key add -\n!apt-get -y -q update\n!apt-get -y -q install gcsfuse\n\n4) Create a directory for mount with any name you like / mount the GCS path with the directory you just created using the gscfuse command.\n\n!mkdir -p tmp\n!gcsfuse --implicit-dirs --limit-bytes-per-sec -1 --limit-ops-per-sec -1 \"｛①path-name｝\" tmp\n```\n",
      "votes": 3
    },
    {
      "id": 1961218,
      "postDate": "2022-09-29T04:42:59.697Z",
      "content": "<p>I had a similar query and with the help of <a href=\"https://www.kaggle.com/icemantd\" target=\"_blank\">@icemantd</a> I am experimenting on the below points<br>\n1) Downscale whole slide images - which will likely result in loss of training signal but you can build and test end to end models faster.<br>\n2) Tile creation and selection - in this method you create tiles from images and have criteria for selecting tiles with maximum signal, and then train an MIL based method which is more complicated but may do better (or not, but there's a better chance). Tile creation and selection is key, this will reduce your training time based on pre-selected best guess part of the data.<br>\nCheers!</p>",
      "rawMarkdown": "I had a similar query and with the help of @icemantd I am experimenting on the below points\n1) Downscale whole slide images - which will likely result in loss of training signal but you can build and test end to end models faster.\n2) Tile creation and selection - in this method you create tiles from images and have criteria for selecting tiles with maximum signal, and then train an MIL based method which is more complicated but may do better (or not, but there's a better chance). Tile creation and selection is key, this will reduce your training time based on pre-selected best guess part of the data.\nCheers!",
      "votes": 3,
      "replies": [
        {
          "id": 1961386,
          "postDate": "2022-09-29T06:14:29.627Z",
          "content": "<p>Yeah actually I am aware of the idea of efficient data processing but actually my question was, is there any specific way or 'hack' to load the entire dataset to colab maybe using gcp bucket or something like that.</p>",
          "rawMarkdown": "Yeah actually I am aware of the idea of efficient data processing but actually my question was, is there any specific way or 'hack' to load the entire dataset to colab maybe using gcp bucket or something like that."
        }
      ]
    },
    {
      "id": 1962349,
      "postDate": "2022-09-29T16:43:28.840Z",
      "content": "<p>I recommend using png data instead of DICOM. Depending on image size you should be able to compress the entire training image data to 10-20 GB. Few people have already shared the png dataset, so you can directly use them to start with. </p>",
      "rawMarkdown": "I recommend using png data instead of DICOM. Depending on image size you should be able to compress the entire training image data to 10-20 GB. Few people have already shared the png dataset, so you can directly use them to start with. ",
      "votes": 2
    },
    {
      "id": 1961639,
      "postDate": "2022-09-29T09:29:22.613Z",
      "content": "<p>I have experimented previously with large datasets in mayo clinic competition and must say conversion to tf record actually does the job. It will compress the size of the dataset a little bit and will make sure the ram crashing issue does not happen on colab.</p>",
      "rawMarkdown": "I have experimented previously with large datasets in mayo clinic competition and must say conversion to tf record actually does the job. It will compress the size of the dataset a little bit and will make sure the ram crashing issue does not happen on colab.",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 1976374,
      "author_name": "Sergey Kr",
      "author_url": "",
      "post_date": "2022-10-07T10:05:49.940000",
      "content": "<p>I saw a similar discussion in another competition, where @ min fuka described a solution how to access such a large dataset in Colab (how to mount it). Please see below the outline and the link to the discussion if you find it helpful:</p>\n<p><a href=\"https://www.kaggle.com/competitions/mayo-clinic-strip-ai/discussion/343470\" target=\"_blank\">https://www.kaggle.com/competitions/mayo-clinic-strip-ai/discussion/343470</a></p>\n<pre><code>1) Open new notebok in kaggle_notebook and display GCS path in kaggle_notebook.\n\nfrom kaggle_datasets import KaggleDatasets\nKaggleDatasets().get_gcs_path()\n\ngs://{path-name}\n\n2) Start colab and allow authentication of Google Cloud SDK via web browser.\n\nfrom google.colab import auth\nauth.authenticate_user()\n\nPaste the authentication code copied above into the colab blank and enter to authenticate.\n\n3) GPG key acquisition / Google Cloud Strage FUSE installation in colab environment.\n\n!echo \"deb http://packages.cloud.google.com/apt gcsfuse-lsb_release -c -s main\" | sudo tee /etc/apt/sources.list.d/gcsfuse.list\n!curl https://packages.cloud.google.com/apt/doc/apt-key.gpg | sudo apt-key add -\n!apt-get -y -q update\n!apt-get -y -q install gcsfuse\n\n4) Create a directory for mount with any name you like / mount the GCS path with the directory you just created using the gscfuse command.\n\n!mkdir -p tmp\n!gcsfuse --implicit-dirs --limit-bytes-per-sec -1 --limit-ops-per-sec -1 \"｛①path-name｝\" tmp\n</code></pre>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1961218,
      "author_name": "Ab_Rafey",
      "author_url": "",
      "post_date": "2022-09-29T04:42:59.697000",
      "content": "<p>I had a similar query and with the help of <a href=\"https://www.kaggle.com/icemantd\" target=\"_blank\">@icemantd</a> I am experimenting on the below points<br>\n1) Downscale whole slide images - which will likely result in loss of training signal but you can build and test end to end models faster.<br>\n2) Tile creation and selection - in this method you create tiles from images and have criteria for selecting tiles with maximum signal, and then train an MIL based method which is more complicated but may do better (or not, but there's a better chance). Tile creation and selection is key, this will reduce your training time based on pre-selected best guess part of the data.<br>\nCheers!</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1961386,
          "author_name": "th3y3llowbird",
          "author_url": "",
          "post_date": "2022-09-29T06:14:29.627000",
          "content": "<p>Yeah actually I am aware of the idea of efficient data processing but actually my question was, is there any specific way or 'hack' to load the entire dataset to colab maybe using gcp bucket or something like that.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1962349,
      "author_name": "Ranjeet",
      "author_url": "",
      "post_date": "2022-09-29T16:43:28.840000",
      "content": "<p>I recommend using png data instead of DICOM. Depending on image size you should be able to compress the entire training image data to 10-20 GB. Few people have already shared the png dataset, so you can directly use them to start with. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1961639,
      "author_name": "Mrinal Tyagi",
      "author_url": "",
      "post_date": "2022-09-29T09:29:22.613000",
      "content": "<p>I have experimented previously with large datasets in mayo clinic competition and must say conversion to tf record actually does the job. It will compress the size of the dataset a little bit and will make sure the ram crashing issue does not happen on colab.</p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1961193": "Hi kagglers,\nI had a very generic and maybe somewhat stupid question , but I have been trying to use this dataset on colab but its too huge for colab or drive.\nNeed some suggestion if  there is any way to use these kind of huge competition datasets  on colab as it is,\nThanks",
    "1976374": "I saw a similar discussion in another competition, where @ min fuka described a solution how to access such a large dataset in Colab (how to mount it). Please see below the outline and the link to the discussion if you find it helpful:\n\nhttps://www.kaggle.com/competitions/mayo-clinic-strip-ai/discussion/343470\n\n```\n1) Open new notebok in kaggle_notebook and display GCS path in kaggle_notebook.\n\nfrom kaggle_datasets import KaggleDatasets\nKaggleDatasets().get_gcs_path()\n\ngs://{path-name}\n\n2) Start colab and allow authentication of Google Cloud SDK via web browser.\n\nfrom google.colab import auth\nauth.authenticate_user()\n\nPaste the authentication code copied above into the colab blank and enter to authenticate.\n\n3) GPG key acquisition / Google Cloud Strage FUSE installation in colab environment.\n\n!echo \"deb http://packages.cloud.google.com/apt gcsfuse-lsb_release -c -s main\" | sudo tee /etc/apt/sources.list.d/gcsfuse.list\n!curl https://packages.cloud.google.com/apt/doc/apt-key.gpg | sudo apt-key add -\n!apt-get -y -q update\n!apt-get -y -q install gcsfuse\n\n4) Create a directory for mount with any name you like / mount the GCS path with the directory you just created using the gscfuse command.\n\n!mkdir -p tmp\n!gcsfuse --implicit-dirs --limit-bytes-per-sec -1 --limit-ops-per-sec -1 \"｛①path-name｝\" tmp\n```\n",
    "1961218": "I had a similar query and with the help of @icemantd I am experimenting on the below points\n1) Downscale whole slide images - which will likely result in loss of training signal but you can build and test end to end models faster.\n2) Tile creation and selection - in this method you create tiles from images and have criteria for selecting tiles with maximum signal, and then train an MIL based method which is more complicated but may do better (or not, but there's a better chance). Tile creation and selection is key, this will reduce your training time based on pre-selected best guess part of the data.\nCheers!",
    "1962349": "I recommend using png data instead of DICOM. Depending on image size you should be able to compress the entire training image data to 10-20 GB. Few people have already shared the png dataset, so you can directly use them to start with. ",
    "1961639": "I have experimented previously with large datasets in mayo clinic competition and must say conversion to tf record actually does the job. It will compress the size of the dataset a little bit and will make sure the ram crashing issue does not happen on colab."
  }
}