{
  "id": 161500,
  "title": "Create medium resolution dataset(30GB) using kaggle notebook",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/161500",
  "author_name": "Raghawendra Singh",
  "post_date": "2020-06-25T04:54:19.781000",
  "votes": 31,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Since Kaggle working directory allows up to 4GB data, it is very difficult to create medium resolution dataset having total size of 30GB(25 tiles per image with size of 256 each tile). I found a way to do so using kaggle notebook.</p>\n\n<p>By running linux command !df -h, I found that I can use upto 2.1 TB on \"/\" mount point. Therefore I used \"/tmp\" directory to store and upload data.</p>\n\n<p>Filesystem        Size  Used Avail Use% Mounted on\noverlay           4.0T  2.0T  2.1T  50% /\ntmpfs              64M     0   64M   0% /dev\ntmpfs             9.4G     0  9.4G   0% /sys/fs/cgroup\nshm               7.0G     0  7.0G   0% /dev/shm\n/dev/mapper/snap  4.0T  2.0T  2.1T  50% /home\n/dev/loop1        4.9G   21M  4.9G   1% /kaggle/lib\ntmpfs             9.4G     0  9.4G   0% /proc/acpi\ntmpfs             9.4G     0  9.4G   0% /proc/scsi\ntmpfs             9.4G     0  9.4G   0% /sys/firmware</p>\n\n<p>please have a look.....</p>\n\n<p>Notebook URL:\n<a href=\"https://www.kaggle.com/raghaw/panda-medium-resolution-dataset-25x256x256\">https://www.kaggle.com/raghaw/panda-medium-resolution-dataset-25x256x256</a></p>\n\n<p>Created Dataset URL:\n<a href=\"https://www.kaggle.com/raghaw/panda-dataset-medium-25-256-256\">https://www.kaggle.com/raghaw/panda-dataset-medium-25-256-256</a></p>",
  "messages": [
    {
      "id": 900824,
      "postDate": "2020-06-25T04:54:19.783Z",
      "content": "<p>Since Kaggle working directory allows up to 4GB data, it is very difficult to create medium resolution dataset having total size of 30GB(25 tiles per image with size of 256 each tile). I found a way to do so using kaggle notebook.</p>\n\n<p>By running linux command !df -h, I found that I can use upto 2.1 TB on \"/\" mount point. Therefore I used \"/tmp\" directory to store and upload data.</p>\n\n<p>Filesystem        Size  Used Avail Use% Mounted on\noverlay           4.0T  2.0T  2.1T  50% /\ntmpfs              64M     0   64M   0% /dev\ntmpfs             9.4G     0  9.4G   0% /sys/fs/cgroup\nshm               7.0G     0  7.0G   0% /dev/shm\n/dev/mapper/snap  4.0T  2.0T  2.1T  50% /home\n/dev/loop1        4.9G   21M  4.9G   1% /kaggle/lib\ntmpfs             9.4G     0  9.4G   0% /proc/acpi\ntmpfs             9.4G     0  9.4G   0% /proc/scsi\ntmpfs             9.4G     0  9.4G   0% /sys/firmware</p>\n\n<p>please have a look.....</p>\n\n<p>Notebook URL:\n<a href=\"https://www.kaggle.com/raghaw/panda-medium-resolution-dataset-25x256x256\">https://www.kaggle.com/raghaw/panda-medium-resolution-dataset-25x256x256</a></p>\n\n<p>Created Dataset URL:\n<a href=\"https://www.kaggle.com/raghaw/panda-dataset-medium-25-256-256\">https://www.kaggle.com/raghaw/panda-dataset-medium-25-256-256</a></p>",
      "rawMarkdown": "Since Kaggle working directory allows up to 4GB data, it is very difficult to create medium resolution dataset having total size of 30GB(25 tiles per image with size of 256 each tile). I found a way to do so using kaggle notebook.\n\nBy running linux command !df -h, I found that I can use upto 2.1 TB on \"/\" mount point. Therefore I used \"/tmp\" directory to store and upload data.\n\nFilesystem        Size  Used Avail Use% Mounted on\noverlay           4.0T  2.0T  2.1T  50% /\ntmpfs              64M     0   64M   0% /dev\ntmpfs             9.4G     0  9.4G   0% /sys/fs/cgroup\nshm               7.0G     0  7.0G   0% /dev/shm\n/dev/mapper/snap  4.0T  2.0T  2.1T  50% /home\n/dev/loop1        4.9G   21M  4.9G   1% /kaggle/lib\ntmpfs             9.4G     0  9.4G   0% /proc/acpi\ntmpfs             9.4G     0  9.4G   0% /proc/scsi\ntmpfs             9.4G     0  9.4G   0% /sys/firmware\n\nplease have a look.....\n\nNotebook URL:\nhttps://www.kaggle.com/raghaw/panda-medium-resolution-dataset-25x256x256\n\nCreated Dataset URL:\nhttps://www.kaggle.com/raghaw/panda-dataset-medium-25-256-256",
      "votes": 31
    },
    {
      "id": 901731,
      "postDate": "2020-06-25T16:43:55.127Z",
      "content": "<p><a href=\"/raghaw\">@raghaw</a> Just a heads up, you don't actually have access to those 2TB of disk space. The 4TB partition is a read-only disk which is part of the Kaggle environment.</p>\n\n<p>When you write to /tmp you're actually writing to a writable disk which is overlayed on top of the 4TB one. </p>\n\n<p>The writeable space is quite large (though not guaranteed since we may use that space), so it's totally fine to write 30GB of space there, and I often suggest users use that scratch space anytime they need. The only reason we impose a 5GB limit on /kaggle/working, is because when you Save &gt; Run All, or Save &gt; Quick Save + outputs, we upload everything in that directory and save it for viewing/download in your Notebook Version, and put a limit on how much you can save per notebook version.</p>",
      "rawMarkdown": "@raghaw Just a heads up, you don't actually have access to those 2TB of disk space. The 4TB partition is a read-only disk which is part of the Kaggle environment.\n\nWhen you write to /tmp you're actually writing to a writable disk which is overlayed on top of the 4TB one. \n\nThe writeable space is quite large (though not guaranteed since we may use that space), so it's totally fine to write 30GB of space there, and I often suggest users use that scratch space anytime they need. The only reason we impose a 5GB limit on /kaggle/working, is because when you Save &gt; Run All, or Save &gt; Quick Save + outputs, we upload everything in that directory and save it for viewing/download in your Notebook Version, and put a limit on how much you can save per notebook version.",
      "votes": 9,
      "replies": [
        {
          "id": 901746,
          "postDate": "2020-06-25T17:01:16.460Z",
          "content": "<p>Thanks <a href=\"/herbison\">@herbison</a> for clarification.</p>",
          "rawMarkdown": "Thanks @herbison for clarification."
        }
      ]
    },
    {
      "id": 900943,
      "postDate": "2020-06-25T06:49:37.790Z",
      "content": "<p>offline package installation in : pneumothorax,severstal steel defect detection then retraining strategy in bengali.ai and now in pandas effective dataset creation!!\nyou are awesome, i have been learning a lot from you in almost all computer vision competition,thank you for everything,looking forward to see some unique work like this in melanoma competition as well &lt;3 </p>",
      "rawMarkdown": "offline package installation in : pneumothorax,severstal steel defect detection then retraining strategy in bengali.ai and now in pandas effective dataset creation!!\nyou are awesome, i have been learning a lot from you in almost all computer vision competition,thank you for everything,looking forward to see some unique work like this in melanoma competition as well &lt;3 ",
      "votes": 2
    },
    {
      "id": 902136,
      "postDate": "2020-06-26T00:41:15.053Z",
      "content": "<p>Great!</p>",
      "rawMarkdown": "Great!",
      "votes": 1
    },
    {
      "id": 902103,
      "postDate": "2020-06-26T00:01:52.773Z",
      "content": "<p>This will be incredibly useful for me, I was using a very impractical system to upload my tiles as Kaggle datasets. Thanks!</p>",
      "rawMarkdown": "This will be incredibly useful for me, I was using a very impractical system to upload my tiles as Kaggle datasets. Thanks!",
      "votes": 1
    },
    {
      "id": 902530,
      "postDate": "2020-06-26T07:47:30.147Z",
      "content": "<p>Nice solution. Thank you for sharing🙌 </p>",
      "rawMarkdown": "Nice solution. Thank you for sharing🙌 ",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 901731,
      "author_name": "Dustin",
      "author_url": "",
      "post_date": "2020-06-25T16:43:55.127000",
      "content": "<p><a href=\"/raghaw\">@raghaw</a> Just a heads up, you don't actually have access to those 2TB of disk space. The 4TB partition is a read-only disk which is part of the Kaggle environment.</p>\n\n<p>When you write to /tmp you're actually writing to a writable disk which is overlayed on top of the 4TB one. </p>\n\n<p>The writeable space is quite large (though not guaranteed since we may use that space), so it's totally fine to write 30GB of space there, and I often suggest users use that scratch space anytime they need. The only reason we impose a 5GB limit on /kaggle/working, is because when you Save &gt; Run All, or Save &gt; Quick Save + outputs, we upload everything in that directory and save it for viewing/download in your Notebook Version, and put a limit on how much you can save per notebook version.</p>",
      "votes": 9,
      "replies": [
        {
          "id": 901746,
          "author_name": "Raghawendra Singh",
          "author_url": "",
          "post_date": "2020-06-25T17:01:16.460000",
          "content": "<p>Thanks <a href=\"/herbison\">@herbison</a> for clarification.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 900943,
      "author_name": "Mobassir",
      "author_url": "",
      "post_date": "2020-06-25T06:49:37.790000",
      "content": "<p>offline package installation in : pneumothorax,severstal steel defect detection then retraining strategy in bengali.ai and now in pandas effective dataset creation!!\nyou are awesome, i have been learning a lot from you in almost all computer vision competition,thank you for everything,looking forward to see some unique work like this in melanoma competition as well &lt;3 </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 902136,
      "author_name": "Rashidul Hasan Hridoy",
      "author_url": "",
      "post_date": "2020-06-26T00:41:15.053000",
      "content": "<p>Great!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 902103,
      "author_name": "Pasquale",
      "author_url": "",
      "post_date": "2020-06-26T00:01:52.773000",
      "content": "<p>This will be incredibly useful for me, I was using a very impractical system to upload my tiles as Kaggle datasets. Thanks!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 902530,
      "author_name": "Karan",
      "author_url": "",
      "post_date": "2020-06-26T07:47:30.147000",
      "content": "<p>Nice solution. Thank you for sharing🙌 </p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "900824": "Since Kaggle working directory allows up to 4GB data, it is very difficult to create medium resolution dataset having total size of 30GB(25 tiles per image with size of 256 each tile). I found a way to do so using kaggle notebook.\n\nBy running linux command !df -h, I found that I can use upto 2.1 TB on \"/\" mount point. Therefore I used \"/tmp\" directory to store and upload data.\n\nFilesystem        Size  Used Avail Use% Mounted on\noverlay           4.0T  2.0T  2.1T  50% /\ntmpfs              64M     0   64M   0% /dev\ntmpfs             9.4G     0  9.4G   0% /sys/fs/cgroup\nshm               7.0G     0  7.0G   0% /dev/shm\n/dev/mapper/snap  4.0T  2.0T  2.1T  50% /home\n/dev/loop1        4.9G   21M  4.9G   1% /kaggle/lib\ntmpfs             9.4G     0  9.4G   0% /proc/acpi\ntmpfs             9.4G     0  9.4G   0% /proc/scsi\ntmpfs             9.4G     0  9.4G   0% /sys/firmware\n\nplease have a look.....\n\nNotebook URL:\nhttps://www.kaggle.com/raghaw/panda-medium-resolution-dataset-25x256x256\n\nCreated Dataset URL:\nhttps://www.kaggle.com/raghaw/panda-dataset-medium-25-256-256",
    "901731": "@raghaw Just a heads up, you don't actually have access to those 2TB of disk space. The 4TB partition is a read-only disk which is part of the Kaggle environment.\n\nWhen you write to /tmp you're actually writing to a writable disk which is overlayed on top of the 4TB one. \n\nThe writeable space is quite large (though not guaranteed since we may use that space), so it's totally fine to write 30GB of space there, and I often suggest users use that scratch space anytime they need. The only reason we impose a 5GB limit on /kaggle/working, is because when you Save &gt; Run All, or Save &gt; Quick Save + outputs, we upload everything in that directory and save it for viewing/download in your Notebook Version, and put a limit on how much you can save per notebook version.",
    "900943": "offline package installation in : pneumothorax,severstal steel defect detection then retraining strategy in bengali.ai and now in pandas effective dataset creation!!\nyou are awesome, i have been learning a lot from you in almost all computer vision competition,thank you for everything,looking forward to see some unique work like this in melanoma competition as well &lt;3 ",
    "902136": "Great!",
    "902103": "This will be incredibly useful for me, I was using a very impractical system to upload my tiles as Kaggle datasets. Thanks!",
    "902530": "Nice solution. Thank you for sharing🙌 "
  }
}