{
  "id": 111029,
  "title": "Quickstart: Guide to Setup GCP",
  "url": "/competitions/rsna-intracranial-hemorrhage-detection/discussion/111029",
  "author_name": "Tim Yee",
  "post_date": "2019-10-03T04:29:46.873000",
  "votes": 34,
  "comment_count": 17,
  "views": 0,
  "content": "<p><strong>Quick summary of how I got GCP vm up and running</strong> (these are not steps, but guidelines):\n1. Setup billing, add GCP credits\n2. Download &amp; install <a href=\"https://cloud.google.com/sdk/docs/downloads-interactive\">gcloud SDK</a>\n3. Request quota increase - under Metrics, select GPU all regions and request to increase to 1 (step 5 in the medium article linked below) - wait for quota increase e-mail first, then continue\n4. Setup your deep learning VM with pytorch or tensorflow - make sure to select <em>preemptible</em> to save costs - $0.177/hr for a 4vcpu, 15GB RAM, and 1 K80 GPU (so that you're not losing too much while you are just trying to get setup)\n5. Start your instance \n6. Open gcloud console that you just installed - type in <code>gcloud init</code>, follow the prompts until there are no more, then type <code>gcloud compute ssh [instance name] -- -L 8080:localhost:8080</code> (step 10 - method 2 in medium article below)\n7. You should now be ssh'ed into your VM - once you reach step 14, follow my link to jupyter lab below instead.\n8. Connect to <a href=\"https://cloud.google.com/deep-learning-vm/docs/jupyter\">jupyter lab</a> - type <code>http://localhost:8080</code> into your browser (step 2 only)</p>\n\n<p>Now that you setup the VM a ssh connection, you'll need to download data into your VM's hard disk to train and test your model. The two main methods are gsutil and kaggle api</p>\n\n<ol>\n<li>Setup Google Cloud Storage using gsutil - if you've installed gcloud SDK above, you should already have the gsutil installed. Go into google cloud platform and find the storage tab and create a bucket (preferrably with the same region and preferrably same location for fastest download times to your VM). your bucket will be named gs://[bucket name]. You can now go into your local machine's terminal or command prompt and cd into the drive that holds the files you want to upload into the bucket. Type <code>gsutil -m cp foldername(should be zipped) gs://[bucket name]</code>. -m for multi-core and faster file transfer. Once the file is uploaded into your bucket, turn on your VM, go into your PuTTY ssh window and type <code>gsutil -m cp gs://[bucket name]/[folder name] /home/jupyter</code> and you should begin seeing the file transferred over. If not, open up jupyterlab and type the <code>!gsutil -m cp gs://[bucket name]/[folder name] /home/jupyter</code> into your jupyter notebook.</li>\n<li>Setup Kaggle API - first get your kaggle.json file downloaded onto your local machine. Open up the file and copy the contents of the json file. Then in your open PuTTY window, type in <code>sudo su</code> to get root access. type <code>cd ../jupyter/</code> then type <code>nano kaggle.json</code> and paste in the contents of the kaggle.json file you just copied (right click to paste). Press CTRL + O to save, then CTR + X to exit. Now type <code>mkdir .kaggle/</code> followed by <code>mv kaggle.json .kaggle/</code> You can simultaneously install kaggle api by typing !pip install kaggle into your jupyter notebook, once installed type <code>!kaggle competitions list</code> should populate and you should be good to go. <em>You do not need to type in any chmod 600 commands</em>. If you need access to .png image files, I created various resolutions <a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/110840#\">here</a>.</li>\n<li>You can also upload files locally into jupyter lab - see attached image. However, this option is substantially slower then options 1 &amp; 2 above and may have a file size limit. You can use this option to upload small .csv .ipynb files.</li>\n</ol>\n\n<p><a href=\"https://medium.com/&lt;a href=\">@howkhang</a>/ultimate-guide-to-setting-up-a-google-cloud-machine-for-fast-ai-version-2-f374208be43\"&gt;Medium Article - parts of this article are (deprecated) towards the bottom half. </p>\n\n<p><em>side remarks - in your putty terminal, you can type <code>htop</code> to monitor your VM's resources, you can always increase your CPU/RAM or switch GPU's (if available)</em></p>\n\n<p><em>SSD is noticeable faster if you're downloading/unzipping (writing files) to your persistent disk. Unzipping files take 1.55-1.6x longer on HDD</em></p>\n\n<p><strong>IF and when your jupyter kernel dies and cannot restart while using <a href=\"http://localhost:8080\">http://localhost:8080</a>, try  <a href=\"http://127.0.0.1:8080\">http://127.0.0.1:8080</a> - it has worked for me so far.</strong></p>\n\n<p><em>Occasionally, you can try clearing your browser cache, it may or may not help</em></p>\n\n<p>I know this can and will be super confusing. There's no straight forward way to explain or get started without a couple kinks. Please let me know if you have questions or if you get stuck somewhere.</p>",
  "messages": [
    {
      "id": 639323,
      "postDate": "2019-10-03T04:29:46.873Z",
      "content": "<p><strong>Quick summary of how I got GCP vm up and running</strong> (these are not steps, but guidelines):\n1. Setup billing, add GCP credits\n2. Download &amp; install <a href=\"https://cloud.google.com/sdk/docs/downloads-interactive\">gcloud SDK</a>\n3. Request quota increase - under Metrics, select GPU all regions and request to increase to 1 (step 5 in the medium article linked below) - wait for quota increase e-mail first, then continue\n4. Setup your deep learning VM with pytorch or tensorflow - make sure to select <em>preemptible</em> to save costs - $0.177/hr for a 4vcpu, 15GB RAM, and 1 K80 GPU (so that you're not losing too much while you are just trying to get setup)\n5. Start your instance \n6. Open gcloud console that you just installed - type in <code>gcloud init</code>, follow the prompts until there are no more, then type <code>gcloud compute ssh [instance name] -- -L 8080:localhost:8080</code> (step 10 - method 2 in medium article below)\n7. You should now be ssh'ed into your VM - once you reach step 14, follow my link to jupyter lab below instead.\n8. Connect to <a href=\"https://cloud.google.com/deep-learning-vm/docs/jupyter\">jupyter lab</a> - type <code>http://localhost:8080</code> into your browser (step 2 only)</p>\n\n<p>Now that you setup the VM a ssh connection, you'll need to download data into your VM's hard disk to train and test your model. The two main methods are gsutil and kaggle api</p>\n\n<ol>\n<li>Setup Google Cloud Storage using gsutil - if you've installed gcloud SDK above, you should already have the gsutil installed. Go into google cloud platform and find the storage tab and create a bucket (preferrably with the same region and preferrably same location for fastest download times to your VM). your bucket will be named gs://[bucket name]. You can now go into your local machine's terminal or command prompt and cd into the drive that holds the files you want to upload into the bucket. Type <code>gsutil -m cp foldername(should be zipped) gs://[bucket name]</code>. -m for multi-core and faster file transfer. Once the file is uploaded into your bucket, turn on your VM, go into your PuTTY ssh window and type <code>gsutil -m cp gs://[bucket name]/[folder name] /home/jupyter</code> and you should begin seeing the file transferred over. If not, open up jupyterlab and type the <code>!gsutil -m cp gs://[bucket name]/[folder name] /home/jupyter</code> into your jupyter notebook.</li>\n<li>Setup Kaggle API - first get your kaggle.json file downloaded onto your local machine. Open up the file and copy the contents of the json file. Then in your open PuTTY window, type in <code>sudo su</code> to get root access. type <code>cd ../jupyter/</code> then type <code>nano kaggle.json</code> and paste in the contents of the kaggle.json file you just copied (right click to paste). Press CTRL + O to save, then CTR + X to exit. Now type <code>mkdir .kaggle/</code> followed by <code>mv kaggle.json .kaggle/</code> You can simultaneously install kaggle api by typing !pip install kaggle into your jupyter notebook, once installed type <code>!kaggle competitions list</code> should populate and you should be good to go. <em>You do not need to type in any chmod 600 commands</em>. If you need access to .png image files, I created various resolutions <a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/110840#\">here</a>.</li>\n<li>You can also upload files locally into jupyter lab - see attached image. However, this option is substantially slower then options 1 &amp; 2 above and may have a file size limit. You can use this option to upload small .csv .ipynb files.</li>\n</ol>\n\n<p><a href=\"https://medium.com/&lt;a href=\">@howkhang</a>/ultimate-guide-to-setting-up-a-google-cloud-machine-for-fast-ai-version-2-f374208be43\"&gt;Medium Article - parts of this article are (deprecated) towards the bottom half. </p>\n\n<p><em>side remarks - in your putty terminal, you can type <code>htop</code> to monitor your VM's resources, you can always increase your CPU/RAM or switch GPU's (if available)</em></p>\n\n<p><em>SSD is noticeable faster if you're downloading/unzipping (writing files) to your persistent disk. Unzipping files take 1.55-1.6x longer on HDD</em></p>\n\n<p><strong>IF and when your jupyter kernel dies and cannot restart while using <a href=\"http://localhost:8080\">http://localhost:8080</a>, try  <a href=\"http://127.0.0.1:8080\">http://127.0.0.1:8080</a> - it has worked for me so far.</strong></p>\n\n<p><em>Occasionally, you can try clearing your browser cache, it may or may not help</em></p>\n\n<p>I know this can and will be super confusing. There's no straight forward way to explain or get started without a couple kinks. Please let me know if you have questions or if you get stuck somewhere.</p>",
      "rawMarkdown": "**Quick summary of how I got GCP vm up and running** (these are not steps, but guidelines):\n1. Setup billing, add GCP credits\n2. Download &amp; install [gcloud SDK](https://cloud.google.com/sdk/docs/downloads-interactive)\n3. Request quota increase - under Metrics, select GPU all regions and request to increase to 1 (step 5 in the medium article linked below) - wait for quota increase e-mail first, then continue\n4. Setup your deep learning VM with pytorch or tensorflow - make sure to select *preemptible* to save costs - $0.177/hr for a 4vcpu, 15GB RAM, and 1 K80 GPU (so that you're not losing too much while you are just trying to get setup)\n5. Start your instance \n6. Open gcloud console that you just installed - type in `gcloud init`, follow the prompts until there are no more, then type `gcloud compute ssh [instance name] -- -L 8080:localhost:8080` (step 10 - method 2 in medium article below)\n7. You should now be ssh'ed into your VM - once you reach step 14, follow my link to jupyter lab below instead.\n8. Connect to [jupyter lab](https://cloud.google.com/deep-learning-vm/docs/jupyter) - type `http://localhost:8080` into your browser (step 2 only)\n\nNow that you setup the VM a ssh connection, you'll need to download data into your VM's hard disk to train and test your model. The two main methods are gsutil and kaggle api\n\n1. Setup Google Cloud Storage using gsutil - if you've installed gcloud SDK above, you should already have the gsutil installed. Go into google cloud platform and find the storage tab and create a bucket (preferrably with the same region and preferrably same location for fastest download times to your VM). your bucket will be named gs://[bucket name]. You can now go into your local machine's terminal or command prompt and cd into the drive that holds the files you want to upload into the bucket. Type `gsutil -m cp foldername(should be zipped) gs://[bucket name]`. -m for multi-core and faster file transfer. Once the file is uploaded into your bucket, turn on your VM, go into your PuTTY ssh window and type `gsutil -m cp gs://[bucket name]/[folder name] /home/jupyter` and you should begin seeing the file transferred over. If not, open up jupyterlab and type the `!gsutil -m cp gs://[bucket name]/[folder name] /home/jupyter` into your jupyter notebook.\n2. Setup Kaggle API - first get your kaggle.json file downloaded onto your local machine. Open up the file and copy the contents of the json file. Then in your open PuTTY window, type in `sudo su` to get root access. type `cd ../jupyter/` then type `nano kaggle.json` and paste in the contents of the kaggle.json file you just copied (right click to paste). Press CTRL + O to save, then CTR + X to exit. Now type `mkdir .kaggle/` followed by `mv kaggle.json .kaggle/ ` You can simultaneously install kaggle api by typing !pip install kaggle into your jupyter notebook, once installed type `!kaggle competitions list` should populate and you should be good to go. *You do not need to type in any chmod 600 commands*. If you need access to .png image files, I created various resolutions [here](https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/110840#).\n3. You can also upload files locally into jupyter lab - see attached image. However, this option is substantially slower then options 1 &amp; 2 above and may have a file size limit. You can use this option to upload small .csv .ipynb files.\n\n[Medium Article](https://medium.com/@howkhang/ultimate-guide-to-setting-up-a-google-cloud-machine-for-fast-ai-version-2-f374208be43) - parts of this article are (deprecated) towards the bottom half. \n\n*side remarks - in your putty terminal, you can type `htop` to monitor your VM's resources, you can always increase your CPU/RAM or switch GPU's (if available)*\n\n*SSD is noticeable faster if you're downloading/unzipping (writing files) to your persistent disk. Unzipping files take 1.55-1.6x longer on HDD*\n\n__IF and when your jupyter kernel dies and cannot restart while using http://localhost:8080, try  http://127.0.0.1:8080 - it has worked for me so far.__\n\n*Occasionally, you can try clearing your browser cache, it may or may not help*\n\nI know this can and will be super confusing. There's no straight forward way to explain or get started without a couple kinks. Please let me know if you have questions or if you get stuck somewhere.",
      "votes": 33
    },
    {
      "id": 639768,
      "postDate": "2019-10-03T14:30:26.297Z",
      "content": "<p>Thank you for the guide!</p>\n\n<p>I just want to note that step 3 took 6 hours for me, so do it as soon as possible.​</p>",
      "rawMarkdown": "Thank you for the guide!\n\nI just want to note that step 3 took 6 hours for me, so do it as soon as possible.​",
      "votes": 1,
      "replies": [
        {
          "id": 639904,
          "postDate": "2019-10-03T17:21:12.817Z",
          "content": "<p>Affirmative.</p>",
          "rawMarkdown": "Affirmative."
        }
      ]
    },
    {
      "id": 640644,
      "postDate": "2019-10-04T06:27:37.283Z",
      "content": "<p>Thanks a lot for providing this information <a href=\"/teeyee314\">@teeyee314</a> .. Saved lot of my time. </p>",
      "rawMarkdown": "Thanks a lot for providing this information @teeyee314 .. Saved lot of my time. ",
      "votes": 2
    },
    {
      "id": 639436,
      "postDate": "2019-10-03T07:51:05.240Z",
      "content": "<p>There is a very helpful guide on getting started with GCP from fast.ai <a href=\"https://course.fast.ai/start_gcp.html\">https://course.fast.ai/start_gcp.html</a>. Even if you don't want to use the fast.ai library and PyTorch this is really useful.</p>\n\n<p>Regarding your permissions problem, all I can suggest is \"User installs are strongly recommended in the case of permissions errors.\" from <a href=\"https://github.com/Kaggle/kaggle-api\">https://github.com/Kaggle/kaggle-api</a>.</p>",
      "rawMarkdown": "There is a very helpful guide on getting started with GCP from fast.ai [https://course.fast.ai/start_gcp.html](https://course.fast.ai/start_gcp.html). Even if you don't want to use the fast.ai library and PyTorch this is really useful.\n\nRegarding your permissions problem, all I can suggest is \"User installs are strongly recommended in the case of permissions errors.\" from [https://github.com/Kaggle/kaggle-api](https://github.com/Kaggle/kaggle-api).",
      "votes": 2,
      "replies": [
        {
          "id": 639538,
          "postDate": "2019-10-03T10:14:03.507Z",
          "content": "<p>Thanks for posting the fast.ai tutorial. I'm sure others may find it helpful.</p>",
          "rawMarkdown": "Thanks for posting the fast.ai tutorial. I'm sure others may find it helpful."
        }
      ]
    },
    {
      "id": 991291,
      "postDate": "2020-08-30T09:49:29.280Z",
      "content": "<p>Its grate </p>",
      "rawMarkdown": "Its grate "
    },
    {
      "id": 644615,
      "postDate": "2019-10-09T04:52:39.727Z",
      "content": "<p>Is there way to load the data directly into a bucket from kaggle?</p>",
      "rawMarkdown": "Is there way to load the data directly into a bucket from kaggle?",
      "replies": [
        {
          "id": 644816,
          "postDate": "2019-10-09T11:41:58.783Z",
          "content": "<p>Not that I know of. Download either locally and then upload into bucket using <code>gsutil -m cp foldername.zip gs://bucketname</code> - make sure you're in the directory of the zipped files you want to upload. or using kaggle api, download the files you want onto your VM and then upload them(or whatever pre-processed versions you need) into your bucket using the gsutil commands. the same command above should work from the VM as if you were on a local machine. Just remember to create a bucket in google cloud storage first.</p>",
          "rawMarkdown": "Not that I know of. Download either locally and then upload into bucket using `gsutil -m cp foldername.zip gs://bucketname` - make sure you're in the directory of the zipped files you want to upload. or using kaggle api, download the files you want onto your VM and then upload them(or whatever pre-processed versions you need) into your bucket using the gsutil commands. the same command above should work from the VM as if you were on a local machine. Just remember to create a bucket in google cloud storage first."
        }
      ]
    },
    {
      "id": 643817,
      "postDate": "2019-10-08T00:52:38.663Z",
      "content": "<p>Thanks for putting this together Tim. I was using kaggle API to download files in GCP jupyter lab, but it runs out of memory when I unzip. There was space on the hard disk, though. Has anyone else had this problem?</p>",
      "rawMarkdown": "Thanks for putting this together Tim. I was using kaggle API to download files in GCP jupyter lab, but it runs out of memory when I unzip. There was space on the hard disk, though. Has anyone else had this problem?",
      "replies": [
        {
          "id": 643824,
          "postDate": "2019-10-08T01:00:50.613Z",
          "content": "<p><a href=\"/twopift\">@twopift</a>  which files, mine? competition data? and how are you unzipping? how large is your vm's hdd? please be specific.</p>",
          "rawMarkdown": "@twopift  which files, mine? competition data? and how are you unzipping? how large is your vm's hdd? please be specific."
        },
        {
          "id": 643828,
          "postDate": "2019-10-08T01:08:46.153Z",
          "content": "<p>I made a 400 GB hard disk VM and downloaded the dataset zip file (~150 GB) using the kaggle API. So far no problem. Then I made a notebook and did !unzip file.zip and it started unzipping (took a long time) and then eventually gave an out of memory error. At first I thought that I had run out of hard disk space, but this was not the case. After that I was not able to get back into the notebook.</p>",
          "rawMarkdown": "I made a 400 GB hard disk VM and downloaded the dataset zip file (~150 GB) using the kaggle API. So far no problem. Then I made a notebook and did !unzip file.zip and it started unzipping (took a long time) and then eventually gave an out of memory error. At first I thought that I had run out of hard disk space, but this was not the case. After that I was not able to get back into the notebook."
        },
        {
          "id": 643829,
          "postDate": "2019-10-08T01:09:52.257Z",
          "content": "<p>The 150GB unzips to ~370GB so therefore the 400GB won't be enough. (150+370 &gt; 520GB)</p>",
          "rawMarkdown": "The 150GB unzips to ~370GB so therefore the 400GB won't be enough. (150+370 &gt; 520GB)"
        },
        {
          "id": 643830,
          "postDate": "2019-10-08T01:14:07.603Z",
          "content": "<p>Thanks Tim! I will make a bigger machine</p>",
          "rawMarkdown": "Thanks Tim! I will make a bigger machine"
        },
        {
          "id": 643832,
          "postDate": "2019-10-08T01:20:35.043Z",
          "content": "<p>Depends on what you want to do. Are you going to store any pre-processed jpgs or pngs locally? and at what resolution sizes? If you just want to read the 370GB dicom files every time, then 600GB would be fine. You can always delete the zipped files when you're done unzipping freeing yourself of 150GB.</p>",
          "rawMarkdown": "Depends on what you want to do. Are you going to store any pre-processed jpgs or pngs locally? and at what resolution sizes? If you just want to read the 370GB dicom files every time, then 600GB would be fine. You can always delete the zipped files when you're done unzipping freeing yourself of 150GB."
        },
        {
          "id": 643837,
          "postDate": "2019-10-08T01:31:33.877Z",
          "content": "<p>Thanks for your help Tim. From what you are saying I can see that it may be better to preprocess the dicom files and then save as png/jpg to feed as training data and that may make things faster and easier, so I should allow space for that, but the 150 GB saved from deleting the zip file should cover this I suppose. </p>",
          "rawMarkdown": "Thanks for your help Tim. From what you are saying I can see that it may be better to preprocess the dicom files and then save as png/jpg to feed as training data and that may make things faster and easier, so I should allow space for that, but the 150 GB saved from deleting the zip file should cover this I suppose. "
        },
        {
          "id": 643849,
          "postDate": "2019-10-08T02:08:53.967Z",
          "content": "<p>Yeah Josh, there is no straight forward answer here. I cannot say reading from dicoms into a dataloader on the fly is better or worse than pre-processing, and saving only pixel array data locally is. There are trade-offs. Just like HDD is cheaper per GB than SSD, but SSD is clearly faster.</p>",
          "rawMarkdown": "Yeah Josh, there is no straight forward answer here. I cannot say reading from dicoms into a dataloader on the fly is better or worse than pre-processing, and saving only pixel array data locally is. There are trade-offs. Just like HDD is cheaper per GB than SSD, but SSD is clearly faster."
        }
      ]
    },
    {
      "id": 1200564,
      "postDate": "2021-02-14T18:50:06.897Z",
      "content": "<p>Thanks a lot !!! nicely detailed 👍</p>",
      "rawMarkdown": "Thanks a lot !!! nicely detailed 👍"
    }
  ],
  "comments": [
    {
      "id": 639768,
      "author_name": "Alexander Abstreiter",
      "author_url": "",
      "post_date": "2019-10-03T14:30:26.297000",
      "content": "<p>Thank you for the guide!</p>\n\n<p>I just want to note that step 3 took 6 hours for me, so do it as soon as possible.​</p>",
      "votes": 1,
      "replies": [
        {
          "id": 639904,
          "author_name": "Tim Yee",
          "author_url": "",
          "post_date": "2019-10-03T17:21:12.817000",
          "content": "<p>Affirmative.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 640644,
      "author_name": "Manoj Prabhakar",
      "author_url": "",
      "post_date": "2019-10-04T06:27:37.283000",
      "content": "<p>Thanks a lot for providing this information <a href=\"/teeyee314\">@teeyee314</a> .. Saved lot of my time. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 639436,
      "author_name": "Alison Davey",
      "author_url": "",
      "post_date": "2019-10-03T07:51:05.240000",
      "content": "<p>There is a very helpful guide on getting started with GCP from fast.ai <a href=\"https://course.fast.ai/start_gcp.html\">https://course.fast.ai/start_gcp.html</a>. Even if you don't want to use the fast.ai library and PyTorch this is really useful.</p>\n\n<p>Regarding your permissions problem, all I can suggest is \"User installs are strongly recommended in the case of permissions errors.\" from <a href=\"https://github.com/Kaggle/kaggle-api\">https://github.com/Kaggle/kaggle-api</a>.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 639538,
          "author_name": "Tim Yee",
          "author_url": "",
          "post_date": "2019-10-03T10:14:03.507000",
          "content": "<p>Thanks for posting the fast.ai tutorial. I'm sure others may find it helpful.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 991291,
      "author_name": "SAHASRANSU KAR",
      "author_url": "",
      "post_date": "2020-08-30T09:49:29.280000",
      "content": "<p>Its grate </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 644615,
      "author_name": "rego-luna",
      "author_url": "",
      "post_date": "2019-10-09T04:52:39.727000",
      "content": "<p>Is there way to load the data directly into a bucket from kaggle?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 644816,
          "author_name": "Tim Yee",
          "author_url": "",
          "post_date": "2019-10-09T11:41:58.783000",
          "content": "<p>Not that I know of. Download either locally and then upload into bucket using <code>gsutil -m cp foldername.zip gs://bucketname</code> - make sure you're in the directory of the zipped files you want to upload. or using kaggle api, download the files you want onto your VM and then upload them(or whatever pre-processed versions you need) into your bucket using the gsutil commands. the same command above should work from the VM as if you were on a local machine. Just remember to create a bucket in google cloud storage first.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 643817,
      "author_name": "Josh Myers",
      "author_url": "",
      "post_date": "2019-10-08T00:52:38.663000",
      "content": "<p>Thanks for putting this together Tim. I was using kaggle API to download files in GCP jupyter lab, but it runs out of memory when I unzip. There was space on the hard disk, though. Has anyone else had this problem?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 643824,
          "author_name": "Tim Yee",
          "author_url": "",
          "post_date": "2019-10-08T01:00:50.613000",
          "content": "<p><a href=\"/twopift\">@twopift</a>  which files, mine? competition data? and how are you unzipping? how large is your vm's hdd? please be specific.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 643828,
          "author_name": "Josh Myers",
          "author_url": "",
          "post_date": "2019-10-08T01:08:46.153000",
          "content": "<p>I made a 400 GB hard disk VM and downloaded the dataset zip file (~150 GB) using the kaggle API. So far no problem. Then I made a notebook and did !unzip file.zip and it started unzipping (took a long time) and then eventually gave an out of memory error. At first I thought that I had run out of hard disk space, but this was not the case. After that I was not able to get back into the notebook.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 643829,
          "author_name": "Tim Yee",
          "author_url": "",
          "post_date": "2019-10-08T01:09:52.257000",
          "content": "<p>The 150GB unzips to ~370GB so therefore the 400GB won't be enough. (150+370 &gt; 520GB)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 643830,
          "author_name": "Josh Myers",
          "author_url": "",
          "post_date": "2019-10-08T01:14:07.603000",
          "content": "<p>Thanks Tim! I will make a bigger machine</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 643832,
          "author_name": "Tim Yee",
          "author_url": "",
          "post_date": "2019-10-08T01:20:35.043000",
          "content": "<p>Depends on what you want to do. Are you going to store any pre-processed jpgs or pngs locally? and at what resolution sizes? If you just want to read the 370GB dicom files every time, then 600GB would be fine. You can always delete the zipped files when you're done unzipping freeing yourself of 150GB.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 643837,
          "author_name": "Josh Myers",
          "author_url": "",
          "post_date": "2019-10-08T01:31:33.877000",
          "content": "<p>Thanks for your help Tim. From what you are saying I can see that it may be better to preprocess the dicom files and then save as png/jpg to feed as training data and that may make things faster and easier, so I should allow space for that, but the 150 GB saved from deleting the zip file should cover this I suppose. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 643849,
          "author_name": "Tim Yee",
          "author_url": "",
          "post_date": "2019-10-08T02:08:53.967000",
          "content": "<p>Yeah Josh, there is no straight forward answer here. I cannot say reading from dicoms into a dataloader on the fly is better or worse than pre-processing, and saving only pixel array data locally is. There are trade-offs. Just like HDD is cheaper per GB than SSD, but SSD is clearly faster.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1200564,
      "author_name": "Tarik Moumini",
      "author_url": "",
      "post_date": "2021-02-14T18:50:06.897000",
      "content": "<p>Thanks a lot !!! nicely detailed 👍</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "639323": "**Quick summary of how I got GCP vm up and running** (these are not steps, but guidelines):\n1. Setup billing, add GCP credits\n2. Download &amp; install [gcloud SDK](https://cloud.google.com/sdk/docs/downloads-interactive)\n3. Request quota increase - under Metrics, select GPU all regions and request to increase to 1 (step 5 in the medium article linked below) - wait for quota increase e-mail first, then continue\n4. Setup your deep learning VM with pytorch or tensorflow - make sure to select *preemptible* to save costs - $0.177/hr for a 4vcpu, 15GB RAM, and 1 K80 GPU (so that you're not losing too much while you are just trying to get setup)\n5. Start your instance \n6. Open gcloud console that you just installed - type in `gcloud init`, follow the prompts until there are no more, then type `gcloud compute ssh [instance name] -- -L 8080:localhost:8080` (step 10 - method 2 in medium article below)\n7. You should now be ssh'ed into your VM - once you reach step 14, follow my link to jupyter lab below instead.\n8. Connect to [jupyter lab](https://cloud.google.com/deep-learning-vm/docs/jupyter) - type `http://localhost:8080` into your browser (step 2 only)\n\nNow that you setup the VM a ssh connection, you'll need to download data into your VM's hard disk to train and test your model. The two main methods are gsutil and kaggle api\n\n1. Setup Google Cloud Storage using gsutil - if you've installed gcloud SDK above, you should already have the gsutil installed. Go into google cloud platform and find the storage tab and create a bucket (preferrably with the same region and preferrably same location for fastest download times to your VM). your bucket will be named gs://[bucket name]. You can now go into your local machine's terminal or command prompt and cd into the drive that holds the files you want to upload into the bucket. Type `gsutil -m cp foldername(should be zipped) gs://[bucket name]`. -m for multi-core and faster file transfer. Once the file is uploaded into your bucket, turn on your VM, go into your PuTTY ssh window and type `gsutil -m cp gs://[bucket name]/[folder name] /home/jupyter` and you should begin seeing the file transferred over. If not, open up jupyterlab and type the `!gsutil -m cp gs://[bucket name]/[folder name] /home/jupyter` into your jupyter notebook.\n2. Setup Kaggle API - first get your kaggle.json file downloaded onto your local machine. Open up the file and copy the contents of the json file. Then in your open PuTTY window, type in `sudo su` to get root access. type `cd ../jupyter/` then type `nano kaggle.json` and paste in the contents of the kaggle.json file you just copied (right click to paste). Press CTRL + O to save, then CTR + X to exit. Now type `mkdir .kaggle/` followed by `mv kaggle.json .kaggle/ ` You can simultaneously install kaggle api by typing !pip install kaggle into your jupyter notebook, once installed type `!kaggle competitions list` should populate and you should be good to go. *You do not need to type in any chmod 600 commands*. If you need access to .png image files, I created various resolutions [here](https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/110840#).\n3. You can also upload files locally into jupyter lab - see attached image. However, this option is substantially slower then options 1 &amp; 2 above and may have a file size limit. You can use this option to upload small .csv .ipynb files.\n\n[Medium Article](https://medium.com/@howkhang/ultimate-guide-to-setting-up-a-google-cloud-machine-for-fast-ai-version-2-f374208be43) - parts of this article are (deprecated) towards the bottom half. \n\n*side remarks - in your putty terminal, you can type `htop` to monitor your VM's resources, you can always increase your CPU/RAM or switch GPU's (if available)*\n\n*SSD is noticeable faster if you're downloading/unzipping (writing files) to your persistent disk. Unzipping files take 1.55-1.6x longer on HDD*\n\n__IF and when your jupyter kernel dies and cannot restart while using http://localhost:8080, try  http://127.0.0.1:8080 - it has worked for me so far.__\n\n*Occasionally, you can try clearing your browser cache, it may or may not help*\n\nI know this can and will be super confusing. There's no straight forward way to explain or get started without a couple kinks. Please let me know if you have questions or if you get stuck somewhere.",
    "639768": "Thank you for the guide!\n\nI just want to note that step 3 took 6 hours for me, so do it as soon as possible.​",
    "640644": "Thanks a lot for providing this information @teeyee314 .. Saved lot of my time. ",
    "639436": "There is a very helpful guide on getting started with GCP from fast.ai [https://course.fast.ai/start_gcp.html](https://course.fast.ai/start_gcp.html). Even if you don't want to use the fast.ai library and PyTorch this is really useful.\n\nRegarding your permissions problem, all I can suggest is \"User installs are strongly recommended in the case of permissions errors.\" from [https://github.com/Kaggle/kaggle-api](https://github.com/Kaggle/kaggle-api).",
    "991291": "Its grate ",
    "644615": "Is there way to load the data directly into a bucket from kaggle?",
    "643817": "Thanks for putting this together Tim. I was using kaggle API to download files in GCP jupyter lab, but it runs out of memory when I unzip. There was space on the hard disk, though. Has anyone else had this problem?",
    "1200564": "Thanks a lot !!! nicely detailed 👍"
  }
}