{
  "id": 109925,
  "title": "Server / VM Recommendation to work on Challenge",
  "url": "/competitions/rsna-intracranial-hemorrhage-detection/discussion/109925",
  "author_name": "RaMa92",
  "post_date": "2019-09-23T14:00:26.734000",
  "votes": 2,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hello, any recommendations on affordable VM/ Server configurations from cloud-service providers to efficiently start working on this challenge? </p>",
  "messages": [
    {
      "id": 632482,
      "postDate": "2019-09-23T16:49:45Z",
      "content": "<p>Hi! <a href=\"/raphael92\">@raphael92</a> I will recommend developing the model on a Small VM  and then train on a bigger machine in the Cloud, for example using Services like Google AI platform and ML Engine or Amazon Sage Maker, this way you only pay for the resources you need.</p>\n\n<p>AWS\n<a href=\"https://aws.amazon.com/sagemaker/\">https://aws.amazon.com/sagemaker/</a>\nGCP\n<a href=\"https://cloud.google.com/ai-platform/\">https://cloud.google.com/ai-platform/</a>\n<a href=\"https://cloud.google.com/ml-engine/\">https://cloud.google.com/ml-engine/</a></p>",
      "rawMarkdown": "Hi! @raphael92 I will recommend developing the model on a Small VM  and then train on a bigger machine in the Cloud, for example using Services like Google AI platform and ML Engine or Amazon Sage Maker, this way you only pay for the resources you need.\n\nAWS\nhttps://aws.amazon.com/sagemaker/\nGCP\nhttps://cloud.google.com/ai-platform/\nhttps://cloud.google.com/ml-engine/",
      "votes": 1,
      "replies": [
        {
          "id": 633245,
          "postDate": "2019-09-24T15:49:20.910Z",
          "content": "<p>Hi <a href=\"/cv13j0\">@cv13j0</a> , same thoughts that I had - is there a way of getting started with the competition on free resources (like Google Codelab, Kaggle Kernel). I believe not, as the data size (~150GB) seems the biggest challenge for beginners, but I would appreciate any hints before starting on SageMaker.</p>",
          "rawMarkdown": "Hi @cv13j0 , same thoughts that I had - is there a way of getting started with the competition on free resources (like Google Codelab, Kaggle Kernel). I believe not, as the data size (~150GB) seems the biggest challenge for beginners, but I would appreciate any hints before starting on SageMaker."
        }
      ]
    },
    {
      "id": 632394,
      "postDate": "2019-09-23T14:20:48.543Z",
      "content": "<p>I am fighting the whole day with GCP preemptible VMs, they keep preempting every 5, 10 or 30 mins, \nsince I migrated to a standard machine everything is fine.</p>\n\n<p>With those big data sets don't think about saving some pennies, because when you have to retrain 1 epoch from start, because of google decided to shut down your VM, you are wasting much more money.</p>\n\n<p>Waiting desperately for those GCP credits ;)</p>",
      "rawMarkdown": "I am fighting the whole day with GCP preemptible VMs, they keep preempting every 5, 10 or 30 mins, \nsince I migrated to a standard machine everything is fine.\n\nWith those big data sets don't think about saving some pennies, because when you have to retrain 1 epoch from start, because of google decided to shut down your VM, you are wasting much more money.\n\nWaiting desperately for those GCP credits ;)\n",
      "votes": 1,
      "replies": [
        {
          "id": 633428,
          "postDate": "2019-09-24T22:53:17.157Z",
          "content": "<p>What size hard disk do you need to unzip the data?</p>",
          "rawMarkdown": "What size hard disk do you need to unzip the data?"
        },
        {
          "id": 650547,
          "postDate": "2019-10-16T13:51:51.567Z",
          "content": "<p>I have ended with a 700GB disk size</p>",
          "rawMarkdown": "I have ended with a 700GB disk size",
          "votes": 1
        },
        {
          "id": 650662,
          "postDate": "2019-10-16T15:25:13.797Z",
          "content": "<p>I am just getting GCP set up - how do people use it mostly?</p>\n\n<p>Where is the best place to store the data? Do people usually just spin up a persistent disk to store their data for the life of the competition?\nI have just spun up a 600GB persistent disk and downloaded/unzipped the data there. </p>\n\n<p>What ML image do you use? I have been trying to use the Deep Learning VM in the market place however haven't been able to spin one up yet as there never seems to be resources available for it ..\nI am thinking I will just create one of the deep learning OS images and either install nvidia-docker on it or just install everything I need and make the disk persist so I can attach it to any VM</p>",
          "rawMarkdown": "I am just getting GCP set up - how do people use it mostly?\n\nWhere is the best place to store the data? Do people usually just spin up a persistent disk to store their data for the life of the competition?\nI have just spun up a 600GB persistent disk and downloaded/unzipped the data there. \n\nWhat ML image do you use? I have been trying to use the Deep Learning VM in the market place however haven't been able to spin one up yet as there never seems to be resources available for it ..\nI am thinking I will just create one of the deep learning OS images and either install nvidia-docker on it or just install everything I need and make the disk persist so I can attach it to any VM"
        },
        {
          "id": 650878,
          "postDate": "2019-10-16T19:31:44.987Z",
          "content": "<p><a href=\"/cherring\">@cherring</a> </p>\n\n<p>Here's my 2c:</p>\n\n<ul>\n<li>Use buckets to store your data, as this is cheap and designed for storage. When you boot your VM, you can copy it across super fast and easy using gsutil</li>\n<li>When creating your VM, look through boot disks to see if there is any to save some time. For example, I tend to use \"Deep Learning Image: PyTorch 1.2.0 and fastai m36\" for pytorch models</li>\n<li>Use pre-emptible machines but try different regions and zones. In some regions you will find it nearly impossible to keep pre-emptible access to e.g. a v-100 for more than 5 minutes at peak times</li>\n<li>Checkpoint models and write to your bucket, and do this regularly. You don't have to wait for a full epoch to do this, just take time to setup your generators and save every N batches</li>\n</ul>",
          "rawMarkdown": "@cherring \n\nHere's my 2c:\n\n- Use buckets to store your data, as this is cheap and designed for storage. When you boot your VM, you can copy it across super fast and easy using gsutil\n- When creating your VM, look through boot disks to see if there is any to save some time. For example, I tend to use \"Deep Learning Image: PyTorch 1.2.0 and fastai m36\" for pytorch models\n- Use pre-emptible machines but try different regions and zones. In some regions you will find it nearly impossible to keep pre-emptible access to e.g. a v-100 for more than 5 minutes at peak times\n- Checkpoint models and write to your bucket, and do this regularly. You don't have to wait for a full epoch to do this, just take time to setup your generators and save every N batches\n",
          "votes": 2
        },
        {
          "id": 650985,
          "postDate": "2019-10-16T23:07:17.633Z",
          "content": "<p>Thanks for the info. I will try using buckets next time. Are buckets available in all regions?\nIf so then I guess having data in a bucket really helps to be able to spin up VM's in different regions. Currently my persistent disk limits me to spinning up in only one region.</p>\n\n<p>I notice that the boot disks seem to have python 2.7. I have tried the TF1.14 ones. Do they all have python 2.7?</p>",
          "rawMarkdown": "Thanks for the info. I will try using buckets next time. Are buckets available in all regions?\nIf so then I guess having data in a bucket really helps to be able to spin up VM's in different regions. Currently my persistent disk limits me to spinning up in only one region.\n\nI notice that the boot disks seem to have python 2.7. I have tried the TF1.14 ones. Do they all have python 2.7?"
        }
      ]
    },
    {
      "id": 632372,
      "postDate": "2019-09-23T14:00:26.733Z",
      "content": "<p>Hello, any recommendations on affordable VM/ Server configurations from cloud-service providers to efficiently start working on this challenge? </p>",
      "rawMarkdown": "Hello, any recommendations on affordable VM/ Server configurations from cloud-service providers to efficiently start working on this challenge? ",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 632482,
      "author_name": "C4rl05/V",
      "author_url": "",
      "post_date": "2019-09-23T16:49:45",
      "content": "<p>Hi! <a href=\"/raphael92\">@raphael92</a> I will recommend developing the model on a Small VM  and then train on a bigger machine in the Cloud, for example using Services like Google AI platform and ML Engine or Amazon Sage Maker, this way you only pay for the resources you need.</p>\n\n<p>AWS\n<a href=\"https://aws.amazon.com/sagemaker/\">https://aws.amazon.com/sagemaker/</a>\nGCP\n<a href=\"https://cloud.google.com/ai-platform/\">https://cloud.google.com/ai-platform/</a>\n<a href=\"https://cloud.google.com/ml-engine/\">https://cloud.google.com/ml-engine/</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 633245,
          "author_name": "RaMa92",
          "author_url": "",
          "post_date": "2019-09-24T15:49:20.910000",
          "content": "<p>Hi <a href=\"/cv13j0\">@cv13j0</a> , same thoughts that I had - is there a way of getting started with the competition on free resources (like Google Codelab, Kaggle Kernel). I believe not, as the data size (~150GB) seems the biggest challenge for beginners, but I would appreciate any hints before starting on SageMaker.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 632394,
      "author_name": "Marc Mauri",
      "author_url": "",
      "post_date": "2019-09-23T14:20:48.543000",
      "content": "<p>I am fighting the whole day with GCP preemptible VMs, they keep preempting every 5, 10 or 30 mins, \nsince I migrated to a standard machine everything is fine.</p>\n\n<p>With those big data sets don't think about saving some pennies, because when you have to retrain 1 epoch from start, because of google decided to shut down your VM, you are wasting much more money.</p>\n\n<p>Waiting desperately for those GCP credits ;)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 633428,
          "author_name": "Josh Myers",
          "author_url": "",
          "post_date": "2019-09-24T22:53:17.157000",
          "content": "<p>What size hard disk do you need to unzip the data?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 650547,
          "author_name": "Marc Mauri",
          "author_url": "",
          "post_date": "2019-10-16T13:51:51.567000",
          "content": "<p>I have ended with a 700GB disk size</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 650662,
          "author_name": "cherring",
          "author_url": "",
          "post_date": "2019-10-16T15:25:13.797000",
          "content": "<p>I am just getting GCP set up - how do people use it mostly?</p>\n\n<p>Where is the best place to store the data? Do people usually just spin up a persistent disk to store their data for the life of the competition?\nI have just spun up a 600GB persistent disk and downloaded/unzipped the data there. </p>\n\n<p>What ML image do you use? I have been trying to use the Deep Learning VM in the market place however haven't been able to spin one up yet as there never seems to be resources available for it ..\nI am thinking I will just create one of the deep learning OS images and either install nvidia-docker on it or just install everything I need and make the disk persist so I can attach it to any VM</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 650878,
          "author_name": "Tom Aindow",
          "author_url": "",
          "post_date": "2019-10-16T19:31:44.987000",
          "content": "<p><a href=\"/cherring\">@cherring</a> </p>\n\n<p>Here's my 2c:</p>\n\n<ul>\n<li>Use buckets to store your data, as this is cheap and designed for storage. When you boot your VM, you can copy it across super fast and easy using gsutil</li>\n<li>When creating your VM, look through boot disks to see if there is any to save some time. For example, I tend to use \"Deep Learning Image: PyTorch 1.2.0 and fastai m36\" for pytorch models</li>\n<li>Use pre-emptible machines but try different regions and zones. In some regions you will find it nearly impossible to keep pre-emptible access to e.g. a v-100 for more than 5 minutes at peak times</li>\n<li>Checkpoint models and write to your bucket, and do this regularly. You don't have to wait for a full epoch to do this, just take time to setup your generators and save every N batches</li>\n</ul>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 650985,
          "author_name": "cherring",
          "author_url": "",
          "post_date": "2019-10-16T23:07:17.633000",
          "content": "<p>Thanks for the info. I will try using buckets next time. Are buckets available in all regions?\nIf so then I guess having data in a bucket really helps to be able to spin up VM's in different regions. Currently my persistent disk limits me to spinning up in only one region.</p>\n\n<p>I notice that the boot disks seem to have python 2.7. I have tried the TF1.14 ones. Do they all have python 2.7?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "632482": "Hi! @raphael92 I will recommend developing the model on a Small VM  and then train on a bigger machine in the Cloud, for example using Services like Google AI platform and ML Engine or Amazon Sage Maker, this way you only pay for the resources you need.\n\nAWS\nhttps://aws.amazon.com/sagemaker/\nGCP\nhttps://cloud.google.com/ai-platform/\nhttps://cloud.google.com/ml-engine/",
    "632394": "I am fighting the whole day with GCP preemptible VMs, they keep preempting every 5, 10 or 30 mins, \nsince I migrated to a standard machine everything is fine.\n\nWith those big data sets don't think about saving some pennies, because when you have to retrain 1 epoch from start, because of google decided to shut down your VM, you are wasting much more money.\n\nWaiting desperately for those GCP credits ;)\n",
    "632372": "Hello, any recommendations on affordable VM/ Server configurations from cloud-service providers to efficiently start working on this challenge? "
  }
}