{
  "id": 111565,
  "title": "GCP - Are all the V100 GPUs taken?",
  "url": "/competitions/rsna-intracranial-hemorrhage-detection/discussion/111565",
  "author_name": "Tim Yee",
  "post_date": "2019-10-07T00:41:50.125000",
  "votes": 2,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Is anyone else noticing that they are unable to attach the V100 GPU to their instances? I can attach other GPU's but the V100's across all the servers seem to be in use - both preempted and non-preempted instances. </p>\n\n<p>If you are able to attach a V100 GPU to your instance, please share the region, location and whether you're using preempted or not. Thanks.\n<br>\nEdit: I just tried the us-west1-a on 10/8 and it works - V100 preemptible. Thanks a lot guys!</p>",
  "messages": [
    {
      "id": 643620,
      "postDate": "2019-10-07T17:36:03.877Z",
      "content": "<p>Got a V100 on <code>us-west1-a</code>.</p>",
      "rawMarkdown": "Got a V100 on `us-west1-a`.",
      "votes": 1
    },
    {
      "id": 643034,
      "postDate": "2019-10-07T01:47:31.233Z",
      "content": "<p>I created a new instance with V100 yesterday. I tried several zones, but V100 wasn't available until I tried <code>us-west1-a</code>.</p>\n\n<ul>\n<li>n1-highmem-8</li>\n<li>us-west1-a</li>\n<li>preempted</li>\n</ul>",
      "rawMarkdown": "I created a new instance with V100 yesterday. I tried several zones, but V100 wasn't available until I tried `us-west1-a`.\n\n- n1-highmem-8\n- us-west1-a\n- preempted",
      "votes": 2
    },
    {
      "id": 643019,
      "postDate": "2019-10-07T00:41:50.127Z",
      "content": "<p>Is anyone else noticing that they are unable to attach the V100 GPU to their instances? I can attach other GPU's but the V100's across all the servers seem to be in use - both preempted and non-preempted instances. </p>\n\n<p>If you are able to attach a V100 GPU to your instance, please share the region, location and whether you're using preempted or not. Thanks.\n<br>\nEdit: I just tried the us-west1-a on 10/8 and it works - V100 preemptible. Thanks a lot guys!</p>",
      "rawMarkdown": "Is anyone else noticing that they are unable to attach the V100 GPU to their instances? I can attach other GPU's but the V100's across all the servers seem to be in use - both preempted and non-preempted instances. \n\nIf you are able to attach a V100 GPU to your instance, please share the region, location and whether you're using preempted or not. Thanks.\n<br>\nEdit: I just tried the us-west1-a on 10/8 and it works - V100 preemptible. Thanks a lot guys!",
      "votes": 2
    },
    {
      "id": 644932,
      "postDate": "2019-10-09T14:36:40.060Z",
      "content": "<p>Try asia-east1-c</p>",
      "rawMarkdown": "Try asia-east1-c"
    },
    {
      "id": 644050,
      "postDate": "2019-10-08T08:48:34.273Z",
      "content": "<p>Yesterday I got a 1 * V100 instance on <code>us-west1-b</code>.</p>",
      "rawMarkdown": "Yesterday I got a 1 * V100 instance on `us-west1-b`."
    },
    {
      "id": 643570,
      "postDate": "2019-10-07T16:25:47.353Z",
      "content": "<p>Hi guys, I’ve never used V100 before, just want to know : what is a training speed boost you get with V100 over P100 for basic baseline models?</p>",
      "rawMarkdown": "Hi guys, I’ve never used V100 before, just want to know : what is a training speed boost you get with V100 over P100 for basic baseline models?",
      "replies": [
        {
          "id": 643614,
          "postDate": "2019-10-07T17:26:20.803Z",
          "content": "<p>I wish I could say based on experience. However, based on the fp32(single precision) TFLOPs, The V100 has 14TFLOPS, while P100 has 9.3TFLOPS. In most cases you'll be comparing fp32, but the V100 does have tensor cores that can run larger batch sizes in mixed precision which would likely speed up training and inference. I'll update you once I get get my hands on the a VM with a V100.</p>",
          "rawMarkdown": "I wish I could say based on experience. However, based on the fp32(single precision) TFLOPs, The V100 has 14TFLOPS, while P100 has 9.3TFLOPS. In most cases you'll be comparing fp32, but the V100 does have tensor cores that can run larger batch sizes in mixed precision which would likely speed up training and inference. I'll update you once I get get my hands on the a VM with a V100.",
          "votes": 1
        },
        {
          "id": 644929,
          "postDate": "2019-10-09T14:33:55.303Z",
          "content": "<p><a href=\"/ratthachat\">@ratthachat</a>  From my testing in mixed precision - the V100 is 2.4-2.8x faster, which is well worth it considering that the cost/hr of a V100 2.35x the cost/hr for the T4. And just running the same tests on the P100, I get same speed as the T4. From what I can tell, the V100 utilizes the same CPU more than the T4 so my guess is that it can iterate over the same batchsize faster so the CPU is idling less.</p>",
          "rawMarkdown": "@ratthachat  From my testing in mixed precision - the V100 is 2.4-2.8x faster, which is well worth it considering that the cost/hr of a V100 2.35x the cost/hr for the T4. And just running the same tests on the P100, I get same speed as the T4. From what I can tell, the V100 utilizes the same CPU more than the T4 so my guess is that it can iterate over the same batchsize faster so the CPU is idling less.",
          "votes": 1
        },
        {
          "id": 644953,
          "postDate": "2019-10-09T15:02:54.250Z",
          "content": "<p><a href=\"/teeyee314\">@teeyee314</a> Are you running preemptible instance or regular instance? At first I tried preemptible instance, but noticed that the performance while training the model isn't consistent. Specially the first epoch takes almost 2x longer time than other epochs. Now I am running a regular V100 instance and the performance is consistent.</p>",
          "rawMarkdown": "@teeyee314 Are you running preemptible instance or regular instance? At first I tried preemptible instance, but noticed that the performance while training the model isn't consistent. Specially the first epoch takes almost 2x longer time than other epochs. Now I am running a regular V100 instance and the performance is consistent.",
          "votes": 1
        },
        {
          "id": 644967,
          "postDate": "2019-10-09T15:26:06.370Z",
          "content": "<p>Strange. I haven't used regular instances, only colab, kaggle kernel and gcp preemptibles. Very consistent pytorch and apex training times. Are you loading from pngs/jpgs or from dicoms?</p>",
          "rawMarkdown": "Strange. I haven't used regular instances, only colab, kaggle kernel and gcp preemptibles. Very consistent pytorch and apex training times. Are you loading from pngs/jpgs or from dicoms?",
          "votes": 1
        },
        {
          "id": 644985,
          "postDate": "2019-10-09T15:43:31.697Z",
          "content": "<p>I am loading png images (the datasets pre-processed by you). I am also using pytorch (fastai). I trained several efficientnet models. Previously I trained on colab and kaggle kernel, where training times were consistent.</p>\n\n<p>On GCP preemptibles, training times were consistent at first. But later I noticed that the first epoch always took more than 2x times. Then I switched to a regular instance. </p>\n\n<p>Can you please share how long does it take to run an epoch on V100? I am currently training an Efficientnet B5 model with 224 png images and it takes around 1h per epoch (batch size 64 &amp; fp16).</p>",
          "rawMarkdown": "I am loading png images (the datasets pre-processed by you). I am also using pytorch (fastai). I trained several efficientnet models. Previously I trained on colab and kaggle kernel, where training times were consistent.\n\nOn GCP preemptibles, training times were consistent at first. But later I noticed that the first epoch always took more than 2x times. Then I switched to a regular instance. \n\nCan you please share how long does it take to run an epoch on V100? I am currently training an Efficientnet B5 model with 224 png images and it takes around 1h per epoch (batch size 64 &amp; fp16).",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 643620,
      "author_name": "Alexander Abstreiter",
      "author_url": "",
      "post_date": "2019-10-07T17:36:03.877000",
      "content": "<p>Got a V100 on <code>us-west1-a</code>.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 643034,
      "author_name": "Atikur Rahman",
      "author_url": "",
      "post_date": "2019-10-07T01:47:31.233000",
      "content": "<p>I created a new instance with V100 yesterday. I tried several zones, but V100 wasn't available until I tried <code>us-west1-a</code>.</p>\n\n<ul>\n<li>n1-highmem-8</li>\n<li>us-west1-a</li>\n<li>preempted</li>\n</ul>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 644932,
      "author_name": "Jayaram",
      "author_url": "",
      "post_date": "2019-10-09T14:36:40.060000",
      "content": "<p>Try asia-east1-c</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 644050,
      "author_name": "kambarakun",
      "author_url": "",
      "post_date": "2019-10-08T08:48:34.273000",
      "content": "<p>Yesterday I got a 1 * V100 instance on <code>us-west1-b</code>.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 643570,
      "author_name": "Neuron Engineer",
      "author_url": "",
      "post_date": "2019-10-07T16:25:47.353000",
      "content": "<p>Hi guys, I’ve never used V100 before, just want to know : what is a training speed boost you get with V100 over P100 for basic baseline models?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 643614,
          "author_name": "Tim Yee",
          "author_url": "",
          "post_date": "2019-10-07T17:26:20.803000",
          "content": "<p>I wish I could say based on experience. However, based on the fp32(single precision) TFLOPs, The V100 has 14TFLOPS, while P100 has 9.3TFLOPS. In most cases you'll be comparing fp32, but the V100 does have tensor cores that can run larger batch sizes in mixed precision which would likely speed up training and inference. I'll update you once I get get my hands on the a VM with a V100.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 644929,
          "author_name": "Tim Yee",
          "author_url": "",
          "post_date": "2019-10-09T14:33:55.303000",
          "content": "<p><a href=\"/ratthachat\">@ratthachat</a>  From my testing in mixed precision - the V100 is 2.4-2.8x faster, which is well worth it considering that the cost/hr of a V100 2.35x the cost/hr for the T4. And just running the same tests on the P100, I get same speed as the T4. From what I can tell, the V100 utilizes the same CPU more than the T4 so my guess is that it can iterate over the same batchsize faster so the CPU is idling less.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 644953,
          "author_name": "Atikur Rahman",
          "author_url": "",
          "post_date": "2019-10-09T15:02:54.250000",
          "content": "<p><a href=\"/teeyee314\">@teeyee314</a> Are you running preemptible instance or regular instance? At first I tried preemptible instance, but noticed that the performance while training the model isn't consistent. Specially the first epoch takes almost 2x longer time than other epochs. Now I am running a regular V100 instance and the performance is consistent.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 644967,
          "author_name": "Tim Yee",
          "author_url": "",
          "post_date": "2019-10-09T15:26:06.370000",
          "content": "<p>Strange. I haven't used regular instances, only colab, kaggle kernel and gcp preemptibles. Very consistent pytorch and apex training times. Are you loading from pngs/jpgs or from dicoms?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 644985,
          "author_name": "Atikur Rahman",
          "author_url": "",
          "post_date": "2019-10-09T15:43:31.697000",
          "content": "<p>I am loading png images (the datasets pre-processed by you). I am also using pytorch (fastai). I trained several efficientnet models. Previously I trained on colab and kaggle kernel, where training times were consistent.</p>\n\n<p>On GCP preemptibles, training times were consistent at first. But later I noticed that the first epoch always took more than 2x times. Then I switched to a regular instance. </p>\n\n<p>Can you please share how long does it take to run an epoch on V100? I am currently training an Efficientnet B5 model with 224 png images and it takes around 1h per epoch (batch size 64 &amp; fp16).</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "643620": "Got a V100 on `us-west1-a`.",
    "643034": "I created a new instance with V100 yesterday. I tried several zones, but V100 wasn't available until I tried `us-west1-a`.\n\n- n1-highmem-8\n- us-west1-a\n- preempted",
    "643019": "Is anyone else noticing that they are unable to attach the V100 GPU to their instances? I can attach other GPU's but the V100's across all the servers seem to be in use - both preempted and non-preempted instances. \n\nIf you are able to attach a V100 GPU to your instance, please share the region, location and whether you're using preempted or not. Thanks.\n<br>\nEdit: I just tried the us-west1-a on 10/8 and it works - V100 preemptible. Thanks a lot guys!",
    "644932": "Try asia-east1-c",
    "644050": "Yesterday I got a 1 * V100 instance on `us-west1-b`.",
    "643570": "Hi guys, I’ve never used V100 before, just want to know : what is a training speed boost you get with V100 over P100 for basic baseline models?"
  }
}