{
  "id": 113284,
  "title": "[WIP] PyTorch TPU pipeline",
  "url": "/competitions/rsna-intracranial-hemorrhage-detection/discussion/113284",
  "author_name": "Konstantin Lopukhin",
  "post_date": "2019-10-18T08:41:27.391000",
  "votes": 26,
  "comment_count": 18,
  "views": 0,
  "content": "<p>I'm sharing my early pipeline which uses TPU + PyTorch <a href=\"https://github.com/lopuhin/kaggle-rsna-2019\">https://github.com/lopuhin/kaggle-rsna-2019</a>. Note that at the moment it's not a complete pipeline, only training is implemented and even then I didn't check anything besides decreasing loss. I'll push updates directly to the repo and will add a comment here once pipeline is more complete. It won't get you 0.066 on LB ;)</p>\n\n<p>My main goal was to check if Pytorch TPU support is good enough for classification. There are some preliminary performance numbers in the repo <a href=\"https://github.com/lopuhin/kaggle-rsna-2019#performance\">https://github.com/lopuhin/kaggle-rsna-2019#performance</a>.</p>\n\n<p>So far working with TPU looks very similar to working with a multi-GPU with distributed data parallel - it needs very little modifications, at least when all ops are supported and shapes are static, like it is for a simple classifications task. It also needs an efficient data pipeline and a powerful machine to feed all 8 TPU core, similar to what you'd need for an 8-GPU machine. So far I didn't see any strange hangs or stability issues.</p>\n\n<p>In terms of the API, what I really like about pytorch TPU support:</p>\n\n<ul>\n<li>you can use any pytorch models as long as all ops are supported, you don't need to convert any weights to/from TPU</li>\n<li>TPU-specicif API is quite small and clear: <a href=\"https://github.com/pytorch/xla/blob/master/API_GUIDE.md\">https://github.com/pytorch/xla/blob/master/API_GUIDE.md</a></li>\n<li>you can use the same data pipeline as for a regular GPU</li>\n</ul>\n\n<p>With TF, recommended approach is to prepare and feed TFRecords which are read by the TPU from Google Cloud Storage. On one hand, it looks much less convenient for quick development and prototyping (e.g. probably you won't be able to use you favorite augmentations library). On the other hand, you don't need a powerful machine to feed the data.</p>\n\n<p>In terms of price effectiveness, it's a hard call and depends very much on how much does TPU cost you and where can you rent GPUs. I also didn't check preemptible TPUs yet.</p>",
  "messages": [
    {
      "id": 652027,
      "postDate": "2019-10-18T08:41:27.390Z",
      "content": "<p>I'm sharing my early pipeline which uses TPU + PyTorch <a href=\"https://github.com/lopuhin/kaggle-rsna-2019\">https://github.com/lopuhin/kaggle-rsna-2019</a>. Note that at the moment it's not a complete pipeline, only training is implemented and even then I didn't check anything besides decreasing loss. I'll push updates directly to the repo and will add a comment here once pipeline is more complete. It won't get you 0.066 on LB ;)</p>\n\n<p>My main goal was to check if Pytorch TPU support is good enough for classification. There are some preliminary performance numbers in the repo <a href=\"https://github.com/lopuhin/kaggle-rsna-2019#performance\">https://github.com/lopuhin/kaggle-rsna-2019#performance</a>.</p>\n\n<p>So far working with TPU looks very similar to working with a multi-GPU with distributed data parallel - it needs very little modifications, at least when all ops are supported and shapes are static, like it is for a simple classifications task. It also needs an efficient data pipeline and a powerful machine to feed all 8 TPU core, similar to what you'd need for an 8-GPU machine. So far I didn't see any strange hangs or stability issues.</p>\n\n<p>In terms of the API, what I really like about pytorch TPU support:</p>\n\n<ul>\n<li>you can use any pytorch models as long as all ops are supported, you don't need to convert any weights to/from TPU</li>\n<li>TPU-specicif API is quite small and clear: <a href=\"https://github.com/pytorch/xla/blob/master/API_GUIDE.md\">https://github.com/pytorch/xla/blob/master/API_GUIDE.md</a></li>\n<li>you can use the same data pipeline as for a regular GPU</li>\n</ul>\n\n<p>With TF, recommended approach is to prepare and feed TFRecords which are read by the TPU from Google Cloud Storage. On one hand, it looks much less convenient for quick development and prototyping (e.g. probably you won't be able to use you favorite augmentations library). On the other hand, you don't need a powerful machine to feed the data.</p>\n\n<p>In terms of price effectiveness, it's a hard call and depends very much on how much does TPU cost you and where can you rent GPUs. I also didn't check preemptible TPUs yet.</p>",
      "rawMarkdown": "I'm sharing my early pipeline which uses TPU + PyTorch https://github.com/lopuhin/kaggle-rsna-2019. Note that at the moment it's not a complete pipeline, only training is implemented and even then I didn't check anything besides decreasing loss. I'll push updates directly to the repo and will add a comment here once pipeline is more complete. It won't get you 0.066 on LB ;)\n\nMy main goal was to check if Pytorch TPU support is good enough for classification. There are some preliminary performance numbers in the repo https://github.com/lopuhin/kaggle-rsna-2019#performance.\n\nSo far working with TPU looks very similar to working with a multi-GPU with distributed data parallel - it needs very little modifications, at least when all ops are supported and shapes are static, like it is for a simple classifications task. It also needs an efficient data pipeline and a powerful machine to feed all 8 TPU core, similar to what you'd need for an 8-GPU machine. So far I didn't see any strange hangs or stability issues.\n\nIn terms of the API, what I really like about pytorch TPU support:\n\n- you can use any pytorch models as long as all ops are supported, you don't need to convert any weights to/from TPU\n- TPU-specicif API is quite small and clear: https://github.com/pytorch/xla/blob/master/API_GUIDE.md\n- you can use the same data pipeline as for a regular GPU\n\nWith TF, recommended approach is to prepare and feed TFRecords which are read by the TPU from Google Cloud Storage. On one hand, it looks much less convenient for quick development and prototyping (e.g. probably you won't be able to use you favorite augmentations library). On the other hand, you don't need a powerful machine to feed the data.\n\nIn terms of price effectiveness, it's a hard call and depends very much on how much does TPU cost you and where can you rent GPUs. I also didn't check preemptible TPUs yet.",
      "votes": 26
    },
    {
      "id": 652040,
      "postDate": "2019-10-18T08:55:41.917Z",
      "content": "<p>Other writeups about using TPUs on Kaggle:\n- <a href=\"https://www.kaggle.com/c/open-images-2019-object-detection/discussion/110941\">https://www.kaggle.com/c/open-images-2019-object-detection/discussion/110941</a> TPU + TF for object detection by Artyom Palvelev\n- <a href=\"https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/102391\">https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/102391</a> TPU + PyTorch for classification by nosound</p>",
      "rawMarkdown": "Other writeups about using TPUs on Kaggle:\n- https://www.kaggle.com/c/open-images-2019-object-detection/discussion/110941 TPU + TF for object detection by Artyom Palvelev\n- https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/102391 TPU + PyTorch for classification by nosound",
      "votes": 6
    },
    {
      "id": 652038,
      "postDate": "2019-10-18T08:52:44.690Z",
      "content": "<p>Same here, I am using TPU+pytorch for this competition. It works smoothly for me now, but the model is straightforward, static shapes, no fancy operations. I noticed that is when the problems start appearing, - dynamic shapes result in slowdowns, and some operations are not supported yet. </p>",
      "rawMarkdown": "Same here, I am using TPU+pytorch for this competition. It works smoothly for me now, but the model is straightforward, static shapes, no fancy operations. I noticed that is when the problems start appearing, - dynamic shapes result in slowdowns, and some operations are not supported yet. ",
      "votes": 6
    },
    {
      "id": 654088,
      "postDate": "2019-10-21T12:27:46.677Z",
      "content": "<p>Hi Konstantin, I am able to run your code to begin training a model, every time I set the device to the tpus it takes more than a minute and I get this message:</p>\n\n<p><code>\n2019-10-21 12:20:25.647590: E tensorflow/core/framework/op_kernel.cc:1579] OpKernel ('op: \"ErfinvGrad\" device_type: \"CPU\" constraint { name: \"T\" allowed_values { list { type: DT_DOUBLE } } }') for unknown op: ErfinvGrad\n2019-10-21 12:20:25.647662: E tensorflow/core/framework/op_kernel.cc:1579] OpKernel ('op: \"ErfinvGrad\" device_type: \"CPU\" constraint { name: \"T\" allowed_values { list { type: DT_FLOAT } } }') for unknown op: ErfinvGrad\n2019-10-21 12:20:25.647695: E tensorflow/core/framework/op_kernel.cc:1579] OpKernel ('op: \"NdtriGrad\" device_type: \"CPU\" constraint { name: \"T\" allowed_values { list { type: DT_DOUBLE } } }') for unknown op: NdtriGrad\n2019-10-21 12:20:25.647722: E tensorflow/core/framework/op_kernel.cc:1579] OpKernel ('op: \"NdtriGrad\" device_type: \"CPU\" constraint { name: \"T\" allowed_values { list { type: DT_FLOAT } } }') for unknown op: NdtriGrad\n</code>\nit still seems to work, but i could not find anything online about this specific message. Have you, or anyone else who has tried using pytorch on tpu seen this?</p>\n\n<p>I assumed it was some error as i saw it also when I tried to use pytorch_xla in the recursion comp after pulling the xla docker image, but there i could not get anything to begin training. Thanks again for the guide!</p>",
      "rawMarkdown": "Hi Konstantin, I am able to run your code to begin training a model, every time I set the device to the tpus it takes more than a minute and I get this message:\n\n```\n2019-10-21 12:20:25.647590: E tensorflow/core/framework/op_kernel.cc:1579] OpKernel ('op: \"ErfinvGrad\" device_type: \"CPU\" constraint { name: \"T\" allowed_values { list { type: DT_DOUBLE } } }') for unknown op: ErfinvGrad\n2019-10-21 12:20:25.647662: E tensorflow/core/framework/op_kernel.cc:1579] OpKernel ('op: \"ErfinvGrad\" device_type: \"CPU\" constraint { name: \"T\" allowed_values { list { type: DT_FLOAT } } }') for unknown op: ErfinvGrad\n2019-10-21 12:20:25.647695: E tensorflow/core/framework/op_kernel.cc:1579] OpKernel ('op: \"NdtriGrad\" device_type: \"CPU\" constraint { name: \"T\" allowed_values { list { type: DT_DOUBLE } } }') for unknown op: NdtriGrad\n2019-10-21 12:20:25.647722: E tensorflow/core/framework/op_kernel.cc:1579] OpKernel ('op: \"NdtriGrad\" device_type: \"CPU\" constraint { name: \"T\" allowed_values { list { type: DT_FLOAT } } }') for unknown op: NdtriGrad\n```\nit still seems to work, but i could not find anything online about this specific message. Have you, or anyone else who has tried using pytorch on tpu seen this?\n\nI assumed it was some error as i saw it also when I tried to use pytorch_xla in the recursion comp after pulling the xla docker image, but there i could not get anything to begin training. Thanks again for the guide!",
      "votes": 1,
      "replies": [
        {
          "id": 654100,
          "postDate": "2019-10-21T12:46:57.957Z",
          "content": "<p>Yes, I get the same message, not sure how important is this. I assumed it's some op which does not have TPU support and would run on CPU.</p>\n\n<p>Regarding training, it's normal that there is no output for some time, especially at first run, because the models are compiled on the TPU. But within a few minutes there should be some output, and some CPU usage. If there is none, then something is not quite right.</p>",
          "rawMarkdown": "Yes, I get the same message, not sure how important is this. I assumed it's some op which does not have TPU support and would run on CPU.\n\nRegarding training, it's normal that there is no output for some time, especially at first run, because the models are compiled on the TPU. But within a few minutes there should be some output, and some CPU usage. If there is none, then something is not quite right.",
          "votes": 2
        }
      ]
    },
    {
      "id": 653216,
      "postDate": "2019-10-20T04:04:38.067Z",
      "content": "<p>How do you upload your files in colab?\nOr do you have access to the GCP?\nThanks</p>",
      "rawMarkdown": "How do you upload your files in colab?\nOr do you have access to the GCP?\nThanks",
      "votes": 1,
      "replies": [
        {
          "id": 653337,
          "postDate": "2019-10-20T09:20:21.727Z",
          "content": "<p>I worked on GCP, yes. Unfortunately with pytorch dynamic loading approach, Colab VM is not powerful enough to fully utilize a TPU. There are some examples though <a href=\"https://github.com/pytorch/xla/tree/master/contrib/colab#get-started-with-our-colab-tutorials\">https://github.com/pytorch/xla/tree/master/contrib/colab#get-started-with-our-colab-tutorials</a> and you can also connect to Colab via ssh.</p>",
          "rawMarkdown": "I worked on GCP, yes. Unfortunately with pytorch dynamic loading approach, Colab VM is not powerful enough to fully utilize a TPU. There are some examples though https://github.com/pytorch/xla/tree/master/contrib/colab#get-started-with-our-colab-tutorials and you can also connect to Colab via ssh."
        }
      ]
    },
    {
      "id": 652956,
      "postDate": "2019-10-19T17:54:12.327Z",
      "content": "<p>Thank you for posting this Konstantin! Your TPU code is very informative, I'd tried a few times to get pytorch_xla working before without success so seeing a pytorch model training on the tpu is most excellent.  </p>",
      "rawMarkdown": "Thank you for posting this Konstantin! Your TPU code is very informative, I'd tried a few times to get pytorch_xla working before without success so seeing a pytorch model training on the tpu is most excellent.  ",
      "votes": 1,
      "replies": [
        {
          "id": 653218,
          "postDate": "2019-10-20T04:16:27.223Z",
          "content": "<p>but the datasets is 150G.how do you manage to upload into GCP or colab?</p>",
          "rawMarkdown": "but the datasets is 150G.how do you manage to upload into GCP or colab?",
          "votes": 1
        },
        {
          "id": 653338,
          "postDate": "2019-10-20T09:22:09.987Z",
          "content": "<p>Good point, large dataset size is another problem for Colab. With GCP it's not as issue (but be sure to use an SSD).</p>",
          "rawMarkdown": "Good point, large dataset size is another problem for Colab. With GCP it's not as issue (but be sure to use an SSD)."
        },
        {
          "id": 653416,
          "postDate": "2019-10-20T12:19:08.170Z",
          "content": "<p>Thanks for your replies.\nCould you tell me how large is your GCP disk space?thanks.\nwhy did you say \"it's not as issue\"?\nThanks</p>",
          "rawMarkdown": "Thanks for your replies.\nCould you tell me how large is your GCP disk space?thanks.\nwhy did you say \"it's not as issue\"?\nThanks"
        },
        {
          "id": 653674,
          "postDate": "2019-10-20T19:50:22.510Z",
          "content": "<p>I created an SSD disk with the size of 400 GB, which is enough to hold the unpacked dataset. I said it's not an issue because you can create large enough disk if you're already on GCP, and it's price is small compared to other components.</p>",
          "rawMarkdown": "I created an SSD disk with the size of 400 GB, which is enough to hold the unpacked dataset. I said it's not an issue because you can create large enough disk if you're already on GCP, and it's price is small compared to other components."
        },
        {
          "id": 653681,
          "postDate": "2019-10-20T20:12:10.857Z",
          "content": "<p>Good point, this is something that took me a few tries to get right. I used the Kaggle api to get the data into my gcp instance. </p>\n\n<p>I had to request a quota increase from the 500GB ssd limit since the zip file and the unzipped data were larger than this. </p>\n\n<p>I agree with Konstatin, you should get the storage you need as it is not the expensive part of this by any means, you can always downsize your gcp disk later. </p>",
          "rawMarkdown": "Good point, this is something that took me a few tries to get right. I used the Kaggle api to get the data into my gcp instance. \n\nI had to request a quota increase from the 500GB ssd limit since the zip file and the unzipped data were larger than this. \n\nI agree with Konstatin, you should get the storage you need as it is not the expensive part of this by any means, you can always downsize your gcp disk later. ",
          "votes": 1
        },
        {
          "id": 653930,
          "postDate": "2019-10-21T07:24:18.610Z",
          "content": "<p>Right, good point that it's better to download with kaggle API. I used a large HDD for download and unzipping, and then copied data to a smaller SSD to avoid quota increase.</p>",
          "rawMarkdown": "Right, good point that it's better to download with kaggle API. I used a large HDD for download and unzipping, and then copied data to a smaller SSD to avoid quota increase.",
          "votes": 1
        },
        {
          "id": 654140,
          "postDate": "2019-10-21T14:02:29.497Z",
          "content": "<p>I was wondering, for training on tpus with Tensorflow we are supposed to convert whatever dataset into tfrecords that live on a gs bucket from which the Estimator pulls batches for training. With pytorch_xla we pull data from the mounted drives as usual using the DataLoader. </p>\n\n<p>The xla parallel loader function just tells the DataLoader how to split up batches to the available tpu cores but still from the attached hard drive. Have you, or anyone, tried or know what it would take to try to allow the loader to pull data from a gs bucket? Would (could) this be any faster than loading directly from a mounted volume?</p>",
          "rawMarkdown": "I was wondering, for training on tpus with Tensorflow we are supposed to convert whatever dataset into tfrecords that live on a gs bucket from which the Estimator pulls batches for training. With pytorch_xla we pull data from the mounted drives as usual using the DataLoader. \n\nThe xla parallel loader function just tells the DataLoader how to split up batches to the available tpu cores but still from the attached hard drive. Have you, or anyone, tried or know what it would take to try to allow the loader to pull data from a gs bucket? Would (could) this be any faster than loading directly from a mounted volume?"
        },
        {
          "id": 654262,
          "postDate": "2019-10-21T16:20:20.377Z",
          "content": "<p>I think currently it's not possible to read data from gs bucket directly with pytorch/xla, but does not hurt to ask in the repo issues (edit: posted <a href=\"https://github.com/pytorch/xla/issues/1221\">https://github.com/pytorch/xla/issues/1221</a>). Regarding speed - I think in theory, it should be the same as with GPUs, meaning that if the data pipeline is fast enough (not bound by network I/O, CPU or disk speed) and can deliver data to TPU fast enough, then there should be no noticeable difference. But I'm not sure it's true in practice, I guess the only way to tell for sure is to have the same code implemented in TF and pytorch/xla and compare the speed.</p>\n\n<p>So if anyone is using TPUs with TF here and has performance numbers, please share :)</p>",
          "rawMarkdown": "I think currently it's not possible to read data from gs bucket directly with pytorch/xla, but does not hurt to ask in the repo issues (edit: posted https://github.com/pytorch/xla/issues/1221). Regarding speed - I think in theory, it should be the same as with GPUs, meaning that if the data pipeline is fast enough (not bound by network I/O, CPU or disk speed) and can deliver data to TPU fast enough, then there should be no noticeable difference. But I'm not sure it's true in practice, I guess the only way to tell for sure is to have the same code implemented in TF and pytorch/xla and compare the speed.\n\nSo if anyone is using TPUs with TF here and has performance numbers, please share :)"
        }
      ]
    },
    {
      "id": 652194,
      "postDate": "2019-10-18T13:52:13.397Z",
      "content": "<p>thanks for sharing! like small and powerful kernels.up</p>",
      "rawMarkdown": "thanks for sharing! like small and powerful kernels.up",
      "votes": 1
    },
    {
      "id": 674782,
      "postDate": "2019-11-17T04:19:04.903Z",
      "content": "<p>could you tell me how to write \"ip address range\"when creating a TPU node?\nThanks</p>",
      "rawMarkdown": "could you tell me how to write \"ip address range\"when creating a TPU node?\nThanks",
      "replies": [
        {
          "id": 675499,
          "postDate": "2019-11-18T06:15:33.637Z",
          "content": "<p>172.16.0.0 worked for me, you can check docs for more details <a href=\"https://cloud.google.com/tpu/docs/internal-ip-blocks\">https://cloud.google.com/tpu/docs/internal-ip-blocks</a></p>",
          "rawMarkdown": "172.16.0.0 worked for me, you can check docs for more details https://cloud.google.com/tpu/docs/internal-ip-blocks"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 652040,
      "author_name": "Konstantin Lopukhin",
      "author_url": "",
      "post_date": "2019-10-18T08:55:41.917000",
      "content": "<p>Other writeups about using TPUs on Kaggle:\n- <a href=\"https://www.kaggle.com/c/open-images-2019-object-detection/discussion/110941\">https://www.kaggle.com/c/open-images-2019-object-detection/discussion/110941</a> TPU + TF for object detection by Artyom Palvelev\n- <a href=\"https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/102391\">https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/102391</a> TPU + PyTorch for classification by nosound</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 652038,
      "author_name": "nosound",
      "author_url": "",
      "post_date": "2019-10-18T08:52:44.690000",
      "content": "<p>Same here, I am using TPU+pytorch for this competition. It works smoothly for me now, but the model is straightforward, static shapes, no fancy operations. I noticed that is when the problems start appearing, - dynamic shapes result in slowdowns, and some operations are not supported yet. </p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 654088,
      "author_name": "interneuron",
      "author_url": "",
      "post_date": "2019-10-21T12:27:46.677000",
      "content": "<p>Hi Konstantin, I am able to run your code to begin training a model, every time I set the device to the tpus it takes more than a minute and I get this message:</p>\n\n<p><code>\n2019-10-21 12:20:25.647590: E tensorflow/core/framework/op_kernel.cc:1579] OpKernel ('op: \"ErfinvGrad\" device_type: \"CPU\" constraint { name: \"T\" allowed_values { list { type: DT_DOUBLE } } }') for unknown op: ErfinvGrad\n2019-10-21 12:20:25.647662: E tensorflow/core/framework/op_kernel.cc:1579] OpKernel ('op: \"ErfinvGrad\" device_type: \"CPU\" constraint { name: \"T\" allowed_values { list { type: DT_FLOAT } } }') for unknown op: ErfinvGrad\n2019-10-21 12:20:25.647695: E tensorflow/core/framework/op_kernel.cc:1579] OpKernel ('op: \"NdtriGrad\" device_type: \"CPU\" constraint { name: \"T\" allowed_values { list { type: DT_DOUBLE } } }') for unknown op: NdtriGrad\n2019-10-21 12:20:25.647722: E tensorflow/core/framework/op_kernel.cc:1579] OpKernel ('op: \"NdtriGrad\" device_type: \"CPU\" constraint { name: \"T\" allowed_values { list { type: DT_FLOAT } } }') for unknown op: NdtriGrad\n</code>\nit still seems to work, but i could not find anything online about this specific message. Have you, or anyone else who has tried using pytorch on tpu seen this?</p>\n\n<p>I assumed it was some error as i saw it also when I tried to use pytorch_xla in the recursion comp after pulling the xla docker image, but there i could not get anything to begin training. Thanks again for the guide!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 654100,
          "author_name": "Konstantin Lopukhin",
          "author_url": "",
          "post_date": "2019-10-21T12:46:57.957000",
          "content": "<p>Yes, I get the same message, not sure how important is this. I assumed it's some op which does not have TPU support and would run on CPU.</p>\n\n<p>Regarding training, it's normal that there is no output for some time, especially at first run, because the models are compiled on the TPU. But within a few minutes there should be some output, and some CPU usage. If there is none, then something is not quite right.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 653216,
      "author_name": "Tian Bingyang",
      "author_url": "",
      "post_date": "2019-10-20T04:04:38.067000",
      "content": "<p>How do you upload your files in colab?\nOr do you have access to the GCP?\nThanks</p>",
      "votes": 1,
      "replies": [
        {
          "id": 653337,
          "author_name": "Konstantin Lopukhin",
          "author_url": "",
          "post_date": "2019-10-20T09:20:21.727000",
          "content": "<p>I worked on GCP, yes. Unfortunately with pytorch dynamic loading approach, Colab VM is not powerful enough to fully utilize a TPU. There are some examples though <a href=\"https://github.com/pytorch/xla/tree/master/contrib/colab#get-started-with-our-colab-tutorials\">https://github.com/pytorch/xla/tree/master/contrib/colab#get-started-with-our-colab-tutorials</a> and you can also connect to Colab via ssh.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 652956,
      "author_name": "interneuron",
      "author_url": "",
      "post_date": "2019-10-19T17:54:12.327000",
      "content": "<p>Thank you for posting this Konstantin! Your TPU code is very informative, I'd tried a few times to get pytorch_xla working before without success so seeing a pytorch model training on the tpu is most excellent.  </p>",
      "votes": 1,
      "replies": [
        {
          "id": 653218,
          "author_name": "Tian Bingyang",
          "author_url": "",
          "post_date": "2019-10-20T04:16:27.223000",
          "content": "<p>but the datasets is 150G.how do you manage to upload into GCP or colab?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 653338,
          "author_name": "Konstantin Lopukhin",
          "author_url": "",
          "post_date": "2019-10-20T09:22:09.987000",
          "content": "<p>Good point, large dataset size is another problem for Colab. With GCP it's not as issue (but be sure to use an SSD).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 653416,
          "author_name": "Tian Bingyang",
          "author_url": "",
          "post_date": "2019-10-20T12:19:08.170000",
          "content": "<p>Thanks for your replies.\nCould you tell me how large is your GCP disk space?thanks.\nwhy did you say \"it's not as issue\"?\nThanks</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 653674,
          "author_name": "Konstantin Lopukhin",
          "author_url": "",
          "post_date": "2019-10-20T19:50:22.510000",
          "content": "<p>I created an SSD disk with the size of 400 GB, which is enough to hold the unpacked dataset. I said it's not an issue because you can create large enough disk if you're already on GCP, and it's price is small compared to other components.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 653681,
          "author_name": "interneuron",
          "author_url": "",
          "post_date": "2019-10-20T20:12:10.857000",
          "content": "<p>Good point, this is something that took me a few tries to get right. I used the Kaggle api to get the data into my gcp instance. </p>\n\n<p>I had to request a quota increase from the 500GB ssd limit since the zip file and the unzipped data were larger than this. </p>\n\n<p>I agree with Konstatin, you should get the storage you need as it is not the expensive part of this by any means, you can always downsize your gcp disk later. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 653930,
          "author_name": "Konstantin Lopukhin",
          "author_url": "",
          "post_date": "2019-10-21T07:24:18.610000",
          "content": "<p>Right, good point that it's better to download with kaggle API. I used a large HDD for download and unzipping, and then copied data to a smaller SSD to avoid quota increase.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 654140,
          "author_name": "interneuron",
          "author_url": "",
          "post_date": "2019-10-21T14:02:29.497000",
          "content": "<p>I was wondering, for training on tpus with Tensorflow we are supposed to convert whatever dataset into tfrecords that live on a gs bucket from which the Estimator pulls batches for training. With pytorch_xla we pull data from the mounted drives as usual using the DataLoader. </p>\n\n<p>The xla parallel loader function just tells the DataLoader how to split up batches to the available tpu cores but still from the attached hard drive. Have you, or anyone, tried or know what it would take to try to allow the loader to pull data from a gs bucket? Would (could) this be any faster than loading directly from a mounted volume?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 654262,
          "author_name": "Konstantin Lopukhin",
          "author_url": "",
          "post_date": "2019-10-21T16:20:20.377000",
          "content": "<p>I think currently it's not possible to read data from gs bucket directly with pytorch/xla, but does not hurt to ask in the repo issues (edit: posted <a href=\"https://github.com/pytorch/xla/issues/1221\">https://github.com/pytorch/xla/issues/1221</a>). Regarding speed - I think in theory, it should be the same as with GPUs, meaning that if the data pipeline is fast enough (not bound by network I/O, CPU or disk speed) and can deliver data to TPU fast enough, then there should be no noticeable difference. But I'm not sure it's true in practice, I guess the only way to tell for sure is to have the same code implemented in TF and pytorch/xla and compare the speed.</p>\n\n<p>So if anyone is using TPUs with TF here and has performance numbers, please share :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 652194,
      "author_name": "Adrian Zinovei",
      "author_url": "",
      "post_date": "2019-10-18T13:52:13.397000",
      "content": "<p>thanks for sharing! like small and powerful kernels.up</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 674782,
      "author_name": "Tian Bingyang",
      "author_url": "",
      "post_date": "2019-11-17T04:19:04.903000",
      "content": "<p>could you tell me how to write \"ip address range\"when creating a TPU node?\nThanks</p>",
      "votes": 0,
      "replies": [
        {
          "id": 675499,
          "author_name": "Konstantin Lopukhin",
          "author_url": "",
          "post_date": "2019-11-18T06:15:33.637000",
          "content": "<p>172.16.0.0 worked for me, you can check docs for more details <a href=\"https://cloud.google.com/tpu/docs/internal-ip-blocks\">https://cloud.google.com/tpu/docs/internal-ip-blocks</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "652027": "I'm sharing my early pipeline which uses TPU + PyTorch https://github.com/lopuhin/kaggle-rsna-2019. Note that at the moment it's not a complete pipeline, only training is implemented and even then I didn't check anything besides decreasing loss. I'll push updates directly to the repo and will add a comment here once pipeline is more complete. It won't get you 0.066 on LB ;)\n\nMy main goal was to check if Pytorch TPU support is good enough for classification. There are some preliminary performance numbers in the repo https://github.com/lopuhin/kaggle-rsna-2019#performance.\n\nSo far working with TPU looks very similar to working with a multi-GPU with distributed data parallel - it needs very little modifications, at least when all ops are supported and shapes are static, like it is for a simple classifications task. It also needs an efficient data pipeline and a powerful machine to feed all 8 TPU core, similar to what you'd need for an 8-GPU machine. So far I didn't see any strange hangs or stability issues.\n\nIn terms of the API, what I really like about pytorch TPU support:\n\n- you can use any pytorch models as long as all ops are supported, you don't need to convert any weights to/from TPU\n- TPU-specicif API is quite small and clear: https://github.com/pytorch/xla/blob/master/API_GUIDE.md\n- you can use the same data pipeline as for a regular GPU\n\nWith TF, recommended approach is to prepare and feed TFRecords which are read by the TPU from Google Cloud Storage. On one hand, it looks much less convenient for quick development and prototyping (e.g. probably you won't be able to use you favorite augmentations library). On the other hand, you don't need a powerful machine to feed the data.\n\nIn terms of price effectiveness, it's a hard call and depends very much on how much does TPU cost you and where can you rent GPUs. I also didn't check preemptible TPUs yet.",
    "652040": "Other writeups about using TPUs on Kaggle:\n- https://www.kaggle.com/c/open-images-2019-object-detection/discussion/110941 TPU + TF for object detection by Artyom Palvelev\n- https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/102391 TPU + PyTorch for classification by nosound",
    "652038": "Same here, I am using TPU+pytorch for this competition. It works smoothly for me now, but the model is straightforward, static shapes, no fancy operations. I noticed that is when the problems start appearing, - dynamic shapes result in slowdowns, and some operations are not supported yet. ",
    "654088": "Hi Konstantin, I am able to run your code to begin training a model, every time I set the device to the tpus it takes more than a minute and I get this message:\n\n```\n2019-10-21 12:20:25.647590: E tensorflow/core/framework/op_kernel.cc:1579] OpKernel ('op: \"ErfinvGrad\" device_type: \"CPU\" constraint { name: \"T\" allowed_values { list { type: DT_DOUBLE } } }') for unknown op: ErfinvGrad\n2019-10-21 12:20:25.647662: E tensorflow/core/framework/op_kernel.cc:1579] OpKernel ('op: \"ErfinvGrad\" device_type: \"CPU\" constraint { name: \"T\" allowed_values { list { type: DT_FLOAT } } }') for unknown op: ErfinvGrad\n2019-10-21 12:20:25.647695: E tensorflow/core/framework/op_kernel.cc:1579] OpKernel ('op: \"NdtriGrad\" device_type: \"CPU\" constraint { name: \"T\" allowed_values { list { type: DT_DOUBLE } } }') for unknown op: NdtriGrad\n2019-10-21 12:20:25.647722: E tensorflow/core/framework/op_kernel.cc:1579] OpKernel ('op: \"NdtriGrad\" device_type: \"CPU\" constraint { name: \"T\" allowed_values { list { type: DT_FLOAT } } }') for unknown op: NdtriGrad\n```\nit still seems to work, but i could not find anything online about this specific message. Have you, or anyone else who has tried using pytorch on tpu seen this?\n\nI assumed it was some error as i saw it also when I tried to use pytorch_xla in the recursion comp after pulling the xla docker image, but there i could not get anything to begin training. Thanks again for the guide!",
    "653216": "How do you upload your files in colab?\nOr do you have access to the GCP?\nThanks",
    "652956": "Thank you for posting this Konstantin! Your TPU code is very informative, I'd tried a few times to get pytorch_xla working before without success so seeing a pytorch model training on the tpu is most excellent.  ",
    "652194": "thanks for sharing! like small and powerful kernels.up",
    "674782": "could you tell me how to write \"ip address range\"when creating a TPU node?\nThanks"
  }
}