{
  "id": 154016,
  "title": "For those with limited GPU resources",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/154016",
  "author_name": "Marco Perini",
  "post_date": "2020-05-26T21:20:31.070000",
  "votes": 5,
  "comment_count": 20,
  "views": 0,
  "content": "<p>Not sure if it has been discussed somewhere else, but for those (like me) who do not have crazy GPU hardware for local training of models I've stumbled upon a couple of good resources so far (on top of the 30 GPU hrs per week offered by Kaggle):\n- <a href=\"https://colab.research.google.com/notebooks/intro.ipynb\">https://colab.research.google.com/notebooks/intro.ipynb</a>  quite well known to run notebooks with GPU and TPU supports, although I haven't tried it for this competition yet, and I'm afraid uploading the necessary data might be a big bottleneck\n- <a href=\"https://devcloud.intel.com/oneapi/home/\">https://devcloud.intel.com/oneapi/home/</a> you can register and run both jobs and jupyter notebooks (so far I've only tested the PyTorch support, but it looks like they support Tensorflow as well). They use Intel Xeon Processors, so it will probably be slower than running on decent GPUs, but could still be useful</p>\n\n<p>Hope it might help some of you to gain a bit more computation power. More suggestions are of course welcome!</p>",
  "messages": [
    {
      "id": 862824,
      "postDate": "2020-05-26T21:20:31.070Z",
      "content": "<p>Not sure if it has been discussed somewhere else, but for those (like me) who do not have crazy GPU hardware for local training of models I've stumbled upon a couple of good resources so far (on top of the 30 GPU hrs per week offered by Kaggle):\n- <a href=\"https://colab.research.google.com/notebooks/intro.ipynb\">https://colab.research.google.com/notebooks/intro.ipynb</a>  quite well known to run notebooks with GPU and TPU supports, although I haven't tried it for this competition yet, and I'm afraid uploading the necessary data might be a big bottleneck\n- <a href=\"https://devcloud.intel.com/oneapi/home/\">https://devcloud.intel.com/oneapi/home/</a> you can register and run both jobs and jupyter notebooks (so far I've only tested the PyTorch support, but it looks like they support Tensorflow as well). They use Intel Xeon Processors, so it will probably be slower than running on decent GPUs, but could still be useful</p>\n\n<p>Hope it might help some of you to gain a bit more computation power. More suggestions are of course welcome!</p>",
      "rawMarkdown": "Not sure if it has been discussed somewhere else, but for those (like me) who do not have crazy GPU hardware for local training of models I've stumbled upon a couple of good resources so far (on top of the 30 GPU hrs per week offered by Kaggle):\n- https://colab.research.google.com/notebooks/intro.ipynb  quite well known to run notebooks with GPU and TPU supports, although I haven't tried it for this competition yet, and I'm afraid uploading the necessary data might be a big bottleneck\n- https://devcloud.intel.com/oneapi/home/ you can register and run both jobs and jupyter notebooks (so far I've only tested the PyTorch support, but it looks like they support Tensorflow as well). They use Intel Xeon Processors, so it will probably be slower than running on decent GPUs, but could still be useful\n\nHope it might help some of you to gain a bit more computation power. More suggestions are of course welcome!",
      "votes": 5
    },
    {
      "id": 865905,
      "postDate": "2020-05-29T01:17:30.813Z",
      "content": "<p>Another option for Chinese: We offer free 1080ti for at most 50 hours. Check it out: <a href=\"https://featurize.cn\">https://featurize.cn</a> </p>",
      "rawMarkdown": "Another option for Chinese: We offer free 1080ti for at most 50 hours. Check it out: https://featurize.cn ",
      "replies": [
        {
          "id": 865936,
          "postDate": "2020-05-29T01:58:43.260Z",
          "content": "<p>hey but its in chinese is there any english version</p>",
          "rawMarkdown": "hey but its in chinese is there any english version"
        }
      ]
    },
    {
      "id": 863157,
      "postDate": "2020-05-27T05:32:36.123Z",
      "content": "<p>My GPU is used up. Have to turn to Colab, it is better than I previously thought.. At least faster than my local cpu. The issue of upload data is solved. You can upload to google drive. and use the following 2 lines to access your data. \n<code>\nfrom google.colab import drive\ndrive.mount('/content/drive')\n</code>\nDon't upload original dataset, that is useless. Just upload your processed training dataset.</p>",
      "rawMarkdown": "My GPU is used up. Have to turn to Colab, it is better than I previously thought.. At least faster than my local cpu. The issue of upload data is solved. You can upload to google drive. and use the following 2 lines to access your data. \n```\nfrom google.colab import drive\ndrive.mount('/content/drive')\n```\nDon't upload original dataset, that is useless. Just upload your processed training dataset.\n",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 863247,
          "postDate": "2020-05-27T07:13:57.330Z",
          "content": "<p>good to hear that, thanks for the tip!</p>",
          "rawMarkdown": "good to hear that, thanks for the tip!"
        },
        {
          "id": 863435,
          "postDate": "2020-05-27T10:14:28.883Z",
          "content": "<p>you are welcome!</p>",
          "rawMarkdown": "you are welcome!",
          "isDeleted": true
        },
        {
          "id": 863539,
          "postDate": "2020-05-27T11:44:50.147Z",
          "content": "<p>There is no need to upload any data if it is already in a public Kaggle dataset: you can use its Google Cloud Storage address (GCS) to access the files remotely and then use them on Colab. To read the GCS address, I import the dataset in a Kaggle notebook and run these 2 lines:\n<code>GCS_DS_PATH = KaggleDatasets().get_gcs_path('dataset_name_here')</code>\n<code>print(GCS_DS_PATH)</code></p>\n\n<p>There may be an easier way to do that but I'm new to Kaggle :D\nAlso, you don't have to use the datasets provided by the competition, you can make your own</p>",
          "rawMarkdown": "There is no need to upload any data if it is already in a public Kaggle dataset: you can use its Google Cloud Storage address (GCS) to access the files remotely and then use them on Colab. To read the GCS address, I import the dataset in a Kaggle notebook and run these 2 lines:\n<code>GCS_DS_PATH = KaggleDatasets().get_gcs_path('dataset_name_here')</code>\n<code>print(GCS_DS_PATH)</code>\n\nThere may be an easier way to do that but I'm new to Kaggle :D\nAlso, you don't have to use the datasets provided by the competition, you can make your own",
          "votes": 6
        },
        {
          "id": 863560,
          "postDate": "2020-05-27T11:56:06.500Z",
          "content": "<p>good idea!!</p>",
          "rawMarkdown": "good idea!!",
          "isDeleted": true
        },
        {
          "id": 873892,
          "postDate": "2020-06-04T13:55:30.913Z",
          "content": "<p><a href=\"/pasqualed\">@pasqualed</a> Sorry to disturb you. While I get the gcs-path in kaggle and use it in colab. I met the problem that I can't get the images from <a href=\"https://www.kaggle.com/lopuhin/panda-2020-level-1-2\">panda-level-1-2</a> dataset. But I can read the train.csv(use path: gcs-path/train.csv) in the panda-level-1-2... The path to read images which I set is gcs-path/train _ images/train _ images/ based on I treat the gcs-path as the rootpath of the dataset? Did i miss something?😭 </p>",
          "rawMarkdown": "@pasqualed Sorry to disturb you. While I get the gcs-path in kaggle and use it in colab. I met the problem that I can't get the images from [panda-level-1-2](https://www.kaggle.com/lopuhin/panda-2020-level-1-2) dataset. But I can read the train.csv(use path: gcs-path/train.csv) in the panda-level-1-2... The path to read images which I set is gcs-path/train _ images/train _ images/ based on I treat the gcs-path as the rootpath of the dataset? Did i miss something?😭 "
        },
        {
          "id": 874242,
          "postDate": "2020-06-04T18:29:24.083Z",
          "content": "<p><a href=\"/cnzengshiyuan\">@cnzengshiyuan</a> how do you load the images? With what library/framework?</p>",
          "rawMarkdown": "@cnzengshiyuan how do you load the images? With what library/framework?"
        },
        {
          "id": 874404,
          "postDate": "2020-06-05T00:28:46.923Z",
          "content": "<p><a href=\"/pasqualed\">@pasqualed</a> Sorry for missing it. I use cv2.imread(f'{path}{image_name} _ 1.jpeg'). And it raise error: ! _ src _ empty(), this error I met is all about root error so I ask you for details😄 \npath is gcs-path/train _ images/train _ images/\nActually there are no blanks, but it will cause diaplay problem here so I add blanks before and after '_'.\nExcept the gcs-path is new, the whole codes I just copy from kaggle to colab.\nWhat's more, is there any way to traverse the gcs-path? I try os.walk, os.listdir, glob.glob, but they all failed to traverse the gcs-path....</p>",
          "rawMarkdown": "@pasqualed Sorry for missing it. I use cv2.imread(f'{path}{image_name} _ 1.jpeg'). And it raise error: ! _ src _ empty(), this error I met is all about root error so I ask you for details😄 \npath is gcs-path/train _ images/train _ images/\nActually there are no blanks, but it will cause diaplay problem here so I add blanks before and after '_'.\nExcept the gcs-path is new, the whole codes I just copy from kaggle to colab.\nWhat's more, is there any way to traverse the gcs-path? I try os.walk, os.listdir, glob.glob, but they all failed to traverse the gcs-path...."
        },
        {
          "id": 875380,
          "postDate": "2020-06-05T18:13:28.427Z",
          "content": "<p><a href=\"/cnzengshiyuan\">@cnzengshiyuan</a> I am not sure but it may be because cv2 can't read files remotely or it can't access GCS directly. You can fix this issue by using a tempfile (the code may not work; anyway that's the general idea :D)\n```\nwith urllib.request.urlopen(IMAGE_URL) as url:\n    fx = io.BytesIO(url.read())</p>\n\n<h1>Change tiff with whatever extension you have</h1>\n\n<p>temp=tempfile.NamedTemporaryFile(delete=True, suffix='.tiff')\ntemp.write(fx.read())\nmy_image=cv2.imread(temp.name)\ntemp.close()</p>\n\n<h1>Not sure this is needed</h1>\n\n<p>os.unlink(temp.name)\n```</p>",
          "rawMarkdown": "@cnzengshiyuan I am not sure but it may be because cv2 can't read files remotely or it can't access GCS directly. You can fix this issue by using a tempfile (the code may not work; anyway that's the general idea :D)\n```\nwith urllib.request.urlopen(IMAGE_URL) as url:\n    fx = io.BytesIO(url.read())\n#Change tiff with whatever extension you have\ntemp=tempfile.NamedTemporaryFile(delete=True, suffix='.tiff')\ntemp.write(fx.read())\nmy_image=cv2.imread(temp.name)\ntemp.close()\n#Not sure this is needed\nos.unlink(temp.name)\n```",
          "votes": 1
        },
        {
          "id": 875575,
          "postDate": "2020-06-06T00:00:11.860Z",
          "content": "<p>Thanks a lot! Your codes are tidy. I will learn it😄 </p>",
          "rawMarkdown": "Thanks a lot! Your codes are tidy. I will learn it😄 ",
          "votes": 1
        },
        {
          "id": 875588,
          "postDate": "2020-06-06T00:51:46.843Z",
          "content": "<p>Try setting delete=False in the NamedTemporaryFile if it still doesn't work and let me know if you eventually manage to read the images from GCS with cv2 :D</p>",
          "rawMarkdown": "Try setting delete=False in the NamedTemporaryFile if it still doesn't work and let me know if you eventually manage to read the images from GCS with cv2 :D",
          "votes": 1
        },
        {
          "id": 876042,
          "postDate": "2020-06-06T11:28:21.233Z",
          "content": "<p>Sorry for late reply(today is a little busy...). I try it but the urllib raises an urlerror, said that gs is an unknown type. I will try some other ways, if they can work, I will give you feedback here. Thank you for your sharing💯 </p>",
          "rawMarkdown": "Sorry for late reply(today is a little busy...). I try it but the urllib raises an urlerror, said that gs is an unknown type. I will try some other ways, if they can work, I will give you feedback here. Thank you for your sharing💯 ",
          "votes": 1
        },
        {
          "id": 876067,
          "postDate": "2020-06-06T12:11:57.320Z",
          "content": "<p>I am sure I managed to make it work somehow, even though I was not using cv2. Try installing and importing gcsfs</p>",
          "rawMarkdown": "I am sure I managed to make it work somehow, even though I was not using cv2. Try installing and importing gcsfs",
          "votes": 1
        },
        {
          "id": 876081,
          "postDate": "2020-06-06T12:27:34.777Z",
          "content": "<p>Yes, I have installed and imported gcsfs. It doesn't matter, as download and unzip the level-1-2 dataset just cost less than 10 mins, it's acceptable. Thank you for your kindness and patience. And good to know a new way to access dataset :D</p>",
          "rawMarkdown": "Yes, I have installed and imported gcsfs. It doesn't matter, as download and unzip the level-1-2 dataset just cost less than 10 mins, it's acceptable. Thank you for your kindness and patience. And good to know a new way to access dataset :D"
        },
        {
          "id": 885175,
          "postDate": "2020-06-14T01:38:39.780Z",
          "content": "<p><a href=\"/cnzengshiyuan\">@cnzengshiyuan</a> you can convert the URL like this <code>http://storage.googleapis.com/'+BUCKET_ID+PATH</code>. <code>BUCKET_ID</code> is the part of the GC URL after <code>gc://</code>, <code>PATH</code> is the path within the bucket, for example <code>/train_images/0005f7aaab2800f6170c399693a96917.tiff</code>. I hope it helps!</p>",
          "rawMarkdown": "@cnzengshiyuan you can convert the URL like this `http://storage.googleapis.com/'+BUCKET_ID+PATH`. `BUCKET_ID` is the part of the GC URL after `gc://`, `PATH ` is the path within the bucket, for example `/train_images/0005f7aaab2800f6170c399693a96917.tiff`. I hope it helps!",
          "votes": 1
        },
        {
          "id": 885196,
          "postDate": "2020-06-14T02:15:36.130Z",
          "content": "<p><a href=\"/pasqualed\">@pasqualed</a> Thanks for your kindness! The whole url 'http://......' you mentioned is able to get the .tiff by clicking it, but still failed in reading the .tiff through skimag.io.MultiImage.\nWhat's more, as the all methods you mentioned above is opening a door for me to know some new ways, and I think it's doesn't need to be stuck in accessing original dataset(of course, knowing the reason of failure is very important) as we can download and unzip the level-1-2 to use.\nThanks again and happy to discuss with you 😄 </p>",
          "rawMarkdown": "@pasqualed Thanks for your kindness! The whole url 'http://......' you mentioned is able to get the .tiff by clicking it, but still failed in reading the .tiff through skimag.io.MultiImage.\nWhat's more, as the all methods you mentioned above is opening a door for me to know some new ways, and I think it's doesn't need to be stuck in accessing original dataset(of course, knowing the reason of failure is very important) as we can download and unzip the level-1-2 to use.\nThanks again and happy to discuss with you 😄 ",
          "votes": 1
        },
        {
          "id": 885206,
          "postDate": "2020-06-14T02:32:57.007Z",
          "content": "<p><a href=\"/cnzengshiyuan\">@cnzengshiyuan</a> I use the http URL to create a tempfile (as in my previous post, <code>IMAGE_URL</code> is an http URL) and then read the image using <code>skimage.io.MultiImage(tempfile.name)</code>. I think at the moment you get an error because <code>skimage.io.MultiImage</code> requires a local path to work.</p>",
          "rawMarkdown": "@cnzengshiyuan I use the http URL to create a tempfile (as in my previous post, `IMAGE_URL` is an http URL) and then read the image using `skimage.io.MultiImage(tempfile.name)`. I think at the moment you get an error because `skimage.io.MultiImage` requires a local path to work.",
          "votes": 1
        },
        {
          "id": 885211,
          "postDate": "2020-06-14T02:46:58.687Z",
          "content": "<p><a href=\"/pasqualed\">@pasqualed</a> Yes, you are right. And what I'm afraid is that\n- 1 if I access the whole dateset and save all of them, the limited memory of colab maybe forbid it.\n- 2 If I access the data when I need it(like call dataset[i]), and them delete it for limited memory. I will do so many times of writing a picture to disk. I think it will waste time because the speed to do io in disj-level is too slow.</p>\n\n<p>way-1 do writing files once but the memory may not enough as whole original dataset is 380g.\nway-2 do so many times of io.\nAnd after thinking about the trade-off above, I use downloading and unzip😂 </p>",
          "rawMarkdown": "@pasqualed Yes, you are right. And what I'm afraid is that\n- 1 if I access the whole dateset and save all of them, the limited memory of colab maybe forbid it.\n- 2 If I access the data when I need it(like call dataset[i]), and them delete it for limited memory. I will do so many times of writing a picture to disk. I think it will waste time because the speed to do io in disj-level is too slow.\n\nway-1 do writing files once but the memory may not enough as whole original dataset is 380g.\nway-2 do so many times of io.\nAnd after thinking about the trade-off above, I use downloading and unzip😂 "
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 865905,
      "author_name": "Chenglu",
      "author_url": "",
      "post_date": "2020-05-29T01:17:30.813000",
      "content": "<p>Another option for Chinese: We offer free 1080ti for at most 50 hours. Check it out: <a href=\"https://featurize.cn\">https://featurize.cn</a> </p>",
      "votes": 0,
      "replies": [
        {
          "id": 865936,
          "author_name": "killermachine_proton",
          "author_url": "",
          "post_date": "2020-05-29T01:58:43.260000",
          "content": "<p>hey but its in chinese is there any english version</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 863157,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-05-27T05:32:36.123000",
      "content": "<p>My GPU is used up. Have to turn to Colab, it is better than I previously thought.. At least faster than my local cpu. The issue of upload data is solved. You can upload to google drive. and use the following 2 lines to access your data. \n<code>\nfrom google.colab import drive\ndrive.mount('/content/drive')\n</code>\nDon't upload original dataset, that is useless. Just upload your processed training dataset.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 863247,
          "author_name": "Marco Perini",
          "author_url": "",
          "post_date": "2020-05-27T07:13:57.330000",
          "content": "<p>good to hear that, thanks for the tip!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 863435,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-05-27T10:14:28.883000",
          "content": "<p>you are welcome!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 863539,
          "author_name": "Pasquale",
          "author_url": "",
          "post_date": "2020-05-27T11:44:50.147000",
          "content": "<p>There is no need to upload any data if it is already in a public Kaggle dataset: you can use its Google Cloud Storage address (GCS) to access the files remotely and then use them on Colab. To read the GCS address, I import the dataset in a Kaggle notebook and run these 2 lines:\n<code>GCS_DS_PATH = KaggleDatasets().get_gcs_path('dataset_name_here')</code>\n<code>print(GCS_DS_PATH)</code></p>\n\n<p>There may be an easier way to do that but I'm new to Kaggle :D\nAlso, you don't have to use the datasets provided by the competition, you can make your own</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 863560,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-05-27T11:56:06.500000",
          "content": "<p>good idea!!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 873892,
          "author_name": "Shiyuan Zeng",
          "author_url": "",
          "post_date": "2020-06-04T13:55:30.913000",
          "content": "<p><a href=\"/pasqualed\">@pasqualed</a> Sorry to disturb you. While I get the gcs-path in kaggle and use it in colab. I met the problem that I can't get the images from <a href=\"https://www.kaggle.com/lopuhin/panda-2020-level-1-2\">panda-level-1-2</a> dataset. But I can read the train.csv(use path: gcs-path/train.csv) in the panda-level-1-2... The path to read images which I set is gcs-path/train _ images/train _ images/ based on I treat the gcs-path as the rootpath of the dataset? Did i miss something?😭 </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 874242,
          "author_name": "Pasquale",
          "author_url": "",
          "post_date": "2020-06-04T18:29:24.083000",
          "content": "<p><a href=\"/cnzengshiyuan\">@cnzengshiyuan</a> how do you load the images? With what library/framework?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 874404,
          "author_name": "Shiyuan Zeng",
          "author_url": "",
          "post_date": "2020-06-05T00:28:46.923000",
          "content": "<p><a href=\"/pasqualed\">@pasqualed</a> Sorry for missing it. I use cv2.imread(f'{path}{image_name} _ 1.jpeg'). And it raise error: ! _ src _ empty(), this error I met is all about root error so I ask you for details😄 \npath is gcs-path/train _ images/train _ images/\nActually there are no blanks, but it will cause diaplay problem here so I add blanks before and after '_'.\nExcept the gcs-path is new, the whole codes I just copy from kaggle to colab.\nWhat's more, is there any way to traverse the gcs-path? I try os.walk, os.listdir, glob.glob, but they all failed to traverse the gcs-path....</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 875380,
          "author_name": "Pasquale",
          "author_url": "",
          "post_date": "2020-06-05T18:13:28.427000",
          "content": "<p><a href=\"/cnzengshiyuan\">@cnzengshiyuan</a> I am not sure but it may be because cv2 can't read files remotely or it can't access GCS directly. You can fix this issue by using a tempfile (the code may not work; anyway that's the general idea :D)\n```\nwith urllib.request.urlopen(IMAGE_URL) as url:\n    fx = io.BytesIO(url.read())</p>\n\n<h1>Change tiff with whatever extension you have</h1>\n\n<p>temp=tempfile.NamedTemporaryFile(delete=True, suffix='.tiff')\ntemp.write(fx.read())\nmy_image=cv2.imread(temp.name)\ntemp.close()</p>\n\n<h1>Not sure this is needed</h1>\n\n<p>os.unlink(temp.name)\n```</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 875575,
          "author_name": "Shiyuan Zeng",
          "author_url": "",
          "post_date": "2020-06-06T00:00:11.860000",
          "content": "<p>Thanks a lot! Your codes are tidy. I will learn it😄 </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 875588,
          "author_name": "Pasquale",
          "author_url": "",
          "post_date": "2020-06-06T00:51:46.843000",
          "content": "<p>Try setting delete=False in the NamedTemporaryFile if it still doesn't work and let me know if you eventually manage to read the images from GCS with cv2 :D</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 876042,
          "author_name": "Shiyuan Zeng",
          "author_url": "",
          "post_date": "2020-06-06T11:28:21.233000",
          "content": "<p>Sorry for late reply(today is a little busy...). I try it but the urllib raises an urlerror, said that gs is an unknown type. I will try some other ways, if they can work, I will give you feedback here. Thank you for your sharing💯 </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 876067,
          "author_name": "Pasquale",
          "author_url": "",
          "post_date": "2020-06-06T12:11:57.320000",
          "content": "<p>I am sure I managed to make it work somehow, even though I was not using cv2. Try installing and importing gcsfs</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 876081,
          "author_name": "Shiyuan Zeng",
          "author_url": "",
          "post_date": "2020-06-06T12:27:34.777000",
          "content": "<p>Yes, I have installed and imported gcsfs. It doesn't matter, as download and unzip the level-1-2 dataset just cost less than 10 mins, it's acceptable. Thank you for your kindness and patience. And good to know a new way to access dataset :D</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 885175,
          "author_name": "Pasquale",
          "author_url": "",
          "post_date": "2020-06-14T01:38:39.780000",
          "content": "<p><a href=\"/cnzengshiyuan\">@cnzengshiyuan</a> you can convert the URL like this <code>http://storage.googleapis.com/'+BUCKET_ID+PATH</code>. <code>BUCKET_ID</code> is the part of the GC URL after <code>gc://</code>, <code>PATH</code> is the path within the bucket, for example <code>/train_images/0005f7aaab2800f6170c399693a96917.tiff</code>. I hope it helps!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 885196,
          "author_name": "Shiyuan Zeng",
          "author_url": "",
          "post_date": "2020-06-14T02:15:36.130000",
          "content": "<p><a href=\"/pasqualed\">@pasqualed</a> Thanks for your kindness! The whole url 'http://......' you mentioned is able to get the .tiff by clicking it, but still failed in reading the .tiff through skimag.io.MultiImage.\nWhat's more, as the all methods you mentioned above is opening a door for me to know some new ways, and I think it's doesn't need to be stuck in accessing original dataset(of course, knowing the reason of failure is very important) as we can download and unzip the level-1-2 to use.\nThanks again and happy to discuss with you 😄 </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 885206,
          "author_name": "Pasquale",
          "author_url": "",
          "post_date": "2020-06-14T02:32:57.007000",
          "content": "<p><a href=\"/cnzengshiyuan\">@cnzengshiyuan</a> I use the http URL to create a tempfile (as in my previous post, <code>IMAGE_URL</code> is an http URL) and then read the image using <code>skimage.io.MultiImage(tempfile.name)</code>. I think at the moment you get an error because <code>skimage.io.MultiImage</code> requires a local path to work.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 885211,
          "author_name": "Shiyuan Zeng",
          "author_url": "",
          "post_date": "2020-06-14T02:46:58.687000",
          "content": "<p><a href=\"/pasqualed\">@pasqualed</a> Yes, you are right. And what I'm afraid is that\n- 1 if I access the whole dateset and save all of them, the limited memory of colab maybe forbid it.\n- 2 If I access the data when I need it(like call dataset[i]), and them delete it for limited memory. I will do so many times of writing a picture to disk. I think it will waste time because the speed to do io in disj-level is too slow.</p>\n\n<p>way-1 do writing files once but the memory may not enough as whole original dataset is 380g.\nway-2 do so many times of io.\nAnd after thinking about the trade-off above, I use downloading and unzip😂 </p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "862824": "Not sure if it has been discussed somewhere else, but for those (like me) who do not have crazy GPU hardware for local training of models I've stumbled upon a couple of good resources so far (on top of the 30 GPU hrs per week offered by Kaggle):\n- https://colab.research.google.com/notebooks/intro.ipynb  quite well known to run notebooks with GPU and TPU supports, although I haven't tried it for this competition yet, and I'm afraid uploading the necessary data might be a big bottleneck\n- https://devcloud.intel.com/oneapi/home/ you can register and run both jobs and jupyter notebooks (so far I've only tested the PyTorch support, but it looks like they support Tensorflow as well). They use Intel Xeon Processors, so it will probably be slower than running on decent GPUs, but could still be useful\n\nHope it might help some of you to gain a bit more computation power. More suggestions are of course welcome!",
    "865905": "Another option for Chinese: We offer free 1080ti for at most 50 hours. Check it out: https://featurize.cn ",
    "863157": "My GPU is used up. Have to turn to Colab, it is better than I previously thought.. At least faster than my local cpu. The issue of upload data is solved. You can upload to google drive. and use the following 2 lines to access your data. \n```\nfrom google.colab import drive\ndrive.mount('/content/drive')\n```\nDon't upload original dataset, that is useless. Just upload your processed training dataset.\n"
  }
}