{
  "id": 111999,
  "title": "Discussing challenges in Training Single Model",
  "url": "/competitions/rsna-intracranial-hemorrhage-detection/discussion/111999",
  "author_name": "Ajay Chauhan",
  "post_date": "2019-10-10T08:09:49.514000",
  "votes": 8,
  "comment_count": 17,
  "views": 0,
  "content": "<p>As per my understanding biggest challenge for this competition is to work around with huge data to train and test.</p>\n\n<p>Converting images to png or jpg cost much time and most of the users are using converted images from other keggle used and that IS THE BEST PART of Kaggle :)</p>\n\n<p>Now about training the model: Data is so huge that even very simple models are taking 3-4 hours for a single epoch. Google Colab does not help much and crashed. GPU is limited.. Bad timing from Kaggle team :)\nAll the suggestions on reducing training part are welcome and appreciated .</p>\n\n<p>Endembling seems far distant as of now because it seems difficult to train a single model.</p>\n\n<p>All suggestions and discussions are welcome..</p>",
  "messages": [
    {
      "id": 646073,
      "postDate": "2019-10-10T20:21:29.513Z",
      "content": "<p>that's easy - start with lightweight models, there are only few of them to try. don't go 512 images - there are already downsampled images for your needs and I can prove they are good enough. find out the best way to preprocess input data and find rich augmentations pipeline for the problem (or take any pipeline from papers mentioned). probably, to win you need more resources but I believe you have everything to achieve good results (lets say top-100)</p>",
      "rawMarkdown": "that's easy - start with lightweight models, there are only few of them to try. don't go 512 images - there are already downsampled images for your needs and I can prove they are good enough. find out the best way to preprocess input data and find rich augmentations pipeline for the problem (or take any pipeline from papers mentioned). probably, to win you need more resources but I believe you have everything to achieve good results (lets say top-100)",
      "votes": 14
    },
    {
      "id": 645554,
      "postDate": "2019-10-10T08:09:49.513Z",
      "content": "<p>As per my understanding biggest challenge for this competition is to work around with huge data to train and test.</p>\n\n<p>Converting images to png or jpg cost much time and most of the users are using converted images from other keggle used and that IS THE BEST PART of Kaggle :)</p>\n\n<p>Now about training the model: Data is so huge that even very simple models are taking 3-4 hours for a single epoch. Google Colab does not help much and crashed. GPU is limited.. Bad timing from Kaggle team :)\nAll the suggestions on reducing training part are welcome and appreciated .</p>\n\n<p>Endembling seems far distant as of now because it seems difficult to train a single model.</p>\n\n<p>All suggestions and discussions are welcome..</p>",
      "rawMarkdown": "As per my understanding biggest challenge for this competition is to work around with huge data to train and test.\n \nConverting images to png or jpg cost much time and most of the users are using converted images from other keggle used and that IS THE BEST PART of Kaggle :)\n\nNow about training the model: Data is so huge that even very simple models are taking 3-4 hours for a single epoch. Google Colab does not help much and crashed. GPU is limited.. Bad timing from Kaggle team :)\nAll the suggestions on reducing training part are welcome and appreciated .\n\nEndembling seems far distant as of now because it seems difficult to train a single model.\n\nAll suggestions and discussions are welcome..",
      "votes": 8
    },
    {
      "id": 645750,
      "postDate": "2019-10-10T12:56:04.793Z",
      "content": "<p>I have found a solution to your 'gpu crash' you may be keen in trying. Actually, the colab K80's are preemptible meaning your work can get interrupted at any time. Even on GCP, if I select a V100 preemptible instance, I get unpredictable instance disconnects.</p>\n\n<p>So my solution to combat the interrupts and perhaps the long training time per epoch is to create a checkpoint that restores my model's training. Particularly, you'd want to restore the checkpoint that gives you the best validation loss. However, I think for your purposes having any checkpoint will be helpful since it takes you 3-4hrs per epoch. So you'll need to connect your gdrive to colab in order to obtain a copy of the checksum files created on colab. Luckily this is quite simple.</p>\n\n<p>&gt; from google.colab import drive\ndrive.mount('/content/gdrive/')</p>\n\n<p>then click on the url and then authorize the connection</p>\n\n<p>&gt; ### save your model\n<code>model_save_name = \"your_model_name.pt\"</code>\n<code>path = f\"content/gdrive/My\\ Drive/{model_save_name}\"</code>\n<code>torch.save(model.state_dict(), path)</code></p>\n\n<p>or that path doesn't work for you, try</p>\n\n<p>&gt; <code>path = \"gdrive/My Drive/your_model_name.pt\"</code></p>\n\n<p>for keras, you can change the file extension to .hdf5</p>\n\n<p><br></p>\n\n<h3>loading the checkpoint</h3>\n\n<p>just mount your gdrive, locate the filepath and then </p>\n\n<p>&gt; ###PyTorch\n<code>model.load.state_dict(torch.load(\"file_path/file_name.pt\"))</code></p>\n\n<p>or</p>\n\n<p>&gt; ###Keras\n<code>model.load_weights(\"file_path/file_name.hdf5\")</code></p>",
      "rawMarkdown": "I have found a solution to your 'gpu crash' you may be keen in trying. Actually, the colab K80's are preemptible meaning your work can get interrupted at any time. Even on GCP, if I select a V100 preemptible instance, I get unpredictable instance disconnects.\n\nSo my solution to combat the interrupts and perhaps the long training time per epoch is to create a checkpoint that restores my model's training. Particularly, you'd want to restore the checkpoint that gives you the best validation loss. However, I think for your purposes having any checkpoint will be helpful since it takes you 3-4hrs per epoch. So you'll need to connect your gdrive to colab in order to obtain a copy of the checksum files created on colab. Luckily this is quite simple.\n\n&gt; from google.colab import drive\ndrive.mount('/content/gdrive/')\n\nthen click on the url and then authorize the connection\n\n&gt; ### save your model\n`model_save_name = \"your_model_name.pt\"`\n`path = f\"content/gdrive/My\\ Drive/{model_save_name}\"`\n`torch.save(model.state_dict(), path)`\n\nor that path doesn't work for you, try\n\n&gt; `path = \"gdrive/My Drive/your_model_name.pt\"`\n\nfor keras, you can change the file extension to .hdf5\n\n<br>\n\n### loading the checkpoint\njust mount your gdrive, locate the filepath and then \n\n&gt; ###PyTorch\n`model.load.state_dict(torch.load(\"file_path/file_name.pt\"))`\n\nor\n\n&gt; ###Keras\n`model.load_weights(\"file_path/file_name.hdf5\")`\n",
      "votes": 3
    },
    {
      "id": 645615,
      "postDate": "2019-10-10T09:41:50.537Z",
      "content": "<p>Don't know how to implement this, but maybe kaggle team could consider redistributing quota from those who don't actually use it. I have local gpu, so I don't need my quota. Maybe they could just create a special thread, where everyone who don't need his quota in a particular competition could declare such intentions. As far as I could roughly estimate, at least 30% of participants are training models locally. This could be a substantial help for others </p>",
      "rawMarkdown": "Don't know how to implement this, but maybe kaggle team could consider redistributing quota from those who don't actually use it. I have local gpu, so I don't need my quota. Maybe they could just create a special thread, where everyone who don't need his quota in a particular competition could declare such intentions. As far as I could roughly estimate, at least 30% of participants are training models locally. This could be a substantial help for others ",
      "votes": 3,
      "replies": [
        {
          "id": 645803,
          "postDate": "2019-10-10T13:54:50.860Z",
          "content": "<p>Strange, why do you assume that unclaimed GPU quotas result in hardware sitting idle? I would assume unclaimed quotas are already factored in, and all GPUs that Kaggle has are already utilized to close to 100%.</p>",
          "rawMarkdown": "Strange, why do you assume that unclaimed GPU quotas result in hardware sitting idle? I would assume unclaimed quotas are already factored in, and all GPUs that Kaggle has are already utilized to close to 100%.",
          "votes": 2
        },
        {
          "id": 645884,
          "postDate": "2019-10-10T15:27:38.193Z",
          "content": "<p>It well may be already accounted for. It is just an assumption, but I still think they should have some reserves to meet demand. Maybe Kaggle team could manage resources more precisely, if they'll have more info.</p>",
          "rawMarkdown": "It well may be already accounted for. It is just an assumption, but I still think they should have some reserves to meet demand. Maybe Kaggle team could manage resources more precisely, if they'll have more info.",
          "votes": 1
        },
        {
          "id": 646038,
          "postDate": "2019-10-10T19:12:48.727Z",
          "content": "<p>They probably took this into account =) If not than its good opportunity to design a challenge which will predict gpu loads =) </p>",
          "rawMarkdown": "They probably took this into account =) If not than its good opportunity to design a challenge which will predict gpu loads =) ",
          "votes": 1
        },
        {
          "id": 646044,
          "postDate": "2019-10-10T19:20:06.637Z",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> :) suggestion to be passed to kaggle team :)</p>",
          "rawMarkdown": "@drhabib :) suggestion to be passed to kaggle team :)",
          "votes": 1
        }
      ]
    },
    {
      "id": 645800,
      "postDate": "2019-10-10T13:51:00.230Z",
      "content": "<p>Even with a GPU training is slow. I am going to start training with all the positive cases and under sample the normal scans.</p>",
      "rawMarkdown": "Even with a GPU training is slow. I am going to start training with all the positive cases and under sample the normal scans.",
      "votes": 1
    },
    {
      "id": 645741,
      "postDate": "2019-10-10T12:48:40.237Z",
      "content": "<p>In this competition you could get credit for Google Cloud. It's late however...\nBut anyway you're right we have a lot of training data. I leave the idea to train by folds, since it's too long. :(\nAlthough it seems you don't need a lot of epochs here.\nThere was a similar topic how to accelerate:\n<a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/110671#latest-638850\">https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/110671#latest-638850</a></p>",
      "rawMarkdown": "In this competition you could get credit for Google Cloud. It's late however...\nBut anyway you're right we have a lot of training data. I leave the idea to train by folds, since it's too long. :(\nAlthough it seems you don't need a lot of epochs here.\nThere was a similar topic how to accelerate:\n[https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/110671#latest-638850](https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/110671#latest-638850)",
      "votes": 1,
      "replies": [
        {
          "id": 645802,
          "postDate": "2019-10-10T13:52:27.710Z",
          "content": "<p>I am training using 5 folds CV and it is super slow but I get nice improvements in the LB score ~0.004</p>",
          "rawMarkdown": "I am training using 5 folds CV and it is super slow but I get nice improvements in the LB score ~0.004",
          "votes": 2
        },
        {
          "id": 645854,
          "postDate": "2019-10-10T14:46:57.897Z",
          "content": "<p>Thanks for info!\nToday I thought I should explore more different models. So probably ensembing (without folds) will be good enough too.\nHowever maybe I'll return to folds later.</p>",
          "rawMarkdown": "Thanks for info!\nToday I thought I should explore more different models. So probably ensembing (without folds) will be good enough too.\nHowever maybe I'll return to folds later.",
          "votes": 1
        }
      ]
    },
    {
      "id": 645568,
      "postDate": "2019-10-10T08:23:59.813Z",
      "content": "<p>I also have lost 3-4 days on Colab and it did not help. Model gets crashed using TPU and GPU even after reducing  to usbatch size..planning to use AWS</p>",
      "rawMarkdown": "I also have lost 3-4 days on Colab and it did not help. Model gets crashed using TPU and GPU even after reducing  to usbatch size..planning to use AWS",
      "votes": 1,
      "replies": [
        {
          "id": 646621,
          "postDate": "2019-10-11T13:41:21.023Z",
          "content": "<p>how do you download datasets to colab ,the colab is very small\nand I can not download datasets directly to google drive.</p>",
          "rawMarkdown": "how do you download datasets to colab ,the colab is very small\nand I can not download datasets directly to google drive."
        }
      ]
    },
    {
      "id": 645563,
      "postDate": "2019-10-10T08:18:54.127Z",
      "content": "<p>i have been trying on colab since last 3 days and it keeps crashing over and over again,i am really very frustrated,model that works in kaggle doesn't work in colab and it keeps crashing,i reduced the batch size and it still keeps crashing,kaggle gpu quota is hurting so badly,i love deep learning a lot,gpu quota is killing me :'(</p>",
      "rawMarkdown": "i have been trying on colab since last 3 days and it keeps crashing over and over again,i am really very frustrated,model that works in kaggle doesn't work in colab and it keeps crashing,i reduced the batch size and it still keeps crashing,kaggle gpu quota is hurting so badly,i love deep learning a lot,gpu quota is killing me :'(",
      "votes": 1
    },
    {
      "id": 646054,
      "postDate": "2019-10-10T19:51:17.713Z",
      "content": "<p>I would like to know which model is better to be used as a single model...efficientNet or Resnet or Densenet or Exception etc.\nDue to big size of data trying everyone is an impossible challenge.\nAny help would be GREAT...</p>",
      "rawMarkdown": "I would like to know which model is better to be used as a single model...efficientNet or Resnet or Densenet or Exception etc.\nDue to big size of data trying everyone is an impossible challenge.\nAny help would be GREAT...\n",
      "replies": [
        {
          "id": 646728,
          "postDate": "2019-10-11T16:30:43.333Z",
          "content": "<p>Try efficientnet b0 ..one epoch would complete within an hour. <br>\nAnd train another model like ResNeXt ( <a href=\"https://www.kaggle.com/taindow/pytorch-resnext-101-32x8d-benchmark\">https://www.kaggle.com/taindow/pytorch-resnext-101-32x8d-benchmark</a> ). Results from these two models differ a lot and good for simple averaging.</p>\n\n<p>Use Image size 224. You can get decent results within 0.075-0.08 with this resolution.     </p>\n\n<p>Regarding hardware..if you have never used GCP...you can get free 300$ worth of compute for one year. \nFire up a pre-emptible P100 GPU and save intermediate checkpoints of the model so even if the GPU is pre-emptied you have still model checkpoints and can recontinue from there. Save model and optimizer states. A pre-emptible GPU would be 3 times cheaper than normal GPU. </p>",
          "rawMarkdown": "Try efficientnet b0 ..one epoch would complete within an hour.  \nAnd train another model like ResNeXt ( https://www.kaggle.com/taindow/pytorch-resnext-101-32x8d-benchmark ). Results from these two models differ a lot and good for simple averaging.\n\nUse Image size 224. You can get decent results within 0.075-0.08 with this resolution.     \n\nRegarding hardware..if you have never used GCP...you can get free 300$ worth of compute for one year. \nFire up a pre-emptible P100 GPU and save intermediate checkpoints of the model so even if the GPU is pre-emptied you have still model checkpoints and can recontinue from there. Save model and optimizer states. A pre-emptible GPU would be 3 times cheaper than normal GPU. ",
          "votes": 3
        }
      ]
    },
    {
      "id": 645620,
      "postDate": "2019-10-10T09:53:01.050Z",
      "content": "<p>Rightly said <a href=\"/cateek\">@cateek</a> </p>",
      "rawMarkdown": "Rightly said @cateek "
    }
  ],
  "comments": [
    {
      "id": 646073,
      "author_name": "Oleg Yaroshevskiy",
      "author_url": "",
      "post_date": "2019-10-10T20:21:29.513000",
      "content": "<p>that's easy - start with lightweight models, there are only few of them to try. don't go 512 images - there are already downsampled images for your needs and I can prove they are good enough. find out the best way to preprocess input data and find rich augmentations pipeline for the problem (or take any pipeline from papers mentioned). probably, to win you need more resources but I believe you have everything to achieve good results (lets say top-100)</p>",
      "votes": 14,
      "replies": []
    },
    {
      "id": 645750,
      "author_name": "Tim Yee",
      "author_url": "",
      "post_date": "2019-10-10T12:56:04.793000",
      "content": "<p>I have found a solution to your 'gpu crash' you may be keen in trying. Actually, the colab K80's are preemptible meaning your work can get interrupted at any time. Even on GCP, if I select a V100 preemptible instance, I get unpredictable instance disconnects.</p>\n\n<p>So my solution to combat the interrupts and perhaps the long training time per epoch is to create a checkpoint that restores my model's training. Particularly, you'd want to restore the checkpoint that gives you the best validation loss. However, I think for your purposes having any checkpoint will be helpful since it takes you 3-4hrs per epoch. So you'll need to connect your gdrive to colab in order to obtain a copy of the checksum files created on colab. Luckily this is quite simple.</p>\n\n<p>&gt; from google.colab import drive\ndrive.mount('/content/gdrive/')</p>\n\n<p>then click on the url and then authorize the connection</p>\n\n<p>&gt; ### save your model\n<code>model_save_name = \"your_model_name.pt\"</code>\n<code>path = f\"content/gdrive/My\\ Drive/{model_save_name}\"</code>\n<code>torch.save(model.state_dict(), path)</code></p>\n\n<p>or that path doesn't work for you, try</p>\n\n<p>&gt; <code>path = \"gdrive/My Drive/your_model_name.pt\"</code></p>\n\n<p>for keras, you can change the file extension to .hdf5</p>\n\n<p><br></p>\n\n<h3>loading the checkpoint</h3>\n\n<p>just mount your gdrive, locate the filepath and then </p>\n\n<p>&gt; ###PyTorch\n<code>model.load.state_dict(torch.load(\"file_path/file_name.pt\"))</code></p>\n\n<p>or</p>\n\n<p>&gt; ###Keras\n<code>model.load_weights(\"file_path/file_name.hdf5\")</code></p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 645615,
      "author_name": "Eek The Cat",
      "author_url": "",
      "post_date": "2019-10-10T09:41:50.537000",
      "content": "<p>Don't know how to implement this, but maybe kaggle team could consider redistributing quota from those who don't actually use it. I have local gpu, so I don't need my quota. Maybe they could just create a special thread, where everyone who don't need his quota in a particular competition could declare such intentions. As far as I could roughly estimate, at least 30% of participants are training models locally. This could be a substantial help for others </p>",
      "votes": 3,
      "replies": [
        {
          "id": 645803,
          "author_name": "nosound",
          "author_url": "",
          "post_date": "2019-10-10T13:54:50.860000",
          "content": "<p>Strange, why do you assume that unclaimed GPU quotas result in hardware sitting idle? I would assume unclaimed quotas are already factored in, and all GPUs that Kaggle has are already utilized to close to 100%.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 645884,
          "author_name": "Eek The Cat",
          "author_url": "",
          "post_date": "2019-10-10T15:27:38.193000",
          "content": "<p>It well may be already accounted for. It is just an assumption, but I still think they should have some reserves to meet demand. Maybe Kaggle team could manage resources more precisely, if they'll have more info.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 646038,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-10-10T19:12:48.727000",
          "content": "<p>They probably took this into account =) If not than its good opportunity to design a challenge which will predict gpu loads =) </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 646044,
          "author_name": "Ajay Chauhan",
          "author_url": "",
          "post_date": "2019-10-10T19:20:06.637000",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> :) suggestion to be passed to kaggle team :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 645800,
      "author_name": "Maria Wellen",
      "author_url": "",
      "post_date": "2019-10-10T13:51:00.230000",
      "content": "<p>Even with a GPU training is slow. I am going to start training with all the positive cases and under sample the normal scans.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 645741,
      "author_name": "Sergey Zlobin",
      "author_url": "",
      "post_date": "2019-10-10T12:48:40.237000",
      "content": "<p>In this competition you could get credit for Google Cloud. It's late however...\nBut anyway you're right we have a lot of training data. I leave the idea to train by folds, since it's too long. :(\nAlthough it seems you don't need a lot of epochs here.\nThere was a similar topic how to accelerate:\n<a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/110671#latest-638850\">https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/110671#latest-638850</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 645802,
          "author_name": "Maria Wellen",
          "author_url": "",
          "post_date": "2019-10-10T13:52:27.710000",
          "content": "<p>I am training using 5 folds CV and it is super slow but I get nice improvements in the LB score ~0.004</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 645854,
          "author_name": "Sergey Zlobin",
          "author_url": "",
          "post_date": "2019-10-10T14:46:57.897000",
          "content": "<p>Thanks for info!\nToday I thought I should explore more different models. So probably ensembing (without folds) will be good enough too.\nHowever maybe I'll return to folds later.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 645568,
      "author_name": "Ajay Chauhan",
      "author_url": "",
      "post_date": "2019-10-10T08:23:59.813000",
      "content": "<p>I also have lost 3-4 days on Colab and it did not help. Model gets crashed using TPU and GPU even after reducing  to usbatch size..planning to use AWS</p>",
      "votes": 1,
      "replies": [
        {
          "id": 646621,
          "author_name": "pupil3",
          "author_url": "",
          "post_date": "2019-10-11T13:41:21.023000",
          "content": "<p>how do you download datasets to colab ,the colab is very small\nand I can not download datasets directly to google drive.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 645563,
      "author_name": "Mobassir",
      "author_url": "",
      "post_date": "2019-10-10T08:18:54.127000",
      "content": "<p>i have been trying on colab since last 3 days and it keeps crashing over and over again,i am really very frustrated,model that works in kaggle doesn't work in colab and it keeps crashing,i reduced the batch size and it still keeps crashing,kaggle gpu quota is hurting so badly,i love deep learning a lot,gpu quota is killing me :'(</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 646054,
      "author_name": "Ajay Chauhan",
      "author_url": "",
      "post_date": "2019-10-10T19:51:17.713000",
      "content": "<p>I would like to know which model is better to be used as a single model...efficientNet or Resnet or Densenet or Exception etc.\nDue to big size of data trying everyone is an impossible challenge.\nAny help would be GREAT...</p>",
      "votes": 0,
      "replies": [
        {
          "id": 646728,
          "author_name": "Abhilash Awasthi",
          "author_url": "",
          "post_date": "2019-10-11T16:30:43.333000",
          "content": "<p>Try efficientnet b0 ..one epoch would complete within an hour. <br>\nAnd train another model like ResNeXt ( <a href=\"https://www.kaggle.com/taindow/pytorch-resnext-101-32x8d-benchmark\">https://www.kaggle.com/taindow/pytorch-resnext-101-32x8d-benchmark</a> ). Results from these two models differ a lot and good for simple averaging.</p>\n\n<p>Use Image size 224. You can get decent results within 0.075-0.08 with this resolution.     </p>\n\n<p>Regarding hardware..if you have never used GCP...you can get free 300$ worth of compute for one year. \nFire up a pre-emptible P100 GPU and save intermediate checkpoints of the model so even if the GPU is pre-emptied you have still model checkpoints and can recontinue from there. Save model and optimizer states. A pre-emptible GPU would be 3 times cheaper than normal GPU. </p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 645620,
      "author_name": "Ajay Chauhan",
      "author_url": "",
      "post_date": "2019-10-10T09:53:01.050000",
      "content": "<p>Rightly said <a href=\"/cateek\">@cateek</a> </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "646073": "that's easy - start with lightweight models, there are only few of them to try. don't go 512 images - there are already downsampled images for your needs and I can prove they are good enough. find out the best way to preprocess input data and find rich augmentations pipeline for the problem (or take any pipeline from papers mentioned). probably, to win you need more resources but I believe you have everything to achieve good results (lets say top-100)",
    "645554": "As per my understanding biggest challenge for this competition is to work around with huge data to train and test.\n \nConverting images to png or jpg cost much time and most of the users are using converted images from other keggle used and that IS THE BEST PART of Kaggle :)\n\nNow about training the model: Data is so huge that even very simple models are taking 3-4 hours for a single epoch. Google Colab does not help much and crashed. GPU is limited.. Bad timing from Kaggle team :)\nAll the suggestions on reducing training part are welcome and appreciated .\n\nEndembling seems far distant as of now because it seems difficult to train a single model.\n\nAll suggestions and discussions are welcome..",
    "645750": "I have found a solution to your 'gpu crash' you may be keen in trying. Actually, the colab K80's are preemptible meaning your work can get interrupted at any time. Even on GCP, if I select a V100 preemptible instance, I get unpredictable instance disconnects.\n\nSo my solution to combat the interrupts and perhaps the long training time per epoch is to create a checkpoint that restores my model's training. Particularly, you'd want to restore the checkpoint that gives you the best validation loss. However, I think for your purposes having any checkpoint will be helpful since it takes you 3-4hrs per epoch. So you'll need to connect your gdrive to colab in order to obtain a copy of the checksum files created on colab. Luckily this is quite simple.\n\n&gt; from google.colab import drive\ndrive.mount('/content/gdrive/')\n\nthen click on the url and then authorize the connection\n\n&gt; ### save your model\n`model_save_name = \"your_model_name.pt\"`\n`path = f\"content/gdrive/My\\ Drive/{model_save_name}\"`\n`torch.save(model.state_dict(), path)`\n\nor that path doesn't work for you, try\n\n&gt; `path = \"gdrive/My Drive/your_model_name.pt\"`\n\nfor keras, you can change the file extension to .hdf5\n\n<br>\n\n### loading the checkpoint\njust mount your gdrive, locate the filepath and then \n\n&gt; ###PyTorch\n`model.load.state_dict(torch.load(\"file_path/file_name.pt\"))`\n\nor\n\n&gt; ###Keras\n`model.load_weights(\"file_path/file_name.hdf5\")`\n",
    "645615": "Don't know how to implement this, but maybe kaggle team could consider redistributing quota from those who don't actually use it. I have local gpu, so I don't need my quota. Maybe they could just create a special thread, where everyone who don't need his quota in a particular competition could declare such intentions. As far as I could roughly estimate, at least 30% of participants are training models locally. This could be a substantial help for others ",
    "645800": "Even with a GPU training is slow. I am going to start training with all the positive cases and under sample the normal scans.",
    "645741": "In this competition you could get credit for Google Cloud. It's late however...\nBut anyway you're right we have a lot of training data. I leave the idea to train by folds, since it's too long. :(\nAlthough it seems you don't need a lot of epochs here.\nThere was a similar topic how to accelerate:\n[https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/110671#latest-638850](https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/110671#latest-638850)",
    "645568": "I also have lost 3-4 days on Colab and it did not help. Model gets crashed using TPU and GPU even after reducing  to usbatch size..planning to use AWS",
    "645563": "i have been trying on colab since last 3 days and it keeps crashing over and over again,i am really very frustrated,model that works in kaggle doesn't work in colab and it keeps crashing,i reduced the batch size and it still keeps crashing,kaggle gpu quota is hurting so badly,i love deep learning a lot,gpu quota is killing me :'(",
    "646054": "I would like to know which model is better to be used as a single model...efficientNet or Resnet or Densenet or Exception etc.\nDue to big size of data trying everyone is an impossible challenge.\nAny help would be GREAT...\n",
    "645620": "Rightly said @cateek "
  }
}