{
  "id": 369155,
  "title": "💡 6 Computer Vision tricks for faster training and better models 🚀🚀🚀",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/369155",
  "author_name": "Radek Osmulski",
  "post_date": "2022-11-29T07:34:07.885000",
  "votes": 85,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Hey!</p>\n<p>Here are a couple of methods for training Computer Vision models that can give you an edge in this competition 🙂 (you will be able to train better models, faster, both locally on your machine and on Kaggle VMs)</p>\n<h1>1. Train with half-precision (FP16)</h1>\n<p>Here is <a href=\"https://docs.fast.ai/callback.fp16.html\" target=\"_blank\">a great, detailed explanation</a> of the technique. But the bottom line is that we can train with our model represented as <code>float16</code> (instead of <code>float32</code>) and gain massive speed up in the process!</p>\n<p>We also often can train on much larger batch sizes, which further accelerates the training.</p>\n<p>Definitely worth trying in this competition.</p>\n<h1>2. Progressive resizing (start with smaller resolution of train images)</h1>\n<p>This method was invented (or at least popularized) by <a href=\"https://www.fast.ai/\" target=\"_blank\">fast.ai</a>. The idea is that training on smaller images is much, much faster than training on larger ones, and additionally -- if we follow the curriculum training perspective -- training on smaller images might be better aligned with what the model needs to learn first!</p>\n<p>It might be beneficial to the model if it learns the big picture first and only then learns to discern the fine details.</p>\n<p>How does the technique work? Begin training on smaller images (say, 64x64) and progressively scale up (-&gt; 128x128, 196x196, 256x256, etc).</p>\n<p>You can read more about progressive resizing <a href=\"https://docs.mosaicml.com/en/latest/method_cards/progressive_resizing.html\" target=\"_blank\">here</a> and <a href=\"https://airctic.com/0.12.0/progressive_resizing/\" target=\"_blank\">here</a>.</p>\n<h1>3. Upsample the lower represented class to improve performance</h1>\n<p>Handling of imbalanced datasets always poses a challenge. A lot has been written about this, people tried to come up with elaborate techniques to address this (SMOTE, for instance).</p>\n<p>But in reality, they rarely work well! On 99% of problems upsampling the lower represented classes tends to work very well.</p>\n<h1>4. Use an LR Scheduler</h1>\n<p>Modulating the <code>learning rate</code> during training can shorten the training time and give your model a performance boost!</p>\n<p>A lot of seminal work on this was done in <a href=\"https://arxiv.org/abs/1708.07120\" target=\"_blank\">Super-Convergence: Very Fast Training of Neural Networks Using Large Learning Rates</a> by Leslie Smith. In fact, I used this technique (in particular, training with cosine annealing) <a href=\"https://www.kaggle.com/competitions/imaterialist-challenge-fashion-2018/discussion/57944\" target=\"_blank\">to win a CV competition on Kaggle</a> some time ago!</p>\n<p>Most frameworks now provide you with the functionality to train with cosine annealing (or another annealing schedule) out of the box! Definitely worth experimenting with. </p>\n<p>BTW another good source of information on this technique is the fabulous <a href=\"https://arxiv.org/abs/1812.01187\" target=\"_blank\">Bag of Tricks for Image Classification with Convolutional Neural Networks</a> paper.</p>\n<h1>5. User LR warmup</h1>\n<p>This is yet another technique covered in the Bag of Tricks paper above! The thinking is this -- if you start with a regular LR, because initially, the weights are random, this can lead to numerical instability in training. This can lead to the model struggling to learn for a non-trivial number of batches thus slowing down training.</p>\n<p>In order to avoid suffering from this problem, train for a couple of batches with an LR an order (or two orders) of magnitude lower than the one you intend to train with! </p>\n<h1>6. Image augmentations</h1>\n<p>Augmenting your train data can go a very long way. In particular given the class imbalance!  You can get very creative on this! For instance, <a href=\"https://paperswithcode.com/method/mixup\" target=\"_blank\">mixup</a> (one of my favorite techniques) -- being very straightforward -- can go a very long way!</p>\n<h1>Summary</h1>\n<p>And that's it 🙂 These are the 7 techniques I would be eager to try out in this competition to significantly improve performance! (both in predictive power and shortening of run time!)</p>\n<p>Hope some of these techniques will speed you along the way! 🙂 </p>\n<h3>Other resources you might find useful:</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/how-to-process-dicom-images-to-pngs\" target=\"_blank\">💡 how to process DICOM images to PNGs</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/eda-training-a-fast-ai-model-submission\" target=\"_blank\">📊 EDA + training a fast.ai model + submission 🚀</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/fast-ai-starter-pack-train-inference\" target=\"_blank\">🤖 [fast.ai starter pack] train + inference 🚀</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/371268\" target=\"_blank\">📸 Over 56GB of processed data, 5 different methods 🥳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369706\" target=\"_blank\">3 resources to get started with Computer Vision in this competition 🚀🚀🚀</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369155\" target=\"_blank\">💡 6 Computer Vision tricks for faster training and better models 🚀🚀🚀</a></li>\n</ul>\n<p>Happy Kaggling! 🥳</p>",
  "messages": [
    {
      "id": 2048018,
      "postDate": "2022-11-29T07:34:07.887Z",
      "content": "<p>Hey!</p>\n<p>Here are a couple of methods for training Computer Vision models that can give you an edge in this competition 🙂 (you will be able to train better models, faster, both locally on your machine and on Kaggle VMs)</p>\n<h1>1. Train with half-precision (FP16)</h1>\n<p>Here is <a href=\"https://docs.fast.ai/callback.fp16.html\" target=\"_blank\">a great, detailed explanation</a> of the technique. But the bottom line is that we can train with our model represented as <code>float16</code> (instead of <code>float32</code>) and gain massive speed up in the process!</p>\n<p>We also often can train on much larger batch sizes, which further accelerates the training.</p>\n<p>Definitely worth trying in this competition.</p>\n<h1>2. Progressive resizing (start with smaller resolution of train images)</h1>\n<p>This method was invented (or at least popularized) by <a href=\"https://www.fast.ai/\" target=\"_blank\">fast.ai</a>. The idea is that training on smaller images is much, much faster than training on larger ones, and additionally -- if we follow the curriculum training perspective -- training on smaller images might be better aligned with what the model needs to learn first!</p>\n<p>It might be beneficial to the model if it learns the big picture first and only then learns to discern the fine details.</p>\n<p>How does the technique work? Begin training on smaller images (say, 64x64) and progressively scale up (-&gt; 128x128, 196x196, 256x256, etc).</p>\n<p>You can read more about progressive resizing <a href=\"https://docs.mosaicml.com/en/latest/method_cards/progressive_resizing.html\" target=\"_blank\">here</a> and <a href=\"https://airctic.com/0.12.0/progressive_resizing/\" target=\"_blank\">here</a>.</p>\n<h1>3. Upsample the lower represented class to improve performance</h1>\n<p>Handling of imbalanced datasets always poses a challenge. A lot has been written about this, people tried to come up with elaborate techniques to address this (SMOTE, for instance).</p>\n<p>But in reality, they rarely work well! On 99% of problems upsampling the lower represented classes tends to work very well.</p>\n<h1>4. Use an LR Scheduler</h1>\n<p>Modulating the <code>learning rate</code> during training can shorten the training time and give your model a performance boost!</p>\n<p>A lot of seminal work on this was done in <a href=\"https://arxiv.org/abs/1708.07120\" target=\"_blank\">Super-Convergence: Very Fast Training of Neural Networks Using Large Learning Rates</a> by Leslie Smith. In fact, I used this technique (in particular, training with cosine annealing) <a href=\"https://www.kaggle.com/competitions/imaterialist-challenge-fashion-2018/discussion/57944\" target=\"_blank\">to win a CV competition on Kaggle</a> some time ago!</p>\n<p>Most frameworks now provide you with the functionality to train with cosine annealing (or another annealing schedule) out of the box! Definitely worth experimenting with. </p>\n<p>BTW another good source of information on this technique is the fabulous <a href=\"https://arxiv.org/abs/1812.01187\" target=\"_blank\">Bag of Tricks for Image Classification with Convolutional Neural Networks</a> paper.</p>\n<h1>5. User LR warmup</h1>\n<p>This is yet another technique covered in the Bag of Tricks paper above! The thinking is this -- if you start with a regular LR, because initially, the weights are random, this can lead to numerical instability in training. This can lead to the model struggling to learn for a non-trivial number of batches thus slowing down training.</p>\n<p>In order to avoid suffering from this problem, train for a couple of batches with an LR an order (or two orders) of magnitude lower than the one you intend to train with! </p>\n<h1>6. Image augmentations</h1>\n<p>Augmenting your train data can go a very long way. In particular given the class imbalance!  You can get very creative on this! For instance, <a href=\"https://paperswithcode.com/method/mixup\" target=\"_blank\">mixup</a> (one of my favorite techniques) -- being very straightforward -- can go a very long way!</p>\n<h1>Summary</h1>\n<p>And that's it 🙂 These are the 7 techniques I would be eager to try out in this competition to significantly improve performance! (both in predictive power and shortening of run time!)</p>\n<p>Hope some of these techniques will speed you along the way! 🙂 </p>\n<h3>Other resources you might find useful:</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/how-to-process-dicom-images-to-pngs\" target=\"_blank\">💡 how to process DICOM images to PNGs</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/eda-training-a-fast-ai-model-submission\" target=\"_blank\">📊 EDA + training a fast.ai model + submission 🚀</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/fast-ai-starter-pack-train-inference\" target=\"_blank\">🤖 [fast.ai starter pack] train + inference 🚀</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/371268\" target=\"_blank\">📸 Over 56GB of processed data, 5 different methods 🥳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369706\" target=\"_blank\">3 resources to get started with Computer Vision in this competition 🚀🚀🚀</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369155\" target=\"_blank\">💡 6 Computer Vision tricks for faster training and better models 🚀🚀🚀</a></li>\n</ul>\n<p>Happy Kaggling! 🥳</p>",
      "rawMarkdown": "Hey!\n\nHere are a couple of methods for training Computer Vision models that can give you an edge in this competition 🙂 (you will be able to train better models, faster, both locally on your machine and on Kaggle VMs)\n\n# 1. Train with half-precision (FP16)\n\nHere is [a great, detailed explanation](https://docs.fast.ai/callback.fp16.html) of the technique. But the bottom line is that we can train with our model represented as `float16` (instead of `float32`) and gain massive speed up in the process!\n\nWe also often can train on much larger batch sizes, which further accelerates the training.\n\nDefinitely worth trying in this competition.\n\n# 2. Progressive resizing (start with smaller resolution of train images)\n\nThis method was invented (or at least popularized) by [fast.ai](https://www.fast.ai/). The idea is that training on smaller images is much, much faster than training on larger ones, and additionally -- if we follow the curriculum training perspective -- training on smaller images might be better aligned with what the model needs to learn first!\n\nIt might be beneficial to the model if it learns the big picture first and only then learns to discern the fine details.\n\nHow does the technique work? Begin training on smaller images (say, 64x64) and progressively scale up (-> 128x128, 196x196, 256x256, etc).\n\nYou can read more about progressive resizing [here](https://docs.mosaicml.com/en/latest/method_cards/progressive_resizing.html) and [here](https://airctic.com/0.12.0/progressive_resizing/).\n\n# 3. Upsample the lower represented class to improve performance\n\nHandling of imbalanced datasets always poses a challenge. A lot has been written about this, people tried to come up with elaborate techniques to address this (SMOTE, for instance).\n\nBut in reality, they rarely work well! On 99% of problems upsampling the lower represented classes tends to work very well.\n\n# 4. Use an LR Scheduler\n\nModulating the `learning rate` during training can shorten the training time and give your model a performance boost!\n\nA lot of seminal work on this was done in [Super-Convergence: Very Fast Training of Neural Networks Using Large Learning Rates](https://arxiv.org/abs/1708.07120) by Leslie Smith. In fact, I used this technique (in particular, training with cosine annealing) [to win a CV competition on Kaggle](https://www.kaggle.com/competitions/imaterialist-challenge-fashion-2018/discussion/57944) some time ago!\n\nMost frameworks now provide you with the functionality to train with cosine annealing (or another annealing schedule) out of the box! Definitely worth experimenting with. \n\nBTW another good source of information on this technique is the fabulous [Bag of Tricks for Image Classification with Convolutional Neural Networks](https://arxiv.org/abs/1812.01187) paper.\n\n# 5. User LR warmup\n\nThis is yet another technique covered in the Bag of Tricks paper above! The thinking is this -- if you start with a regular LR, because initially, the weights are random, this can lead to numerical instability in training. This can lead to the model struggling to learn for a non-trivial number of batches thus slowing down training.\n\nIn order to avoid suffering from this problem, train for a couple of batches with an LR an order (or two orders) of magnitude lower than the one you intend to train with! \n\n# 6. Image augmentations\n\nAugmenting your train data can go a very long way. In particular given the class imbalance!  You can get very creative on this! For instance, [mixup](https://paperswithcode.com/method/mixup) (one of my favorite techniques) -- being very straightforward -- can go a very long way!\n\n# Summary\n\nAnd that's it 🙂 These are the 7 techniques I would be eager to try out in this competition to significantly improve performance! (both in predictive power and shortening of run time!)\n\nHope some of these techniques will speed you along the way! 🙂 \n\n### Other resources you might find useful:\n\n* [💡 how to process DICOM images to PNGs](https://www.kaggle.com/code/radek1/how-to-process-dicom-images-to-pngs)\n* [📊 EDA + training a fast.ai model + submission 🚀](https://www.kaggle.com/code/radek1/eda-training-a-fast-ai-model-submission)\n* [🤖 [fast.ai starter pack] train + inference 🚀](https://www.kaggle.com/code/radek1/fast-ai-starter-pack-train-inference)\n* [📸 Over 56GB of processed data, 5 different methods 🥳](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/371268)\n* [3 resources to get started with Computer Vision in this competition 🚀🚀🚀](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369706)\n* [💡 6 Computer Vision tricks for faster training and better models 🚀🚀🚀](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369155)\n\n\nHappy Kaggling! 🥳",
      "votes": 85
    },
    {
      "id": 2067647,
      "postDate": "2022-12-17T00:19:56.967Z",
      "content": "<p>Thanks for this helpful post! 👍👍</p>\n<p>Regarding the cosine annealing scheduler, do you apply it per batch or just once per epoch? Have you noticed any differences?</p>",
      "rawMarkdown": "Thanks for this helpful post! 👍👍\n\nRegarding the cosine annealing scheduler, do you apply it per batch or just once per epoch? Have you noticed any differences?",
      "votes": 2,
      "replies": [
        {
          "id": 2067752,
          "postDate": "2022-12-17T05:35:43.790Z",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/jnickerson\" target=\"_blank\">@jnickerson</a>!</p>\n<p>Thanks for your kind words 🙏 I apply it on a per batch basis (this is his it is implemented in fastai and also how it works natively in pytorch).</p>\n<p>I haven't tried applying it on a per epoch basis, but as I am training here for very few epochs, I do believe it would led to a worse result.</p>",
          "rawMarkdown": "Hey @jnickerson!\n\nThanks for your kind words 🙏 I apply it on a per batch basis (this is his it is implemented in fastai and also how it works natively in pytorch).\n\nI haven't tried applying it on a per epoch basis, but as I am training here for very few epochs, I do believe it would led to a worse result.",
          "replies": [
            {
              "id": 2068187,
              "postDate": "2022-12-17T15:36:33.250Z",
              "content": "<p>That makes sense. I hadn't thought to check fastai's implementation, so that's good to know!</p>\n<p>Thanks so much for your comments, Radek! 😄</p>",
              "rawMarkdown": "That makes sense. I hadn't thought to check fastai's implementation, so that's good to know!\n\nThanks so much for your comments, Radek! 😄"
            }
          ]
        }
      ]
    },
    {
      "id": 2048644,
      "postDate": "2022-11-29T15:33:02.257Z",
      "content": "<p>I have had this question for a long time, when someone says \"train your model locally\" it means that in order to train the&nbsp;model your computer&nbsp;needs to have a GPU, doesn't it?&nbsp;<br>\nAnother question, as my computer doesn't have a GPU, for this competition, would it be enough the GPU (and TPU) of&nbsp; Kaggle for the training, or do you recommend doing it somewhere else, for example, I've heard that Paperspace also offers free access and rents GPUs.&nbsp;<br>\nPS: This will be my first competition&nbsp;in Kaggle. I'm not sure If I'm prepared as&nbsp; I'm currently in lesson 2  of the fastai course, but as I've read in your book \"Use what you already know. Claiming you need to learn more is just a crutch that prevents you from&nbsp;starting\". I'm passionate about the subject of this competition, so I think it would be a perfect option for my DL project.&nbsp;</p>\n<p>Thank you very much for the tips!&nbsp;</p>",
      "rawMarkdown": "I have had this question for a long time, when someone says \"train your model locally\" it means that in order to train the model your computer needs to have a GPU, doesn't it? \nAnother question, as my computer doesn't have a GPU, for this competition, would it be enough the GPU (and TPU) of  Kaggle for the training, or do you recommend doing it somewhere else, for example, I've heard that Paperspace also offers free access and rents GPUs. \nPS: This will be my first competition in Kaggle. I'm not sure If I'm prepared as  I'm currently in lesson 2  of the fastai course, but as I've read in your book \"Use what you already know. Claiming you need to learn more is just a crutch that prevents you from starting\". I'm passionate about the subject of this competition, so I think it would be a perfect option for my DL project. \n\nThank you very much for the tips! ",
      "votes": 2,
      "replies": [
        {
          "id": 2048654,
          "postDate": "2022-11-29T15:37:38.573Z",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/karensnchez\" target=\"_blank\">@karensnchez</a>!</p>\n<p>By training locally I mean training on a machine that has a GPU. In many cases this will be a DL rig with just a single GPU, but some people (including many Kagglers 🙂) build more elaborate setups.</p>\n<p>I think you are in a perfect spot 🙂 Over the next couple of days people will most likely share a lot of notebooks on how to get started in this competition.</p>\n<p>What is a bit tricky here is the size of the data and the fact that inference needs to happen in Kaggle notebooks (it is a code competition). But I am quite certain there will be enough information provided to figure all of this out 🙂</p>\n<p>Thanks for reading my book! Hoping you will have a lot of fun in the competition!</p>",
          "rawMarkdown": "Hey @karensnchez!\n\nBy training locally I mean training on a machine that has a GPU. In many cases this will be a DL rig with just a single GPU, but some people (including many Kagglers 🙂) build more elaborate setups.\n\nI think you are in a perfect spot 🙂 Over the next couple of days people will most likely share a lot of notebooks on how to get started in this competition.\n\nWhat is a bit tricky here is the size of the data and the fact that inference needs to happen in Kaggle notebooks (it is a code competition). But I am quite certain there will be enough information provided to figure all of this out 🙂\n\nThanks for reading my book! Hoping you will have a lot of fun in the competition!",
          "votes": 2
        }
      ]
    },
    {
      "id": 2160614,
      "postDate": "2023-02-26T19:58:40.873Z",
      "content": "<p>This is so useful, thank you!</p>\n<p>Trying to run with FP16 I run into NaN outputs immediately in training. I think it's zero-division error in Adam due to the reduced float precision. Increasing the eps fixes the issue for anyone else coming across this.<br>\n<code>optimizer = torch.optim.Adam(model.parameters(), lr=LR, weight_decay=WD, eps=1e-3)</code></p>\n<p>Discussed here: <a href=\"https://discuss.pytorch.org/t/adam-half-precision-nans/1765\" target=\"_blank\">https://discuss.pytorch.org/t/adam-half-precision-nans/1765</a></p>",
      "rawMarkdown": "This is so useful, thank you!\n\nTrying to run with FP16 I run into NaN outputs immediately in training. I think it's zero-division error in Adam due to the reduced float precision. Increasing the eps fixes the issue for anyone else coming across this.\n`optimizer = torch.optim.Adam(model.parameters(), lr=LR, weight_decay=WD, eps=1e-3)`\n\nDiscussed here: https://discuss.pytorch.org/t/adam-half-precision-nans/1765"
    },
    {
      "id": 2145309,
      "postDate": "2023-02-14T23:28:16.913Z",
      "content": "<p>Thank you! Now I want to try progressive resizing. It's very interesting.</p>",
      "rawMarkdown": "Thank you! Now I want to try progressive resizing. It's very interesting."
    },
    {
      "id": 2117136,
      "postDate": "2023-01-27T03:54:21.577Z",
      "content": "<p>Thanks for the awesome resources! </p>\n<p>Would you recommend TPU training to speed up experiments, or would the extra overhead in writing code for TPU slow down the experimentation process?</p>",
      "rawMarkdown": "Thanks for the awesome resources! \n\nWould you recommend TPU training to speed up experiments, or would the extra overhead in writing code for TPU slow down the experimentation process?"
    },
    {
      "id": 2056320,
      "postDate": "2022-12-06T02:06:25.447Z",
      "content": "<p>Does half-precision affect model accuracy to provide the speedup? Thanks for the awesome tips!</p>",
      "rawMarkdown": "Does half-precision affect model accuracy to provide the speedup? Thanks for the awesome tips!",
      "replies": [
        {
          "id": 2066837,
          "postDate": "2022-12-16T05:24:25.950Z",
          "content": "<p>It affects slightly but is a great way to test ideas. If your new idea trained with fp16 is better than the baseline you can consider training the new idea on fp32.</p>",
          "rawMarkdown": "It affects slightly but is a great way to test ideas. If your new idea trained with fp16 is better than the baseline you can consider training the new idea on fp32.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2055188,
      "postDate": "2022-12-04T20:23:07.703Z",
      "content": "<p>Thank  you for sharing!</p>",
      "rawMarkdown": "Thank  you for sharing!"
    }
  ],
  "comments": [
    {
      "id": 2067647,
      "author_name": "Julia Nickerson",
      "author_url": "",
      "post_date": "2022-12-17T00:19:56.967000",
      "content": "<p>Thanks for this helpful post! 👍👍</p>\n<p>Regarding the cosine annealing scheduler, do you apply it per batch or just once per epoch? Have you noticed any differences?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2067752,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-12-17T05:35:43.790000",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/jnickerson\" target=\"_blank\">@jnickerson</a>!</p>\n<p>Thanks for your kind words 🙏 I apply it on a per batch basis (this is his it is implemented in fastai and also how it works natively in pytorch).</p>\n<p>I haven't tried applying it on a per epoch basis, but as I am training here for very few epochs, I do believe it would led to a worse result.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2068187,
              "author_name": "Julia Nickerson",
              "author_url": "",
              "post_date": "2022-12-17T15:36:33.250000",
              "content": "<p>That makes sense. I hadn't thought to check fastai's implementation, so that's good to know!</p>\n<p>Thanks so much for your comments, Radek! 😄</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2048644,
      "author_name": "Karen Sánchez",
      "author_url": "",
      "post_date": "2022-11-29T15:33:02.257000",
      "content": "<p>I have had this question for a long time, when someone says \"train your model locally\" it means that in order to train the&nbsp;model your computer&nbsp;needs to have a GPU, doesn't it?&nbsp;<br>\nAnother question, as my computer doesn't have a GPU, for this competition, would it be enough the GPU (and TPU) of&nbsp; Kaggle for the training, or do you recommend doing it somewhere else, for example, I've heard that Paperspace also offers free access and rents GPUs.&nbsp;<br>\nPS: This will be my first competition&nbsp;in Kaggle. I'm not sure If I'm prepared as&nbsp; I'm currently in lesson 2  of the fastai course, but as I've read in your book \"Use what you already know. Claiming you need to learn more is just a crutch that prevents you from&nbsp;starting\". I'm passionate about the subject of this competition, so I think it would be a perfect option for my DL project.&nbsp;</p>\n<p>Thank you very much for the tips!&nbsp;</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2048654,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-11-29T15:37:38.573000",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/karensnchez\" target=\"_blank\">@karensnchez</a>!</p>\n<p>By training locally I mean training on a machine that has a GPU. In many cases this will be a DL rig with just a single GPU, but some people (including many Kagglers 🙂) build more elaborate setups.</p>\n<p>I think you are in a perfect spot 🙂 Over the next couple of days people will most likely share a lot of notebooks on how to get started in this competition.</p>\n<p>What is a bit tricky here is the size of the data and the fact that inference needs to happen in Kaggle notebooks (it is a code competition). But I am quite certain there will be enough information provided to figure all of this out 🙂</p>\n<p>Thanks for reading my book! Hoping you will have a lot of fun in the competition!</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2160614,
      "author_name": "Andy Everall",
      "author_url": "",
      "post_date": "2023-02-26T19:58:40.873000",
      "content": "<p>This is so useful, thank you!</p>\n<p>Trying to run with FP16 I run into NaN outputs immediately in training. I think it's zero-division error in Adam due to the reduced float precision. Increasing the eps fixes the issue for anyone else coming across this.<br>\n<code>optimizer = torch.optim.Adam(model.parameters(), lr=LR, weight_decay=WD, eps=1e-3)</code></p>\n<p>Discussed here: <a href=\"https://discuss.pytorch.org/t/adam-half-precision-nans/1765\" target=\"_blank\">https://discuss.pytorch.org/t/adam-half-precision-nans/1765</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2145309,
      "author_name": "junseonglee11",
      "author_url": "",
      "post_date": "2023-02-14T23:28:16.913000",
      "content": "<p>Thank you! Now I want to try progressive resizing. It's very interesting.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2117136,
      "author_name": "William Zhou",
      "author_url": "",
      "post_date": "2023-01-27T03:54:21.577000",
      "content": "<p>Thanks for the awesome resources! </p>\n<p>Would you recommend TPU training to speed up experiments, or would the extra overhead in writing code for TPU slow down the experimentation process?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2056320,
      "author_name": "Abhishek Shah",
      "author_url": "",
      "post_date": "2022-12-06T02:06:25.447000",
      "content": "<p>Does half-precision affect model accuracy to provide the speedup? Thanks for the awesome tips!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2066837,
          "author_name": "Ayush Thakur",
          "author_url": "",
          "post_date": "2022-12-16T05:24:25.950000",
          "content": "<p>It affects slightly but is a great way to test ideas. If your new idea trained with fp16 is better than the baseline you can consider training the new idea on fp32.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2055188,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-12-04T20:23:07.703000",
      "content": "<p>Thank  you for sharing!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2048018": "Hey!\n\nHere are a couple of methods for training Computer Vision models that can give you an edge in this competition 🙂 (you will be able to train better models, faster, both locally on your machine and on Kaggle VMs)\n\n# 1. Train with half-precision (FP16)\n\nHere is [a great, detailed explanation](https://docs.fast.ai/callback.fp16.html) of the technique. But the bottom line is that we can train with our model represented as `float16` (instead of `float32`) and gain massive speed up in the process!\n\nWe also often can train on much larger batch sizes, which further accelerates the training.\n\nDefinitely worth trying in this competition.\n\n# 2. Progressive resizing (start with smaller resolution of train images)\n\nThis method was invented (or at least popularized) by [fast.ai](https://www.fast.ai/). The idea is that training on smaller images is much, much faster than training on larger ones, and additionally -- if we follow the curriculum training perspective -- training on smaller images might be better aligned with what the model needs to learn first!\n\nIt might be beneficial to the model if it learns the big picture first and only then learns to discern the fine details.\n\nHow does the technique work? Begin training on smaller images (say, 64x64) and progressively scale up (-> 128x128, 196x196, 256x256, etc).\n\nYou can read more about progressive resizing [here](https://docs.mosaicml.com/en/latest/method_cards/progressive_resizing.html) and [here](https://airctic.com/0.12.0/progressive_resizing/).\n\n# 3. Upsample the lower represented class to improve performance\n\nHandling of imbalanced datasets always poses a challenge. A lot has been written about this, people tried to come up with elaborate techniques to address this (SMOTE, for instance).\n\nBut in reality, they rarely work well! On 99% of problems upsampling the lower represented classes tends to work very well.\n\n# 4. Use an LR Scheduler\n\nModulating the `learning rate` during training can shorten the training time and give your model a performance boost!\n\nA lot of seminal work on this was done in [Super-Convergence: Very Fast Training of Neural Networks Using Large Learning Rates](https://arxiv.org/abs/1708.07120) by Leslie Smith. In fact, I used this technique (in particular, training with cosine annealing) [to win a CV competition on Kaggle](https://www.kaggle.com/competitions/imaterialist-challenge-fashion-2018/discussion/57944) some time ago!\n\nMost frameworks now provide you with the functionality to train with cosine annealing (or another annealing schedule) out of the box! Definitely worth experimenting with. \n\nBTW another good source of information on this technique is the fabulous [Bag of Tricks for Image Classification with Convolutional Neural Networks](https://arxiv.org/abs/1812.01187) paper.\n\n# 5. User LR warmup\n\nThis is yet another technique covered in the Bag of Tricks paper above! The thinking is this -- if you start with a regular LR, because initially, the weights are random, this can lead to numerical instability in training. This can lead to the model struggling to learn for a non-trivial number of batches thus slowing down training.\n\nIn order to avoid suffering from this problem, train for a couple of batches with an LR an order (or two orders) of magnitude lower than the one you intend to train with! \n\n# 6. Image augmentations\n\nAugmenting your train data can go a very long way. In particular given the class imbalance!  You can get very creative on this! For instance, [mixup](https://paperswithcode.com/method/mixup) (one of my favorite techniques) -- being very straightforward -- can go a very long way!\n\n# Summary\n\nAnd that's it 🙂 These are the 7 techniques I would be eager to try out in this competition to significantly improve performance! (both in predictive power and shortening of run time!)\n\nHope some of these techniques will speed you along the way! 🙂 \n\n### Other resources you might find useful:\n\n* [💡 how to process DICOM images to PNGs](https://www.kaggle.com/code/radek1/how-to-process-dicom-images-to-pngs)\n* [📊 EDA + training a fast.ai model + submission 🚀](https://www.kaggle.com/code/radek1/eda-training-a-fast-ai-model-submission)\n* [🤖 [fast.ai starter pack] train + inference 🚀](https://www.kaggle.com/code/radek1/fast-ai-starter-pack-train-inference)\n* [📸 Over 56GB of processed data, 5 different methods 🥳](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/371268)\n* [3 resources to get started with Computer Vision in this competition 🚀🚀🚀](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369706)\n* [💡 6 Computer Vision tricks for faster training and better models 🚀🚀🚀](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369155)\n\n\nHappy Kaggling! 🥳",
    "2067647": "Thanks for this helpful post! 👍👍\n\nRegarding the cosine annealing scheduler, do you apply it per batch or just once per epoch? Have you noticed any differences?",
    "2048644": "I have had this question for a long time, when someone says \"train your model locally\" it means that in order to train the model your computer needs to have a GPU, doesn't it? \nAnother question, as my computer doesn't have a GPU, for this competition, would it be enough the GPU (and TPU) of  Kaggle for the training, or do you recommend doing it somewhere else, for example, I've heard that Paperspace also offers free access and rents GPUs. \nPS: This will be my first competition in Kaggle. I'm not sure If I'm prepared as  I'm currently in lesson 2  of the fastai course, but as I've read in your book \"Use what you already know. Claiming you need to learn more is just a crutch that prevents you from starting\". I'm passionate about the subject of this competition, so I think it would be a perfect option for my DL project. \n\nThank you very much for the tips! ",
    "2160614": "This is so useful, thank you!\n\nTrying to run with FP16 I run into NaN outputs immediately in training. I think it's zero-division error in Adam due to the reduced float precision. Increasing the eps fixes the issue for anyone else coming across this.\n`optimizer = torch.optim.Adam(model.parameters(), lr=LR, weight_decay=WD, eps=1e-3)`\n\nDiscussed here: https://discuss.pytorch.org/t/adam-half-precision-nans/1765",
    "2145309": "Thank you! Now I want to try progressive resizing. It's very interesting.",
    "2117136": "Thanks for the awesome resources! \n\nWould you recommend TPU training to speed up experiments, or would the extra overhead in writing code for TPU slow down the experimentation process?",
    "2056320": "Does half-precision affect model accuracy to provide the speedup? Thanks for the awesome tips!",
    "2055188": "Thank  you for sharing!"
  }
}