{
  "id": 372550,
  "title": "Freezing Layers to Avoid OOM (1024px)",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/372550",
  "author_name": "moth",
  "post_date": "2022-12-16T15:32:18.499000",
  "votes": 19,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Thanks <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for pointing to this alternative to reduce memory and avoid OOM errors when training on 1024.</p>\n<h3>Why freezing layers?</h3>\n<p>It has been pointed in this competition that training on higher resolutions leads to better model performance. However, when we increase image resolution (say from 512px to 1024px) we may encounter OOM errors. There are many ways to solve OOM issues. See this <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/372192\" target=\"_blank\">topic</a>.</p>\n<p>Freezing layers help you train faster and reduce VRAM when training, thus avoiding OOM errors. When you freeze a layer, backpropagation do not affect frozen layers. Since these are not trained this <strong>might</strong> lead to poorer performance. It's best to check via experimentation. Even if it leads to poorer performance it may not be that significant.</p>\n<h3>How to freeze layers</h3>\n<p>First determine how many layers your model has. You can do so by:</p>\n<pre><code>for i,(name, param) in enumerate(list(model.named_parameters())):\n    print(i,name)\n</code></pre>\n<p>This will print the layers of your model. Then determine how many layers you want to freeze. You can experiment by freezing 100% model, 50%, 33%, 20%, etc.</p>\n<p>You can do so by:</p>\n<pre><code>NUM_FROZEN_LAYERS = 100 # how many layers you want to freeze\n\nfor i,(name, param) in enumerate(list(model.named_parameters())\\\n                                 [0:NUM_FROZEN_LAYERS]):\n    param.requires_grad = False\n</code></pre>\n<p><strong>Note:</strong> in the code above we freeze the <strong>first</strong> N layers. This is because the first layers capture low-level image features and we want to fine-tune on high-level image features.</p>\n<h3>My experiments:</h3>\n<p>Experiment 1:</p>\n<ul>\n<li><strong>Architecture:</strong> EfficientNetB2.</li>\n<li><strong>Image resolution:</strong> 1024px.</li>\n<li><strong>Train Batch Size:</strong> 8 images per batch.</li>\n<li><strong>Number of Frozen Layers:</strong> 100 out of 300 (~33% of the model).</li>\n<li><strong>Consumed GPU's VRAM:</strong> 4.7 out of 15.9 GB.</li>\n<li><strong>Training time for 1 EPOCH:</strong> 43k images on ~40 minutes.</li>\n<li><strong>Gradient Accumulation:</strong> True. Every 4 steps.</li>\n</ul>\n<p>Experiment 2:</p>\n<ul>\n<li><strong>Architecture:</strong> EfficientNetB2.</li>\n<li><strong>Image resolution:</strong> 1024px.</li>\n<li><strong>Train Batch Size:</strong> 8 images per batch.</li>\n<li><strong>Number of Frozen Layers:</strong> 166 out of 300 (~50% of the model).</li>\n<li><strong>Consumed GPU's VRAM:</strong> 3.7 out of 15.9 GB.</li>\n<li><strong>Training time for 1 EPOCH:</strong> 43k images on ~40 min.</li>\n<li><strong>Gradient Accumulation:</strong> True. Every 4 steps.</li>\n</ul>\n<p>Experiment 3:</p>\n<ul>\n<li><strong>Architecture:</strong> EfficientNetB2.</li>\n<li><strong>Image resolution:</strong> 1024px.</li>\n<li><strong>Train Batch Size:</strong> 16 images per batch.</li>\n<li><strong>Number of Frozen Layers:</strong> 100 out of 300 (~33% of the model).</li>\n<li><strong>Consumed GPU's VRAM:</strong> 8.9 out of 15.9 GB.</li>\n<li><strong>Training time for 1 EPOCH:</strong> 43k images on ~35 min.</li>\n<li><strong>Gradient Accumulation:</strong> True. Every 4 steps.</li>\n</ul>\n<p>Experiment 4:</p>\n<ul>\n<li><strong>Architecture:</strong> EfficientNetB2.</li>\n<li><strong>Image resolution:</strong> 1024px.</li>\n<li><strong>Train Batch Size:</strong> 16 images per batch.</li>\n<li><strong>Number of Frozen Layers:</strong> 166 out of 300 (~50% of the model).</li>\n<li><strong>Consumed GPU's VRAM:</strong> 7.1 out of 15.9 GB.</li>\n<li><strong>Training time for 1 EPOCH:</strong> 43k images on ~35 min.</li>\n<li><strong>Gradient Accumulation:</strong> True. Every 4 steps.</li>\n</ul>",
  "messages": [
    {
      "id": 2067313,
      "postDate": "2022-12-16T15:32:18.500Z",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for pointing to this alternative to reduce memory and avoid OOM errors when training on 1024.</p>\n<h3>Why freezing layers?</h3>\n<p>It has been pointed in this competition that training on higher resolutions leads to better model performance. However, when we increase image resolution (say from 512px to 1024px) we may encounter OOM errors. There are many ways to solve OOM issues. See this <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/372192\" target=\"_blank\">topic</a>.</p>\n<p>Freezing layers help you train faster and reduce VRAM when training, thus avoiding OOM errors. When you freeze a layer, backpropagation do not affect frozen layers. Since these are not trained this <strong>might</strong> lead to poorer performance. It's best to check via experimentation. Even if it leads to poorer performance it may not be that significant.</p>\n<h3>How to freeze layers</h3>\n<p>First determine how many layers your model has. You can do so by:</p>\n<pre><code>for i,(name, param) in enumerate(list(model.named_parameters())):\n    print(i,name)\n</code></pre>\n<p>This will print the layers of your model. Then determine how many layers you want to freeze. You can experiment by freezing 100% model, 50%, 33%, 20%, etc.</p>\n<p>You can do so by:</p>\n<pre><code>NUM_FROZEN_LAYERS = 100 # how many layers you want to freeze\n\nfor i,(name, param) in enumerate(list(model.named_parameters())\\\n                                 [0:NUM_FROZEN_LAYERS]):\n    param.requires_grad = False\n</code></pre>\n<p><strong>Note:</strong> in the code above we freeze the <strong>first</strong> N layers. This is because the first layers capture low-level image features and we want to fine-tune on high-level image features.</p>\n<h3>My experiments:</h3>\n<p>Experiment 1:</p>\n<ul>\n<li><strong>Architecture:</strong> EfficientNetB2.</li>\n<li><strong>Image resolution:</strong> 1024px.</li>\n<li><strong>Train Batch Size:</strong> 8 images per batch.</li>\n<li><strong>Number of Frozen Layers:</strong> 100 out of 300 (~33% of the model).</li>\n<li><strong>Consumed GPU's VRAM:</strong> 4.7 out of 15.9 GB.</li>\n<li><strong>Training time for 1 EPOCH:</strong> 43k images on ~40 minutes.</li>\n<li><strong>Gradient Accumulation:</strong> True. Every 4 steps.</li>\n</ul>\n<p>Experiment 2:</p>\n<ul>\n<li><strong>Architecture:</strong> EfficientNetB2.</li>\n<li><strong>Image resolution:</strong> 1024px.</li>\n<li><strong>Train Batch Size:</strong> 8 images per batch.</li>\n<li><strong>Number of Frozen Layers:</strong> 166 out of 300 (~50% of the model).</li>\n<li><strong>Consumed GPU's VRAM:</strong> 3.7 out of 15.9 GB.</li>\n<li><strong>Training time for 1 EPOCH:</strong> 43k images on ~40 min.</li>\n<li><strong>Gradient Accumulation:</strong> True. Every 4 steps.</li>\n</ul>\n<p>Experiment 3:</p>\n<ul>\n<li><strong>Architecture:</strong> EfficientNetB2.</li>\n<li><strong>Image resolution:</strong> 1024px.</li>\n<li><strong>Train Batch Size:</strong> 16 images per batch.</li>\n<li><strong>Number of Frozen Layers:</strong> 100 out of 300 (~33% of the model).</li>\n<li><strong>Consumed GPU's VRAM:</strong> 8.9 out of 15.9 GB.</li>\n<li><strong>Training time for 1 EPOCH:</strong> 43k images on ~35 min.</li>\n<li><strong>Gradient Accumulation:</strong> True. Every 4 steps.</li>\n</ul>\n<p>Experiment 4:</p>\n<ul>\n<li><strong>Architecture:</strong> EfficientNetB2.</li>\n<li><strong>Image resolution:</strong> 1024px.</li>\n<li><strong>Train Batch Size:</strong> 16 images per batch.</li>\n<li><strong>Number of Frozen Layers:</strong> 166 out of 300 (~50% of the model).</li>\n<li><strong>Consumed GPU's VRAM:</strong> 7.1 out of 15.9 GB.</li>\n<li><strong>Training time for 1 EPOCH:</strong> 43k images on ~35 min.</li>\n<li><strong>Gradient Accumulation:</strong> True. Every 4 steps.</li>\n</ul>",
      "rawMarkdown": "Thanks @cdeotte for pointing to this alternative to reduce memory and avoid OOM errors when training on 1024.\n\n### Why freezing layers?\n\nIt has been pointed in this competition that training on higher resolutions leads to better model performance. However, when we increase image resolution (say from 512px to 1024px) we may encounter OOM errors. There are many ways to solve OOM issues. See this [topic](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/372192).\n\nFreezing layers help you train faster and reduce VRAM when training, thus avoiding OOM errors. When you freeze a layer, backpropagation do not affect frozen layers. Since these are not trained this **might** lead to poorer performance. It's best to check via experimentation. Even if it leads to poorer performance it may not be that significant.\n\n### How to freeze layers\n\nFirst determine how many layers your model has. You can do so by:\n```\nfor i,(name, param) in enumerate(list(model.named_parameters())):\n    print(i,name)\n```\nThis will print the layers of your model. Then determine how many layers you want to freeze. You can experiment by freezing 100% model, 50%, 33%, 20%, etc.\n\nYou can do so by:\n```\nNUM_FROZEN_LAYERS = 100 # how many layers you want to freeze\n\nfor i,(name, param) in enumerate(list(model.named_parameters())\\\n                                 [0:NUM_FROZEN_LAYERS]):\n    param.requires_grad = False\n```\n\n**Note:** in the code above we freeze the **first** N layers. This is because the first layers capture low-level image features and we want to fine-tune on high-level image features.\n\n### My experiments:\n\nExperiment 1:\n- **Architecture:** EfficientNetB2.\n- **Image resolution:** 1024px.\n- **Train Batch Size:** 8 images per batch.\n- **Number of Frozen Layers:** 100 out of 300 (~33% of the model).\n- **Consumed GPU's VRAM:** 4.7 out of 15.9 GB.\n- **Training time for 1 EPOCH:** 43k images on ~40 minutes.\n- **Gradient Accumulation:** True. Every 4 steps.\n\nExperiment 2:\n- **Architecture:** EfficientNetB2.\n- **Image resolution:** 1024px.\n- **Train Batch Size:** 8 images per batch.\n- **Number of Frozen Layers:** 166 out of 300 (~50% of the model).\n- **Consumed GPU's VRAM:** 3.7 out of 15.9 GB.\n- **Training time for 1 EPOCH:** 43k images on ~40 min.\n- **Gradient Accumulation:** True. Every 4 steps.\n\nExperiment 3:\n- **Architecture:** EfficientNetB2.\n- **Image resolution:** 1024px.\n- **Train Batch Size:** 16 images per batch.\n- **Number of Frozen Layers:** 100 out of 300 (~33% of the model).\n- **Consumed GPU's VRAM:** 8.9 out of 15.9 GB.\n- **Training time for 1 EPOCH:** 43k images on ~35 min.\n- **Gradient Accumulation:** True. Every 4 steps.\n\nExperiment 4:\n- **Architecture:** EfficientNetB2.\n- **Image resolution:** 1024px.\n- **Train Batch Size:** 16 images per batch.\n- **Number of Frozen Layers:** 166 out of 300 (~50% of the model).\n- **Consumed GPU's VRAM:** 7.1 out of 15.9 GB.\n- **Training time for 1 EPOCH:** 43k images on ~35 min.\n- **Gradient Accumulation:** True. Every 4 steps.",
      "votes": 18
    },
    {
      "id": 2067469,
      "postDate": "2022-12-16T18:34:49.853Z",
      "content": "<p>Experiments visualized on a table:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3197853%2F4f111c9bd07ab7c543cdf0abc8ceaba3%2FScreen%20Shot%202022-12-16%20at%2015.30.44.png?generation=1671215554177960&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li><strong>Resolution:</strong> pixels.</li>\n<li><strong>Batch size:</strong> number of images.</li>\n<li><strong>Frozen %:</strong> percentage of frozen layers in the architecture.</li>\n<li><strong>Time:</strong> minutes of compute time for training one epoch.</li>\n</ul>",
      "rawMarkdown": "Experiments visualized on a table:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3197853%2F4f111c9bd07ab7c543cdf0abc8ceaba3%2FScreen%20Shot%202022-12-16%20at%2015.30.44.png?generation=1671215554177960&alt=media)\n\n- **Resolution:** pixels.\n- **Batch size:** number of images.\n- **Frozen %:** percentage of frozen layers in the architecture.\n- **Time:** minutes of compute time for training one epoch.",
      "votes": 4,
      "replies": [
        {
          "id": 2067598,
          "postDate": "2022-12-16T21:40:11.073Z",
          "content": "<p>Good! 0% frozen should be included in test. Thank you for test.</p>",
          "rawMarkdown": "Good! 0% frozen should be included in test. Thank you for test.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2067356,
      "postDate": "2022-12-16T16:19:21.453Z",
      "content": "<p>Thank you for tests - what about model perofmance (score result)?</p>",
      "rawMarkdown": "Thank you for tests - what about model perofmance (score result)?",
      "votes": 1,
      "replies": [
        {
          "id": 2067417,
          "postDate": "2022-12-16T17:20:14.900Z",
          "content": "<p>Unfortunately I only have 3 hrs left of this weeks compute time so I'll have to wait. I run one epoch with 1/3 of the layers frozen though and results were promising. I'll update next week.</p>",
          "rawMarkdown": "Unfortunately I only have 3 hrs left of this weeks compute time so I'll have to wait. I run one epoch with 1/3 of the layers frozen though and results were promising. I'll update next week.",
          "replies": [
            {
              "id": 2067460,
              "postDate": "2022-12-16T18:18:33.897Z",
              "content": "<p>I understand. I freeze feature extractor weights during warm up phase (only classifier is unfreezed). Then unfreeze all. I will try your idea. Thank you for sharing.</p>",
              "rawMarkdown": "I understand. I freeze feature extractor weights during warm up phase (only classifier is unfreezed). Then unfreeze all. I will try your idea. Thank you for sharing."
            },
            {
              "id": 2067638,
              "postDate": "2022-12-16T23:43:02.243Z",
              "content": "<p>Hmm, if you wanted to share your CV / loss that would be great too :)</p>",
              "rawMarkdown": "Hmm, if you wanted to share your CV / loss that would be great too :)"
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2067469,
      "author_name": "moth",
      "author_url": "",
      "post_date": "2022-12-16T18:34:49.853000",
      "content": "<p>Experiments visualized on a table:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3197853%2F4f111c9bd07ab7c543cdf0abc8ceaba3%2FScreen%20Shot%202022-12-16%20at%2015.30.44.png?generation=1671215554177960&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li><strong>Resolution:</strong> pixels.</li>\n<li><strong>Batch size:</strong> number of images.</li>\n<li><strong>Frozen %:</strong> percentage of frozen layers in the architecture.</li>\n<li><strong>Time:</strong> minutes of compute time for training one epoch.</li>\n</ul>",
      "votes": 4,
      "replies": [
        {
          "id": 2067598,
          "author_name": "Remek Kinas",
          "author_url": "",
          "post_date": "2022-12-16T21:40:11.073000",
          "content": "<p>Good! 0% frozen should be included in test. Thank you for test.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2067356,
      "author_name": "Remek Kinas",
      "author_url": "",
      "post_date": "2022-12-16T16:19:21.453000",
      "content": "<p>Thank you for tests - what about model perofmance (score result)?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2067417,
          "author_name": "moth",
          "author_url": "",
          "post_date": "2022-12-16T17:20:14.900000",
          "content": "<p>Unfortunately I only have 3 hrs left of this weeks compute time so I'll have to wait. I run one epoch with 1/3 of the layers frozen though and results were promising. I'll update next week.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2067460,
              "author_name": "Remek Kinas",
              "author_url": "",
              "post_date": "2022-12-16T18:18:33.897000",
              "content": "<p>I understand. I freeze feature extractor weights during warm up phase (only classifier is unfreezed). Then unfreeze all. I will try your idea. Thank you for sharing.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2067638,
              "author_name": "@kaggleqrdl",
              "author_url": "",
              "post_date": "2022-12-16T23:43:02.243000",
              "content": "<p>Hmm, if you wanted to share your CV / loss that would be great too :)</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2067313": "Thanks @cdeotte for pointing to this alternative to reduce memory and avoid OOM errors when training on 1024.\n\n### Why freezing layers?\n\nIt has been pointed in this competition that training on higher resolutions leads to better model performance. However, when we increase image resolution (say from 512px to 1024px) we may encounter OOM errors. There are many ways to solve OOM issues. See this [topic](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/372192).\n\nFreezing layers help you train faster and reduce VRAM when training, thus avoiding OOM errors. When you freeze a layer, backpropagation do not affect frozen layers. Since these are not trained this **might** lead to poorer performance. It's best to check via experimentation. Even if it leads to poorer performance it may not be that significant.\n\n### How to freeze layers\n\nFirst determine how many layers your model has. You can do so by:\n```\nfor i,(name, param) in enumerate(list(model.named_parameters())):\n    print(i,name)\n```\nThis will print the layers of your model. Then determine how many layers you want to freeze. You can experiment by freezing 100% model, 50%, 33%, 20%, etc.\n\nYou can do so by:\n```\nNUM_FROZEN_LAYERS = 100 # how many layers you want to freeze\n\nfor i,(name, param) in enumerate(list(model.named_parameters())\\\n                                 [0:NUM_FROZEN_LAYERS]):\n    param.requires_grad = False\n```\n\n**Note:** in the code above we freeze the **first** N layers. This is because the first layers capture low-level image features and we want to fine-tune on high-level image features.\n\n### My experiments:\n\nExperiment 1:\n- **Architecture:** EfficientNetB2.\n- **Image resolution:** 1024px.\n- **Train Batch Size:** 8 images per batch.\n- **Number of Frozen Layers:** 100 out of 300 (~33% of the model).\n- **Consumed GPU's VRAM:** 4.7 out of 15.9 GB.\n- **Training time for 1 EPOCH:** 43k images on ~40 minutes.\n- **Gradient Accumulation:** True. Every 4 steps.\n\nExperiment 2:\n- **Architecture:** EfficientNetB2.\n- **Image resolution:** 1024px.\n- **Train Batch Size:** 8 images per batch.\n- **Number of Frozen Layers:** 166 out of 300 (~50% of the model).\n- **Consumed GPU's VRAM:** 3.7 out of 15.9 GB.\n- **Training time for 1 EPOCH:** 43k images on ~40 min.\n- **Gradient Accumulation:** True. Every 4 steps.\n\nExperiment 3:\n- **Architecture:** EfficientNetB2.\n- **Image resolution:** 1024px.\n- **Train Batch Size:** 16 images per batch.\n- **Number of Frozen Layers:** 100 out of 300 (~33% of the model).\n- **Consumed GPU's VRAM:** 8.9 out of 15.9 GB.\n- **Training time for 1 EPOCH:** 43k images on ~35 min.\n- **Gradient Accumulation:** True. Every 4 steps.\n\nExperiment 4:\n- **Architecture:** EfficientNetB2.\n- **Image resolution:** 1024px.\n- **Train Batch Size:** 16 images per batch.\n- **Number of Frozen Layers:** 166 out of 300 (~50% of the model).\n- **Consumed GPU's VRAM:** 7.1 out of 15.9 GB.\n- **Training time for 1 EPOCH:** 43k images on ~35 min.\n- **Gradient Accumulation:** True. Every 4 steps.",
    "2067469": "Experiments visualized on a table:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3197853%2F4f111c9bd07ab7c543cdf0abc8ceaba3%2FScreen%20Shot%202022-12-16%20at%2015.30.44.png?generation=1671215554177960&alt=media)\n\n- **Resolution:** pixels.\n- **Batch size:** number of images.\n- **Frozen %:** percentage of frozen layers in the architecture.\n- **Time:** minutes of compute time for training one epoch.",
    "2067356": "Thank you for tests - what about model perofmance (score result)?"
  }
}