{
  "id": 111292,
  "title": "EfficientNet-PyTorch Speed up and Memory usage",
  "url": "/competitions/rsna-intracranial-hemorrhage-detection/discussion/111292",
  "author_name": "DrHB",
  "post_date": "2019-10-04T14:53:28.592000",
  "votes": 80,
  "comment_count": 19,
  "views": 0,
  "content": "<p>I have been struggling with Memory issues and slow speed when using <code>EfficientNet</code>. After reading and roaming on dark web I found this GitHub disccusion.</p>\n\n<p><a href=\"https://github.com/lukemelas/EfficientNet-PyTorch/issues/18\">https://github.com/lukemelas/EfficientNet-PyTorch/issues/18</a></p>\n\n<p>TLDR: if you want to lower your memory and fit more images in the batch do following:</p>\n\n<p>1) Go to site package (or clone git hub repo)\n2) Go to utils and modify following line (<a href=\"https://github.com/lukemelas/EfficientNet-PyTorch/blob/de40cbfec8244a6ddbb367fd491d700ecc2eef85/efficientnet_pytorch/utils.py#L39\">https://github.com/lukemelas/EfficientNet-PyTorch/blob/de40cbfec8244a6ddbb367fd491d700ecc2eef85/efficientnet_pytorch/utils.py#L39</a>)</p>\n\n<p>from:\n<code>\ndef relu_fn(x):\n    \"\"\" Swish activation function \"\"\"\n    return x * torch.sigmoid(x)\n</code></p>\n\n<p>to:\n```\nsigmoid = torch.nn.Sigmoid()\nclass Swish(torch.autograd.Function):\n    @staticmethod\n    def forward(ctx, i):\n        result = i * sigmoid(i)\n        ctx.save_for_backward(i)\n        return result</p>\n\n<pre><code>@staticmethod\ndef backward(ctx, grad_output):\n    i = ctx.saved_variables[0]\n    sigmoid_i = sigmoid(i)\n    return grad_output * (sigmoid_i * (1 + i * (1 - sigmoid_i)))\n</code></pre>\n\n<p>swish = Swish.apply</p>\n\n<p>class Swish_module(nn.Module):\n    def forward(self, x):\n        return swish(x)</p>\n\n<p>swish_layer = Swish_module()\ndef relu_fn(x):\n    \"\"\" Swish activation function \"\"\"\n    # return x * torch.sigmoid(x)\n    return swish_layer(x)\n```</p>\n\n<p>Enjoy your 20-30% Memory reduction </p>",
  "messages": [
    {
      "id": 641275,
      "postDate": "2019-10-04T14:53:28.593Z",
      "content": "<p>I have been struggling with Memory issues and slow speed when using <code>EfficientNet</code>. After reading and roaming on dark web I found this GitHub disccusion.</p>\n\n<p><a href=\"https://github.com/lukemelas/EfficientNet-PyTorch/issues/18\">https://github.com/lukemelas/EfficientNet-PyTorch/issues/18</a></p>\n\n<p>TLDR: if you want to lower your memory and fit more images in the batch do following:</p>\n\n<p>1) Go to site package (or clone git hub repo)\n2) Go to utils and modify following line (<a href=\"https://github.com/lukemelas/EfficientNet-PyTorch/blob/de40cbfec8244a6ddbb367fd491d700ecc2eef85/efficientnet_pytorch/utils.py#L39\">https://github.com/lukemelas/EfficientNet-PyTorch/blob/de40cbfec8244a6ddbb367fd491d700ecc2eef85/efficientnet_pytorch/utils.py#L39</a>)</p>\n\n<p>from:\n<code>\ndef relu_fn(x):\n    \"\"\" Swish activation function \"\"\"\n    return x * torch.sigmoid(x)\n</code></p>\n\n<p>to:\n```\nsigmoid = torch.nn.Sigmoid()\nclass Swish(torch.autograd.Function):\n    @staticmethod\n    def forward(ctx, i):\n        result = i * sigmoid(i)\n        ctx.save_for_backward(i)\n        return result</p>\n\n<pre><code>@staticmethod\ndef backward(ctx, grad_output):\n    i = ctx.saved_variables[0]\n    sigmoid_i = sigmoid(i)\n    return grad_output * (sigmoid_i * (1 + i * (1 - sigmoid_i)))\n</code></pre>\n\n<p>swish = Swish.apply</p>\n\n<p>class Swish_module(nn.Module):\n    def forward(self, x):\n        return swish(x)</p>\n\n<p>swish_layer = Swish_module()\ndef relu_fn(x):\n    \"\"\" Swish activation function \"\"\"\n    # return x * torch.sigmoid(x)\n    return swish_layer(x)\n```</p>\n\n<p>Enjoy your 20-30% Memory reduction </p>",
      "rawMarkdown": "I have been struggling with Memory issues and slow speed when using `EfficientNet`. After reading and roaming on dark web I found this GitHub disccusion.\n\nhttps://github.com/lukemelas/EfficientNet-PyTorch/issues/18\n\nTLDR: if you want to lower your memory and fit more images in the batch do following:\n\n1) Go to site package (or clone git hub repo)\n2) Go to utils and modify following line (https://github.com/lukemelas/EfficientNet-PyTorch/blob/de40cbfec8244a6ddbb367fd491d700ecc2eef85/efficientnet_pytorch/utils.py#L39)\n\nfrom:\n```\ndef relu_fn(x):\n    \"\"\" Swish activation function \"\"\"\n    return x * torch.sigmoid(x)\n```\n\nto:\n```\nsigmoid = torch.nn.Sigmoid()\nclass Swish(torch.autograd.Function):\n    @staticmethod\n    def forward(ctx, i):\n        result = i * sigmoid(i)\n        ctx.save_for_backward(i)\n        return result\n\n    @staticmethod\n    def backward(ctx, grad_output):\n        i = ctx.saved_variables[0]\n        sigmoid_i = sigmoid(i)\n        return grad_output * (sigmoid_i * (1 + i * (1 - sigmoid_i)))\n\nswish = Swish.apply\n\nclass Swish_module(nn.Module):\n    def forward(self, x):\n        return swish(x)\n\nswish_layer = Swish_module()\ndef relu_fn(x):\n    \"\"\" Swish activation function \"\"\"\n    # return x * torch.sigmoid(x)\n    return swish_layer(x)\n```\n\nEnjoy your 20-30% Memory reduction ",
      "votes": 79
    },
    {
      "id": 659090,
      "postDate": "2019-10-27T04:13:06.253Z",
      "content": "<p>Jeremy Howard has been on the case with another implementation of Swish: <a href=\"https://twitter.com/jeremyphoward/status/1188304165971218432?s=20\">https://twitter.com/jeremyphoward/status/1188304165971218432?s=20</a> </p>",
      "rawMarkdown": "Jeremy Howard has been on the case with another implementation of Swish: https://twitter.com/jeremyphoward/status/1188304165971218432?s=20 ",
      "votes": 1,
      "replies": [
        {
          "id": 660391,
          "postDate": "2019-10-29T04:12:45.147Z",
          "content": "<p>I think updated library has this correction my training tine reduced from 52 to 39 minutes and model.summary says memory efficient swish </p>",
          "rawMarkdown": "I think updated library has this correction my training tine reduced from 52 to 39 minutes and model.summary says memory efficient swish "
        }
      ]
    },
    {
      "id": 641815,
      "postDate": "2019-10-05T06:30:50.910Z",
      "content": "<p>Yes it did lower my efficientnet memory usage from 9100mb to 7600mb with exactly the same setting.\nThanks for sharing!</p>",
      "rawMarkdown": "Yes it did lower my efficientnet memory usage from 9100mb to 7600mb with exactly the same setting.\nThanks for sharing!",
      "votes": 1
    },
    {
      "id": 644320,
      "postDate": "2019-10-08T16:01:18.153Z",
      "content": "<p>I just made experiment with original and corrected EfficientNet's. \nThat's my results on V100 GPU:\n- original EffNet b2 takes 15.75Gb and 17.2 min to train one epoch;\n- corrected EffNet b2 takes 13.5Gb and 18.3 min to train one epoch.</p>",
      "rawMarkdown": "I just made experiment with original and corrected EfficientNet's. \nThat's my results on V100 GPU:\n- original EffNet b2 takes 15.75Gb and 17.2 min to train one epoch;\n- corrected EffNet b2 takes 13.5Gb and 18.3 min to train one epoch.",
      "votes": 2,
      "replies": [
        {
          "id": 644325,
          "postDate": "2019-10-08T16:03:28.077Z",
          "content": "<p>Interesting. Same batch size? Have you tried it more than once? perhaps there is variance?</p>",
          "rawMarkdown": "Interesting. Same batch size? Have you tried it more than once? perhaps there is variance?"
        },
        {
          "id": 644336,
          "postDate": "2019-10-08T16:30:36.730Z",
          "content": "<p>all the same parameters except proposed correction. \nfor sure it isnt variance, because i averaged across some folds.</p>",
          "rawMarkdown": "all the same parameters except proposed correction. \nfor sure it isnt variance, because i averaged across some folds."
        },
        {
          "id": 644353,
          "postDate": "2019-10-08T17:07:00.363Z",
          "content": "<p>I have similar results. However, if you adjust BS to fill all the memory GPU has, you will gain a small speed-up.</p>",
          "rawMarkdown": "I have similar results. However, if you adjust BS to fill all the memory GPU has, you will gain a small speed-up."
        }
      ]
    },
    {
      "id": 642275,
      "postDate": "2019-10-05T19:41:32.397Z",
      "content": "<p>Can anyone explain why the original function was using so much more memory? </p>\n\n<p>PS. Thanks DrHB, gonna try this :)</p>",
      "rawMarkdown": "Can anyone explain why the original function was using so much more memory? \n\nPS. Thanks DrHB, gonna try this :)",
      "votes": 2,
      "replies": [
        {
          "id": 642388,
          "postDate": "2019-10-06T01:39:56.843Z",
          "content": "<p>Thanks to the author!\nI believe that when using <code>x * torch.sigmoid(x)</code>, the autograd will remember both input to <code>torch.sigmoid</code> and the two inputs to <code>torch.mul</code> as separate tensors. There are two graph nodes: sigmoid and multiplication. This can be optimized if we define <code>Swish</code> as a separate graph node with its own gradient computation.\nSo, the main reason is that autograd is not very smart :)</p>",
          "rawMarkdown": "Thanks to the author!\nI believe that when using `x * torch.sigmoid(x)`, the autograd will remember both input to `torch.sigmoid` and the two inputs to `torch.mul` as separate tensors. There are two graph nodes: sigmoid and multiplication. This can be optimized if we define `Swish` as a separate graph node with its own gradient computation.\nSo, the main reason is that autograd is not very smart :)",
          "votes": 4
        }
      ]
    },
    {
      "id": 641293,
      "postDate": "2019-10-04T15:03:45.750Z",
      "content": "<p>thanks for your sharing.\nWill this modification potentially affect the final training result?\nIf not, the 20-30% memory reduction is dope!</p>",
      "rawMarkdown": "thanks for your sharing.\nWill this modification potentially affect the final training result?\nIf not, the 20-30% memory reduction is dope!",
      "votes": 2
    },
    {
      "id": 644209,
      "postDate": "2019-10-08T13:44:27.510Z",
      "content": "<p>I made a fork with a fix for convenience purposes: <a href=\"https://github.com/hokmund/EfficientNet-PyTorch\">https://github.com/hokmund/EfficientNet-PyTorch</a>\nYou are welcome to use it via <code>pip install git+https://github.com/hokmund/EfficientNet-PyTorch</code></p>\n\n<p>However, it is still slower and more memory-consuming than the original TF implementation.</p>",
      "rawMarkdown": "I made a fork with a fix for convenience purposes: https://github.com/hokmund/EfficientNet-PyTorch\nYou are welcome to use it via `pip install git+https://github.com/hokmund/EfficientNet-PyTorch`\n\nHowever, it is still slower and more memory-consuming than the original TF implementation."
    },
    {
      "id": 643452,
      "postDate": "2019-10-07T14:18:36.080Z",
      "content": "<p>This might be a silly question, but how would you do this on a Linux VM on GCP? Specific commands?</p>",
      "rawMarkdown": "This might be a silly question, but how would you do this on a Linux VM on GCP? Specific commands?",
      "replies": [
        {
          "id": 644255,
          "postDate": "2019-10-08T14:37:04.230Z",
          "content": "<h3>Locate utils.py file</h3>\n\n<p>First you need to find path of the <code>site packages</code> directory.  If you run <code>pip install efficientnet-pytorch</code> command, then you will find exact location of site packages directory. From there, you will find <code>efficientnet_pytorch</code> package directory and under that directory, the <code>utils.py</code> file is located.  On my GCP instance, the <code>utils.py</code> file is located here -</p>\n\n<p><code>/opt/anaconda3/lib/python3.7/site-packages/efficientnet_pytorch/utils.py</code></p>\n\n<p>Once you find <code>utils.py</code> file location, open the file and replace the code mentioned in the original post. </p>\n\n<h3>Replace <code>relu_fn</code> definition [using terminal]</h3>\n\n<p>If you are comfortable using command line editor, then open the <code>utils.py</code> file using the following command (otherwise skip to the next section) -</p>\n\n<p><code>\nvi /path/to/site-packages/efficientnet_pytorch/utils.py\n</code></p>\n\n<h3>Replace <code>relu_fn</code> definition [from Jupyter notebook]</h3>\n\n<p>If you prefer to use Jupyter environment, you can run the following command from a cell -</p>\n\n<p><code>\n%load /path/to/site-packages/efficientnet_pytorch/utils.py\n</code></p>\n\n<p>The <code>utils.py</code> file content will be loaded to the notebook cell. Edit the code and then run the cell. Finally to overwrite the code to the <code>utils.py</code> file, add the <code>%%writefile</code> cell magic command at the <code>top</code> of the code cell and run the cell again -</p>\n\n<p><code>\n%%writefile /path/to/site-packages/efficientnet_pytorch/utils.py\n// utils.py file content goes here\n</code></p>",
          "rawMarkdown": "### Locate utils.py file\nFirst you need to find path of the `site packages` directory.  If you run `pip install efficientnet-pytorch` command, then you will find exact location of site packages directory. From there, you will find `efficientnet_pytorch` package directory and under that directory, the `utils.py` file is located.  On my GCP instance, the `utils.py` file is located here -\n\n`/opt/anaconda3/lib/python3.7/site-packages/efficientnet_pytorch/utils.py`\n\nOnce you find `utils.py` file location, open the file and replace the code mentioned in the original post. \n\n### Replace `relu_fn` definition [using terminal]\nIf you are comfortable using command line editor, then open the `utils.py` file using the following command (otherwise skip to the next section) -\n\n```\nvi /path/to/site-packages/efficientnet_pytorch/utils.py\n```\n\n### Replace `relu_fn` definition [from Jupyter notebook]\nIf you prefer to use Jupyter environment, you can run the following command from a cell -\n\n```\n%load /path/to/site-packages/efficientnet_pytorch/utils.py\n```\n\nThe `utils.py` file content will be loaded to the notebook cell. Edit the code and then run the cell. Finally to overwrite the code to the `utils.py` file, add the `%%writefile` cell magic command at the `top` of the code cell and run the cell again -\n\n```\n%%writefile /path/to/site-packages/efficientnet_pytorch/utils.py\n// utils.py file content goes here\n```\n",
          "votes": 3
        }
      ]
    },
    {
      "id": 643307,
      "postDate": "2019-10-07T11:39:38.567Z",
      "content": "<p>Nice work.</p>\n\n<blockquote>\n  <p>Will this modification potentially affect the final training result?</p>\n</blockquote>\n\n<p>It shouldn't. That looks like a perfectly valid implementation of the gradient for Swish, probably the same thing PyTorch does when applying the chain rule,  so should be identical.</p>\n\n<p>It will likely perform a little slower as it needs more CUDA kernel launches but the increased batch sizes should mean it still performs better than the original.</p>",
      "rawMarkdown": "Nice work.\n\n&gt; Will this modification potentially affect the final training result?\n\nIt shouldn't. That looks like a perfectly valid implementation of the gradient for Swish, probably the same thing PyTorch does when applying the chain rule,  so should be identical.\n\nIt will likely perform a little slower as it needs more CUDA kernel launches but the increased batch sizes should mean it still performs better than the original."
    },
    {
      "id": 643280,
      "postDate": "2019-10-07T11:18:30.323Z",
      "content": "<p>@DrHb when we were using effnet in aptos,we dint face this issue any time...\nWhat is different now out here</p>",
      "rawMarkdown": "@DrHb when we were using effnet in aptos,we dint face this issue any time...\nWhat is different now out here"
    },
    {
      "id": 642966,
      "postDate": "2019-10-06T21:28:19.907Z",
      "content": "<p>Intresting idea!  Great Work...!! Thanks  <a href=\"/drhabib\">@drhabib</a> </p>",
      "rawMarkdown": "Intresting idea!  Great Work...!! Thanks  @drhabib "
    },
    {
      "id": 641437,
      "postDate": "2019-10-04T17:29:00.133Z",
      "content": "<p>Very Helpful Share,...\nThanks <a href=\"/drhabib\">@drhabib</a> </p>",
      "rawMarkdown": "Very Helpful Share,...\nThanks @drhabib ",
      "votes": 1
    },
    {
      "id": 641378,
      "postDate": "2019-10-04T16:23:24.483Z",
      "content": "<p>Thank you very much for sharing! :-D</p>",
      "rawMarkdown": "Thank you very much for sharing! :-D",
      "votes": 1
    },
    {
      "id": 644648,
      "postDate": "2019-10-09T06:14:32.710Z",
      "content": "<p>Thank you for sharing! Very useful !</p>",
      "rawMarkdown": "Thank you for sharing! Very useful !"
    }
  ],
  "comments": [
    {
      "id": 659090,
      "author_name": "datasaurus",
      "author_url": "",
      "post_date": "2019-10-27T04:13:06.253000",
      "content": "<p>Jeremy Howard has been on the case with another implementation of Swish: <a href=\"https://twitter.com/jeremyphoward/status/1188304165971218432?s=20\">https://twitter.com/jeremyphoward/status/1188304165971218432?s=20</a> </p>",
      "votes": 1,
      "replies": [
        {
          "id": 660391,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2019-10-29T04:12:45.147000",
          "content": "<p>I think updated library has this correction my training tine reduced from 52 to 39 minutes and model.summary says memory efficient swish </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 641815,
      "author_name": "Qishen Ha",
      "author_url": "",
      "post_date": "2019-10-05T06:30:50.910000",
      "content": "<p>Yes it did lower my efficientnet memory usage from 9100mb to 7600mb with exactly the same setting.\nThanks for sharing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 644320,
      "author_name": "Ivan V.",
      "author_url": "",
      "post_date": "2019-10-08T16:01:18.153000",
      "content": "<p>I just made experiment with original and corrected EfficientNet's. \nThat's my results on V100 GPU:\n- original EffNet b2 takes 15.75Gb and 17.2 min to train one epoch;\n- corrected EffNet b2 takes 13.5Gb and 18.3 min to train one epoch.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 644325,
          "author_name": "Tim Yee",
          "author_url": "",
          "post_date": "2019-10-08T16:03:28.077000",
          "content": "<p>Interesting. Same batch size? Have you tried it more than once? perhaps there is variance?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 644336,
          "author_name": "Ivan V.",
          "author_url": "",
          "post_date": "2019-10-08T16:30:36.730000",
          "content": "<p>all the same parameters except proposed correction. \nfor sure it isnt variance, because i averaged across some folds.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 644353,
          "author_name": "Dmytro Panchenko",
          "author_url": "",
          "post_date": "2019-10-08T17:07:00.363000",
          "content": "<p>I have similar results. However, if you adjust BS to fill all the memory GPU has, you will gain a small speed-up.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 642275,
      "author_name": "Tom Aindow",
      "author_url": "",
      "post_date": "2019-10-05T19:41:32.397000",
      "content": "<p>Can anyone explain why the original function was using so much more memory? </p>\n\n<p>PS. Thanks DrHB, gonna try this :)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 642388,
          "author_name": "Ruslan Baynazarov",
          "author_url": "",
          "post_date": "2019-10-06T01:39:56.843000",
          "content": "<p>Thanks to the author!\nI believe that when using <code>x * torch.sigmoid(x)</code>, the autograd will remember both input to <code>torch.sigmoid</code> and the two inputs to <code>torch.mul</code> as separate tensors. There are two graph nodes: sigmoid and multiplication. This can be optimized if we define <code>Swish</code> as a separate graph node with its own gradient computation.\nSo, the main reason is that autograd is not very smart :)</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 641293,
      "author_name": "Salaryman",
      "author_url": "",
      "post_date": "2019-10-04T15:03:45.750000",
      "content": "<p>thanks for your sharing.\nWill this modification potentially affect the final training result?\nIf not, the 20-30% memory reduction is dope!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 644209,
      "author_name": "Dmytro Panchenko",
      "author_url": "",
      "post_date": "2019-10-08T13:44:27.510000",
      "content": "<p>I made a fork with a fix for convenience purposes: <a href=\"https://github.com/hokmund/EfficientNet-PyTorch\">https://github.com/hokmund/EfficientNet-PyTorch</a>\nYou are welcome to use it via <code>pip install git+https://github.com/hokmund/EfficientNet-PyTorch</code></p>\n\n<p>However, it is still slower and more memory-consuming than the original TF implementation.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 643452,
      "author_name": "Tim Yee",
      "author_url": "",
      "post_date": "2019-10-07T14:18:36.080000",
      "content": "<p>This might be a silly question, but how would you do this on a Linux VM on GCP? Specific commands?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 644255,
          "author_name": "Atikur Rahman",
          "author_url": "",
          "post_date": "2019-10-08T14:37:04.230000",
          "content": "<h3>Locate utils.py file</h3>\n\n<p>First you need to find path of the <code>site packages</code> directory.  If you run <code>pip install efficientnet-pytorch</code> command, then you will find exact location of site packages directory. From there, you will find <code>efficientnet_pytorch</code> package directory and under that directory, the <code>utils.py</code> file is located.  On my GCP instance, the <code>utils.py</code> file is located here -</p>\n\n<p><code>/opt/anaconda3/lib/python3.7/site-packages/efficientnet_pytorch/utils.py</code></p>\n\n<p>Once you find <code>utils.py</code> file location, open the file and replace the code mentioned in the original post. </p>\n\n<h3>Replace <code>relu_fn</code> definition [using terminal]</h3>\n\n<p>If you are comfortable using command line editor, then open the <code>utils.py</code> file using the following command (otherwise skip to the next section) -</p>\n\n<p><code>\nvi /path/to/site-packages/efficientnet_pytorch/utils.py\n</code></p>\n\n<h3>Replace <code>relu_fn</code> definition [from Jupyter notebook]</h3>\n\n<p>If you prefer to use Jupyter environment, you can run the following command from a cell -</p>\n\n<p><code>\n%load /path/to/site-packages/efficientnet_pytorch/utils.py\n</code></p>\n\n<p>The <code>utils.py</code> file content will be loaded to the notebook cell. Edit the code and then run the cell. Finally to overwrite the code to the <code>utils.py</code> file, add the <code>%%writefile</code> cell magic command at the <code>top</code> of the code cell and run the cell again -</p>\n\n<p><code>\n%%writefile /path/to/site-packages/efficientnet_pytorch/utils.py\n// utils.py file content goes here\n</code></p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 643307,
      "author_name": "Thomas Brandon",
      "author_url": "",
      "post_date": "2019-10-07T11:39:38.567000",
      "content": "<p>Nice work.</p>\n\n<blockquote>\n  <p>Will this modification potentially affect the final training result?</p>\n</blockquote>\n\n<p>It shouldn't. That looks like a perfectly valid implementation of the gradient for Swish, probably the same thing PyTorch does when applying the chain rule,  so should be identical.</p>\n\n<p>It will likely perform a little slower as it needs more CUDA kernel launches but the increased batch sizes should mean it still performs better than the original.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 643280,
      "author_name": "Jaideep",
      "author_url": "",
      "post_date": "2019-10-07T11:18:30.323000",
      "content": "<p>@DrHb when we were using effnet in aptos,we dint face this issue any time...\nWhat is different now out here</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 642966,
      "author_name": "Evgeny Shtepin",
      "author_url": "",
      "post_date": "2019-10-06T21:28:19.907000",
      "content": "<p>Intresting idea!  Great Work...!! Thanks  <a href=\"/drhabib\">@drhabib</a> </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 641437,
      "author_name": "Ailurophile",
      "author_url": "",
      "post_date": "2019-10-04T17:29:00.133000",
      "content": "<p>Very Helpful Share,...\nThanks <a href=\"/drhabib\">@drhabib</a> </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 641378,
      "author_name": "Michael Pieler",
      "author_url": "",
      "post_date": "2019-10-04T16:23:24.483000",
      "content": "<p>Thank you very much for sharing! :-D</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 644648,
      "author_name": "seefun",
      "author_url": "",
      "post_date": "2019-10-09T06:14:32.710000",
      "content": "<p>Thank you for sharing! Very useful !</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "641275": "I have been struggling with Memory issues and slow speed when using `EfficientNet`. After reading and roaming on dark web I found this GitHub disccusion.\n\nhttps://github.com/lukemelas/EfficientNet-PyTorch/issues/18\n\nTLDR: if you want to lower your memory and fit more images in the batch do following:\n\n1) Go to site package (or clone git hub repo)\n2) Go to utils and modify following line (https://github.com/lukemelas/EfficientNet-PyTorch/blob/de40cbfec8244a6ddbb367fd491d700ecc2eef85/efficientnet_pytorch/utils.py#L39)\n\nfrom:\n```\ndef relu_fn(x):\n    \"\"\" Swish activation function \"\"\"\n    return x * torch.sigmoid(x)\n```\n\nto:\n```\nsigmoid = torch.nn.Sigmoid()\nclass Swish(torch.autograd.Function):\n    @staticmethod\n    def forward(ctx, i):\n        result = i * sigmoid(i)\n        ctx.save_for_backward(i)\n        return result\n\n    @staticmethod\n    def backward(ctx, grad_output):\n        i = ctx.saved_variables[0]\n        sigmoid_i = sigmoid(i)\n        return grad_output * (sigmoid_i * (1 + i * (1 - sigmoid_i)))\n\nswish = Swish.apply\n\nclass Swish_module(nn.Module):\n    def forward(self, x):\n        return swish(x)\n\nswish_layer = Swish_module()\ndef relu_fn(x):\n    \"\"\" Swish activation function \"\"\"\n    # return x * torch.sigmoid(x)\n    return swish_layer(x)\n```\n\nEnjoy your 20-30% Memory reduction ",
    "659090": "Jeremy Howard has been on the case with another implementation of Swish: https://twitter.com/jeremyphoward/status/1188304165971218432?s=20 ",
    "641815": "Yes it did lower my efficientnet memory usage from 9100mb to 7600mb with exactly the same setting.\nThanks for sharing!",
    "644320": "I just made experiment with original and corrected EfficientNet's. \nThat's my results on V100 GPU:\n- original EffNet b2 takes 15.75Gb and 17.2 min to train one epoch;\n- corrected EffNet b2 takes 13.5Gb and 18.3 min to train one epoch.",
    "642275": "Can anyone explain why the original function was using so much more memory? \n\nPS. Thanks DrHB, gonna try this :)",
    "641293": "thanks for your sharing.\nWill this modification potentially affect the final training result?\nIf not, the 20-30% memory reduction is dope!",
    "644209": "I made a fork with a fix for convenience purposes: https://github.com/hokmund/EfficientNet-PyTorch\nYou are welcome to use it via `pip install git+https://github.com/hokmund/EfficientNet-PyTorch`\n\nHowever, it is still slower and more memory-consuming than the original TF implementation.",
    "643452": "This might be a silly question, but how would you do this on a Linux VM on GCP? Specific commands?",
    "643307": "Nice work.\n\n&gt; Will this modification potentially affect the final training result?\n\nIt shouldn't. That looks like a perfectly valid implementation of the gradient for Swish, probably the same thing PyTorch does when applying the chain rule,  so should be identical.\n\nIt will likely perform a little slower as it needs more CUDA kernel launches but the increased batch sizes should mean it still performs better than the original.",
    "643280": "@DrHb when we were using effnet in aptos,we dint face this issue any time...\nWhat is different now out here",
    "642966": "Intresting idea!  Great Work...!! Thanks  @drhabib ",
    "641437": "Very Helpful Share,...\nThanks @drhabib ",
    "641378": "Thank you very much for sharing! :-D",
    "644648": "Thank you for sharing! Very useful !"
  }
}