{
  "id": 156258,
  "title": "CUDA out of memory",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/156258",
  "author_name": "Son of Anton v3.0",
  "post_date": "2020-06-05T06:43:16.715000",
  "votes": 1,
  "comment_count": 14,
  "views": 0,
  "content": "<p>I use a blend of PyTorch and FastAI for this contest. During training, I got NaN for validation loss. So I decided to venture in and understand what was going wrong. I wrote a simple snippet to print the validation loss.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2262353%2F643cc2b2155f5d122f824042eb8b657e%2Fqwf.PNG?generation=1591339207287062&amp;alt=media\" alt=\"\">\n I am getting a CUDA memory error after 1 iteration<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2262353%2F34b282d4f2a3c4c8dfad04e7eaa5abdc%2Feror.PNG?generation=1591339394146588&amp;alt=media\" alt=\"\"></p>\n\n<p>What is wrong?</p>",
  "messages": [
    {
      "id": 874940,
      "postDate": "2020-06-05T11:29:28.450Z",
      "content": "<p>try to run </p>\n\n<p><code>\ngpu_devices = tf.config.experimental.list_physical_devices('GPU')\ntf.config.experimental.set_memory_growth(gpu_devices[0], True)\n</code>\nand see if it works.\nyou are out of memory.</p>",
      "rawMarkdown": "try to run \n\n```\ngpu_devices = tf.config.experimental.list_physical_devices('GPU')\ntf.config.experimental.set_memory_growth(gpu_devices[0], True)\n```\nand see if it works.\nyou are out of memory.",
      "votes": 1,
      "replies": [
        {
          "id": 875060,
          "postDate": "2020-06-05T13:56:02.800Z",
          "content": "<p>What will it return? I use PyTorch btw</p>",
          "rawMarkdown": "What will it return? I use PyTorch btw\n"
        },
        {
          "id": 875109,
          "postDate": "2020-06-05T14:24:26.283Z",
          "content": "<p>That's for Tensorflow, and it lists the gpu devices, and manages memory growth.</p>\n\n<p>More OT: are you using a softmax activation as the last layer (or something else that squashed the output to [0,1])? \nYou can also try running some predictions on cpu, since they don't take <em>that</em> much time, especially if the error happens early on. In that case, you can just look at the output predictions </p>",
          "rawMarkdown": "That's for Tensorflow, and it lists the gpu devices, and manages memory growth.\n\nMore OT: are you using a softmax activation as the last layer (or something else that squashed the output to [0,1])? \nYou can also try running some predictions on cpu, since they don't take *that* much time, especially if the error happens early on. In that case, you can just look at the output predictions \n",
          "votes": 1
        },
        {
          "id": 875682,
          "postDate": "2020-06-06T04:53:00.440Z",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2262353%2F00322064e6146effdea84abd3ba2f3c1%2Foko.PNG?generation=1591419241906110&amp;alt=media\" alt=\"\">\nI am trying it with the CPU as you suggested but it has led to the kernel to crash and displays the above error message. Meanwhile, nn.CrossEntropyLoss() applies softmax to the output vectors. So I dint use it separately.</p>",
          "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2262353%2F00322064e6146effdea84abd3ba2f3c1%2Foko.PNG?generation=1591419241906110&amp;alt=media)\nI am trying it with the CPU as you suggested but it has led to the kernel to crash and displays the above error message. Meanwhile, nn.CrossEntropyLoss() applies softmax to the output vectors. So I dint use it separately.",
          "votes": 1
        },
        {
          "id": 876009,
          "postDate": "2020-06-06T11:06:50.017Z",
          "content": "<p>That already narrows to the DataLoader I guess. </p>\n\n<p>What kind of approach are you using (e.g. tiling or something else)?\nDid you decrease the batch size to see if that works?\nIs there any operation that might be persistent in memory throughout data processing, such that memory accumulates and it crashes?\nCan you try it with code that has proven to work by copying it from another kernel?</p>",
          "rawMarkdown": "That already narrows to the DataLoader I guess. \n\nWhat kind of approach are you using (e.g. tiling or something else)?\nDid you decrease the batch size to see if that works?\nIs there any operation that might be persistent in memory throughout data processing, such that memory accumulates and it crashes?\nCan you try it with code that has proven to work by copying it from another kernel?",
          "votes": 2
        },
        {
          "id": 876185,
          "postDate": "2020-06-06T14:07:55.227Z",
          "content": "<p>I use iafoss' tilling approach. But the model architecture is kinda different. </p>",
          "rawMarkdown": "I use iafoss' tilling approach. But the model architecture is kinda different. "
        },
        {
          "id": 876239,
          "postDate": "2020-06-06T14:58:51.867Z",
          "content": "<p>I reduced the batch_size and it did the trick. The validation losses seem to be fine when I print them but during training, it's still returning NaN. What might be the reason?</p>",
          "rawMarkdown": "I reduced the batch_size and it did the trick. The validation losses seem to be fine when I print them but during training, it's still returning NaN. What might be the reason?"
        },
        {
          "id": 876274,
          "postDate": "2020-06-06T15:38:39.580Z",
          "content": "<p>Good to hear!</p>\n\n<p>About the validation loss, is it evaluated after training the epoch?\nIf so, you might want to run through the entire validation set with the model and print out the outputs of:\n1. the network, <br>\n2. the labels, and <br>\n3. the loss <br>\nall per batch.\nIt only takes a single NaN after all to make the aggregate go NaN as well.</p>\n\n<p>This is kind of a brute force approach, but you'll be sure to find it that way. It can occur in all the small things: a small unintented data transformation, a mismatch between the loss and labels, e.g. given one hot encoded labels to a function that expects integers etc. Sorry I can't be more specific, but it's quite difficult for me to exactly pinpoint the problem without the code (and not being an expert in Pytorch also does not help).</p>",
          "rawMarkdown": "Good to hear!\n\nAbout the validation loss, is it evaluated after training the epoch?\nIf so, you might want to run through the entire validation set with the model and print out the outputs of:\n1. the network,   \n2. the labels, and    \n3. the loss   \nall per batch.\nIt only takes a single NaN after all to make the aggregate go NaN as well.\n\nThis is kind of a brute force approach, but you'll be sure to find it that way. It can occur in all the small things: a small unintented data transformation, a mismatch between the loss and labels, e.g. given one hot encoded labels to a function that expects integers etc. Sorry I can't be more specific, but it's quite difficult for me to exactly pinpoint the problem without the code (and not being an expert in Pytorch also does not help).",
          "votes": 1
        },
        {
          "id": 876869,
          "postDate": "2020-06-07T05:37:01.907Z",
          "content": "<p>Yes. It is evaluated after training an epoch.\nThis is the training code:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2262353%2F461f2f8e47bf8d5364da4addc35bc99f%2F2q4g.PNG?generation=1591508218230615&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "Yes. It is evaluated after training an epoch.\nThis is the training code:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2262353%2F461f2f8e47bf8d5364da4addc35bc99f%2F2q4g.PNG?generation=1591508218230615&amp;alt=media)\n",
          "votes": 1
        },
        {
          "id": 877766,
          "postDate": "2020-06-08T00:21:47.467Z",
          "content": "<p>I really think that the best way to solve this is to do the following procedure:\n1. Let the model train for a single epoch on the training data and save its weights\n2. For each sample/batch in the validation set:\n    -  Predict the batch using the model with weights after one round of training and print its prediction\n    -  Print the labels of the batch \n    -  Compute the loss for the validation batch and print it. I see you're using cross entropy so this should work (it's not a macro loss).</p>\n\n<p>This should allow you to see when, what, and where it is going wrong, since you're printing all batches and all possible inputs to the cross entropy loss. \nIt might also be a good idea to print their types while you're at it, just to be sure.</p>",
          "rawMarkdown": "I really think that the best way to solve this is to do the following procedure:\n1. Let the model train for a single epoch on the training data and save its weights\n2. For each sample/batch in the validation set:\n    -  Predict the batch using the model with weights after one round of training and print its prediction\n    -  Print the labels of the batch \n    -  Compute the loss for the validation batch and print it. I see you're using cross entropy so this should work (it's not a macro loss).\n\nThis should allow you to see when, what, and where it is going wrong, since you're printing all batches and all possible inputs to the cross entropy loss. \nIt might also be a good idea to print their types while you're at it, just to be sure.",
          "votes": 1
        },
        {
          "id": 878115,
          "postDate": "2020-06-08T08:59:45.677Z",
          "content": "<p>I did it as you said. My one epoch model returns some NaN values for some images whereas my untrained model returned floating-points. Does this mean the gradient has exploded?<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2262353%2F174b7efaa33aa706ffb8ba4f542c922b%2FCapture.PNG?generation=1591606778715320&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "I did it as you said. My one epoch model returns some NaN values for some images whereas my untrained model returned floating-points. Does this mean the gradient has exploded?![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2262353%2F174b7efaa33aa706ffb8ba4f542c922b%2FCapture.PNG?generation=1591606778715320&amp;alt=media)\n"
        },
        {
          "id": 879052,
          "postDate": "2020-06-09T06:33:57.373Z",
          "content": "<p>I printed the tensor outputs after each layer. I am getting Nan after some layers. When does such a situation occur? <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2262353%2Fcbbbd2d6a448a9f648d9477b6a8b25f3%2FCapture.PNG?generation=1591684301949292&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "I printed the tensor outputs after each layer. I am getting Nan after some layers. When does such a situation occur? ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2262353%2Fcbbbd2d6a448a9f648d9477b6a8b25f3%2FCapture.PNG?generation=1591684301949292&amp;alt=media)\n"
        },
        {
          "id": 879109,
          "postDate": "2020-06-09T07:39:13.603Z",
          "content": "<p>Hi, check this may help you.</p>\n\n<p><a href=\"https://www.kaggle.com/c/global-wheat-detection/discussion/156906\">https://www.kaggle.com/c/global-wheat-detection/discussion/156906</a></p>",
          "rawMarkdown": "Hi, check this may help you.\n\n[https://www.kaggle.com/c/global-wheat-detection/discussion/156906](https://www.kaggle.com/c/global-wheat-detection/discussion/156906)",
          "votes": 1
        },
        {
          "id": 879432,
          "postDate": "2020-06-09T13:22:02.817Z",
          "content": "<p>What I would check is if Nan appear right away =&gt; Likely to be a calculation error or if the model trains fine for a while an then nan appear =&gt; Gradient explosion.</p>",
          "rawMarkdown": "What I would check is if Nan appear right away =&gt; Likely to be a calculation error or if the model trains fine for a while an then nan appear =&gt; Gradient explosion."
        }
      ]
    },
    {
      "id": 874641,
      "postDate": "2020-06-05T06:43:16.717Z",
      "content": "<p>I use a blend of PyTorch and FastAI for this contest. During training, I got NaN for validation loss. So I decided to venture in and understand what was going wrong. I wrote a simple snippet to print the validation loss.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2262353%2F643cc2b2155f5d122f824042eb8b657e%2Fqwf.PNG?generation=1591339207287062&amp;alt=media\" alt=\"\">\n I am getting a CUDA memory error after 1 iteration<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2262353%2F34b282d4f2a3c4c8dfad04e7eaa5abdc%2Feror.PNG?generation=1591339394146588&amp;alt=media\" alt=\"\"></p>\n\n<p>What is wrong?</p>",
      "rawMarkdown": "I use a blend of PyTorch and FastAI for this contest. During training, I got NaN for validation loss. So I decided to venture in and understand what was going wrong. I wrote a simple snippet to print the validation loss.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2262353%2F643cc2b2155f5d122f824042eb8b657e%2Fqwf.PNG?generation=1591339207287062&amp;alt=media)\n I am getting a CUDA memory error after 1 iteration![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2262353%2F34b282d4f2a3c4c8dfad04e7eaa5abdc%2Feror.PNG?generation=1591339394146588&amp;alt=media)\n\nWhat is wrong?",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 874940,
      "author_name": "George",
      "author_url": "",
      "post_date": "2020-06-05T11:29:28.450000",
      "content": "<p>try to run </p>\n\n<p><code>\ngpu_devices = tf.config.experimental.list_physical_devices('GPU')\ntf.config.experimental.set_memory_growth(gpu_devices[0], True)\n</code>\nand see if it works.\nyou are out of memory.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 875060,
          "author_name": "Son of Anton v3.0",
          "author_url": "",
          "post_date": "2020-06-05T13:56:02.800000",
          "content": "<p>What will it return? I use PyTorch btw</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 875109,
          "author_name": "Stephan",
          "author_url": "",
          "post_date": "2020-06-05T14:24:26.283000",
          "content": "<p>That's for Tensorflow, and it lists the gpu devices, and manages memory growth.</p>\n\n<p>More OT: are you using a softmax activation as the last layer (or something else that squashed the output to [0,1])? \nYou can also try running some predictions on cpu, since they don't take <em>that</em> much time, especially if the error happens early on. In that case, you can just look at the output predictions </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 875682,
          "author_name": "Son of Anton v3.0",
          "author_url": "",
          "post_date": "2020-06-06T04:53:00.440000",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2262353%2F00322064e6146effdea84abd3ba2f3c1%2Foko.PNG?generation=1591419241906110&amp;alt=media\" alt=\"\">\nI am trying it with the CPU as you suggested but it has led to the kernel to crash and displays the above error message. Meanwhile, nn.CrossEntropyLoss() applies softmax to the output vectors. So I dint use it separately.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 876009,
          "author_name": "Stephan",
          "author_url": "",
          "post_date": "2020-06-06T11:06:50.017000",
          "content": "<p>That already narrows to the DataLoader I guess. </p>\n\n<p>What kind of approach are you using (e.g. tiling or something else)?\nDid you decrease the batch size to see if that works?\nIs there any operation that might be persistent in memory throughout data processing, such that memory accumulates and it crashes?\nCan you try it with code that has proven to work by copying it from another kernel?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 876185,
          "author_name": "Son of Anton v3.0",
          "author_url": "",
          "post_date": "2020-06-06T14:07:55.227000",
          "content": "<p>I use iafoss' tilling approach. But the model architecture is kinda different. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 876239,
          "author_name": "Son of Anton v3.0",
          "author_url": "",
          "post_date": "2020-06-06T14:58:51.867000",
          "content": "<p>I reduced the batch_size and it did the trick. The validation losses seem to be fine when I print them but during training, it's still returning NaN. What might be the reason?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 876274,
          "author_name": "Stephan",
          "author_url": "",
          "post_date": "2020-06-06T15:38:39.580000",
          "content": "<p>Good to hear!</p>\n\n<p>About the validation loss, is it evaluated after training the epoch?\nIf so, you might want to run through the entire validation set with the model and print out the outputs of:\n1. the network, <br>\n2. the labels, and <br>\n3. the loss <br>\nall per batch.\nIt only takes a single NaN after all to make the aggregate go NaN as well.</p>\n\n<p>This is kind of a brute force approach, but you'll be sure to find it that way. It can occur in all the small things: a small unintented data transformation, a mismatch between the loss and labels, e.g. given one hot encoded labels to a function that expects integers etc. Sorry I can't be more specific, but it's quite difficult for me to exactly pinpoint the problem without the code (and not being an expert in Pytorch also does not help).</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 876869,
          "author_name": "Son of Anton v3.0",
          "author_url": "",
          "post_date": "2020-06-07T05:37:01.907000",
          "content": "<p>Yes. It is evaluated after training an epoch.\nThis is the training code:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2262353%2F461f2f8e47bf8d5364da4addc35bc99f%2F2q4g.PNG?generation=1591508218230615&amp;alt=media\" alt=\"\"></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 877766,
          "author_name": "Stephan",
          "author_url": "",
          "post_date": "2020-06-08T00:21:47.467000",
          "content": "<p>I really think that the best way to solve this is to do the following procedure:\n1. Let the model train for a single epoch on the training data and save its weights\n2. For each sample/batch in the validation set:\n    -  Predict the batch using the model with weights after one round of training and print its prediction\n    -  Print the labels of the batch \n    -  Compute the loss for the validation batch and print it. I see you're using cross entropy so this should work (it's not a macro loss).</p>\n\n<p>This should allow you to see when, what, and where it is going wrong, since you're printing all batches and all possible inputs to the cross entropy loss. \nIt might also be a good idea to print their types while you're at it, just to be sure.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 878115,
          "author_name": "Son of Anton v3.0",
          "author_url": "",
          "post_date": "2020-06-08T08:59:45.677000",
          "content": "<p>I did it as you said. My one epoch model returns some NaN values for some images whereas my untrained model returned floating-points. Does this mean the gradient has exploded?<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2262353%2F174b7efaa33aa706ffb8ba4f542c922b%2FCapture.PNG?generation=1591606778715320&amp;alt=media\" alt=\"\"></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 879052,
          "author_name": "Son of Anton v3.0",
          "author_url": "",
          "post_date": "2020-06-09T06:33:57.373000",
          "content": "<p>I printed the tensor outputs after each layer. I am getting Nan after some layers. When does such a situation occur? <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2262353%2Fcbbbd2d6a448a9f648d9477b6a8b25f3%2FCapture.PNG?generation=1591684301949292&amp;alt=media\" alt=\"\"></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 879109,
          "author_name": "George",
          "author_url": "",
          "post_date": "2020-06-09T07:39:13.603000",
          "content": "<p>Hi, check this may help you.</p>\n\n<p><a href=\"https://www.kaggle.com/c/global-wheat-detection/discussion/156906\">https://www.kaggle.com/c/global-wheat-detection/discussion/156906</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 879432,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2020-06-09T13:22:02.817000",
          "content": "<p>What I would check is if Nan appear right away =&gt; Likely to be a calculation error or if the model trains fine for a while an then nan appear =&gt; Gradient explosion.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "874940": "try to run \n\n```\ngpu_devices = tf.config.experimental.list_physical_devices('GPU')\ntf.config.experimental.set_memory_growth(gpu_devices[0], True)\n```\nand see if it works.\nyou are out of memory.",
    "874641": "I use a blend of PyTorch and FastAI for this contest. During training, I got NaN for validation loss. So I decided to venture in and understand what was going wrong. I wrote a simple snippet to print the validation loss.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2262353%2F643cc2b2155f5d122f824042eb8b657e%2Fqwf.PNG?generation=1591339207287062&amp;alt=media)\n I am getting a CUDA memory error after 1 iteration![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2262353%2F34b282d4f2a3c4c8dfad04e7eaa5abdc%2Feror.PNG?generation=1591339394146588&amp;alt=media)\n\nWhat is wrong?"
  }
}