{
  "id": 420980,
  "title": "Training at Paperspace Gradient (NaN loss values) [Solved]",
  "url": "/competitions/google-research-identify-contrails-reduce-global-warming/discussion/420980",
  "author_name": "Maximiliano Diaz Battan",
  "post_date": "2023-07-03T13:24:07.563000",
  "votes": 8,
  "comment_count": 24,
  "views": 0,
  "content": "<p>Hi Kagglers, is anyone here training at Paperspace? Something very strange is happening to me, I'm using the Kaggle docker image, therefore the same libraries, same Cuda drivers, anyway, all the same, including the notebook and configs. But when I train with PyTorch Lightning in fp16, the losses become NaN, but in fp32 everything works fine. While in Kaggle everything works correctly in fp16, and as I mentioned previously, they are the same versions of libraries, PL, Pytorch, SMP, cuda drivers, etc. Has the same thing happened to anyone? Or do you know why something like this might happen?</p>",
  "messages": [
    {
      "id": 2328288,
      "postDate": "2023-07-03T13:24:07.563Z",
      "content": "<p>Hi Kagglers, is anyone here training at Paperspace? Something very strange is happening to me, I'm using the Kaggle docker image, therefore the same libraries, same Cuda drivers, anyway, all the same, including the notebook and configs. But when I train with PyTorch Lightning in fp16, the losses become NaN, but in fp32 everything works fine. While in Kaggle everything works correctly in fp16, and as I mentioned previously, they are the same versions of libraries, PL, Pytorch, SMP, cuda drivers, etc. Has the same thing happened to anyone? Or do you know why something like this might happen?</p>",
      "rawMarkdown": "Hi Kagglers, is anyone here training at Paperspace? Something very strange is happening to me, I'm using the Kaggle docker image, therefore the same libraries, same Cuda drivers, anyway, all the same, including the notebook and configs. But when I train with PyTorch Lightning in fp16, the losses become NaN, but in fp32 everything works fine. While in Kaggle everything works correctly in fp16, and as I mentioned previously, they are the same versions of libraries, PL, Pytorch, SMP, cuda drivers, etc. Has the same thing happened to anyone? Or do you know why something like this might happen?",
      "votes": 7
    },
    {
      "id": 2370974,
      "postDate": "2023-08-02T18:40:15.897Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/maxdiazbattan\" target=\"_blank\">@maxdiazbattan</a>,</p>\n<p>I had the same problem, but I fixed it by setting the epsilon in the Adam optimizer to 1e-4 (eps=1e-4). The default epsilon of Adam is set to <a href=\"https://pytorch.org/docs/stable/generated/torch.optim.Adam.html\" target=\"_blank\">1e-8 </a>, which can cause some numerical instability after some epochs with fp16 I read. </p>\n<p>I hope this will also help you! Let me know!</p>",
      "rawMarkdown": "Hi @maxdiazbattan,\n\nI had the same problem, but I fixed it by setting the epsilon in the Adam optimizer to 1e-4 (eps=1e-4). The default epsilon of Adam is set to [1e-8 ](https://pytorch.org/docs/stable/generated/torch.optim.Adam.html), which can cause some numerical instability after some epochs with fp16 I read. \n\nI hope this will also help you! Let me know!\n",
      "votes": 6,
      "replies": [
        {
          "id": 2371896,
          "postDate": "2023-08-03T10:35:12.993Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/menno1111\" target=\"_blank\">@menno1111</a>, thanks, it worked! This could be the solution if you don't want to change the lr. Good luck with the competition!</p>",
          "rawMarkdown": "Hi @menno1111, thanks, it worked! This could be the solution if you don't want to change the lr. Good luck with the competition!",
          "votes": 2,
          "replies": [
            {
              "id": 2372589,
              "postDate": "2023-08-03T19:00:22.663Z",
              "content": "<p>That is very nice to hear! Thank you and good luck with the final days of the competition!</p>",
              "rawMarkdown": "That is very nice to hear! Thank you and good luck with the final days of the competition!",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2329277,
      "postDate": "2023-07-04T07:17:00.017Z",
      "content": "<p>Probably you should use bfloat16 instead of float16. Probably …</p>\n<p>with torch.amp.autocast(device_type=“cuda”, dtype=torch.bfloat16):</p>\n<p>Here you can find some useful tips: <a href=\"https://lightning.ai/pages/community/tutorial/pytorch-memory-vit-llm/\" target=\"_blank\">https://lightning.ai/pages/community/tutorial/pytorch-memory-vit-llm/</a></p>",
      "rawMarkdown": "Probably you should use bfloat16 instead of float16. Probably …\n\nwith torch.amp.autocast(device_type=“cuda”, dtype=torch.bfloat16):\n\nHere you can find some useful tips: https://lightning.ai/pages/community/tutorial/pytorch-memory-vit-llm/",
      "votes": 3,
      "replies": [
        {
          "id": 2329603,
          "postDate": "2023-07-04T11:15:55.437Z",
          "content": "<p>Thank you Sir, I've tried bf16, but unfortunately gives me some errors, apparently, some Pytorch functions do not support that precision. I've managed to solve it for now, by tweaking a tinny bit the lr. I guess I have to check the code again for bugs. </p>",
          "rawMarkdown": "Thank you Sir, I've tried bf16, but unfortunately gives me some errors, apparently, some Pytorch functions do not support that precision. I've managed to solve it for now, by tweaking a tinny bit the lr. I guess I have to check the code again for bugs. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 2331454,
      "postDate": "2023-07-05T14:56:23.010Z",
      "content": "<p>Did you try grad clipping along with lower LR? It helped in my case. PL has an easy callback for that. </p>",
      "rawMarkdown": "Did you try grad clipping along with lower LR? It helped in my case. PL has an easy callback for that. ",
      "votes": 1,
      "replies": [
        {
          "id": 2331739,
          "postDate": "2023-07-05T18:34:44.190Z",
          "content": "<p>Thank you very much Phaedrus. I haven't tried the 2 things together.</p>",
          "rawMarkdown": "Thank you very much Phaedrus. I haven't tried the 2 things together."
        }
      ]
    },
    {
      "id": 2330168,
      "postDate": "2023-07-04T19:12:25.807Z",
      "content": "<p>I think I spoke too soon, as I'm once again seeing NaN values in the loss with fp16. Let the debugging begin.</p>",
      "rawMarkdown": "I think I spoke too soon, as I'm once again seeing NaN values in the loss with fp16. Let the debugging begin.",
      "votes": 1,
      "replies": [
        {
          "id": 2330276,
          "postDate": "2023-07-04T21:17:08.200Z",
          "content": "<p>When I use any learning rate value greater than 5e-4, my loss function occasionally goes to NAN as well.  </p>\n<p>Any learning rate value below this has not gone to NAN so far.</p>",
          "rawMarkdown": "When I use any learning rate value greater than 5e-4, my loss function occasionally goes to NAN as well.  \n\nAny learning rate value below this has not gone to NAN so far.",
          "votes": 1,
          "replies": [
            {
              "id": 2331131,
              "postDate": "2023-07-05T11:37:48.953Z",
              "content": "<p>Thank you very much Bartley. By following your advice, I have been able to solve it for now. With 5e-4 missing values kept appearing, but on the epochs 8-10. I continued testing the lr, and with 1e-4 for the moment in 20 epochs they disappeared. It's quite strange because I was using much high on Kaggle, and I didn't encounter a single NaN value. Thank you!</p>",
              "rawMarkdown": "Thank you very much Bartley. By following your advice, I have been able to solve it for now. With 5e-4 missing values kept appearing, but on the epochs 8-10. I continued testing the lr, and with 1e-4 for the moment in 20 epochs they disappeared. It's quite strange because I was using much high on Kaggle, and I didn't encounter a single NaN value. Thank you!",
              "votes": 1
            },
            {
              "id": 2331480,
              "postDate": "2023-07-05T15:19:54.150Z",
              "content": "<p>Are you using SGD ? if so do you clip the gradients ?</p>",
              "rawMarkdown": "Are you using SGD ? if so do you clip the gradients ?",
              "votes": 1
            },
            {
              "id": 2331744,
              "postDate": "2023-07-05T18:39:30.550Z",
              "content": "<p>Thanks Janmpia, but no, I'm using Adam. I've tried gradient clipping, but since the problem, it's more focused on the validation loss, I guess that's why didn't work for me. Reducing the lr it's the only solution until now. Thanks anyway man, really appreciated it.</p>",
              "rawMarkdown": "Thanks Janmpia, but no, I'm using Adam. I've tried gradient clipping, but since the problem, it's more focused on the validation loss, I guess that's why didn't work for me. Reducing the lr it's the only solution until now. Thanks anyway man, really appreciated it."
            }
          ]
        }
      ]
    },
    {
      "id": 2329306,
      "postDate": "2023-07-04T07:37:36.057Z",
      "content": "<p>Hi! What are host system CUDA versions? I have seen strange behaviour running inference on kaggle 11.3 vs training on local 12.0, got it fixed when updated inference kaggle env to 11.7.</p>\n<p>One easiest thing to try is to rebuild the container without cache in paperclip env, at least that what I do when something is messed up with docker.</p>",
      "rawMarkdown": "Hi! What are host system CUDA versions? I have seen strange behaviour running inference on kaggle 11.3 vs training on local 12.0, got it fixed when updated inference kaggle env to 11.7.\n\nOne easiest thing to try is to rebuild the container without cache in paperclip env, at least that what I do when something is messed up with docker.",
      "votes": 1,
      "replies": [
        {
          "id": 2329597,
          "postDate": "2023-07-04T11:09:12.010Z",
          "content": "<p>Thanks Mikhail, very useful info. Both environments now have CUDA 11.8. After double-checking everything and confirming that everything is the same, I managed to \"solve\" the issue, at least partially for now. I'm not sure why, but by slightly tweaking the learning rate, the NaN values have disappeared. I trained for 20 epochs, and so far, there are no more NaN values. I'm not sure why the same learning rate and parameters work fine on Kaggle but not on Paperspace. I suppose I need to check the code for any bugs.</p>",
          "rawMarkdown": "Thanks Mikhail, very useful info. Both environments now have CUDA 11.8. After double-checking everything and confirming that everything is the same, I managed to \"solve\" the issue, at least partially for now. I'm not sure why, but by slightly tweaking the learning rate, the NaN values have disappeared. I trained for 20 epochs, and so far, there are no more NaN values. I'm not sure why the same learning rate and parameters work fine on Kaggle but not on Paperspace. I suppose I need to check the code for any bugs.",
          "votes": 1,
          "replies": [
            {
              "id": 2330487,
              "postDate": "2023-07-05T03:56:17.940Z",
              "content": "<p>Try to skip steps which have NaN gradients and see if it is occasional or training just breaks at some point. PL allows to return None from <code>LightningModule.training_step</code> to skip the step. I also have seen both NaNs and Infs in gradients at the beginning and with skipping / clipping model trains ok.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2Fa9d806c5a20c993119143d1658f5e9d4%2Ftrain_grads_nan.png?generation=1688529168741579&amp;alt=media\" alt=\"grads vs loss\"></p>",
              "rawMarkdown": "Try to skip steps which have NaN gradients and see if it is occasional or training just breaks at some point. PL allows to return None from `LightningModule.training_step` to skip the step. I also have seen both NaNs and Infs in gradients at the beginning and with skipping / clipping model trains ok.\n\n![grads vs loss](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2Fa9d806c5a20c993119143d1658f5e9d4%2Ftrain_grads_nan.png?generation=1688529168741579&alt=media)",
              "votes": 1
            },
            {
              "id": 2331136,
              "postDate": "2023-07-05T11:40:58.573Z",
              "content": "<p>I didn't know that. Thanks a lot, Mikhail. I'm going to test this as well.</p>",
              "rawMarkdown": "I didn't know that. Thanks a lot, Mikhail. I'm going to test this as well."
            }
          ]
        }
      ]
    },
    {
      "id": 2328298,
      "postDate": "2023-07-03T13:35:32.593Z",
      "content": "<p>are you training with MAnet or DeepLabv3+ by any chance ? if so the same happens to me on fp16</p>",
      "rawMarkdown": "are you training with MAnet or DeepLabv3+ by any chance ? if so the same happens to me on fp16",
      "votes": 2,
      "replies": [
        {
          "id": 2328451,
          "postDate": "2023-07-03T15:27:47.837Z",
          "content": "<p>I had the same issue with MAnet; the loss became NaN after 3-4 epochs.</p>",
          "rawMarkdown": "I had the same issue with MAnet; the loss became NaN after 3-4 epochs.",
          "votes": 1
        },
        {
          "id": 2328608,
          "postDate": "2023-07-03T17:51:22.630Z",
          "content": "<p>Thanks man, but no, it's just an Unet + Resnest101, works perfectly fine with fp32, but with fp16 don't. But at Kaggle did.</p>",
          "rawMarkdown": "Thanks man, but no, it's just an Unet + Resnest101, works perfectly fine with fp32, but with fp16 don't. But at Kaggle did.",
          "votes": 1,
          "replies": [
            {
              "id": 2330674,
              "postDate": "2023-07-05T06:19:21.163Z",
              "content": "<p>Can restnest101 be sustained in Caggle notebook?</p>",
              "rawMarkdown": "Can restnest101 be sustained in Caggle notebook?"
            },
            {
              "id": 2331257,
              "postDate": "2023-07-05T13:05:59.150Z",
              "content": "<p>Yes man, you can, just adjust the batch size.</p>",
              "rawMarkdown": "Yes man, you can, just adjust the batch size."
            },
            {
              "id": 2336574,
              "postDate": "2023-07-09T12:44:42.803Z",
              "content": "<p>Thanks very much, i will try it.</p>",
              "rawMarkdown": "Thanks very much, i will try it."
            },
            {
              "id": 2341260,
              "postDate": "2023-07-12T02:03:51.943Z",
              "content": "<p>Tried resnext101_32x4d, trained seperately first , got lower credit than 26d…………..</p>",
              "rawMarkdown": "Tried resnext101_32x4d, trained seperately first , got lower credit than 26d.............."
            },
            {
              "id": 2341916,
              "postDate": "2023-07-12T11:11:41.823Z",
              "content": "<p>try the resnest101e</p>",
              "rawMarkdown": "try the resnest101e"
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2370974,
      "author_name": "menno",
      "author_url": "",
      "post_date": "2023-08-02T18:40:15.897000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/maxdiazbattan\" target=\"_blank\">@maxdiazbattan</a>,</p>\n<p>I had the same problem, but I fixed it by setting the epsilon in the Adam optimizer to 1e-4 (eps=1e-4). The default epsilon of Adam is set to <a href=\"https://pytorch.org/docs/stable/generated/torch.optim.Adam.html\" target=\"_blank\">1e-8 </a>, which can cause some numerical instability after some epochs with fp16 I read. </p>\n<p>I hope this will also help you! Let me know!</p>",
      "votes": 6,
      "replies": [
        {
          "id": 2371896,
          "author_name": "Maximiliano Diaz Battan",
          "author_url": "",
          "post_date": "2023-08-03T10:35:12.993000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/menno1111\" target=\"_blank\">@menno1111</a>, thanks, it worked! This could be the solution if you don't want to change the lr. Good luck with the competition!</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2372589,
              "author_name": "menno",
              "author_url": "",
              "post_date": "2023-08-03T19:00:22.663000",
              "content": "<p>That is very nice to hear! Thank you and good luck with the final days of the competition!</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2329277,
      "author_name": "Remek Kinas",
      "author_url": "",
      "post_date": "2023-07-04T07:17:00.017000",
      "content": "<p>Probably you should use bfloat16 instead of float16. Probably …</p>\n<p>with torch.amp.autocast(device_type=“cuda”, dtype=torch.bfloat16):</p>\n<p>Here you can find some useful tips: <a href=\"https://lightning.ai/pages/community/tutorial/pytorch-memory-vit-llm/\" target=\"_blank\">https://lightning.ai/pages/community/tutorial/pytorch-memory-vit-llm/</a></p>",
      "votes": 3,
      "replies": [
        {
          "id": 2329603,
          "author_name": "Maximiliano Diaz Battan",
          "author_url": "",
          "post_date": "2023-07-04T11:15:55.437000",
          "content": "<p>Thank you Sir, I've tried bf16, but unfortunately gives me some errors, apparently, some Pytorch functions do not support that precision. I've managed to solve it for now, by tweaking a tinny bit the lr. I guess I have to check the code again for bugs. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2331454,
      "author_name": "Phaedrus",
      "author_url": "",
      "post_date": "2023-07-05T14:56:23.010000",
      "content": "<p>Did you try grad clipping along with lower LR? It helped in my case. PL has an easy callback for that. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2331739,
          "author_name": "Maximiliano Diaz Battan",
          "author_url": "",
          "post_date": "2023-07-05T18:34:44.190000",
          "content": "<p>Thank you very much Phaedrus. I haven't tried the 2 things together.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2330168,
      "author_name": "Maximiliano Diaz Battan",
      "author_url": "",
      "post_date": "2023-07-04T19:12:25.807000",
      "content": "<p>I think I spoke too soon, as I'm once again seeing NaN values in the loss with fp16. Let the debugging begin.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2330276,
          "author_name": "Bartley",
          "author_url": "",
          "post_date": "2023-07-04T21:17:08.200000",
          "content": "<p>When I use any learning rate value greater than 5e-4, my loss function occasionally goes to NAN as well.  </p>\n<p>Any learning rate value below this has not gone to NAN so far.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2331131,
              "author_name": "Maximiliano Diaz Battan",
              "author_url": "",
              "post_date": "2023-07-05T11:37:48.953000",
              "content": "<p>Thank you very much Bartley. By following your advice, I have been able to solve it for now. With 5e-4 missing values kept appearing, but on the epochs 8-10. I continued testing the lr, and with 1e-4 for the moment in 20 epochs they disappeared. It's quite strange because I was using much high on Kaggle, and I didn't encounter a single NaN value. Thank you!</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2331480,
              "author_name": "JEANMPIA",
              "author_url": "",
              "post_date": "2023-07-05T15:19:54.150000",
              "content": "<p>Are you using SGD ? if so do you clip the gradients ?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2331744,
              "author_name": "Maximiliano Diaz Battan",
              "author_url": "",
              "post_date": "2023-07-05T18:39:30.550000",
              "content": "<p>Thanks Janmpia, but no, I'm using Adam. I've tried gradient clipping, but since the problem, it's more focused on the validation loss, I guess that's why didn't work for me. Reducing the lr it's the only solution until now. Thanks anyway man, really appreciated it.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2329306,
      "author_name": "Mikhail Kotyushev",
      "author_url": "",
      "post_date": "2023-07-04T07:37:36.057000",
      "content": "<p>Hi! What are host system CUDA versions? I have seen strange behaviour running inference on kaggle 11.3 vs training on local 12.0, got it fixed when updated inference kaggle env to 11.7.</p>\n<p>One easiest thing to try is to rebuild the container without cache in paperclip env, at least that what I do when something is messed up with docker.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2329597,
          "author_name": "Maximiliano Diaz Battan",
          "author_url": "",
          "post_date": "2023-07-04T11:09:12.010000",
          "content": "<p>Thanks Mikhail, very useful info. Both environments now have CUDA 11.8. After double-checking everything and confirming that everything is the same, I managed to \"solve\" the issue, at least partially for now. I'm not sure why, but by slightly tweaking the learning rate, the NaN values have disappeared. I trained for 20 epochs, and so far, there are no more NaN values. I'm not sure why the same learning rate and parameters work fine on Kaggle but not on Paperspace. I suppose I need to check the code for any bugs.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2330487,
              "author_name": "Mikhail Kotyushev",
              "author_url": "",
              "post_date": "2023-07-05T03:56:17.940000",
              "content": "<p>Try to skip steps which have NaN gradients and see if it is occasional or training just breaks at some point. PL allows to return None from <code>LightningModule.training_step</code> to skip the step. I also have seen both NaNs and Infs in gradients at the beginning and with skipping / clipping model trains ok.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2Fa9d806c5a20c993119143d1658f5e9d4%2Ftrain_grads_nan.png?generation=1688529168741579&amp;alt=media\" alt=\"grads vs loss\"></p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2331136,
              "author_name": "Maximiliano Diaz Battan",
              "author_url": "",
              "post_date": "2023-07-05T11:40:58.573000",
              "content": "<p>I didn't know that. Thanks a lot, Mikhail. I'm going to test this as well.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2328298,
      "author_name": "JEANMPIA",
      "author_url": "",
      "post_date": "2023-07-03T13:35:32.593000",
      "content": "<p>are you training with MAnet or DeepLabv3+ by any chance ? if so the same happens to me on fp16</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2328451,
          "author_name": "june",
          "author_url": "",
          "post_date": "2023-07-03T15:27:47.837000",
          "content": "<p>I had the same issue with MAnet; the loss became NaN after 3-4 epochs.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2328608,
          "author_name": "Maximiliano Diaz Battan",
          "author_url": "",
          "post_date": "2023-07-03T17:51:22.630000",
          "content": "<p>Thanks man, but no, it's just an Unet + Resnest101, works perfectly fine with fp32, but with fp16 don't. But at Kaggle did.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2330674,
              "author_name": "Wood Carter",
              "author_url": "",
              "post_date": "2023-07-05T06:19:21.163000",
              "content": "<p>Can restnest101 be sustained in Caggle notebook?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2331257,
              "author_name": "Maximiliano Diaz Battan",
              "author_url": "",
              "post_date": "2023-07-05T13:05:59.150000",
              "content": "<p>Yes man, you can, just adjust the batch size.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2336574,
              "author_name": "Wood Carter",
              "author_url": "",
              "post_date": "2023-07-09T12:44:42.803000",
              "content": "<p>Thanks very much, i will try it.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2341260,
              "author_name": "Wood Carter",
              "author_url": "",
              "post_date": "2023-07-12T02:03:51.943000",
              "content": "<p>Tried resnext101_32x4d, trained seperately first , got lower credit than 26d…………..</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2341916,
              "author_name": "Maximiliano Diaz Battan",
              "author_url": "",
              "post_date": "2023-07-12T11:11:41.823000",
              "content": "<p>try the resnest101e</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2328288": "Hi Kagglers, is anyone here training at Paperspace? Something very strange is happening to me, I'm using the Kaggle docker image, therefore the same libraries, same Cuda drivers, anyway, all the same, including the notebook and configs. But when I train with PyTorch Lightning in fp16, the losses become NaN, but in fp32 everything works fine. While in Kaggle everything works correctly in fp16, and as I mentioned previously, they are the same versions of libraries, PL, Pytorch, SMP, cuda drivers, etc. Has the same thing happened to anyone? Or do you know why something like this might happen?",
    "2370974": "Hi @maxdiazbattan,\n\nI had the same problem, but I fixed it by setting the epsilon in the Adam optimizer to 1e-4 (eps=1e-4). The default epsilon of Adam is set to [1e-8 ](https://pytorch.org/docs/stable/generated/torch.optim.Adam.html), which can cause some numerical instability after some epochs with fp16 I read. \n\nI hope this will also help you! Let me know!\n",
    "2329277": "Probably you should use bfloat16 instead of float16. Probably …\n\nwith torch.amp.autocast(device_type=“cuda”, dtype=torch.bfloat16):\n\nHere you can find some useful tips: https://lightning.ai/pages/community/tutorial/pytorch-memory-vit-llm/",
    "2331454": "Did you try grad clipping along with lower LR? It helped in my case. PL has an easy callback for that. ",
    "2330168": "I think I spoke too soon, as I'm once again seeing NaN values in the loss with fp16. Let the debugging begin.",
    "2329306": "Hi! What are host system CUDA versions? I have seen strange behaviour running inference on kaggle 11.3 vs training on local 12.0, got it fixed when updated inference kaggle env to 11.7.\n\nOne easiest thing to try is to rebuild the container without cache in paperclip env, at least that what I do when something is messed up with docker.",
    "2328298": "are you training with MAnet or DeepLabv3+ by any chance ? if so the same happens to me on fp16"
  }
}