{
  "id": 419784,
  "title": "train with AMP or not? ",
  "url": "/competitions/google-research-identify-contrails-reduce-global-warming/discussion/419784",
  "author_name": "Sasha Mogilevskii",
  "post_date": "2023-06-27T14:47:49.346000",
  "votes": 3,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hello. Can you please tell me, do you train models with AMP or without?<br>\nPerhaps someone can recommend an article comparing the training methods?</p>",
  "messages": [
    {
      "id": 2320348,
      "postDate": "2023-06-27T17:58:06.750Z",
      "content": "<p>Are you training in Kaggle? FP16 works with the T4, but with P100 doesn't, slow down the training. Perhaps this article can clarify some doubts you have, <a href=\"https://magazine.sebastianraschka.com/p/accelerating-pytorch-model-training\" target=\"_blank\">https://magazine.sebastianraschka.com/p/accelerating-pytorch-model-training</a>.</p>",
      "rawMarkdown": "Are you training in Kaggle? FP16 works with the T4, but with P100 doesn't, slow down the training. Perhaps this article can clarify some doubts you have, https://magazine.sebastianraschka.com/p/accelerating-pytorch-model-training.",
      "votes": 4,
      "replies": [
        {
          "id": 2320354,
          "postDate": "2023-06-27T18:05:52.833Z",
          "content": "<p>No, I train models locally on A6000 x2 </p>",
          "rawMarkdown": "No, I train models locally on A6000 x2 ",
          "votes": 1
        }
      ]
    },
    {
      "id": 2320446,
      "postDate": "2023-06-27T19:33:28.170Z",
      "content": "<p>yes!!  I train with mixed precision.</p>",
      "rawMarkdown": "yes!!  I train with mixed precision.",
      "votes": 1,
      "replies": [
        {
          "id": 2320568,
          "postDate": "2023-06-27T22:12:28.650Z",
          "content": "<p>Have you probelms with NaN values in the loss? If yes, how did you solve this problem? Thank you.</p>",
          "rawMarkdown": "Have you probelms with NaN values in the loss? If yes, how did you solve this problem? Thank you.",
          "votes": 1,
          "replies": [
            {
              "id": 2320644,
              "postDate": "2023-06-28T02:03:01.997Z",
              "content": "<p>You can try gradient clipping and see if it helps.</p>",
              "rawMarkdown": "You can try gradient clipping and see if it helps.",
              "votes": 2
            },
            {
              "id": 2320756,
              "postDate": "2023-06-28T04:26:28.517Z",
              "content": "<p>exactly ! I clip the gradients</p>",
              "rawMarkdown": "exactly ! I clip the gradients",
              "votes": 1
            },
            {
              "id": 2320994,
              "postDate": "2023-06-28T07:59:27.367Z",
              "content": "<p>NaN values could come from different places in training:</p>\n<ul>\n<li>too high LR</li>\n<li>gradient exploding (can be connected with optimization, model architecture etc.)</li>\n<li>bad input (bad data, problem with augmentations) etc.</li>\n<li>data normalization</li>\n<li>etc.</li>\n</ul>\n<p>Before clipping gradient check your training pipeline.</p>",
              "rawMarkdown": "NaN values could come from different places in training:\n- too high LR\n- gradient exploding (can be connected with optimization, model architecture etc.)\n- bad input (bad data, problem with augmentations) etc.\n- data normalization\n- etc.\n\nBefore clipping gradient check your training pipeline.",
              "votes": 4
            },
            {
              "id": 2321038,
              "postDate": "2023-06-28T08:36:43.520Z",
              "content": "<p>I have this problem but only in the valid loss and very rarely. i'll gave to look it up more thouroughly</p>",
              "rawMarkdown": "I have this problem but only in the valid loss and very rarely. i'll gave to look it up more thouroughly",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2320177,
      "postDate": "2023-06-27T14:47:49.347Z",
      "content": "<p>Hello. Can you please tell me, do you train models with AMP or without?<br>\nPerhaps someone can recommend an article comparing the training methods?</p>",
      "rawMarkdown": "Hello. Can you please tell me, do you train models with AMP or without?\nPerhaps someone can recommend an article comparing the training methods?",
      "votes": 2
    },
    {
      "id": 2351232,
      "postDate": "2023-07-20T00:48:19.913Z",
      "content": "<p>Depends whether there's NAN or not. According to my experience, if you come across NAN loss for that model, it's hard to fix. So I go back to fp32</p>",
      "rawMarkdown": "Depends whether there's NAN or not. According to my experience, if you come across NAN loss for that model, it's hard to fix. So I go back to fp32"
    }
  ],
  "comments": [
    {
      "id": 2320348,
      "author_name": "Maximiliano Diaz Battan",
      "author_url": "",
      "post_date": "2023-06-27T17:58:06.750000",
      "content": "<p>Are you training in Kaggle? FP16 works with the T4, but with P100 doesn't, slow down the training. Perhaps this article can clarify some doubts you have, <a href=\"https://magazine.sebastianraschka.com/p/accelerating-pytorch-model-training\" target=\"_blank\">https://magazine.sebastianraschka.com/p/accelerating-pytorch-model-training</a>.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2320354,
          "author_name": "Sasha Mogilevskii",
          "author_url": "",
          "post_date": "2023-06-27T18:05:52.833000",
          "content": "<p>No, I train models locally on A6000 x2 </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2320446,
      "author_name": "Reacher",
      "author_url": "",
      "post_date": "2023-06-27T19:33:28.170000",
      "content": "<p>yes!!  I train with mixed precision.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2320568,
          "author_name": "Sasha Mogilevskii",
          "author_url": "",
          "post_date": "2023-06-27T22:12:28.650000",
          "content": "<p>Have you probelms with NaN values in the loss? If yes, how did you solve this problem? Thank you.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2320644,
              "author_name": "Phaedrus",
              "author_url": "",
              "post_date": "2023-06-28T02:03:01.997000",
              "content": "<p>You can try gradient clipping and see if it helps.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2320756,
              "author_name": "Reacher",
              "author_url": "",
              "post_date": "2023-06-28T04:26:28.517000",
              "content": "<p>exactly ! I clip the gradients</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2320994,
              "author_name": "Remek Kinas",
              "author_url": "",
              "post_date": "2023-06-28T07:59:27.367000",
              "content": "<p>NaN values could come from different places in training:</p>\n<ul>\n<li>too high LR</li>\n<li>gradient exploding (can be connected with optimization, model architecture etc.)</li>\n<li>bad input (bad data, problem with augmentations) etc.</li>\n<li>data normalization</li>\n<li>etc.</li>\n</ul>\n<p>Before clipping gradient check your training pipeline.</p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 2321038,
              "author_name": "JEANMPIA",
              "author_url": "",
              "post_date": "2023-06-28T08:36:43.520000",
              "content": "<p>I have this problem but only in the valid loss and very rarely. i'll gave to look it up more thouroughly</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2351232,
      "author_name": "william.wu",
      "author_url": "",
      "post_date": "2023-07-20T00:48:19.913000",
      "content": "<p>Depends whether there's NAN or not. According to my experience, if you come across NAN loss for that model, it's hard to fix. So I go back to fp32</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2320348": "Are you training in Kaggle? FP16 works with the T4, but with P100 doesn't, slow down the training. Perhaps this article can clarify some doubts you have, https://magazine.sebastianraschka.com/p/accelerating-pytorch-model-training.",
    "2320446": "yes!!  I train with mixed precision.",
    "2320177": "Hello. Can you please tell me, do you train models with AMP or without?\nPerhaps someone can recommend an article comparing the training methods?",
    "2351232": "Depends whether there's NAN or not. According to my experience, if you come across NAN loss for that model, it's hard to fix. So I go back to fp32"
  }
}