{
  "id": 507854,
  "title": "Train Loss spikes",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/507854",
  "author_name": "Vasilis",
  "post_date": "2024-05-27T14:05:19.638000",
  "votes": 2,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hello everybody, i notice that when i multiply the y with the weights before training my train loss has big spikes</p>\n<p>epoch: 0 , chunk: 0, Step 100, Training Loss: 0.5917 iterations: 100 time: 2783.370905160904<br>\nepoch: 0 , chunk: 0, Step 200, Training Loss: 0.5177 iterations: 200 time: 3947.4694044589996<br>\nepoch: 0 , chunk: 0, Step 300, Training Loss: 0.5088 iterations: 300 time: 5116.634320020676<br>\nepoch: 0 , chunk: 0, Step 400, Training Loss: 0.5138 iterations: 400 time: 6274.054301977158<br>\nepoch: 0 model saved chunk: 0 iterations: 486 val loss: 0.5341707844085616 time: 3782.210193157196</p>\n<p>tensor_data input shape: torch.Size([497123, 60, 25])<br>\nepoch: 0 chunk: 1 prep chunk time: 15.340328693389893<br>\nepoch: 0 , chunk: 1, Step 100, Training Loss: 0.4475 iterations: 586 time: 2813.811069250107<br>\nepoch: 0 , chunk: 1, Step 200, Training Loss: 0.4961 iterations: 686 time: 6220.386344671249<br>\nepoch: 0 , chunk: 1, Step 300, Training Loss: 0.4782 iterations: 786 time: 7684.73356628418<br>\nepoch: 0 , chunk: 1, Step 400, Training Loss: 1.3982 iterations: 886 time: 9230.30635881424</p>\n<p>This happens regardless if i use float32 or float64, any idea why that happens?<br>\nThis does not occur if i dont multiply y with the weights before training.</p>\n<p>This is what i do:</p>\n<pre><code>def calc:\n     y = pl.read\n\n     # Convert the DataFrame  a NumPy \n     y_np = y.\n\n     # Apply weights\n     y_weighted = y_npTARGET_WEIGHTS\n\n     # Calculate mean  standard deviation\n     mean_y = np.mean(y_weighted, axis=)\n     std_y = np.std(y_weighted, axis=)\n</code></pre>\n<p>And then later i use the mean and std like that:</p>\n<pre><code> = y.to_numpy()\n = y * TARGET_WEIGHTS\n = (y - mean_y) / std_y\n = torch.tensor(y, dtype=torch.float64).cuda()\n</code></pre>",
  "messages": [
    {
      "id": 2839405,
      "postDate": "2024-05-27T14:05:19.637Z",
      "content": "<p>Hello everybody, i notice that when i multiply the y with the weights before training my train loss has big spikes</p>\n<p>epoch: 0 , chunk: 0, Step 100, Training Loss: 0.5917 iterations: 100 time: 2783.370905160904<br>\nepoch: 0 , chunk: 0, Step 200, Training Loss: 0.5177 iterations: 200 time: 3947.4694044589996<br>\nepoch: 0 , chunk: 0, Step 300, Training Loss: 0.5088 iterations: 300 time: 5116.634320020676<br>\nepoch: 0 , chunk: 0, Step 400, Training Loss: 0.5138 iterations: 400 time: 6274.054301977158<br>\nepoch: 0 model saved chunk: 0 iterations: 486 val loss: 0.5341707844085616 time: 3782.210193157196</p>\n<p>tensor_data input shape: torch.Size([497123, 60, 25])<br>\nepoch: 0 chunk: 1 prep chunk time: 15.340328693389893<br>\nepoch: 0 , chunk: 1, Step 100, Training Loss: 0.4475 iterations: 586 time: 2813.811069250107<br>\nepoch: 0 , chunk: 1, Step 200, Training Loss: 0.4961 iterations: 686 time: 6220.386344671249<br>\nepoch: 0 , chunk: 1, Step 300, Training Loss: 0.4782 iterations: 786 time: 7684.73356628418<br>\nepoch: 0 , chunk: 1, Step 400, Training Loss: 1.3982 iterations: 886 time: 9230.30635881424</p>\n<p>This happens regardless if i use float32 or float64, any idea why that happens?<br>\nThis does not occur if i dont multiply y with the weights before training.</p>\n<p>This is what i do:</p>\n<pre><code>def calc:\n     y = pl.read\n\n     # Convert the DataFrame  a NumPy \n     y_np = y.\n\n     # Apply weights\n     y_weighted = y_npTARGET_WEIGHTS\n\n     # Calculate mean  standard deviation\n     mean_y = np.mean(y_weighted, axis=)\n     std_y = np.std(y_weighted, axis=)\n</code></pre>\n<p>And then later i use the mean and std like that:</p>\n<pre><code> = y.to_numpy()\n = y * TARGET_WEIGHTS\n = (y - mean_y) / std_y\n = torch.tensor(y, dtype=torch.float64).cuda()\n</code></pre>",
      "rawMarkdown": "Hello everybody, i notice that when i multiply the y with the weights before training my train loss has big spikes\n\nepoch: 0 , chunk: 0, Step 100, Training Loss: 0.5917 iterations: 100 time: 2783.370905160904\nepoch: 0 , chunk: 0, Step 200, Training Loss: 0.5177 iterations: 200 time: 3947.4694044589996\nepoch: 0 , chunk: 0, Step 300, Training Loss: 0.5088 iterations: 300 time: 5116.634320020676\nepoch: 0 , chunk: 0, Step 400, Training Loss: 0.5138 iterations: 400 time: 6274.054301977158\nepoch: 0 model saved chunk: 0 iterations: 486 val loss: 0.5341707844085616 time: 3782.210193157196\n\ntensor_data input shape: torch.Size([497123, 60, 25])\nepoch: 0 chunk: 1 prep chunk time: 15.340328693389893\nepoch: 0 , chunk: 1, Step 100, Training Loss: 0.4475 iterations: 586 time: 2813.811069250107\nepoch: 0 , chunk: 1, Step 200, Training Loss: 0.4961 iterations: 686 time: 6220.386344671249\nepoch: 0 , chunk: 1, Step 300, Training Loss: 0.4782 iterations: 786 time: 7684.73356628418\nepoch: 0 , chunk: 1, Step 400, Training Loss: 1.3982 iterations: 886 time: 9230.30635881424\n\nThis happens regardless if i use float32 or float64, any idea why that happens?\nThis does not occur if i dont multiply y with the weights before training.\n\nThis is what i do:\n```\ndef calc_mean_std():\n     y = pl.read_csv(file_path, columns=TARGET_COLS, n_rows=3000000)\n\n     # Convert the DataFrame to a NumPy array\n     y_np = y.to_numpy()\n\n     # Apply weights\n     y_weighted = y_np * TARGET_WEIGHTS\n\n     # Calculate mean and standard deviation\n     mean_y = np.mean(y_weighted, axis=0)\n     std_y = np.std(y_weighted, axis=0)\n```\nAnd then later i use the mean and std like that:\n```\ny = y.to_numpy()\ny = y * TARGET_WEIGHTS\ny = (y - mean_y) / std_y\ntensor_target = torch.tensor(y, dtype=torch.float64).cuda()\n```\n \n",
      "votes": 2
    },
    {
      "id": 2839965,
      "postDate": "2024-05-27T20:35:42.963Z",
      "content": "<p>Try adding something like that in the first function</p>\n<p><code>std_y = np.clip(std_y, 1e-12)</code></p>",
      "rawMarkdown": "Try adding something like that in the first function\n\n`std_y = np.clip(std_y, 1e-12)`",
      "votes": 1,
      "replies": [
        {
          "id": 2840332,
          "postDate": "2024-05-28T04:37:28.693Z",
          "content": "<p>oh sorry i forgot to mention that i was already doing this </p>\n<pre><code>  e-\nstd_y[std_y  ]  epsilon\n</code></pre>",
          "rawMarkdown": "oh sorry i forgot to mention that i was already doing this \n```\nepsilon = 1e-10\nstd_y[std_y == 0] = epsilon\n```",
          "replies": [
            {
              "id": 2840655,
              "postDate": "2024-05-28T07:31:00.760Z",
              "content": "<p>This is different from clipping</p>",
              "rawMarkdown": "This is different from clipping",
              "votes": 3
            }
          ]
        },
        {
          "id": 2841252,
          "postDate": "2024-05-28T13:34:07.580Z",
          "content": "<p>Thanks both for your input, i have added the clip but i think another step needed to stabilise it was to calculate the mean and std from the whole train dataset. I dont get this, why calculating mean and std using 3  million rows ( so approx 30%) of the dataset can give so much instability? By the way although much much more stable now i still see some small bumps upwards like from step 200 to 300 but i hope that is normal 😅</p>\n<pre><code> input shape: torch.Size([, , ])\n:  chunk:  prep chunk time: .\n:  , chunk: , Step , Training Loss: . iterations:  time: .\n:  , chunk: , Step , Training Loss: . iterations:  time: .\n:  , chunk: , Step , Training Loss: . iterations:  time: .\n:  , chunk: , Step , Training Loss: . iterations:  time: .\n</code></pre>",
          "rawMarkdown": "Thanks both for your input, i have added the clip but i think another step needed to stabilise it was to calculate the mean and std from the whole train dataset. I dont get this, why calculating mean and std using 3  million rows ( so approx 30%) of the dataset can give so much instability? By the way although much much more stable now i still see some small bumps upwards like from step 200 to 300 but i hope that is normal 😅\n```\ntensor_data input shape: torch.Size([497111, 60, 25])\nepoch: 0 chunk: 2 prep chunk time: 9.979077577590942\nepoch: 0 , chunk: 2, Step 100, Training Loss: 0.4969 iterations: 1072 time: 344.9132242202759\nepoch: 0 , chunk: 2, Step 200, Training Loss: 0.4819 iterations: 1172 time: 678.7866270542145\nepoch: 0 , chunk: 2, Step 300, Training Loss: 0.5447 iterations: 1272 time: 1014.9243459701538\nepoch: 0 , chunk: 2, Step 400, Training Loss: 0.4761 iterations: 1372 time: 1346.2622275352478\n```",
          "votes": 1,
          "replies": [
            {
              "id": 2841325,
              "postDate": "2024-05-28T14:25:21.763Z",
              "content": "<p>It can happen, sometimes warmup up in learning rate scheduling and/or norm clipping the gradient can help</p>",
              "rawMarkdown": "It can happen, sometimes warmup up in learning rate scheduling and/or norm clipping the gradient can help",
              "votes": 2
            },
            {
              "id": 2841493,
              "postDate": "2024-05-28T15:40:15.717Z",
              "content": "<p>thanks i will try them. I use scheduler = optim.lr_scheduler.PolynomialLR(optimizer, power=1.0, total_iters=50) with a LR of 1e-3. If i understand correctly you suggest that this can be also cause of the unstability?</p>",
              "rawMarkdown": "thanks i will try them. I use scheduler = optim.lr_scheduler.PolynomialLR(optimizer, power=1.0, total_iters=50) with a LR of 1e-3. If i understand correctly you suggest that this can be also cause of the unstability?"
            },
            {
              "id": 2842866,
              "postDate": "2024-05-29T09:46:00.347Z",
              "content": "<p>I would recommend reading on transformer training stability (paper inside).<br>\n<a href=\"https://github.com/LiyuanLucasLiu/Transformer-Clinic\" target=\"_blank\">https://github.com/LiyuanLucasLiu/Transformer-Clinic</a></p>",
              "rawMarkdown": "I would recommend reading on transformer training stability (paper inside).\nhttps://github.com/LiyuanLucasLiu/Transformer-Clinic",
              "votes": 3
            },
            {
              "id": 2843089,
              "postDate": "2024-05-29T12:06:23.960Z",
              "content": "<p>thank you, i will have a look!</p>",
              "rawMarkdown": "thank you, i will have a look!"
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2839965,
      "author_name": "Amedeo Biolatti",
      "author_url": "",
      "post_date": "2024-05-27T20:35:42.963000",
      "content": "<p>Try adding something like that in the first function</p>\n<p><code>std_y = np.clip(std_y, 1e-12)</code></p>",
      "votes": 1,
      "replies": [
        {
          "id": 2840332,
          "author_name": "Vasilis",
          "author_url": "",
          "post_date": "2024-05-28T04:37:28.693000",
          "content": "<p>oh sorry i forgot to mention that i was already doing this </p>\n<pre><code>  e-\nstd_y[std_y  ]  epsilon\n</code></pre>",
          "votes": 0,
          "replies": [
            {
              "id": 2840655,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-05-28T07:31:00.760000",
              "content": "<p>This is different from clipping</p>",
              "votes": 3,
              "replies": []
            }
          ]
        },
        {
          "id": 2841252,
          "author_name": "Vasilis",
          "author_url": "",
          "post_date": "2024-05-28T13:34:07.580000",
          "content": "<p>Thanks both for your input, i have added the clip but i think another step needed to stabilise it was to calculate the mean and std from the whole train dataset. I dont get this, why calculating mean and std using 3  million rows ( so approx 30%) of the dataset can give so much instability? By the way although much much more stable now i still see some small bumps upwards like from step 200 to 300 but i hope that is normal 😅</p>\n<pre><code> input shape: torch.Size([, , ])\n:  chunk:  prep chunk time: .\n:  , chunk: , Step , Training Loss: . iterations:  time: .\n:  , chunk: , Step , Training Loss: . iterations:  time: .\n:  , chunk: , Step , Training Loss: . iterations:  time: .\n:  , chunk: , Step , Training Loss: . iterations:  time: .\n</code></pre>",
          "votes": 1,
          "replies": [
            {
              "id": 2841325,
              "author_name": "Amedeo Biolatti",
              "author_url": "",
              "post_date": "2024-05-28T14:25:21.763000",
              "content": "<p>It can happen, sometimes warmup up in learning rate scheduling and/or norm clipping the gradient can help</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2841493,
              "author_name": "Vasilis",
              "author_url": "",
              "post_date": "2024-05-28T15:40:15.717000",
              "content": "<p>thanks i will try them. I use scheduler = optim.lr_scheduler.PolynomialLR(optimizer, power=1.0, total_iters=50) with a LR of 1e-3. If i understand correctly you suggest that this can be also cause of the unstability?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2842866,
              "author_name": "sroger",
              "author_url": "",
              "post_date": "2024-05-29T09:46:00.347000",
              "content": "<p>I would recommend reading on transformer training stability (paper inside).<br>\n<a href=\"https://github.com/LiyuanLucasLiu/Transformer-Clinic\" target=\"_blank\">https://github.com/LiyuanLucasLiu/Transformer-Clinic</a></p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2843089,
              "author_name": "Vasilis",
              "author_url": "",
              "post_date": "2024-05-29T12:06:23.960000",
              "content": "<p>thank you, i will have a look!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2839405": "Hello everybody, i notice that when i multiply the y with the weights before training my train loss has big spikes\n\nepoch: 0 , chunk: 0, Step 100, Training Loss: 0.5917 iterations: 100 time: 2783.370905160904\nepoch: 0 , chunk: 0, Step 200, Training Loss: 0.5177 iterations: 200 time: 3947.4694044589996\nepoch: 0 , chunk: 0, Step 300, Training Loss: 0.5088 iterations: 300 time: 5116.634320020676\nepoch: 0 , chunk: 0, Step 400, Training Loss: 0.5138 iterations: 400 time: 6274.054301977158\nepoch: 0 model saved chunk: 0 iterations: 486 val loss: 0.5341707844085616 time: 3782.210193157196\n\ntensor_data input shape: torch.Size([497123, 60, 25])\nepoch: 0 chunk: 1 prep chunk time: 15.340328693389893\nepoch: 0 , chunk: 1, Step 100, Training Loss: 0.4475 iterations: 586 time: 2813.811069250107\nepoch: 0 , chunk: 1, Step 200, Training Loss: 0.4961 iterations: 686 time: 6220.386344671249\nepoch: 0 , chunk: 1, Step 300, Training Loss: 0.4782 iterations: 786 time: 7684.73356628418\nepoch: 0 , chunk: 1, Step 400, Training Loss: 1.3982 iterations: 886 time: 9230.30635881424\n\nThis happens regardless if i use float32 or float64, any idea why that happens?\nThis does not occur if i dont multiply y with the weights before training.\n\nThis is what i do:\n```\ndef calc_mean_std():\n     y = pl.read_csv(file_path, columns=TARGET_COLS, n_rows=3000000)\n\n     # Convert the DataFrame to a NumPy array\n     y_np = y.to_numpy()\n\n     # Apply weights\n     y_weighted = y_np * TARGET_WEIGHTS\n\n     # Calculate mean and standard deviation\n     mean_y = np.mean(y_weighted, axis=0)\n     std_y = np.std(y_weighted, axis=0)\n```\nAnd then later i use the mean and std like that:\n```\ny = y.to_numpy()\ny = y * TARGET_WEIGHTS\ny = (y - mean_y) / std_y\ntensor_target = torch.tensor(y, dtype=torch.float64).cuda()\n```\n \n",
    "2839965": "Try adding something like that in the first function\n\n`std_y = np.clip(std_y, 1e-12)`"
  }
}