{
  "id": 518599,
  "title": "Learning rate schedule",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/518599",
  "author_name": "Vasilis",
  "post_date": "2024-07-07T11:16:39.474000",
  "votes": 4,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hello guys, in this competition i realised how important LR is. Currently i use a linear LR with a warmup. So it has 2 parameters num_warmup_steps and num_trainin_steps. In the warm_up steps it raise liniearly from 0 to the starting LR, then from the starting LR it goes to 0 in num_training_steps. The LR in linear scheduler change after every iteration (and not after every eval or epoch). The results are decent but i struggle a bit restarting it and that has cost me computation time. Is there any other more intelligent scheduler that you suggest? </p>",
  "messages": [
    {
      "id": 2909896,
      "postDate": "2024-07-07T11:51:33.923Z",
      "content": "<p>It's hard to beat a fine-tuned cosine decay (and it's also very easy to fine-tune).</p>",
      "rawMarkdown": "It's hard to beat a fine-tuned cosine decay (and it's also very easy to fine-tune).",
      "votes": 3
    },
    {
      "id": 2909847,
      "postDate": "2024-07-07T11:16:39.473Z",
      "content": "<p>Hello guys, in this competition i realised how important LR is. Currently i use a linear LR with a warmup. So it has 2 parameters num_warmup_steps and num_trainin_steps. In the warm_up steps it raise liniearly from 0 to the starting LR, then from the starting LR it goes to 0 in num_training_steps. The LR in linear scheduler change after every iteration (and not after every eval or epoch). The results are decent but i struggle a bit restarting it and that has cost me computation time. Is there any other more intelligent scheduler that you suggest? </p>",
      "rawMarkdown": "Hello guys, in this competition i realised how important LR is. Currently i use a linear LR with a warmup. So it has 2 parameters num_warmup_steps and num_trainin_steps. In the warm_up steps it raise liniearly from 0 to the starting LR, then from the starting LR it goes to 0 in num_training_steps. The LR in linear scheduler change after every iteration (and not after every eval or epoch). The results are decent but i struggle a bit restarting it and that has cost me computation time. Is there any other more intelligent scheduler that you suggest? ",
      "votes": 4
    },
    {
      "id": 2910112,
      "postDate": "2024-07-07T14:09:20.167Z",
      "content": "<p>Thanks for the suggestions, looks like in pytorch there is CosineAnnealingLR and CosineAnnealingLRwithRestarts.  In both of them at some point the scheduler raises the LR back to the starting level, isnt this bad for the model? since if it has been training for quite some time and the LR becomes big it will start overshoot going away from the optimal values.</p>",
      "rawMarkdown": "Thanks for the suggestions, looks like in pytorch there is CosineAnnealingLR and CosineAnnealingLRwithRestarts.  In both of them at some point the scheduler raises the LR back to the starting level, isnt this bad for the model? since if it has been training for quite some time and the LR becomes big it will start overshoot going away from the optimal values.",
      "votes": 1,
      "replies": [
        {
          "id": 2910150,
          "postDate": "2024-07-07T14:46:18.630Z",
          "content": "<p>There are some papers suggesting multiple cosines are better. But the usual thing to do is just one half cycle. This is the SOTA. I don't know how to do it in pytorch. For tensorflow implementation see <a href=\"https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place\" target=\"_blank\">here</a>.</p>",
          "rawMarkdown": "There are some papers suggesting multiple cosines are better. But the usual thing to do is just one half cycle. This is the SOTA. I don't know how to do it in pytorch. For tensorflow implementation see [here](https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place).",
          "votes": 2
        }
      ]
    },
    {
      "id": 2909979,
      "postDate": "2024-07-07T12:33:22.017Z",
      "content": "<p>There are some papers in the area, on how to probe the required time for training of nn without re-training (e.g. schedule-free/constant lr + rapid decay in a separate process), but so far cosine is SOTA.</p>",
      "rawMarkdown": "There are some papers in the area, on how to probe the required time for training of nn without re-training (e.g. schedule-free/constant lr + rapid decay in a separate process), but so far cosine is SOTA.",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 2909896,
      "author_name": "greySnow",
      "author_url": "",
      "post_date": "2024-07-07T11:51:33.923000",
      "content": "<p>It's hard to beat a fine-tuned cosine decay (and it's also very easy to fine-tune).</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2910112,
      "author_name": "Vasilis",
      "author_url": "",
      "post_date": "2024-07-07T14:09:20.167000",
      "content": "<p>Thanks for the suggestions, looks like in pytorch there is CosineAnnealingLR and CosineAnnealingLRwithRestarts.  In both of them at some point the scheduler raises the LR back to the starting level, isnt this bad for the model? since if it has been training for quite some time and the LR becomes big it will start overshoot going away from the optimal values.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2910150,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2024-07-07T14:46:18.630000",
          "content": "<p>There are some papers suggesting multiple cosines are better. But the usual thing to do is just one half cycle. This is the SOTA. I don't know how to do it in pytorch. For tensorflow implementation see <a href=\"https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place\" target=\"_blank\">here</a>.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2909979,
      "author_name": "slime",
      "author_url": "",
      "post_date": "2024-07-07T12:33:22.017000",
      "content": "<p>There are some papers in the area, on how to probe the required time for training of nn without re-training (e.g. schedule-free/constant lr + rapid decay in a separate process), but so far cosine is SOTA.</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2909896": "It's hard to beat a fine-tuned cosine decay (and it's also very easy to fine-tune).",
    "2909847": "Hello guys, in this competition i realised how important LR is. Currently i use a linear LR with a warmup. So it has 2 parameters num_warmup_steps and num_trainin_steps. In the warm_up steps it raise liniearly from 0 to the starting LR, then from the starting LR it goes to 0 in num_training_steps. The LR in linear scheduler change after every iteration (and not after every eval or epoch). The results are decent but i struggle a bit restarting it and that has cost me computation time. Is there any other more intelligent scheduler that you suggest? ",
    "2910112": "Thanks for the suggestions, looks like in pytorch there is CosineAnnealingLR and CosineAnnealingLRwithRestarts.  In both of them at some point the scheduler raises the LR back to the starting level, isnt this bad for the model? since if it has been training for quite some time and the LR becomes big it will start overshoot going away from the optimal values.",
    "2909979": "There are some papers in the area, on how to probe the required time for training of nn without re-training (e.g. schedule-free/constant lr + rapid decay in a separate process), but so far cosine is SOTA."
  }
}