{
  "id": 420926,
  "title": "Differential Learning rate for Transfer Learning",
  "url": "/competitions/google-research-identify-contrails-reduce-global-warming/discussion/420926",
  "author_name": "DeepUnderstanding",
  "post_date": "2023-07-03T08:42:02.724000",
  "votes": 2,
  "comment_count": 7,
  "views": 0,
  "content": "<p>To speed up the training process, we can use different learning rates for different parts of the model during fine-tuning. When we fine-tune a pretrained model, we want to preserve the knowledge it has already learned while allowing new layers to quickly adapt to the new task. To achieve this, we set a higher learning rate for the new layers, so they can adjust faster, and a lower learning rate for the pretrained layers, so their existing knowledge is not drastically changed. This helps strike a balance and ensures efficient training of the model.</p>\n<pre><code> ():\n    \n    param_optimizer = (model.named_parameters())\n\n    \n    encoder_parameters = [p  n, p  param_optimizer    n]\n    decoder_parameters = [p  n, p  param_optimizer     n]\n\n    optimizer_parameters = [  {: encoder_parameters, : encoder_lr},\n        {: decoder_parameters, : decoder_lr} ]\n\n      optimizer_parameters\n</code></pre>\n<p>here we initialized two different learning rates for encoder and decoder.</p>\n<p>Inside the training loop use this</p>\n<pre><code>lr = \noptimizer_parameters= get_optimizer_params(model)\noptimizer = torch.optim.Adam(optimizer_parameters, lr=lr)\n</code></pre>",
  "messages": [
    {
      "id": 2327917,
      "postDate": "2023-07-03T08:42:02.723Z",
      "content": "<p>To speed up the training process, we can use different learning rates for different parts of the model during fine-tuning. When we fine-tune a pretrained model, we want to preserve the knowledge it has already learned while allowing new layers to quickly adapt to the new task. To achieve this, we set a higher learning rate for the new layers, so they can adjust faster, and a lower learning rate for the pretrained layers, so their existing knowledge is not drastically changed. This helps strike a balance and ensures efficient training of the model.</p>\n<pre><code> ():\n    \n    param_optimizer = (model.named_parameters())\n\n    \n    encoder_parameters = [p  n, p  param_optimizer    n]\n    decoder_parameters = [p  n, p  param_optimizer     n]\n\n    optimizer_parameters = [  {: encoder_parameters, : encoder_lr},\n        {: decoder_parameters, : decoder_lr} ]\n\n      optimizer_parameters\n</code></pre>\n<p>here we initialized two different learning rates for encoder and decoder.</p>\n<p>Inside the training loop use this</p>\n<pre><code>lr = \noptimizer_parameters= get_optimizer_params(model)\noptimizer = torch.optim.Adam(optimizer_parameters, lr=lr)\n</code></pre>",
      "rawMarkdown": "To speed up the training process, we can use different learning rates for different parts of the model during fine-tuning. When we fine-tune a pretrained model, we want to preserve the knowledge it has already learned while allowing new layers to quickly adapt to the new task. To achieve this, we set a higher learning rate for the new layers, so they can adjust faster, and a lower learning rate for the pretrained layers, so their existing knowledge is not drastically changed. This helps strike a balance and ensures efficient training of the model.\n\n```python\ndef get_optimizer_params(model, encoder_lr = lr, decoder_lr = lr*10):\n    # Get the named parameters of the model\n    param_optimizer = list(model.named_parameters())\n\n    # Create separate parameter groups for the encoder and decoder\n    encoder_parameters = [p for n, p in param_optimizer if \"encoder\" in n]\n    decoder_parameters = [p for n, p in param_optimizer if \"encoder\" not in n]\n\n    optimizer_parameters = [  {'params': encoder_parameters, 'lr': encoder_lr},\n        {'params': decoder_parameters, 'lr': decoder_lr} ]\n    \n     return optimizer_parameters\n\n```\nhere we initialized two different learning rates for encoder and decoder.\n\nInside the training loop use this\n```python\nlr = 3e-4\noptimizer_parameters= get_optimizer_params(model)\noptimizer = torch.optim.Adam(optimizer_parameters, lr=lr)\n\n```",
      "votes": 1
    },
    {
      "id": 2328036,
      "postDate": "2023-07-03T10:15:36.600Z",
      "content": "<p>Is this ChatGPT ?</p>",
      "rawMarkdown": "Is this ChatGPT ?",
      "votes": -3,
      "replies": [
        {
          "id": 2328090,
          "postDate": "2023-07-03T10:58:05.117Z",
          "content": "<p>real human here</p>",
          "rawMarkdown": "real human here",
          "replies": [
            {
              "id": 2328871,
              "postDate": "2023-07-03T23:52:57.580Z",
              "content": "<p>Mea culpa sir</p>",
              "rawMarkdown": "Mea culpa sir"
            },
            {
              "id": 2329133,
              "postDate": "2023-07-04T06:06:49.027Z",
              "content": "<p>sorry for the bad joke lol, no its not chat gpt, although,  I tried to take inspiration from old competitions, my code was based on this notebook <a href=\"https://www.kaggle.com/code/kashiwaba/train-deberta-v3-large-with-optimization-approach\" target=\"_blank\">https://www.kaggle.com/code/kashiwaba/train-deberta-v3-large-with-optimization-approach</a>, differential learning rate worked really well for NLP models, I wanted to test it here, and also using it while training has given me a slight CV boost. Hope it heps</p>",
              "rawMarkdown": "sorry for the bad joke lol, no its not chat gpt, although,  I tried to take inspiration from old competitions, my code was based on this notebook https://www.kaggle.com/code/kashiwaba/train-deberta-v3-large-with-optimization-approach, differential learning rate worked really well for NLP models, I wanted to test it here, and also using it while training has given me a slight CV boost. Hope it heps",
              "votes": 1
            },
            {
              "id": 2329870,
              "postDate": "2023-07-04T14:59:36.753Z",
              "content": "<p>no hard feeling dw :)</p>",
              "rawMarkdown": "no hard feeling dw :)",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2328949,
      "postDate": "2023-07-04T01:42:16.580Z",
      "content": "<p>May I know which pretrained model are you using?</p>",
      "rawMarkdown": "May I know which pretrained model are you using?",
      "replies": [
        {
          "id": 2329127,
          "postDate": "2023-07-04T05:59:38.240Z",
          "content": "<p>I used this on <a href=\"https://www.kaggle.com/code/shashwatraman/simple-unet-baseline-train-lb-0-580\" target=\"_blank\">https://www.kaggle.com/code/shashwatraman/simple-unet-baseline-train-lb-0-580</a> and got validation around 0.61 for threshold 0.01.</p>",
          "rawMarkdown": "I used this on https://www.kaggle.com/code/shashwatraman/simple-unet-baseline-train-lb-0-580 and got validation around 0.61 for threshold 0.01.\n\n"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2328036,
      "author_name": "JEANMPIA",
      "author_url": "",
      "post_date": "2023-07-03T10:15:36.600000",
      "content": "<p>Is this ChatGPT ?</p>",
      "votes": -3,
      "replies": [
        {
          "id": 2328090,
          "author_name": "DeepUnderstanding",
          "author_url": "",
          "post_date": "2023-07-03T10:58:05.117000",
          "content": "<p>real human here</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2328871,
              "author_name": "JEANMPIA",
              "author_url": "",
              "post_date": "2023-07-03T23:52:57.580000",
              "content": "<p>Mea culpa sir</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2329133,
              "author_name": "DeepUnderstanding",
              "author_url": "",
              "post_date": "2023-07-04T06:06:49.027000",
              "content": "<p>sorry for the bad joke lol, no its not chat gpt, although,  I tried to take inspiration from old competitions, my code was based on this notebook <a href=\"https://www.kaggle.com/code/kashiwaba/train-deberta-v3-large-with-optimization-approach\" target=\"_blank\">https://www.kaggle.com/code/kashiwaba/train-deberta-v3-large-with-optimization-approach</a>, differential learning rate worked really well for NLP models, I wanted to test it here, and also using it while training has given me a slight CV boost. Hope it heps</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2329870,
              "author_name": "JEANMPIA",
              "author_url": "",
              "post_date": "2023-07-04T14:59:36.753000",
              "content": "<p>no hard feeling dw :)</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2328949,
      "author_name": "william.wu",
      "author_url": "",
      "post_date": "2023-07-04T01:42:16.580000",
      "content": "<p>May I know which pretrained model are you using?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2329127,
          "author_name": "DeepUnderstanding",
          "author_url": "",
          "post_date": "2023-07-04T05:59:38.240000",
          "content": "<p>I used this on <a href=\"https://www.kaggle.com/code/shashwatraman/simple-unet-baseline-train-lb-0-580\" target=\"_blank\">https://www.kaggle.com/code/shashwatraman/simple-unet-baseline-train-lb-0-580</a> and got validation around 0.61 for threshold 0.01.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2327917": "To speed up the training process, we can use different learning rates for different parts of the model during fine-tuning. When we fine-tune a pretrained model, we want to preserve the knowledge it has already learned while allowing new layers to quickly adapt to the new task. To achieve this, we set a higher learning rate for the new layers, so they can adjust faster, and a lower learning rate for the pretrained layers, so their existing knowledge is not drastically changed. This helps strike a balance and ensures efficient training of the model.\n\n```python\ndef get_optimizer_params(model, encoder_lr = lr, decoder_lr = lr*10):\n    # Get the named parameters of the model\n    param_optimizer = list(model.named_parameters())\n\n    # Create separate parameter groups for the encoder and decoder\n    encoder_parameters = [p for n, p in param_optimizer if \"encoder\" in n]\n    decoder_parameters = [p for n, p in param_optimizer if \"encoder\" not in n]\n\n    optimizer_parameters = [  {'params': encoder_parameters, 'lr': encoder_lr},\n        {'params': decoder_parameters, 'lr': decoder_lr} ]\n    \n     return optimizer_parameters\n\n```\nhere we initialized two different learning rates for encoder and decoder.\n\nInside the training loop use this\n```python\nlr = 3e-4\noptimizer_parameters= get_optimizer_params(model)\noptimizer = torch.optim.Adam(optimizer_parameters, lr=lr)\n\n```",
    "2328036": "Is this ChatGPT ?",
    "2328949": "May I know which pretrained model are you using?"
  }
}