{
  "id": 388258,
  "title": "Lion Optimizer",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/388258",
  "author_name": "Gunes Evitan",
  "post_date": "2023-02-16T17:08:35.054000",
  "votes": 21,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I was trying Google's new optimizer Lion. According to its paper, it beats AdamW on ImageNet. I'm using lucidrains' version and it doesn't work better than AdamW for me, maybe it needs further tuning. I'll update this post after some trials.</p>\n<ul>\n<li>Paper: <a href=\"https://arxiv.org/abs/2302.06675\" target=\"_blank\">https://arxiv.org/abs/2302.06675</a></li>\n<li>Google repository: <a href=\"https://github.com/google/automl/tree/master/lion\" target=\"_blank\">https://github.com/google/automl/tree/master/lion</a></li>\n<li>lucidrains repository: <a href=\"https://github.com/lucidrains/lion-pytorch\" target=\"_blank\">https://github.com/lucidrains/lion-pytorch</a></li>\n</ul>\n<p>Update 1: When I use the same or 3x smaller learning rate compared to AdamW, loss diverges. 10x smaller learning rate looks okay but val score isn't as good as AdamW.<br>\nUpdate 2: I tried 10x, 20x, 100x smaller learning rates with higher weight decay but none of the trials worked better than AdamW.</p>",
  "messages": [
    {
      "id": 2147513,
      "postDate": "2023-02-16T17:08:35.053Z",
      "content": "<p>I was trying Google's new optimizer Lion. According to its paper, it beats AdamW on ImageNet. I'm using lucidrains' version and it doesn't work better than AdamW for me, maybe it needs further tuning. I'll update this post after some trials.</p>\n<ul>\n<li>Paper: <a href=\"https://arxiv.org/abs/2302.06675\" target=\"_blank\">https://arxiv.org/abs/2302.06675</a></li>\n<li>Google repository: <a href=\"https://github.com/google/automl/tree/master/lion\" target=\"_blank\">https://github.com/google/automl/tree/master/lion</a></li>\n<li>lucidrains repository: <a href=\"https://github.com/lucidrains/lion-pytorch\" target=\"_blank\">https://github.com/lucidrains/lion-pytorch</a></li>\n</ul>\n<p>Update 1: When I use the same or 3x smaller learning rate compared to AdamW, loss diverges. 10x smaller learning rate looks okay but val score isn't as good as AdamW.<br>\nUpdate 2: I tried 10x, 20x, 100x smaller learning rates with higher weight decay but none of the trials worked better than AdamW.</p>",
      "rawMarkdown": "I was trying Google's new optimizer Lion. According to its paper, it beats AdamW on ImageNet. I'm using lucidrains' version and it doesn't work better than AdamW for me, maybe it needs further tuning. I'll update this post after some trials.\n\n* Paper: https://arxiv.org/abs/2302.06675\n* Google repository: https://github.com/google/automl/tree/master/lion\n* lucidrains repository: https://github.com/lucidrains/lion-pytorch\n\nUpdate 1: When I use the same or 3x smaller learning rate compared to AdamW, loss diverges. 10x smaller learning rate looks okay but val score isn't as good as AdamW.\nUpdate 2: I tried 10x, 20x, 100x smaller learning rates with higher weight decay but none of the trials worked better than AdamW.",
      "votes": 21
    },
    {
      "id": 2150045,
      "postDate": "2023-02-18T22:15:26.120Z",
      "content": "<p>Lion is now part of timm (optim). </p>\n<p>I tested Lion and have not got good result on this dataset and selected nn architecture. Probably more experiments are required to say something about Lion.</p>",
      "rawMarkdown": "Lion is now part of timm (optim). \n\nI tested Lion and have not got good result on this dataset and selected nn architecture. Probably more experiments are required to say something about Lion.",
      "votes": 2
    },
    {
      "id": 2149400,
      "postDate": "2023-02-18T09:10:15.427Z",
      "content": "<p>I tried Lion with ConvNextV2, but I didn't get better results. It might need some tuning, though.</p>",
      "rawMarkdown": "I tried Lion with ConvNextV2, but I didn't get better results. It might need some tuning, though."
    },
    {
      "id": 2149174,
      "postDate": "2023-02-18T03:18:49.347Z",
      "content": "<p>I tried it, but loss diverges with small learing rate (1E-4). Maybe it cannot be used with fp16 mixed precision.</p>",
      "rawMarkdown": "I tried it, but loss diverges with small learing rate (1E-4). Maybe it cannot be used with fp16 mixed precision."
    },
    {
      "id": 2149141,
      "postDate": "2023-02-18T01:50:35.100Z",
      "content": "<p>anyone get it to work for vision transformer in this competitor?<br>\n(i haven't got good results yet …. but i haven't tune yet)</p>",
      "rawMarkdown": "anyone get it to work for vision transformer in this competitor?\n(i haven't got good results yet .... but i haven't tune yet)\n"
    },
    {
      "id": 2147571,
      "postDate": "2023-02-16T18:00:52.167Z",
      "content": "<p>Just replaced with adamw with lion (tf 2), looks promising. Thanks for informing. </p>",
      "rawMarkdown": "Just replaced with adamw with lion (tf 2), looks promising. Thanks for informing. "
    }
  ],
  "comments": [
    {
      "id": 2150045,
      "author_name": "Remek Kinas",
      "author_url": "",
      "post_date": "2023-02-18T22:15:26.120000",
      "content": "<p>Lion is now part of timm (optim). </p>\n<p>I tested Lion and have not got good result on this dataset and selected nn architecture. Probably more experiments are required to say something about Lion.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2149400,
      "author_name": "Rasoul Mojtahedzadeh",
      "author_url": "",
      "post_date": "2023-02-18T09:10:15.427000",
      "content": "<p>I tried Lion with ConvNextV2, but I didn't get better results. It might need some tuning, though.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2149174,
      "author_name": "ECO",
      "author_url": "",
      "post_date": "2023-02-18T03:18:49.347000",
      "content": "<p>I tried it, but loss diverges with small learing rate (1E-4). Maybe it cannot be used with fp16 mixed precision.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2149141,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-02-18T01:50:35.100000",
      "content": "<p>anyone get it to work for vision transformer in this competitor?<br>\n(i haven't got good results yet …. but i haven't tune yet)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2147571,
      "author_name": "Innat",
      "author_url": "",
      "post_date": "2023-02-16T18:00:52.167000",
      "content": "<p>Just replaced with adamw with lion (tf 2), looks promising. Thanks for informing. </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2147513": "I was trying Google's new optimizer Lion. According to its paper, it beats AdamW on ImageNet. I'm using lucidrains' version and it doesn't work better than AdamW for me, maybe it needs further tuning. I'll update this post after some trials.\n\n* Paper: https://arxiv.org/abs/2302.06675\n* Google repository: https://github.com/google/automl/tree/master/lion\n* lucidrains repository: https://github.com/lucidrains/lion-pytorch\n\nUpdate 1: When I use the same or 3x smaller learning rate compared to AdamW, loss diverges. 10x smaller learning rate looks okay but val score isn't as good as AdamW.\nUpdate 2: I tried 10x, 20x, 100x smaller learning rates with higher weight decay but none of the trials worked better than AdamW.",
    "2150045": "Lion is now part of timm (optim). \n\nI tested Lion and have not got good result on this dataset and selected nn architecture. Probably more experiments are required to say something about Lion.",
    "2149400": "I tried Lion with ConvNextV2, but I didn't get better results. It might need some tuning, though.",
    "2149174": "I tried it, but loss diverges with small learing rate (1E-4). Maybe it cannot be used with fp16 mixed precision.",
    "2149141": "anyone get it to work for vision transformer in this competitor?\n(i haven't got good results yet .... but i haven't tune yet)\n",
    "2147571": "Just replaced with adamw with lion (tf 2), looks promising. Thanks for informing. "
  }
}