{
  "id": 375791,
  "title": "ConvNeXt-V2 : Arrival of new SOTA CNN architectures",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/375791",
  "author_name": "Nischay Dhankhar",
  "post_date": "2023-01-03T14:20:37.678000",
  "votes": 18,
  "comment_count": 7,
  "views": 0,
  "content": "<p><strong>ConvNeXt V2: Co-designing and Scaling ConvNets with Masked Autoencoders</strong></p>\n<p>An improvised version of popular ConvNeXt models which were widely used in several past Kaggle competitions specially Segmentation ones, is published today by Facebook research. This makes me wonder how well they could be utilised in this competition. </p>\n<p><strong>Arxiv link:</strong> <a href=\"https://arxiv.org/pdf/2301.00808v1.pdf\" target=\"_blank\">https://arxiv.org/pdf/2301.00808v1.pdf</a></p>\n<p><strong>Pytorch code and pre-trained weights are also available :</strong> <a href=\"https://github.com/facebookresearch/ConvNeXt-V2\" target=\"_blank\">https://github.com/facebookresearch/ConvNeXt-V2</a> </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2Fd4b66fa511ded2a42f952015f81acaae%2FScreenshot%202023-01-03%20at%207.42.25%20PM.png?generation=1672755176468434&amp;alt=media\" alt=\"\"></p>\n<p>Comparing the results, there is around a 1% boost in accuracy scores over V1 models. Also to mention,  ConvNeXt V2-H is introduced which is pretrained on 512 x 512 images and could be impactful in this competition. </p>\n<p>Two major changes I could notice in the new architectures are self-supervised learning with the concept of <strong>Masked AutoEncoders</strong> and the introduction of a new normalization technique named <strong>Global Response Normalization</strong>.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2F665fc1056951260cd5d2a048bfbe10ae%2FScreenshot%202023-01-03%20at%207.48.29%20PM.png?generation=1672755529498810&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 2084471,
      "postDate": "2023-01-03T14:20:37.680Z",
      "content": "<p><strong>ConvNeXt V2: Co-designing and Scaling ConvNets with Masked Autoencoders</strong></p>\n<p>An improvised version of popular ConvNeXt models which were widely used in several past Kaggle competitions specially Segmentation ones, is published today by Facebook research. This makes me wonder how well they could be utilised in this competition. </p>\n<p><strong>Arxiv link:</strong> <a href=\"https://arxiv.org/pdf/2301.00808v1.pdf\" target=\"_blank\">https://arxiv.org/pdf/2301.00808v1.pdf</a></p>\n<p><strong>Pytorch code and pre-trained weights are also available :</strong> <a href=\"https://github.com/facebookresearch/ConvNeXt-V2\" target=\"_blank\">https://github.com/facebookresearch/ConvNeXt-V2</a> </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2Fd4b66fa511ded2a42f952015f81acaae%2FScreenshot%202023-01-03%20at%207.42.25%20PM.png?generation=1672755176468434&amp;alt=media\" alt=\"\"></p>\n<p>Comparing the results, there is around a 1% boost in accuracy scores over V1 models. Also to mention,  ConvNeXt V2-H is introduced which is pretrained on 512 x 512 images and could be impactful in this competition. </p>\n<p>Two major changes I could notice in the new architectures are self-supervised learning with the concept of <strong>Masked AutoEncoders</strong> and the introduction of a new normalization technique named <strong>Global Response Normalization</strong>.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2F665fc1056951260cd5d2a048bfbe10ae%2FScreenshot%202023-01-03%20at%207.48.29%20PM.png?generation=1672755529498810&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "**ConvNeXt V2: Co-designing and Scaling ConvNets with Masked Autoencoders**\n\nAn improvised version of popular ConvNeXt models which were widely used in several past Kaggle competitions specially Segmentation ones, is published today by Facebook research. This makes me wonder how well they could be utilised in this competition. \n\n**Arxiv link:** https://arxiv.org/pdf/2301.00808v1.pdf\n\n**Pytorch code and pre-trained weights are also available :** https://github.com/facebookresearch/ConvNeXt-V2 \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2Fd4b66fa511ded2a42f952015f81acaae%2FScreenshot%202023-01-03%20at%207.42.25%20PM.png?generation=1672755176468434&alt=media)\n\nComparing the results, there is around a 1% boost in accuracy scores over V1 models. Also to mention,  ConvNeXt V2-H is introduced which is pretrained on 512 x 512 images and could be impactful in this competition. \n\nTwo major changes I could notice in the new architectures are self-supervised learning with the concept of **Masked AutoEncoders** and the introduction of a new normalization technique named **Global Response Normalization**.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2F665fc1056951260cd5d2a048bfbe10ae%2FScreenshot%202023-01-03%20at%207.48.29%20PM.png?generation=1672755529498810&alt=media)",
      "votes": 18
    },
    {
      "id": 2084889,
      "postDate": "2023-01-03T20:26:59.540Z",
      "content": "<p>The original code and weights appear to be CC - non commercial though, jfyi.</p>",
      "rawMarkdown": "The original code and weights appear to be CC - non commercial though, jfyi.",
      "votes": 11,
      "replies": [
        {
          "id": 2085094,
          "postDate": "2023-01-03T23:37:53.047Z",
          "content": "<p>great point <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a>! thx for bringing this up, I would have not been aware of this for sure</p>\n<p>I wonder -- does this mean we cannot use this arch at Kaggle? Thinking kaggle might be okay (it might fall under research?) but commercial use might not be feasible?</p>\n<p>Seems a bit murky though </p>",
          "rawMarkdown": "great point @philippsinger! thx for bringing this up, I would have not been aware of this for sure\n\n\nI wonder -- does this mean we cannot use this arch at Kaggle? Thinking kaggle might be okay (it might fall under research?) but commercial use might not be feasible?\n\nSeems a bit murky though ",
          "votes": 2,
          "replies": [
            {
              "id": 2085843,
              "postDate": "2023-01-04T12:38:14.640Z",
              "content": "<p>I am personally not risking non-commercial datasets/models - and according to the rules they are not allowed.</p>",
              "rawMarkdown": "I am personally not risking non-commercial datasets/models - and according to the rules they are not allowed.",
              "votes": 5
            },
            {
              "id": 2085852,
              "postDate": "2023-01-04T12:42:01.250Z",
              "content": "<p>thank you very much for the reply, <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a>! appreciate it! 🙂</p>",
              "rawMarkdown": "thank you very much for the reply, @philippsinger! appreciate it! 🙂"
            },
            {
              "id": 2086083,
              "postDate": "2023-01-04T14:55:36.870Z",
              "content": "<p>I missed the license part, thanks for mentioning that.  </p>",
              "rawMarkdown": "I missed the license part, thanks for mentioning that.  "
            }
          ]
        },
        {
          "id": 2096113,
          "postDate": "2023-01-11T20:26:38.203Z",
          "content": "<p>amazing very interesting and details! thanks</p>",
          "rawMarkdown": "amazing very interesting and details! thanks"
        }
      ]
    },
    {
      "id": 2095554,
      "postDate": "2023-01-11T12:57:49.597Z",
      "content": "<p>All ConvNextV2 Tensorflow models are now available in the <a href=\"https://github.com/leondgarse/keras_cv_attention_models/tree/ecbd324103e116a6e28cb6033d6bb4bd60b92e68\" target=\"_blank\">keras-cv-attention-models</a> package, which can be installed with <code>pip install keras-cv-attention-models</code>.</p>\n<p>Examples can be found on the <a href=\"https://github.com/leondgarse/keras_cv_attention_models/tree/ecbd324103e116a6e28cb6033d6bb4bd60b92e68/keras_cv_attention_models/convnext\" target=\"_blank\">GitHub ConvNext page</a></p>\n<pre><code>from keras_cv_attention_models import convnext\nmm = convnext.ConvNeXtV2Nano(input_shape=(480, 480, 3), pretrained='imagenet21k-ft1k')\n# &gt;&gt;&gt;&gt; Load pretrained from: ~/.keras/models/convnext_v2_nano_384_imagenet21k-ft1k.h5\n\nfrom skimage.data import chelsea\npreds = mm(mm.preprocess_input(chelsea()))\nprint(mm.decode_predictions(preds)[0])\n# [('n02124075', 'Egyptian_cat', 0.7427755), ('n02123159', 'tiger_cat', 0.092012934), ...]\n</code></pre>",
      "rawMarkdown": "All ConvNextV2 Tensorflow models are now available in the [keras-cv-attention-models](https://github.com/leondgarse/keras_cv_attention_models/tree/ecbd324103e116a6e28cb6033d6bb4bd60b92e68) package, which can be installed with `pip install keras-cv-attention-models`.\n\nExamples can be found on the [GitHub ConvNext page](https://github.com/leondgarse/keras_cv_attention_models/tree/ecbd324103e116a6e28cb6033d6bb4bd60b92e68/keras_cv_attention_models/convnext)\n\n```\nfrom keras_cv_attention_models import convnext\nmm = convnext.ConvNeXtV2Nano(input_shape=(480, 480, 3), pretrained='imagenet21k-ft1k')\n# >>>> Load pretrained from: ~/.keras/models/convnext_v2_nano_384_imagenet21k-ft1k.h5\n\nfrom skimage.data import chelsea\npreds = mm(mm.preprocess_input(chelsea()))\nprint(mm.decode_predictions(preds)[0])\n# [('n02124075', 'Egyptian_cat', 0.7427755), ('n02123159', 'tiger_cat', 0.092012934), ...]\n```",
      "votes": 7
    }
  ],
  "comments": [
    {
      "id": 2084889,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2023-01-03T20:26:59.540000",
      "content": "<p>The original code and weights appear to be CC - non commercial though, jfyi.</p>",
      "votes": 11,
      "replies": [
        {
          "id": 2085094,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2023-01-03T23:37:53.047000",
          "content": "<p>great point <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a>! thx for bringing this up, I would have not been aware of this for sure</p>\n<p>I wonder -- does this mean we cannot use this arch at Kaggle? Thinking kaggle might be okay (it might fall under research?) but commercial use might not be feasible?</p>\n<p>Seems a bit murky though </p>",
          "votes": 2,
          "replies": [
            {
              "id": 2085843,
              "author_name": "Psi",
              "author_url": "",
              "post_date": "2023-01-04T12:38:14.640000",
              "content": "<p>I am personally not risking non-commercial datasets/models - and according to the rules they are not allowed.</p>",
              "votes": 5,
              "replies": []
            },
            {
              "id": 2085852,
              "author_name": "Radek Osmulski",
              "author_url": "",
              "post_date": "2023-01-04T12:42:01.250000",
              "content": "<p>thank you very much for the reply, <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a>! appreciate it! 🙂</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2086083,
              "author_name": "Nischay Dhankhar",
              "author_url": "",
              "post_date": "2023-01-04T14:55:36.870000",
              "content": "<p>I missed the license part, thanks for mentioning that.  </p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2096113,
          "author_name": "Matrix Terminator",
          "author_url": "",
          "post_date": "2023-01-11T20:26:38.203000",
          "content": "<p>amazing very interesting and details! thanks</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2095554,
      "author_name": "Mark Wijkhuizen",
      "author_url": "",
      "post_date": "2023-01-11T12:57:49.597000",
      "content": "<p>All ConvNextV2 Tensorflow models are now available in the <a href=\"https://github.com/leondgarse/keras_cv_attention_models/tree/ecbd324103e116a6e28cb6033d6bb4bd60b92e68\" target=\"_blank\">keras-cv-attention-models</a> package, which can be installed with <code>pip install keras-cv-attention-models</code>.</p>\n<p>Examples can be found on the <a href=\"https://github.com/leondgarse/keras_cv_attention_models/tree/ecbd324103e116a6e28cb6033d6bb4bd60b92e68/keras_cv_attention_models/convnext\" target=\"_blank\">GitHub ConvNext page</a></p>\n<pre><code>from keras_cv_attention_models import convnext\nmm = convnext.ConvNeXtV2Nano(input_shape=(480, 480, 3), pretrained='imagenet21k-ft1k')\n# &gt;&gt;&gt;&gt; Load pretrained from: ~/.keras/models/convnext_v2_nano_384_imagenet21k-ft1k.h5\n\nfrom skimage.data import chelsea\npreds = mm(mm.preprocess_input(chelsea()))\nprint(mm.decode_predictions(preds)[0])\n# [('n02124075', 'Egyptian_cat', 0.7427755), ('n02123159', 'tiger_cat', 0.092012934), ...]\n</code></pre>",
      "votes": 7,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2084471": "**ConvNeXt V2: Co-designing and Scaling ConvNets with Masked Autoencoders**\n\nAn improvised version of popular ConvNeXt models which were widely used in several past Kaggle competitions specially Segmentation ones, is published today by Facebook research. This makes me wonder how well they could be utilised in this competition. \n\n**Arxiv link:** https://arxiv.org/pdf/2301.00808v1.pdf\n\n**Pytorch code and pre-trained weights are also available :** https://github.com/facebookresearch/ConvNeXt-V2 \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2Fd4b66fa511ded2a42f952015f81acaae%2FScreenshot%202023-01-03%20at%207.42.25%20PM.png?generation=1672755176468434&alt=media)\n\nComparing the results, there is around a 1% boost in accuracy scores over V1 models. Also to mention,  ConvNeXt V2-H is introduced which is pretrained on 512 x 512 images and could be impactful in this competition. \n\nTwo major changes I could notice in the new architectures are self-supervised learning with the concept of **Masked AutoEncoders** and the introduction of a new normalization technique named **Global Response Normalization**.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2F665fc1056951260cd5d2a048bfbe10ae%2FScreenshot%202023-01-03%20at%207.48.29%20PM.png?generation=1672755529498810&alt=media)",
    "2084889": "The original code and weights appear to be CC - non commercial though, jfyi.",
    "2095554": "All ConvNextV2 Tensorflow models are now available in the [keras-cv-attention-models](https://github.com/leondgarse/keras_cv_attention_models/tree/ecbd324103e116a6e28cb6033d6bb4bd60b92e68) package, which can be installed with `pip install keras-cv-attention-models`.\n\nExamples can be found on the [GitHub ConvNext page](https://github.com/leondgarse/keras_cv_attention_models/tree/ecbd324103e116a6e28cb6033d6bb4bd60b92e68/keras_cv_attention_models/convnext)\n\n```\nfrom keras_cv_attention_models import convnext\nmm = convnext.ConvNeXtV2Nano(input_shape=(480, 480, 3), pretrained='imagenet21k-ft1k')\n# >>>> Load pretrained from: ~/.keras/models/convnext_v2_nano_384_imagenet21k-ft1k.h5\n\nfrom skimage.data import chelsea\npreds = mm(mm.preprocess_input(chelsea()))\nprint(mm.decode_predictions(preds)[0])\n# [('n02124075', 'Egyptian_cat', 0.7427755), ('n02123159', 'tiger_cat', 0.092012934), ...]\n```"
  }
}