{
  "id": 344301,
  "title": "GCViT: Global Context Vision Transformer",
  "url": "/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/344301",
  "author_name": "Awsaf",
  "post_date": "2022-08-14T18:50:34.025000",
  "votes": 13,
  "comment_count": 4,
  "views": 0,
  "content": "<p><img src=\"https://raw.githubusercontent.com/awsaf49/gcvit-tf/main/image/lvg_arch.PNG\"></p>\n<p>Hello everyone,<br>\nAs you may already know, NVIDIA has recently published (20 June 2022) its paper, <strong>GCViT: Global Context Vision Transformer</strong> which outperforms <strong>ConvNeXt</strong> and <strong>SwinTransformer</strong> . As this competition is a Classification Problem (give or take), I think <strong>GCViT</strong> can be useful here. You are welcome to explore this model for this competition.</p>\n<blockquote>\n  <p>There was some issue with their weights and they recently released their <strong>final version of ImageNet pretrained weights</strong> which brought over <code>~1%</code> improvement over older weights.</p>\n</blockquote>\n<p>To ease things up, I've implemented this model using <strong>TensorFlow</strong> and created an open-source library <a href=\"https://github.com/awsaf49/gcvit-tf\" target=\"_blank\">gcvit-tf</a>. I've also made a notebook to help you get started with <strong>GCViT</strong>,. </p>\n<h2>Resources:</h2>\n<ul>\n<li><strong>Notebook:</strong> <a href=\"https://www.kaggle.com/code/awsaf49/gcvit-global-context-vision-transformer\" target=\"_blank\">GCViT: Global Context Vision Transformer</a><ul>\n<li>Refer to <a href=\"https://www.kaggle.com/code/awsaf49/gcvit-global-context-vision-transformer#6.-Reusability\" target=\"_blank\">Reusability</a> section if are interested in using this model for some other tasks (classification or feature_extraction) on another dataset.</li></ul></li>\n<li><strong>GitHub Code:</strong> <a href=\"https://github.com/awsaf49/gcvit-tf\" target=\"_blank\">gcvit-tf</a><ul>\n<li>Feel free to send <strong>Pull-Request</strong> if you are interested in contributing to this project.</li></ul></li>\n</ul>\n<h2>Features of <a href=\"library\" target=\"_blank\">gcvit-tf</a>:</h2>\n<ul>\n<li>This library loads <strong>ImageNet</strong> weights from the official repo.</li>\n<li>Also, it has <code>timm</code> like features such as<code>forward_features</code>, <code>forward_head</code>, and <code>reset_classifier</code> which might come in handy.</li>\n<li>It can be used in both <strong>GPU</strong> and <strong>TPU</strong>.</li>\n</ul>\n<h2>Supported Models</h2>\n<p>The official codebase had some issue which has been fixed recently (27 July 2022). Here's the result of ported weights on <strong>ImageNetV2-Test</strong> data,</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Acc@1</th>\n<th>Acc@5</th>\n<th>#Params</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>GCViT-XXTiny</td>\n<td>66</td>\n<td>87</td>\n<td>12M</td>\n</tr>\n<tr>\n<td>GCViT-XTiny</td>\n<td>69</td>\n<td>88</td>\n<td>20M</td>\n</tr>\n<tr>\n<td>GCViT-Tiny</td>\n<td>71</td>\n<td>90</td>\n<td>28M</td>\n</tr>\n<tr>\n<td>GCViT-Small</td>\n<td>72</td>\n<td>90</td>\n<td>51M</td>\n</tr>\n<tr>\n<td>GCViT-Base</td>\n<td>73</td>\n<td>91</td>\n<td>90M</td>\n</tr>\n</tbody>\n</table>\n<h2>Usage</h2>\n<p>Install Library</p>\n<pre><code>pip install gcvit\n</code></pre>\n<p>Load model using the following codes,</p>\n<pre><code>from gcvit import GCViTTiny\nmodel = GCViTTiny(pretrain=True)\n</code></pre>\n<p>For feature extraction:</p>\n<pre><code>model.reset_classifier(num_classes=0, head_act=None)\nfeature = model(img)\nprint(feature.shape)\n</code></pre>\n<p>Feature:</p>\n<pre><code>(None, 512)\n</code></pre>\n<p>For feature map:</p>\n<pre><code>feature = model.forward_features(img)\nprint(feature.shape)\n</code></pre>\n<p>Feature map:</p>\n<pre><code>(None, 7, 7, 512)\n</code></pre>",
  "messages": [
    {
      "id": 1898717,
      "postDate": "2022-08-14T18:50:34.027Z",
      "content": "<p><img src=\"https://raw.githubusercontent.com/awsaf49/gcvit-tf/main/image/lvg_arch.PNG\"></p>\n<p>Hello everyone,<br>\nAs you may already know, NVIDIA has recently published (20 June 2022) its paper, <strong>GCViT: Global Context Vision Transformer</strong> which outperforms <strong>ConvNeXt</strong> and <strong>SwinTransformer</strong> . As this competition is a Classification Problem (give or take), I think <strong>GCViT</strong> can be useful here. You are welcome to explore this model for this competition.</p>\n<blockquote>\n  <p>There was some issue with their weights and they recently released their <strong>final version of ImageNet pretrained weights</strong> which brought over <code>~1%</code> improvement over older weights.</p>\n</blockquote>\n<p>To ease things up, I've implemented this model using <strong>TensorFlow</strong> and created an open-source library <a href=\"https://github.com/awsaf49/gcvit-tf\" target=\"_blank\">gcvit-tf</a>. I've also made a notebook to help you get started with <strong>GCViT</strong>,. </p>\n<h2>Resources:</h2>\n<ul>\n<li><strong>Notebook:</strong> <a href=\"https://www.kaggle.com/code/awsaf49/gcvit-global-context-vision-transformer\" target=\"_blank\">GCViT: Global Context Vision Transformer</a><ul>\n<li>Refer to <a href=\"https://www.kaggle.com/code/awsaf49/gcvit-global-context-vision-transformer#6.-Reusability\" target=\"_blank\">Reusability</a> section if are interested in using this model for some other tasks (classification or feature_extraction) on another dataset.</li></ul></li>\n<li><strong>GitHub Code:</strong> <a href=\"https://github.com/awsaf49/gcvit-tf\" target=\"_blank\">gcvit-tf</a><ul>\n<li>Feel free to send <strong>Pull-Request</strong> if you are interested in contributing to this project.</li></ul></li>\n</ul>\n<h2>Features of <a href=\"library\" target=\"_blank\">gcvit-tf</a>:</h2>\n<ul>\n<li>This library loads <strong>ImageNet</strong> weights from the official repo.</li>\n<li>Also, it has <code>timm</code> like features such as<code>forward_features</code>, <code>forward_head</code>, and <code>reset_classifier</code> which might come in handy.</li>\n<li>It can be used in both <strong>GPU</strong> and <strong>TPU</strong>.</li>\n</ul>\n<h2>Supported Models</h2>\n<p>The official codebase had some issue which has been fixed recently (27 July 2022). Here's the result of ported weights on <strong>ImageNetV2-Test</strong> data,</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Acc@1</th>\n<th>Acc@5</th>\n<th>#Params</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>GCViT-XXTiny</td>\n<td>66</td>\n<td>87</td>\n<td>12M</td>\n</tr>\n<tr>\n<td>GCViT-XTiny</td>\n<td>69</td>\n<td>88</td>\n<td>20M</td>\n</tr>\n<tr>\n<td>GCViT-Tiny</td>\n<td>71</td>\n<td>90</td>\n<td>28M</td>\n</tr>\n<tr>\n<td>GCViT-Small</td>\n<td>72</td>\n<td>90</td>\n<td>51M</td>\n</tr>\n<tr>\n<td>GCViT-Base</td>\n<td>73</td>\n<td>91</td>\n<td>90M</td>\n</tr>\n</tbody>\n</table>\n<h2>Usage</h2>\n<p>Install Library</p>\n<pre><code>pip install gcvit\n</code></pre>\n<p>Load model using the following codes,</p>\n<pre><code>from gcvit import GCViTTiny\nmodel = GCViTTiny(pretrain=True)\n</code></pre>\n<p>For feature extraction:</p>\n<pre><code>model.reset_classifier(num_classes=0, head_act=None)\nfeature = model(img)\nprint(feature.shape)\n</code></pre>\n<p>Feature:</p>\n<pre><code>(None, 512)\n</code></pre>\n<p>For feature map:</p>\n<pre><code>feature = model.forward_features(img)\nprint(feature.shape)\n</code></pre>\n<p>Feature map:</p>\n<pre><code>(None, 7, 7, 512)\n</code></pre>",
      "rawMarkdown": "<img src=\"https://raw.githubusercontent.com/awsaf49/gcvit-tf/main/image/lvg_arch.PNG\" width=800>\n\nHello everyone,\nAs you may already know, NVIDIA has recently published (20 June 2022) its paper, **GCViT: Global Context Vision Transformer** which outperforms **ConvNeXt** and **SwinTransformer** . As this competition is a Classification Problem (give or take), I think **GCViT** can be useful here. You are welcome to explore this model for this competition.\n\n> There was some issue with their weights and they recently released their **final version of ImageNet pretrained weights** which brought over `~1%` improvement over older weights.\n\nTo ease things up, I've implemented this model using **TensorFlow** and created an open-source library [gcvit-tf](https://github.com/awsaf49/gcvit-tf). I've also made a notebook to help you get started with **GCViT**,. \n\n## Resources:\n* **Notebook:** [GCViT: Global Context Vision Transformer](https://www.kaggle.com/code/awsaf49/gcvit-global-context-vision-transformer)\n    * Refer to [Reusability](https://www.kaggle.com/code/awsaf49/gcvit-global-context-vision-transformer#6.-Reusability) section if are interested in using this model for some other tasks (classification or feature_extraction) on another dataset.\n* **GitHub Code:** [gcvit-tf](https://github.com/awsaf49/gcvit-tf)\n    * Feel free to send **Pull-Request** if you are interested in contributing to this project.\n\n## Features of [gcvit-tf](library):\n* This library loads **ImageNet** weights from the official repo.\n* Also, it has `timm` like features such as`forward_features`, `forward_head`, and `reset_classifier` which might come in handy.\n* It can be used in both **GPU** and **TPU**.\n\n## Supported Models\nThe official codebase had some issue which has been fixed recently (27 July 2022). Here's the result of ported weights on **ImageNetV2-Test** data,\n\n| Model        | Acc@1 | Acc@5 | #Params |\n|--------------|-------|-------|---------|\n| GCViT-XXTiny | 66    | 87    | 12M     |\n| GCViT-XTiny  | 69    | 88    | 20M     |\n| GCViT-Tiny   | 71    | 90    | 28M     |\n| GCViT-Small  | 72    | 90    | 51M     |\n| GCViT-Base   | 73    | 91    | 90M     |\n\n## Usage\nInstall Library\n```shell\npip install gcvit\n```\nLoad model using the following codes,\n```py\nfrom gcvit import GCViTTiny\nmodel = GCViTTiny(pretrain=True)\n```\n\nFor feature extraction:\n```py\nmodel.reset_classifier(num_classes=0, head_act=None)\nfeature = model(img)\nprint(feature.shape)\n```\nFeature:\n```py\n(None, 512)\n```\nFor feature map:\n```py\nfeature = model.forward_features(img)\nprint(feature.shape)\n```\nFeature map:\n```py\n(None, 7, 7, 512)\n```\n\n",
      "votes": 13
    },
    {
      "id": 1904577,
      "postDate": "2022-08-18T09:56:29.127Z",
      "content": "<p>Repeat. Same discussion thread, at a same on going competition.</p>\n<p>4 days ago and 22 days ago</p>\n<p><img src=\"https://user-images.githubusercontent.com/45315076/185367129-4e4eaf21-59e3-4b7e-a91f-786de2144b30.png\" alt=\"image\"></p>\n<p><img src=\"https://user-images.githubusercontent.com/45315076/185367451-99916fc8-8bc6-4ffd-91f0-45bcad57a752.png\" alt=\"image\"></p>",
      "rawMarkdown": "Repeat. Same discussion thread, at a same on going competition.\n\n4 days ago and 22 days ago\n\n![image](https://user-images.githubusercontent.com/45315076/185367129-4e4eaf21-59e3-4b7e-a91f-786de2144b30.png)\n\n![image](https://user-images.githubusercontent.com/45315076/185367451-99916fc8-8bc6-4ffd-91f0-45bcad57a752.png)",
      "replies": [
        {
          "id": 1908134,
          "postDate": "2022-08-21T11:46:04.800Z",
          "content": "<p>The title same but the purpose is different. I've mentioned the purpose in this post,</p>\n<blockquote>\n  <p>There was some issue with their weights and they recently released their final version of ImageNet pretrained weights which brought over ~1% improvement over older weights.</p>\n</blockquote>",
          "rawMarkdown": "The title same but the purpose is different. I've mentioned the purpose in this post,\n> There was some issue with their weights and they recently released their final version of ImageNet pretrained weights which brought over ~1% improvement over older weights."
        },
        {
          "id": 1908140,
          "postDate": "2022-08-21T11:50:07.857Z",
          "content": "<p>Another purpose is to make people in this competition aware of <strong>GCViT</strong>. People often publish similar publish similar notebook just changing the dataset for a new competition so I don't see any problem here.</p>",
          "rawMarkdown": "Another purpose is to make people in this competition aware of **GCViT**. People often publish similar publish similar notebook just changing the dataset for a new competition so I don't see any problem here."
        }
      ]
    },
    {
      "id": 1900952,
      "postDate": "2022-08-16T11:28:33.240Z",
      "content": "<p>It is incredible that each time a new paper comes out we can just find it on your github. <br>\n<strong>Your work is impressive as hell.</strong></p>\n<p>As for the model itself: Those performance plots of the models are just SICK! They outperform everyone.</p>\n<p>The Devastator.</p>",
      "rawMarkdown": "It is incredible that each time a new paper comes out we can just find it on your github. \n**Your work is impressive as hell.**\n\nAs for the model itself: Those performance plots of the models are just SICK! They outperform everyone.\n\n\nThe Devastator.\n"
    }
  ],
  "comments": [
    {
      "id": 1904577,
      "author_name": "Simon Alerdic",
      "author_url": "",
      "post_date": "2022-08-18T09:56:29.127000",
      "content": "<p>Repeat. Same discussion thread, at a same on going competition.</p>\n<p>4 days ago and 22 days ago</p>\n<p><img src=\"https://user-images.githubusercontent.com/45315076/185367129-4e4eaf21-59e3-4b7e-a91f-786de2144b30.png\" alt=\"image\"></p>\n<p><img src=\"https://user-images.githubusercontent.com/45315076/185367451-99916fc8-8bc6-4ffd-91f0-45bcad57a752.png\" alt=\"image\"></p>",
      "votes": 0,
      "replies": [
        {
          "id": 1908134,
          "author_name": "Awsaf",
          "author_url": "",
          "post_date": "2022-08-21T11:46:04.800000",
          "content": "<p>The title same but the purpose is different. I've mentioned the purpose in this post,</p>\n<blockquote>\n  <p>There was some issue with their weights and they recently released their final version of ImageNet pretrained weights which brought over ~1% improvement over older weights.</p>\n</blockquote>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1908140,
          "author_name": "Awsaf",
          "author_url": "",
          "post_date": "2022-08-21T11:50:07.857000",
          "content": "<p>Another purpose is to make people in this competition aware of <strong>GCViT</strong>. People often publish similar publish similar notebook just changing the dataset for a new competition so I don't see any problem here.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1900952,
      "author_name": "The Devastator",
      "author_url": "",
      "post_date": "2022-08-16T11:28:33.240000",
      "content": "<p>It is incredible that each time a new paper comes out we can just find it on your github. <br>\n<strong>Your work is impressive as hell.</strong></p>\n<p>As for the model itself: Those performance plots of the models are just SICK! They outperform everyone.</p>\n<p>The Devastator.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1898717": "<img src=\"https://raw.githubusercontent.com/awsaf49/gcvit-tf/main/image/lvg_arch.PNG\" width=800>\n\nHello everyone,\nAs you may already know, NVIDIA has recently published (20 June 2022) its paper, **GCViT: Global Context Vision Transformer** which outperforms **ConvNeXt** and **SwinTransformer** . As this competition is a Classification Problem (give or take), I think **GCViT** can be useful here. You are welcome to explore this model for this competition.\n\n> There was some issue with their weights and they recently released their **final version of ImageNet pretrained weights** which brought over `~1%` improvement over older weights.\n\nTo ease things up, I've implemented this model using **TensorFlow** and created an open-source library [gcvit-tf](https://github.com/awsaf49/gcvit-tf). I've also made a notebook to help you get started with **GCViT**,. \n\n## Resources:\n* **Notebook:** [GCViT: Global Context Vision Transformer](https://www.kaggle.com/code/awsaf49/gcvit-global-context-vision-transformer)\n    * Refer to [Reusability](https://www.kaggle.com/code/awsaf49/gcvit-global-context-vision-transformer#6.-Reusability) section if are interested in using this model for some other tasks (classification or feature_extraction) on another dataset.\n* **GitHub Code:** [gcvit-tf](https://github.com/awsaf49/gcvit-tf)\n    * Feel free to send **Pull-Request** if you are interested in contributing to this project.\n\n## Features of [gcvit-tf](library):\n* This library loads **ImageNet** weights from the official repo.\n* Also, it has `timm` like features such as`forward_features`, `forward_head`, and `reset_classifier` which might come in handy.\n* It can be used in both **GPU** and **TPU**.\n\n## Supported Models\nThe official codebase had some issue which has been fixed recently (27 July 2022). Here's the result of ported weights on **ImageNetV2-Test** data,\n\n| Model        | Acc@1 | Acc@5 | #Params |\n|--------------|-------|-------|---------|\n| GCViT-XXTiny | 66    | 87    | 12M     |\n| GCViT-XTiny  | 69    | 88    | 20M     |\n| GCViT-Tiny   | 71    | 90    | 28M     |\n| GCViT-Small  | 72    | 90    | 51M     |\n| GCViT-Base   | 73    | 91    | 90M     |\n\n## Usage\nInstall Library\n```shell\npip install gcvit\n```\nLoad model using the following codes,\n```py\nfrom gcvit import GCViTTiny\nmodel = GCViTTiny(pretrain=True)\n```\n\nFor feature extraction:\n```py\nmodel.reset_classifier(num_classes=0, head_act=None)\nfeature = model(img)\nprint(feature.shape)\n```\nFeature:\n```py\n(None, 512)\n```\nFor feature map:\n```py\nfeature = model.forward_features(img)\nprint(feature.shape)\n```\nFeature map:\n```py\n(None, 7, 7, 512)\n```\n\n",
    "1904577": "Repeat. Same discussion thread, at a same on going competition.\n\n4 days ago and 22 days ago\n\n![image](https://user-images.githubusercontent.com/45315076/185367129-4e4eaf21-59e3-4b7e-a91f-786de2144b30.png)\n\n![image](https://user-images.githubusercontent.com/45315076/185367451-99916fc8-8bc6-4ffd-91f0-45bcad57a752.png)",
    "1900952": "It is incredible that each time a new paper comes out we can just find it on your github. \n**Your work is impressive as hell.**\n\nAs for the model itself: Those performance plots of the models are just SICK! They outperform everyone.\n\n\nThe Devastator.\n"
  }
}