{
  "id": 375169,
  "title": "State-of-the-art papers of 2022",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/375169",
  "author_name": "The Devastator",
  "post_date": "2022-12-30T17:29:54.678000",
  "votes": 30,
  "comment_count": 6,
  "views": 0,
  "content": "<h5>EVERYTHING. is. state-of-the-art.</h5>\n<p>Over the past year of computer vision, we have seen everything!</p>\n<ul>\n<li><a href=\"https://arxiv.org/abs/2204.07118\" target=\"_blank\">Transformers?</a> State-of-the-art!</li>\n<li><a href=\"https://arxiv.org/abs/2201.03545\" target=\"_blank\">Convolutional Networks?</a> State-of-the-art!</li>\n<li><a href=\"https://arxiv.org/pdf/2110.00476.pdf\" target=\"_blank\">ResNet-50?</a> - State-of-the-art!</li>\n<li><a href=\"https://arxiv.org/abs/2105.01601\" target=\"_blank\">Freaking MLP?!?!!</a> - State-of-the-art!</li>\n</ul>\n<p>Everything is state-of-the-art. <br>\nIt doesn't matter what it is!<br>\nIf it relates to computer vision, it has to be the best of the best! \"Everything is the current State-of-the-Art\" - and that's a fact!</p>\n<p>In this short summary, I go through some of the notable papers claiming state-of-the-art performance on 2022. I attempt to explain each paper's main contributions of each paper while providing links to each paper's pretrained weights on <code>timm</code>.</p>\n<p>Enjoy!</p>\n<hr>\n<h4>MaxViT - Google</h4>\n<p><strong>Paper:</strong> <a href=\"https://arxiv.org/abs/2204.01697\" target=\"_blank\">MaxViT: Multi-Axis Vision Transformer</a><br>\n<strong>TIMM:</strong> <a href=\"https://github.com/rwightman/pytorch-image-models/blob/main/timm/models/maxxvit.py\" target=\"_blank\">source</a></p>\n<p>Google presents a new multi-axis approach that improves on the original ViT and MLP models, can better adapt to high-resolution, dense prediction tasks, and can naturally adapt to different input sizes with high flexibility and low complexity. </p>\n<p>The new approach is based on multi-axis attention, which decomposes the full-size attention (each pixel attends to all the pixels) used in ViT into two sparse forms — local and (sparse) global. </p>\n<p>The multi-axis attention contains a sequential stack of block attention and grid attention. The block attention works within non-overlapping windows (small patches in intermediate feature maps) to capture local patterns, while the grid attention works on a sparsely sampled uniform grid for long-range (global) interactions. The window sizes of grid and block attentions can be fully controlled as hyperparameters to ensure a linear computational complexity to the input size.<br>\nThe proposed multi-axis attention conducts blocked local and dilated global attention sequentially followed by a FFN, with only a linear complexity. The pixels in the same colors are attended together.</p>\n<p><img src=\"https://i.ibb.co/CvFWcP4/image4-2.png\" alt=\"\"></p>\n<p>Such low-complexity attention can significantly improve its wide applicability to many vision tasks, especially for high-resolution visual predictions, demonstrating greater generality than the original attention used in ViT.</p>\n<p>The MaxViT is built by concatenating MBConv (EfficientNet) with multi-axis attention. This single block can encode local and global visual information regardless of input resolution. We then simply stack repeated blocks composed of attention and convolutions in a hierarchical architecture (ResNet, CoAtNet, etc), </p>\n<p><img src=\"https://i.ibb.co/Vv815YM/image6.png\" alt=\"\"></p>\n<p><strong>MaxViT is distinguished from previous hierarchical approaches as it can “see” globally throughout the entire network, even in earlier, high-resolution stages, demonstrating stronger model capacity on various tasks.</strong></p>\n<p><strong>The \"we are the state-of-the-art\" claim:</strong></p>\n<p><img src=\"https://i.ibb.co/6BWNdk5/maxvit.png\" alt=\"\"></p>\n<blockquote>\n  <p><strong>TiMMs full list of pretrained weights</strong></p>\n  <ul>\n  <li><p><code>maxvit_pico_rw_256</code></p></li>\n  <li><p><code>maxvit_nano_rw_256</code></p></li>\n  <li><p><code>maxvit_tiny_rw_224</code></p></li>\n  <li><p><code>maxvit_tiny_rw_256</code></p></li>\n  <li><p><code>maxvit_rmlp_pico_rw_256</code></p></li>\n  <li><p><code>maxvit_rmlp_nano_rw_256</code></p></li>\n  <li><p><code>maxvit_rmlp_tiny_rw_256</code></p></li>\n  <li><p><code>maxvit_rmlp_small_rw_224</code></p></li>\n  <li><p><code>maxvit_rmlp_small_rw_256</code></p></li>\n  <li><p><code>maxvit_tiny_pm_256</code></p></li>\n  <li><p><code>maxxvit_rmlp_nano_rw_256</code></p></li>\n  <li><p><code>maxxvit_rmlp_tiny_rw_256</code></p></li>\n  <li><p><code>maxxvit_rmlp_small_rw_256</code></p></li>\n  <li><p><code>maxxvit_rmlp_base_rw_224</code></p></li>\n  <li><p><code>maxxvit_rmlp_large_rw_224</code></p></li>\n  <li><p><code>maxvit_tiny_tf_224.in1k</code></p></li>\n  <li><p><code>maxvit_tiny_tf_384.in1k</code></p></li>\n  <li><p><code>maxvit_tiny_tf_512.in1k</code></p></li>\n  <li><p><code>maxvit_small_tf_224.in1k</code></p></li>\n  <li><p><code>maxvit_small_tf_384.in1k</code></p></li>\n  <li><p><code>maxvit_small_tf_512.in1k</code></p></li>\n  <li><p><code>maxvit_base_tf_224.in1k</code></p></li>\n  <li><p><code>maxvit_base_tf_384.in1k</code></p></li>\n  <li><p><code>maxvit_base_tf_512.in1k</code></p></li>\n  <li><p><code>maxvit_large_tf_224.in1k</code></p></li>\n  <li><p><code>maxvit_large_tf_384.in1k</code></p></li>\n  <li><p><code>maxvit_large_tf_512.in1k</code></p></li>\n  <li><p><code>maxvit_base_tf_224.in21k</code></p></li>\n  <li><p><code>maxvit_base_tf_384.in21k_ft_in1k</code></p></li>\n  <li><p><code>maxvit_base_tf_512.in21k_ft_in1k</code></p></li>\n  <li><p><code>maxvit_large_tf_224.in21k</code></p></li>\n  <li><p><code>maxvit_large_tf_384.in21k_ft_in1k</code></p></li>\n  <li><p><code>maxvit_large_tf_512.in21k_ft_in1k</code></p></li>\n  <li><p><code>maxvit_xlarge_tf_224.in21k</code></p></li>\n  <li><p><code>maxvit_xlarge_tf_384.in21k_ft_in1k</code></p></li>\n  <li><p><code>maxvit_xlarge_tf_512.in21k_ft_in1k</code></p></li>\n  </ul>\n</blockquote>\n<hr>\n<h4>CoAtNet - Google</h4>\n<p><strong>Paper:</strong> <a href=\"https://arxiv.org/abs/2106.04803\" target=\"_blank\">CoAtNet: Marrying Convolution and Attention for All Data Sizes</a><br>\n<strong>TIMM:</strong> <a href=\"https://github.com/rwightman/pytorch-image-models/blob/6902c48a5f0637c8155c1c4bc10ad35930f3e772/timm/models/maxxvit.py#L1927\" target=\"_blank\">source</a><br>\n<strong>Kaggle Notebook:</strong> <a href=\"https://www.kaggle.com/code/thedevastator/train-infer-coatnet-efficientnet\" target=\"_blank\">code</a></p>\n<p>The CoAtNet paper attempts to effectively combine the strengths from both convolutional and transformers architectures, they present CoAtNets(pronounced \"coat\" nets), a family of hybrid models built from two key insights:</p>\n<ul>\n<li>Depthwise Convolution and self-Attention can be naturally unified via simple relative attention</li>\n<li>Vertically stacking convolution layers and attention layers in a principled way is surprisingly effective in improving generalization, capacity and efficiency.</li>\n</ul>\n<p><img src=\"https://i.ibb.co/Sd6wj7D/Selection-998.png\" alt=\"\"></p>\n<p><strong>The \"we are the state-of-the-art\" claim:</strong></p>\n<p><img src=\"https://i.ibb.co/djDn4Y8/coatnet.png\" alt=\"\"></p>\n<blockquote>\n  <p><strong>TiMMs full list of pretrained weights</strong></p>\n  <ul>\n  <li><code>coatnet_pico_rw_224</code></li>\n  <li><code>coatnet_nano_rw_224</code></li>\n  <li><code>coatnet_0_rw_224</code></li>\n  <li><code>coatnet_1_rw_224</code></li>\n  <li><code>coatnet_2_rw_224</code></li>\n  <li><code>coatnet_3_rw_224</code></li>\n  <li><code>coatnet_bn_0_rw_224</code></li>\n  <li><code>coatnet_rmlp_nano_rw_224</code></li>\n  <li><code>coatnet_rmlp_0_rw_224</code></li>\n  <li><code>coatnet_rmlp_1_rw_224</code></li>\n  <li><code>coatnet_rmlp_1_rw2_224</code></li>\n  <li><code>coatnet_rmlp_2_rw_224</code></li>\n  <li><code>coatnet_rmlp_3_rw_224</code></li>\n  <li><code>coatnet_nano_cc_224</code></li>\n  <li><code>coatnext_nano_rw_224</code></li>\n  <li><code>coatnet_0_224</code></li>\n  <li><code>coatnet_1_224</code></li>\n  <li><code>coatnet_2_224</code></li>\n  <li><code>coatnet_3_224</code></li>\n  <li><code>coatnet_4_224</code></li>\n  <li><code>coatnet_5_224</code></li>\n  </ul>\n</blockquote>\n<hr>\n<h4>DeiT III - Meta (Facebook)</h4>\n<p><strong>Paper:</strong> <a href=\"https://arxiv.org/abs/2204.07118\" target=\"_blank\">DeiT III: Revenge of the ViT</a><br>\n<strong>TIMM:</strong> <a href=\"https://github.com/rwightman/pytorch-image-models/blob/6902c48a5f0637c8155c1c4bc10ad35930f3e772/timm/models/deit.py#L65\" target=\"_blank\">source</a><br>\n<strong>Note:</strong> Used on a <a href=\"https://www.kaggle.com/competitions/herbarium-2022-fgvc9/discussion/329299\" target=\"_blank\">winning solution</a> on Herbarium 2022 - FGVC9 </p>\n<p>In this work, the authors used several techniques to improve the training of large vision transformers (ViT) on the ImageNet dataset. These techniques include:</p>\n<ul>\n<li><strong>Stochastic depth</strong>, a regularization method that is especially useful for training deep networks.</li>\n<li><strong>LayerScale</strong>, a method introduced to facilitate the convergence of deep transformers.</li>\n<li><strong>Binary cross entropy loss</strong> (In contrast to categorical cross entropy) for Imagenet1k training, which was found to provide a significant improvement in performance for larger ViTs.</li>\n</ul>\n<p>They also found that a simple data augmentation method called 3-Augment worked better than other methods, and that using <strong>simple random cropping</strong> was more effective than random resize cropping when pre-training on a larger dataset like ImageNet-21k. They observed that using a lower resolution at training time had a regularizing effect and helped prevent overfitting, especially for the largest models.</p>\n<p>Additionally, they found that using the binary cross entropy loss provided a significant improvement in performance for larger ViTs trained on ImageNet-1k, but that using the cross entropy loss was more effective when pre-training with ImageNet-21k or for fine-tuning.</p>\n<p><strong>The \"we are the state-of-the-art\" claim:</strong></p>\n<p><img src=\"https://i.ibb.co/6Nqkqt9/deit.png\" alt=\"\"></p>\n<blockquote>\n  <p><strong>TiMMs full list of pretrained weights</strong></p>\n  <ul>\n  <li><p><code>deit3_small_patch16_224</code></p></li>\n  <li><p><code>deit3_small_patch16_384</code></p></li>\n  <li><p><code>deit3_medium_patch16_224</code></p></li>\n  <li><p><code>deit3_base_patch16_224</code></p></li>\n  <li><p><code>deit3_base_patch16_384</code></p></li>\n  <li><p><code>deit3_large_patch16_224</code></p></li>\n  <li><p><code>deit3_large_patch16_384</code></p></li>\n  <li><p><code>deit3_huge_patch14_224</code></p></li>\n  <li><p><code>deit3_small_patch16_224_in21ft1k</code></p></li>\n  <li><p><code>deit3_small_patch16_384_in21ft1k</code></p></li>\n  <li><p><code>deit3_medium_patch16_224_in21ft1k</code></p></li>\n  <li><p><code>deit3_base_patch16_224_in21ft1k</code></p></li>\n  <li><p><code>deit3_base_patch16_384_in21ft1k</code></p></li>\n  <li><p><code>deit3_large_patch16_224_in21ft1k</code></p></li>\n  <li><p><code>deit3_large_patch16_384_in21ft1k</code></p></li>\n  <li><p><code>deit3_huge_patch14_224_in21ft1k</code></p></li>\n  </ul>\n</blockquote>\n<hr>\n<h4>FlexiViT - Google</h4>\n<p><strong>Paper:</strong> <a href=\"https://arxiv.org/abs/2212.08013\" target=\"_blank\">FlexiViT: One Model for All Patch Sizes</a><br>\n<strong>TIMM:</strong> <a href=\"https://github.com/rwightman/pytorch-image-models/blob/6902c48a5f0637c8155c1c4bc10ad35930f3e772/timm/models/vision_transformer.py#L1579\" target=\"_blank\">source</a></p>\n<p>\"Simply randomizing the patch size at training time leads to a single set of weights that performs well across a wide range of patch sizes, making it possible to tailor the model to different compute budgets at deployment time.\"</p>\n<p><img src=\"https://i.ibb.co/MN2n0QH/Selection-1118.png\" alt=\"\"></p>\n<p><strong>The \"we are the state-of-the-art\" claim:</strong></p>\n<p><img src=\"https://i.ibb.co/wR9YC4H/flexivit.png\" alt=\"\"></p>\n<blockquote>\n  <p><strong>TiMMs full list of pretrained weights</strong></p>\n  <ul>\n  <li><code>flexivit_small.1200ep_in1k</code></li>\n  <li><code>flexivit_small.600ep_in1k</code></li>\n  <li><code>flexivit_small.300ep_in1k</code></li>\n  <li><code>flexivit_base.1200ep_in1k</code></li>\n  <li><code>flexivit_base.600ep_in1k</code></li>\n  <li><code>flexivit_base.300ep_in1k</code></li>\n  <li><code>flexivit_base.1000ep_in21k</code></li>\n  <li><code>flexivit_base.300ep_in21k</code></li>\n  <li><code>flexivit_large.1200ep_in1k</code></li>\n  <li><code>flexivit_large.600ep_in1k</code></li>\n  <li><code>flexivit_large.300ep_in1k</code></li>\n  <li><code>flexivit_base.patch16_in21k</code></li>\n  <li><code>flexivit_base.patch30_in21k</code></li>\n  </ul>\n</blockquote>\n<hr>\n<h4>GC ViT</h4>\n<p><strong>Paper:</strong> <a href=\"https://arxiv.org/pdf/2206.09959.pdf\" target=\"_blank\">Global Context Vision Transformers</a><br>\n<strong>TIMM:</strong> <a href=\"https://github.com/rwightman/pytorch-image-models/blob/6902c48a5f0637c8155c1c4bc10ad35930f3e772/timm/models/gcvit.py#L564\" target=\"_blank\">source</a><br>\n<strong>Kaggle Notebook:</strong> <a href=\"https://www.kaggle.com/code/awsaf49/gcvit-global-context-vision-transformer\" target=\"_blank\">code</a></p>\n<blockquote>\n  <p><strong>Credit:</strong> The summary is based on the <a href=\"https://www.kaggle.com/code/awsaf49/gcvit-global-context-vision-transformer\" target=\"_blank\">amazing notebook</a> by <a href=\"https://www.kaggle.com/awsaf49\" target=\"_blank\">Awsaf</a></p>\n</blockquote>\n<p>The paper proposes a novel architecture namely, <strong>Global Context Vision Transformer (GCViT)</strong> that utilizes window attention mechanism similar to <strong>Swin Transformer</strong>.<br>\nUnlike <strong>Swin Transformer</strong> this paper uses <strong>global context self-attention</strong>, with local self-attention, rather than <strong>shifted window self-attention</strong>, to model both long and short-range dependencies. Even though <strong>global-window-attention</strong> is a window-attention but it takes leverage of <strong>global query</strong> which contains global information hence captures long-range information. This paper compensates for the lack of the <strong>inductive bias</strong> that exists in both ViTs and Swin Transformer by utilizing a <strong>CNN</strong> based module. </p>\n<p><strong>GCViT achieves state-of-the-art results across image classification, object detection and semantic segmentation tasks.</strong></p>\n<blockquote>\n  <p><strong>TL;DR:</strong> Global Context ViT (<strong>GCViT</strong>) is a hierarchical architecture like <strong>Swin Transformer</strong> but utilizes<code>global-window-attention</code> instead of <code>shifted-window-attention</code> for effectively capturing long-range information. It also introduces <strong>CNN</strong> based module to include <strong>inductive-bias</strong> a useful feature for image that has been missing in both <strong>ViT</strong> and <strong>Swin Transformer</strong>.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/awsaf49\" target=\"_blank\">Awsaf</a> annotated the architecture figure to make it easier to digest:<br>\n<img src=\"https://raw.githubusercontent.com/awsaf49/gcvit-tf/main/image/arch_annot.png\" alt=\"\"></p>\n<ul>\n<li><p><code>SE</code>: <strong>Squeeze-Excitation (SE)</strong> aka <strong>Bottleneck</strong> module acts sd kind of <strong>channel attention</strong>. It consits of <strong>AvgPooling</strong>, <strong>Dense/FullyConnected (FC)/Linear</strong> , <strong>GELU</strong> and <strong>Sigmoid</strong> module.<br>\n<img src=\"https://raw.githubusercontent.com/awsaf49/gcvit-tf/main/image/se_annot.png\"></p></li>\n<li><p><code>Fused-MBConv:</code> This is similar to the one used in <strong>EfficientNetV2</strong>. It uses <strong>Depthwise-Conv</strong>, <strong>GELU</strong>, <strong>SE</strong>, <strong>Conv</strong>, to extract feature with a resiudal connection. Note that, no new module is declared for this one, we simply applied corresponding modules directly.<br>\n<img src=\"https://raw.githubusercontent.com/awsaf49/gcvit-tf/main/image/fmb_annot.png\"></p></li>\n<li><p><code>ReduceSize</code>: It is a <strong>CNN</strong> based <strong>downsample</strong> module which abvobe mentioned <code>Fused-MBConv</code> module to extract feature, <strong>Strided Conv</strong> to simultaneously reduce spatial dimension and increse channelwise dimention of the features and finally <strong>LayerNormalization</strong> module to normalize features. In the paper/figure this module is referred as <strong>downsample</strong> module. I think it is mention worthy that <strong>SwniTransformer</strong> used <code>PatchMerging</code> module instead of <code>ReduceSize</code> to reduce the spatial dimention and increase channelwise dimension which uses <strong>fully-connected/dense/linear</strong> module. According to the <strong>GCViT</strong> paper, one of the purposes of using <code>ReduceSize</code> is to add inductive bias through <strong>CNN</strong> module.<br>\n<img src=\"https://raw.githubusercontent.com/awsaf49/gcvit-tf/main/image/down_annot.png\"></p></li>\n</ul>\n<p><strong>The \"we are the state-of-the-art\" claim:</strong></p>\n<p><img src=\"https://i.ibb.co/thd6Gcn/gcvit.png\" alt=\"\"></p>\n<blockquote>\n  <p>**TIMM's full list of pretrained </p>\n  <ul>\n  <li><code>gcvit_xxtiny</code></li>\n  <li><code>gcvit_xtiny</code></li>\n  <li><code>gcvit_tiny</code></li>\n  <li><code>gcvit_small</code></li>\n  <li><code>gcvit_base</code></li>\n  </ul>\n</blockquote>\n<hr>\n<h4>EVA-L</h4>\n<p><strong>Paper:</strong> <a href=\"https://arxiv.org/abs/2211.07636\" target=\"_blank\">EVA: An Open Billion-Scale Vision Foundation Model</a><br>\n<strong>TIMM:</strong> <a href=\"https://github.com/rwightman/pytorch-image-models/blob/6902c48a5f0637c8155c1c4bc10ad35930f3e772/timm/models/vision_transformer.py#L1579\" target=\"_blank\">source</a></p>\n<p>EVA is a vanilla ViT pre-trained to reconstruct the masked out image-text aligned vision features (i.e., CLIP features) conditioned on visible image patches. Via this pretext task, the authors can efficiently scale up EVA to one billion parameters, and sets new records on a broad range of representative vision downstream tasks.</p>\n<p><strong>The \"we are the state-of-the-art\" claim:</strong></p>\n<ul>\n<li><strong>EVA-L is the best ViT-L (304M) to date</strong> that can reach up to 89.2 top-1 acc on IN-1K (weights &amp; logs) by leveraging vision features from EVA-CLIP.</li>\n</ul>\n<blockquote>\n  <p><strong>TiMMs full list of pretrained weights</strong></p>\n  <ul>\n  <li><p><code>eva_large_patch14_336.in22k_ft_in22k_in1k</code></p></li>\n  <li><p><code>eva_large_patch14_336.in22k_ft_in1k</code></p></li>\n  <li><p><code>eva_large_patch14_196.in22k_ft_in22k_in1k</code></p></li>\n  <li><p><code>eva_large_patch14_196.in22k_ft_in1k</code></p></li>\n  <li><p><code>eva_giant_patch14_560.m30m_ft_in22k_in1k</code></p></li>\n  <li><p><code>eva_giant_patch14_336.m30m_ft_in22k_in1k</code></p></li>\n  <li><p><code>eva_giant_patch14_336.clip_ft_in1k</code></p></li>\n  <li><p><code>eva_giant_patch14_224.clip_ft_in1k</code></p></li>\n  </ul>\n</blockquote>",
  "messages": [
    {
      "id": 2080985,
      "postDate": "2022-12-30T17:29:54.680Z",
      "content": "<h5>EVERYTHING. is. state-of-the-art.</h5>\n<p>Over the past year of computer vision, we have seen everything!</p>\n<ul>\n<li><a href=\"https://arxiv.org/abs/2204.07118\" target=\"_blank\">Transformers?</a> State-of-the-art!</li>\n<li><a href=\"https://arxiv.org/abs/2201.03545\" target=\"_blank\">Convolutional Networks?</a> State-of-the-art!</li>\n<li><a href=\"https://arxiv.org/pdf/2110.00476.pdf\" target=\"_blank\">ResNet-50?</a> - State-of-the-art!</li>\n<li><a href=\"https://arxiv.org/abs/2105.01601\" target=\"_blank\">Freaking MLP?!?!!</a> - State-of-the-art!</li>\n</ul>\n<p>Everything is state-of-the-art. <br>\nIt doesn't matter what it is!<br>\nIf it relates to computer vision, it has to be the best of the best! \"Everything is the current State-of-the-Art\" - and that's a fact!</p>\n<p>In this short summary, I go through some of the notable papers claiming state-of-the-art performance on 2022. I attempt to explain each paper's main contributions of each paper while providing links to each paper's pretrained weights on <code>timm</code>.</p>\n<p>Enjoy!</p>\n<hr>\n<h4>MaxViT - Google</h4>\n<p><strong>Paper:</strong> <a href=\"https://arxiv.org/abs/2204.01697\" target=\"_blank\">MaxViT: Multi-Axis Vision Transformer</a><br>\n<strong>TIMM:</strong> <a href=\"https://github.com/rwightman/pytorch-image-models/blob/main/timm/models/maxxvit.py\" target=\"_blank\">source</a></p>\n<p>Google presents a new multi-axis approach that improves on the original ViT and MLP models, can better adapt to high-resolution, dense prediction tasks, and can naturally adapt to different input sizes with high flexibility and low complexity. </p>\n<p>The new approach is based on multi-axis attention, which decomposes the full-size attention (each pixel attends to all the pixels) used in ViT into two sparse forms — local and (sparse) global. </p>\n<p>The multi-axis attention contains a sequential stack of block attention and grid attention. The block attention works within non-overlapping windows (small patches in intermediate feature maps) to capture local patterns, while the grid attention works on a sparsely sampled uniform grid for long-range (global) interactions. The window sizes of grid and block attentions can be fully controlled as hyperparameters to ensure a linear computational complexity to the input size.<br>\nThe proposed multi-axis attention conducts blocked local and dilated global attention sequentially followed by a FFN, with only a linear complexity. The pixels in the same colors are attended together.</p>\n<p><img src=\"https://i.ibb.co/CvFWcP4/image4-2.png\" alt=\"\"></p>\n<p>Such low-complexity attention can significantly improve its wide applicability to many vision tasks, especially for high-resolution visual predictions, demonstrating greater generality than the original attention used in ViT.</p>\n<p>The MaxViT is built by concatenating MBConv (EfficientNet) with multi-axis attention. This single block can encode local and global visual information regardless of input resolution. We then simply stack repeated blocks composed of attention and convolutions in a hierarchical architecture (ResNet, CoAtNet, etc), </p>\n<p><img src=\"https://i.ibb.co/Vv815YM/image6.png\" alt=\"\"></p>\n<p><strong>MaxViT is distinguished from previous hierarchical approaches as it can “see” globally throughout the entire network, even in earlier, high-resolution stages, demonstrating stronger model capacity on various tasks.</strong></p>\n<p><strong>The \"we are the state-of-the-art\" claim:</strong></p>\n<p><img src=\"https://i.ibb.co/6BWNdk5/maxvit.png\" alt=\"\"></p>\n<blockquote>\n  <p><strong>TiMMs full list of pretrained weights</strong></p>\n  <ul>\n  <li><p><code>maxvit_pico_rw_256</code></p></li>\n  <li><p><code>maxvit_nano_rw_256</code></p></li>\n  <li><p><code>maxvit_tiny_rw_224</code></p></li>\n  <li><p><code>maxvit_tiny_rw_256</code></p></li>\n  <li><p><code>maxvit_rmlp_pico_rw_256</code></p></li>\n  <li><p><code>maxvit_rmlp_nano_rw_256</code></p></li>\n  <li><p><code>maxvit_rmlp_tiny_rw_256</code></p></li>\n  <li><p><code>maxvit_rmlp_small_rw_224</code></p></li>\n  <li><p><code>maxvit_rmlp_small_rw_256</code></p></li>\n  <li><p><code>maxvit_tiny_pm_256</code></p></li>\n  <li><p><code>maxxvit_rmlp_nano_rw_256</code></p></li>\n  <li><p><code>maxxvit_rmlp_tiny_rw_256</code></p></li>\n  <li><p><code>maxxvit_rmlp_small_rw_256</code></p></li>\n  <li><p><code>maxxvit_rmlp_base_rw_224</code></p></li>\n  <li><p><code>maxxvit_rmlp_large_rw_224</code></p></li>\n  <li><p><code>maxvit_tiny_tf_224.in1k</code></p></li>\n  <li><p><code>maxvit_tiny_tf_384.in1k</code></p></li>\n  <li><p><code>maxvit_tiny_tf_512.in1k</code></p></li>\n  <li><p><code>maxvit_small_tf_224.in1k</code></p></li>\n  <li><p><code>maxvit_small_tf_384.in1k</code></p></li>\n  <li><p><code>maxvit_small_tf_512.in1k</code></p></li>\n  <li><p><code>maxvit_base_tf_224.in1k</code></p></li>\n  <li><p><code>maxvit_base_tf_384.in1k</code></p></li>\n  <li><p><code>maxvit_base_tf_512.in1k</code></p></li>\n  <li><p><code>maxvit_large_tf_224.in1k</code></p></li>\n  <li><p><code>maxvit_large_tf_384.in1k</code></p></li>\n  <li><p><code>maxvit_large_tf_512.in1k</code></p></li>\n  <li><p><code>maxvit_base_tf_224.in21k</code></p></li>\n  <li><p><code>maxvit_base_tf_384.in21k_ft_in1k</code></p></li>\n  <li><p><code>maxvit_base_tf_512.in21k_ft_in1k</code></p></li>\n  <li><p><code>maxvit_large_tf_224.in21k</code></p></li>\n  <li><p><code>maxvit_large_tf_384.in21k_ft_in1k</code></p></li>\n  <li><p><code>maxvit_large_tf_512.in21k_ft_in1k</code></p></li>\n  <li><p><code>maxvit_xlarge_tf_224.in21k</code></p></li>\n  <li><p><code>maxvit_xlarge_tf_384.in21k_ft_in1k</code></p></li>\n  <li><p><code>maxvit_xlarge_tf_512.in21k_ft_in1k</code></p></li>\n  </ul>\n</blockquote>\n<hr>\n<h4>CoAtNet - Google</h4>\n<p><strong>Paper:</strong> <a href=\"https://arxiv.org/abs/2106.04803\" target=\"_blank\">CoAtNet: Marrying Convolution and Attention for All Data Sizes</a><br>\n<strong>TIMM:</strong> <a href=\"https://github.com/rwightman/pytorch-image-models/blob/6902c48a5f0637c8155c1c4bc10ad35930f3e772/timm/models/maxxvit.py#L1927\" target=\"_blank\">source</a><br>\n<strong>Kaggle Notebook:</strong> <a href=\"https://www.kaggle.com/code/thedevastator/train-infer-coatnet-efficientnet\" target=\"_blank\">code</a></p>\n<p>The CoAtNet paper attempts to effectively combine the strengths from both convolutional and transformers architectures, they present CoAtNets(pronounced \"coat\" nets), a family of hybrid models built from two key insights:</p>\n<ul>\n<li>Depthwise Convolution and self-Attention can be naturally unified via simple relative attention</li>\n<li>Vertically stacking convolution layers and attention layers in a principled way is surprisingly effective in improving generalization, capacity and efficiency.</li>\n</ul>\n<p><img src=\"https://i.ibb.co/Sd6wj7D/Selection-998.png\" alt=\"\"></p>\n<p><strong>The \"we are the state-of-the-art\" claim:</strong></p>\n<p><img src=\"https://i.ibb.co/djDn4Y8/coatnet.png\" alt=\"\"></p>\n<blockquote>\n  <p><strong>TiMMs full list of pretrained weights</strong></p>\n  <ul>\n  <li><code>coatnet_pico_rw_224</code></li>\n  <li><code>coatnet_nano_rw_224</code></li>\n  <li><code>coatnet_0_rw_224</code></li>\n  <li><code>coatnet_1_rw_224</code></li>\n  <li><code>coatnet_2_rw_224</code></li>\n  <li><code>coatnet_3_rw_224</code></li>\n  <li><code>coatnet_bn_0_rw_224</code></li>\n  <li><code>coatnet_rmlp_nano_rw_224</code></li>\n  <li><code>coatnet_rmlp_0_rw_224</code></li>\n  <li><code>coatnet_rmlp_1_rw_224</code></li>\n  <li><code>coatnet_rmlp_1_rw2_224</code></li>\n  <li><code>coatnet_rmlp_2_rw_224</code></li>\n  <li><code>coatnet_rmlp_3_rw_224</code></li>\n  <li><code>coatnet_nano_cc_224</code></li>\n  <li><code>coatnext_nano_rw_224</code></li>\n  <li><code>coatnet_0_224</code></li>\n  <li><code>coatnet_1_224</code></li>\n  <li><code>coatnet_2_224</code></li>\n  <li><code>coatnet_3_224</code></li>\n  <li><code>coatnet_4_224</code></li>\n  <li><code>coatnet_5_224</code></li>\n  </ul>\n</blockquote>\n<hr>\n<h4>DeiT III - Meta (Facebook)</h4>\n<p><strong>Paper:</strong> <a href=\"https://arxiv.org/abs/2204.07118\" target=\"_blank\">DeiT III: Revenge of the ViT</a><br>\n<strong>TIMM:</strong> <a href=\"https://github.com/rwightman/pytorch-image-models/blob/6902c48a5f0637c8155c1c4bc10ad35930f3e772/timm/models/deit.py#L65\" target=\"_blank\">source</a><br>\n<strong>Note:</strong> Used on a <a href=\"https://www.kaggle.com/competitions/herbarium-2022-fgvc9/discussion/329299\" target=\"_blank\">winning solution</a> on Herbarium 2022 - FGVC9 </p>\n<p>In this work, the authors used several techniques to improve the training of large vision transformers (ViT) on the ImageNet dataset. These techniques include:</p>\n<ul>\n<li><strong>Stochastic depth</strong>, a regularization method that is especially useful for training deep networks.</li>\n<li><strong>LayerScale</strong>, a method introduced to facilitate the convergence of deep transformers.</li>\n<li><strong>Binary cross entropy loss</strong> (In contrast to categorical cross entropy) for Imagenet1k training, which was found to provide a significant improvement in performance for larger ViTs.</li>\n</ul>\n<p>They also found that a simple data augmentation method called 3-Augment worked better than other methods, and that using <strong>simple random cropping</strong> was more effective than random resize cropping when pre-training on a larger dataset like ImageNet-21k. They observed that using a lower resolution at training time had a regularizing effect and helped prevent overfitting, especially for the largest models.</p>\n<p>Additionally, they found that using the binary cross entropy loss provided a significant improvement in performance for larger ViTs trained on ImageNet-1k, but that using the cross entropy loss was more effective when pre-training with ImageNet-21k or for fine-tuning.</p>\n<p><strong>The \"we are the state-of-the-art\" claim:</strong></p>\n<p><img src=\"https://i.ibb.co/6Nqkqt9/deit.png\" alt=\"\"></p>\n<blockquote>\n  <p><strong>TiMMs full list of pretrained weights</strong></p>\n  <ul>\n  <li><p><code>deit3_small_patch16_224</code></p></li>\n  <li><p><code>deit3_small_patch16_384</code></p></li>\n  <li><p><code>deit3_medium_patch16_224</code></p></li>\n  <li><p><code>deit3_base_patch16_224</code></p></li>\n  <li><p><code>deit3_base_patch16_384</code></p></li>\n  <li><p><code>deit3_large_patch16_224</code></p></li>\n  <li><p><code>deit3_large_patch16_384</code></p></li>\n  <li><p><code>deit3_huge_patch14_224</code></p></li>\n  <li><p><code>deit3_small_patch16_224_in21ft1k</code></p></li>\n  <li><p><code>deit3_small_patch16_384_in21ft1k</code></p></li>\n  <li><p><code>deit3_medium_patch16_224_in21ft1k</code></p></li>\n  <li><p><code>deit3_base_patch16_224_in21ft1k</code></p></li>\n  <li><p><code>deit3_base_patch16_384_in21ft1k</code></p></li>\n  <li><p><code>deit3_large_patch16_224_in21ft1k</code></p></li>\n  <li><p><code>deit3_large_patch16_384_in21ft1k</code></p></li>\n  <li><p><code>deit3_huge_patch14_224_in21ft1k</code></p></li>\n  </ul>\n</blockquote>\n<hr>\n<h4>FlexiViT - Google</h4>\n<p><strong>Paper:</strong> <a href=\"https://arxiv.org/abs/2212.08013\" target=\"_blank\">FlexiViT: One Model for All Patch Sizes</a><br>\n<strong>TIMM:</strong> <a href=\"https://github.com/rwightman/pytorch-image-models/blob/6902c48a5f0637c8155c1c4bc10ad35930f3e772/timm/models/vision_transformer.py#L1579\" target=\"_blank\">source</a></p>\n<p>\"Simply randomizing the patch size at training time leads to a single set of weights that performs well across a wide range of patch sizes, making it possible to tailor the model to different compute budgets at deployment time.\"</p>\n<p><img src=\"https://i.ibb.co/MN2n0QH/Selection-1118.png\" alt=\"\"></p>\n<p><strong>The \"we are the state-of-the-art\" claim:</strong></p>\n<p><img src=\"https://i.ibb.co/wR9YC4H/flexivit.png\" alt=\"\"></p>\n<blockquote>\n  <p><strong>TiMMs full list of pretrained weights</strong></p>\n  <ul>\n  <li><code>flexivit_small.1200ep_in1k</code></li>\n  <li><code>flexivit_small.600ep_in1k</code></li>\n  <li><code>flexivit_small.300ep_in1k</code></li>\n  <li><code>flexivit_base.1200ep_in1k</code></li>\n  <li><code>flexivit_base.600ep_in1k</code></li>\n  <li><code>flexivit_base.300ep_in1k</code></li>\n  <li><code>flexivit_base.1000ep_in21k</code></li>\n  <li><code>flexivit_base.300ep_in21k</code></li>\n  <li><code>flexivit_large.1200ep_in1k</code></li>\n  <li><code>flexivit_large.600ep_in1k</code></li>\n  <li><code>flexivit_large.300ep_in1k</code></li>\n  <li><code>flexivit_base.patch16_in21k</code></li>\n  <li><code>flexivit_base.patch30_in21k</code></li>\n  </ul>\n</blockquote>\n<hr>\n<h4>GC ViT</h4>\n<p><strong>Paper:</strong> <a href=\"https://arxiv.org/pdf/2206.09959.pdf\" target=\"_blank\">Global Context Vision Transformers</a><br>\n<strong>TIMM:</strong> <a href=\"https://github.com/rwightman/pytorch-image-models/blob/6902c48a5f0637c8155c1c4bc10ad35930f3e772/timm/models/gcvit.py#L564\" target=\"_blank\">source</a><br>\n<strong>Kaggle Notebook:</strong> <a href=\"https://www.kaggle.com/code/awsaf49/gcvit-global-context-vision-transformer\" target=\"_blank\">code</a></p>\n<blockquote>\n  <p><strong>Credit:</strong> The summary is based on the <a href=\"https://www.kaggle.com/code/awsaf49/gcvit-global-context-vision-transformer\" target=\"_blank\">amazing notebook</a> by <a href=\"https://www.kaggle.com/awsaf49\" target=\"_blank\">Awsaf</a></p>\n</blockquote>\n<p>The paper proposes a novel architecture namely, <strong>Global Context Vision Transformer (GCViT)</strong> that utilizes window attention mechanism similar to <strong>Swin Transformer</strong>.<br>\nUnlike <strong>Swin Transformer</strong> this paper uses <strong>global context self-attention</strong>, with local self-attention, rather than <strong>shifted window self-attention</strong>, to model both long and short-range dependencies. Even though <strong>global-window-attention</strong> is a window-attention but it takes leverage of <strong>global query</strong> which contains global information hence captures long-range information. This paper compensates for the lack of the <strong>inductive bias</strong> that exists in both ViTs and Swin Transformer by utilizing a <strong>CNN</strong> based module. </p>\n<p><strong>GCViT achieves state-of-the-art results across image classification, object detection and semantic segmentation tasks.</strong></p>\n<blockquote>\n  <p><strong>TL;DR:</strong> Global Context ViT (<strong>GCViT</strong>) is a hierarchical architecture like <strong>Swin Transformer</strong> but utilizes<code>global-window-attention</code> instead of <code>shifted-window-attention</code> for effectively capturing long-range information. It also introduces <strong>CNN</strong> based module to include <strong>inductive-bias</strong> a useful feature for image that has been missing in both <strong>ViT</strong> and <strong>Swin Transformer</strong>.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/awsaf49\" target=\"_blank\">Awsaf</a> annotated the architecture figure to make it easier to digest:<br>\n<img src=\"https://raw.githubusercontent.com/awsaf49/gcvit-tf/main/image/arch_annot.png\" alt=\"\"></p>\n<ul>\n<li><p><code>SE</code>: <strong>Squeeze-Excitation (SE)</strong> aka <strong>Bottleneck</strong> module acts sd kind of <strong>channel attention</strong>. It consits of <strong>AvgPooling</strong>, <strong>Dense/FullyConnected (FC)/Linear</strong> , <strong>GELU</strong> and <strong>Sigmoid</strong> module.<br>\n<img src=\"https://raw.githubusercontent.com/awsaf49/gcvit-tf/main/image/se_annot.png\"></p></li>\n<li><p><code>Fused-MBConv:</code> This is similar to the one used in <strong>EfficientNetV2</strong>. It uses <strong>Depthwise-Conv</strong>, <strong>GELU</strong>, <strong>SE</strong>, <strong>Conv</strong>, to extract feature with a resiudal connection. Note that, no new module is declared for this one, we simply applied corresponding modules directly.<br>\n<img src=\"https://raw.githubusercontent.com/awsaf49/gcvit-tf/main/image/fmb_annot.png\"></p></li>\n<li><p><code>ReduceSize</code>: It is a <strong>CNN</strong> based <strong>downsample</strong> module which abvobe mentioned <code>Fused-MBConv</code> module to extract feature, <strong>Strided Conv</strong> to simultaneously reduce spatial dimension and increse channelwise dimention of the features and finally <strong>LayerNormalization</strong> module to normalize features. In the paper/figure this module is referred as <strong>downsample</strong> module. I think it is mention worthy that <strong>SwniTransformer</strong> used <code>PatchMerging</code> module instead of <code>ReduceSize</code> to reduce the spatial dimention and increase channelwise dimension which uses <strong>fully-connected/dense/linear</strong> module. According to the <strong>GCViT</strong> paper, one of the purposes of using <code>ReduceSize</code> is to add inductive bias through <strong>CNN</strong> module.<br>\n<img src=\"https://raw.githubusercontent.com/awsaf49/gcvit-tf/main/image/down_annot.png\"></p></li>\n</ul>\n<p><strong>The \"we are the state-of-the-art\" claim:</strong></p>\n<p><img src=\"https://i.ibb.co/thd6Gcn/gcvit.png\" alt=\"\"></p>\n<blockquote>\n  <p>**TIMM's full list of pretrained </p>\n  <ul>\n  <li><code>gcvit_xxtiny</code></li>\n  <li><code>gcvit_xtiny</code></li>\n  <li><code>gcvit_tiny</code></li>\n  <li><code>gcvit_small</code></li>\n  <li><code>gcvit_base</code></li>\n  </ul>\n</blockquote>\n<hr>\n<h4>EVA-L</h4>\n<p><strong>Paper:</strong> <a href=\"https://arxiv.org/abs/2211.07636\" target=\"_blank\">EVA: An Open Billion-Scale Vision Foundation Model</a><br>\n<strong>TIMM:</strong> <a href=\"https://github.com/rwightman/pytorch-image-models/blob/6902c48a5f0637c8155c1c4bc10ad35930f3e772/timm/models/vision_transformer.py#L1579\" target=\"_blank\">source</a></p>\n<p>EVA is a vanilla ViT pre-trained to reconstruct the masked out image-text aligned vision features (i.e., CLIP features) conditioned on visible image patches. Via this pretext task, the authors can efficiently scale up EVA to one billion parameters, and sets new records on a broad range of representative vision downstream tasks.</p>\n<p><strong>The \"we are the state-of-the-art\" claim:</strong></p>\n<ul>\n<li><strong>EVA-L is the best ViT-L (304M) to date</strong> that can reach up to 89.2 top-1 acc on IN-1K (weights &amp; logs) by leveraging vision features from EVA-CLIP.</li>\n</ul>\n<blockquote>\n  <p><strong>TiMMs full list of pretrained weights</strong></p>\n  <ul>\n  <li><p><code>eva_large_patch14_336.in22k_ft_in22k_in1k</code></p></li>\n  <li><p><code>eva_large_patch14_336.in22k_ft_in1k</code></p></li>\n  <li><p><code>eva_large_patch14_196.in22k_ft_in22k_in1k</code></p></li>\n  <li><p><code>eva_large_patch14_196.in22k_ft_in1k</code></p></li>\n  <li><p><code>eva_giant_patch14_560.m30m_ft_in22k_in1k</code></p></li>\n  <li><p><code>eva_giant_patch14_336.m30m_ft_in22k_in1k</code></p></li>\n  <li><p><code>eva_giant_patch14_336.clip_ft_in1k</code></p></li>\n  <li><p><code>eva_giant_patch14_224.clip_ft_in1k</code></p></li>\n  </ul>\n</blockquote>",
      "rawMarkdown": "##### EVERYTHING. is. state-of-the-art.\n\nOver the past year of computer vision, we have seen everything!\n\n- [Transformers?](https://arxiv.org/abs/2204.07118) State-of-the-art!\n- [Convolutional Networks?](https://arxiv.org/abs/2201.03545) State-of-the-art!\n- [ResNet-50?](https://arxiv.org/pdf/2110.00476.pdf) - State-of-the-art!\n- [Freaking MLP?!?!!](https://arxiv.org/abs/2105.01601) - State-of-the-art!\n\nEverything is state-of-the-art. \nIt doesn't matter what it is!\nIf it relates to computer vision, it has to be the best of the best! \"Everything is the current State-of-the-Art\" - and that's a fact!\n\nIn this short summary, I go through some of the notable papers claiming state-of-the-art performance on 2022. I attempt to explain each paper's main contributions of each paper while providing links to each paper's pretrained weights on `timm`.\n\nEnjoy!\n\n_____\n\n\n#### MaxViT - Google\n\n**Paper:** [MaxViT: Multi-Axis Vision Transformer](https://arxiv.org/abs/2204.01697)\n**TIMM:** [source](https://github.com/rwightman/pytorch-image-models/blob/main/timm/models/maxxvit.py)\n\nGoogle presents a new multi-axis approach that improves on the original ViT and MLP models, can better adapt to high-resolution, dense prediction tasks, and can naturally adapt to different input sizes with high flexibility and low complexity. \n\nThe new approach is based on multi-axis attention, which decomposes the full-size attention (each pixel attends to all the pixels) used in ViT into two sparse forms — local and (sparse) global. \n\nThe multi-axis attention contains a sequential stack of block attention and grid attention. The block attention works within non-overlapping windows (small patches in intermediate feature maps) to capture local patterns, while the grid attention works on a sparsely sampled uniform grid for long-range (global) interactions. The window sizes of grid and block attentions can be fully controlled as hyperparameters to ensure a linear computational complexity to the input size.\nThe proposed multi-axis attention conducts blocked local and dilated global attention sequentially followed by a FFN, with only a linear complexity. The pixels in the same colors are attended together.\n\n\n![](https://i.ibb.co/CvFWcP4/image4-2.png)\n\n\nSuch low-complexity attention can significantly improve its wide applicability to many vision tasks, especially for high-resolution visual predictions, demonstrating greater generality than the original attention used in ViT.\n\nThe MaxViT is built by concatenating MBConv (EfficientNet) with multi-axis attention. This single block can encode local and global visual information regardless of input resolution. We then simply stack repeated blocks composed of attention and convolutions in a hierarchical architecture (ResNet, CoAtNet, etc), \n\n\n![](https://i.ibb.co/Vv815YM/image6.png)\n\n\n**MaxViT is distinguished from previous hierarchical approaches as it can “see” globally throughout the entire network, even in earlier, high-resolution stages, demonstrating stronger model capacity on various tasks.**\n\n\n**The \"we are the state-of-the-art\" claim:**\n\n![](https://i.ibb.co/6BWNdk5/maxvit.png)\n\n\n\n> **TiMMs full list of pretrained weights**\n>\n>- `maxvit_pico_rw_256`\n>- `maxvit_nano_rw_256`\n>- `maxvit_tiny_rw_224`\n>- `maxvit_tiny_rw_256`\n>- `maxvit_rmlp_pico_rw_256`\n>- `maxvit_rmlp_nano_rw_256`\n>- `maxvit_rmlp_tiny_rw_256`\n>- `maxvit_rmlp_small_rw_224`\n>- `maxvit_rmlp_small_rw_256`\n>- `maxvit_tiny_pm_256`\n>- `maxxvit_rmlp_nano_rw_256`\n>- `maxxvit_rmlp_tiny_rw_256`\n>- `maxxvit_rmlp_small_rw_256`\n>- `maxxvit_rmlp_base_rw_224`\n>- `maxxvit_rmlp_large_rw_224`\n>\n>- `maxvit_tiny_tf_224.in1k`\n>- `maxvit_tiny_tf_384.in1k`\n>- `maxvit_tiny_tf_512.in1k`\n>- `maxvit_small_tf_224.in1k`\n>- `maxvit_small_tf_384.in1k`\n>- `maxvit_small_tf_512.in1k`\n>- `maxvit_base_tf_224.in1k`\n>- `maxvit_base_tf_384.in1k`\n>- `maxvit_base_tf_512.in1k`\n>- `maxvit_large_tf_224.in1k`\n>- `maxvit_large_tf_384.in1k`\n>- `maxvit_large_tf_512.in1k`\n>- `maxvit_base_tf_224.in21k`\n>- `maxvit_base_tf_384.in21k_ft_in1k`\n>- `maxvit_base_tf_512.in21k_ft_in1k`\n>- `maxvit_large_tf_224.in21k`\n>- `maxvit_large_tf_384.in21k_ft_in1k`\n>- `maxvit_large_tf_512.in21k_ft_in1k`\n>- `maxvit_xlarge_tf_224.in21k`\n>- `maxvit_xlarge_tf_384.in21k_ft_in1k`\n>- `maxvit_xlarge_tf_512.in21k_ft_in1k`\n\n\n_____\n\n\n\n#### CoAtNet - Google\n\n**Paper:** [CoAtNet: Marrying Convolution and Attention for All Data Sizes](https://arxiv.org/abs/2106.04803)\n**TIMM:** [source](https://github.com/rwightman/pytorch-image-models/blob/6902c48a5f0637c8155c1c4bc10ad35930f3e772/timm/models/maxxvit.py#L1927)\n**Kaggle Notebook:** [code](https://www.kaggle.com/code/thedevastator/train-infer-coatnet-efficientnet)\n\nThe CoAtNet paper attempts to effectively combine the strengths from both convolutional and transformers architectures, they present CoAtNets(pronounced \"coat\" nets), a family of hybrid models built from two key insights:\n\n- Depthwise Convolution and self-Attention can be naturally unified via simple relative attention\n- Vertically stacking convolution layers and attention layers in a principled way is surprisingly effective in improving generalization, capacity and efficiency.\n\n\n![](https://i.ibb.co/Sd6wj7D/Selection-998.png)\n\n\n**The \"we are the state-of-the-art\" claim:**\n\n![](https://i.ibb.co/djDn4Y8/coatnet.png)\n\n\n\n\n> **TiMMs full list of pretrained weights**\n>\n>- `coatnet_pico_rw_224`\n>- `coatnet_nano_rw_224`\n>- `coatnet_0_rw_224`\n>- `coatnet_1_rw_224`\n>- `coatnet_2_rw_224`\n>- `coatnet_3_rw_224`\n>- `coatnet_bn_0_rw_224`\n>- `coatnet_rmlp_nano_rw_224`\n>- `coatnet_rmlp_0_rw_224`\n>- `coatnet_rmlp_1_rw_224`\n>- `coatnet_rmlp_1_rw2_224`\n>- `coatnet_rmlp_2_rw_224`\n>- `coatnet_rmlp_3_rw_224`\n>- `coatnet_nano_cc_224`\n>- `coatnext_nano_rw_224`\n>- `coatnet_0_224`\n>- `coatnet_1_224`\n>- `coatnet_2_224`\n>- `coatnet_3_224`\n>- `coatnet_4_224`\n>- `coatnet_5_224`\n\n\n\n\n_____\n\n\n\n#### DeiT III - Meta (Facebook)\n\n**Paper:** [DeiT III: Revenge of the ViT](https://arxiv.org/abs/2204.07118)\n**TIMM:** [source](https://github.com/rwightman/pytorch-image-models/blob/6902c48a5f0637c8155c1c4bc10ad35930f3e772/timm/models/deit.py#L65)\n**Note:** Used on a [winning solution](https://www.kaggle.com/competitions/herbarium-2022-fgvc9/discussion/329299) on Herbarium 2022 - FGVC9 \n\n\nIn this work, the authors used several techniques to improve the training of large vision transformers (ViT) on the ImageNet dataset. These techniques include:\n\n- **Stochastic depth**, a regularization method that is especially useful for training deep networks.\n- **LayerScale**, a method introduced to facilitate the convergence of deep transformers.\n- **Binary cross entropy loss** (In contrast to categorical cross entropy) for Imagenet1k training, which was found to provide a significant improvement in performance for larger ViTs.\n\nThey also found that a simple data augmentation method called 3-Augment worked better than other methods, and that using **simple random cropping** was more effective than random resize cropping when pre-training on a larger dataset like ImageNet-21k. They observed that using a lower resolution at training time had a regularizing effect and helped prevent overfitting, especially for the largest models.\n\nAdditionally, they found that using the binary cross entropy loss provided a significant improvement in performance for larger ViTs trained on ImageNet-1k, but that using the cross entropy loss was more effective when pre-training with ImageNet-21k or for fine-tuning.\n\n\n**The \"we are the state-of-the-art\" claim:**\n\n![](https://i.ibb.co/6Nqkqt9/deit.png)\n\n\n\n> **TiMMs full list of pretrained weights**\n>\n>- `deit3_small_patch16_224`\n>- `deit3_small_patch16_384`\n>- `deit3_medium_patch16_224`\n>- `deit3_base_patch16_224`\n>- `deit3_base_patch16_384`\n>- `deit3_large_patch16_224`\n>- `deit3_large_patch16_384`\n>- `deit3_huge_patch14_224`\n\n>- `deit3_small_patch16_224_in21ft1k`\n>- `deit3_small_patch16_384_in21ft1k`\n>- `deit3_medium_patch16_224_in21ft1k`\n>- `deit3_base_patch16_224_in21ft1k`\n>- `deit3_base_patch16_384_in21ft1k`\n>- `deit3_large_patch16_224_in21ft1k`\n>- `deit3_large_patch16_384_in21ft1k`\n>- `deit3_huge_patch14_224_in21ft1k`\n\n\n\n_____\n\n\n\n\n#### FlexiViT - Google\n\n**Paper:** [FlexiViT: One Model for All Patch Sizes](https://arxiv.org/abs/2212.08013)\n**TIMM:** [source](https://github.com/rwightman/pytorch-image-models/blob/6902c48a5f0637c8155c1c4bc10ad35930f3e772/timm/models/vision_transformer.py#L1579)\n\n\"Simply randomizing the patch size at training time leads to a single set of weights that performs well across a wide range of patch sizes, making it possible to tailor the model to different compute budgets at deployment time.\"\n\n![](https://i.ibb.co/MN2n0QH/Selection-1118.png)\n\n\n**The \"we are the state-of-the-art\" claim:**\n\n![](https://i.ibb.co/wR9YC4H/flexivit.png)\n\n\n\n> **TiMMs full list of pretrained weights**\n>\n>- `flexivit_small.1200ep_in1k`\n>- `flexivit_small.600ep_in1k`\n>- `flexivit_small.300ep_in1k`\n>- `flexivit_base.1200ep_in1k`\n>- `flexivit_base.600ep_in1k`\n>- `flexivit_base.300ep_in1k`\n>- `flexivit_base.1000ep_in21k`\n>- `flexivit_base.300ep_in21k`\n>- `flexivit_large.1200ep_in1k`\n>- `flexivit_large.600ep_in1k`\n>- `flexivit_large.300ep_in1k`\n>- `flexivit_base.patch16_in21k`\n>- `flexivit_base.patch30_in21k`\n\n\n\n_____\n\n\n\n\n\n#### GC ViT\n\n**Paper:** [Global Context Vision Transformers](https://arxiv.org/pdf/2206.09959.pdf)\n**TIMM:** [source](https://github.com/rwightman/pytorch-image-models/blob/6902c48a5f0637c8155c1c4bc10ad35930f3e772/timm/models/gcvit.py#L564)\n**Kaggle Notebook:** [code](https://www.kaggle.com/code/awsaf49/gcvit-global-context-vision-transformer)\n\n\n> **Credit:** The summary is based on the [amazing notebook](https://www.kaggle.com/code/awsaf49/gcvit-global-context-vision-transformer) by [Awsaf](https://www.kaggle.com/awsaf49)\n\n\nThe paper proposes a novel architecture namely, **Global Context Vision Transformer (GCViT)** that utilizes window attention mechanism similar to **Swin Transformer**.\nUnlike **Swin Transformer** this paper uses **global context self-attention**, with local self-attention, rather than **shifted window self-attention**, to model both long and short-range dependencies. Even though **global-window-attention** is a window-attention but it takes leverage of **global query** which contains global information hence captures long-range information. This paper compensates for the lack of the **inductive bias** that exists in both ViTs and Swin Transformer by utilizing a **CNN** based module. \n\n**GCViT achieves state-of-the-art results across image classification, object detection and semantic segmentation tasks.**\n\n> **TL;DR:** Global Context ViT (**GCViT**) is a hierarchical architecture like **Swin Transformer** but utilizes`global-window-attention` instead of `shifted-window-attention` for effectively capturing long-range information. It also introduces **CNN** based module to include **inductive-bias** a useful feature for image that has been missing in both **ViT** and **Swin Transformer**.\n\n\n[Awsaf](https://www.kaggle.com/awsaf49) annotated the architecture figure to make it easier to digest:\n![](https://raw.githubusercontent.com/awsaf49/gcvit-tf/main/image/arch_annot.png)\n\n- `SE`: **Squeeze-Excitation (SE)** aka **Bottleneck** module acts sd kind of **channel attention**. It consits of **AvgPooling**, **Dense/FullyConnected (FC)/Linear** , **GELU** and **Sigmoid** module.\n<img src=\"https://raw.githubusercontent.com/awsaf49/gcvit-tf/main/image/se_annot.png\" width=400>\n\n- `Fused-MBConv:` This is similar to the one used in **EfficientNetV2**. It uses **Depthwise-Conv**, **GELU**, **SE**, **Conv**, to extract feature with a resiudal connection. Note that, no new module is declared for this one, we simply applied corresponding modules directly.\n<img src=\"https://raw.githubusercontent.com/awsaf49/gcvit-tf/main/image/fmb_annot.png\" width=350>\n\n\n- `ReduceSize`: It is a **CNN** based **downsample** module which abvobe mentioned `Fused-MBConv` module to extract feature, **Strided Conv** to simultaneously reduce spatial dimension and increse channelwise dimention of the features and finally **LayerNormalization** module to normalize features. In the paper/figure this module is referred as **downsample** module. I think it is mention worthy that **SwniTransformer** used `PatchMerging` module instead of `ReduceSize` to reduce the spatial dimention and increase channelwise dimension which uses **fully-connected/dense/linear** module. According to the **GCViT** paper, one of the purposes of using `ReduceSize` is to add inductive bias through **CNN** module.\n<img src=\"https://raw.githubusercontent.com/awsaf49/gcvit-tf/main/image/down_annot.png\" width=300>\n\n\n**The \"we are the state-of-the-art\" claim:**\n\n![](https://i.ibb.co/thd6Gcn/gcvit.png)\n\n\n> **TIMM's full list of pretrained \n>\n>- `gcvit_xxtiny`\n>- `gcvit_xtiny`\n>- `gcvit_tiny`\n>- `gcvit_small`\n>- `gcvit_base`\n\n\n\n_____\n\n\n\n\n\n#### EVA-L\n**Paper:** [EVA: An Open Billion-Scale Vision Foundation Model](https://arxiv.org/abs/2211.07636)\n**TIMM:** [source](https://github.com/rwightman/pytorch-image-models/blob/6902c48a5f0637c8155c1c4bc10ad35930f3e772/timm/models/vision_transformer.py#L1579)\n\nEVA is a vanilla ViT pre-trained to reconstruct the masked out image-text aligned vision features (i.e., CLIP features) conditioned on visible image patches. Via this pretext task, the authors can efficiently scale up EVA to one billion parameters, and sets new records on a broad range of representative vision downstream tasks.\n\n\n**The \"we are the state-of-the-art\" claim:**\n\n- **EVA-L is the best ViT-L (304M) to date** that can reach up to 89.2 top-1 acc on IN-1K (weights & logs) by leveraging vision features from EVA-CLIP.\n\n\n\n> **TiMMs full list of pretrained weights**\n>\n>- `eva_large_patch14_336.in22k_ft_in22k_in1k`\n>- `eva_large_patch14_336.in22k_ft_in1k`\n>- `eva_large_patch14_196.in22k_ft_in22k_in1k`\n>- `eva_large_patch14_196.in22k_ft_in1k`\n>\n>- `eva_giant_patch14_560.m30m_ft_in22k_in1k`\n>- `eva_giant_patch14_336.m30m_ft_in22k_in1k`\n>- `eva_giant_patch14_336.clip_ft_in1k`\n>- `eva_giant_patch14_224.clip_ft_in1k`\n",
      "votes": 30
    },
    {
      "id": 2084247,
      "postDate": "2023-01-03T10:42:06.267Z",
      "content": "<p>Very informative documentation, Thanks for your efforts.</p>",
      "rawMarkdown": "Very informative documentation, Thanks for your efforts.",
      "votes": 1
    },
    {
      "id": 2081162,
      "postDate": "2022-12-30T21:39:45.893Z",
      "content": "<p>haha, fun and informative read, <a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a>! 😁 the first paragraph made me chuckle! but yeah, lovely summary, how much keeps happening in CV sort of escapes us sometimes given how the debate is mostly taken over by generative models and NLP</p>\n<p>thx for the summary! 🙌</p>",
      "rawMarkdown": "haha, fun and informative read, @thedevastator! 😁 the first paragraph made me chuckle! but yeah, lovely summary, how much keeps happening in CV sort of escapes us sometimes given how the debate is mostly taken over by generative models and NLP\n\nthx for the summary! 🙌",
      "votes": 1
    },
    {
      "id": 2090264,
      "postDate": "2023-01-07T07:12:28.097Z",
      "content": "<p>appreciate your effort to put it all together.</p>",
      "rawMarkdown": "appreciate your effort to put it all together."
    },
    {
      "id": 2090185,
      "postDate": "2023-01-07T04:23:07.550Z",
      "content": "<p>Thanks, really helpful to draw some inspiration from some SOTA methods!</p>",
      "rawMarkdown": "Thanks, really helpful to draw some inspiration from some SOTA methods!"
    },
    {
      "id": 2086665,
      "postDate": "2023-01-04T23:16:02.413Z",
      "content": "<p>Thanks. These papers are very interesting!</p>",
      "rawMarkdown": "Thanks. These papers are very interesting!"
    },
    {
      "id": 2084902,
      "postDate": "2023-01-03T20:35:58.950Z",
      "content": "<p>Very nice work! thks👊</p>",
      "rawMarkdown": "Very nice work! thks👊"
    }
  ],
  "comments": [
    {
      "id": 2084247,
      "author_name": "Gaju Ahmed",
      "author_url": "",
      "post_date": "2023-01-03T10:42:06.267000",
      "content": "<p>Very informative documentation, Thanks for your efforts.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2081162,
      "author_name": "Radek Osmulski",
      "author_url": "",
      "post_date": "2022-12-30T21:39:45.893000",
      "content": "<p>haha, fun and informative read, <a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a>! 😁 the first paragraph made me chuckle! but yeah, lovely summary, how much keeps happening in CV sort of escapes us sometimes given how the debate is mostly taken over by generative models and NLP</p>\n<p>thx for the summary! 🙌</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2090264,
      "author_name": "tau__tsm1",
      "author_url": "",
      "post_date": "2023-01-07T07:12:28.097000",
      "content": "<p>appreciate your effort to put it all together.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2090185,
      "author_name": "Scott Campit",
      "author_url": "",
      "post_date": "2023-01-07T04:23:07.550000",
      "content": "<p>Thanks, really helpful to draw some inspiration from some SOTA methods!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2086665,
      "author_name": "Fabio Souza",
      "author_url": "",
      "post_date": "2023-01-04T23:16:02.413000",
      "content": "<p>Thanks. These papers are very interesting!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2084902,
      "author_name": "starsnew",
      "author_url": "",
      "post_date": "2023-01-03T20:35:58.950000",
      "content": "<p>Very nice work! thks👊</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2080985": "##### EVERYTHING. is. state-of-the-art.\n\nOver the past year of computer vision, we have seen everything!\n\n- [Transformers?](https://arxiv.org/abs/2204.07118) State-of-the-art!\n- [Convolutional Networks?](https://arxiv.org/abs/2201.03545) State-of-the-art!\n- [ResNet-50?](https://arxiv.org/pdf/2110.00476.pdf) - State-of-the-art!\n- [Freaking MLP?!?!!](https://arxiv.org/abs/2105.01601) - State-of-the-art!\n\nEverything is state-of-the-art. \nIt doesn't matter what it is!\nIf it relates to computer vision, it has to be the best of the best! \"Everything is the current State-of-the-Art\" - and that's a fact!\n\nIn this short summary, I go through some of the notable papers claiming state-of-the-art performance on 2022. I attempt to explain each paper's main contributions of each paper while providing links to each paper's pretrained weights on `timm`.\n\nEnjoy!\n\n_____\n\n\n#### MaxViT - Google\n\n**Paper:** [MaxViT: Multi-Axis Vision Transformer](https://arxiv.org/abs/2204.01697)\n**TIMM:** [source](https://github.com/rwightman/pytorch-image-models/blob/main/timm/models/maxxvit.py)\n\nGoogle presents a new multi-axis approach that improves on the original ViT and MLP models, can better adapt to high-resolution, dense prediction tasks, and can naturally adapt to different input sizes with high flexibility and low complexity. \n\nThe new approach is based on multi-axis attention, which decomposes the full-size attention (each pixel attends to all the pixels) used in ViT into two sparse forms — local and (sparse) global. \n\nThe multi-axis attention contains a sequential stack of block attention and grid attention. The block attention works within non-overlapping windows (small patches in intermediate feature maps) to capture local patterns, while the grid attention works on a sparsely sampled uniform grid for long-range (global) interactions. The window sizes of grid and block attentions can be fully controlled as hyperparameters to ensure a linear computational complexity to the input size.\nThe proposed multi-axis attention conducts blocked local and dilated global attention sequentially followed by a FFN, with only a linear complexity. The pixels in the same colors are attended together.\n\n\n![](https://i.ibb.co/CvFWcP4/image4-2.png)\n\n\nSuch low-complexity attention can significantly improve its wide applicability to many vision tasks, especially for high-resolution visual predictions, demonstrating greater generality than the original attention used in ViT.\n\nThe MaxViT is built by concatenating MBConv (EfficientNet) with multi-axis attention. This single block can encode local and global visual information regardless of input resolution. We then simply stack repeated blocks composed of attention and convolutions in a hierarchical architecture (ResNet, CoAtNet, etc), \n\n\n![](https://i.ibb.co/Vv815YM/image6.png)\n\n\n**MaxViT is distinguished from previous hierarchical approaches as it can “see” globally throughout the entire network, even in earlier, high-resolution stages, demonstrating stronger model capacity on various tasks.**\n\n\n**The \"we are the state-of-the-art\" claim:**\n\n![](https://i.ibb.co/6BWNdk5/maxvit.png)\n\n\n\n> **TiMMs full list of pretrained weights**\n>\n>- `maxvit_pico_rw_256`\n>- `maxvit_nano_rw_256`\n>- `maxvit_tiny_rw_224`\n>- `maxvit_tiny_rw_256`\n>- `maxvit_rmlp_pico_rw_256`\n>- `maxvit_rmlp_nano_rw_256`\n>- `maxvit_rmlp_tiny_rw_256`\n>- `maxvit_rmlp_small_rw_224`\n>- `maxvit_rmlp_small_rw_256`\n>- `maxvit_tiny_pm_256`\n>- `maxxvit_rmlp_nano_rw_256`\n>- `maxxvit_rmlp_tiny_rw_256`\n>- `maxxvit_rmlp_small_rw_256`\n>- `maxxvit_rmlp_base_rw_224`\n>- `maxxvit_rmlp_large_rw_224`\n>\n>- `maxvit_tiny_tf_224.in1k`\n>- `maxvit_tiny_tf_384.in1k`\n>- `maxvit_tiny_tf_512.in1k`\n>- `maxvit_small_tf_224.in1k`\n>- `maxvit_small_tf_384.in1k`\n>- `maxvit_small_tf_512.in1k`\n>- `maxvit_base_tf_224.in1k`\n>- `maxvit_base_tf_384.in1k`\n>- `maxvit_base_tf_512.in1k`\n>- `maxvit_large_tf_224.in1k`\n>- `maxvit_large_tf_384.in1k`\n>- `maxvit_large_tf_512.in1k`\n>- `maxvit_base_tf_224.in21k`\n>- `maxvit_base_tf_384.in21k_ft_in1k`\n>- `maxvit_base_tf_512.in21k_ft_in1k`\n>- `maxvit_large_tf_224.in21k`\n>- `maxvit_large_tf_384.in21k_ft_in1k`\n>- `maxvit_large_tf_512.in21k_ft_in1k`\n>- `maxvit_xlarge_tf_224.in21k`\n>- `maxvit_xlarge_tf_384.in21k_ft_in1k`\n>- `maxvit_xlarge_tf_512.in21k_ft_in1k`\n\n\n_____\n\n\n\n#### CoAtNet - Google\n\n**Paper:** [CoAtNet: Marrying Convolution and Attention for All Data Sizes](https://arxiv.org/abs/2106.04803)\n**TIMM:** [source](https://github.com/rwightman/pytorch-image-models/blob/6902c48a5f0637c8155c1c4bc10ad35930f3e772/timm/models/maxxvit.py#L1927)\n**Kaggle Notebook:** [code](https://www.kaggle.com/code/thedevastator/train-infer-coatnet-efficientnet)\n\nThe CoAtNet paper attempts to effectively combine the strengths from both convolutional and transformers architectures, they present CoAtNets(pronounced \"coat\" nets), a family of hybrid models built from two key insights:\n\n- Depthwise Convolution and self-Attention can be naturally unified via simple relative attention\n- Vertically stacking convolution layers and attention layers in a principled way is surprisingly effective in improving generalization, capacity and efficiency.\n\n\n![](https://i.ibb.co/Sd6wj7D/Selection-998.png)\n\n\n**The \"we are the state-of-the-art\" claim:**\n\n![](https://i.ibb.co/djDn4Y8/coatnet.png)\n\n\n\n\n> **TiMMs full list of pretrained weights**\n>\n>- `coatnet_pico_rw_224`\n>- `coatnet_nano_rw_224`\n>- `coatnet_0_rw_224`\n>- `coatnet_1_rw_224`\n>- `coatnet_2_rw_224`\n>- `coatnet_3_rw_224`\n>- `coatnet_bn_0_rw_224`\n>- `coatnet_rmlp_nano_rw_224`\n>- `coatnet_rmlp_0_rw_224`\n>- `coatnet_rmlp_1_rw_224`\n>- `coatnet_rmlp_1_rw2_224`\n>- `coatnet_rmlp_2_rw_224`\n>- `coatnet_rmlp_3_rw_224`\n>- `coatnet_nano_cc_224`\n>- `coatnext_nano_rw_224`\n>- `coatnet_0_224`\n>- `coatnet_1_224`\n>- `coatnet_2_224`\n>- `coatnet_3_224`\n>- `coatnet_4_224`\n>- `coatnet_5_224`\n\n\n\n\n_____\n\n\n\n#### DeiT III - Meta (Facebook)\n\n**Paper:** [DeiT III: Revenge of the ViT](https://arxiv.org/abs/2204.07118)\n**TIMM:** [source](https://github.com/rwightman/pytorch-image-models/blob/6902c48a5f0637c8155c1c4bc10ad35930f3e772/timm/models/deit.py#L65)\n**Note:** Used on a [winning solution](https://www.kaggle.com/competitions/herbarium-2022-fgvc9/discussion/329299) on Herbarium 2022 - FGVC9 \n\n\nIn this work, the authors used several techniques to improve the training of large vision transformers (ViT) on the ImageNet dataset. These techniques include:\n\n- **Stochastic depth**, a regularization method that is especially useful for training deep networks.\n- **LayerScale**, a method introduced to facilitate the convergence of deep transformers.\n- **Binary cross entropy loss** (In contrast to categorical cross entropy) for Imagenet1k training, which was found to provide a significant improvement in performance for larger ViTs.\n\nThey also found that a simple data augmentation method called 3-Augment worked better than other methods, and that using **simple random cropping** was more effective than random resize cropping when pre-training on a larger dataset like ImageNet-21k. They observed that using a lower resolution at training time had a regularizing effect and helped prevent overfitting, especially for the largest models.\n\nAdditionally, they found that using the binary cross entropy loss provided a significant improvement in performance for larger ViTs trained on ImageNet-1k, but that using the cross entropy loss was more effective when pre-training with ImageNet-21k or for fine-tuning.\n\n\n**The \"we are the state-of-the-art\" claim:**\n\n![](https://i.ibb.co/6Nqkqt9/deit.png)\n\n\n\n> **TiMMs full list of pretrained weights**\n>\n>- `deit3_small_patch16_224`\n>- `deit3_small_patch16_384`\n>- `deit3_medium_patch16_224`\n>- `deit3_base_patch16_224`\n>- `deit3_base_patch16_384`\n>- `deit3_large_patch16_224`\n>- `deit3_large_patch16_384`\n>- `deit3_huge_patch14_224`\n\n>- `deit3_small_patch16_224_in21ft1k`\n>- `deit3_small_patch16_384_in21ft1k`\n>- `deit3_medium_patch16_224_in21ft1k`\n>- `deit3_base_patch16_224_in21ft1k`\n>- `deit3_base_patch16_384_in21ft1k`\n>- `deit3_large_patch16_224_in21ft1k`\n>- `deit3_large_patch16_384_in21ft1k`\n>- `deit3_huge_patch14_224_in21ft1k`\n\n\n\n_____\n\n\n\n\n#### FlexiViT - Google\n\n**Paper:** [FlexiViT: One Model for All Patch Sizes](https://arxiv.org/abs/2212.08013)\n**TIMM:** [source](https://github.com/rwightman/pytorch-image-models/blob/6902c48a5f0637c8155c1c4bc10ad35930f3e772/timm/models/vision_transformer.py#L1579)\n\n\"Simply randomizing the patch size at training time leads to a single set of weights that performs well across a wide range of patch sizes, making it possible to tailor the model to different compute budgets at deployment time.\"\n\n![](https://i.ibb.co/MN2n0QH/Selection-1118.png)\n\n\n**The \"we are the state-of-the-art\" claim:**\n\n![](https://i.ibb.co/wR9YC4H/flexivit.png)\n\n\n\n> **TiMMs full list of pretrained weights**\n>\n>- `flexivit_small.1200ep_in1k`\n>- `flexivit_small.600ep_in1k`\n>- `flexivit_small.300ep_in1k`\n>- `flexivit_base.1200ep_in1k`\n>- `flexivit_base.600ep_in1k`\n>- `flexivit_base.300ep_in1k`\n>- `flexivit_base.1000ep_in21k`\n>- `flexivit_base.300ep_in21k`\n>- `flexivit_large.1200ep_in1k`\n>- `flexivit_large.600ep_in1k`\n>- `flexivit_large.300ep_in1k`\n>- `flexivit_base.patch16_in21k`\n>- `flexivit_base.patch30_in21k`\n\n\n\n_____\n\n\n\n\n\n#### GC ViT\n\n**Paper:** [Global Context Vision Transformers](https://arxiv.org/pdf/2206.09959.pdf)\n**TIMM:** [source](https://github.com/rwightman/pytorch-image-models/blob/6902c48a5f0637c8155c1c4bc10ad35930f3e772/timm/models/gcvit.py#L564)\n**Kaggle Notebook:** [code](https://www.kaggle.com/code/awsaf49/gcvit-global-context-vision-transformer)\n\n\n> **Credit:** The summary is based on the [amazing notebook](https://www.kaggle.com/code/awsaf49/gcvit-global-context-vision-transformer) by [Awsaf](https://www.kaggle.com/awsaf49)\n\n\nThe paper proposes a novel architecture namely, **Global Context Vision Transformer (GCViT)** that utilizes window attention mechanism similar to **Swin Transformer**.\nUnlike **Swin Transformer** this paper uses **global context self-attention**, with local self-attention, rather than **shifted window self-attention**, to model both long and short-range dependencies. Even though **global-window-attention** is a window-attention but it takes leverage of **global query** which contains global information hence captures long-range information. This paper compensates for the lack of the **inductive bias** that exists in both ViTs and Swin Transformer by utilizing a **CNN** based module. \n\n**GCViT achieves state-of-the-art results across image classification, object detection and semantic segmentation tasks.**\n\n> **TL;DR:** Global Context ViT (**GCViT**) is a hierarchical architecture like **Swin Transformer** but utilizes`global-window-attention` instead of `shifted-window-attention` for effectively capturing long-range information. It also introduces **CNN** based module to include **inductive-bias** a useful feature for image that has been missing in both **ViT** and **Swin Transformer**.\n\n\n[Awsaf](https://www.kaggle.com/awsaf49) annotated the architecture figure to make it easier to digest:\n![](https://raw.githubusercontent.com/awsaf49/gcvit-tf/main/image/arch_annot.png)\n\n- `SE`: **Squeeze-Excitation (SE)** aka **Bottleneck** module acts sd kind of **channel attention**. It consits of **AvgPooling**, **Dense/FullyConnected (FC)/Linear** , **GELU** and **Sigmoid** module.\n<img src=\"https://raw.githubusercontent.com/awsaf49/gcvit-tf/main/image/se_annot.png\" width=400>\n\n- `Fused-MBConv:` This is similar to the one used in **EfficientNetV2**. It uses **Depthwise-Conv**, **GELU**, **SE**, **Conv**, to extract feature with a resiudal connection. Note that, no new module is declared for this one, we simply applied corresponding modules directly.\n<img src=\"https://raw.githubusercontent.com/awsaf49/gcvit-tf/main/image/fmb_annot.png\" width=350>\n\n\n- `ReduceSize`: It is a **CNN** based **downsample** module which abvobe mentioned `Fused-MBConv` module to extract feature, **Strided Conv** to simultaneously reduce spatial dimension and increse channelwise dimention of the features and finally **LayerNormalization** module to normalize features. In the paper/figure this module is referred as **downsample** module. I think it is mention worthy that **SwniTransformer** used `PatchMerging` module instead of `ReduceSize` to reduce the spatial dimention and increase channelwise dimension which uses **fully-connected/dense/linear** module. According to the **GCViT** paper, one of the purposes of using `ReduceSize` is to add inductive bias through **CNN** module.\n<img src=\"https://raw.githubusercontent.com/awsaf49/gcvit-tf/main/image/down_annot.png\" width=300>\n\n\n**The \"we are the state-of-the-art\" claim:**\n\n![](https://i.ibb.co/thd6Gcn/gcvit.png)\n\n\n> **TIMM's full list of pretrained \n>\n>- `gcvit_xxtiny`\n>- `gcvit_xtiny`\n>- `gcvit_tiny`\n>- `gcvit_small`\n>- `gcvit_base`\n\n\n\n_____\n\n\n\n\n\n#### EVA-L\n**Paper:** [EVA: An Open Billion-Scale Vision Foundation Model](https://arxiv.org/abs/2211.07636)\n**TIMM:** [source](https://github.com/rwightman/pytorch-image-models/blob/6902c48a5f0637c8155c1c4bc10ad35930f3e772/timm/models/vision_transformer.py#L1579)\n\nEVA is a vanilla ViT pre-trained to reconstruct the masked out image-text aligned vision features (i.e., CLIP features) conditioned on visible image patches. Via this pretext task, the authors can efficiently scale up EVA to one billion parameters, and sets new records on a broad range of representative vision downstream tasks.\n\n\n**The \"we are the state-of-the-art\" claim:**\n\n- **EVA-L is the best ViT-L (304M) to date** that can reach up to 89.2 top-1 acc on IN-1K (weights & logs) by leveraging vision features from EVA-CLIP.\n\n\n\n> **TiMMs full list of pretrained weights**\n>\n>- `eva_large_patch14_336.in22k_ft_in22k_in1k`\n>- `eva_large_patch14_336.in22k_ft_in1k`\n>- `eva_large_patch14_196.in22k_ft_in22k_in1k`\n>- `eva_large_patch14_196.in22k_ft_in1k`\n>\n>- `eva_giant_patch14_560.m30m_ft_in22k_in1k`\n>- `eva_giant_patch14_336.m30m_ft_in22k_in1k`\n>- `eva_giant_patch14_336.clip_ft_in1k`\n>- `eva_giant_patch14_224.clip_ft_in1k`\n",
    "2084247": "Very informative documentation, Thanks for your efforts.",
    "2081162": "haha, fun and informative read, @thedevastator! 😁 the first paragraph made me chuckle! but yeah, lovely summary, how much keeps happening in CV sort of escapes us sometimes given how the debate is mostly taken over by generative models and NLP\n\nthx for the summary! 🙌",
    "2090264": "appreciate your effort to put it all together.",
    "2090185": "Thanks, really helpful to draw some inspiration from some SOTA methods!",
    "2086665": "Thanks. These papers are very interesting!",
    "2084902": "Very nice work! thks👊"
  }
}