{
  "id": 374259,
  "title": "Use self-supervision for training on unlabeled external datasets",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/374259",
  "author_name": "The Devastator",
  "post_date": "2022-12-26T08:02:41.057000",
  "votes": 12,
  "comment_count": 2,
  "views": 0,
  "content": "<h4>Self-supervision</h4>\n<p><strong>Should we try it?</strong></p>\n<p>Many fellow Kagglers are using external datasets, some of them are unlabeled (or labeled for a different task).<br>\nSelf supervision methods can help capture information from unlabeled datasets. </p>\n<blockquote>\n  <p><strong>Also:</strong> They are a really good \"secret weapon\" to have in your back pocket since if they are implemented in the correct way they can \"just work\" all the time and boost the performance of every model <br>\n  <strong>(Like we treat MLM in NLP nowadays because huggingface did all the heavy lifting for us)</strong></p>\n</blockquote>\n<hr>\n<p><strong>Self Supervision Baselines</strong></p>\n<p>There had been many self-supervision papers released recently, today we will focus on the following: </p>\n<ul>\n<li><strong>BYOL:</strong> <a href=\"https://arxiv.org/abs/2006.07733\" target=\"_blank\">Bootstrap Your Own Latent</a> - </li>\n<li><strong>SwAV﻿:</strong> <a href=\"https://arxiv.org/abs/2006.09882\" target=\"_blank\">Unsupervised Learning of Visual Features by Contrasting Cluster Assignments</a></li>\n<li><strong>SimCLR:</strong> <a href=\"https://arxiv.org/abs/2002.05709\" target=\"_blank\">A Simple Framework for Contrastive Learning of Visual Representations﻿</a></li>\n</ul>\n<hr>\n<h3>SimCLR</h3>\n<h5>A Simple Framework for Contrastive Learning of Visual Representations﻿</h5>\n<blockquote>\n  <ul>\n  <li><a href=\"https://www.kaggle.com/code/aritrag/simclr\" target=\"_blank\">Kaggle Notebook</a></li>\n  <li><a href=\"https://arxiv.org/abs/2002.05709\" target=\"_blank\">Paper</a></li>\n  <li><a href=\"https://github.com/google-research/simclr\" target=\"_blank\">Official Implementation</a></li>\n  </ul>\n</blockquote>\n<p>A Simple framework for contrastive learning of visual representation</p>\n<p><img src=\"https://i.ibb.co/dMFvndK/image4.gif\" alt=\"\"></p>\n<p>SimCLR provides a great platform (and an easy one too) to help achieve a good representation out of unlabelled images. <br>\nThe itutions and conjectures provided by the paper comes quite naturally to one's mind. The ease of the concepts will be portrayed in the kernel along with some comments on the same.</p>\n<p><strong>SimCLR is based out of the following simplified modules:</strong></p>\n<ul>\n<li>A stochastic data augmentation module.</li>\n<li>A neural network base encoder  𝑓(.) .</li>\n<li>A neural network projection head  𝑔(.) </li>\n<li>A contrastive loss function.</li>\n</ul>\n<p><strong>Stochastic Augmentation</strong></p>\n<p>The authors suggest that a strong data augmentation is useful for unsupervised learning. <br>\nThe following augmentation are suggested by the authors:</p>\n<ul>\n<li>Random Crop with Resize</li>\n<li>Random Horizontal Flip with 50% probability</li>\n<li>Random Color Distortion</li>\n<li>Random Color Jitter with 80% probability</li>\n<li>Random Color Drop with 20% probability</li>\n<li>Random Gaussian Blur with 50% probability</li>\n</ul>\n<p><img src=\"https://i.ibb.co/xDnJZZp/simclr-random-transformation-function.gif\" alt=\"\"></p>\n<p>The data pipeline does not take an image and output a single augmented view, but on the contrary, outputs two randomly augmented views of the original image.</p>\n<p><strong>The SimCLR Model</strong><br>\nThe data pipeline outputs two augmented views of an image. The views go into a neural network encoder  𝑓(.)  that gives us the corresponding representation of the augmented views.</p>\n<p><strong>Our objective is to maximise the similarity quotient of the two distinct learned representations.</strong></p>\n<p>The idea here is to force the model to learn a general representation of an object from two distinct augmented views of it.<br>\nThe intuition is quite similar to viewing an object from different perspectives and gaining a better understanding.</p>\n<p>The authors <strong>do not put constraint</strong> on the encoder model. </p>\n<p><strong>Contrastive Loss</strong></p>\n<p>We use the projected vectors and run cosine similarity function to check how similar they are.<br>\nWe run the cosine similarity on both the positive and negative pairs. After we have the similarity matrix we apply softmax on to it to get the probability distribution of the entire model.</p>\n<p><strong>Our objective is</strong> to tune the parameters so that the softmax distribution is peaked on the positive pair.</p>\n<p><strong>Code:</strong> You can find a really good implementation in a Kaggle notebook of SimCLR <a href=\"https://www.kaggle.com/code/aritrag/simclr\" target=\"_blank\">here</a></p>\n<hr>\n<h3>SwAV</h3>\n<h5>Unsupervised Learning of Visual Features by Contrasting Cluster Assignments</h5>\n<blockquote>\n  <ul>\n  <li><a href=\"https://www.kaggle.com/code/ayuraj/v2-self-supervised-pretraining-with-swav/notebook\" target=\"_blank\">Kaggle Notebook</a></li>\n  <li><a href=\"https://arxiv.org/abs/2006.09882\" target=\"_blank\">Paper</a> </li>\n  <li><a href=\"https://github.com/facebookresearch/swav\" target=\"_blank\">PyTorch Implementation</a></li>\n  <li><a href=\"https://github.com/ayulockin/SwAV-TF\" target=\"_blank\">TensorFlow Implementation</a></li>\n  </ul>\n</blockquote>\n<p>Unsupervised visual representation learning is progressing at an exceptionally fast pace. Most of the modern training frameworks (SimCLR, BYOL, SwAV) in this area make use of a self-supervised model pre-trained with some contrastive learning objective. Saying these frameworks perform great w.r.t supervised model pre-training would be an understatement, as evident from the figure below -</p>\n<h2>How can SwAV be helpful in this competition?</h2>\n<p>​</p>\n<h2>What's SwAV?</h2>\n<p>​<br>\n<img src=\"https://i.ibb.co/JBgfJx0/download-11.png\" alt=\"image.png\"><br>\n​<br>\nThe authors of this paper investigated a question:<br>\n​</p>\n<blockquote>\n  <p><strong>Can we learn a meaningful metric that reflects apparent similarity among instances via pure discriminative learning?</strong><br>\n  ​</p>\n</blockquote>\n<p>To answer this, they devised a novel unsupervised feature learning algorithm called instance-level discrimination. Here each image and its transformations/views are treated as two separate instances. Each image instance is treated as a separate class. The aim is to learn an embedding, mapping $x$ (image) to $v$ (feature) such that semantically similar instances(images) are closer in the embedding space.<br>\n​</p>\n<p><strong>Code:</strong> You can find a really good implementation in a Kaggle notebook of SwAV <a href=\"https://www.kaggle.com/code/ayuraj/v2-self-supervised-pretraining-with-swav/notebook\" target=\"_blank\">here</a></p>\n<hr>\n<h3>BYOL</h3>\n<h5>Bootstrap Your Own Latent</h5>\n<blockquote>\n  <ul>\n  <li><a href=\"https://www.kaggle.com/code/nikhilpandey360/contrastive-learning-using-byol\" target=\"_blank\">Kaggle Notebook</a></li>\n  <li><a href=\"https://arxiv.org/abs/2006.07733\" target=\"_blank\">Paper</a></li>\n  <li><a href=\"https://github.com/lucidrains/byol-pytorch\" target=\"_blank\">Official Implementation</a></li>\n  </ul>\n</blockquote>\n<p><strong>Main Idea</strong></p>\n<p><img src=\"https://i.ibb.co/8nCh0mv/Selection-1110.png\" alt=\"\"></p>\n<p>BYOL could be summarized in the following 5 straightforward steps.</p>\n<ul>\n<li>Given an input image <code>x</code>, two views of the same image <code>v</code> and <code>v</code> are generated by applying two random augmentations to <code>x</code>.</li>\n<li>Given <code>v</code> and <code>v</code> to online and target encoders in order, vector representations <code>y'_θ</code> and <code>y'_ϵ</code> are obtained.</li>\n<li>Now, these representations are projected to another subspace z. These projected representations are indicated by <code>z_θ</code> and <code>z’_ϵ</code> in the image below.</li>\n<li>Since the target network is the slow moving average of the online network, the online representations should be predictive of the target representations, i.e. <code>z_θ</code> should predict <code>z’_ϵ</code> and hence another predictor(<code>q_θ</code>) is put on top of <code>z_θ</code>.</li>\n<li>Contrastive loss is reduced between &lt;<code>q_θ(z_θ)</code>, <code>z’_ϵ</code>&gt;.</li>\n</ul>\n<p><strong>Implementation</strong></p>\n<p><strong>Augmentations</strong></p>\n<p>To make oue implementation easy we use the following set of augmentations are used.</p>\n<pre><code> torchvision  transforms  tfms\nbyol_tfms = tfms.Compose([\n    tfms.RandomResizedCrop(size=, scale=(, )),\n    tfms.RandomHorizontalFlip(),\n    tfms.ToPILImage(),\n    tfms.RandomApply([\n            tfms.ColorJitter(, , , )\n    ], p=),\n   tfms.RandomGrayscale(p=),\n   tfms.ToTensor()\n])\n</code></pre>\n<p>We then set up a pipeline to run our models through for training with contrastive learning. <br>\nBelow PyTorch snippet implements the an encoder based BYOL network, but it could also be used in conjunction with any arbitrary encoder network such as VGG, InceptionNet, etc. without any significant change.</p>\n<p><strong>BYOL Module</strong></p>\n<pre><code> (nn.Module):\n     ():\n        ().__init__()\n        self.base_ema = base_target_ema        \n         backbone  :\n            backbone = models.resnet50(pretrained=)\n            backbone.output_dim = backbone.fc.in_features\n            backbone.fc = torch.nn.Identity()\n        projector = MLPHead(in_dim=backbone.output_dim)       \n        self.online_encoder = nn.Sequential(backbone, projector)        \n        self.target_encoder = copy.deepcopy(self.online_encoder)\n        self.online_predictor = MLPHead(in_dim=,hidden_size=, projection_size=)                 \n\n\n     ():        \n        tau = - (( - self.base_ema)* (cos(pi*global_step/max_steps)+)/)         \n         online, target  (self.online_encoder.parameters(), self.target_encoder.parameters()):\n            target.data = tau * target.data + ( - tau) * online.data     \n\n     ():        \n        z1 = self.online_encoder(x1)\n        z2 = self.online_encoder(x2)        \n        q1 = self.online_predictor(z1)\n        q2 = self.online_predictor(z2)        \n         torch.no_grad():\n            z1_t = self.target_encoder(x1)\n            z2_t = self.target_encoder(x2)       \n        loss = loss_fn(q1, q2, z1_t, z2_t)        \n         loss\n</code></pre>\n<p><strong>Training Loop</strong></p>\n<p>You then can use a simple training loop to train your model.</p>\n<pre><code> epoch  global_progress:\n    model.train()     \n\n     idx, (image, label)  (local_progress):\n        image = image.to(device)\n        aug_image = train_transform(image)\n\n        model.zero_grad()\n        loss = model.forward(image.to(device, non_blocking=), aug_image.to(device, non_blocking=))       \n        loss.backward()\n\n        optimizer.step()\n        model.update_moving_average(epoch, epochs)\n\n        scheduler.step()                             \n</code></pre>\n<pre><code> \n</code></pre>\n<p><strong>Code:</strong> You can find a good implementation in a Kaggle notebook of BYOL <a href=\"https://www.kaggle.com/code/nikhilpandey360/contrastive-learning-using-byol\" target=\"_blank\">here</a></p>",
  "messages": [
    {
      "id": 2076194,
      "postDate": "2022-12-26T08:02:41.057Z",
      "content": "<h4>Self-supervision</h4>\n<p><strong>Should we try it?</strong></p>\n<p>Many fellow Kagglers are using external datasets, some of them are unlabeled (or labeled for a different task).<br>\nSelf supervision methods can help capture information from unlabeled datasets. </p>\n<blockquote>\n  <p><strong>Also:</strong> They are a really good \"secret weapon\" to have in your back pocket since if they are implemented in the correct way they can \"just work\" all the time and boost the performance of every model <br>\n  <strong>(Like we treat MLM in NLP nowadays because huggingface did all the heavy lifting for us)</strong></p>\n</blockquote>\n<hr>\n<p><strong>Self Supervision Baselines</strong></p>\n<p>There had been many self-supervision papers released recently, today we will focus on the following: </p>\n<ul>\n<li><strong>BYOL:</strong> <a href=\"https://arxiv.org/abs/2006.07733\" target=\"_blank\">Bootstrap Your Own Latent</a> - </li>\n<li><strong>SwAV﻿:</strong> <a href=\"https://arxiv.org/abs/2006.09882\" target=\"_blank\">Unsupervised Learning of Visual Features by Contrasting Cluster Assignments</a></li>\n<li><strong>SimCLR:</strong> <a href=\"https://arxiv.org/abs/2002.05709\" target=\"_blank\">A Simple Framework for Contrastive Learning of Visual Representations﻿</a></li>\n</ul>\n<hr>\n<h3>SimCLR</h3>\n<h5>A Simple Framework for Contrastive Learning of Visual Representations﻿</h5>\n<blockquote>\n  <ul>\n  <li><a href=\"https://www.kaggle.com/code/aritrag/simclr\" target=\"_blank\">Kaggle Notebook</a></li>\n  <li><a href=\"https://arxiv.org/abs/2002.05709\" target=\"_blank\">Paper</a></li>\n  <li><a href=\"https://github.com/google-research/simclr\" target=\"_blank\">Official Implementation</a></li>\n  </ul>\n</blockquote>\n<p>A Simple framework for contrastive learning of visual representation</p>\n<p><img src=\"https://i.ibb.co/dMFvndK/image4.gif\" alt=\"\"></p>\n<p>SimCLR provides a great platform (and an easy one too) to help achieve a good representation out of unlabelled images. <br>\nThe itutions and conjectures provided by the paper comes quite naturally to one's mind. The ease of the concepts will be portrayed in the kernel along with some comments on the same.</p>\n<p><strong>SimCLR is based out of the following simplified modules:</strong></p>\n<ul>\n<li>A stochastic data augmentation module.</li>\n<li>A neural network base encoder  𝑓(.) .</li>\n<li>A neural network projection head  𝑔(.) </li>\n<li>A contrastive loss function.</li>\n</ul>\n<p><strong>Stochastic Augmentation</strong></p>\n<p>The authors suggest that a strong data augmentation is useful for unsupervised learning. <br>\nThe following augmentation are suggested by the authors:</p>\n<ul>\n<li>Random Crop with Resize</li>\n<li>Random Horizontal Flip with 50% probability</li>\n<li>Random Color Distortion</li>\n<li>Random Color Jitter with 80% probability</li>\n<li>Random Color Drop with 20% probability</li>\n<li>Random Gaussian Blur with 50% probability</li>\n</ul>\n<p><img src=\"https://i.ibb.co/xDnJZZp/simclr-random-transformation-function.gif\" alt=\"\"></p>\n<p>The data pipeline does not take an image and output a single augmented view, but on the contrary, outputs two randomly augmented views of the original image.</p>\n<p><strong>The SimCLR Model</strong><br>\nThe data pipeline outputs two augmented views of an image. The views go into a neural network encoder  𝑓(.)  that gives us the corresponding representation of the augmented views.</p>\n<p><strong>Our objective is to maximise the similarity quotient of the two distinct learned representations.</strong></p>\n<p>The idea here is to force the model to learn a general representation of an object from two distinct augmented views of it.<br>\nThe intuition is quite similar to viewing an object from different perspectives and gaining a better understanding.</p>\n<p>The authors <strong>do not put constraint</strong> on the encoder model. </p>\n<p><strong>Contrastive Loss</strong></p>\n<p>We use the projected vectors and run cosine similarity function to check how similar they are.<br>\nWe run the cosine similarity on both the positive and negative pairs. After we have the similarity matrix we apply softmax on to it to get the probability distribution of the entire model.</p>\n<p><strong>Our objective is</strong> to tune the parameters so that the softmax distribution is peaked on the positive pair.</p>\n<p><strong>Code:</strong> You can find a really good implementation in a Kaggle notebook of SimCLR <a href=\"https://www.kaggle.com/code/aritrag/simclr\" target=\"_blank\">here</a></p>\n<hr>\n<h3>SwAV</h3>\n<h5>Unsupervised Learning of Visual Features by Contrasting Cluster Assignments</h5>\n<blockquote>\n  <ul>\n  <li><a href=\"https://www.kaggle.com/code/ayuraj/v2-self-supervised-pretraining-with-swav/notebook\" target=\"_blank\">Kaggle Notebook</a></li>\n  <li><a href=\"https://arxiv.org/abs/2006.09882\" target=\"_blank\">Paper</a> </li>\n  <li><a href=\"https://github.com/facebookresearch/swav\" target=\"_blank\">PyTorch Implementation</a></li>\n  <li><a href=\"https://github.com/ayulockin/SwAV-TF\" target=\"_blank\">TensorFlow Implementation</a></li>\n  </ul>\n</blockquote>\n<p>Unsupervised visual representation learning is progressing at an exceptionally fast pace. Most of the modern training frameworks (SimCLR, BYOL, SwAV) in this area make use of a self-supervised model pre-trained with some contrastive learning objective. Saying these frameworks perform great w.r.t supervised model pre-training would be an understatement, as evident from the figure below -</p>\n<h2>How can SwAV be helpful in this competition?</h2>\n<p>​</p>\n<h2>What's SwAV?</h2>\n<p>​<br>\n<img src=\"https://i.ibb.co/JBgfJx0/download-11.png\" alt=\"image.png\"><br>\n​<br>\nThe authors of this paper investigated a question:<br>\n​</p>\n<blockquote>\n  <p><strong>Can we learn a meaningful metric that reflects apparent similarity among instances via pure discriminative learning?</strong><br>\n  ​</p>\n</blockquote>\n<p>To answer this, they devised a novel unsupervised feature learning algorithm called instance-level discrimination. Here each image and its transformations/views are treated as two separate instances. Each image instance is treated as a separate class. The aim is to learn an embedding, mapping $x$ (image) to $v$ (feature) such that semantically similar instances(images) are closer in the embedding space.<br>\n​</p>\n<p><strong>Code:</strong> You can find a really good implementation in a Kaggle notebook of SwAV <a href=\"https://www.kaggle.com/code/ayuraj/v2-self-supervised-pretraining-with-swav/notebook\" target=\"_blank\">here</a></p>\n<hr>\n<h3>BYOL</h3>\n<h5>Bootstrap Your Own Latent</h5>\n<blockquote>\n  <ul>\n  <li><a href=\"https://www.kaggle.com/code/nikhilpandey360/contrastive-learning-using-byol\" target=\"_blank\">Kaggle Notebook</a></li>\n  <li><a href=\"https://arxiv.org/abs/2006.07733\" target=\"_blank\">Paper</a></li>\n  <li><a href=\"https://github.com/lucidrains/byol-pytorch\" target=\"_blank\">Official Implementation</a></li>\n  </ul>\n</blockquote>\n<p><strong>Main Idea</strong></p>\n<p><img src=\"https://i.ibb.co/8nCh0mv/Selection-1110.png\" alt=\"\"></p>\n<p>BYOL could be summarized in the following 5 straightforward steps.</p>\n<ul>\n<li>Given an input image <code>x</code>, two views of the same image <code>v</code> and <code>v</code> are generated by applying two random augmentations to <code>x</code>.</li>\n<li>Given <code>v</code> and <code>v</code> to online and target encoders in order, vector representations <code>y'_θ</code> and <code>y'_ϵ</code> are obtained.</li>\n<li>Now, these representations are projected to another subspace z. These projected representations are indicated by <code>z_θ</code> and <code>z’_ϵ</code> in the image below.</li>\n<li>Since the target network is the slow moving average of the online network, the online representations should be predictive of the target representations, i.e. <code>z_θ</code> should predict <code>z’_ϵ</code> and hence another predictor(<code>q_θ</code>) is put on top of <code>z_θ</code>.</li>\n<li>Contrastive loss is reduced between &lt;<code>q_θ(z_θ)</code>, <code>z’_ϵ</code>&gt;.</li>\n</ul>\n<p><strong>Implementation</strong></p>\n<p><strong>Augmentations</strong></p>\n<p>To make oue implementation easy we use the following set of augmentations are used.</p>\n<pre><code> torchvision  transforms  tfms\nbyol_tfms = tfms.Compose([\n    tfms.RandomResizedCrop(size=, scale=(, )),\n    tfms.RandomHorizontalFlip(),\n    tfms.ToPILImage(),\n    tfms.RandomApply([\n            tfms.ColorJitter(, , , )\n    ], p=),\n   tfms.RandomGrayscale(p=),\n   tfms.ToTensor()\n])\n</code></pre>\n<p>We then set up a pipeline to run our models through for training with contrastive learning. <br>\nBelow PyTorch snippet implements the an encoder based BYOL network, but it could also be used in conjunction with any arbitrary encoder network such as VGG, InceptionNet, etc. without any significant change.</p>\n<p><strong>BYOL Module</strong></p>\n<pre><code> (nn.Module):\n     ():\n        ().__init__()\n        self.base_ema = base_target_ema        \n         backbone  :\n            backbone = models.resnet50(pretrained=)\n            backbone.output_dim = backbone.fc.in_features\n            backbone.fc = torch.nn.Identity()\n        projector = MLPHead(in_dim=backbone.output_dim)       \n        self.online_encoder = nn.Sequential(backbone, projector)        \n        self.target_encoder = copy.deepcopy(self.online_encoder)\n        self.online_predictor = MLPHead(in_dim=,hidden_size=, projection_size=)                 \n\n\n     ():        \n        tau = - (( - self.base_ema)* (cos(pi*global_step/max_steps)+)/)         \n         online, target  (self.online_encoder.parameters(), self.target_encoder.parameters()):\n            target.data = tau * target.data + ( - tau) * online.data     \n\n     ():        \n        z1 = self.online_encoder(x1)\n        z2 = self.online_encoder(x2)        \n        q1 = self.online_predictor(z1)\n        q2 = self.online_predictor(z2)        \n         torch.no_grad():\n            z1_t = self.target_encoder(x1)\n            z2_t = self.target_encoder(x2)       \n        loss = loss_fn(q1, q2, z1_t, z2_t)        \n         loss\n</code></pre>\n<p><strong>Training Loop</strong></p>\n<p>You then can use a simple training loop to train your model.</p>\n<pre><code> epoch  global_progress:\n    model.train()     \n\n     idx, (image, label)  (local_progress):\n        image = image.to(device)\n        aug_image = train_transform(image)\n\n        model.zero_grad()\n        loss = model.forward(image.to(device, non_blocking=), aug_image.to(device, non_blocking=))       \n        loss.backward()\n\n        optimizer.step()\n        model.update_moving_average(epoch, epochs)\n\n        scheduler.step()                             \n</code></pre>\n<pre><code> \n</code></pre>\n<p><strong>Code:</strong> You can find a good implementation in a Kaggle notebook of BYOL <a href=\"https://www.kaggle.com/code/nikhilpandey360/contrastive-learning-using-byol\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "#### Self-supervision\n**Should we try it?**\n\nMany fellow Kagglers are using external datasets, some of them are unlabeled (or labeled for a different task).\nSelf supervision methods can help capture information from unlabeled datasets. \n\n> **Also:** They are a really good \"secret weapon\" to have in your back pocket since if they are implemented in the correct way they can \"just work\" all the time and boost the performance of every model \n> **(Like we treat MLM in NLP nowadays because huggingface did all the heavy lifting for us)**\n\n\n_____\n\n\n**Self Supervision Baselines**\n\nThere had been many self-supervision papers released recently, today we will focus on the following: \n\n- **BYOL:** [Bootstrap Your Own Latent](https://arxiv.org/abs/2006.07733) - \n- **SwAV﻿:** [Unsupervised Learning of Visual Features by Contrasting Cluster Assignments](https://arxiv.org/abs/2006.09882)\n- **SimCLR:** [A Simple Framework for Contrastive Learning of Visual Representations﻿](https://arxiv.org/abs/2002.05709)\n\n_____\n\n\n### SimCLR\n##### A Simple Framework for Contrastive Learning of Visual Representations﻿\n\n> - [Kaggle Notebook](https://www.kaggle.com/code/aritrag/simclr)\n> - [Paper](https://arxiv.org/abs/2002.05709)\n> - [Official Implementation](https://github.com/google-research/simclr)\n\nA Simple framework for contrastive learning of visual representation\n\n![](https://i.ibb.co/dMFvndK/image4.gif)\n\n\nSimCLR provides a great platform (and an easy one too) to help achieve a good representation out of unlabelled images. \nThe itutions and conjectures provided by the paper comes quite naturally to one's mind. The ease of the concepts will be portrayed in the kernel along with some comments on the same.\n\n\n**SimCLR is based out of the following simplified modules:**\n\n- A stochastic data augmentation module.\n- A neural network base encoder  𝑓(.) .\n- A neural network projection head  𝑔(.) \n- A contrastive loss function.\n\n**Stochastic Augmentation**\n\nThe authors suggest that a strong data augmentation is useful for unsupervised learning. \nThe following augmentation are suggested by the authors:\n\n- Random Crop with Resize\n- Random Horizontal Flip with 50% probability\n- Random Color Distortion\n- Random Color Jitter with 80% probability\n- Random Color Drop with 20% probability\n- Random Gaussian Blur with 50% probability\n\n![](https://i.ibb.co/xDnJZZp/simclr-random-transformation-function.gif)\n\n\n\nThe data pipeline does not take an image and output a single augmented view, but on the contrary, outputs two randomly augmented views of the original image.\n\n**The SimCLR Model**\nThe data pipeline outputs two augmented views of an image. The views go into a neural network encoder  𝑓(.)  that gives us the corresponding representation of the augmented views.\n\n**Our objective is to maximise the similarity quotient of the two distinct learned representations.**\n\nThe idea here is to force the model to learn a general representation of an object from two distinct augmented views of it.\nThe intuition is quite similar to viewing an object from different perspectives and gaining a better understanding.\n\nThe authors **do not put constraint** on the encoder model. \n\n\n**Contrastive Loss**\n\nWe use the projected vectors and run cosine similarity function to check how similar they are.\nWe run the cosine similarity on both the positive and negative pairs. After we have the similarity matrix we apply softmax on to it to get the probability distribution of the entire model.\n\n**Our objective is** to tune the parameters so that the softmax distribution is peaked on the positive pair.\n\n\n**Code:** You can find a really good implementation in a Kaggle notebook of SimCLR [here](https://www.kaggle.com/code/aritrag/simclr)\n\n\n_____\n\n\n\n### SwAV\n##### Unsupervised Learning of Visual Features by Contrasting Cluster Assignments\n\n> - [Kaggle Notebook](https://www.kaggle.com/code/ayuraj/v2-self-supervised-pretraining-with-swav/notebook)\n> - [Paper](https://arxiv.org/abs/2006.09882) \n> - [PyTorch Implementation](https://github.com/facebookresearch/swav)\n> - [TensorFlow Implementation](https://github.com/ayulockin/SwAV-TF)\n\n\nUnsupervised visual representation learning is progressing at an exceptionally fast pace. Most of the modern training frameworks (SimCLR, BYOL, SwAV) in this area make use of a self-supervised model pre-trained with some contrastive learning objective. Saying these frameworks perform great w.r.t supervised model pre-training would be an understatement, as evident from the figure below -\n## How can SwAV be helpful in this competition?\n​\n\n## What's SwAV?\n​\n![image.png](https://i.ibb.co/JBgfJx0/download-11.png)\n​\nThe authors of this paper investigated a question:\n​\n\n> **Can we learn a meaningful metric that reflects apparent similarity among instances via pure discriminative learning?**\n​\n\n\nTo answer this, they devised a novel unsupervised feature learning algorithm called instance-level discrimination. Here each image and its transformations/views are treated as two separate instances. Each image instance is treated as a separate class. The aim is to learn an embedding, mapping $x$ (image) to $v$ (feature) such that semantically similar instances(images) are closer in the embedding space.\n​\n\n**Code:** You can find a really good implementation in a Kaggle notebook of SwAV [here](https://www.kaggle.com/code/ayuraj/v2-self-supervised-pretraining-with-swav/notebook)\n\n\n_____\n\n\n\n### BYOL\n##### Bootstrap Your Own Latent\n\n> - [Kaggle Notebook](https://www.kaggle.com/code/nikhilpandey360/contrastive-learning-using-byol)\n> - [Paper](https://arxiv.org/abs/2006.07733)\n> - [Official Implementation](https://github.com/lucidrains/byol-pytorch)\n\n\n**Main Idea**\n\n\n![](https://i.ibb.co/8nCh0mv/Selection-1110.png)\n\n\nBYOL could be summarized in the following 5 straightforward steps.\n\n- Given an input image `x`, two views of the same image `v` and `v` are generated by applying two random augmentations to `x`.\n- Given `v` and `v` to online and target encoders in order, vector representations `y'_θ` and `y'_ϵ` are obtained.\n- Now, these representations are projected to another subspace z. These projected representations are indicated by `z_θ` and `z’_ϵ` in the image below.\n- Since the target network is the slow moving average of the online network, the online representations should be predictive of the target representations, i.e. `z_θ` should predict `z’_ϵ` and hence another predictor(`q_θ`) is put on top of `z_θ`.\n- Contrastive loss is reduced between <`q_θ(z_θ)`, `z’_ϵ`>.\n\n\n**Implementation**\n\n**Augmentations**\n\nTo make oue implementation easy we use the following set of augmentations are used.\n\n```python\nfrom torchvision import transforms as tfms\nbyol_tfms = tfms.Compose([\n    tfms.RandomResizedCrop(size=512, scale=(0.3, 1)),\n    tfms.RandomHorizontalFlip(),\n    tfms.ToPILImage(),\n    tfms.RandomApply([\n            tfms.ColorJitter(0.4, 0.4, 0.4, 0.1)\n    ], p=0.8),\n   tfms.RandomGrayscale(p=0.2),\n   tfms.ToTensor()\n])\n```\n\nWe then set up a pipeline to run our models through for training with contrastive learning. \nBelow PyTorch snippet implements the an encoder based BYOL network, but it could also be used in conjunction with any arbitrary encoder network such as VGG, InceptionNet, etc. without any significant change.\n\n**BYOL Module**\n\n\n```python\nclass BYOL(nn.Module):\n    def __init__(self, backbone=None,base_target_ema=0.996,**kwargs):\n        super().__init__()\n        self.base_ema = base_target_ema        \n        if backbone is None:\n            backbone = models.resnet50(pretrained=False)\n            backbone.output_dim = backbone.fc.in_features\n            backbone.fc = torch.nn.Identity()\n        projector = MLPHead(in_dim=backbone.output_dim)       \n        self.online_encoder = nn.Sequential(backbone, projector)        \n        self.target_encoder = copy.deepcopy(self.online_encoder)\n        self.online_predictor = MLPHead(in_dim=256,hidden_size=1024, projection_size=256)                 \n\n    @torch.no_grad()\n    def update_moving_average(self, global_step, max_steps):        \n        tau = 1- ((1 - self.base_ema)* (cos(pi*global_step/max_steps)+1)/2)         \n        for online, target in zip(self.online_encoder.parameters(), self.target_encoder.parameters()):\n            target.data = tau * target.data + (1 - tau) * online.data     \n    \n    def forward(self,x1,x2):        \n        z1 = self.online_encoder(x1)\n        z2 = self.online_encoder(x2)        \n        q1 = self.online_predictor(z1)\n        q2 = self.online_predictor(z2)        \n        with torch.no_grad():\n            z1_t = self.target_encoder(x1)\n            z2_t = self.target_encoder(x2)       \n        loss = loss_fn(q1, q2, z1_t, z2_t)        \n        return loss\n```\n\n\n**Training Loop**\n\nYou then can use a simple training loop to train your model.\n\n\n```python\nfor epoch in global_progress:\n    model.train()     \n    \n    for idx, (image, label) in enumerate(local_progress):\n        image = image.to(device)\n        aug_image = train_transform(image)\n \n        model.zero_grad()\n        loss = model.forward(image.to(device, non_blocking=True), aug_image.to(device, non_blocking=True))       \n        loss.backward()\n        \n        optimizer.step()\n        model.update_moving_average(epoch, epochs)\n        \n        scheduler.step()                             \n```     \n\n\n\n**Code:** You can find a good implementation in a Kaggle notebook of BYOL [here](https://www.kaggle.com/code/nikhilpandey360/contrastive-learning-using-byol)",
      "votes": 12
    },
    {
      "id": 2079078,
      "postDate": "2022-12-29T00:43:45.987Z",
      "content": "<p>Self-Supervised Deep Learning to Enhance Breast Cancer Detection on Screening Mammograph<br>\n<a href=\"https://arxiv.org/pdf/2203.08812.pdf\" target=\"_blank\">https://arxiv.org/pdf/2203.08812.pdf</a></p>\n<p><a href=\"https://ibb.co/fHc2mpS\"><img src=\"https://i.ibb.co/zZkrKR5/Selection-329.png\" alt=\"Selection-329\"></a><br>\n<a href=\"https://ibb.co/Jzsp58H\"><img src=\"https://i.ibb.co/P56WwJg/Selection-328.png\" alt=\"Selection-328\"></a></p>\n<hr>\n<p>fyi:</p>\n<p>A built in MIP (multiple instance pooling) network is facebook's PatchConvnet[1]</p>\n<p>\" We replace the final average pooling by an attention-based aggregation layer akin to a single transformer block, that weights how the patches are involved in the classification decision.\"</p>\n<p>[1] Augmenting Convolutional networks with attention-based aggregation <br>\n[2] <a href=\"https://github.com/facebookresearch/deit/blob/main/README_patchconvnet.md\" target=\"_blank\">https://github.com/facebookresearch/deit/blob/main/README_patchconvnet.md</a></p>\n<p><img src=\"https://img-blog.csdnimg.cn/img_convert/08ae270d6a9106aaa869b20964acd435.png\" alt=\"https://img-blog.csdnimg.cn/img_convert/08ae270d6a9106aaa869b20964acd435.png\"></p>",
      "rawMarkdown": "Self-Supervised Deep Learning to Enhance Breast Cancer Detection on Screening Mammograph\nhttps://arxiv.org/pdf/2203.08812.pdf\n\n<a href=\"https://ibb.co/fHc2mpS\"><img src=\"https://i.ibb.co/zZkrKR5/Selection-329.png\" alt=\"Selection-329\" border=\"0\"></a>\n<a href=\"https://ibb.co/Jzsp58H\"><img src=\"https://i.ibb.co/P56WwJg/Selection-328.png\" alt=\"Selection-328\" border=\"0\"></a>\n\n---\n\nfyi:\n\nA built in MIP (multiple instance pooling) network is facebook's PatchConvnet[1]\n\n\" We replace the final average pooling by an attention-based aggregation layer akin to a single transformer block, that weights how the patches are involved in the classification decision.\"\n\n[1] Augmenting Convolutional networks with attention-based aggregation \n[2] https://github.com/facebookresearch/deit/blob/main/README_patchconvnet.md\n\n\n![https://img-blog.csdnimg.cn/img_convert/08ae270d6a9106aaa869b20964acd435.png](https://img-blog.csdnimg.cn/img_convert/08ae270d6a9106aaa869b20964acd435.png)",
      "votes": 3
    },
    {
      "id": 2076861,
      "postDate": "2022-12-27T00:02:27.693Z",
      "content": "<p>Great thread! I also feel self-supervision is very much worth exploring in this competition.</p>\n<p>The dataset is unbalanced, there are not that many examples in total and we are using models that have been pretrained on Imagenet, images that are vastly different from the DICOMs we get to work with in this competition.</p>\n<p>Really nice write-up of the available techniques!</p>",
      "rawMarkdown": "Great thread! I also feel self-supervision is very much worth exploring in this competition.\n\nThe dataset is unbalanced, there are not that many examples in total and we are using models that have been pretrained on Imagenet, images that are vastly different from the DICOMs we get to work with in this competition.\n\nReally nice write-up of the available techniques!",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 2079078,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-12-29T00:43:45.987000",
      "content": "<p>Self-Supervised Deep Learning to Enhance Breast Cancer Detection on Screening Mammograph<br>\n<a href=\"https://arxiv.org/pdf/2203.08812.pdf\" target=\"_blank\">https://arxiv.org/pdf/2203.08812.pdf</a></p>\n<p><a href=\"https://ibb.co/fHc2mpS\"><img src=\"https://i.ibb.co/zZkrKR5/Selection-329.png\" alt=\"Selection-329\"></a><br>\n<a href=\"https://ibb.co/Jzsp58H\"><img src=\"https://i.ibb.co/P56WwJg/Selection-328.png\" alt=\"Selection-328\"></a></p>\n<hr>\n<p>fyi:</p>\n<p>A built in MIP (multiple instance pooling) network is facebook's PatchConvnet[1]</p>\n<p>\" We replace the final average pooling by an attention-based aggregation layer akin to a single transformer block, that weights how the patches are involved in the classification decision.\"</p>\n<p>[1] Augmenting Convolutional networks with attention-based aggregation <br>\n[2] <a href=\"https://github.com/facebookresearch/deit/blob/main/README_patchconvnet.md\" target=\"_blank\">https://github.com/facebookresearch/deit/blob/main/README_patchconvnet.md</a></p>\n<p><img src=\"https://img-blog.csdnimg.cn/img_convert/08ae270d6a9106aaa869b20964acd435.png\" alt=\"https://img-blog.csdnimg.cn/img_convert/08ae270d6a9106aaa869b20964acd435.png\"></p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2076861,
      "author_name": "Radek Osmulski",
      "author_url": "",
      "post_date": "2022-12-27T00:02:27.693000",
      "content": "<p>Great thread! I also feel self-supervision is very much worth exploring in this competition.</p>\n<p>The dataset is unbalanced, there are not that many examples in total and we are using models that have been pretrained on Imagenet, images that are vastly different from the DICOMs we get to work with in this competition.</p>\n<p>Really nice write-up of the available techniques!</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2076194": "#### Self-supervision\n**Should we try it?**\n\nMany fellow Kagglers are using external datasets, some of them are unlabeled (or labeled for a different task).\nSelf supervision methods can help capture information from unlabeled datasets. \n\n> **Also:** They are a really good \"secret weapon\" to have in your back pocket since if they are implemented in the correct way they can \"just work\" all the time and boost the performance of every model \n> **(Like we treat MLM in NLP nowadays because huggingface did all the heavy lifting for us)**\n\n\n_____\n\n\n**Self Supervision Baselines**\n\nThere had been many self-supervision papers released recently, today we will focus on the following: \n\n- **BYOL:** [Bootstrap Your Own Latent](https://arxiv.org/abs/2006.07733) - \n- **SwAV﻿:** [Unsupervised Learning of Visual Features by Contrasting Cluster Assignments](https://arxiv.org/abs/2006.09882)\n- **SimCLR:** [A Simple Framework for Contrastive Learning of Visual Representations﻿](https://arxiv.org/abs/2002.05709)\n\n_____\n\n\n### SimCLR\n##### A Simple Framework for Contrastive Learning of Visual Representations﻿\n\n> - [Kaggle Notebook](https://www.kaggle.com/code/aritrag/simclr)\n> - [Paper](https://arxiv.org/abs/2002.05709)\n> - [Official Implementation](https://github.com/google-research/simclr)\n\nA Simple framework for contrastive learning of visual representation\n\n![](https://i.ibb.co/dMFvndK/image4.gif)\n\n\nSimCLR provides a great platform (and an easy one too) to help achieve a good representation out of unlabelled images. \nThe itutions and conjectures provided by the paper comes quite naturally to one's mind. The ease of the concepts will be portrayed in the kernel along with some comments on the same.\n\n\n**SimCLR is based out of the following simplified modules:**\n\n- A stochastic data augmentation module.\n- A neural network base encoder  𝑓(.) .\n- A neural network projection head  𝑔(.) \n- A contrastive loss function.\n\n**Stochastic Augmentation**\n\nThe authors suggest that a strong data augmentation is useful for unsupervised learning. \nThe following augmentation are suggested by the authors:\n\n- Random Crop with Resize\n- Random Horizontal Flip with 50% probability\n- Random Color Distortion\n- Random Color Jitter with 80% probability\n- Random Color Drop with 20% probability\n- Random Gaussian Blur with 50% probability\n\n![](https://i.ibb.co/xDnJZZp/simclr-random-transformation-function.gif)\n\n\n\nThe data pipeline does not take an image and output a single augmented view, but on the contrary, outputs two randomly augmented views of the original image.\n\n**The SimCLR Model**\nThe data pipeline outputs two augmented views of an image. The views go into a neural network encoder  𝑓(.)  that gives us the corresponding representation of the augmented views.\n\n**Our objective is to maximise the similarity quotient of the two distinct learned representations.**\n\nThe idea here is to force the model to learn a general representation of an object from two distinct augmented views of it.\nThe intuition is quite similar to viewing an object from different perspectives and gaining a better understanding.\n\nThe authors **do not put constraint** on the encoder model. \n\n\n**Contrastive Loss**\n\nWe use the projected vectors and run cosine similarity function to check how similar they are.\nWe run the cosine similarity on both the positive and negative pairs. After we have the similarity matrix we apply softmax on to it to get the probability distribution of the entire model.\n\n**Our objective is** to tune the parameters so that the softmax distribution is peaked on the positive pair.\n\n\n**Code:** You can find a really good implementation in a Kaggle notebook of SimCLR [here](https://www.kaggle.com/code/aritrag/simclr)\n\n\n_____\n\n\n\n### SwAV\n##### Unsupervised Learning of Visual Features by Contrasting Cluster Assignments\n\n> - [Kaggle Notebook](https://www.kaggle.com/code/ayuraj/v2-self-supervised-pretraining-with-swav/notebook)\n> - [Paper](https://arxiv.org/abs/2006.09882) \n> - [PyTorch Implementation](https://github.com/facebookresearch/swav)\n> - [TensorFlow Implementation](https://github.com/ayulockin/SwAV-TF)\n\n\nUnsupervised visual representation learning is progressing at an exceptionally fast pace. Most of the modern training frameworks (SimCLR, BYOL, SwAV) in this area make use of a self-supervised model pre-trained with some contrastive learning objective. Saying these frameworks perform great w.r.t supervised model pre-training would be an understatement, as evident from the figure below -\n## How can SwAV be helpful in this competition?\n​\n\n## What's SwAV?\n​\n![image.png](https://i.ibb.co/JBgfJx0/download-11.png)\n​\nThe authors of this paper investigated a question:\n​\n\n> **Can we learn a meaningful metric that reflects apparent similarity among instances via pure discriminative learning?**\n​\n\n\nTo answer this, they devised a novel unsupervised feature learning algorithm called instance-level discrimination. Here each image and its transformations/views are treated as two separate instances. Each image instance is treated as a separate class. The aim is to learn an embedding, mapping $x$ (image) to $v$ (feature) such that semantically similar instances(images) are closer in the embedding space.\n​\n\n**Code:** You can find a really good implementation in a Kaggle notebook of SwAV [here](https://www.kaggle.com/code/ayuraj/v2-self-supervised-pretraining-with-swav/notebook)\n\n\n_____\n\n\n\n### BYOL\n##### Bootstrap Your Own Latent\n\n> - [Kaggle Notebook](https://www.kaggle.com/code/nikhilpandey360/contrastive-learning-using-byol)\n> - [Paper](https://arxiv.org/abs/2006.07733)\n> - [Official Implementation](https://github.com/lucidrains/byol-pytorch)\n\n\n**Main Idea**\n\n\n![](https://i.ibb.co/8nCh0mv/Selection-1110.png)\n\n\nBYOL could be summarized in the following 5 straightforward steps.\n\n- Given an input image `x`, two views of the same image `v` and `v` are generated by applying two random augmentations to `x`.\n- Given `v` and `v` to online and target encoders in order, vector representations `y'_θ` and `y'_ϵ` are obtained.\n- Now, these representations are projected to another subspace z. These projected representations are indicated by `z_θ` and `z’_ϵ` in the image below.\n- Since the target network is the slow moving average of the online network, the online representations should be predictive of the target representations, i.e. `z_θ` should predict `z’_ϵ` and hence another predictor(`q_θ`) is put on top of `z_θ`.\n- Contrastive loss is reduced between <`q_θ(z_θ)`, `z’_ϵ`>.\n\n\n**Implementation**\n\n**Augmentations**\n\nTo make oue implementation easy we use the following set of augmentations are used.\n\n```python\nfrom torchvision import transforms as tfms\nbyol_tfms = tfms.Compose([\n    tfms.RandomResizedCrop(size=512, scale=(0.3, 1)),\n    tfms.RandomHorizontalFlip(),\n    tfms.ToPILImage(),\n    tfms.RandomApply([\n            tfms.ColorJitter(0.4, 0.4, 0.4, 0.1)\n    ], p=0.8),\n   tfms.RandomGrayscale(p=0.2),\n   tfms.ToTensor()\n])\n```\n\nWe then set up a pipeline to run our models through for training with contrastive learning. \nBelow PyTorch snippet implements the an encoder based BYOL network, but it could also be used in conjunction with any arbitrary encoder network such as VGG, InceptionNet, etc. without any significant change.\n\n**BYOL Module**\n\n\n```python\nclass BYOL(nn.Module):\n    def __init__(self, backbone=None,base_target_ema=0.996,**kwargs):\n        super().__init__()\n        self.base_ema = base_target_ema        \n        if backbone is None:\n            backbone = models.resnet50(pretrained=False)\n            backbone.output_dim = backbone.fc.in_features\n            backbone.fc = torch.nn.Identity()\n        projector = MLPHead(in_dim=backbone.output_dim)       \n        self.online_encoder = nn.Sequential(backbone, projector)        \n        self.target_encoder = copy.deepcopy(self.online_encoder)\n        self.online_predictor = MLPHead(in_dim=256,hidden_size=1024, projection_size=256)                 \n\n    @torch.no_grad()\n    def update_moving_average(self, global_step, max_steps):        \n        tau = 1- ((1 - self.base_ema)* (cos(pi*global_step/max_steps)+1)/2)         \n        for online, target in zip(self.online_encoder.parameters(), self.target_encoder.parameters()):\n            target.data = tau * target.data + (1 - tau) * online.data     \n    \n    def forward(self,x1,x2):        \n        z1 = self.online_encoder(x1)\n        z2 = self.online_encoder(x2)        \n        q1 = self.online_predictor(z1)\n        q2 = self.online_predictor(z2)        \n        with torch.no_grad():\n            z1_t = self.target_encoder(x1)\n            z2_t = self.target_encoder(x2)       \n        loss = loss_fn(q1, q2, z1_t, z2_t)        \n        return loss\n```\n\n\n**Training Loop**\n\nYou then can use a simple training loop to train your model.\n\n\n```python\nfor epoch in global_progress:\n    model.train()     \n    \n    for idx, (image, label) in enumerate(local_progress):\n        image = image.to(device)\n        aug_image = train_transform(image)\n \n        model.zero_grad()\n        loss = model.forward(image.to(device, non_blocking=True), aug_image.to(device, non_blocking=True))       \n        loss.backward()\n        \n        optimizer.step()\n        model.update_moving_average(epoch, epochs)\n        \n        scheduler.step()                             \n```     \n\n\n\n**Code:** You can find a good implementation in a Kaggle notebook of BYOL [here](https://www.kaggle.com/code/nikhilpandey360/contrastive-learning-using-byol)",
    "2079078": "Self-Supervised Deep Learning to Enhance Breast Cancer Detection on Screening Mammograph\nhttps://arxiv.org/pdf/2203.08812.pdf\n\n<a href=\"https://ibb.co/fHc2mpS\"><img src=\"https://i.ibb.co/zZkrKR5/Selection-329.png\" alt=\"Selection-329\" border=\"0\"></a>\n<a href=\"https://ibb.co/Jzsp58H\"><img src=\"https://i.ibb.co/P56WwJg/Selection-328.png\" alt=\"Selection-328\" border=\"0\"></a>\n\n---\n\nfyi:\n\nA built in MIP (multiple instance pooling) network is facebook's PatchConvnet[1]\n\n\" We replace the final average pooling by an attention-based aggregation layer akin to a single transformer block, that weights how the patches are involved in the classification decision.\"\n\n[1] Augmenting Convolutional networks with attention-based aggregation \n[2] https://github.com/facebookresearch/deit/blob/main/README_patchconvnet.md\n\n\n![https://img-blog.csdnimg.cn/img_convert/08ae270d6a9106aaa869b20964acd435.png](https://img-blog.csdnimg.cn/img_convert/08ae270d6a9106aaa869b20964acd435.png)",
    "2076861": "Great thread! I also feel self-supervision is very much worth exploring in this competition.\n\nThe dataset is unbalanced, there are not that many examples in total and we are using models that have been pretrained on Imagenet, images that are vastly different from the DICOMs we get to work with in this competition.\n\nReally nice write-up of the available techniques!"
  }
}