{
  "id": 169437,
  "title": "3rd public/20th private solution-segmentation + simple tiles and multiheaded attention",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/169437",
  "author_name": "Shujun",
  "post_date": "2020-07-23T20:32:43.241000",
  "votes": 21,
  "comment_count": 3,
  "views": 0,
  "content": "<p>First, thanks to the organizers <a href=\"/wouterbulten\">@wouterbulten</a> as i understand it is not easy to collect a dataset like this. Second, thanks to my teammates <a href=\"/rvslight\">@rvslight</a> <a href=\"/aksell7\">@aksell7</a>  <a href=\"/ruozha001\">@ruozha001</a>  who worked hard with me. Third, congrats to <a href=\"/iafoss\">@iafoss</a> for his solo gold and thanks to him for sharing the incredible tile idea and also all the participants who worked hard on this competition.</p>\n\n<p>We suffered in the shakeup, dropping from 3rd to 20th place, but i think our approach is quite interesting, and our selected sub was pretty good and balanced at both public (0.921) and private (0.927) with just 4 models.</p>\n\n<p>First I will briefly describe important details of my method using simple tiles which can generate a 0.927 single model single fold private score. My main idea is to keep things simple, apply attention, and use enough augmentation to avoid overfitting to label noise. My pure pytorch code is released on github at <a href=\"https://github.com/Shujun-He/PANDA\">https://github.com/Shujun-He/PANDA</a> (see folder layer1test4maxmeanwuncertainty for the pure pytorch pipeline and I will clearn up and update later). Later, I will detail the segmentation part of our solution. </p>\n\n<p>Our best private score (not selected) was achieved by ensembling 5 models (2 simple tiles and 3 segmented tiles) and using median avg (middle 3). Best simple tile (given by iafoss' tile function) setting was 36x256x256, and any number above 36 also works.</p>\n\n<h1>Model architecture</h1>\n\n<p>Since iafoss released his tile idea, I immediately thought of using attention so the network can learn importance of different tiles and make predictions based on the set of tiles for each WSI.  Here sometimes a full blown transformer encoder layer is used and sometimes just nn.MultiheadAttention + Mish activation. Also, resnext50 proved to be much better in this competition.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3355848%2F6845534ba574bd21bfa2006e0f7471b1%2Farch.PNG?generation=1595536578888614&amp;alt=media\" alt=\"\">. </p>\n\n<p>Mathematically, each tile becomes a feature vector after passed through the backbone, and the transformer encoder layer just operates on the set of feature vectors. 2D positional encoding can be added here but I did not think it was important based on reading about prostate cancer diagnosis. Here we usually used model=512 and nhead=8 same as the default setting of original transformer paper.</p>\n\n<p>I actually used multitasking learning by adding multiple attention classifiers on top the backbone:</p>\n\n<p>```python\nclass MultiheadAttentionClassifier(nn.Module):\n    def <strong>init</strong>(self,num_classes,out_features,ninp,nhead,dropout,attention_dropout=0.1):\n        super(MultiheadAttentionClassifier, self).<strong>init</strong>()\n        self.attention=nn.MultiheadAttention(ninp, nhead, dropout=attention_dropout)\n        self.classifier=nn.Linear(ninp*2,num_classes)\n        self.dropout=nn.Dropout(dropout)\n        self.mish=Mish()</p>\n\n<pre><code>def forward(self,x):\n    x=x.permute(1,0,2)\n    x,_=self.attention(x,x,x)\n    x=self.mish(x)\n    x=x.permute(1,0,2)\n    max_x,_=torch.max(x,dim=1)\n    x=torch.cat([torch.mean(x,dim=1),max_x],dim=-1)\n    x=self.dropout(x)\n    x=self.classifier(x)\n    return x\n</code></pre>\n\n<p>```</p>\n\n<p>This always resulted in much better CV convergence than just using isup grade, and lb was alway higher than CV so I stuck with the multitasking learning.</p>\n\n<h1>Augmentation</h1>\n\n<p>Augmentation wise I use cutout (replacing cutout region with just white pixels) 50% of\nthe time and the other 50% I change the gamma. The tiles always have 50% chance of being\nrotated/ flipped/transposed. In our N=64 runs, I used a new augmentation which I call\nwhiteout, where I simply turn some tiles white so the model can learn to be invariant to white\ntiles. </p>\n\n<p><code>python\ndef whiteout(tensor,n=6):\n    sample_shape=tensor.shape\n    to_drop=np.random.choice(tensor.shape[1],size=n,replace=False)\n    tensor[:,to_drop]=1\n    return tensor\n</code></p>\n\n<p>Later I found that after whiteout, even when using masked pooling (blocking white tiles), the model gives almost identical results, indicating that our model is invariant to white tiles.</p>\n\n<p>```python\nMultiheadAttentionClassifier with masked pooling and masked attention:\nclass MultiheadAttentionClassifier(nn.Module):\n    def <strong>init</strong>(self,num_classes,out_features,ninp,nhead,dropout,nlayers=1,attention_dropout=0.1):\n        super(MultiheadAttentionClassifier, self).<strong>init</strong>()\n        encoder_layers = nn.TransformerEncoderLayer(ninp, nhead, ninp*2, attention_dropout)\n        self.attention = nn.TransformerEncoder(encoder_layers, nlayers)\n        self.classifier=nn.Linear(ninp*2,num_classes)\n        self.dropout=nn.Dropout(dropout)</p>\n\n<pre><code>def forward(self,x,mask):\n    x=self.dropout(x)\n    x=x.permute(1,0,2)\n    src_key_padding_mask=mask==0\n    x=self.attention(x,src_key_padding_mask=src_key_padding_mask)\n    x=x.permute(1,0,2)\n    max_x,_=torch.max(x+src_key_padding_mask.unsqueeze(-1)*(-1e-9),dim=1)\n    mean_x=torch.sum(x*mask.unsqueeze(-1),dim=1)\n    tile_count=torch.sum(mask,dim=1).unsqueeze(-1)\n    mean_x=mean_x/tile_count\n    x=torch.cat([torch.mean(x,dim=1),max_x],dim=-1)\n    x=self.dropout(x)\n    x=self.classifier(x)\n    return x\n</code></pre>\n\n<p>```</p>\n\n<h1>Progressive upsampling</h1>\n\n<p>One thing that really sped up my training was the usage of progressive upsampling. Training is\nusually 45 epochs with first ten epochs on half resolution tiles (downsized with cv2.resize). At\n25 and 36 epochs, learning rate is reduced 10 times. This is a cool idea for people with limited computing power and for people who have a lot, it speed up training even more.</p>\n\n<h1>Segmentation model</h1>\n\n<p>To be updated. But to put it simply, we basically used masks on lowest resolution images to train a segmentation model distinguishing if the particular tile has cancer in it or not. Subsequently tiles were selected based on which ones are more likely to contain cancer based on trained segmentation model. This method should be better at predicting class 2,3,4,5, which was the case in CV at least. Somehow this method worked not so well in the private test set; however, ensembling this method with my models which use simple tiles still gave a boost.</p>\n\n<p>Combining the segmentation tiles with simple tiles worked well in public and also in private (just not as much as the boost in private given by denoising). Based on lb and cv, we thought that segmentation tiles would have better performance on 2,3,4,5 while simple tiles would be better at 0, 1, so the combination logically made sense.</p>\n\n<p>I had some worries that this method may be too biased towards predicting cancer, which is probably the reason it did not work well in private (judging from single model scores of segmentation tiles). Surprisingly, in private test set, when we made a mistake, where we used simple tiles instead of segmentation tiles on a model trained on segmentation tiles, we received a higher private lb in that submission than using segmentation tiles, which was of course not submitted since we identified that error.</p>\n\n<h1>Conclusion</h1>\n\n<p>In the end, we had multiple moments where we selected a 0.932 run, which would have resulted in a gold medal rather than a high silver. However, we changed it based on some reasoning that i still don't think is wrong. So just unlucky. </p>\n\n<p>About top solutions, I see most of them using some type of denoising method or just getting lucky based on some public kernels. Of course, using a large ensemble (~10 models) helps as well. What is really surprising to me is how denoising did not bring any recognizable improvement on public lb. I cannot help but think that there is some unintended difference between public and test set, because there is no reason denoising shouldn't work for public lb. In fact, I tried to do some denoising, but the results were not convincing and I stopped, which I do not consider a mistake, because there was no way to validate that anyone's denoising method was indeed working properly.</p>",
  "messages": [
    {
      "id": 942567,
      "postDate": "2020-07-23T20:32:43.240Z",
      "content": "<p>First, thanks to the organizers <a href=\"/wouterbulten\">@wouterbulten</a> as i understand it is not easy to collect a dataset like this. Second, thanks to my teammates <a href=\"/rvslight\">@rvslight</a> <a href=\"/aksell7\">@aksell7</a>  <a href=\"/ruozha001\">@ruozha001</a>  who worked hard with me. Third, congrats to <a href=\"/iafoss\">@iafoss</a> for his solo gold and thanks to him for sharing the incredible tile idea and also all the participants who worked hard on this competition.</p>\n\n<p>We suffered in the shakeup, dropping from 3rd to 20th place, but i think our approach is quite interesting, and our selected sub was pretty good and balanced at both public (0.921) and private (0.927) with just 4 models.</p>\n\n<p>First I will briefly describe important details of my method using simple tiles which can generate a 0.927 single model single fold private score. My main idea is to keep things simple, apply attention, and use enough augmentation to avoid overfitting to label noise. My pure pytorch code is released on github at <a href=\"https://github.com/Shujun-He/PANDA\">https://github.com/Shujun-He/PANDA</a> (see folder layer1test4maxmeanwuncertainty for the pure pytorch pipeline and I will clearn up and update later). Later, I will detail the segmentation part of our solution. </p>\n\n<p>Our best private score (not selected) was achieved by ensembling 5 models (2 simple tiles and 3 segmented tiles) and using median avg (middle 3). Best simple tile (given by iafoss' tile function) setting was 36x256x256, and any number above 36 also works.</p>\n\n<h1>Model architecture</h1>\n\n<p>Since iafoss released his tile idea, I immediately thought of using attention so the network can learn importance of different tiles and make predictions based on the set of tiles for each WSI.  Here sometimes a full blown transformer encoder layer is used and sometimes just nn.MultiheadAttention + Mish activation. Also, resnext50 proved to be much better in this competition.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3355848%2F6845534ba574bd21bfa2006e0f7471b1%2Farch.PNG?generation=1595536578888614&amp;alt=media\" alt=\"\">. </p>\n\n<p>Mathematically, each tile becomes a feature vector after passed through the backbone, and the transformer encoder layer just operates on the set of feature vectors. 2D positional encoding can be added here but I did not think it was important based on reading about prostate cancer diagnosis. Here we usually used model=512 and nhead=8 same as the default setting of original transformer paper.</p>\n\n<p>I actually used multitasking learning by adding multiple attention classifiers on top the backbone:</p>\n\n<p>```python\nclass MultiheadAttentionClassifier(nn.Module):\n    def <strong>init</strong>(self,num_classes,out_features,ninp,nhead,dropout,attention_dropout=0.1):\n        super(MultiheadAttentionClassifier, self).<strong>init</strong>()\n        self.attention=nn.MultiheadAttention(ninp, nhead, dropout=attention_dropout)\n        self.classifier=nn.Linear(ninp*2,num_classes)\n        self.dropout=nn.Dropout(dropout)\n        self.mish=Mish()</p>\n\n<pre><code>def forward(self,x):\n    x=x.permute(1,0,2)\n    x,_=self.attention(x,x,x)\n    x=self.mish(x)\n    x=x.permute(1,0,2)\n    max_x,_=torch.max(x,dim=1)\n    x=torch.cat([torch.mean(x,dim=1),max_x],dim=-1)\n    x=self.dropout(x)\n    x=self.classifier(x)\n    return x\n</code></pre>\n\n<p>```</p>\n\n<p>This always resulted in much better CV convergence than just using isup grade, and lb was alway higher than CV so I stuck with the multitasking learning.</p>\n\n<h1>Augmentation</h1>\n\n<p>Augmentation wise I use cutout (replacing cutout region with just white pixels) 50% of\nthe time and the other 50% I change the gamma. The tiles always have 50% chance of being\nrotated/ flipped/transposed. In our N=64 runs, I used a new augmentation which I call\nwhiteout, where I simply turn some tiles white so the model can learn to be invariant to white\ntiles. </p>\n\n<p><code>python\ndef whiteout(tensor,n=6):\n    sample_shape=tensor.shape\n    to_drop=np.random.choice(tensor.shape[1],size=n,replace=False)\n    tensor[:,to_drop]=1\n    return tensor\n</code></p>\n\n<p>Later I found that after whiteout, even when using masked pooling (blocking white tiles), the model gives almost identical results, indicating that our model is invariant to white tiles.</p>\n\n<p>```python\nMultiheadAttentionClassifier with masked pooling and masked attention:\nclass MultiheadAttentionClassifier(nn.Module):\n    def <strong>init</strong>(self,num_classes,out_features,ninp,nhead,dropout,nlayers=1,attention_dropout=0.1):\n        super(MultiheadAttentionClassifier, self).<strong>init</strong>()\n        encoder_layers = nn.TransformerEncoderLayer(ninp, nhead, ninp*2, attention_dropout)\n        self.attention = nn.TransformerEncoder(encoder_layers, nlayers)\n        self.classifier=nn.Linear(ninp*2,num_classes)\n        self.dropout=nn.Dropout(dropout)</p>\n\n<pre><code>def forward(self,x,mask):\n    x=self.dropout(x)\n    x=x.permute(1,0,2)\n    src_key_padding_mask=mask==0\n    x=self.attention(x,src_key_padding_mask=src_key_padding_mask)\n    x=x.permute(1,0,2)\n    max_x,_=torch.max(x+src_key_padding_mask.unsqueeze(-1)*(-1e-9),dim=1)\n    mean_x=torch.sum(x*mask.unsqueeze(-1),dim=1)\n    tile_count=torch.sum(mask,dim=1).unsqueeze(-1)\n    mean_x=mean_x/tile_count\n    x=torch.cat([torch.mean(x,dim=1),max_x],dim=-1)\n    x=self.dropout(x)\n    x=self.classifier(x)\n    return x\n</code></pre>\n\n<p>```</p>\n\n<h1>Progressive upsampling</h1>\n\n<p>One thing that really sped up my training was the usage of progressive upsampling. Training is\nusually 45 epochs with first ten epochs on half resolution tiles (downsized with cv2.resize). At\n25 and 36 epochs, learning rate is reduced 10 times. This is a cool idea for people with limited computing power and for people who have a lot, it speed up training even more.</p>\n\n<h1>Segmentation model</h1>\n\n<p>To be updated. But to put it simply, we basically used masks on lowest resolution images to train a segmentation model distinguishing if the particular tile has cancer in it or not. Subsequently tiles were selected based on which ones are more likely to contain cancer based on trained segmentation model. This method should be better at predicting class 2,3,4,5, which was the case in CV at least. Somehow this method worked not so well in the private test set; however, ensembling this method with my models which use simple tiles still gave a boost.</p>\n\n<p>Combining the segmentation tiles with simple tiles worked well in public and also in private (just not as much as the boost in private given by denoising). Based on lb and cv, we thought that segmentation tiles would have better performance on 2,3,4,5 while simple tiles would be better at 0, 1, so the combination logically made sense.</p>\n\n<p>I had some worries that this method may be too biased towards predicting cancer, which is probably the reason it did not work well in private (judging from single model scores of segmentation tiles). Surprisingly, in private test set, when we made a mistake, where we used simple tiles instead of segmentation tiles on a model trained on segmentation tiles, we received a higher private lb in that submission than using segmentation tiles, which was of course not submitted since we identified that error.</p>\n\n<h1>Conclusion</h1>\n\n<p>In the end, we had multiple moments where we selected a 0.932 run, which would have resulted in a gold medal rather than a high silver. However, we changed it based on some reasoning that i still don't think is wrong. So just unlucky. </p>\n\n<p>About top solutions, I see most of them using some type of denoising method or just getting lucky based on some public kernels. Of course, using a large ensemble (~10 models) helps as well. What is really surprising to me is how denoising did not bring any recognizable improvement on public lb. I cannot help but think that there is some unintended difference between public and test set, because there is no reason denoising shouldn't work for public lb. In fact, I tried to do some denoising, but the results were not convincing and I stopped, which I do not consider a mistake, because there was no way to validate that anyone's denoising method was indeed working properly.</p>",
      "rawMarkdown": "First, thanks to the organizers @wouterbulten as i understand it is not easy to collect a dataset like this. Second, thanks to my teammates @rvslight @aksell7  @ruozha001  who worked hard with me. Third, congrats to @iafoss for his solo gold and thanks to him for sharing the incredible tile idea and also all the participants who worked hard on this competition.\n\nWe suffered in the shakeup, dropping from 3rd to 20th place, but i think our approach is quite interesting, and our selected sub was pretty good and balanced at both public (0.921) and private (0.927) with just 4 models.\n\nFirst I will briefly describe important details of my method using simple tiles which can generate a 0.927 single model single fold private score. My main idea is to keep things simple, apply attention, and use enough augmentation to avoid overfitting to label noise. My pure pytorch code is released on github at https://github.com/Shujun-He/PANDA (see folder layer1test4maxmeanwuncertainty for the pure pytorch pipeline and I will clearn up and update later). Later, I will detail the segmentation part of our solution. \n\nOur best private score (not selected) was achieved by ensembling 5 models (2 simple tiles and 3 segmented tiles) and using median avg (middle 3). Best simple tile (given by iafoss' tile function) setting was 36x256x256, and any number above 36 also works.\n\n# Model architecture\n\nSince iafoss released his tile idea, I immediately thought of using attention so the network can learn importance of different tiles and make predictions based on the set of tiles for each WSI.  Here sometimes a full blown transformer encoder layer is used and sometimes just nn.MultiheadAttention + Mish activation. Also, resnext50 proved to be much better in this competition.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3355848%2F6845534ba574bd21bfa2006e0f7471b1%2Farch.PNG?generation=1595536578888614&amp;alt=media). \n\nMathematically, each tile becomes a feature vector after passed through the backbone, and the transformer encoder layer just operates on the set of feature vectors. 2D positional encoding can be added here but I did not think it was important based on reading about prostate cancer diagnosis. Here we usually used model=512 and nhead=8 same as the default setting of original transformer paper.\n\nI actually used multitasking learning by adding multiple attention classifiers on top the backbone:\n\n```python\nclass MultiheadAttentionClassifier(nn.Module):\n    def __init__(self,num_classes,out_features,ninp,nhead,dropout,attention_dropout=0.1):\n        super(MultiheadAttentionClassifier, self).__init__()\n        self.attention=nn.MultiheadAttention(ninp, nhead, dropout=attention_dropout)\n        self.classifier=nn.Linear(ninp*2,num_classes)\n        self.dropout=nn.Dropout(dropout)\n        self.mish=Mish()\n\n    def forward(self,x):\n        x=x.permute(1,0,2)\n        x,_=self.attention(x,x,x)\n        x=self.mish(x)\n        x=x.permute(1,0,2)\n        max_x,_=torch.max(x,dim=1)\n        x=torch.cat([torch.mean(x,dim=1),max_x],dim=-1)\n        x=self.dropout(x)\n        x=self.classifier(x)\n        return x\n```\n\n\nThis always resulted in much better CV convergence than just using isup grade, and lb was alway higher than CV so I stuck with the multitasking learning.\n\n\n# Augmentation\n\nAugmentation wise I use cutout (replacing cutout region with just white pixels) 50% of\nthe time and the other 50% I change the gamma. The tiles always have 50% chance of being\nrotated/ flipped/transposed. In our N=64 runs, I used a new augmentation which I call\nwhiteout, where I simply turn some tiles white so the model can learn to be invariant to white\ntiles. \n\n```python\ndef whiteout(tensor,n=6):\n    sample_shape=tensor.shape\n    to_drop=np.random.choice(tensor.shape[1],size=n,replace=False)\n    tensor[:,to_drop]=1\n    return tensor\n```\n\nLater I found that after whiteout, even when using masked pooling (blocking white tiles), the model gives almost identical results, indicating that our model is invariant to white tiles.\n\n```python\nMultiheadAttentionClassifier with masked pooling and masked attention:\nclass MultiheadAttentionClassifier(nn.Module):\n    def __init__(self,num_classes,out_features,ninp,nhead,dropout,nlayers=1,attention_dropout=0.1):\n        super(MultiheadAttentionClassifier, self).__init__()\n        encoder_layers = nn.TransformerEncoderLayer(ninp, nhead, ninp*2, attention_dropout)\n        self.attention = nn.TransformerEncoder(encoder_layers, nlayers)\n        self.classifier=nn.Linear(ninp*2,num_classes)\n        self.dropout=nn.Dropout(dropout)\n\n    def forward(self,x,mask):\n        x=self.dropout(x)\n        x=x.permute(1,0,2)\n        src_key_padding_mask=mask==0\n        x=self.attention(x,src_key_padding_mask=src_key_padding_mask)\n        x=x.permute(1,0,2)\n        max_x,_=torch.max(x+src_key_padding_mask.unsqueeze(-1)*(-1e-9),dim=1)\n        mean_x=torch.sum(x*mask.unsqueeze(-1),dim=1)\n        tile_count=torch.sum(mask,dim=1).unsqueeze(-1)\n        mean_x=mean_x/tile_count\n        x=torch.cat([torch.mean(x,dim=1),max_x],dim=-1)\n        x=self.dropout(x)\n        x=self.classifier(x)\n        return x\n```\n\n# Progressive upsampling\n\nOne thing that really sped up my training was the usage of progressive upsampling. Training is\nusually 45 epochs with first ten epochs on half resolution tiles (downsized with cv2.resize). At\n25 and 36 epochs, learning rate is reduced 10 times. This is a cool idea for people with limited computing power and for people who have a lot, it speed up training even more.\n\n# Segmentation model\n\nTo be updated. But to put it simply, we basically used masks on lowest resolution images to train a segmentation model distinguishing if the particular tile has cancer in it or not. Subsequently tiles were selected based on which ones are more likely to contain cancer based on trained segmentation model. This method should be better at predicting class 2,3,4,5, which was the case in CV at least. Somehow this method worked not so well in the private test set; however, ensembling this method with my models which use simple tiles still gave a boost.\n\nCombining the segmentation tiles with simple tiles worked well in public and also in private (just not as much as the boost in private given by denoising). Based on lb and cv, we thought that segmentation tiles would have better performance on 2,3,4,5 while simple tiles would be better at 0, 1, so the combination logically made sense.\n\nI had some worries that this method may be too biased towards predicting cancer, which is probably the reason it did not work well in private (judging from single model scores of segmentation tiles). Surprisingly, in private test set, when we made a mistake, where we used simple tiles instead of segmentation tiles on a model trained on segmentation tiles, we received a higher private lb in that submission than using segmentation tiles, which was of course not submitted since we identified that error.\n\n# Conclusion\n\nIn the end, we had multiple moments where we selected a 0.932 run, which would have resulted in a gold medal rather than a high silver. However, we changed it based on some reasoning that i still don't think is wrong. So just unlucky. \n\nAbout top solutions, I see most of them using some type of denoising method or just getting lucky based on some public kernels. Of course, using a large ensemble (~10 models) helps as well. What is really surprising to me is how denoising did not bring any recognizable improvement on public lb. I cannot help but think that there is some unintended difference between public and test set, because there is no reason denoising shouldn't work for public lb. In fact, I tried to do some denoising, but the results were not convincing and I stopped, which I do not consider a mistake, because there was no way to validate that anyone's denoising method was indeed working properly.\n",
      "votes": 21
    },
    {
      "id": 942762,
      "postDate": "2020-07-24T01:38:31.133Z",
      "content": "<p>Interesting and unique approach. Thank you for sharing it!</p>",
      "rawMarkdown": "Interesting and unique approach. Thank you for sharing it!",
      "replies": [
        {
          "id": 942790,
          "postDate": "2020-07-24T02:13:36.077Z",
          "content": "<p>Thank you and no problem!</p>",
          "rawMarkdown": "Thank you and no problem!",
          "votes": 1
        }
      ]
    },
    {
      "id": 3177968,
      "postDate": "2025-04-13T14:47:04.227Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 942762,
      "author_name": "عثمان",
      "author_url": "",
      "post_date": "2020-07-24T01:38:31.133000",
      "content": "<p>Interesting and unique approach. Thank you for sharing it!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 942790,
          "author_name": "Shujun",
          "author_url": "",
          "post_date": "2020-07-24T02:13:36.077000",
          "content": "<p>Thank you and no problem!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3177968,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-04-13T14:47:04.227000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "942567": "First, thanks to the organizers @wouterbulten as i understand it is not easy to collect a dataset like this. Second, thanks to my teammates @rvslight @aksell7  @ruozha001  who worked hard with me. Third, congrats to @iafoss for his solo gold and thanks to him for sharing the incredible tile idea and also all the participants who worked hard on this competition.\n\nWe suffered in the shakeup, dropping from 3rd to 20th place, but i think our approach is quite interesting, and our selected sub was pretty good and balanced at both public (0.921) and private (0.927) with just 4 models.\n\nFirst I will briefly describe important details of my method using simple tiles which can generate a 0.927 single model single fold private score. My main idea is to keep things simple, apply attention, and use enough augmentation to avoid overfitting to label noise. My pure pytorch code is released on github at https://github.com/Shujun-He/PANDA (see folder layer1test4maxmeanwuncertainty for the pure pytorch pipeline and I will clearn up and update later). Later, I will detail the segmentation part of our solution. \n\nOur best private score (not selected) was achieved by ensembling 5 models (2 simple tiles and 3 segmented tiles) and using median avg (middle 3). Best simple tile (given by iafoss' tile function) setting was 36x256x256, and any number above 36 also works.\n\n# Model architecture\n\nSince iafoss released his tile idea, I immediately thought of using attention so the network can learn importance of different tiles and make predictions based on the set of tiles for each WSI.  Here sometimes a full blown transformer encoder layer is used and sometimes just nn.MultiheadAttention + Mish activation. Also, resnext50 proved to be much better in this competition.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3355848%2F6845534ba574bd21bfa2006e0f7471b1%2Farch.PNG?generation=1595536578888614&amp;alt=media). \n\nMathematically, each tile becomes a feature vector after passed through the backbone, and the transformer encoder layer just operates on the set of feature vectors. 2D positional encoding can be added here but I did not think it was important based on reading about prostate cancer diagnosis. Here we usually used model=512 and nhead=8 same as the default setting of original transformer paper.\n\nI actually used multitasking learning by adding multiple attention classifiers on top the backbone:\n\n```python\nclass MultiheadAttentionClassifier(nn.Module):\n    def __init__(self,num_classes,out_features,ninp,nhead,dropout,attention_dropout=0.1):\n        super(MultiheadAttentionClassifier, self).__init__()\n        self.attention=nn.MultiheadAttention(ninp, nhead, dropout=attention_dropout)\n        self.classifier=nn.Linear(ninp*2,num_classes)\n        self.dropout=nn.Dropout(dropout)\n        self.mish=Mish()\n\n    def forward(self,x):\n        x=x.permute(1,0,2)\n        x,_=self.attention(x,x,x)\n        x=self.mish(x)\n        x=x.permute(1,0,2)\n        max_x,_=torch.max(x,dim=1)\n        x=torch.cat([torch.mean(x,dim=1),max_x],dim=-1)\n        x=self.dropout(x)\n        x=self.classifier(x)\n        return x\n```\n\n\nThis always resulted in much better CV convergence than just using isup grade, and lb was alway higher than CV so I stuck with the multitasking learning.\n\n\n# Augmentation\n\nAugmentation wise I use cutout (replacing cutout region with just white pixels) 50% of\nthe time and the other 50% I change the gamma. The tiles always have 50% chance of being\nrotated/ flipped/transposed. In our N=64 runs, I used a new augmentation which I call\nwhiteout, where I simply turn some tiles white so the model can learn to be invariant to white\ntiles. \n\n```python\ndef whiteout(tensor,n=6):\n    sample_shape=tensor.shape\n    to_drop=np.random.choice(tensor.shape[1],size=n,replace=False)\n    tensor[:,to_drop]=1\n    return tensor\n```\n\nLater I found that after whiteout, even when using masked pooling (blocking white tiles), the model gives almost identical results, indicating that our model is invariant to white tiles.\n\n```python\nMultiheadAttentionClassifier with masked pooling and masked attention:\nclass MultiheadAttentionClassifier(nn.Module):\n    def __init__(self,num_classes,out_features,ninp,nhead,dropout,nlayers=1,attention_dropout=0.1):\n        super(MultiheadAttentionClassifier, self).__init__()\n        encoder_layers = nn.TransformerEncoderLayer(ninp, nhead, ninp*2, attention_dropout)\n        self.attention = nn.TransformerEncoder(encoder_layers, nlayers)\n        self.classifier=nn.Linear(ninp*2,num_classes)\n        self.dropout=nn.Dropout(dropout)\n\n    def forward(self,x,mask):\n        x=self.dropout(x)\n        x=x.permute(1,0,2)\n        src_key_padding_mask=mask==0\n        x=self.attention(x,src_key_padding_mask=src_key_padding_mask)\n        x=x.permute(1,0,2)\n        max_x,_=torch.max(x+src_key_padding_mask.unsqueeze(-1)*(-1e-9),dim=1)\n        mean_x=torch.sum(x*mask.unsqueeze(-1),dim=1)\n        tile_count=torch.sum(mask,dim=1).unsqueeze(-1)\n        mean_x=mean_x/tile_count\n        x=torch.cat([torch.mean(x,dim=1),max_x],dim=-1)\n        x=self.dropout(x)\n        x=self.classifier(x)\n        return x\n```\n\n# Progressive upsampling\n\nOne thing that really sped up my training was the usage of progressive upsampling. Training is\nusually 45 epochs with first ten epochs on half resolution tiles (downsized with cv2.resize). At\n25 and 36 epochs, learning rate is reduced 10 times. This is a cool idea for people with limited computing power and for people who have a lot, it speed up training even more.\n\n# Segmentation model\n\nTo be updated. But to put it simply, we basically used masks on lowest resolution images to train a segmentation model distinguishing if the particular tile has cancer in it or not. Subsequently tiles were selected based on which ones are more likely to contain cancer based on trained segmentation model. This method should be better at predicting class 2,3,4,5, which was the case in CV at least. Somehow this method worked not so well in the private test set; however, ensembling this method with my models which use simple tiles still gave a boost.\n\nCombining the segmentation tiles with simple tiles worked well in public and also in private (just not as much as the boost in private given by denoising). Based on lb and cv, we thought that segmentation tiles would have better performance on 2,3,4,5 while simple tiles would be better at 0, 1, so the combination logically made sense.\n\nI had some worries that this method may be too biased towards predicting cancer, which is probably the reason it did not work well in private (judging from single model scores of segmentation tiles). Surprisingly, in private test set, when we made a mistake, where we used simple tiles instead of segmentation tiles on a model trained on segmentation tiles, we received a higher private lb in that submission than using segmentation tiles, which was of course not submitted since we identified that error.\n\n# Conclusion\n\nIn the end, we had multiple moments where we selected a 0.932 run, which would have resulted in a gold medal rather than a high silver. However, we changed it based on some reasoning that i still don't think is wrong. So just unlucky. \n\nAbout top solutions, I see most of them using some type of denoising method or just getting lucky based on some public kernels. Of course, using a large ensemble (~10 models) helps as well. What is really surprising to me is how denoising did not bring any recognizable improvement on public lb. I cannot help but think that there is some unintended difference between public and test set, because there is no reason denoising shouldn't work for public lb. In fact, I tried to do some denoising, but the results were not convincing and I stopped, which I do not consider a mistake, because there was no way to validate that anyone's denoising method was indeed working properly.\n",
    "942762": "Interesting and unique approach. Thank you for sharing it!",
    "3177968": ""
  }
}