{
  "id": 169222,
  "title": "Attention based sampling single fold only 4th public / 49 private ",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/169222",
  "author_name": "abzaliev",
  "post_date": "2020-07-23T08:59:02.782000",
  "votes": 14,
  "comment_count": 0,
  "views": 0,
  "content": "<p>First of all, many thanks to the organizers and to kaggle for such a interesting and hard problem. Great thanks to the HOSTKEY service provider  (<a href=\"https://www.hostkey.com/gpu-servers\">https://www.hostkey.com/gpu-servers</a>) for their grant and excellent support. Without 2 1080Ti's it would be very hard to compete. </p>\n\n<h1>Solution overview</h1>\n\n<p>My solution is based on attentions. It is a two-step approach:\n1. Train a simple resnet18 network with attention mechanism on all 128x128 tiles from medium resolution. The small size of the network turned out to be very important - first, because of the memory constraints, and second, more important, because of the label noise. \n2. Use obtained attentions to train a full b3 network on the most interesting 16x256x256 tiles to predict the isup_grade.</p>\n\n<h1>Attention network</h1>\n\n<p>First, I created a 256x256 tiles from the middle resolution. To cut down uninformative tiles, I use commonly reported in the literature blue ratio. This resulted in the following mask: </p>\n\n<p>![](<a href=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F393423%2F6b2b871462da89ded82d4c0b8e17a1ed%2Foriginal_img.png?generation=1595491956906370&amp;alt=media\">https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F393423%2F6b2b871462da89ded82d4c0b8e17a1ed%2Foriginal_img.png?generation=1595491956906370&amp;alt=media</a> =150x150) ![](<a href=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F393423%2F206975d1d7509145a27299f61e8fda9f%2Fmask.png?generation=1595491973353519&amp;alt=media\">https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F393423%2F206975d1d7509145a27299f61e8fda9f%2Fmask.png?generation=1595491973353519&amp;alt=media</a> =150x150)</p>\n\n<p>I selected only the tiles with at least 10% informative pixels on them. All those tiles were fed to the resnet18 with 1-dimentional 16-head attention layer before the head. The performance of this network on the public leaderboard  was around 0.84. After the training for each image in the training data I had the score from 0 to 1 for each tile. Here is the visualization, where brightness represents the importance of the tile:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F393423%2Fc3f6a1c33a679bf96334e809d92bb8b8%2Fattentions.png?generation=1595492360671689&amp;alt=media\" alt=\"\"></p>\n\n<p>Here is the code I used during the prediction for the attention network. The training one is almost the same with the head on top. </p>\n\n<p>class CoolTilesExtractor(nn.Module):</p>\n\n<pre><code>def __init__(self, n=6, n_filters=16, n_features=512):\n\n    super().__init__()\n\n    self.n_features = n_features\n\n    self.n_filters = n_filters\n\n    m = torch_models.resnet18(pretrained=False)\n\n    self.enc = nn.Sequential(*list(m.children())[:-1])\n\n    self.attention1d = nn.Conv1d(self.n_features, self.n_filters, kernel_size=(1), padding=(0))\n\n    self.softmax = torch.nn.Softmax(dim=2)\n\ndef forward(self, x): \n\n    res = self.enc(x)\n\n    # after the pooling split back to batches \n    res_with_batch_dim = res.view(1, len(res), self.n_features, 1)\n\n    # (2, 1, n_features, seq_len)\n    # we permute in order to correctly go through the attention\n    r_c_k = res_with_batch_dim.permute(0, 3, 2, 1)\n\n    att_raw = self.attention1d(r_c_k.squeeze(1))\n\n    # now normalize\n    after_softmax = self.softmax(att_raw)\n\n    return after_softmax, res\n</code></pre>\n\n<p>I formulated problem as regression with SmoothL1Loss. Batch size was 4, optimizer RAdam with 1e-4 lr with plain ReduceLROnPlateau. Trained for 20 epochs, best one was around 15. </p>\n\n<h1>Prediction network</h1>\n\n<p>After obtaining the attentions (averaged from all 16 heads) for each tile in the image, I used most informative tiles and combined them into single image, as majority of the competitors did. I used the following sampling during training: randomly select the tiles with the probability equal to the attention score. I was hoping that this way there will be less effect of noisy labels on the predictions. It turned out that it helped (at least on public leaderboard). During the prediction I just selected the tiles with the most attentions. </p>\n\n<p>The network I used was efficientnet b3, formulated as ordinal regression problem. Optimizer was RAdam with 1e-4 lr, batch size 4. No fancy schedule, just decreasing the lr by 2 every 5 epochs. Best epoch was 12. I also augumented the images during the prediction by randomly shuffling and rotating them. </p>\n\n<h1>Didn't work</h1>\n\n<p>This is a big section for each participant I suppose. Here is mine:\n- co-teaching. I put big efforts into making it work, but the results were not satisfying. \n- using slides from TCGA. Maybe the resolution was the issue, couldn't get any improvements.\n- run segmentation on PESO dataset and the use pre-trained encoder as a better starting point. Didn't help\n- Stain normalization. Also invested quite some time here, but nothing worked. \n- bigger networks for attentions - they started to overfit very quickly and picked up wrong patterns from mislabelled data. \n- manually constructed features. I implemented some features from here: <a href=\"https://github.com/hwanglab/tcga-prad-cslbp\">https://github.com/hwanglab/tcga-prad-cslbp</a>, but they didn't bring any improvement. Probably missed some details. </p>\n\n<p>I plan to release the code as soon as I will clean it. Once again, thanks to everyone and good luck next time!</p>",
  "messages": [
    {
      "id": 941510,
      "postDate": "2020-07-23T08:59:02.783Z",
      "content": "<p>First of all, many thanks to the organizers and to kaggle for such a interesting and hard problem. Great thanks to the HOSTKEY service provider  (<a href=\"https://www.hostkey.com/gpu-servers\">https://www.hostkey.com/gpu-servers</a>) for their grant and excellent support. Without 2 1080Ti's it would be very hard to compete. </p>\n\n<h1>Solution overview</h1>\n\n<p>My solution is based on attentions. It is a two-step approach:\n1. Train a simple resnet18 network with attention mechanism on all 128x128 tiles from medium resolution. The small size of the network turned out to be very important - first, because of the memory constraints, and second, more important, because of the label noise. \n2. Use obtained attentions to train a full b3 network on the most interesting 16x256x256 tiles to predict the isup_grade.</p>\n\n<h1>Attention network</h1>\n\n<p>First, I created a 256x256 tiles from the middle resolution. To cut down uninformative tiles, I use commonly reported in the literature blue ratio. This resulted in the following mask: </p>\n\n<p>![](<a href=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F393423%2F6b2b871462da89ded82d4c0b8e17a1ed%2Foriginal_img.png?generation=1595491956906370&amp;alt=media\">https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F393423%2F6b2b871462da89ded82d4c0b8e17a1ed%2Foriginal_img.png?generation=1595491956906370&amp;alt=media</a> =150x150) ![](<a href=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F393423%2F206975d1d7509145a27299f61e8fda9f%2Fmask.png?generation=1595491973353519&amp;alt=media\">https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F393423%2F206975d1d7509145a27299f61e8fda9f%2Fmask.png?generation=1595491973353519&amp;alt=media</a> =150x150)</p>\n\n<p>I selected only the tiles with at least 10% informative pixels on them. All those tiles were fed to the resnet18 with 1-dimentional 16-head attention layer before the head. The performance of this network on the public leaderboard  was around 0.84. After the training for each image in the training data I had the score from 0 to 1 for each tile. Here is the visualization, where brightness represents the importance of the tile:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F393423%2Fc3f6a1c33a679bf96334e809d92bb8b8%2Fattentions.png?generation=1595492360671689&amp;alt=media\" alt=\"\"></p>\n\n<p>Here is the code I used during the prediction for the attention network. The training one is almost the same with the head on top. </p>\n\n<p>class CoolTilesExtractor(nn.Module):</p>\n\n<pre><code>def __init__(self, n=6, n_filters=16, n_features=512):\n\n    super().__init__()\n\n    self.n_features = n_features\n\n    self.n_filters = n_filters\n\n    m = torch_models.resnet18(pretrained=False)\n\n    self.enc = nn.Sequential(*list(m.children())[:-1])\n\n    self.attention1d = nn.Conv1d(self.n_features, self.n_filters, kernel_size=(1), padding=(0))\n\n    self.softmax = torch.nn.Softmax(dim=2)\n\ndef forward(self, x): \n\n    res = self.enc(x)\n\n    # after the pooling split back to batches \n    res_with_batch_dim = res.view(1, len(res), self.n_features, 1)\n\n    # (2, 1, n_features, seq_len)\n    # we permute in order to correctly go through the attention\n    r_c_k = res_with_batch_dim.permute(0, 3, 2, 1)\n\n    att_raw = self.attention1d(r_c_k.squeeze(1))\n\n    # now normalize\n    after_softmax = self.softmax(att_raw)\n\n    return after_softmax, res\n</code></pre>\n\n<p>I formulated problem as regression with SmoothL1Loss. Batch size was 4, optimizer RAdam with 1e-4 lr with plain ReduceLROnPlateau. Trained for 20 epochs, best one was around 15. </p>\n\n<h1>Prediction network</h1>\n\n<p>After obtaining the attentions (averaged from all 16 heads) for each tile in the image, I used most informative tiles and combined them into single image, as majority of the competitors did. I used the following sampling during training: randomly select the tiles with the probability equal to the attention score. I was hoping that this way there will be less effect of noisy labels on the predictions. It turned out that it helped (at least on public leaderboard). During the prediction I just selected the tiles with the most attentions. </p>\n\n<p>The network I used was efficientnet b3, formulated as ordinal regression problem. Optimizer was RAdam with 1e-4 lr, batch size 4. No fancy schedule, just decreasing the lr by 2 every 5 epochs. Best epoch was 12. I also augumented the images during the prediction by randomly shuffling and rotating them. </p>\n\n<h1>Didn't work</h1>\n\n<p>This is a big section for each participant I suppose. Here is mine:\n- co-teaching. I put big efforts into making it work, but the results were not satisfying. \n- using slides from TCGA. Maybe the resolution was the issue, couldn't get any improvements.\n- run segmentation on PESO dataset and the use pre-trained encoder as a better starting point. Didn't help\n- Stain normalization. Also invested quite some time here, but nothing worked. \n- bigger networks for attentions - they started to overfit very quickly and picked up wrong patterns from mislabelled data. \n- manually constructed features. I implemented some features from here: <a href=\"https://github.com/hwanglab/tcga-prad-cslbp\">https://github.com/hwanglab/tcga-prad-cslbp</a>, but they didn't bring any improvement. Probably missed some details. </p>\n\n<p>I plan to release the code as soon as I will clean it. Once again, thanks to everyone and good luck next time!</p>",
      "rawMarkdown": "First of all, many thanks to the organizers and to kaggle for such a interesting and hard problem. Great thanks to the HOSTKEY service provider  (https://www.hostkey.com/gpu-servers) for their grant and excellent support. Without 2 1080Ti's it would be very hard to compete. \n\n# Solution overview\nMy solution is based on attentions. It is a two-step approach:\n1. Train a simple resnet18 network with attention mechanism on all 128x128 tiles from medium resolution. The small size of the network turned out to be very important - first, because of the memory constraints, and second, more important, because of the label noise. \n2. Use obtained attentions to train a full b3 network on the most interesting 16x256x256 tiles to predict the isup_grade.\n\n# Attention network\nFirst, I created a 256x256 tiles from the middle resolution. To cut down uninformative tiles, I use commonly reported in the literature blue ratio. This resulted in the following mask: \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F393423%2F6b2b871462da89ded82d4c0b8e17a1ed%2Foriginal_img.png?generation=1595491956906370&amp;alt=media =150x150) ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F393423%2F206975d1d7509145a27299f61e8fda9f%2Fmask.png?generation=1595491973353519&amp;alt=media =150x150)\n\nI selected only the tiles with at least 10% informative pixels on them. All those tiles were fed to the resnet18 with 1-dimentional 16-head attention layer before the head. The performance of this network on the public leaderboard  was around 0.84. After the training for each image in the training data I had the score from 0 to 1 for each tile. Here is the visualization, where brightness represents the importance of the tile:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F393423%2Fc3f6a1c33a679bf96334e809d92bb8b8%2Fattentions.png?generation=1595492360671689&amp;alt=media)\n\nHere is the code I used during the prediction for the attention network. The training one is almost the same with the head on top. \n\nclass CoolTilesExtractor(nn.Module):\n\n    def __init__(self, n=6, n_filters=16, n_features=512):\n\n        super().__init__()\n\n        self.n_features = n_features\n\n        self.n_filters = n_filters\n\n        m = torch_models.resnet18(pretrained=False)\n\n        self.enc = nn.Sequential(*list(m.children())[:-1])\n\n        self.attention1d = nn.Conv1d(self.n_features, self.n_filters, kernel_size=(1), padding=(0))\n\n        self.softmax = torch.nn.Softmax(dim=2)\n\n    def forward(self, x): \n    \n        res = self.enc(x)\n        \n        # after the pooling split back to batches \n        res_with_batch_dim = res.view(1, len(res), self.n_features, 1)\n        \n        # (2, 1, n_features, seq_len)\n        # we permute in order to correctly go through the attention\n        r_c_k = res_with_batch_dim.permute(0, 3, 2, 1)\n        \n        att_raw = self.attention1d(r_c_k.squeeze(1))\n        \n        # now normalize\n        after_softmax = self.softmax(att_raw)\n        \n        return after_softmax, res\n\nI formulated problem as regression with SmoothL1Loss. Batch size was 4, optimizer RAdam with 1e-4 lr with plain ReduceLROnPlateau. Trained for 20 epochs, best one was around 15. \n\n# Prediction network\nAfter obtaining the attentions (averaged from all 16 heads) for each tile in the image, I used most informative tiles and combined them into single image, as majority of the competitors did. I used the following sampling during training: randomly select the tiles with the probability equal to the attention score. I was hoping that this way there will be less effect of noisy labels on the predictions. It turned out that it helped (at least on public leaderboard). During the prediction I just selected the tiles with the most attentions. \n\nThe network I used was efficientnet b3, formulated as ordinal regression problem. Optimizer was RAdam with 1e-4 lr, batch size 4. No fancy schedule, just decreasing the lr by 2 every 5 epochs. Best epoch was 12. I also augumented the images during the prediction by randomly shuffling and rotating them. \n\n# Didn't work\nThis is a big section for each participant I suppose. Here is mine:\n- co-teaching. I put big efforts into making it work, but the results were not satisfying. \n- using slides from TCGA. Maybe the resolution was the issue, couldn't get any improvements.\n- run segmentation on PESO dataset and the use pre-trained encoder as a better starting point. Didn't help\n- Stain normalization. Also invested quite some time here, but nothing worked. \n- bigger networks for attentions - they started to overfit very quickly and picked up wrong patterns from mislabelled data. \n- manually constructed features. I implemented some features from here: https://github.com/hwanglab/tcga-prad-cslbp, but they didn't bring any improvement. Probably missed some details. \n\nI plan to release the code as soon as I will clean it. Once again, thanks to everyone and good luck next time!",
      "votes": 14
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "941510": "First of all, many thanks to the organizers and to kaggle for such a interesting and hard problem. Great thanks to the HOSTKEY service provider  (https://www.hostkey.com/gpu-servers) for their grant and excellent support. Without 2 1080Ti's it would be very hard to compete. \n\n# Solution overview\nMy solution is based on attentions. It is a two-step approach:\n1. Train a simple resnet18 network with attention mechanism on all 128x128 tiles from medium resolution. The small size of the network turned out to be very important - first, because of the memory constraints, and second, more important, because of the label noise. \n2. Use obtained attentions to train a full b3 network on the most interesting 16x256x256 tiles to predict the isup_grade.\n\n# Attention network\nFirst, I created a 256x256 tiles from the middle resolution. To cut down uninformative tiles, I use commonly reported in the literature blue ratio. This resulted in the following mask: \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F393423%2F6b2b871462da89ded82d4c0b8e17a1ed%2Foriginal_img.png?generation=1595491956906370&amp;alt=media =150x150) ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F393423%2F206975d1d7509145a27299f61e8fda9f%2Fmask.png?generation=1595491973353519&amp;alt=media =150x150)\n\nI selected only the tiles with at least 10% informative pixels on them. All those tiles were fed to the resnet18 with 1-dimentional 16-head attention layer before the head. The performance of this network on the public leaderboard  was around 0.84. After the training for each image in the training data I had the score from 0 to 1 for each tile. Here is the visualization, where brightness represents the importance of the tile:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F393423%2Fc3f6a1c33a679bf96334e809d92bb8b8%2Fattentions.png?generation=1595492360671689&amp;alt=media)\n\nHere is the code I used during the prediction for the attention network. The training one is almost the same with the head on top. \n\nclass CoolTilesExtractor(nn.Module):\n\n    def __init__(self, n=6, n_filters=16, n_features=512):\n\n        super().__init__()\n\n        self.n_features = n_features\n\n        self.n_filters = n_filters\n\n        m = torch_models.resnet18(pretrained=False)\n\n        self.enc = nn.Sequential(*list(m.children())[:-1])\n\n        self.attention1d = nn.Conv1d(self.n_features, self.n_filters, kernel_size=(1), padding=(0))\n\n        self.softmax = torch.nn.Softmax(dim=2)\n\n    def forward(self, x): \n    \n        res = self.enc(x)\n        \n        # after the pooling split back to batches \n        res_with_batch_dim = res.view(1, len(res), self.n_features, 1)\n        \n        # (2, 1, n_features, seq_len)\n        # we permute in order to correctly go through the attention\n        r_c_k = res_with_batch_dim.permute(0, 3, 2, 1)\n        \n        att_raw = self.attention1d(r_c_k.squeeze(1))\n        \n        # now normalize\n        after_softmax = self.softmax(att_raw)\n        \n        return after_softmax, res\n\nI formulated problem as regression with SmoothL1Loss. Batch size was 4, optimizer RAdam with 1e-4 lr with plain ReduceLROnPlateau. Trained for 20 epochs, best one was around 15. \n\n# Prediction network\nAfter obtaining the attentions (averaged from all 16 heads) for each tile in the image, I used most informative tiles and combined them into single image, as majority of the competitors did. I used the following sampling during training: randomly select the tiles with the probability equal to the attention score. I was hoping that this way there will be less effect of noisy labels on the predictions. It turned out that it helped (at least on public leaderboard). During the prediction I just selected the tiles with the most attentions. \n\nThe network I used was efficientnet b3, formulated as ordinal regression problem. Optimizer was RAdam with 1e-4 lr, batch size 4. No fancy schedule, just decreasing the lr by 2 every 5 epochs. Best epoch was 12. I also augumented the images during the prediction by randomly shuffling and rotating them. \n\n# Didn't work\nThis is a big section for each participant I suppose. Here is mine:\n- co-teaching. I put big efforts into making it work, but the results were not satisfying. \n- using slides from TCGA. Maybe the resolution was the issue, couldn't get any improvements.\n- run segmentation on PESO dataset and the use pre-trained encoder as a better starting point. Didn't help\n- Stain normalization. Also invested quite some time here, but nothing worked. \n- bigger networks for attentions - they started to overfit very quickly and picked up wrong patterns from mislabelled data. \n- manually constructed features. I implemented some features from here: https://github.com/hwanglab/tcga-prad-cslbp, but they didn't bring any improvement. Probably missed some details. \n\nI plan to release the code as soon as I will clean it. Once again, thanks to everyone and good luck next time!"
  }
}