{
  "id": 169108,
  "title": "2nd Place Solution [Save the Prostate]",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/169108",
  "author_name": "DrHB",
  "post_date": "2020-07-23T00:20:44.227000",
  "votes": 115,
  "comment_count": 48,
  "views": 0,
  "content": "<p>First of all, Thank you very much to organizers.\nSecond I would like to thanks my team. We had such a positive, encouraging working environment. Our team contribution generates most of the ideas which you will read below and are shared by members. </p>\n\n<h1>Simple Resnet34 (DrHB)</h1>\n\n<h1>Image Preprocessing</h1>\n\n<p>I used medium resolution, the only preprocessing I did was to remove the white background and store medium resolution on SSD drive: </p>\n\n<p>```</p>\n\n<h1>function taken from R Guo</h1>\n\n<p>def crop_white(image, value: int = 255):\n    assert image.shape[2] == 3\n    assert image.dtype == np.uint8\n    ys, = (image.min((1, 2)) &lt; value).nonzero()\n    xs, = (image.min(0).min(1) &lt; value).nonzero()\n    if len(xs) == 0 or len(ys) == 0:\n        return image\n    return image[ys.min():ys.max() + 1, xs.min():xs.max() + 1]</p>\n\n<p>```</p>\n\n<h1>Cleaning data</h1>\n\n<p>Like in APTOS competition, it was essential to clean images from pen marks, etc. I have used excellent work from this post: <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/151323\">https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/151323</a> This also reduced the gap between CV and LB</p>\n\n<h1>Image Augmenatiosn</h1>\n\n<p>Augmentation occurred at two levels.  (Slide and Tile): </p>\n\n<h3>1) Full slide</h3>\n\n<p>After the biopsy slide is open, we do random padding and applying one of the following transformations (similar to R Guo). </p>\n\n<p><code>\ndef get_transforms_train():\n    transforms=A.Compose(\n        [\n            A.ShiftScaleRotate(shift_limit=0.1, scale_limit=0.05, rotate_limit=10, border_mode=cv2.BORDER_CONSTANT, p=0.5,value=(255,255,255)),\n            A.OneOf([\n                A.Flip(p=0.5),\n                A.RandomRotate90(p=0.5),\n            ], p=0.3\n            )\n        ]\n    )\n    return transforms\n</code></p>\n\n<h3>2) Tile</h3>\n\n<p>For each tile I used standard fastai <code>GPU</code> augmentations: <code>rotate=(-10, 10</code>),<code>flip vertically (p=0.5)</code>. For the padding I used <code>reflection</code>  it gave a slight boost on CV </p>\n\n<h1>Model</h1>\n\n<p>I decided to use a very simple model resnet34 but throughout competitions ended up doing few modifications </p>\n\n<h3>1) Making square features:</h3>\n\n<p>The main idea is builds up on <a href=\"/iafoss\">@iafoss</a>. Aftter resnet enccoder we reshape features to look like a square in a following way: <code>x = x.view(x.shape[0], x.shape[1], x.shape[2]//int(np.sqrt(N)), -1)</code> . Here <code>N</code> represents Number of Tiles.  After this we pass all the features to <code>SqueezeExcite</code> Block</p>\n\n2) SqueezeExcite block\n\n<p>After reshaping features, we added 1 SE block to enable the network to learn features for individual slides based on tiles. </p>\n\n<p>experiment done by <a href=\"/cateek\">@cateek</a> \n```</p>\n\n<h1>code adopted</h1>\n\n<h1><a href=\"https://github.com/rwightman/pytorch-image-models/tree/master/timm/models\">https://github.com/rwightman/pytorch-image-models/tree/master/timm/models</a></h1>\n\n<p>def make_divisible(v, divisor=8, min_value=None):\n   min_value = min_value or divisor\n   new_v = max(min_value, int(v + divisor / 2) // divisor * divisor)\n   # Make sure that round down does not go down by more than 10%.\n   if new_v &lt; 0.9 * v:\n      new_v += divisor\n   return new_v\ndef sigmoid(x, inplace: bool = False):\n   return x.sigmoid_() if inplace else x.sigmoid()\nclass SqueezeExcite(nn.Module):\n   def <strong>init</strong>(self, in_chs, se_ratio=0.25, reduced_base_chs=None,\n             act_layer=nn.ReLU, gate_fn=sigmoid, divisor=1, **_):\n      super(SqueezeExcite, self).<strong>init</strong>()\n      self.gate_fn = gate_fn\n      reduced_chs = make_divisible((reduced_base_chs or in_chs) * se_ratio, divisor)\n      self.avg_pool = nn.AdaptiveAvgPool2d(1)\n      self.conv_reduce = nn.Conv2d(in_chs, reduced_chs, 1, bias=True)\n      self.act1 = act_layer(inplace=True)\n      self.conv_expand = nn.Conv2d(reduced_chs, in_chs, 1, bias=True)\n   def forward(self, x):\n      x_se = self.avg_pool(x)\n      x_se = self.conv_reduce(x_se)\n      x_se = self.act1(x_se)\n      x_se = self.conv_expand(x_se)\n      x = x * self.gate_fn(x_se)\n      return x\n```</p>\n\n<h3>3) Pooling Layer</h3>\n\n<p>Once the feature passed thru SqueezeExcite Layer, I did Normal pooling. Our experiment showed that the batch normalization layer was messing with the last layer's features, so we removed it and saw a slight jump on local cv. </p>\n\n<p><code>\nself.pool = nn.Sequential(AdaptiveConcatPool2d(),\n                          Flatten(),\n                          nn.Linear(2*nc,512),\n                          nn.ReLU(inplace=True),\n                          nn.Dropout(0.4),\n                          nn.Linear(512,7), \n</code></p>\n\n<h3>4) Final Head</h3>\n\n<p>I used two heads. One head was for classification second was for regression. I noticed that training with two looses makes training much smoother (with sigmoid trick below) and yields higher local CV (0.88 -&gt; 0.90). In the final prediction, I use output only for the regression head. </p>\n\n<p>One small modification that I did before calculating loss is that the regression head used sigmoid to scale outputs between (-1. 6.). This enables much smoother training without bumps and faster convergence.</p>\n\n<p>```</p>\n\n<h1>idea taken from fastai</h1>\n\n<p>def sigmoid_range(x, low, high):\n    return torch.sigmoid(x) * (high - low) + low \n```</p>\n\n<h1>Training</h1>\n\n<p>I trained in two phases. In the First phase was trained with 49 tiles and later finetuned with 81 tiles. Both phases were using standard one cycle.</p>\n\n<h1>Final Model</h1>\n\n<p>I trained 5 fold wich resulted on the CV of 0.911 and PB: 0.922.</p>\n\n<p>Our Best Ensemble was simple average. Of 4 models. </p>\n\n<p><code>\n@drhb resnet34 5 FOLD (CV -0.911) + \n<a href=\"/rguo97\">@rguo97</a>  5 FOLD (two stage attention model CV 0.92 ) + \n<a href=\"/xiejialun\">@xiejialun</a>  FOLD (EFNET) (CV 0.915-0.917) +\n<a href=\"/cateek\">@cateek</a>  Se 1 FOLD (CV -0.91) Final Standing\n</code></p>\n\n<p><a href=\"/xiejialun\">@xiejialun</a> <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169303\">https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169303</a>\n<a href=\"/rguo97\">@rguo97</a> <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169108#940504\">https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169108#940504</a></p>\n\n<p><code>LB: 0.914 PB: 0.937</code></p>",
  "messages": [
    {
      "id": 940434,
      "postDate": "2020-07-23T00:20:44.227Z",
      "content": "<p>First of all, Thank you very much to organizers.\nSecond I would like to thanks my team. We had such a positive, encouraging working environment. Our team contribution generates most of the ideas which you will read below and are shared by members. </p>\n\n<h1>Simple Resnet34 (DrHB)</h1>\n\n<h1>Image Preprocessing</h1>\n\n<p>I used medium resolution, the only preprocessing I did was to remove the white background and store medium resolution on SSD drive: </p>\n\n<p>```</p>\n\n<h1>function taken from R Guo</h1>\n\n<p>def crop_white(image, value: int = 255):\n    assert image.shape[2] == 3\n    assert image.dtype == np.uint8\n    ys, = (image.min((1, 2)) &lt; value).nonzero()\n    xs, = (image.min(0).min(1) &lt; value).nonzero()\n    if len(xs) == 0 or len(ys) == 0:\n        return image\n    return image[ys.min():ys.max() + 1, xs.min():xs.max() + 1]</p>\n\n<p>```</p>\n\n<h1>Cleaning data</h1>\n\n<p>Like in APTOS competition, it was essential to clean images from pen marks, etc. I have used excellent work from this post: <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/151323\">https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/151323</a> This also reduced the gap between CV and LB</p>\n\n<h1>Image Augmenatiosn</h1>\n\n<p>Augmentation occurred at two levels.  (Slide and Tile): </p>\n\n<h3>1) Full slide</h3>\n\n<p>After the biopsy slide is open, we do random padding and applying one of the following transformations (similar to R Guo). </p>\n\n<p><code>\ndef get_transforms_train():\n    transforms=A.Compose(\n        [\n            A.ShiftScaleRotate(shift_limit=0.1, scale_limit=0.05, rotate_limit=10, border_mode=cv2.BORDER_CONSTANT, p=0.5,value=(255,255,255)),\n            A.OneOf([\n                A.Flip(p=0.5),\n                A.RandomRotate90(p=0.5),\n            ], p=0.3\n            )\n        ]\n    )\n    return transforms\n</code></p>\n\n<h3>2) Tile</h3>\n\n<p>For each tile I used standard fastai <code>GPU</code> augmentations: <code>rotate=(-10, 10</code>),<code>flip vertically (p=0.5)</code>. For the padding I used <code>reflection</code>  it gave a slight boost on CV </p>\n\n<h1>Model</h1>\n\n<p>I decided to use a very simple model resnet34 but throughout competitions ended up doing few modifications </p>\n\n<h3>1) Making square features:</h3>\n\n<p>The main idea is builds up on <a href=\"/iafoss\">@iafoss</a>. Aftter resnet enccoder we reshape features to look like a square in a following way: <code>x = x.view(x.shape[0], x.shape[1], x.shape[2]//int(np.sqrt(N)), -1)</code> . Here <code>N</code> represents Number of Tiles.  After this we pass all the features to <code>SqueezeExcite</code> Block</p>\n\n2) SqueezeExcite block\n\n<p>After reshaping features, we added 1 SE block to enable the network to learn features for individual slides based on tiles. </p>\n\n<p>experiment done by <a href=\"/cateek\">@cateek</a> \n```</p>\n\n<h1>code adopted</h1>\n\n<h1><a href=\"https://github.com/rwightman/pytorch-image-models/tree/master/timm/models\">https://github.com/rwightman/pytorch-image-models/tree/master/timm/models</a></h1>\n\n<p>def make_divisible(v, divisor=8, min_value=None):\n   min_value = min_value or divisor\n   new_v = max(min_value, int(v + divisor / 2) // divisor * divisor)\n   # Make sure that round down does not go down by more than 10%.\n   if new_v &lt; 0.9 * v:\n      new_v += divisor\n   return new_v\ndef sigmoid(x, inplace: bool = False):\n   return x.sigmoid_() if inplace else x.sigmoid()\nclass SqueezeExcite(nn.Module):\n   def <strong>init</strong>(self, in_chs, se_ratio=0.25, reduced_base_chs=None,\n             act_layer=nn.ReLU, gate_fn=sigmoid, divisor=1, **_):\n      super(SqueezeExcite, self).<strong>init</strong>()\n      self.gate_fn = gate_fn\n      reduced_chs = make_divisible((reduced_base_chs or in_chs) * se_ratio, divisor)\n      self.avg_pool = nn.AdaptiveAvgPool2d(1)\n      self.conv_reduce = nn.Conv2d(in_chs, reduced_chs, 1, bias=True)\n      self.act1 = act_layer(inplace=True)\n      self.conv_expand = nn.Conv2d(reduced_chs, in_chs, 1, bias=True)\n   def forward(self, x):\n      x_se = self.avg_pool(x)\n      x_se = self.conv_reduce(x_se)\n      x_se = self.act1(x_se)\n      x_se = self.conv_expand(x_se)\n      x = x * self.gate_fn(x_se)\n      return x\n```</p>\n\n<h3>3) Pooling Layer</h3>\n\n<p>Once the feature passed thru SqueezeExcite Layer, I did Normal pooling. Our experiment showed that the batch normalization layer was messing with the last layer's features, so we removed it and saw a slight jump on local cv. </p>\n\n<p><code>\nself.pool = nn.Sequential(AdaptiveConcatPool2d(),\n                          Flatten(),\n                          nn.Linear(2*nc,512),\n                          nn.ReLU(inplace=True),\n                          nn.Dropout(0.4),\n                          nn.Linear(512,7), \n</code></p>\n\n<h3>4) Final Head</h3>\n\n<p>I used two heads. One head was for classification second was for regression. I noticed that training with two looses makes training much smoother (with sigmoid trick below) and yields higher local CV (0.88 -&gt; 0.90). In the final prediction, I use output only for the regression head. </p>\n\n<p>One small modification that I did before calculating loss is that the regression head used sigmoid to scale outputs between (-1. 6.). This enables much smoother training without bumps and faster convergence.</p>\n\n<p>```</p>\n\n<h1>idea taken from fastai</h1>\n\n<p>def sigmoid_range(x, low, high):\n    return torch.sigmoid(x) * (high - low) + low \n```</p>\n\n<h1>Training</h1>\n\n<p>I trained in two phases. In the First phase was trained with 49 tiles and later finetuned with 81 tiles. Both phases were using standard one cycle.</p>\n\n<h1>Final Model</h1>\n\n<p>I trained 5 fold wich resulted on the CV of 0.911 and PB: 0.922.</p>\n\n<p>Our Best Ensemble was simple average. Of 4 models. </p>\n\n<p><code>\n@drhb resnet34 5 FOLD (CV -0.911) + \n<a href=\"/rguo97\">@rguo97</a>  5 FOLD (two stage attention model CV 0.92 ) + \n<a href=\"/xiejialun\">@xiejialun</a>  FOLD (EFNET) (CV 0.915-0.917) +\n<a href=\"/cateek\">@cateek</a>  Se 1 FOLD (CV -0.91) Final Standing\n</code></p>\n\n<p><a href=\"/xiejialun\">@xiejialun</a> <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169303\">https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169303</a>\n<a href=\"/rguo97\">@rguo97</a> <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169108#940504\">https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169108#940504</a></p>\n\n<p><code>LB: 0.914 PB: 0.937</code></p>",
      "rawMarkdown": "First of all, Thank you very much to organizers.\nSecond I would like to thanks my team. We had such a positive, encouraging working environment. Our team contribution generates most of the ideas which you will read below and are shared by members. \n\n\n#Simple Resnet34 (DrHB)\n\n# Image Preprocessing \nI used medium resolution, the only preprocessing I did was to remove the white background and store medium resolution on SSD drive: \n\n```\n#function taken from R Guo\ndef crop_white(image, value: int = 255):\n    assert image.shape[2] == 3\n    assert image.dtype == np.uint8\n    ys, = (image.min((1, 2)) &lt; value).nonzero()\n    xs, = (image.min(0).min(1) &lt; value).nonzero()\n    if len(xs) == 0 or len(ys) == 0:\n        return image\n    return image[ys.min():ys.max() + 1, xs.min():xs.max() + 1]\n\n```\n\n# Cleaning data\nLike in APTOS competition, it was essential to clean images from pen marks, etc. I have used excellent work from this post: https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/151323 This also reduced the gap between CV and LB\n\n\n# Image Augmenatiosn\nAugmentation occurred at two levels.  (Slide and Tile): \n### 1) Full slide\nAfter the biopsy slide is open, we do random padding and applying one of the following transformations (similar to R Guo). \n\n```\ndef get_transforms_train():\n    transforms=A.Compose(\n        [\n            A.ShiftScaleRotate(shift_limit=0.1, scale_limit=0.05, rotate_limit=10, border_mode=cv2.BORDER_CONSTANT, p=0.5,value=(255,255,255)),\n            A.OneOf([\n                A.Flip(p=0.5),\n                A.RandomRotate90(p=0.5),\n            ], p=0.3\n            )\n        ]\n    )\n    return transforms\n```\n\n### 2) Tile\nFor each tile I used standard fastai `GPU` augmentations: `rotate=(-10, 10`),` flip vertically (p=0.5)`. For the padding I used `reflection`  it gave a slight boost on CV \n\n# Model \n\nI decided to use a very simple model resnet34 but throughout competitions ended up doing few modifications \n\n### 1) Making square features:\nThe main idea is builds up on @iafoss. Aftter resnet enccoder we reshape features to look like a square in a following way: `x = x.view(x.shape[0], x.shape[1], x.shape[2]//int(np.sqrt(N)), -1)` . Here `N` represents Number of Tiles.  After this we pass all the features to `SqueezeExcite` Block\n\n#### 2) SqueezeExcite block\nAfter reshaping features, we added 1 SE block to enable the network to learn features for individual slides based on tiles. \n\nexperiment done by @cateek \n```\n#code adopted \n#https://github.com/rwightman/pytorch-image-models/tree/master/timm/models\ndef make_divisible(v, divisor=8, min_value=None):\n   min_value = min_value or divisor\n   new_v = max(min_value, int(v + divisor / 2) // divisor * divisor)\n   # Make sure that round down does not go down by more than 10%.\n   if new_v &lt; 0.9 * v:\n      new_v += divisor\n   return new_v\ndef sigmoid(x, inplace: bool = False):\n   return x.sigmoid_() if inplace else x.sigmoid()\nclass SqueezeExcite(nn.Module):\n   def __init__(self, in_chs, se_ratio=0.25, reduced_base_chs=None,\n             act_layer=nn.ReLU, gate_fn=sigmoid, divisor=1, **_):\n      super(SqueezeExcite, self).__init__()\n      self.gate_fn = gate_fn\n      reduced_chs = make_divisible((reduced_base_chs or in_chs) * se_ratio, divisor)\n      self.avg_pool = nn.AdaptiveAvgPool2d(1)\n      self.conv_reduce = nn.Conv2d(in_chs, reduced_chs, 1, bias=True)\n      self.act1 = act_layer(inplace=True)\n      self.conv_expand = nn.Conv2d(reduced_chs, in_chs, 1, bias=True)\n   def forward(self, x):\n      x_se = self.avg_pool(x)\n      x_se = self.conv_reduce(x_se)\n      x_se = self.act1(x_se)\n      x_se = self.conv_expand(x_se)\n      x = x * self.gate_fn(x_se)\n      return x\n```\n\n### 3) Pooling Layer \nOnce the feature passed thru SqueezeExcite Layer, I did Normal pooling. Our experiment showed that the batch normalization layer was messing with the last layer's features, so we removed it and saw a slight jump on local cv. \n\n\n\n```\nself.pool = nn.Sequential(AdaptiveConcatPool2d(),\n                          Flatten(),\n                          nn.Linear(2*nc,512),\n                          nn.ReLU(inplace=True),\n                          nn.Dropout(0.4),\n                          nn.Linear(512,7), \n```\n\n\n### 4) Final Head\nI used two heads. One head was for classification second was for regression. I noticed that training with two looses makes training much smoother (with sigmoid trick below) and yields higher local CV (0.88 -&gt; 0.90). In the final prediction, I use output only for the regression head. \n\nOne small modification that I did before calculating loss is that the regression head used sigmoid to scale outputs between (-1. 6.). This enables much smoother training without bumps and faster convergence.\n\n\n```\n#idea taken from fastai\ndef sigmoid_range(x, low, high):\n    return torch.sigmoid(x) * (high - low) + low \n```\n\n# Training \nI trained in two phases. In the First phase was trained with 49 tiles and later finetuned with 81 tiles. Both phases were using standard one cycle.\n\n# Final Model \nI trained 5 fold wich resulted on the CV of 0.911 and PB: 0.922.\n\nOur Best Ensemble was simple average. Of 4 models. \n\n```\n@drhb resnet34 5 FOLD (CV -0.911) + \n@rguo97  5 FOLD (two stage attention model CV 0.92 ) + \n@xiejialun  FOLD (EFNET) (CV 0.915-0.917) +\n@cateek  Se 1 FOLD (CV -0.91) Final Standing\n```\n\n@xiejialun https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169303\n@rguo97 https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169108#940504\n\n `LB: 0.914 PB: 0.937`\n\n\n\n",
      "votes": 115
    },
    {
      "id": 940504,
      "postDate": "2020-07-23T01:48:44.607Z",
      "content": "<p>Thanks to all our teammates and congratulations to all winners. \nI provided two models in our final solution, and they work on median resolution and half-scale high resolution (2x median).</p>\n\n<p><strong>Median Resolution Model</strong></p>\n\n<p>The median level model takes tiles from <a href=\"/iafoss\">@iafoss</a> tiling method as input. Instead of caching all tiles, I give a random offset to the grid during training, this gives around 0.005 boost in CV. </p>\n\n<p>I used a concatenation of <a href=\"http://proceedings.mlr.press/v80/ilse18a/ilse18a.pdf\">attention pooling</a> and max pooling in my model. I also added two auxiliary tasks, one for classification of isup_grade and another for classifying benign or malicious tiles. I used 48 192x192 tiles for median resolution.</p>\n\n<p>5 fold median level models with seresnext50 backbone scores around 0.92 in private LB.</p>\n\n<p><strong>High Resolution Model</strong></p>\n\n<p>In the second month, I was working on higher resolution. My initial trial was to enlarge tile size and tile number. But even with half-scale high resolution images, my poor machine(2080) can hardly hold for 1 batch size. So I was trying to reduce input dimensions. </p>\n\n<p>My median resolution model has an attention pooling, which automatically calculated weights for each tile during pooling. After visualizing it, I found it fantastic to remove useless tiles. Tiles with high attention are often malicious.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3601023%2F4eb536a6ced46b5fc7189bca9a476b08%2Fimage%20(2\" alt=\"\">.png?generation=1595468713784626&amp;alt=media)</p>\n\n<p>Therefore, I input all tiles in median resolution with foreground into the network, calculate their attentions and save corresponding tiles in higher resolution according to the attention. I saved top 64 tiles with 384x384 size. In order to deal with overfitting, I saved several group of tiles for each slide, each was padded  with different offsets in tiling. </p>\n\n<p>During training, I randomly sampled 32 tiles in these top 64 tiles according to their attention and do validation/inference with top 32 tiles. </p>\n\n<p>This model scores over 0.92 in CV and around 0.93 in private LB.</p>\n\n<p><strong>Things worked</strong>\n- Add random padding during tiling</p>\n\n<ul>\n<li><p>Enlarge batchsize. While training high resolution models, I used gradient checkpoint and apex to reduce GPU memory usage. Finally with RTX Titan, I can train high resolution efficientnet-b0 with batch size 8.</p></li>\n<li><p>Deal with noise. The label is noisy, I used mse loss for karolinska and huber loss for radboud. I also tried using 0.3* oof prediction + 0.7* groundtruth as new label and train with mse loss. Both strategy improves CV by 0.005-0.01.</p></li>\n</ul>",
      "rawMarkdown": "Thanks to all our teammates and congratulations to all winners. \nI provided two models in our final solution, and they work on median resolution and half-scale high resolution (2x median).\n\n **Median Resolution Model**\n\nThe median level model takes tiles from @iafoss tiling method as input. Instead of caching all tiles, I give a random offset to the grid during training, this gives around 0.005 boost in CV. \n\nI used a concatenation of [attention pooling](http://proceedings.mlr.press/v80/ilse18a/ilse18a.pdf) and max pooling in my model. I also added two auxiliary tasks, one for classification of isup_grade and another for classifying benign or malicious tiles. I used 48 192x192 tiles for median resolution.\n\n5 fold median level models with seresnext50 backbone scores around 0.92 in private LB.\n\n**High Resolution Model**\n\nIn the second month, I was working on higher resolution. My initial trial was to enlarge tile size and tile number. But even with half-scale high resolution images, my poor machine(2080) can hardly hold for 1 batch size. So I was trying to reduce input dimensions. \n\nMy median resolution model has an attention pooling, which automatically calculated weights for each tile during pooling. After visualizing it, I found it fantastic to remove useless tiles. Tiles with high attention are often malicious.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3601023%2F4eb536a6ced46b5fc7189bca9a476b08%2Fimage%20(2).png?generation=1595468713784626&amp;alt=media)\n\nTherefore, I input all tiles in median resolution with foreground into the network, calculate their attentions and save corresponding tiles in higher resolution according to the attention. I saved top 64 tiles with 384x384 size. In order to deal with overfitting, I saved several group of tiles for each slide, each was padded  with different offsets in tiling. \n\nDuring training, I randomly sampled 32 tiles in these top 64 tiles according to their attention and do validation/inference with top 32 tiles. \n\nThis model scores over 0.92 in CV and around 0.93 in private LB.\n\n**Things worked**\n- Add random padding during tiling\n\n- Enlarge batchsize. While training high resolution models, I used gradient checkpoint and apex to reduce GPU memory usage. Finally with RTX Titan, I can train high resolution efficientnet-b0 with batch size 8.\n\n- Deal with noise. The label is noisy, I used mse loss for karolinska and huber loss for radboud. I also tried using 0.3* oof prediction + 0.7* groundtruth as new label and train with mse loss. Both strategy improves CV by 0.005-0.01.",
      "votes": 17,
      "replies": [
        {
          "id": 940795,
          "postDate": "2020-07-23T05:34:01.570Z",
          "content": "<p>congrates!!👍 \nHigh Resolution Model during inference, you input all tiles in median resolution then getting top 32 tiles. Top 32 tiles Input into high resolution model get final results. So the median resolution model plays  a preprocessing role that attention top tiles, and high resolution model predicts isup.\nDid I get it right?</p>",
          "rawMarkdown": "congrates!!👍 \nHigh Resolution Model during inference, you input all tiles in median resolution then getting top 32 tiles. Top 32 tiles Input into high resolution model get final results. So the median resolution model plays  a preprocessing role that attention top tiles, and high resolution model predicts isup.\nDid I get it right?",
          "votes": 1
        },
        {
          "id": 940798,
          "postDate": "2020-07-23T05:37:04.543Z",
          "content": "<p>Yes.</p>",
          "rawMarkdown": "Yes.",
          "votes": 1
        },
        {
          "id": 941819,
          "postDate": "2020-07-23T12:33:08.353Z",
          "content": "<p>I look forward to seeing the code, to see how that attention thing works ! The image above is quite appealing :) </p>",
          "rawMarkdown": "I look forward to seeing the code, to see how that attention thing works ! The image above is quite appealing :) ",
          "votes": 1
        },
        {
          "id": 942327,
          "postDate": "2020-07-23T17:43:59.273Z",
          "content": "<p>```python\nclass AttentionPool(nn.Module):\n    def <strong>init</strong>(self,in_ch,hidden=512,dropout=0.25):\n        super().<strong>init</strong>()\n        self.in_ch=in_ch</p>\n\n<pre><code>    module=[nn.Linear(in_ch,hidden,bias=True),\n            nn.Tanh()\n            ]\n    if dropout&gt;0:\n        module.append(nn.Dropout(dropout))\n    module.append(nn.Linear(hidden,1,bias=True))\n    self.attention=nn.Sequential(*module)\n\ndef forward(self,x):\n    num_patch=x.size(1)\n    x=x.view(-1,x.size(2))\n    A=self.attention(x)\n    A=A.view(-1,num_patch,1)\n    wt=F.softmax(A,dim=1)\n    return (x.view(-1,num_patch,self.in_ch)*wt).sum(dim=1),A\n</code></pre>\n\n<p>```</p>",
          "rawMarkdown": "```python\nclass AttentionPool(nn.Module):\n    def __init__(self,in_ch,hidden=512,dropout=0.25):\n        super().__init__()\n        self.in_ch=in_ch\n\n        module=[nn.Linear(in_ch,hidden,bias=True),\n                nn.Tanh()\n                ]\n        if dropout&gt;0:\n            module.append(nn.Dropout(dropout))\n        module.append(nn.Linear(hidden,1,bias=True))\n        self.attention=nn.Sequential(*module)\n\n    def forward(self,x):\n        num_patch=x.size(1)\n        x=x.view(-1,x.size(2))\n        A=self.attention(x)\n        A=A.view(-1,num_patch,1)\n        wt=F.softmax(A,dim=1)\n        return (x.view(-1,num_patch,self.in_ch)*wt).sum(dim=1),A\n```\n\n",
          "votes": 3
        },
        {
          "id": 944902,
          "postDate": "2020-07-25T12:45:01.127Z",
          "content": "<p><a href=\"/rguo97\">@rguo97</a> I see that this layer has two outputs, does this relate to what you said above about adding \n two auxiliary tasks, one for classifying the isup grade, and the other for classifying benign/malicious tiles? </p>\n\n<p>If so, could you tell which output is for which task (I am guessing the first output is for classifying tiles, but I am not sure), and how you went on about classifying tiles into benign/malicious categories? Did you simply resize the masks and use those as labels?</p>",
          "rawMarkdown": "@rguo97 I see that this layer has two outputs, does this relate to what you said above about adding \n two auxiliary tasks, one for classifying the isup grade, and the other for classifying benign/malicious tiles? \n\nIf so, could you tell which output is for which task (I am guessing the first output is for classifying tiles, but I am not sure), and how you went on about classifying tiles into benign/malicious categories? Did you simply resize the masks and use those as labels?"
        },
        {
          "id": 945510,
          "postDate": "2020-07-25T22:33:03.267Z",
          "content": "<p>The first output is used for following regression and classification, second one is just attention. I used the attention output to sample higher resolution tiles. The tile prediction head is before this attention pooling. I used mask to generate labels, and will not use tiles for back propagation if their mask values are all 0(unknown).\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3601023%2Fa67618109df67663c6afcf3a873054d2%2F.PNG?generation=1595716262962256&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "The first output is used for following regression and classification, second one is just attention. I used the attention output to sample higher resolution tiles. The tile prediction head is before this attention pooling. I used mask to generate labels, and will not use tiles for back propagation if their mask values are all 0(unknown).\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3601023%2Fa67618109df67663c6afcf3a873054d2%2F.PNG?generation=1595716262962256&amp;alt=media)\n",
          "votes": 1
        },
        {
          "id": 945528,
          "postDate": "2020-07-25T23:12:46.910Z",
          "content": "<p>Thank you, this definitely makes things much clearer. To fully see if I understand it, the general schedule looks something like this, correct?:</p>\n\n<ol>\n<li>Train the above schematic model on median tiles, if the mask are all unknown, exclude that specific tile from backprop.</li>\n<li>Use the trained model's attention layer to extract tiles with high attention values, and upscale these coordinates.</li>\n</ol>\n\n<p>The tile prediction head can then be seen as a raw classifier that should give a better distinction between isup grades of 0-2 (benign) and malignant (3-5), and the actual regression/classification heads are then the usual heads that predict the actual isup grades. Is that correct?</p>",
          "rawMarkdown": "Thank you, this definitely makes things much clearer. To fully see if I understand it, the general schedule looks something like this, correct?:\n\n1. Train the above schematic model on median tiles, if the mask are all unknown, exclude that specific tile from backprop.\n2. Use the trained model's attention layer to extract tiles with high attention values, and upscale these coordinates.\n\nThe tile prediction head can then be seen as a raw classifier that should give a better distinction between isup grades of 0-2 (benign) and malignant (3-5), and the actual regression/classification heads are then the usual heads that predict the actual isup grades. Is that correct?"
        },
        {
          "id": 945596,
          "postDate": "2020-07-26T01:57:16.447Z",
          "content": "<p>Yes. Except tiles are cropped from highest resolution in tiff file then downsampled by 2.</p>",
          "rawMarkdown": "Yes. Except tiles are cropped from highest resolution in tiff file then downsampled by 2.\n",
          "votes": 1
        },
        {
          "id": 946018,
          "postDate": "2020-07-26T09:32:33.117Z",
          "content": "<p>I see, thank you very much! </p>",
          "rawMarkdown": "I see, thank you very much! "
        },
        {
          "id": 965163,
          "postDate": "2020-08-10T12:29:41.453Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/rguo97\" target=\"_blank\">@rguo97</a>!<br>\nwhat is shown in the second of three picture? </p>",
          "rawMarkdown": "Hi @rguo97!\nwhat is shown in the second of three picture? "
        },
        {
          "id": 965533,
          "postDate": "2020-08-10T17:43:28.467Z",
          "content": "<p>Masks, gray and green are benign (isup 1/2).</p>",
          "rawMarkdown": "Masks, gray and green are benign (isup 1/2)."
        }
      ]
    },
    {
      "id": 941984,
      "postDate": "2020-07-23T14:26:25.893Z",
      "content": "<p>Thanks to Kaggle team and organizers for the interesting competition! Also big thanks to my teammates! \nAfter some experiments with low resolution and attempts to specifically address label noise issue, I concentrated on designing a generally robust solution. I found training to be very sensitive to batch size and number of tiles. Therefore I decided to stick with resnext50 with groupnorm and <a href=\"https://arxiv.org/abs/1903.10520\">weights stantardization</a>, which allowed to train with batch size 1. </p>\n\n<p>My work is based on <a href=\"/iafoss\">@iafoss</a> 's kernel. I extracted tiles from WSI using <code>skimage view_as_windows</code>, removing completely white tiles. I applied very mild random_crop+rotate to whole slide, which allowed to get slightly different tiles during each iteration. I used 49 tiles of shape 224x224, taking ~1/3 of tiles from an array, sorted by tissue amount, and the rest of tiles were taken at random.</p>\n\n<p>Model was trained for 50 epochs with reduce on plateau scheduler, and finetuned for another 20 epochs with cosine decay. Although the latter resulted in nothing more, than overfitting. \nI used regression formulation with a single output and smoothL1 loss.</p>\n\n<p>Tile level augmentations used(courtesy of albumentations library):\n<code>\nOneOf([\n    RGBShift(p=1),\n    RandomGamma(p=1),\n], p=0.5),\nRandomBrightnessContrast(p=0.7),\nOneOf([\n    RandomRotate90(p=1),\n    Flip(p=1),\n    Rotate(limit=10, border_mode=0, value=(255, 255, 255), p=1),\n    ShiftScaleRotate(shift_limit=0.15, scale_limit=0.1, rotate_limit=10, border_mode=0, value=(255, 255, 255), p=1),\n], p=0.25),\nOneOf([\n    Cutout(num_holes=50, max_h_size=10, max_w_size=10, fill_value=0, p=1),\n    Cutout(num_holes=70, max_h_size=7, max_w_size=7, fill_value=0, p=1),\n    Cutout(num_holes=100, max_h_size=5, max_w_size=5, fill_value=0, p=1),\n], p=0.2),\n</code>\nOut of necessity to save submits, I really did only 3 submission of this model\nLocal / Public / Private:\n0.896 / 0.885 / 0.922\n0.903 / 0.892 / 0.920\n0.904 / 0.893 / 0.914\nOne with the highest local/public score was finetuned for another 20 epochs, as described earlier, and led to overfitting. </p>\n\n<p>Things that didn't work:\n1. Various stain normalization methods\n2. Online erroneous labels mining\n3. Overlapping tiles \n4. smaller models with batchnorm</p>",
      "rawMarkdown": "Thanks to Kaggle team and organizers for the interesting competition! Also big thanks to my teammates! \nAfter some experiments with low resolution and attempts to specifically address label noise issue, I concentrated on designing a generally robust solution. I found training to be very sensitive to batch size and number of tiles. Therefore I decided to stick with resnext50 with groupnorm and [weights stantardization](https://arxiv.org/abs/1903.10520), which allowed to train with batch size 1. \n\nMy work is based on @iafoss 's kernel. I extracted tiles from WSI using `skimage view_as_windows`, removing completely white tiles. I applied very mild random_crop+rotate to whole slide, which allowed to get slightly different tiles during each iteration. I used 49 tiles of shape 224x224, taking ~1/3 of tiles from an array, sorted by tissue amount, and the rest of tiles were taken at random.\n \nModel was trained for 50 epochs with reduce on plateau scheduler, and finetuned for another 20 epochs with cosine decay. Although the latter resulted in nothing more, than overfitting. \nI used regression formulation with a single output and smoothL1 loss.\n\nTile level augmentations used(courtesy of albumentations library):\n```\nOneOf([\n\tRGBShift(p=1),\n\tRandomGamma(p=1),\n], p=0.5),\nRandomBrightnessContrast(p=0.7),\nOneOf([\n\tRandomRotate90(p=1),\n\tFlip(p=1),\n\tRotate(limit=10, border_mode=0, value=(255, 255, 255), p=1),\n\tShiftScaleRotate(shift_limit=0.15, scale_limit=0.1, rotate_limit=10, border_mode=0, value=(255, 255, 255), p=1),\n], p=0.25),\nOneOf([\n\tCutout(num_holes=50, max_h_size=10, max_w_size=10, fill_value=0, p=1),\n\tCutout(num_holes=70, max_h_size=7, max_w_size=7, fill_value=0, p=1),\n\tCutout(num_holes=100, max_h_size=5, max_w_size=5, fill_value=0, p=1),\n], p=0.2),\n```\nOut of necessity to save submits, I really did only 3 submission of this model\nLocal / Public / Private:\n0.896 / 0.885 / 0.922\n0.903 / 0.892 / 0.920\n0.904 / 0.893 / 0.914\nOne with the highest local/public score was finetuned for another 20 epochs, as described earlier, and led to overfitting. \n\nThings that didn't work:\n1. Various stain normalization methods\n2. Online erroneous labels mining\n3. Overlapping tiles \n4. smaller models with batchnorm\n ",
      "votes": 7
    },
    {
      "id": 944247,
      "postDate": "2020-07-25T00:48:16.613Z",
      "content": "<p><a href=\"/drhabib\">@drhabib</a> , Congratulations. It seems that you didn't remove noise from labels, so I wanted to ask what was the spread of your LB result for different seeds if u checked it. I have a hypotheses that noise removal reduces it making the results more stable, but most of my experiments ran with noise removal, so I cannot do a comparison. In my case the std is ~0.003 for 5 models (4-folds) trained with different seeds, while I saw some people at the forum reported really huge differences of 0.03. I just wanted to check how it looks for your models?</p>",
      "rawMarkdown": "@drhabib , Congratulations. It seems that you didn't remove noise from labels, so I wanted to ask what was the spread of your LB result for different seeds if u checked it. I have a hypotheses that noise removal reduces it making the results more stable, but most of my experiments ran with noise removal, so I cannot do a comparison. In my case the std is ~0.003 for 5 models (4-folds) trained with different seeds, while I saw some people at the forum reported really huge differences of 0.03. I just wanted to check how it looks for your models?",
      "votes": 6,
      "replies": [
        {
          "id": 945534,
          "postDate": "2020-07-25T23:35:25.830Z",
          "content": "<p>Thank you for the questions. If I understand your question correctly, we have not tried to do diffrent seeds me and <a href=\"/cateek\">@cateek</a> had split which we created when we merged and we used this split to run all our experiments... </p>\n\n<p>By the way. I just want to say again thank you very much for your wonderful kernels and insight during the competition. </p>",
          "rawMarkdown": "Thank you for the questions. If I understand your question correctly, we have not tried to do diffrent seeds me and @cateek had split which we created when we merged and we used this split to run all our experiments... \n\nBy the way. I just want to say again thank you very much for your wonderful kernels and insight during the competition. ",
          "votes": 1
        },
        {
          "id": 945536,
          "postDate": "2020-07-25T23:38:39.077Z",
          "content": "<p>You are very welcome </p>",
          "rawMarkdown": "You are very welcome ",
          "votes": 1
        }
      ]
    },
    {
      "id": 979773,
      "postDate": "2020-08-21T05:38:47.167Z",
      "content": "<p>Great work <a href=\"https://www.kaggle.com/drhabib\" target=\"_blank\">@drhabib</a> and an amazing write up! Great to see Fastai implementations on top of the LB! </p>",
      "rawMarkdown": "Great work @drhabib and an amazing write up! Great to see Fastai implementations on top of the LB! ",
      "votes": 1
    },
    {
      "id": 945741,
      "postDate": "2020-07-26T05:34:46.317Z",
      "content": "<p>Cool!</p>",
      "rawMarkdown": "Cool!",
      "votes": 1
    },
    {
      "id": 945036,
      "postDate": "2020-07-25T14:25:40.343Z",
      "content": "<p>Congratulation!!. thank you for your insight!. I learned a lot ~ </p>",
      "rawMarkdown": "Congratulation!!. thank you for your insight!. I learned a lot ~ ",
      "votes": 1
    },
    {
      "id": 944962,
      "postDate": "2020-07-25T13:30:13.717Z",
      "content": "<p><a href=\"/drhabib\">@drhabib</a> \nCongratulation. and thanks for sharing. I didn't know about such a trick on sigmoid. Happy to learn. =)</p>",
      "rawMarkdown": "@drhabib \nCongratulation. and thanks for sharing. I didn't know about such a trick on sigmoid. Happy to learn. =)",
      "votes": 1
    },
    {
      "id": 941547,
      "postDate": "2020-07-23T09:39:59.217Z",
      "content": "<p>Training\n\"I trained in two phases. In the First phase was trained with 49 tiles and later finetuned with 81 tiles. Both phases were using standard one cycle.\"\nI don't quite understand your statement, could you elaborate on it?</p>",
      "rawMarkdown": "Training\n\"I trained in two phases. In the First phase was trained with 49 tiles and later finetuned with 81 tiles. Both phases were using standard one cycle.\"\nI don't quite understand your statement, could you elaborate on it?",
      "votes": 1,
      "replies": [
        {
          "id": 941806,
          "postDate": "2020-07-23T12:27:15.980Z",
          "content": "<p>I train first using 49 Tiles. Save The Weights of the model. Change Number of Tiles to 81, Load the saved weights and continue training. Both 49 and 82 Tiles use one cycle learning rate schedule. </p>",
          "rawMarkdown": "I train first using 49 Tiles. Save The Weights of the model. Change Number of Tiles to 81, Load the saved weights and continue training. Both 49 and 82 Tiles use one cycle learning rate schedule. ",
          "votes": 2
        },
        {
          "id": 942795,
          "postDate": "2020-07-24T02:15:45.333Z",
          "content": "<p>Have you ever tried using only one stage? Will the performance make a difference？</p>",
          "rawMarkdown": "Have you ever tried using only one stage? Will the performance make a difference？"
        }
      ]
    },
    {
      "id": 941235,
      "postDate": "2020-07-23T06:24:12.820Z",
      "content": "<p>congrats <a href=\"/drhabib\">@drhabib</a> and thanks for good explanation!</p>\n\n<p>I didn’t participate in this competition.\nThanks to a good explanation, I learned a lot. 👍 </p>",
      "rawMarkdown": "congrats @drhabib and thanks for good explanation!\n\nI didn’t participate in this competition.\nThanks to a good explanation, I learned a lot. 👍 ",
      "votes": 1
    },
    {
      "id": 940811,
      "postDate": "2020-07-23T05:49:30.240Z",
      "content": "<p>Congratulations, very successful</p>",
      "rawMarkdown": "Congratulations, very successful\n",
      "votes": 1
    },
    {
      "id": 940711,
      "postDate": "2020-07-23T05:12:54.667Z",
      "content": "<p>Congrats! Do you plan to share your code?</p>\n\n<p>Also, this may be the first time fastai has been used to win a competition! cc: <a href=\"/jhoward\">@jhoward</a> </p>",
      "rawMarkdown": "Congrats! Do you plan to share your code?\n\nAlso, this may be the first time fastai has been used to win a competition! cc: @jhoward ",
      "votes": 1,
      "replies": [
        {
          "id": 940808,
          "postDate": "2020-07-23T05:46:28.330Z",
          "content": "<p>Thank you very much! Yes I am planing to share code :) </p>",
          "rawMarkdown": "Thank you very much! Yes I am planing to share code :) ",
          "votes": 1
        }
      ]
    },
    {
      "id": 940610,
      "postDate": "2020-07-23T03:48:32.943Z",
      "content": "<p>Great solution and thanks for the explanation. Can you say how many epochs you trained for ?</p>",
      "rawMarkdown": "Great solution and thanks for the explanation. Can you say how many epochs you trained for ?",
      "votes": 1,
      "replies": [
        {
          "id": 940613,
          "postDate": "2020-07-23T03:50:50.297Z",
          "content": "<p>In my case is it was 60=) For other it was between (10-40).</p>",
          "rawMarkdown": "In my case is it was 60=) For other it was between (10-40).",
          "votes": 2
        },
        {
          "id": 940801,
          "postDate": "2020-07-23T05:41:14.747Z",
          "content": "<p>Thanks for the reply. How long approx did it take to run an epoch ? </p>",
          "rawMarkdown": "Thanks for the reply. How long approx did it take to run an epoch ? ",
          "votes": 1
        },
        {
          "id": 940806,
          "postDate": "2020-07-23T05:45:07.990Z",
          "content": "<p>for me 11 min. For <a href=\"/rguo97\">@rguo97</a> on high resolution 2nd level takes 30–40 min . For rest of the team between 7-20 min. </p>",
          "rawMarkdown": "for me 11 min. For @rguo97 on high resolution 2nd level takes 30–40 min . For rest of the team between 7-20 min. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 940576,
      "postDate": "2020-07-23T03:12:19.547Z",
      "content": "<p>Congratulations for the 2nd place!</p>\n\n<p>Your team is high on both public and private LB, so I want to ask: do you have stable cv - LB relationship? How did you chose the final subs? Thanks.</p>",
      "rawMarkdown": "Congratulations for the 2nd place!\n\nYour team is high on both public and private LB, so I want to ask: do you have stable cv - LB relationship? How did you chose the final subs? Thanks.",
      "votes": 1,
      "replies": [
        {
          "id": 940580,
          "postDate": "2020-07-23T03:18:46.483Z",
          "content": "<p>Our best CV ensemble is a little bit higher than best LB ensemble. CV-LB  relationship is still not very reliable but in my case, CV and private-LB has good relationship.</p>",
          "rawMarkdown": "Our best CV ensemble is a little bit higher than best LB ensemble. CV-LB  relationship is still not very reliable but in my case, CV and private-LB has good relationship.",
          "votes": 2
        },
        {
          "id": 940597,
          "postDate": "2020-07-23T03:39:27.133Z",
          "content": "<p>Thanks. So I assume you chose best cv sub and best public LB sub as final subs. And the former turned out to be 2nd place private?</p>",
          "rawMarkdown": "Thanks. So I assume you chose best cv sub and best public LB sub as final subs. And the former turned out to be 2nd place private?",
          "votes": 1
        },
        {
          "id": 940600,
          "postDate": "2020-07-23T03:40:36.577Z",
          "content": "<p>Yes, but best CV scores 0.937 and best LB scores 0.935</p>",
          "rawMarkdown": "Yes, but best CV scores 0.937 and best LB scores 0.935",
          "votes": 2
        },
        {
          "id": 940625,
          "postDate": "2020-07-23T03:59:37.283Z",
          "content": "<p>wow that's very stable. Both of your subs would be 2nd private. Looks like diverse ensembles can mitigate the cv-LB inconsistency to some extent 👍 </p>",
          "rawMarkdown": "wow that's very stable. Both of your subs would be 2nd private. Looks like diverse ensembles can mitigate the cv-LB inconsistency to some extent 👍 ",
          "votes": 3
        },
        {
          "id": 941781,
          "postDate": "2020-07-23T12:12:12.137Z",
          "content": "<p>&gt; Yes, but best CV scores 0.937 and best LB scores 0.935</p>\n\n<p>That's an incredible feat ! </p>",
          "rawMarkdown": "&gt; Yes, but best CV scores 0.937 and best LB scores 0.935\n\nThat's an incredible feat ! "
        }
      ]
    },
    {
      "id": 940532,
      "postDate": "2020-07-23T02:18:50.043Z",
      "content": "<p>Congrats!!👍 </p>",
      "rawMarkdown": "Congrats!!👍 ",
      "votes": 1
    },
    {
      "id": 940500,
      "postDate": "2020-07-23T01:45:24.337Z",
      "content": "<p>Congrats!👍👍   I'm so curious about your idea. </p>",
      "rawMarkdown": "Congrats!👍👍   I'm so curious about your idea. ",
      "votes": 1
    },
    {
      "id": 940470,
      "postDate": "2020-07-23T01:09:13.353Z",
      "content": "<p>Great game!\nLike APTOS, I guess ensambling was a key solution. Looking forward to specific solutions.</p>",
      "rawMarkdown": "Great game!\nLike APTOS, I guess ensambling was a key solution. Looking forward to specific solutions.",
      "votes": 1,
      "replies": [
        {
          "id": 940657,
          "postDate": "2020-07-23T04:27:26.640Z",
          "content": "<p>yes ensembling helped us a lot =)  Congratulations on 1st place =)</p>",
          "rawMarkdown": "yes ensembling helped us a lot =)  Congratulations on 1st place =)",
          "votes": 1
        },
        {
          "id": 944351,
          "postDate": "2020-07-25T03:35:52.450Z",
          "content": "<p>Thank you! Your solution was a lot more stable both on LB/PB and I learned a lot from it!</p>",
          "rawMarkdown": "Thank you! Your solution was a lot more stable both on LB/PB and I learned a lot from it!",
          "votes": 1
        }
      ]
    },
    {
      "id": 940454,
      "postDate": "2020-07-23T00:38:59.567Z",
      "content": "<p>Congrats and wait for your update</p>",
      "rawMarkdown": "Congrats and wait for your update\n\n",
      "votes": 1
    },
    {
      "id": 941901,
      "postDate": "2020-07-23T13:38:44.137Z",
      "content": "<ol>\n<li>For final head, how do you weight the losses? equally 50/50?</li>\n<li>For final head, what loss did you use for regression? MSE? MAE? I like that sigmoid trick, have never seen before.</li>\n</ol>",
      "rawMarkdown": "1. For final head, how do you weight the losses? equally 50/50?\n2. For final head, what loss did you use for regression? MSE? MAE? I like that sigmoid trick, have never seen before.",
      "votes": 2,
      "replies": [
        {
          "id": 941926,
          "postDate": "2020-07-23T13:54:11.713Z",
          "content": "<p>1) Yes 50, 50\n2) <code>nn.MSELoss()</code> </p>",
          "rawMarkdown": "1) Yes 50, 50\n2) `nn.MSELoss()` ",
          "votes": 1
        }
      ]
    },
    {
      "id": 943738,
      "postDate": "2020-07-24T14:43:31.600Z",
      "content": "<p>Very cool.</p>",
      "rawMarkdown": "Very cool."
    },
    {
      "id": 941944,
      "postDate": "2020-07-23T14:03:28.400Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 977015,
      "postDate": "2020-08-19T08:27:55.960Z",
      "content": "<p>Thanks for sharing and congrats! </p>",
      "rawMarkdown": "Thanks for sharing and congrats! ",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 940504,
      "author_name": "R Guo",
      "author_url": "",
      "post_date": "2020-07-23T01:48:44.607000",
      "content": "<p>Thanks to all our teammates and congratulations to all winners. \nI provided two models in our final solution, and they work on median resolution and half-scale high resolution (2x median).</p>\n\n<p><strong>Median Resolution Model</strong></p>\n\n<p>The median level model takes tiles from <a href=\"/iafoss\">@iafoss</a> tiling method as input. Instead of caching all tiles, I give a random offset to the grid during training, this gives around 0.005 boost in CV. </p>\n\n<p>I used a concatenation of <a href=\"http://proceedings.mlr.press/v80/ilse18a/ilse18a.pdf\">attention pooling</a> and max pooling in my model. I also added two auxiliary tasks, one for classification of isup_grade and another for classifying benign or malicious tiles. I used 48 192x192 tiles for median resolution.</p>\n\n<p>5 fold median level models with seresnext50 backbone scores around 0.92 in private LB.</p>\n\n<p><strong>High Resolution Model</strong></p>\n\n<p>In the second month, I was working on higher resolution. My initial trial was to enlarge tile size and tile number. But even with half-scale high resolution images, my poor machine(2080) can hardly hold for 1 batch size. So I was trying to reduce input dimensions. </p>\n\n<p>My median resolution model has an attention pooling, which automatically calculated weights for each tile during pooling. After visualizing it, I found it fantastic to remove useless tiles. Tiles with high attention are often malicious.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3601023%2F4eb536a6ced46b5fc7189bca9a476b08%2Fimage%20(2\" alt=\"\">.png?generation=1595468713784626&amp;alt=media)</p>\n\n<p>Therefore, I input all tiles in median resolution with foreground into the network, calculate their attentions and save corresponding tiles in higher resolution according to the attention. I saved top 64 tiles with 384x384 size. In order to deal with overfitting, I saved several group of tiles for each slide, each was padded  with different offsets in tiling. </p>\n\n<p>During training, I randomly sampled 32 tiles in these top 64 tiles according to their attention and do validation/inference with top 32 tiles. </p>\n\n<p>This model scores over 0.92 in CV and around 0.93 in private LB.</p>\n\n<p><strong>Things worked</strong>\n- Add random padding during tiling</p>\n\n<ul>\n<li><p>Enlarge batchsize. While training high resolution models, I used gradient checkpoint and apex to reduce GPU memory usage. Finally with RTX Titan, I can train high resolution efficientnet-b0 with batch size 8.</p></li>\n<li><p>Deal with noise. The label is noisy, I used mse loss for karolinska and huber loss for radboud. I also tried using 0.3* oof prediction + 0.7* groundtruth as new label and train with mse loss. Both strategy improves CV by 0.005-0.01.</p></li>\n</ul>",
      "votes": 17,
      "replies": [
        {
          "id": 940795,
          "author_name": "ecnu_ybx",
          "author_url": "",
          "post_date": "2020-07-23T05:34:01.570000",
          "content": "<p>congrates!!👍 \nHigh Resolution Model during inference, you input all tiles in median resolution then getting top 32 tiles. Top 32 tiles Input into high resolution model get final results. So the median resolution model plays  a preprocessing role that attention top tiles, and high resolution model predicts isup.\nDid I get it right?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 940798,
          "author_name": "R Guo",
          "author_url": "",
          "post_date": "2020-07-23T05:37:04.543000",
          "content": "<p>Yes.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 941819,
          "author_name": "Benjamin Dubreu",
          "author_url": "",
          "post_date": "2020-07-23T12:33:08.353000",
          "content": "<p>I look forward to seeing the code, to see how that attention thing works ! The image above is quite appealing :) </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 942327,
          "author_name": "R Guo",
          "author_url": "",
          "post_date": "2020-07-23T17:43:59.273000",
          "content": "<p>```python\nclass AttentionPool(nn.Module):\n    def <strong>init</strong>(self,in_ch,hidden=512,dropout=0.25):\n        super().<strong>init</strong>()\n        self.in_ch=in_ch</p>\n\n<pre><code>    module=[nn.Linear(in_ch,hidden,bias=True),\n            nn.Tanh()\n            ]\n    if dropout&gt;0:\n        module.append(nn.Dropout(dropout))\n    module.append(nn.Linear(hidden,1,bias=True))\n    self.attention=nn.Sequential(*module)\n\ndef forward(self,x):\n    num_patch=x.size(1)\n    x=x.view(-1,x.size(2))\n    A=self.attention(x)\n    A=A.view(-1,num_patch,1)\n    wt=F.softmax(A,dim=1)\n    return (x.view(-1,num_patch,self.in_ch)*wt).sum(dim=1),A\n</code></pre>\n\n<p>```</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 944902,
          "author_name": "Stephan",
          "author_url": "",
          "post_date": "2020-07-25T12:45:01.127000",
          "content": "<p><a href=\"/rguo97\">@rguo97</a> I see that this layer has two outputs, does this relate to what you said above about adding \n two auxiliary tasks, one for classifying the isup grade, and the other for classifying benign/malicious tiles? </p>\n\n<p>If so, could you tell which output is for which task (I am guessing the first output is for classifying tiles, but I am not sure), and how you went on about classifying tiles into benign/malicious categories? Did you simply resize the masks and use those as labels?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 945510,
          "author_name": "R Guo",
          "author_url": "",
          "post_date": "2020-07-25T22:33:03.267000",
          "content": "<p>The first output is used for following regression and classification, second one is just attention. I used the attention output to sample higher resolution tiles. The tile prediction head is before this attention pooling. I used mask to generate labels, and will not use tiles for back propagation if their mask values are all 0(unknown).\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3601023%2Fa67618109df67663c6afcf3a873054d2%2F.PNG?generation=1595716262962256&amp;alt=media\" alt=\"\"></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 945528,
          "author_name": "Stephan",
          "author_url": "",
          "post_date": "2020-07-25T23:12:46.910000",
          "content": "<p>Thank you, this definitely makes things much clearer. To fully see if I understand it, the general schedule looks something like this, correct?:</p>\n\n<ol>\n<li>Train the above schematic model on median tiles, if the mask are all unknown, exclude that specific tile from backprop.</li>\n<li>Use the trained model's attention layer to extract tiles with high attention values, and upscale these coordinates.</li>\n</ol>\n\n<p>The tile prediction head can then be seen as a raw classifier that should give a better distinction between isup grades of 0-2 (benign) and malignant (3-5), and the actual regression/classification heads are then the usual heads that predict the actual isup grades. Is that correct?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 945596,
          "author_name": "R Guo",
          "author_url": "",
          "post_date": "2020-07-26T01:57:16.447000",
          "content": "<p>Yes. Except tiles are cropped from highest resolution in tiff file then downsampled by 2.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 946018,
          "author_name": "Stephan",
          "author_url": "",
          "post_date": "2020-07-26T09:32:33.117000",
          "content": "<p>I see, thank you very much! </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 965163,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-10T12:29:41.453000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/rguo97\" target=\"_blank\">@rguo97</a>!<br>\nwhat is shown in the second of three picture? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 965533,
          "author_name": "R Guo",
          "author_url": "",
          "post_date": "2020-08-10T17:43:28.467000",
          "content": "<p>Masks, gray and green are benign (isup 1/2).</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 941984,
      "author_name": "Eek The Cat",
      "author_url": "",
      "post_date": "2020-07-23T14:26:25.893000",
      "content": "<p>Thanks to Kaggle team and organizers for the interesting competition! Also big thanks to my teammates! \nAfter some experiments with low resolution and attempts to specifically address label noise issue, I concentrated on designing a generally robust solution. I found training to be very sensitive to batch size and number of tiles. Therefore I decided to stick with resnext50 with groupnorm and <a href=\"https://arxiv.org/abs/1903.10520\">weights stantardization</a>, which allowed to train with batch size 1. </p>\n\n<p>My work is based on <a href=\"/iafoss\">@iafoss</a> 's kernel. I extracted tiles from WSI using <code>skimage view_as_windows</code>, removing completely white tiles. I applied very mild random_crop+rotate to whole slide, which allowed to get slightly different tiles during each iteration. I used 49 tiles of shape 224x224, taking ~1/3 of tiles from an array, sorted by tissue amount, and the rest of tiles were taken at random.</p>\n\n<p>Model was trained for 50 epochs with reduce on plateau scheduler, and finetuned for another 20 epochs with cosine decay. Although the latter resulted in nothing more, than overfitting. \nI used regression formulation with a single output and smoothL1 loss.</p>\n\n<p>Tile level augmentations used(courtesy of albumentations library):\n<code>\nOneOf([\n    RGBShift(p=1),\n    RandomGamma(p=1),\n], p=0.5),\nRandomBrightnessContrast(p=0.7),\nOneOf([\n    RandomRotate90(p=1),\n    Flip(p=1),\n    Rotate(limit=10, border_mode=0, value=(255, 255, 255), p=1),\n    ShiftScaleRotate(shift_limit=0.15, scale_limit=0.1, rotate_limit=10, border_mode=0, value=(255, 255, 255), p=1),\n], p=0.25),\nOneOf([\n    Cutout(num_holes=50, max_h_size=10, max_w_size=10, fill_value=0, p=1),\n    Cutout(num_holes=70, max_h_size=7, max_w_size=7, fill_value=0, p=1),\n    Cutout(num_holes=100, max_h_size=5, max_w_size=5, fill_value=0, p=1),\n], p=0.2),\n</code>\nOut of necessity to save submits, I really did only 3 submission of this model\nLocal / Public / Private:\n0.896 / 0.885 / 0.922\n0.903 / 0.892 / 0.920\n0.904 / 0.893 / 0.914\nOne with the highest local/public score was finetuned for another 20 epochs, as described earlier, and led to overfitting. </p>\n\n<p>Things that didn't work:\n1. Various stain normalization methods\n2. Online erroneous labels mining\n3. Overlapping tiles \n4. smaller models with batchnorm</p>",
      "votes": 7,
      "replies": []
    },
    {
      "id": 944247,
      "author_name": "Iafoss",
      "author_url": "",
      "post_date": "2020-07-25T00:48:16.613000",
      "content": "<p><a href=\"/drhabib\">@drhabib</a> , Congratulations. It seems that you didn't remove noise from labels, so I wanted to ask what was the spread of your LB result for different seeds if u checked it. I have a hypotheses that noise removal reduces it making the results more stable, but most of my experiments ran with noise removal, so I cannot do a comparison. In my case the std is ~0.003 for 5 models (4-folds) trained with different seeds, while I saw some people at the forum reported really huge differences of 0.03. I just wanted to check how it looks for your models?</p>",
      "votes": 6,
      "replies": [
        {
          "id": 945534,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2020-07-25T23:35:25.830000",
          "content": "<p>Thank you for the questions. If I understand your question correctly, we have not tried to do diffrent seeds me and <a href=\"/cateek\">@cateek</a> had split which we created when we merged and we used this split to run all our experiments... </p>\n\n<p>By the way. I just want to say again thank you very much for your wonderful kernels and insight during the competition. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 945536,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-07-25T23:38:39.077000",
          "content": "<p>You are very welcome </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 979773,
      "author_name": "Harish Vadlamani",
      "author_url": "",
      "post_date": "2020-08-21T05:38:47.167000",
      "content": "<p>Great work <a href=\"https://www.kaggle.com/drhabib\" target=\"_blank\">@drhabib</a> and an amazing write up! Great to see Fastai implementations on top of the LB! </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 945741,
      "author_name": "Vincent Junitio Ungu",
      "author_url": "",
      "post_date": "2020-07-26T05:34:46.317000",
      "content": "<p>Cool!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 945036,
      "author_name": "JJShadow",
      "author_url": "",
      "post_date": "2020-07-25T14:25:40.343000",
      "content": "<p>Congratulation!!. thank you for your insight!. I learned a lot ~ </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 944962,
      "author_name": "Innat",
      "author_url": "",
      "post_date": "2020-07-25T13:30:13.717000",
      "content": "<p><a href=\"/drhabib\">@drhabib</a> \nCongratulation. and thanks for sharing. I didn't know about such a trick on sigmoid. Happy to learn. =)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 941547,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-23T09:39:59.217000",
      "content": "<p>Training\n\"I trained in two phases. In the First phase was trained with 49 tiles and later finetuned with 81 tiles. Both phases were using standard one cycle.\"\nI don't quite understand your statement, could you elaborate on it?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 941806,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2020-07-23T12:27:15.980000",
          "content": "<p>I train first using 49 Tiles. Save The Weights of the model. Change Number of Tiles to 81, Load the saved weights and continue training. Both 49 and 82 Tiles use one cycle learning rate schedule. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 942795,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-24T02:15:45.333000",
          "content": "<p>Have you ever tried using only one stage? Will the performance make a difference？</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 941235,
      "author_name": "Heroseo",
      "author_url": "",
      "post_date": "2020-07-23T06:24:12.820000",
      "content": "<p>congrats <a href=\"/drhabib\">@drhabib</a> and thanks for good explanation!</p>\n\n<p>I didn’t participate in this competition.\nThanks to a good explanation, I learned a lot. 👍 </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 940811,
      "author_name": "Hamza Tanç",
      "author_url": "",
      "post_date": "2020-07-23T05:49:30.240000",
      "content": "<p>Congratulations, very successful</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 940711,
      "author_name": "ilovescience",
      "author_url": "",
      "post_date": "2020-07-23T05:12:54.667000",
      "content": "<p>Congrats! Do you plan to share your code?</p>\n\n<p>Also, this may be the first time fastai has been used to win a competition! cc: <a href=\"/jhoward\">@jhoward</a> </p>",
      "votes": 1,
      "replies": [
        {
          "id": 940808,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2020-07-23T05:46:28.330000",
          "content": "<p>Thank you very much! Yes I am planing to share code :) </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 940610,
      "author_name": "jwelliav",
      "author_url": "",
      "post_date": "2020-07-23T03:48:32.943000",
      "content": "<p>Great solution and thanks for the explanation. Can you say how many epochs you trained for ?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 940613,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2020-07-23T03:50:50.297000",
          "content": "<p>In my case is it was 60=) For other it was between (10-40).</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 940801,
          "author_name": "jwelliav",
          "author_url": "",
          "post_date": "2020-07-23T05:41:14.747000",
          "content": "<p>Thanks for the reply. How long approx did it take to run an epoch ? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 940806,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2020-07-23T05:45:07.990000",
          "content": "<p>for me 11 min. For <a href=\"/rguo97\">@rguo97</a> on high resolution 2nd level takes 30–40 min . For rest of the team between 7-20 min. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 940576,
      "author_name": "Bo",
      "author_url": "",
      "post_date": "2020-07-23T03:12:19.547000",
      "content": "<p>Congratulations for the 2nd place!</p>\n\n<p>Your team is high on both public and private LB, so I want to ask: do you have stable cv - LB relationship? How did you chose the final subs? Thanks.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 940580,
          "author_name": "R Guo",
          "author_url": "",
          "post_date": "2020-07-23T03:18:46.483000",
          "content": "<p>Our best CV ensemble is a little bit higher than best LB ensemble. CV-LB  relationship is still not very reliable but in my case, CV and private-LB has good relationship.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 940597,
          "author_name": "Bo",
          "author_url": "",
          "post_date": "2020-07-23T03:39:27.133000",
          "content": "<p>Thanks. So I assume you chose best cv sub and best public LB sub as final subs. And the former turned out to be 2nd place private?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 940600,
          "author_name": "R Guo",
          "author_url": "",
          "post_date": "2020-07-23T03:40:36.577000",
          "content": "<p>Yes, but best CV scores 0.937 and best LB scores 0.935</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 940625,
          "author_name": "Bo",
          "author_url": "",
          "post_date": "2020-07-23T03:59:37.283000",
          "content": "<p>wow that's very stable. Both of your subs would be 2nd private. Looks like diverse ensembles can mitigate the cv-LB inconsistency to some extent 👍 </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 941781,
          "author_name": "Benjamin Dubreu",
          "author_url": "",
          "post_date": "2020-07-23T12:12:12.137000",
          "content": "<p>&gt; Yes, but best CV scores 0.937 and best LB scores 0.935</p>\n\n<p>That's an incredible feat ! </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 940532,
      "author_name": "wangyunpeng_bio",
      "author_url": "",
      "post_date": "2020-07-23T02:18:50.043000",
      "content": "<p>Congrats!!👍 </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 940500,
      "author_name": "Aroddary",
      "author_url": "",
      "post_date": "2020-07-23T01:45:24.337000",
      "content": "<p>Congrats!👍👍   I'm so curious about your idea. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 940470,
      "author_name": "arutema47",
      "author_url": "",
      "post_date": "2020-07-23T01:09:13.353000",
      "content": "<p>Great game!\nLike APTOS, I guess ensambling was a key solution. Looking forward to specific solutions.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 940657,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2020-07-23T04:27:26.640000",
          "content": "<p>yes ensembling helped us a lot =)  Congratulations on 1st place =)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 944351,
          "author_name": "arutema47",
          "author_url": "",
          "post_date": "2020-07-25T03:35:52.450000",
          "content": "<p>Thank you! Your solution was a lot more stable both on LB/PB and I learned a lot from it!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 940454,
      "author_name": "Youngno Yoon",
      "author_url": "",
      "post_date": "2020-07-23T00:38:59.567000",
      "content": "<p>Congrats and wait for your update</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 941901,
      "author_name": "CoreyJamesLevinson",
      "author_url": "",
      "post_date": "2020-07-23T13:38:44.137000",
      "content": "<ol>\n<li>For final head, how do you weight the losses? equally 50/50?</li>\n<li>For final head, what loss did you use for regression? MSE? MAE? I like that sigmoid trick, have never seen before.</li>\n</ol>",
      "votes": 2,
      "replies": [
        {
          "id": 941926,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2020-07-23T13:54:11.713000",
          "content": "<p>1) Yes 50, 50\n2) <code>nn.MSELoss()</code> </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 943738,
      "author_name": "Mitchell Zufelt",
      "author_url": "",
      "post_date": "2020-07-24T14:43:31.600000",
      "content": "<p>Very cool.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 941944,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-23T14:03:28.400000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 977015,
      "author_name": "Abishek Sudarshan",
      "author_url": "",
      "post_date": "2020-08-19T08:27:55.960000",
      "content": "<p>Thanks for sharing and congrats! </p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "940434": "First of all, Thank you very much to organizers.\nSecond I would like to thanks my team. We had such a positive, encouraging working environment. Our team contribution generates most of the ideas which you will read below and are shared by members. \n\n\n#Simple Resnet34 (DrHB)\n\n# Image Preprocessing \nI used medium resolution, the only preprocessing I did was to remove the white background and store medium resolution on SSD drive: \n\n```\n#function taken from R Guo\ndef crop_white(image, value: int = 255):\n    assert image.shape[2] == 3\n    assert image.dtype == np.uint8\n    ys, = (image.min((1, 2)) &lt; value).nonzero()\n    xs, = (image.min(0).min(1) &lt; value).nonzero()\n    if len(xs) == 0 or len(ys) == 0:\n        return image\n    return image[ys.min():ys.max() + 1, xs.min():xs.max() + 1]\n\n```\n\n# Cleaning data\nLike in APTOS competition, it was essential to clean images from pen marks, etc. I have used excellent work from this post: https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/151323 This also reduced the gap between CV and LB\n\n\n# Image Augmenatiosn\nAugmentation occurred at two levels.  (Slide and Tile): \n### 1) Full slide\nAfter the biopsy slide is open, we do random padding and applying one of the following transformations (similar to R Guo). \n\n```\ndef get_transforms_train():\n    transforms=A.Compose(\n        [\n            A.ShiftScaleRotate(shift_limit=0.1, scale_limit=0.05, rotate_limit=10, border_mode=cv2.BORDER_CONSTANT, p=0.5,value=(255,255,255)),\n            A.OneOf([\n                A.Flip(p=0.5),\n                A.RandomRotate90(p=0.5),\n            ], p=0.3\n            )\n        ]\n    )\n    return transforms\n```\n\n### 2) Tile\nFor each tile I used standard fastai `GPU` augmentations: `rotate=(-10, 10`),` flip vertically (p=0.5)`. For the padding I used `reflection`  it gave a slight boost on CV \n\n# Model \n\nI decided to use a very simple model resnet34 but throughout competitions ended up doing few modifications \n\n### 1) Making square features:\nThe main idea is builds up on @iafoss. Aftter resnet enccoder we reshape features to look like a square in a following way: `x = x.view(x.shape[0], x.shape[1], x.shape[2]//int(np.sqrt(N)), -1)` . Here `N` represents Number of Tiles.  After this we pass all the features to `SqueezeExcite` Block\n\n#### 2) SqueezeExcite block\nAfter reshaping features, we added 1 SE block to enable the network to learn features for individual slides based on tiles. \n\nexperiment done by @cateek \n```\n#code adopted \n#https://github.com/rwightman/pytorch-image-models/tree/master/timm/models\ndef make_divisible(v, divisor=8, min_value=None):\n   min_value = min_value or divisor\n   new_v = max(min_value, int(v + divisor / 2) // divisor * divisor)\n   # Make sure that round down does not go down by more than 10%.\n   if new_v &lt; 0.9 * v:\n      new_v += divisor\n   return new_v\ndef sigmoid(x, inplace: bool = False):\n   return x.sigmoid_() if inplace else x.sigmoid()\nclass SqueezeExcite(nn.Module):\n   def __init__(self, in_chs, se_ratio=0.25, reduced_base_chs=None,\n             act_layer=nn.ReLU, gate_fn=sigmoid, divisor=1, **_):\n      super(SqueezeExcite, self).__init__()\n      self.gate_fn = gate_fn\n      reduced_chs = make_divisible((reduced_base_chs or in_chs) * se_ratio, divisor)\n      self.avg_pool = nn.AdaptiveAvgPool2d(1)\n      self.conv_reduce = nn.Conv2d(in_chs, reduced_chs, 1, bias=True)\n      self.act1 = act_layer(inplace=True)\n      self.conv_expand = nn.Conv2d(reduced_chs, in_chs, 1, bias=True)\n   def forward(self, x):\n      x_se = self.avg_pool(x)\n      x_se = self.conv_reduce(x_se)\n      x_se = self.act1(x_se)\n      x_se = self.conv_expand(x_se)\n      x = x * self.gate_fn(x_se)\n      return x\n```\n\n### 3) Pooling Layer \nOnce the feature passed thru SqueezeExcite Layer, I did Normal pooling. Our experiment showed that the batch normalization layer was messing with the last layer's features, so we removed it and saw a slight jump on local cv. \n\n\n\n```\nself.pool = nn.Sequential(AdaptiveConcatPool2d(),\n                          Flatten(),\n                          nn.Linear(2*nc,512),\n                          nn.ReLU(inplace=True),\n                          nn.Dropout(0.4),\n                          nn.Linear(512,7), \n```\n\n\n### 4) Final Head\nI used two heads. One head was for classification second was for regression. I noticed that training with two looses makes training much smoother (with sigmoid trick below) and yields higher local CV (0.88 -&gt; 0.90). In the final prediction, I use output only for the regression head. \n\nOne small modification that I did before calculating loss is that the regression head used sigmoid to scale outputs between (-1. 6.). This enables much smoother training without bumps and faster convergence.\n\n\n```\n#idea taken from fastai\ndef sigmoid_range(x, low, high):\n    return torch.sigmoid(x) * (high - low) + low \n```\n\n# Training \nI trained in two phases. In the First phase was trained with 49 tiles and later finetuned with 81 tiles. Both phases were using standard one cycle.\n\n# Final Model \nI trained 5 fold wich resulted on the CV of 0.911 and PB: 0.922.\n\nOur Best Ensemble was simple average. Of 4 models. \n\n```\n@drhb resnet34 5 FOLD (CV -0.911) + \n@rguo97  5 FOLD (two stage attention model CV 0.92 ) + \n@xiejialun  FOLD (EFNET) (CV 0.915-0.917) +\n@cateek  Se 1 FOLD (CV -0.91) Final Standing\n```\n\n@xiejialun https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169303\n@rguo97 https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169108#940504\n\n `LB: 0.914 PB: 0.937`\n\n\n\n",
    "940504": "Thanks to all our teammates and congratulations to all winners. \nI provided two models in our final solution, and they work on median resolution and half-scale high resolution (2x median).\n\n **Median Resolution Model**\n\nThe median level model takes tiles from @iafoss tiling method as input. Instead of caching all tiles, I give a random offset to the grid during training, this gives around 0.005 boost in CV. \n\nI used a concatenation of [attention pooling](http://proceedings.mlr.press/v80/ilse18a/ilse18a.pdf) and max pooling in my model. I also added two auxiliary tasks, one for classification of isup_grade and another for classifying benign or malicious tiles. I used 48 192x192 tiles for median resolution.\n\n5 fold median level models with seresnext50 backbone scores around 0.92 in private LB.\n\n**High Resolution Model**\n\nIn the second month, I was working on higher resolution. My initial trial was to enlarge tile size and tile number. But even with half-scale high resolution images, my poor machine(2080) can hardly hold for 1 batch size. So I was trying to reduce input dimensions. \n\nMy median resolution model has an attention pooling, which automatically calculated weights for each tile during pooling. After visualizing it, I found it fantastic to remove useless tiles. Tiles with high attention are often malicious.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3601023%2F4eb536a6ced46b5fc7189bca9a476b08%2Fimage%20(2).png?generation=1595468713784626&amp;alt=media)\n\nTherefore, I input all tiles in median resolution with foreground into the network, calculate their attentions and save corresponding tiles in higher resolution according to the attention. I saved top 64 tiles with 384x384 size. In order to deal with overfitting, I saved several group of tiles for each slide, each was padded  with different offsets in tiling. \n\nDuring training, I randomly sampled 32 tiles in these top 64 tiles according to their attention and do validation/inference with top 32 tiles. \n\nThis model scores over 0.92 in CV and around 0.93 in private LB.\n\n**Things worked**\n- Add random padding during tiling\n\n- Enlarge batchsize. While training high resolution models, I used gradient checkpoint and apex to reduce GPU memory usage. Finally with RTX Titan, I can train high resolution efficientnet-b0 with batch size 8.\n\n- Deal with noise. The label is noisy, I used mse loss for karolinska and huber loss for radboud. I also tried using 0.3* oof prediction + 0.7* groundtruth as new label and train with mse loss. Both strategy improves CV by 0.005-0.01.",
    "941984": "Thanks to Kaggle team and organizers for the interesting competition! Also big thanks to my teammates! \nAfter some experiments with low resolution and attempts to specifically address label noise issue, I concentrated on designing a generally robust solution. I found training to be very sensitive to batch size and number of tiles. Therefore I decided to stick with resnext50 with groupnorm and [weights stantardization](https://arxiv.org/abs/1903.10520), which allowed to train with batch size 1. \n\nMy work is based on @iafoss 's kernel. I extracted tiles from WSI using `skimage view_as_windows`, removing completely white tiles. I applied very mild random_crop+rotate to whole slide, which allowed to get slightly different tiles during each iteration. I used 49 tiles of shape 224x224, taking ~1/3 of tiles from an array, sorted by tissue amount, and the rest of tiles were taken at random.\n \nModel was trained for 50 epochs with reduce on plateau scheduler, and finetuned for another 20 epochs with cosine decay. Although the latter resulted in nothing more, than overfitting. \nI used regression formulation with a single output and smoothL1 loss.\n\nTile level augmentations used(courtesy of albumentations library):\n```\nOneOf([\n\tRGBShift(p=1),\n\tRandomGamma(p=1),\n], p=0.5),\nRandomBrightnessContrast(p=0.7),\nOneOf([\n\tRandomRotate90(p=1),\n\tFlip(p=1),\n\tRotate(limit=10, border_mode=0, value=(255, 255, 255), p=1),\n\tShiftScaleRotate(shift_limit=0.15, scale_limit=0.1, rotate_limit=10, border_mode=0, value=(255, 255, 255), p=1),\n], p=0.25),\nOneOf([\n\tCutout(num_holes=50, max_h_size=10, max_w_size=10, fill_value=0, p=1),\n\tCutout(num_holes=70, max_h_size=7, max_w_size=7, fill_value=0, p=1),\n\tCutout(num_holes=100, max_h_size=5, max_w_size=5, fill_value=0, p=1),\n], p=0.2),\n```\nOut of necessity to save submits, I really did only 3 submission of this model\nLocal / Public / Private:\n0.896 / 0.885 / 0.922\n0.903 / 0.892 / 0.920\n0.904 / 0.893 / 0.914\nOne with the highest local/public score was finetuned for another 20 epochs, as described earlier, and led to overfitting. \n\nThings that didn't work:\n1. Various stain normalization methods\n2. Online erroneous labels mining\n3. Overlapping tiles \n4. smaller models with batchnorm\n ",
    "944247": "@drhabib , Congratulations. It seems that you didn't remove noise from labels, so I wanted to ask what was the spread of your LB result for different seeds if u checked it. I have a hypotheses that noise removal reduces it making the results more stable, but most of my experiments ran with noise removal, so I cannot do a comparison. In my case the std is ~0.003 for 5 models (4-folds) trained with different seeds, while I saw some people at the forum reported really huge differences of 0.03. I just wanted to check how it looks for your models?",
    "979773": "Great work @drhabib and an amazing write up! Great to see Fastai implementations on top of the LB! ",
    "945741": "Cool!",
    "945036": "Congratulation!!. thank you for your insight!. I learned a lot ~ ",
    "944962": "@drhabib \nCongratulation. and thanks for sharing. I didn't know about such a trick on sigmoid. Happy to learn. =)",
    "941547": "Training\n\"I trained in two phases. In the First phase was trained with 49 tiles and later finetuned with 81 tiles. Both phases were using standard one cycle.\"\nI don't quite understand your statement, could you elaborate on it?",
    "941235": "congrats @drhabib and thanks for good explanation!\n\nI didn’t participate in this competition.\nThanks to a good explanation, I learned a lot. 👍 ",
    "940811": "Congratulations, very successful\n",
    "940711": "Congrats! Do you plan to share your code?\n\nAlso, this may be the first time fastai has been used to win a competition! cc: @jhoward ",
    "940610": "Great solution and thanks for the explanation. Can you say how many epochs you trained for ?",
    "940576": "Congratulations for the 2nd place!\n\nYour team is high on both public and private LB, so I want to ask: do you have stable cv - LB relationship? How did you chose the final subs? Thanks.",
    "940532": "Congrats!!👍 ",
    "940500": "Congrats!👍👍   I'm so curious about your idea. ",
    "940470": "Great game!\nLike APTOS, I guess ensambling was a key solution. Looking forward to specific solutions.",
    "940454": "Congrats and wait for your update\n\n",
    "941901": "1. For final head, how do you weight the losses? equally 50/50?\n2. For final head, what loss did you use for regression? MSE? MAE? I like that sigmoid trick, have never seen before.",
    "943738": "Very cool.",
    "941944": "",
    "977015": "Thanks for sharing and congrats! "
  }
}