{
  "id": 169225,
  "title": "7th Place Solution（simple but messy）",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/169225",
  "author_name": "ctrasd123",
  "post_date": "2020-07-23T09:03:45.153000",
  "votes": 9,
  "comment_count": 2,
  "views": 0,
  "content": "<p>First of all, Thank you very much to organizers.</p>\n\n<p>This challenge is very similar to APTOS-2019 which I have worked for months as a course assignment. So I simlpy used the pipeline of my course assignment(Based on <a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/107947\">Lex Toumbourou‘s solution</a>, thanks a lot) with some revised details. Also thanks a lot to <a href=\"https://www.kaggle.com/iafoss/panda-16x128x128-tiles\">Iafoss</a> and <a href=\"https://www.kaggle.com/haqishen/train-efficientnet-b0-w-36-tiles-256-lb0-87\">Qishen Ha</a> for their useful notebooks.</p>\n\n<p>Our model is simple but messy.</p>\n\n<h2>Tiles</h2>\n\n<p>We tried 256x256x32, 192x192x64, 154x154x100 but they didn't show some difference on Public LB. </p>\n\n<p>We also propoesd a new tile approach. It can contain more pathological parts without destroying the shape features. Large size tile ensures that the shape features will not be lost, while small size tiles ensure that the blank area is not that large.</p>\n\n<pre><code>def get_tiles_combine(img,mode=0):\n    images = np.ones((1536, 1536, 3))*255\n    h, w, c = img.shape\n    result_all=[]\n    pad_h = (256 - h % 256) % 256 + ((256 * mode) // 2)\n    pad_w = (256 - w % 256) % 256 + ((256 * mode) // 2)\n    #print(pad_h,pad_w,c)\n    img2 = np.pad(img,[[pad_h // 2, pad_h - pad_h // 2], [pad_w // 2,pad_w - pad_w//2], [0,0]], 'constant',constant_values=255)\n    windows=[256,256,256,256,192,192,128]\n    x_start=0\n    for i in range(len(windows)):\n        result = []\n        window_size=windows[i]\n        for x in range((h+pad_h)//window_size):\n            for y in range((w+pad_w)//window_size):\n                tile=img2[x*window_size:(x+1)*window_size,y*window_size:(y+1)*window_size]\n                result.append([x,y,tile.sum()])\n        #print(len(result))\n        result.sort(key=lambda ele:ele[2])\n        result=result[:1536//window_size]\n        #print(len(result),result)\n        for y in range(min(1536//window_size,len(result))):\n            xx=result[y][0]\n            yy=result[y][1]\n            result_all.append([xx,yy])\n            images[x_start:x_start+window_size,y*window_size:(y+1)*window_size]=\\\n                img2[xx*window_size:(xx+1)*window_size,yy*window_size:(yy+1)*window_size].copy()\n            img2[xx*window_size:(xx+1)*window_size,yy*window_size:(yy+1)*window_size]=255\n        x_start=x_start+windows[i]\n    return images \n</code></pre>\n\n<h2>Models</h2>\n\n<p>We simply used Efficientnet-B0. We tried B1-B3,Densenet and Resnext, but they didn't show some difference on Public LB and need more GPU memory.</p>\n\n<p>According to APTOS-2019, we used the GeM pooling:</p>\n\n<pre><code>def gem(x, p=3, eps=1e-6):\n    return F.avg_pool2d(x.clamp(min=eps).pow(p), (x.size(-2), x.size(-1))).pow(1./p)\n\nclass GeM(nn.Module):\n    def __init__(self, p=3, eps=1e-6):\n        super(GeM,self).__init__()\n        self.p = Parameter(torch.ones(1)*p)\n        self.eps = eps\n    def forward(self, x):\n        return gem(x, p=self.p, eps=self.eps)       \n    def __repr__(self):\n        return self.__class__.__name__ + '(' + 'p=' + '{:.4f}'.format(self.p.data.tolist()[0]) + ', ' + 'eps=' + str(self.eps) + ')'\n</code></pre>\n\n<p>We also designed a more complex model based on b0. It has lower Public LB(average about 0.85) so we didn't add it to our final models. But it got the highest single-model Private LB (max 0.926, average about 0.920, What a pitty!). We will do some more experiments on this model.</p>\n\n<h2>Loss and label</h2>\n\n<p>BCE Loss and label smoothing: 3-&gt;[0.95,0.95,0.95,0.95,0.05,0.05]</p>\n\n<p>We also tried regression with mse loss and smooth L1 loss, but they didn't show any improvement.</p>\n\n<h2>Ensemble</h2>\n\n<p>8 models with 6 * TTA:</p>\n\n<blockquote>\n  <p>1: fold_1 b0 256-tile  Public LB:0.879, Private LB:0.904.</p>\n  \n  <p>2: fold_3 b0 256-tile Public LB:0.879, Private LB:0.899.</p>\n  \n  <p>3: fold_4 b0 256-tile Public LB:0.886, Private LB:0.883.</p>\n  \n  <p>4: fold_4 b0 combine-tile Public LB:0.879, Private LB:0.910.</p>\n  \n  <p>5: fold_4 b0 256-tile Public LB:0.880, Private LB:0.909.</p>\n  \n  <p>6: fold_4 b0 256-tile Public LB:0.891, Private LB:0.920.</p>\n  \n  <p>7: fold_4 b0 combine-tile Public LB:0.881 Private LB:0.917.</p>\n  \n  <p>8: fold_0 b0 256-tile Public LB:0.872, Private LB:0.906.</p>\n</blockquote>\n\n<p>The final model has Public LB:0.894, Private LB:0.932.</p>",
  "messages": [
    {
      "id": 941515,
      "postDate": "2020-07-23T09:03:45.153Z",
      "content": "<p>First of all, Thank you very much to organizers.</p>\n\n<p>This challenge is very similar to APTOS-2019 which I have worked for months as a course assignment. So I simlpy used the pipeline of my course assignment(Based on <a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/107947\">Lex Toumbourou‘s solution</a>, thanks a lot) with some revised details. Also thanks a lot to <a href=\"https://www.kaggle.com/iafoss/panda-16x128x128-tiles\">Iafoss</a> and <a href=\"https://www.kaggle.com/haqishen/train-efficientnet-b0-w-36-tiles-256-lb0-87\">Qishen Ha</a> for their useful notebooks.</p>\n\n<p>Our model is simple but messy.</p>\n\n<h2>Tiles</h2>\n\n<p>We tried 256x256x32, 192x192x64, 154x154x100 but they didn't show some difference on Public LB. </p>\n\n<p>We also propoesd a new tile approach. It can contain more pathological parts without destroying the shape features. Large size tile ensures that the shape features will not be lost, while small size tiles ensure that the blank area is not that large.</p>\n\n<pre><code>def get_tiles_combine(img,mode=0):\n    images = np.ones((1536, 1536, 3))*255\n    h, w, c = img.shape\n    result_all=[]\n    pad_h = (256 - h % 256) % 256 + ((256 * mode) // 2)\n    pad_w = (256 - w % 256) % 256 + ((256 * mode) // 2)\n    #print(pad_h,pad_w,c)\n    img2 = np.pad(img,[[pad_h // 2, pad_h - pad_h // 2], [pad_w // 2,pad_w - pad_w//2], [0,0]], 'constant',constant_values=255)\n    windows=[256,256,256,256,192,192,128]\n    x_start=0\n    for i in range(len(windows)):\n        result = []\n        window_size=windows[i]\n        for x in range((h+pad_h)//window_size):\n            for y in range((w+pad_w)//window_size):\n                tile=img2[x*window_size:(x+1)*window_size,y*window_size:(y+1)*window_size]\n                result.append([x,y,tile.sum()])\n        #print(len(result))\n        result.sort(key=lambda ele:ele[2])\n        result=result[:1536//window_size]\n        #print(len(result),result)\n        for y in range(min(1536//window_size,len(result))):\n            xx=result[y][0]\n            yy=result[y][1]\n            result_all.append([xx,yy])\n            images[x_start:x_start+window_size,y*window_size:(y+1)*window_size]=\\\n                img2[xx*window_size:(xx+1)*window_size,yy*window_size:(yy+1)*window_size].copy()\n            img2[xx*window_size:(xx+1)*window_size,yy*window_size:(yy+1)*window_size]=255\n        x_start=x_start+windows[i]\n    return images \n</code></pre>\n\n<h2>Models</h2>\n\n<p>We simply used Efficientnet-B0. We tried B1-B3,Densenet and Resnext, but they didn't show some difference on Public LB and need more GPU memory.</p>\n\n<p>According to APTOS-2019, we used the GeM pooling:</p>\n\n<pre><code>def gem(x, p=3, eps=1e-6):\n    return F.avg_pool2d(x.clamp(min=eps).pow(p), (x.size(-2), x.size(-1))).pow(1./p)\n\nclass GeM(nn.Module):\n    def __init__(self, p=3, eps=1e-6):\n        super(GeM,self).__init__()\n        self.p = Parameter(torch.ones(1)*p)\n        self.eps = eps\n    def forward(self, x):\n        return gem(x, p=self.p, eps=self.eps)       \n    def __repr__(self):\n        return self.__class__.__name__ + '(' + 'p=' + '{:.4f}'.format(self.p.data.tolist()[0]) + ', ' + 'eps=' + str(self.eps) + ')'\n</code></pre>\n\n<p>We also designed a more complex model based on b0. It has lower Public LB(average about 0.85) so we didn't add it to our final models. But it got the highest single-model Private LB (max 0.926, average about 0.920, What a pitty!). We will do some more experiments on this model.</p>\n\n<h2>Loss and label</h2>\n\n<p>BCE Loss and label smoothing: 3-&gt;[0.95,0.95,0.95,0.95,0.05,0.05]</p>\n\n<p>We also tried regression with mse loss and smooth L1 loss, but they didn't show any improvement.</p>\n\n<h2>Ensemble</h2>\n\n<p>8 models with 6 * TTA:</p>\n\n<blockquote>\n  <p>1: fold_1 b0 256-tile  Public LB:0.879, Private LB:0.904.</p>\n  \n  <p>2: fold_3 b0 256-tile Public LB:0.879, Private LB:0.899.</p>\n  \n  <p>3: fold_4 b0 256-tile Public LB:0.886, Private LB:0.883.</p>\n  \n  <p>4: fold_4 b0 combine-tile Public LB:0.879, Private LB:0.910.</p>\n  \n  <p>5: fold_4 b0 256-tile Public LB:0.880, Private LB:0.909.</p>\n  \n  <p>6: fold_4 b0 256-tile Public LB:0.891, Private LB:0.920.</p>\n  \n  <p>7: fold_4 b0 combine-tile Public LB:0.881 Private LB:0.917.</p>\n  \n  <p>8: fold_0 b0 256-tile Public LB:0.872, Private LB:0.906.</p>\n</blockquote>\n\n<p>The final model has Public LB:0.894, Private LB:0.932.</p>",
      "rawMarkdown": "First of all, Thank you very much to organizers.\n\nThis challenge is very similar to APTOS-2019 which I have worked for months as a course assignment. So I simlpy used the pipeline of my course assignment(Based on [Lex Toumbourou‘s solution](https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/107947), thanks a lot) with some revised details. Also thanks a lot to [Iafoss](https://www.kaggle.com/iafoss/panda-16x128x128-tiles) and [Qishen Ha](https://www.kaggle.com/haqishen/train-efficientnet-b0-w-36-tiles-256-lb0-87) for their useful notebooks.\n\nOur model is simple but messy.\n\n## Tiles\nWe tried 256x256x32, 192x192x64, 154x154x100 but they didn't show some difference on Public LB. \n\nWe also propoesd a new tile approach. It can contain more pathological parts without destroying the shape features. Large size tile ensures that the shape features will not be lost, while small size tiles ensure that the blank area is not that large.\n\n    def get_tiles_combine(img,mode=0):\n        images = np.ones((1536, 1536, 3))*255\n        h, w, c = img.shape\n        result_all=[]\n        pad_h = (256 - h % 256) % 256 + ((256 * mode) // 2)\n        pad_w = (256 - w % 256) % 256 + ((256 * mode) // 2)\n        #print(pad_h,pad_w,c)\n        img2 = np.pad(img,[[pad_h // 2, pad_h - pad_h // 2], [pad_w // 2,pad_w - pad_w//2], [0,0]], 'constant',constant_values=255)\n        windows=[256,256,256,256,192,192,128]\n        x_start=0\n        for i in range(len(windows)):\n            result = []\n            window_size=windows[i]\n            for x in range((h+pad_h)//window_size):\n                for y in range((w+pad_w)//window_size):\n                    tile=img2[x*window_size:(x+1)*window_size,y*window_size:(y+1)*window_size]\n                    result.append([x,y,tile.sum()])\n            #print(len(result))\n            result.sort(key=lambda ele:ele[2])\n            result=result[:1536//window_size]\n            #print(len(result),result)\n            for y in range(min(1536//window_size,len(result))):\n                xx=result[y][0]\n                yy=result[y][1]\n                result_all.append([xx,yy])\n                images[x_start:x_start+window_size,y*window_size:(y+1)*window_size]=\\\n                    img2[xx*window_size:(xx+1)*window_size,yy*window_size:(yy+1)*window_size].copy()\n                img2[xx*window_size:(xx+1)*window_size,yy*window_size:(yy+1)*window_size]=255\n            x_start=x_start+windows[i]\n        return images \n \n##Models\n\nWe simply used Efficientnet-B0. We tried B1-B3,Densenet and Resnext, but they didn't show some difference on Public LB and need more GPU memory.\n\nAccording to APTOS-2019, we used the GeM pooling:\n\n\n    def gem(x, p=3, eps=1e-6):\n        return F.avg_pool2d(x.clamp(min=eps).pow(p), (x.size(-2), x.size(-1))).pow(1./p)\n\n    class GeM(nn.Module):\n        def __init__(self, p=3, eps=1e-6):\n            super(GeM,self).__init__()\n            self.p = Parameter(torch.ones(1)*p)\n            self.eps = eps\n        def forward(self, x):\n            return gem(x, p=self.p, eps=self.eps)       \n        def __repr__(self):\n            return self.__class__.__name__ + '(' + 'p=' + '{:.4f}'.format(self.p.data.tolist()[0]) + ', ' + 'eps=' + str(self.eps) + ')'\n\nWe also designed a more complex model based on b0. It has lower Public LB(average about 0.85) so we didn't add it to our final models. But it got the highest single-model Private LB (max 0.926, average about 0.920, What a pitty!). We will do some more experiments on this model.\n\n##Loss and label\n\nBCE Loss and label smoothing: 3-&gt;[0.95,0.95,0.95,0.95,0.05,0.05]\n\nWe also tried regression with mse loss and smooth L1 loss, but they didn't show any improvement.\n\n##Ensemble\n8 models with 6 * TTA:\n\n&gt; 1: fold_1 b0 256-tile  Public LB:0.879, Private LB:0.904.\n\n&gt; 2: fold_3 b0 256-tile Public LB:0.879, Private LB:0.899.\n\n&gt; 3: fold_4 b0 256-tile Public LB:0.886, Private LB:0.883.\n\n&gt; 4: fold_4 b0 combine-tile Public LB:0.879, Private LB:0.910.\n\n&gt; 5: fold_4 b0 256-tile Public LB:0.880, Private LB:0.909.\n\n&gt; 6: fold_4 b0 256-tile Public LB:0.891, Private LB:0.920.\n\n&gt; 7: fold_4 b0 combine-tile Public LB:0.881 Private LB:0.917.\n\n&gt; 8: fold_0 b0 256-tile Public LB:0.872, Private LB:0.906.\n\n\nThe final model has Public LB:0.894, Private LB:0.932.\n\n",
      "votes": 8
    },
    {
      "id": 941537,
      "postDate": "2020-07-23T09:33:27.530Z",
      "content": "<p>Interesting, congrats for the result.  I used the same label smoothing than you and it performed badly for me ( actually I just checked, public LB was bad but private was pretty good).  I am not sure why, given it worked fine for you.</p>\n\n<p>What fold_x  means?  You selected only some folds from you models?</p>",
      "rawMarkdown": "Interesting, congrats for the result.  I used the same label smoothing than you and it performed badly for me ( actually I just checked, public LB was bad but private was pretty good).  I am not sure why, given it worked fine for you.\n\n\nWhat fold_x  means?  You selected only some folds from you models?",
      "replies": [
        {
          "id": 941688,
          "postDate": "2020-07-23T10:57:43.330Z",
          "content": "<p>We split the dataset into 5 folds. Fold_x above means we use the xth fold as the test dataset and other folds as the training dataset. \nI don't know why your label smoothing doesn't work either😂. I think it is not strange we have different performance between public LB and Private LB since the dataset is noisy(Which lead to the shakeup).  Maybe label smoothing is not the real factor. I use the label smoothing just because there is a deviation between CV and LB, and I want to avoid overfit. </p>",
          "rawMarkdown": "We split the dataset into 5 folds. Fold_x above means we use the xth fold as the test dataset and other folds as the training dataset. \nI don't know why your label smoothing doesn't work either😂. I think it is not strange we have different performance between public LB and Private LB since the dataset is noisy(Which lead to the shakeup).  Maybe label smoothing is not the real factor. I use the label smoothing just because there is a deviation between CV and LB, and I want to avoid overfit. "
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 941537,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2020-07-23T09:33:27.530000",
      "content": "<p>Interesting, congrats for the result.  I used the same label smoothing than you and it performed badly for me ( actually I just checked, public LB was bad but private was pretty good).  I am not sure why, given it worked fine for you.</p>\n\n<p>What fold_x  means?  You selected only some folds from you models?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 941688,
          "author_name": "ctrasd123",
          "author_url": "",
          "post_date": "2020-07-23T10:57:43.330000",
          "content": "<p>We split the dataset into 5 folds. Fold_x above means we use the xth fold as the test dataset and other folds as the training dataset. \nI don't know why your label smoothing doesn't work either😂. I think it is not strange we have different performance between public LB and Private LB since the dataset is noisy(Which lead to the shakeup).  Maybe label smoothing is not the real factor. I use the label smoothing just because there is a deviation between CV and LB, and I want to avoid overfit. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "941515": "First of all, Thank you very much to organizers.\n\nThis challenge is very similar to APTOS-2019 which I have worked for months as a course assignment. So I simlpy used the pipeline of my course assignment(Based on [Lex Toumbourou‘s solution](https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/107947), thanks a lot) with some revised details. Also thanks a lot to [Iafoss](https://www.kaggle.com/iafoss/panda-16x128x128-tiles) and [Qishen Ha](https://www.kaggle.com/haqishen/train-efficientnet-b0-w-36-tiles-256-lb0-87) for their useful notebooks.\n\nOur model is simple but messy.\n\n## Tiles\nWe tried 256x256x32, 192x192x64, 154x154x100 but they didn't show some difference on Public LB. \n\nWe also propoesd a new tile approach. It can contain more pathological parts without destroying the shape features. Large size tile ensures that the shape features will not be lost, while small size tiles ensure that the blank area is not that large.\n\n    def get_tiles_combine(img,mode=0):\n        images = np.ones((1536, 1536, 3))*255\n        h, w, c = img.shape\n        result_all=[]\n        pad_h = (256 - h % 256) % 256 + ((256 * mode) // 2)\n        pad_w = (256 - w % 256) % 256 + ((256 * mode) // 2)\n        #print(pad_h,pad_w,c)\n        img2 = np.pad(img,[[pad_h // 2, pad_h - pad_h // 2], [pad_w // 2,pad_w - pad_w//2], [0,0]], 'constant',constant_values=255)\n        windows=[256,256,256,256,192,192,128]\n        x_start=0\n        for i in range(len(windows)):\n            result = []\n            window_size=windows[i]\n            for x in range((h+pad_h)//window_size):\n                for y in range((w+pad_w)//window_size):\n                    tile=img2[x*window_size:(x+1)*window_size,y*window_size:(y+1)*window_size]\n                    result.append([x,y,tile.sum()])\n            #print(len(result))\n            result.sort(key=lambda ele:ele[2])\n            result=result[:1536//window_size]\n            #print(len(result),result)\n            for y in range(min(1536//window_size,len(result))):\n                xx=result[y][0]\n                yy=result[y][1]\n                result_all.append([xx,yy])\n                images[x_start:x_start+window_size,y*window_size:(y+1)*window_size]=\\\n                    img2[xx*window_size:(xx+1)*window_size,yy*window_size:(yy+1)*window_size].copy()\n                img2[xx*window_size:(xx+1)*window_size,yy*window_size:(yy+1)*window_size]=255\n            x_start=x_start+windows[i]\n        return images \n \n##Models\n\nWe simply used Efficientnet-B0. We tried B1-B3,Densenet and Resnext, but they didn't show some difference on Public LB and need more GPU memory.\n\nAccording to APTOS-2019, we used the GeM pooling:\n\n\n    def gem(x, p=3, eps=1e-6):\n        return F.avg_pool2d(x.clamp(min=eps).pow(p), (x.size(-2), x.size(-1))).pow(1./p)\n\n    class GeM(nn.Module):\n        def __init__(self, p=3, eps=1e-6):\n            super(GeM,self).__init__()\n            self.p = Parameter(torch.ones(1)*p)\n            self.eps = eps\n        def forward(self, x):\n            return gem(x, p=self.p, eps=self.eps)       \n        def __repr__(self):\n            return self.__class__.__name__ + '(' + 'p=' + '{:.4f}'.format(self.p.data.tolist()[0]) + ', ' + 'eps=' + str(self.eps) + ')'\n\nWe also designed a more complex model based on b0. It has lower Public LB(average about 0.85) so we didn't add it to our final models. But it got the highest single-model Private LB (max 0.926, average about 0.920, What a pitty!). We will do some more experiments on this model.\n\n##Loss and label\n\nBCE Loss and label smoothing: 3-&gt;[0.95,0.95,0.95,0.95,0.05,0.05]\n\nWe also tried regression with mse loss and smooth L1 loss, but they didn't show any improvement.\n\n##Ensemble\n8 models with 6 * TTA:\n\n&gt; 1: fold_1 b0 256-tile  Public LB:0.879, Private LB:0.904.\n\n&gt; 2: fold_3 b0 256-tile Public LB:0.879, Private LB:0.899.\n\n&gt; 3: fold_4 b0 256-tile Public LB:0.886, Private LB:0.883.\n\n&gt; 4: fold_4 b0 combine-tile Public LB:0.879, Private LB:0.910.\n\n&gt; 5: fold_4 b0 256-tile Public LB:0.880, Private LB:0.909.\n\n&gt; 6: fold_4 b0 256-tile Public LB:0.891, Private LB:0.920.\n\n&gt; 7: fold_4 b0 combine-tile Public LB:0.881 Private LB:0.917.\n\n&gt; 8: fold_0 b0 256-tile Public LB:0.872, Private LB:0.906.\n\n\nThe final model has Public LB:0.894, Private LB:0.932.\n\n",
    "941537": "Interesting, congrats for the result.  I used the same label smoothing than you and it performed badly for me ( actually I just checked, public LB was bad but private was pretty good).  I am not sure why, given it worked fine for you.\n\n\nWhat fold_x  means?  You selected only some folds from you models?"
  }
}