{
  "id": 171065,
  "title": "25th place solution [Kaggle_gaggle]",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/171065",
  "author_name": "Sangwon Lee",
  "post_date": "2020-07-30T08:34:09.983000",
  "votes": 6,
  "comment_count": 1,
  "views": 0,
  "content": "<p>First of all Thank you very much to organizers and thanks to <a href=\"/kyunghoonhur\">@kyunghoonhur</a> for collaborating with me!</p>\n\n<p>We are happy to get unexpected medal(Our PB rank is 124th). Individually, this is my first medal on kaggle and this medal made me even more into kaggle.</p>\n\n<p>Based on <a href=\"/haqishen\">@haqishen</a>  train &amp; inference notebook, we will mention some of things we tried and a comment about how those works affect on our final result.</p>\n\n<h2>Tile Size Selection</h2>\n\n<ul>\n<li>256x256x36 tiles : best results</li>\n<li>128x128x16 tiles : lower result than 256x256x36 tiles. It seemed necessary to raise the image resolution.</li>\n<li>256x256x16 tiles : better result than 128x128x16 tiles, but not satisfactory.</li>\n<li>256x256x36 tiles with little white as possible : Since 256x256x36 tiles have lots of white spaces, we tried to remove white spaces based on <a href=\"/rftexas\">@rftexas</a>’s <a href=\"https://www.kaggle.com/rftexas/better-image-tiles-removing-white-spaces\">notebook</a>. It achieves lower train loss than simple 256x256x36 tiles but quadratic weighted kappa score did not improved.</li>\n</ul>\n\n<h2>Augmentation</h2>\n\n<p>Several different augmentation were tested (Transpose, VerticalFlip, HorizontalFlip, RandomRotate, Blur, etc), but not much performance improvement was seen.\nJust taking basic augmenation configuration based on <a href=\"/haqishen\">@haqishen</a> ’s notebook.\nAlbumentation library</p>\n\n<blockquote>\n  <p>Transpose(p=0.5)\n  VerticalFlip(p=0.5)\n  HorizontalFlip(p=0.5)\n  All the augmentation were made at 2 levels: tile level + after the tile concatenated</p>\n</blockquote>\n\n<h2>Model</h2>\n\n<p>Similar to other competition (Deep learning for image classification), the most popular model architecture (Resnet, efficientnet) we tried.\nAmong many several Resnet model structure,  SE_Resnext50 was shown the highest score (except more than 50 model because our GPU limitation).\nEfficientnet showed stable and high score at CV.\nWe couldn't get high level of efficientnet model due to our GPU unfortunately , but some discussion let us know that deep and heavy size model will lead to overfit (Effnet b6)\nSo we focus on Efficientnet B0 and B1, between them not much difference shown.</p>\n\n<h2>Optimizer &amp; schedular</h2>\n\n<p>Adam optimzer \nAdam + GradualWarmupScheduler + CosineAnnealingLR</p>\n\n<h2>Inference</h2>\n\n<p>a) TTA(Test Time Augmentation)</p>\n\n<p>Based on tile generation method from Quishen Ha kernel, slight augmentation was added when conducting tile extraction\nThat code is at mode=0 or mode=1 option of PANDA dataset generation class.\nDifference between mode =0 and mode1 is the sequence of tile into the concatenated input (36 x tile).\nSo, when inferencing model, mode1 tile and mode 2 tile were considered as augmented data for test time augmentation(TTA).\nAdditionally, we added transform augmentation in the same way  of train (2 levels, tile + concatenated input).\nFrom several experiments, TTA showed quite positive effects on our public score when increasing the number of augmentation data.\nHowever, considering this competition is code competition which limits the submission time below 9 hours,  we made intermediate number of TTA not as much like more than 100 TTA for preventing over of regular submission time.</p>\n\n<blockquote>\n  <p>16TTA(mode=0) + 16TTA(mode=1)\n  Transpose(p=0.5)\n  VerticalFlip(p=0.5)\n  HorizontalFlip(p=0.5)</p>\n</blockquote>\n\n<p>b) Model Ensemble</p>\n\n<p>The hardest part in this competition was how consider overfit on our training data and how predict shake up from private data.\nWe carefully watched our CV score and LB score and continuously compere them.\nAt last, from the comparison CV and LB for each fold, we got the fold which had the most similar result between CV and LB score.</p>\n\n<p>Ensemble result [Efficient net b0(fold0) and Efficient net b1 (fold0 and fold1)] showed the best score at public score and final(private) score both.</p>",
  "messages": [
    {
      "id": 951587,
      "postDate": "2020-07-30T08:34:09.983Z",
      "content": "<p>First of all Thank you very much to organizers and thanks to <a href=\"/kyunghoonhur\">@kyunghoonhur</a> for collaborating with me!</p>\n\n<p>We are happy to get unexpected medal(Our PB rank is 124th). Individually, this is my first medal on kaggle and this medal made me even more into kaggle.</p>\n\n<p>Based on <a href=\"/haqishen\">@haqishen</a>  train &amp; inference notebook, we will mention some of things we tried and a comment about how those works affect on our final result.</p>\n\n<h2>Tile Size Selection</h2>\n\n<ul>\n<li>256x256x36 tiles : best results</li>\n<li>128x128x16 tiles : lower result than 256x256x36 tiles. It seemed necessary to raise the image resolution.</li>\n<li>256x256x16 tiles : better result than 128x128x16 tiles, but not satisfactory.</li>\n<li>256x256x36 tiles with little white as possible : Since 256x256x36 tiles have lots of white spaces, we tried to remove white spaces based on <a href=\"/rftexas\">@rftexas</a>’s <a href=\"https://www.kaggle.com/rftexas/better-image-tiles-removing-white-spaces\">notebook</a>. It achieves lower train loss than simple 256x256x36 tiles but quadratic weighted kappa score did not improved.</li>\n</ul>\n\n<h2>Augmentation</h2>\n\n<p>Several different augmentation were tested (Transpose, VerticalFlip, HorizontalFlip, RandomRotate, Blur, etc), but not much performance improvement was seen.\nJust taking basic augmenation configuration based on <a href=\"/haqishen\">@haqishen</a> ’s notebook.\nAlbumentation library</p>\n\n<blockquote>\n  <p>Transpose(p=0.5)\n  VerticalFlip(p=0.5)\n  HorizontalFlip(p=0.5)\n  All the augmentation were made at 2 levels: tile level + after the tile concatenated</p>\n</blockquote>\n\n<h2>Model</h2>\n\n<p>Similar to other competition (Deep learning for image classification), the most popular model architecture (Resnet, efficientnet) we tried.\nAmong many several Resnet model structure,  SE_Resnext50 was shown the highest score (except more than 50 model because our GPU limitation).\nEfficientnet showed stable and high score at CV.\nWe couldn't get high level of efficientnet model due to our GPU unfortunately , but some discussion let us know that deep and heavy size model will lead to overfit (Effnet b6)\nSo we focus on Efficientnet B0 and B1, between them not much difference shown.</p>\n\n<h2>Optimizer &amp; schedular</h2>\n\n<p>Adam optimzer \nAdam + GradualWarmupScheduler + CosineAnnealingLR</p>\n\n<h2>Inference</h2>\n\n<p>a) TTA(Test Time Augmentation)</p>\n\n<p>Based on tile generation method from Quishen Ha kernel, slight augmentation was added when conducting tile extraction\nThat code is at mode=0 or mode=1 option of PANDA dataset generation class.\nDifference between mode =0 and mode1 is the sequence of tile into the concatenated input (36 x tile).\nSo, when inferencing model, mode1 tile and mode 2 tile were considered as augmented data for test time augmentation(TTA).\nAdditionally, we added transform augmentation in the same way  of train (2 levels, tile + concatenated input).\nFrom several experiments, TTA showed quite positive effects on our public score when increasing the number of augmentation data.\nHowever, considering this competition is code competition which limits the submission time below 9 hours,  we made intermediate number of TTA not as much like more than 100 TTA for preventing over of regular submission time.</p>\n\n<blockquote>\n  <p>16TTA(mode=0) + 16TTA(mode=1)\n  Transpose(p=0.5)\n  VerticalFlip(p=0.5)\n  HorizontalFlip(p=0.5)</p>\n</blockquote>\n\n<p>b) Model Ensemble</p>\n\n<p>The hardest part in this competition was how consider overfit on our training data and how predict shake up from private data.\nWe carefully watched our CV score and LB score and continuously compere them.\nAt last, from the comparison CV and LB for each fold, we got the fold which had the most similar result between CV and LB score.</p>\n\n<p>Ensemble result [Efficient net b0(fold0) and Efficient net b1 (fold0 and fold1)] showed the best score at public score and final(private) score both.</p>",
      "rawMarkdown": "First of all Thank you very much to organizers and thanks to @kyunghoonhur for collaborating with me!\n\nWe are happy to get unexpected medal(Our PB rank is 124th). Individually, this is my first medal on kaggle and this medal made me even more into kaggle.\n\n\nBased on @haqishen  train &amp; inference notebook, we will mention some of things we tried and a comment about how those works affect on our final result.\n\n## Tile Size Selection \n \n- 256x256x36 tiles : best results\n- 128x128x16 tiles : lower result than 256x256x36 tiles. It seemed necessary to raise the image resolution.\n- 256x256x16 tiles : better result than 128x128x16 tiles, but not satisfactory.\n- 256x256x36 tiles with little white as possible : Since 256x256x36 tiles have lots of white spaces, we tried to remove white spaces based on @rftexas’s [notebook](https://www.kaggle.com/rftexas/better-image-tiles-removing-white-spaces). It achieves lower train loss than simple 256x256x36 tiles but quadratic weighted kappa score did not improved.\n\n## Augmentation\n\nSeveral different augmentation were tested (Transpose, VerticalFlip, HorizontalFlip, RandomRotate, Blur, etc), but not much performance improvement was seen.\nJust taking basic augmenation configuration based on @haqishen ’s notebook.\nAlbumentation library\n&gt; Transpose(p=0.5)\nVerticalFlip(p=0.5)\nHorizontalFlip(p=0.5)\nAll the augmentation were made at 2 levels: tile level + after the tile concatenated\n\n## Model\n\nSimilar to other competition (Deep learning for image classification), the most popular model architecture (Resnet, efficientnet) we tried.\nAmong many several Resnet model structure,  SE_Resnext50 was shown the highest score (except more than 50 model because our GPU limitation).\nEfficientnet showed stable and high score at CV.\nWe couldn't get high level of efficientnet model due to our GPU unfortunately , but some discussion let us know that deep and heavy size model will lead to overfit (Effnet b6)\nSo we focus on Efficientnet B0 and B1, between them not much difference shown.\n\n## Optimizer &amp; schedular\n\nAdam optimzer \nAdam + GradualWarmupScheduler + CosineAnnealingLR\n\n## Inference\n\na) TTA(Test Time Augmentation)\n\nBased on tile generation method from Quishen Ha kernel, slight augmentation was added when conducting tile extraction\nThat code is at mode=0 or mode=1 option of PANDA dataset generation class.\nDifference between mode =0 and mode1 is the sequence of tile into the concatenated input (36 x tile).\nSo, when inferencing model, mode1 tile and mode 2 tile were considered as augmented data for test time augmentation(TTA).\nAdditionally, we added transform augmentation in the same way  of train (2 levels, tile + concatenated input).\nFrom several experiments, TTA showed quite positive effects on our public score when increasing the number of augmentation data.\nHowever, considering this competition is code competition which limits the submission time below 9 hours,  we made intermediate number of TTA not as much like more than 100 TTA for preventing over of regular submission time.\n\n&gt; 16TTA(mode=0) + 16TTA(mode=1)\nTranspose(p=0.5)\nVerticalFlip(p=0.5)\nHorizontalFlip(p=0.5)\n\nb) Model Ensemble\n\nThe hardest part in this competition was how consider overfit on our training data and how predict shake up from private data.\nWe carefully watched our CV score and LB score and continuously compere them.\nAt last, from the comparison CV and LB for each fold, we got the fold which had the most similar result between CV and LB score.\n\nEnsemble result [Efficient net b0(fold0) and Efficient net b1 (fold0 and fold1)] showed the best score at public score and final(private) score both.\n",
      "votes": 6
    },
    {
      "id": 951600,
      "postDate": "2020-07-30T08:50:37.253Z",
      "content": "<p>Big wave shake up led us to the silver medal!! </p>\n\n<p>We will continuously keep going on other competitions too! </p>",
      "rawMarkdown": "Big wave shake up led us to the silver medal!! \n\nWe will continuously keep going on other competitions too! ",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 951600,
      "author_name": "kyunghoon Hur",
      "author_url": "",
      "post_date": "2020-07-30T08:50:37.253000",
      "content": "<p>Big wave shake up led us to the silver medal!! </p>\n\n<p>We will continuously keep going on other competitions too! </p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "951587": "First of all Thank you very much to organizers and thanks to @kyunghoonhur for collaborating with me!\n\nWe are happy to get unexpected medal(Our PB rank is 124th). Individually, this is my first medal on kaggle and this medal made me even more into kaggle.\n\n\nBased on @haqishen  train &amp; inference notebook, we will mention some of things we tried and a comment about how those works affect on our final result.\n\n## Tile Size Selection \n \n- 256x256x36 tiles : best results\n- 128x128x16 tiles : lower result than 256x256x36 tiles. It seemed necessary to raise the image resolution.\n- 256x256x16 tiles : better result than 128x128x16 tiles, but not satisfactory.\n- 256x256x36 tiles with little white as possible : Since 256x256x36 tiles have lots of white spaces, we tried to remove white spaces based on @rftexas’s [notebook](https://www.kaggle.com/rftexas/better-image-tiles-removing-white-spaces). It achieves lower train loss than simple 256x256x36 tiles but quadratic weighted kappa score did not improved.\n\n## Augmentation\n\nSeveral different augmentation were tested (Transpose, VerticalFlip, HorizontalFlip, RandomRotate, Blur, etc), but not much performance improvement was seen.\nJust taking basic augmenation configuration based on @haqishen ’s notebook.\nAlbumentation library\n&gt; Transpose(p=0.5)\nVerticalFlip(p=0.5)\nHorizontalFlip(p=0.5)\nAll the augmentation were made at 2 levels: tile level + after the tile concatenated\n\n## Model\n\nSimilar to other competition (Deep learning for image classification), the most popular model architecture (Resnet, efficientnet) we tried.\nAmong many several Resnet model structure,  SE_Resnext50 was shown the highest score (except more than 50 model because our GPU limitation).\nEfficientnet showed stable and high score at CV.\nWe couldn't get high level of efficientnet model due to our GPU unfortunately , but some discussion let us know that deep and heavy size model will lead to overfit (Effnet b6)\nSo we focus on Efficientnet B0 and B1, between them not much difference shown.\n\n## Optimizer &amp; schedular\n\nAdam optimzer \nAdam + GradualWarmupScheduler + CosineAnnealingLR\n\n## Inference\n\na) TTA(Test Time Augmentation)\n\nBased on tile generation method from Quishen Ha kernel, slight augmentation was added when conducting tile extraction\nThat code is at mode=0 or mode=1 option of PANDA dataset generation class.\nDifference between mode =0 and mode1 is the sequence of tile into the concatenated input (36 x tile).\nSo, when inferencing model, mode1 tile and mode 2 tile were considered as augmented data for test time augmentation(TTA).\nAdditionally, we added transform augmentation in the same way  of train (2 levels, tile + concatenated input).\nFrom several experiments, TTA showed quite positive effects on our public score when increasing the number of augmentation data.\nHowever, considering this competition is code competition which limits the submission time below 9 hours,  we made intermediate number of TTA not as much like more than 100 TTA for preventing over of regular submission time.\n\n&gt; 16TTA(mode=0) + 16TTA(mode=1)\nTranspose(p=0.5)\nVerticalFlip(p=0.5)\nHorizontalFlip(p=0.5)\n\nb) Model Ensemble\n\nThe hardest part in this competition was how consider overfit on our training data and how predict shake up from private data.\nWe carefully watched our CV score and LB score and continuously compere them.\nAt last, from the comparison CV and LB for each fold, we got the fold which had the most similar result between CV and LB score.\n\nEnsemble result [Efficient net b0(fold0) and Efficient net b1 (fold0 and fold1)] showed the best score at public score and final(private) score both.\n",
    "951600": "Big wave shake up led us to the silver medal!! \n\nWe will continuously keep going on other competitions too! "
  }
}