{
  "id": 169113,
  "title": "4th place solution [NS Pathology]",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/169113",
  "author_name": "hirune924",
  "post_date": "2020-07-23T00:47:08.331000",
  "votes": 40,
  "comment_count": 15,
  "views": 0,
  "content": "<p>First of all Thank you very much to organizers and thanks to <a href=\"/sinpcw\">@sinpcw</a>  for fighting with me!\nOur solution is simple.\nWe ended up with the following models in our final ensemble submission.</p>\n\n<p>&gt; efficientnet_b5 x 4\n&gt; seresnext101-64x4d x2\n&gt; seresnext101-32x4d x2\n&gt; resnest101e x2\n&gt; gem+efficientnet-b3 x 1</p>\n\n<p>We used hard voting for the ensemble method, not soft voting. This definitely improved our score!\nWe also used the average if the one getting the most votes in our hard voting did not get more than 1/3 of the total votes. But this method only worked on publicLB.</p>\n\n<p>In training, We use the technique of tiling the images in the following links. This allows us to ensure that the tissues are evenly distributed across all tiles.It is also able to perform Data Augmentation by changing the scaling factor. We use 512x512x16 from middle layer. We've also tried using 1024x1024x16 from highest resolution layer, but there was no improvement.\n<a href=\"https://www.kaggle.com/hirune924/image-loader-test\">https://www.kaggle.com/hirune924/image-loader-test</a></p>\n\n<p>We also use syncBN. This was important when training the larger models.\nblue line is normal BN, brown line is syncBN.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1626653%2F3adeecd96e2f4fa58855db72e7a93999%2FsyncBN.png?generation=1595467440685526&amp;alt=media\" alt=\"\"></p>\n\n<h3>O2U-Net</h3>\n\n<p>(This method seemed to work in the privateLB, but we didn't use it in the end because we couldn't see the effect in the publicLB)\nWe also tried using O2U-Net to remove the data noise, but didn't work in publicLB.\nBut data cleansing of radboud only seemed to work for privateLB.\nseresnext50 trained on noise removed dataset for only radboud achieves 0.933 in privateLB.(if without data cleansing privateLB 0.915)\n<a href=\"https://openaccess.thecvf.com/content_ICCV_2019/papers/Huang_O2U-Net_A_Simple_Noisy_Label_Detection_Approach_for_Deep_Neural_ICCV_2019_paper.pdf\">https://openaccess.thecvf.com/content_ICCV_2019/papers/Huang_O2U-Net_A_Simple_Noisy_Label_Detection_Approach_for_Deep_Neural_ICCV_2019_paper.pdf</a></p>\n\n<p>I share a notebook that calculates the noise level based on the recorded loss by O2UNet.\n<a href=\"https://www.kaggle.com/hirune924/o2unet-loss-aggregate\">https://www.kaggle.com/hirune924/o2unet-loss-aggregate</a>\nThe effect of data cleansing on private LB is also described in the 1st place solution.\n<a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169143\">https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169143</a></p>\n\n<h3>Usefull tools</h3>\n\n<p>PyTorch Lightning <a href=\"https://github.com/PyTorchLightning/pytorch-lightning\">https://github.com/PyTorchLightning/pytorch-lightning</a>\nHydra <a href=\"https://hydra.cc/\">https://hydra.cc/</a>\nNeptune ai <a href=\"https://neptune.ai/\">https://neptune.ai/</a>\nKAMONOHASHI  <a href=\"https://github.com/KAMONOHASHI\">https://github.com/KAMONOHASHI</a></p>",
  "messages": [
    {
      "id": 940460,
      "postDate": "2020-07-23T00:47:08.330Z",
      "content": "<p>First of all Thank you very much to organizers and thanks to <a href=\"/sinpcw\">@sinpcw</a>  for fighting with me!\nOur solution is simple.\nWe ended up with the following models in our final ensemble submission.</p>\n\n<p>&gt; efficientnet_b5 x 4\n&gt; seresnext101-64x4d x2\n&gt; seresnext101-32x4d x2\n&gt; resnest101e x2\n&gt; gem+efficientnet-b3 x 1</p>\n\n<p>We used hard voting for the ensemble method, not soft voting. This definitely improved our score!\nWe also used the average if the one getting the most votes in our hard voting did not get more than 1/3 of the total votes. But this method only worked on publicLB.</p>\n\n<p>In training, We use the technique of tiling the images in the following links. This allows us to ensure that the tissues are evenly distributed across all tiles.It is also able to perform Data Augmentation by changing the scaling factor. We use 512x512x16 from middle layer. We've also tried using 1024x1024x16 from highest resolution layer, but there was no improvement.\n<a href=\"https://www.kaggle.com/hirune924/image-loader-test\">https://www.kaggle.com/hirune924/image-loader-test</a></p>\n\n<p>We also use syncBN. This was important when training the larger models.\nblue line is normal BN, brown line is syncBN.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1626653%2F3adeecd96e2f4fa58855db72e7a93999%2FsyncBN.png?generation=1595467440685526&amp;alt=media\" alt=\"\"></p>\n\n<h3>O2U-Net</h3>\n\n<p>(This method seemed to work in the privateLB, but we didn't use it in the end because we couldn't see the effect in the publicLB)\nWe also tried using O2U-Net to remove the data noise, but didn't work in publicLB.\nBut data cleansing of radboud only seemed to work for privateLB.\nseresnext50 trained on noise removed dataset for only radboud achieves 0.933 in privateLB.(if without data cleansing privateLB 0.915)\n<a href=\"https://openaccess.thecvf.com/content_ICCV_2019/papers/Huang_O2U-Net_A_Simple_Noisy_Label_Detection_Approach_for_Deep_Neural_ICCV_2019_paper.pdf\">https://openaccess.thecvf.com/content_ICCV_2019/papers/Huang_O2U-Net_A_Simple_Noisy_Label_Detection_Approach_for_Deep_Neural_ICCV_2019_paper.pdf</a></p>\n\n<p>I share a notebook that calculates the noise level based on the recorded loss by O2UNet.\n<a href=\"https://www.kaggle.com/hirune924/o2unet-loss-aggregate\">https://www.kaggle.com/hirune924/o2unet-loss-aggregate</a>\nThe effect of data cleansing on private LB is also described in the 1st place solution.\n<a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169143\">https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169143</a></p>\n\n<h3>Usefull tools</h3>\n\n<p>PyTorch Lightning <a href=\"https://github.com/PyTorchLightning/pytorch-lightning\">https://github.com/PyTorchLightning/pytorch-lightning</a>\nHydra <a href=\"https://hydra.cc/\">https://hydra.cc/</a>\nNeptune ai <a href=\"https://neptune.ai/\">https://neptune.ai/</a>\nKAMONOHASHI  <a href=\"https://github.com/KAMONOHASHI\">https://github.com/KAMONOHASHI</a></p>",
      "rawMarkdown": "First of all Thank you very much to organizers and thanks to @sinpcw  for fighting with me!\nOur solution is simple.\nWe ended up with the following models in our final ensemble submission.\n\n&gt; efficientnet_b5 x 4\n&gt; seresnext101-64x4d x2\n&gt; seresnext101-32x4d x2\n&gt; resnest101e x2\n&gt; gem+efficientnet-b3 x 1\n\nWe used hard voting for the ensemble method, not soft voting. This definitely improved our score!\nWe also used the average if the one getting the most votes in our hard voting did not get more than 1/3 of the total votes. But this method only worked on publicLB.\n\nIn training, We use the technique of tiling the images in the following links. This allows us to ensure that the tissues are evenly distributed across all tiles.It is also able to perform Data Augmentation by changing the scaling factor. We use 512x512x16 from middle layer. We've also tried using 1024x1024x16 from highest resolution layer, but there was no improvement.\nhttps://www.kaggle.com/hirune924/image-loader-test\n\nWe also use syncBN. This was important when training the larger models.\nblue line is normal BN, brown line is syncBN.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1626653%2F3adeecd96e2f4fa58855db72e7a93999%2FsyncBN.png?generation=1595467440685526&amp;alt=media)\n\n### O2U-Net\n(This method seemed to work in the privateLB, but we didn't use it in the end because we couldn't see the effect in the publicLB)\nWe also tried using O2U-Net to remove the data noise, but didn't work in publicLB.\nBut data cleansing of radboud only seemed to work for privateLB.\nseresnext50 trained on noise removed dataset for only radboud achieves 0.933 in privateLB.(if without data cleansing privateLB 0.915)\nhttps://openaccess.thecvf.com/content_ICCV_2019/papers/Huang_O2U-Net_A_Simple_Noisy_Label_Detection_Approach_for_Deep_Neural_ICCV_2019_paper.pdf\n\nI share a notebook that calculates the noise level based on the recorded loss by O2UNet.\nhttps://www.kaggle.com/hirune924/o2unet-loss-aggregate\nThe effect of data cleansing on private LB is also described in the 1st place solution.\nhttps://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169143\n\n### Usefull tools\nPyTorch Lightning https://github.com/PyTorchLightning/pytorch-lightning\nHydra https://hydra.cc/\nNeptune ai https://neptune.ai/\nKAMONOHASHI  https://github.com/KAMONOHASHI",
      "votes": 39
    },
    {
      "id": 940467,
      "postDate": "2020-07-23T01:05:03.093Z",
      "content": "<p>I'm very happy to be in 4th place. I am deeply grateful to my teammate <a href=\"/hirune924\">@hirune924</a> .\nI share list some other topic than the above.</p>\n\n<ul>\n<li><p>segmentation mask\nIt was known that the radboud data has some kind of noise. And our model was overfitting against noise, so we tried to improve using segmentation. we actually learned lightGBM using radboud mask and got validation score 0.94 as in the public kernel. However, we suspected radboud used the Gleason score to change the mask and thought this was a leak.</p></li>\n<li><p>hard-vote\nSince there was no improvement due to the avg ensemble, we spent most of the last week investigating the methods of ensembles. Both hard-vote and the predicted average seemed to have good points, so in the ensemble model uses both methods with threshold switching.</p></li>\n<li><p>model archtecture\nWe also tried other models, but the one in the list written by <a href=\"/hirune924\">@hirune924</a>  was good.</p></li>\n<li><p>pair-coding\nThis competition the test data was invisible and it was difficult to debug. Therefore, we implemented code with the same purpose for each other and checked each other to prevent bug mixing.</p></li>\n</ul>\n\n<p><strong>Not effective TIPS</strong>\n* smoothL1 etc.\nMSE was enough for our training.\n* subtile augmentation\nwe made an augments for each tile, but there was not much difference. However, it may have affected robustness.\n* mixtile\nwe made an augment that mixed some tiles that expected an effect like mixup, but it was not effective.\n* classification\nPrediction by classification, it din't work.\n* use gleason score\n* train/multi-head for each data provider\n* 512 pixels or more per tile</p>",
      "rawMarkdown": "I'm very happy to be in 4th place. I am deeply grateful to my teammate @hirune924 .\nI share list some other topic than the above.\n\n* segmentation mask\nIt was known that the radboud data has some kind of noise. And our model was overfitting against noise, so we tried to improve using segmentation. we actually learned lightGBM using radboud mask and got validation score 0.94 as in the public kernel. However, we suspected radboud used the Gleason score to change the mask and thought this was a leak.\n\n* hard-vote\nSince there was no improvement due to the avg ensemble, we spent most of the last week investigating the methods of ensembles. Both hard-vote and the predicted average seemed to have good points, so in the ensemble model uses both methods with threshold switching.\n\n* model archtecture\nWe also tried other models, but the one in the list written by @hirune924  was good.\n\n* pair-coding\nThis competition the test data was invisible and it was difficult to debug. Therefore, we implemented code with the same purpose for each other and checked each other to prevent bug mixing.\n\n**Not effective TIPS**\n* smoothL1 etc.\nMSE was enough for our training.\n* subtile augmentation\nwe made an augments for each tile, but there was not much difference. However, it may have affected robustness.\n* mixtile\nwe made an augment that mixed some tiles that expected an effect like mixup, but it was not effective.\n* classification\nPrediction by classification, it din't work.\n* use gleason score\n* train/multi-head for each data provider\n* 512 pixels or more per tile\n",
      "votes": 3,
      "replies": [
        {
          "id": 941771,
          "postDate": "2020-07-23T12:05:48.067Z",
          "content": "<p>So when you say hard voting, it's like [5,5,5,3,3] will give 5, while soft-voting is just an average hence 21/5, right ?</p>\n\n<p>\"However, we suspected radboud used the Gleason score to change the mask and thought this was a leak.\" : \nI strongly suspect you are right with this, I changed Iafoss' approach to tile selection to use the masks for selecting the tiles, and I got a huge boost on CV, but not at all on LB (where I had to use a segmentation model to get then mask and then only pick the tiles)</p>",
          "rawMarkdown": "So when you say hard voting, it's like [5,5,5,3,3] will give 5, while soft-voting is just an average hence 21/5, right ?\n\n\"However, we suspected radboud used the Gleason score to change the mask and thought this was a leak.\" : \nI strongly suspect you are right with this, I changed Iafoss' approach to tile selection to use the masks for selecting the tiles, and I got a huge boost on CV, but not at all on LB (where I had to use a segmentation model to get then mask and then only pick the tiles)"
        },
        {
          "id": 942168,
          "postDate": "2020-07-23T16:06:17.573Z",
          "content": "<p>Yes, that's right. We are observing that some models prediction has outlier in same data, such as a very different inference value (e.g. pred 0 / label 5). It was assumed that these effects would result in poor results if we used the mean value.</p>",
          "rawMarkdown": "Yes, that's right. We are observing that some models prediction has outlier in same data, such as a very different inference value (e.g. pred 0 / label 5). It was assumed that these effects would result in poor results if we used the mean value."
        }
      ]
    },
    {
      "id": 941378,
      "postDate": "2020-07-23T07:37:02.717Z",
      "content": "<h2>Robustness strategy for shake</h2>\n\n<p>We were afraid to shake in this competition from the beginning. So we were very careful in our choice of methods.\nFirst of all, we worked hard to create a good validation, but no matter what we did, we couldn't create a validation that would work with LB. We were able to get our local CV score up to over 0.95.\nWe figured this was due to the fact that there are many similar images in the train data. And there are so many of them that I've given up on removing them. So we decided to trust publicLB.\nIt's easy to create an over-fitted submission to the publicLB by submitting a similar approach over and over again, but that's not the true ability of that approach.\nWe have tried many techniques such as tiling methods, data cleansing, use of masks, per-provider learning, loss, custom architectures, etc., but none of them clearly improved the publicLB.\nHowever, I noticed that using a model larger than seresnext50, I was able to consistently exceed 0.90 at publicLB.\nSo we finally thought that we could achieve consistently high scores in public and private LB by using multiple large models. I also noticed that using hardvoting instead of avg at the end can improve the stability of publicLB.\nThis simple idea seemed to be correct.\nWe also took the lessons learned from the APTOS shake down and made sure that the tiles are distributed evenly without tissues being cut off.</p>",
      "rawMarkdown": "## Robustness strategy for shake\nWe were afraid to shake in this competition from the beginning. So we were very careful in our choice of methods.\nFirst of all, we worked hard to create a good validation, but no matter what we did, we couldn't create a validation that would work with LB. We were able to get our local CV score up to over 0.95.\nWe figured this was due to the fact that there are many similar images in the train data. And there are so many of them that I've given up on removing them. So we decided to trust publicLB.\nIt's easy to create an over-fitted submission to the publicLB by submitting a similar approach over and over again, but that's not the true ability of that approach.\nWe have tried many techniques such as tiling methods, data cleansing, use of masks, per-provider learning, loss, custom architectures, etc., but none of them clearly improved the publicLB.\nHowever, I noticed that using a model larger than seresnext50, I was able to consistently exceed 0.90 at publicLB.\nSo we finally thought that we could achieve consistently high scores in public and private LB by using multiple large models. I also noticed that using hardvoting instead of avg at the end can improve the stability of publicLB.\nThis simple idea seemed to be correct.\nWe also took the lessons learned from the APTOS shake down and made sure that the tiles are distributed evenly without tissues being cut off.",
      "votes": 2
    },
    {
      "id": 942991,
      "postDate": "2020-07-24T05:07:01.670Z",
      "content": "<p>Congrats and thanks for sharing your solution!\nHow did you do the Train/Val split?</p>",
      "rawMarkdown": "Congrats and thanks for sharing your solution!\nHow did you do the Train/Val split?",
      "replies": [
        {
          "id": 943059,
          "postDate": "2020-07-24T06:00:04Z",
          "content": "<p>This is code snipet for Train Val spilit\n ``` \nkf = sklearn.model_selection.StratifiedKFold(n_splits=10, shuffle=true, random_state=2020)</p>\n\n<p>for fold, (train_index, val_index) in enumerate(kf.split(df.values, df[\"isup_grade\"].astype(str) + df[\"data_provider\"],)):\n　　df.loc[val_index, \"fold\"] = int(fold)\ndf[\"fold\"] = df[\"fold\"].astype(int)</p>\n\n<p>train_df = df[df[\"fold\"] != 1]\nvalid_df = df[df[\"fold\"] == 1]\n<code>\nWe also calculated many types of validations by trimming the validation dataset, which helped us to estimate the performance of the model.\n</code>\navg_val_loss (simple val loss)\nval_acc (simple val accuracy)\nval_qwk (simple val qwk)\nkarolinska_qwk (val qwk using only karolinska)\nradboud_qwk (val qwk using only radboud)\nsample_qwk (val qwk using only isup&gt;1 )\nval_qwk_o (observed of simple val qwk)\nval_qwk_e (expected of simple val qwk)\npublic_sim_qwk (val qwk using only ((data_provider == 'karolinska') &amp; (isup &gt; 2.5)) | ((data_provider == 'radboud') &amp; (isup &lt; 2.5)))\nprivate_sim_qwk (val qwk using only ((data_provider == 'radboud') &amp; (isup &gt; 2.5)) | ((data_provider == 'karolinska') &amp; (isup &lt; 2.5)))\n```</p>",
          "rawMarkdown": "This is code snipet for Train Val spilit\n ``` \nkf = sklearn.model_selection.StratifiedKFold(n_splits=10, shuffle=true, random_state=2020)\n\nfor fold, (train_index, val_index) in enumerate(kf.split(df.values, df[\"isup_grade\"].astype(str) + df[\"data_provider\"],)):\n　　df.loc[val_index, \"fold\"] = int(fold)\ndf[\"fold\"] = df[\"fold\"].astype(int)\n\ntrain_df = df[df[\"fold\"] != 1]\nvalid_df = df[df[\"fold\"] == 1]\n```\nWe also calculated many types of validations by trimming the validation dataset, which helped us to estimate the performance of the model.\n```\navg_val_loss (simple val loss)\nval_acc (simple val accuracy)\nval_qwk (simple val qwk)\nkarolinska_qwk (val qwk using only karolinska)\nradboud_qwk (val qwk using only radboud)\nsample_qwk (val qwk using only isup&gt;1 )\nval_qwk_o (observed of simple val qwk)\nval_qwk_e (expected of simple val qwk)\npublic_sim_qwk (val qwk using only ((data_provider == 'karolinska') &amp; (isup &gt; 2.5)) | ((data_provider == 'radboud') &amp; (isup &lt; 2.5)))\nprivate_sim_qwk (val qwk using only ((data_provider == 'radboud') &amp; (isup &gt; 2.5)) | ((data_provider == 'karolinska') &amp; (isup &lt; 2.5)))\n```"
        },
        {
          "id": 943066,
          "postDate": "2020-07-24T06:05:44.620Z",
          "content": "<p><a href=\"/hirune924\">@hirune924</a> thx! It's very helpful.</p>",
          "rawMarkdown": "@hirune924 thx! It's very helpful."
        }
      ]
    },
    {
      "id": 941386,
      "postDate": "2020-07-23T07:42:44.620Z",
      "content": "<p>Interesting that you used hard voting for ensamble! That will greatly reduce variances.\nHow much did LB/PB change with that technique?</p>",
      "rawMarkdown": "Interesting that you used hard voting for ensamble! That will greatly reduce variances.\nHow much did LB/PB change with that technique?",
      "replies": [
        {
          "id": 941410,
          "postDate": "2020-07-23T08:00:00.213Z",
          "content": "<p>Using hardvoting improved pubLB about 0.004. The effect on privateLB seems to be a bit less, but at least it didn't have a negative effect.\nI think this is probably because using hardvoting is less susceptible to the negative effects of lower-performing models than using avg.</p>",
          "rawMarkdown": "Using hardvoting improved pubLB about 0.004. The effect on privateLB seems to be a bit less, but at least it didn't have a negative effect.\nI think this is probably because using hardvoting is less susceptible to the negative effects of lower-performing models than using avg.",
          "replies": [
            {
              "id": 941532,
              "postDate": "2020-07-23T09:23:44.690Z",
              "content": "<p>Congrats! Hard voting also worked quite good for me on public, and equally stable or a bit less on private.</p>",
              "rawMarkdown": "Congrats! Hard voting also worked quite good for me on public, and equally stable or a bit less on private."
            }
          ]
        }
      ]
    },
    {
      "id": 940586,
      "postDate": "2020-07-23T03:24:43.503Z",
      "content": "<p>congrats for sharing your solution approach.  can you discuss more about O2U net performance. how many noisy labels were you able to detect and remove using it.</p>",
      "rawMarkdown": "congrats for sharing your solution approach.  can you discuss more about O2U net performance. how many noisy labels were you able to detect and remove using it.",
      "replies": [
        {
          "id": 940611,
          "postDate": "2020-07-23T03:49:35.090Z",
          "content": "<p>I added about O2UNet details.</p>",
          "rawMarkdown": "I added about O2UNet details.",
          "votes": 1
        }
      ]
    },
    {
      "id": 940499,
      "postDate": "2020-07-23T01:44:42.787Z",
      "content": "<p>Congrats and thanks for sharing!\nDo you have any insights about why syncBN is so important?</p>",
      "rawMarkdown": "Congrats and thanks for sharing!\nDo you have any insights about why syncBN is so important?\n",
      "replies": [
        {
          "id": 940503,
          "postDate": "2020-07-23T01:47:01.253Z",
          "content": "<p>We trained with mini batch size=2x4 using 4 GPUs.\nIn this case, it seems that normal BN can not accurately estimate the statistics of the batches</p>",
          "rawMarkdown": "We trained with mini batch size=2x4 using 4 GPUs.\nIn this case, it seems that normal BN can not accurately estimate the statistics of the batches",
          "votes": 1
        }
      ]
    },
    {
      "id": 940492,
      "postDate": "2020-07-23T01:33:10.143Z",
      "content": "<p>Great solutions! Thanks for sharing. I am looking forward your O2Unet.</p>",
      "rawMarkdown": "Great solutions! Thanks for sharing. I am looking forward your O2Unet."
    }
  ],
  "comments": [
    {
      "id": 940467,
      "author_name": "SiNpcw",
      "author_url": "",
      "post_date": "2020-07-23T01:05:03.093000",
      "content": "<p>I'm very happy to be in 4th place. I am deeply grateful to my teammate <a href=\"/hirune924\">@hirune924</a> .\nI share list some other topic than the above.</p>\n\n<ul>\n<li><p>segmentation mask\nIt was known that the radboud data has some kind of noise. And our model was overfitting against noise, so we tried to improve using segmentation. we actually learned lightGBM using radboud mask and got validation score 0.94 as in the public kernel. However, we suspected radboud used the Gleason score to change the mask and thought this was a leak.</p></li>\n<li><p>hard-vote\nSince there was no improvement due to the avg ensemble, we spent most of the last week investigating the methods of ensembles. Both hard-vote and the predicted average seemed to have good points, so in the ensemble model uses both methods with threshold switching.</p></li>\n<li><p>model archtecture\nWe also tried other models, but the one in the list written by <a href=\"/hirune924\">@hirune924</a>  was good.</p></li>\n<li><p>pair-coding\nThis competition the test data was invisible and it was difficult to debug. Therefore, we implemented code with the same purpose for each other and checked each other to prevent bug mixing.</p></li>\n</ul>\n\n<p><strong>Not effective TIPS</strong>\n* smoothL1 etc.\nMSE was enough for our training.\n* subtile augmentation\nwe made an augments for each tile, but there was not much difference. However, it may have affected robustness.\n* mixtile\nwe made an augment that mixed some tiles that expected an effect like mixup, but it was not effective.\n* classification\nPrediction by classification, it din't work.\n* use gleason score\n* train/multi-head for each data provider\n* 512 pixels or more per tile</p>",
      "votes": 3,
      "replies": [
        {
          "id": 941771,
          "author_name": "Benjamin Dubreu",
          "author_url": "",
          "post_date": "2020-07-23T12:05:48.067000",
          "content": "<p>So when you say hard voting, it's like [5,5,5,3,3] will give 5, while soft-voting is just an average hence 21/5, right ?</p>\n\n<p>\"However, we suspected radboud used the Gleason score to change the mask and thought this was a leak.\" : \nI strongly suspect you are right with this, I changed Iafoss' approach to tile selection to use the masks for selecting the tiles, and I got a huge boost on CV, but not at all on LB (where I had to use a segmentation model to get then mask and then only pick the tiles)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 942168,
          "author_name": "SiNpcw",
          "author_url": "",
          "post_date": "2020-07-23T16:06:17.573000",
          "content": "<p>Yes, that's right. We are observing that some models prediction has outlier in same data, such as a very different inference value (e.g. pred 0 / label 5). It was assumed that these effects would result in poor results if we used the mean value.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 941378,
      "author_name": "hirune924",
      "author_url": "",
      "post_date": "2020-07-23T07:37:02.717000",
      "content": "<h2>Robustness strategy for shake</h2>\n\n<p>We were afraid to shake in this competition from the beginning. So we were very careful in our choice of methods.\nFirst of all, we worked hard to create a good validation, but no matter what we did, we couldn't create a validation that would work with LB. We were able to get our local CV score up to over 0.95.\nWe figured this was due to the fact that there are many similar images in the train data. And there are so many of them that I've given up on removing them. So we decided to trust publicLB.\nIt's easy to create an over-fitted submission to the publicLB by submitting a similar approach over and over again, but that's not the true ability of that approach.\nWe have tried many techniques such as tiling methods, data cleansing, use of masks, per-provider learning, loss, custom architectures, etc., but none of them clearly improved the publicLB.\nHowever, I noticed that using a model larger than seresnext50, I was able to consistently exceed 0.90 at publicLB.\nSo we finally thought that we could achieve consistently high scores in public and private LB by using multiple large models. I also noticed that using hardvoting instead of avg at the end can improve the stability of publicLB.\nThis simple idea seemed to be correct.\nWe also took the lessons learned from the APTOS shake down and made sure that the tiles are distributed evenly without tissues being cut off.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 942991,
      "author_name": "fam_taro",
      "author_url": "",
      "post_date": "2020-07-24T05:07:01.670000",
      "content": "<p>Congrats and thanks for sharing your solution!\nHow did you do the Train/Val split?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 943059,
          "author_name": "hirune924",
          "author_url": "",
          "post_date": "2020-07-24T06:00:04",
          "content": "<p>This is code snipet for Train Val spilit\n ``` \nkf = sklearn.model_selection.StratifiedKFold(n_splits=10, shuffle=true, random_state=2020)</p>\n\n<p>for fold, (train_index, val_index) in enumerate(kf.split(df.values, df[\"isup_grade\"].astype(str) + df[\"data_provider\"],)):\n　　df.loc[val_index, \"fold\"] = int(fold)\ndf[\"fold\"] = df[\"fold\"].astype(int)</p>\n\n<p>train_df = df[df[\"fold\"] != 1]\nvalid_df = df[df[\"fold\"] == 1]\n<code>\nWe also calculated many types of validations by trimming the validation dataset, which helped us to estimate the performance of the model.\n</code>\navg_val_loss (simple val loss)\nval_acc (simple val accuracy)\nval_qwk (simple val qwk)\nkarolinska_qwk (val qwk using only karolinska)\nradboud_qwk (val qwk using only radboud)\nsample_qwk (val qwk using only isup&gt;1 )\nval_qwk_o (observed of simple val qwk)\nval_qwk_e (expected of simple val qwk)\npublic_sim_qwk (val qwk using only ((data_provider == 'karolinska') &amp; (isup &gt; 2.5)) | ((data_provider == 'radboud') &amp; (isup &lt; 2.5)))\nprivate_sim_qwk (val qwk using only ((data_provider == 'radboud') &amp; (isup &gt; 2.5)) | ((data_provider == 'karolinska') &amp; (isup &lt; 2.5)))\n```</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 943066,
          "author_name": "fam_taro",
          "author_url": "",
          "post_date": "2020-07-24T06:05:44.620000",
          "content": "<p><a href=\"/hirune924\">@hirune924</a> thx! It's very helpful.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 941386,
      "author_name": "arutema47",
      "author_url": "",
      "post_date": "2020-07-23T07:42:44.620000",
      "content": "<p>Interesting that you used hard voting for ensamble! That will greatly reduce variances.\nHow much did LB/PB change with that technique?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 941410,
          "author_name": "hirune924",
          "author_url": "",
          "post_date": "2020-07-23T08:00:00.213000",
          "content": "<p>Using hardvoting improved pubLB about 0.004. The effect on privateLB seems to be a bit less, but at least it didn't have a negative effect.\nI think this is probably because using hardvoting is less susceptible to the negative effects of lower-performing models than using avg.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 941532,
              "author_name": "Vladimir Groza",
              "author_url": "",
              "post_date": "2020-07-23T09:23:44.690000",
              "content": "<p>Congrats! Hard voting also worked quite good for me on public, and equally stable or a bit less on private.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 940586,
      "author_name": "yuvaramsingh",
      "author_url": "",
      "post_date": "2020-07-23T03:24:43.503000",
      "content": "<p>congrats for sharing your solution approach.  can you discuss more about O2U net performance. how many noisy labels were you able to detect and remove using it.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 940611,
          "author_name": "hirune924",
          "author_url": "",
          "post_date": "2020-07-23T03:49:35.090000",
          "content": "<p>I added about O2UNet details.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 940499,
      "author_name": "Camaro",
      "author_url": "",
      "post_date": "2020-07-23T01:44:42.787000",
      "content": "<p>Congrats and thanks for sharing!\nDo you have any insights about why syncBN is so important?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 940503,
          "author_name": "hirune924",
          "author_url": "",
          "post_date": "2020-07-23T01:47:01.253000",
          "content": "<p>We trained with mini batch size=2x4 using 4 GPUs.\nIn this case, it seems that normal BN can not accurately estimate the statistics of the batches</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 940492,
      "author_name": "Nishi yasu",
      "author_url": "",
      "post_date": "2020-07-23T01:33:10.143000",
      "content": "<p>Great solutions! Thanks for sharing. I am looking forward your O2Unet.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "940460": "First of all Thank you very much to organizers and thanks to @sinpcw  for fighting with me!\nOur solution is simple.\nWe ended up with the following models in our final ensemble submission.\n\n&gt; efficientnet_b5 x 4\n&gt; seresnext101-64x4d x2\n&gt; seresnext101-32x4d x2\n&gt; resnest101e x2\n&gt; gem+efficientnet-b3 x 1\n\nWe used hard voting for the ensemble method, not soft voting. This definitely improved our score!\nWe also used the average if the one getting the most votes in our hard voting did not get more than 1/3 of the total votes. But this method only worked on publicLB.\n\nIn training, We use the technique of tiling the images in the following links. This allows us to ensure that the tissues are evenly distributed across all tiles.It is also able to perform Data Augmentation by changing the scaling factor. We use 512x512x16 from middle layer. We've also tried using 1024x1024x16 from highest resolution layer, but there was no improvement.\nhttps://www.kaggle.com/hirune924/image-loader-test\n\nWe also use syncBN. This was important when training the larger models.\nblue line is normal BN, brown line is syncBN.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1626653%2F3adeecd96e2f4fa58855db72e7a93999%2FsyncBN.png?generation=1595467440685526&amp;alt=media)\n\n### O2U-Net\n(This method seemed to work in the privateLB, but we didn't use it in the end because we couldn't see the effect in the publicLB)\nWe also tried using O2U-Net to remove the data noise, but didn't work in publicLB.\nBut data cleansing of radboud only seemed to work for privateLB.\nseresnext50 trained on noise removed dataset for only radboud achieves 0.933 in privateLB.(if without data cleansing privateLB 0.915)\nhttps://openaccess.thecvf.com/content_ICCV_2019/papers/Huang_O2U-Net_A_Simple_Noisy_Label_Detection_Approach_for_Deep_Neural_ICCV_2019_paper.pdf\n\nI share a notebook that calculates the noise level based on the recorded loss by O2UNet.\nhttps://www.kaggle.com/hirune924/o2unet-loss-aggregate\nThe effect of data cleansing on private LB is also described in the 1st place solution.\nhttps://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169143\n\n### Usefull tools\nPyTorch Lightning https://github.com/PyTorchLightning/pytorch-lightning\nHydra https://hydra.cc/\nNeptune ai https://neptune.ai/\nKAMONOHASHI  https://github.com/KAMONOHASHI",
    "940467": "I'm very happy to be in 4th place. I am deeply grateful to my teammate @hirune924 .\nI share list some other topic than the above.\n\n* segmentation mask\nIt was known that the radboud data has some kind of noise. And our model was overfitting against noise, so we tried to improve using segmentation. we actually learned lightGBM using radboud mask and got validation score 0.94 as in the public kernel. However, we suspected radboud used the Gleason score to change the mask and thought this was a leak.\n\n* hard-vote\nSince there was no improvement due to the avg ensemble, we spent most of the last week investigating the methods of ensembles. Both hard-vote and the predicted average seemed to have good points, so in the ensemble model uses both methods with threshold switching.\n\n* model archtecture\nWe also tried other models, but the one in the list written by @hirune924  was good.\n\n* pair-coding\nThis competition the test data was invisible and it was difficult to debug. Therefore, we implemented code with the same purpose for each other and checked each other to prevent bug mixing.\n\n**Not effective TIPS**\n* smoothL1 etc.\nMSE was enough for our training.\n* subtile augmentation\nwe made an augments for each tile, but there was not much difference. However, it may have affected robustness.\n* mixtile\nwe made an augment that mixed some tiles that expected an effect like mixup, but it was not effective.\n* classification\nPrediction by classification, it din't work.\n* use gleason score\n* train/multi-head for each data provider\n* 512 pixels or more per tile\n",
    "941378": "## Robustness strategy for shake\nWe were afraid to shake in this competition from the beginning. So we were very careful in our choice of methods.\nFirst of all, we worked hard to create a good validation, but no matter what we did, we couldn't create a validation that would work with LB. We were able to get our local CV score up to over 0.95.\nWe figured this was due to the fact that there are many similar images in the train data. And there are so many of them that I've given up on removing them. So we decided to trust publicLB.\nIt's easy to create an over-fitted submission to the publicLB by submitting a similar approach over and over again, but that's not the true ability of that approach.\nWe have tried many techniques such as tiling methods, data cleansing, use of masks, per-provider learning, loss, custom architectures, etc., but none of them clearly improved the publicLB.\nHowever, I noticed that using a model larger than seresnext50, I was able to consistently exceed 0.90 at publicLB.\nSo we finally thought that we could achieve consistently high scores in public and private LB by using multiple large models. I also noticed that using hardvoting instead of avg at the end can improve the stability of publicLB.\nThis simple idea seemed to be correct.\nWe also took the lessons learned from the APTOS shake down and made sure that the tiles are distributed evenly without tissues being cut off.",
    "942991": "Congrats and thanks for sharing your solution!\nHow did you do the Train/Val split?",
    "941386": "Interesting that you used hard voting for ensamble! That will greatly reduce variances.\nHow much did LB/PB change with that technique?",
    "940586": "congrats for sharing your solution approach.  can you discuss more about O2U net performance. how many noisy labels were you able to detect and remove using it.",
    "940499": "Congrats and thanks for sharing!\nDo you have any insights about why syncBN is so important?\n",
    "940492": "Great solutions! Thanks for sharing. I am looking forward your O2Unet."
  }
}