{
  "id": 169107,
  "title": "This is a complete fiasco - N place solution [17th Public]",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/169107",
  "author_name": "Vladimir Groza",
  "post_date": "2020-07-23T00:20:37.941000",
  "votes": 10,
  "comment_count": 1,
  "views": 0,
  "content": "<p><strong>Despite the final standing, I'd like to share this post anyway..</strong>\n*<em>Everything below was written before this collapse was revealed lol</em>*</p>\n\n<p>First thanks to the hosts for making this competition, the most unpredictable and intriguing one in my experience.</p>\n\n<p>Additionally, congratulations to the winners and thanks to all kagglers who competed hard and made me worry about staying in public top10 all the time. Unfortunately, around 2 last weeks before the competition ends all new ideas stopped working for me (but not for many of others), so my concerns about top10 became reality.. So, it was tough but fun!</p>\n\n<p>One more thanks goes to the Dutch service provider HOSTKEY (<a href=\"https://www.hostkey.com/gpu-servers\">https://www.hostkey.com/gpu-servers</a>) that granted me with the free GPU server access during this challenge. This really helped me to investigate and to experiment more broadly .</p>\n\n<h2><strong>Details:</strong></h2>\n\n<p>In the end my solution is quite simple one and based only on the classification, without using the masks.\nMy 2 final solutions are the average and voting ensembles of several models, all trained mostly on the single fold (in fact I added only 2 other models from different folds to increase a bit the stability) with the D4 TTA.</p>\n\n<p><em><strong>Data:</strong></em>\na) Remove several suspicious samples (somewhere from discussions), rearrange similar samples to the same folds (didn't bring too much) &amp; 5-folds\nb) I used only medium resolution and tiles 256x256 (I was really upset when the combination 36x256x256 was announced on the forum lol)\nc) I used slightly different strategy of the tiles sampling : if number of tiles in the image is less than 36 -&gt; add missing tiles by random selection of existing ones. Otherwise, select 36x1.3 = 46 most informative tiles (to engage more available information), take always first 24 of them. And always take random 12 from the rest. It was considered as an additional \"data augmentation / generalisation\" method.\n<strong><em>Augmentations:</em></strong>\nHflip, Vflip, RandRotate, RandBrightnessContrast, ShiftScaleRotate\n<strong><em>Networks architecture:</em></strong>\na) At some point I found that have almost no impact of heavy and fancy backbones (i was also quite limited in resources to train with big BS), so stayed mostly with ResNet18/34 with AdaptiveAvgPool2d. (Included one good ResNeXt50 model in the ensemble)\nb) I extended the baseline networks with additional convolutional block and spatial attention (this constantly improved my CV but made convergence longer). Before this block tensors were reshaped back to the \"tile-level\" as [BS, C, H*sqrt(N_tiles), W*sqrt(N_tiles)]\nc) Gated Attention module after the Avgpool - most important customisation on the architecture level that always helped.\nd) \n<strong><em>Training:</em></strong>\na) big enough batch - 24 or 32\nb) Adam + ReduceOnPlateau or MultiStep\nc) CrossEntropy / Focal Loss\nd) Baseline networks ResNet-18 / 34 / ResNext50_32x4d_swsl\ne) warmup\nf) apex O1</p>\n\n<h3><strong>What didn't work (completely, almost or was the same):</strong></h3>\n\n<ul>\n<li>Regression</li>\n<li>Optimising of Gleason scores directly (N+N) - 10 classes</li>\n<li>Effnets / RegNets / Inception</li>\n<li>MaxPool / GeM</li>\n<li>patch sizes of 128/224/384/512</li>\n<li>Fancy optimisers as Radam/Ralamb/Ranger</li>\n<li>Custom Loss / Combination of Losses / OHEM / LabelSmoothing / HybridCappaLoss</li>\n<li>Stain normalisation (tried to process all training with Vahadane and Macenko from staintools)</li>\n<li>Highest resolution didn't improve the performance</li>\n<li>Tiles to the single big image</li>\n<li>Multi-task learning (tried: 1 - split classes 0/1 vs 2/3/4/5 -&gt; binary + multiclass on each; 2 split -&gt; binary + regression on each)</li>\n<li>Sequence models such as LSTM extension on the features like <a href=\"https://www.nature.com/articles/s41598-020-58467-9\">here</a> or <a href=\"https://www.researchgate.net/publication/323591215_Differentiation_among_prostate_cancer_patients_with_Gleason_score_of_7_using_histopathology_image_and_genomic_data\">here</a>. However, one such model was included in the final ensemble.</li>\n<li>I also tried to cut 4 sets of tiles with the vertical and horizontal shift of 1/2 of tiles size. Observed no improvement either by using them during the training (tried to select the most informative / to increase the tiles number / to periodically select of such different sets by epochs) or inference. I also tried to cut 4 different sets of tiles for each image and create 4 samples from each one on the fly and feed it as independent samples - no improvement</li>\n<li>Use only non-empty tiles / Balance all samples as 90% of non-empty + 10% always empty to standardize the input type</li>\n<li>Inference and averaging on the 36 \"standard\" tiles + most important 18 with 2 variants of shift worked quite good for many of the model on CV but didn't help on LB</li>\n<li>In some experiments I noticed that 36 is not always the optimal tiles number, but it didn't always work</li>\n<li>TTA with brightness/contrast/scaling</li>\n<li>Use segmentation masks (tried to use only the original masks, didn't try to pseudo-label karolinska with radboud-like type - maybe that was the key)</li>\n<li>Could forget to mention something else that I tried, but for sure I just didn't the golden seed!</li>\n</ul>\n\n<p>What was particularly annoying is that common ensembling didn't really work and often inclusion of strong single models didn't improve neither CV nor LB, but who would expect that with QWK!</p>\n\n<p>And what is surprising my best private submission (<strong>0.932</strong>) is the single model that even wasn't included in the ensemble.. </p>\n\n<p>At least i'm happy with the stable and robust solution, that gives constantly 0.91 on local CV, public and private LB.</p>\n\n<p>Cheers, peace, bisou!</p>",
  "messages": [
    {
      "id": 940433,
      "postDate": "2020-07-23T00:20:37.940Z",
      "content": "<p><strong>Despite the final standing, I'd like to share this post anyway..</strong>\n*<em>Everything below was written before this collapse was revealed lol</em>*</p>\n\n<p>First thanks to the hosts for making this competition, the most unpredictable and intriguing one in my experience.</p>\n\n<p>Additionally, congratulations to the winners and thanks to all kagglers who competed hard and made me worry about staying in public top10 all the time. Unfortunately, around 2 last weeks before the competition ends all new ideas stopped working for me (but not for many of others), so my concerns about top10 became reality.. So, it was tough but fun!</p>\n\n<p>One more thanks goes to the Dutch service provider HOSTKEY (<a href=\"https://www.hostkey.com/gpu-servers\">https://www.hostkey.com/gpu-servers</a>) that granted me with the free GPU server access during this challenge. This really helped me to investigate and to experiment more broadly .</p>\n\n<h2><strong>Details:</strong></h2>\n\n<p>In the end my solution is quite simple one and based only on the classification, without using the masks.\nMy 2 final solutions are the average and voting ensembles of several models, all trained mostly on the single fold (in fact I added only 2 other models from different folds to increase a bit the stability) with the D4 TTA.</p>\n\n<p><em><strong>Data:</strong></em>\na) Remove several suspicious samples (somewhere from discussions), rearrange similar samples to the same folds (didn't bring too much) &amp; 5-folds\nb) I used only medium resolution and tiles 256x256 (I was really upset when the combination 36x256x256 was announced on the forum lol)\nc) I used slightly different strategy of the tiles sampling : if number of tiles in the image is less than 36 -&gt; add missing tiles by random selection of existing ones. Otherwise, select 36x1.3 = 46 most informative tiles (to engage more available information), take always first 24 of them. And always take random 12 from the rest. It was considered as an additional \"data augmentation / generalisation\" method.\n<strong><em>Augmentations:</em></strong>\nHflip, Vflip, RandRotate, RandBrightnessContrast, ShiftScaleRotate\n<strong><em>Networks architecture:</em></strong>\na) At some point I found that have almost no impact of heavy and fancy backbones (i was also quite limited in resources to train with big BS), so stayed mostly with ResNet18/34 with AdaptiveAvgPool2d. (Included one good ResNeXt50 model in the ensemble)\nb) I extended the baseline networks with additional convolutional block and spatial attention (this constantly improved my CV but made convergence longer). Before this block tensors were reshaped back to the \"tile-level\" as [BS, C, H*sqrt(N_tiles), W*sqrt(N_tiles)]\nc) Gated Attention module after the Avgpool - most important customisation on the architecture level that always helped.\nd) \n<strong><em>Training:</em></strong>\na) big enough batch - 24 or 32\nb) Adam + ReduceOnPlateau or MultiStep\nc) CrossEntropy / Focal Loss\nd) Baseline networks ResNet-18 / 34 / ResNext50_32x4d_swsl\ne) warmup\nf) apex O1</p>\n\n<h3><strong>What didn't work (completely, almost or was the same):</strong></h3>\n\n<ul>\n<li>Regression</li>\n<li>Optimising of Gleason scores directly (N+N) - 10 classes</li>\n<li>Effnets / RegNets / Inception</li>\n<li>MaxPool / GeM</li>\n<li>patch sizes of 128/224/384/512</li>\n<li>Fancy optimisers as Radam/Ralamb/Ranger</li>\n<li>Custom Loss / Combination of Losses / OHEM / LabelSmoothing / HybridCappaLoss</li>\n<li>Stain normalisation (tried to process all training with Vahadane and Macenko from staintools)</li>\n<li>Highest resolution didn't improve the performance</li>\n<li>Tiles to the single big image</li>\n<li>Multi-task learning (tried: 1 - split classes 0/1 vs 2/3/4/5 -&gt; binary + multiclass on each; 2 split -&gt; binary + regression on each)</li>\n<li>Sequence models such as LSTM extension on the features like <a href=\"https://www.nature.com/articles/s41598-020-58467-9\">here</a> or <a href=\"https://www.researchgate.net/publication/323591215_Differentiation_among_prostate_cancer_patients_with_Gleason_score_of_7_using_histopathology_image_and_genomic_data\">here</a>. However, one such model was included in the final ensemble.</li>\n<li>I also tried to cut 4 sets of tiles with the vertical and horizontal shift of 1/2 of tiles size. Observed no improvement either by using them during the training (tried to select the most informative / to increase the tiles number / to periodically select of such different sets by epochs) or inference. I also tried to cut 4 different sets of tiles for each image and create 4 samples from each one on the fly and feed it as independent samples - no improvement</li>\n<li>Use only non-empty tiles / Balance all samples as 90% of non-empty + 10% always empty to standardize the input type</li>\n<li>Inference and averaging on the 36 \"standard\" tiles + most important 18 with 2 variants of shift worked quite good for many of the model on CV but didn't help on LB</li>\n<li>In some experiments I noticed that 36 is not always the optimal tiles number, but it didn't always work</li>\n<li>TTA with brightness/contrast/scaling</li>\n<li>Use segmentation masks (tried to use only the original masks, didn't try to pseudo-label karolinska with radboud-like type - maybe that was the key)</li>\n<li>Could forget to mention something else that I tried, but for sure I just didn't the golden seed!</li>\n</ul>\n\n<p>What was particularly annoying is that common ensembling didn't really work and often inclusion of strong single models didn't improve neither CV nor LB, but who would expect that with QWK!</p>\n\n<p>And what is surprising my best private submission (<strong>0.932</strong>) is the single model that even wasn't included in the ensemble.. </p>\n\n<p>At least i'm happy with the stable and robust solution, that gives constantly 0.91 on local CV, public and private LB.</p>\n\n<p>Cheers, peace, bisou!</p>",
      "rawMarkdown": "**Despite the final standing, I'd like to share this post anyway..**\n**Everything below was written before this collapse was revealed lol**\n\nFirst thanks to the hosts for making this competition, the most unpredictable and intriguing one in my experience.\n\nAdditionally, congratulations to the winners and thanks to all kagglers who competed hard and made me worry about staying in public top10 all the time. Unfortunately, around 2 last weeks before the competition ends all new ideas stopped working for me (but not for many of others), so my concerns about top10 became reality.. So, it was tough but fun!\n\nOne more thanks goes to the Dutch service provider HOSTKEY (https://www.hostkey.com/gpu-servers) that granted me with the free GPU server access during this challenge. This really helped me to investigate and to experiment more broadly .\n\n## **Details:**\nIn the end my solution is quite simple one and based only on the classification, without using the masks.\nMy 2 final solutions are the average and voting ensembles of several models, all trained mostly on the single fold (in fact I added only 2 other models from different folds to increase a bit the stability) with the D4 TTA.\n\n***Data:***\na) Remove several suspicious samples (somewhere from discussions), rearrange similar samples to the same folds (didn't bring too much) &amp; 5-folds\nb) I used only medium resolution and tiles 256x256 (I was really upset when the combination 36x256x256 was announced on the forum lol)\nc) I used slightly different strategy of the tiles sampling : if number of tiles in the image is less than 36 -&gt; add missing tiles by random selection of existing ones. Otherwise, select 36x1.3 = 46 most informative tiles (to engage more available information), take always first 24 of them. And always take random 12 from the rest. It was considered as an additional \"data augmentation / generalisation\" method.\n***Augmentations:***\nHflip, Vflip, RandRotate, RandBrightnessContrast, ShiftScaleRotate\n***Networks architecture:***\na) At some point I found that have almost no impact of heavy and fancy backbones (i was also quite limited in resources to train with big BS), so stayed mostly with ResNet18/34 with AdaptiveAvgPool2d. (Included one good ResNeXt50 model in the ensemble)\nb) I extended the baseline networks with additional convolutional block and spatial attention (this constantly improved my CV but made convergence longer). Before this block tensors were reshaped back to the \"tile-level\" as [BS, C, H\\*sqrt(N\\_tiles), W\\*sqrt(N\\_tiles)]\nc) Gated Attention module after the Avgpool - most important customisation on the architecture level that always helped.\nd) \n***Training:***\na) big enough batch - 24 or 32\nb) Adam + ReduceOnPlateau or MultiStep\nc) CrossEntropy / Focal Loss\nd) Baseline networks ResNet-18 / 34 / ResNext50\\_32x4d\\_swsl\ne) warmup\nf) apex O1\n\n###**What didn't work (completely, almost or was the same):**\n- Regression\n- Optimising of Gleason scores directly (N+N) - 10 classes\n- Effnets / RegNets / Inception\n- MaxPool / GeM\n- patch sizes of 128/224/384/512\n- Fancy optimisers as Radam/Ralamb/Ranger\n- Custom Loss / Combination of Losses / OHEM / LabelSmoothing / HybridCappaLoss\n- Stain normalisation (tried to process all training with Vahadane and Macenko from staintools)\n- Highest resolution didn't improve the performance\n- Tiles to the single big image\n- Multi-task learning (tried: 1 - split classes 0/1 vs 2/3/4/5 -&gt; binary + multiclass on each; 2 split -&gt; binary + regression on each)\n- Sequence models such as LSTM extension on the features like [here](https://www.nature.com/articles/s41598-020-58467-9) or [here](https://www.researchgate.net/publication/323591215_Differentiation_among_prostate_cancer_patients_with_Gleason_score_of_7_using_histopathology_image_and_genomic_data). However, one such model was included in the final ensemble.\n- I also tried to cut 4 sets of tiles with the vertical and horizontal shift of 1/2 of tiles size. Observed no improvement either by using them during the training (tried to select the most informative / to increase the tiles number / to periodically select of such different sets by epochs) or inference. I also tried to cut 4 different sets of tiles for each image and create 4 samples from each one on the fly and feed it as independent samples - no improvement\n- Use only non-empty tiles / Balance all samples as 90% of non-empty + 10% always empty to standardize the input type\n- Inference and averaging on the 36 \"standard\" tiles + most important 18 with 2 variants of shift worked quite good for many of the model on CV but didn't help on LB\n- In some experiments I noticed that 36 is not always the optimal tiles number, but it didn't always work\n- TTA with brightness/contrast/scaling\n- Use segmentation masks (tried to use only the original masks, didn't try to pseudo-label karolinska with radboud-like type - maybe that was the key)\n- Could forget to mention something else that I tried, but for sure I just didn't the golden seed!\n\nWhat was particularly annoying is that common ensembling didn't really work and often inclusion of strong single models didn't improve neither CV nor LB, but who would expect that with QWK!\n\nAnd what is surprising my best private submission (**0.932**) is the single model that even wasn't included in the ensemble.. \n\nAt least i'm happy with the stable and robust solution, that gives constantly 0.91 on local CV, public and private LB.\n\nCheers, peace, bisou!",
      "votes": 11
    },
    {
      "id": 940459,
      "postDate": "2020-07-23T00:46:20.423Z",
      "content": "<p>I have to agree here. This competition is too random due to the small test set</p>",
      "rawMarkdown": "I have to agree here. This competition is too random due to the small test set",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 940459,
      "author_name": "Shujun",
      "author_url": "",
      "post_date": "2020-07-23T00:46:20.423000",
      "content": "<p>I have to agree here. This competition is too random due to the small test set</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "940433": "**Despite the final standing, I'd like to share this post anyway..**\n**Everything below was written before this collapse was revealed lol**\n\nFirst thanks to the hosts for making this competition, the most unpredictable and intriguing one in my experience.\n\nAdditionally, congratulations to the winners and thanks to all kagglers who competed hard and made me worry about staying in public top10 all the time. Unfortunately, around 2 last weeks before the competition ends all new ideas stopped working for me (but not for many of others), so my concerns about top10 became reality.. So, it was tough but fun!\n\nOne more thanks goes to the Dutch service provider HOSTKEY (https://www.hostkey.com/gpu-servers) that granted me with the free GPU server access during this challenge. This really helped me to investigate and to experiment more broadly .\n\n## **Details:**\nIn the end my solution is quite simple one and based only on the classification, without using the masks.\nMy 2 final solutions are the average and voting ensembles of several models, all trained mostly on the single fold (in fact I added only 2 other models from different folds to increase a bit the stability) with the D4 TTA.\n\n***Data:***\na) Remove several suspicious samples (somewhere from discussions), rearrange similar samples to the same folds (didn't bring too much) &amp; 5-folds\nb) I used only medium resolution and tiles 256x256 (I was really upset when the combination 36x256x256 was announced on the forum lol)\nc) I used slightly different strategy of the tiles sampling : if number of tiles in the image is less than 36 -&gt; add missing tiles by random selection of existing ones. Otherwise, select 36x1.3 = 46 most informative tiles (to engage more available information), take always first 24 of them. And always take random 12 from the rest. It was considered as an additional \"data augmentation / generalisation\" method.\n***Augmentations:***\nHflip, Vflip, RandRotate, RandBrightnessContrast, ShiftScaleRotate\n***Networks architecture:***\na) At some point I found that have almost no impact of heavy and fancy backbones (i was also quite limited in resources to train with big BS), so stayed mostly with ResNet18/34 with AdaptiveAvgPool2d. (Included one good ResNeXt50 model in the ensemble)\nb) I extended the baseline networks with additional convolutional block and spatial attention (this constantly improved my CV but made convergence longer). Before this block tensors were reshaped back to the \"tile-level\" as [BS, C, H\\*sqrt(N\\_tiles), W\\*sqrt(N\\_tiles)]\nc) Gated Attention module after the Avgpool - most important customisation on the architecture level that always helped.\nd) \n***Training:***\na) big enough batch - 24 or 32\nb) Adam + ReduceOnPlateau or MultiStep\nc) CrossEntropy / Focal Loss\nd) Baseline networks ResNet-18 / 34 / ResNext50\\_32x4d\\_swsl\ne) warmup\nf) apex O1\n\n###**What didn't work (completely, almost or was the same):**\n- Regression\n- Optimising of Gleason scores directly (N+N) - 10 classes\n- Effnets / RegNets / Inception\n- MaxPool / GeM\n- patch sizes of 128/224/384/512\n- Fancy optimisers as Radam/Ralamb/Ranger\n- Custom Loss / Combination of Losses / OHEM / LabelSmoothing / HybridCappaLoss\n- Stain normalisation (tried to process all training with Vahadane and Macenko from staintools)\n- Highest resolution didn't improve the performance\n- Tiles to the single big image\n- Multi-task learning (tried: 1 - split classes 0/1 vs 2/3/4/5 -&gt; binary + multiclass on each; 2 split -&gt; binary + regression on each)\n- Sequence models such as LSTM extension on the features like [here](https://www.nature.com/articles/s41598-020-58467-9) or [here](https://www.researchgate.net/publication/323591215_Differentiation_among_prostate_cancer_patients_with_Gleason_score_of_7_using_histopathology_image_and_genomic_data). However, one such model was included in the final ensemble.\n- I also tried to cut 4 sets of tiles with the vertical and horizontal shift of 1/2 of tiles size. Observed no improvement either by using them during the training (tried to select the most informative / to increase the tiles number / to periodically select of such different sets by epochs) or inference. I also tried to cut 4 different sets of tiles for each image and create 4 samples from each one on the fly and feed it as independent samples - no improvement\n- Use only non-empty tiles / Balance all samples as 90% of non-empty + 10% always empty to standardize the input type\n- Inference and averaging on the 36 \"standard\" tiles + most important 18 with 2 variants of shift worked quite good for many of the model on CV but didn't help on LB\n- In some experiments I noticed that 36 is not always the optimal tiles number, but it didn't always work\n- TTA with brightness/contrast/scaling\n- Use segmentation masks (tried to use only the original masks, didn't try to pseudo-label karolinska with radboud-like type - maybe that was the key)\n- Could forget to mention something else that I tried, but for sure I just didn't the golden seed!\n\nWhat was particularly annoying is that common ensembling didn't really work and often inclusion of strong single models didn't improve neither CV nor LB, but who would expect that with QWK!\n\nAnd what is surprising my best private submission (**0.932**) is the single model that even wasn't included in the ensemble.. \n\nAt least i'm happy with the stable and robust solution, that gives constantly 0.91 on local CV, public and private LB.\n\nCheers, peace, bisou!",
    "940459": "I have to agree here. This competition is too random due to the small test set"
  }
}