{
  "id": 169205,
  "title": "[11th place solution] I have survived in this storm",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/169205",
  "author_name": "Iafoss",
  "post_date": "2020-07-23T07:56:22.284000",
  "votes": 125,
  "comment_count": 62,
  "views": 0,
  "content": "<h2>Summary</h2>\n<ul>\n<li>Tile extraction is based on my <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/146855\" target=\"_blank\">public pipeline</a> with <strong>128x128x128</strong> tiles from intermediate resolution layer</li>\n<li>Label nose removal gives <strong>~0.005 public and 0.01+ private LB boost</strong></li>\n<li>tile cutout + tile selection augmentations</li>\n<li><a href=\"https://arxiv.org/pdf/1509.07107v2.pdf\" target=\"_blank\">kappa loss</a></li>\n<li>majority voting ensemble of 8 <strong>ResNeXt50</strong> based models (<strong>0.917 public and 0.930 private LB</strong>)</li>\n<li>more advanced tile selection could give <strong>~0.004 boost</strong> at private LB on average (and the maximum private LB score of <strong>0.941</strong>)</li>\n</ul>\n<h2>Introduction</h2>\n<p>To begin with, I would really like to express my gratitude to organizes and kaggle team for making this competition possible. It was really enjoying working on it and learned many new things. By sharing some of my ideas in this competition, such as <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/146855\" target=\"_blank\">tile pooling base pipeline</a> used by many participants, gave me 3 kernel gold medals, so I have reached the kernel grand-master rank. And I have received my first solo competition gold medal. Also, I would like to say congratulations to all winners and people who received medals.</p>\n<p>However, this day is quite sad for many participants, especially ones who worked very hard throughout the entire competition and got down at private LB. The <strong>chose of the metric by organizers could be done more wisely</strong>: 500+500 test set is definitely not enough for QWK. It is not really normal when LB score is changing by 0.005+ when a different seed is used. The things got deteriorated when the third digit became available for LB score: many people got seduced by overfitting LB noise.</p>\n<p>Below I outline the main things that worked for me. I have tried many more, but most of them have never worked, and I couldn't get any further improvement of my LB during the last month.</p>\n<h2>Main challenges</h2>\n<p>This competition to a large extent was about dealing with noisy data and train/test bias: as reported by organizers, the Redbound train data has only about <strong>0.853</strong> QWK, and I expect that Karolinska train data has 0.95-0.96 QWK. Beyond this, since Redbound data is graded by students, and Karolinska data is graded by only a single expert, while the test data is graded by 3 experts, there could be train/test bias because of the subjective opinion of people performing grading the train set. Therefore, <strong>solely relaying on CV was not really good strategy in this competition</strong>: at some point I saw a consistent decrease (~10 different models) of LB score when I ran training for longer, while CV was increasing. It confirms the hypothesis about the bias, and the trick was to train models only for limited number of epochs (even if CV could be increased), 32-48 depending on the setup, to <strong>prevent learning the bias</strong>.</p>\n<p>Meanwhile, LB was also not the best thing to trust because of severe noise, but some ppl tried to fit random seed as a hyperparameter 😄. The right thing, in my opinion, in this competition was to find the balance between CV and LB, and <strong>trust to your intuition and the experience gained in the previous competitions</strong>.</p>\n<h2>Noise</h2>\n<p>It is the most important part of this competition, in my opinion. After organizers have disclosed that there is a substantial level of noise, especially in Redbound train data, I have explored a number of techniques to deal with the noise: progressive label distillation, JoCoR (Joint Training with Co-Regularization), Co-teaching, negative learning, excluding hard examples from the batch, etc. However, most of them didn’t really work well here. The additional challenge is the bias between train and test and unstable LB. The thing I found to be the best for this data is removal of the uncertain examples from training set based on the out of fold predictions. I excluded ~1400 Redbound and 300 Karalinska data, so my clean training set contains about 8700 items. Relabeling the excluded images didn’t improve the performance. At the end of the competition <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/161909\" target=\"_blank\">some ppl have discovered this trick as well</a>, so I got nervous about my LB position 😬 <br>\n<strong>This trick gave ~0.005 public LB boost and 0.01+ private LB boost.</strong></p>\n<h2>Pipeline</h2>\n<p>The method I have used is mainly based on my <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/146855\" target=\"_blank\">tile pooling pipeline</a> with several additional tricks:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1212661%2Fe6fe32d759a28480343001aa3c661723%2FTILE.png?generation=1588094975239255&amp;alt=media\" alt=\"\"></p>\n<p>Based on my public kernel, one could reach ~0.90 public and 0.91 private LB averaged (over different submissions) using 36x256x256 tile setup and the kappa loss (see below) without any other changes.</p>\n<p><a href=\"https://arxiv.org/pdf/1509.07107v2.pdf\" target=\"_blank\"><strong>kappa loss</strong></a>: I have used one minus<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1212661%2F99945116e2c9228e352645ee5f0bdfcc%2F2.png?generation=1595472932908419&amp;alt=media\" alt=\"\"></p>\n<p>(both predictions and labels are centered based on the mean value of labels). In my experiments I found that kappa loss &gt; sorted <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/155424\" target=\"_blank\">binning loss</a> &gt; binning loss &gt; MSE &gt; CE. The only issue with the loss is that bs should be sufficiently large: I needed to pretrain models on low resolution and then continue training on intermediate resolution with bs =6-8 (<strong>progressive resizing</strong>), while at bs 1-2 I couldn't get convergence. The predicted value is limited within [-0.5,5.5] as <code>yp = 6*sigmoid(p) - 0.5</code>. In addition, I have CE aux for prediction of the Gleason score with 0.08 weight.</p>\n<p><strong>tile cutout</strong>:  <strong>Instead of using all tile tiles, why not to randomly select part of them</strong> (let's say 96 out of 128). So, I can use large bs and the model is regularized in the same way as if cutout is used. It gave me quite good boost for CV and quite fair boost at LB.</p>\n<p><strong>128x128x128 tiles from intermediate resolution</strong>: It appeared that many smaller tiles work better than 36x256x256. I think that it helps to select the tissue areas more effectively and at the same time prevents overfitting. </p>\n<p><strong>tile selection augmentation</strong> The idea is quite simple: instead of <a href=\"https://www.kaggle.com/iafoss/panda-16x128x128-tiles\" target=\"_blank\">generating a single tile set</a>, I can generate 4 with adding sz/2 padding to x, y, or both before cutting the image into tiles and selecting ones having the most of the tissue. So, each tile in these 4 datasets will be different, but it is important not to mix tiles from them. During training I select the dataset by random, so effectively I have x4 data. It is an approximation of tile selection with random offset each time, which would be even more effective (based on my experience in Severstal competition), though, too slow to be used with intermediate res images. I also tried TTA based on tile selection, (as well as selection of the tile set with the largest tissue area out of 4), but I couldn't get any statistically significant improvement.</p>\n<p><strong>The above tricks gave ~0.005 boost over my baseline if I consider multiple submissions</strong>. Though, score from submission to submission could change quite a bit.</p>\n<p><strong>Advanced tile selection</strong>: In addition to my main pipeline I also tried to use the method proposed by <a href=\"https://www.kaggle.com/akensert\" target=\"_blank\">@akensert</a> <a href=\"https://www.kaggle.com/akensert/panda-optimized-tiling-tf-data-dataset\" target=\"_blank\">here</a> with 128x128 tiles (but didn't use it for my final sub). It gave ~0.934 private LB single model performance on average for 8 different single 4 fold model subs (<strong>and maximum 0.941 private LB</strong>) and ~0.910 average public LB (~0.912 maximum). Too bad that I didn't create an ensemble based on this method. The trick that I have used in the model for training with such tiles is <strong>n-pooling</strong>: at test I apply the pooling only to nonempty tiles (n for a particular image, while the batch may contain some extra empty tiles for padding), and training is done with random selection of 96 tiles with repetitions (so I don't consider white tiles, which can change the mean statistics at pooling).</p>\n<p><strong>High resolution</strong>: I tried to train several models on high res/2 resolution, 128x256x256 tiles. With tile cutout I could use batches of size 4 (and include 64 random tiles). However, the results were slightly worse than ones for 128x128x128 tiles from the intermediate resolution layer. It indicates that <strong>going to higher resolution would likely provide only a minor boost</strong>, even if I try to optimize my pipeline for training with small batches. Some idea I had is based on having two conv parts for intermediate and high res tiles. First pass through the low res model selects tiles having the highest uncertainty. Next, the selected tiles (but in high res) are passed through the second conv part. The produced feature maps are downscaled twice and replace the low res feature maps that had high uncertainty. Finally, pooling and head are applied to produce the final prediction. This method would allow to keep overall statistics of tiles with only correcting ones that model is not confident about. However, too large level of noise in the training set, noisy LB inconsistent with train labeling, and small potential gain, which would likely be overshadowed by the noise, have prevented me from going into this direction. Also, more complicated pipeline is more likely to be broken under such competition, where there is no certain way to evaluate the performance.</p>\n<p><strong>Augmentation</strong>: I have used Albumentations with the following parameters:</p>\n<pre><code>Compose([\n        HorizontalFlip(),\n        VerticalFlip(),\n        RandomRotate90(),\n        ShiftScaleRotate(shift_limit=0.0625, scale_limit=0.3, rotate_limit=15, p=0.9, \n                         border_mode=cv2.BORDER_CONSTANT),\n        OneOf([#off in most cases\n            MotionBlur(blur_limit=3, p=0.1),\n            MedianBlur(blur_limit=3, p=0.1),\n            Blur(blur_limit=3, p=0.1),\n        ], p=0.2),\n        OneOf([#off in most cases\n            OpticalDistortion(p=0.3),\n            GridDistortion(p=.1),\n            IAAPiecewiseAffine(p=0.3),\n        ], p=0.3),\n        OneOf([\n            HueSaturationValue(10,15,10),\n            CLAHE(clip_limit=2),\n            RandomBrightnessContrast(),            \n        ], p=0.3),\n    ], p=1)\n</code></pre>\n<p><strong>Model</strong>: All my models are based on <strong>ResNeXt50</strong>, similar to my public kernel, with batch norm in the head replaced with Group-norm. The optimizer, best model selection based on CV, and other things are similar to my public kernel, and I was using 32-48 epochs, depending on the setup. In addition, I tried ResNet34, ResNeXt101, and EfficientNet, while all of them were performing worse. I think ResNet34 may be not capable enough for this task, while ResNeXt101 is too large to do training on my computer with sufficient bs. However, I would say that <strong>the model is the minor thing in this competition, and the main role is played by considering the noise and by optimizing the pipeline: there is no magic model, but there are hard work and solid understanding of the task and the data</strong>.</p>\n<h2>Final ensemble</h2>\n<p>The submission that gave me the 11th place (<strong>0.930 private LB/0.917 public LB</strong>) is based on a majority voting ensemble of 8 models (4 fold) with 6 TTA. They are trained with different train/val splits and other modifications in the training procedure. On average each of the models trained in such manner gave <strong>~0.930 private and ~0.910 public LB</strong> single model 4 fold performance (with the <strong>maximum of 0.938 and 0.916</strong>, respectively). So, I got quite fair score, not good or bad luck (and my LB position almost haven’t changed). However, the large number of models was a way to survive in this storm. My another ensemble of 11 models, not selected as a final one, got 0.934 private LB. And as I mentioned above, more advanced tiling gives about <strong>0.004 boost</strong> on private LB (while similar public score as my main approach based on 128x128x128 tiles), with the average of <strong>~0.934</strong> and the maximum of <strong>0.941 private LB</strong>, but unfortunately, I haven’t built an ensemble based on them for my final submissions.</p>\n<p><strong>the code snippets are available at:</strong> <a href=\"https://github.com/iafoss/PANDA\" target=\"_blank\">https://github.com/iafoss/PANDA</a></p>\n<p>And I would like to congratulate all participants and wish the best luck in the next competitions. I hope some of my tricks would be useful to you.</p>",
  "messages": [
    {
      "id": 941402,
      "postDate": "2020-07-23T07:56:22.283Z",
      "content": "<h2>Summary</h2>\n<ul>\n<li>Tile extraction is based on my <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/146855\" target=\"_blank\">public pipeline</a> with <strong>128x128x128</strong> tiles from intermediate resolution layer</li>\n<li>Label nose removal gives <strong>~0.005 public and 0.01+ private LB boost</strong></li>\n<li>tile cutout + tile selection augmentations</li>\n<li><a href=\"https://arxiv.org/pdf/1509.07107v2.pdf\" target=\"_blank\">kappa loss</a></li>\n<li>majority voting ensemble of 8 <strong>ResNeXt50</strong> based models (<strong>0.917 public and 0.930 private LB</strong>)</li>\n<li>more advanced tile selection could give <strong>~0.004 boost</strong> at private LB on average (and the maximum private LB score of <strong>0.941</strong>)</li>\n</ul>\n<h2>Introduction</h2>\n<p>To begin with, I would really like to express my gratitude to organizes and kaggle team for making this competition possible. It was really enjoying working on it and learned many new things. By sharing some of my ideas in this competition, such as <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/146855\" target=\"_blank\">tile pooling base pipeline</a> used by many participants, gave me 3 kernel gold medals, so I have reached the kernel grand-master rank. And I have received my first solo competition gold medal. Also, I would like to say congratulations to all winners and people who received medals.</p>\n<p>However, this day is quite sad for many participants, especially ones who worked very hard throughout the entire competition and got down at private LB. The <strong>chose of the metric by organizers could be done more wisely</strong>: 500+500 test set is definitely not enough for QWK. It is not really normal when LB score is changing by 0.005+ when a different seed is used. The things got deteriorated when the third digit became available for LB score: many people got seduced by overfitting LB noise.</p>\n<p>Below I outline the main things that worked for me. I have tried many more, but most of them have never worked, and I couldn't get any further improvement of my LB during the last month.</p>\n<h2>Main challenges</h2>\n<p>This competition to a large extent was about dealing with noisy data and train/test bias: as reported by organizers, the Redbound train data has only about <strong>0.853</strong> QWK, and I expect that Karolinska train data has 0.95-0.96 QWK. Beyond this, since Redbound data is graded by students, and Karolinska data is graded by only a single expert, while the test data is graded by 3 experts, there could be train/test bias because of the subjective opinion of people performing grading the train set. Therefore, <strong>solely relaying on CV was not really good strategy in this competition</strong>: at some point I saw a consistent decrease (~10 different models) of LB score when I ran training for longer, while CV was increasing. It confirms the hypothesis about the bias, and the trick was to train models only for limited number of epochs (even if CV could be increased), 32-48 depending on the setup, to <strong>prevent learning the bias</strong>.</p>\n<p>Meanwhile, LB was also not the best thing to trust because of severe noise, but some ppl tried to fit random seed as a hyperparameter 😄. The right thing, in my opinion, in this competition was to find the balance between CV and LB, and <strong>trust to your intuition and the experience gained in the previous competitions</strong>.</p>\n<h2>Noise</h2>\n<p>It is the most important part of this competition, in my opinion. After organizers have disclosed that there is a substantial level of noise, especially in Redbound train data, I have explored a number of techniques to deal with the noise: progressive label distillation, JoCoR (Joint Training with Co-Regularization), Co-teaching, negative learning, excluding hard examples from the batch, etc. However, most of them didn’t really work well here. The additional challenge is the bias between train and test and unstable LB. The thing I found to be the best for this data is removal of the uncertain examples from training set based on the out of fold predictions. I excluded ~1400 Redbound and 300 Karalinska data, so my clean training set contains about 8700 items. Relabeling the excluded images didn’t improve the performance. At the end of the competition <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/161909\" target=\"_blank\">some ppl have discovered this trick as well</a>, so I got nervous about my LB position 😬 <br>\n<strong>This trick gave ~0.005 public LB boost and 0.01+ private LB boost.</strong></p>\n<h2>Pipeline</h2>\n<p>The method I have used is mainly based on my <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/146855\" target=\"_blank\">tile pooling pipeline</a> with several additional tricks:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1212661%2Fe6fe32d759a28480343001aa3c661723%2FTILE.png?generation=1588094975239255&amp;alt=media\" alt=\"\"></p>\n<p>Based on my public kernel, one could reach ~0.90 public and 0.91 private LB averaged (over different submissions) using 36x256x256 tile setup and the kappa loss (see below) without any other changes.</p>\n<p><a href=\"https://arxiv.org/pdf/1509.07107v2.pdf\" target=\"_blank\"><strong>kappa loss</strong></a>: I have used one minus<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1212661%2F99945116e2c9228e352645ee5f0bdfcc%2F2.png?generation=1595472932908419&amp;alt=media\" alt=\"\"></p>\n<p>(both predictions and labels are centered based on the mean value of labels). In my experiments I found that kappa loss &gt; sorted <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/155424\" target=\"_blank\">binning loss</a> &gt; binning loss &gt; MSE &gt; CE. The only issue with the loss is that bs should be sufficiently large: I needed to pretrain models on low resolution and then continue training on intermediate resolution with bs =6-8 (<strong>progressive resizing</strong>), while at bs 1-2 I couldn't get convergence. The predicted value is limited within [-0.5,5.5] as <code>yp = 6*sigmoid(p) - 0.5</code>. In addition, I have CE aux for prediction of the Gleason score with 0.08 weight.</p>\n<p><strong>tile cutout</strong>:  <strong>Instead of using all tile tiles, why not to randomly select part of them</strong> (let's say 96 out of 128). So, I can use large bs and the model is regularized in the same way as if cutout is used. It gave me quite good boost for CV and quite fair boost at LB.</p>\n<p><strong>128x128x128 tiles from intermediate resolution</strong>: It appeared that many smaller tiles work better than 36x256x256. I think that it helps to select the tissue areas more effectively and at the same time prevents overfitting. </p>\n<p><strong>tile selection augmentation</strong> The idea is quite simple: instead of <a href=\"https://www.kaggle.com/iafoss/panda-16x128x128-tiles\" target=\"_blank\">generating a single tile set</a>, I can generate 4 with adding sz/2 padding to x, y, or both before cutting the image into tiles and selecting ones having the most of the tissue. So, each tile in these 4 datasets will be different, but it is important not to mix tiles from them. During training I select the dataset by random, so effectively I have x4 data. It is an approximation of tile selection with random offset each time, which would be even more effective (based on my experience in Severstal competition), though, too slow to be used with intermediate res images. I also tried TTA based on tile selection, (as well as selection of the tile set with the largest tissue area out of 4), but I couldn't get any statistically significant improvement.</p>\n<p><strong>The above tricks gave ~0.005 boost over my baseline if I consider multiple submissions</strong>. Though, score from submission to submission could change quite a bit.</p>\n<p><strong>Advanced tile selection</strong>: In addition to my main pipeline I also tried to use the method proposed by <a href=\"https://www.kaggle.com/akensert\" target=\"_blank\">@akensert</a> <a href=\"https://www.kaggle.com/akensert/panda-optimized-tiling-tf-data-dataset\" target=\"_blank\">here</a> with 128x128 tiles (but didn't use it for my final sub). It gave ~0.934 private LB single model performance on average for 8 different single 4 fold model subs (<strong>and maximum 0.941 private LB</strong>) and ~0.910 average public LB (~0.912 maximum). Too bad that I didn't create an ensemble based on this method. The trick that I have used in the model for training with such tiles is <strong>n-pooling</strong>: at test I apply the pooling only to nonempty tiles (n for a particular image, while the batch may contain some extra empty tiles for padding), and training is done with random selection of 96 tiles with repetitions (so I don't consider white tiles, which can change the mean statistics at pooling).</p>\n<p><strong>High resolution</strong>: I tried to train several models on high res/2 resolution, 128x256x256 tiles. With tile cutout I could use batches of size 4 (and include 64 random tiles). However, the results were slightly worse than ones for 128x128x128 tiles from the intermediate resolution layer. It indicates that <strong>going to higher resolution would likely provide only a minor boost</strong>, even if I try to optimize my pipeline for training with small batches. Some idea I had is based on having two conv parts for intermediate and high res tiles. First pass through the low res model selects tiles having the highest uncertainty. Next, the selected tiles (but in high res) are passed through the second conv part. The produced feature maps are downscaled twice and replace the low res feature maps that had high uncertainty. Finally, pooling and head are applied to produce the final prediction. This method would allow to keep overall statistics of tiles with only correcting ones that model is not confident about. However, too large level of noise in the training set, noisy LB inconsistent with train labeling, and small potential gain, which would likely be overshadowed by the noise, have prevented me from going into this direction. Also, more complicated pipeline is more likely to be broken under such competition, where there is no certain way to evaluate the performance.</p>\n<p><strong>Augmentation</strong>: I have used Albumentations with the following parameters:</p>\n<pre><code>Compose([\n        HorizontalFlip(),\n        VerticalFlip(),\n        RandomRotate90(),\n        ShiftScaleRotate(shift_limit=0.0625, scale_limit=0.3, rotate_limit=15, p=0.9, \n                         border_mode=cv2.BORDER_CONSTANT),\n        OneOf([#off in most cases\n            MotionBlur(blur_limit=3, p=0.1),\n            MedianBlur(blur_limit=3, p=0.1),\n            Blur(blur_limit=3, p=0.1),\n        ], p=0.2),\n        OneOf([#off in most cases\n            OpticalDistortion(p=0.3),\n            GridDistortion(p=.1),\n            IAAPiecewiseAffine(p=0.3),\n        ], p=0.3),\n        OneOf([\n            HueSaturationValue(10,15,10),\n            CLAHE(clip_limit=2),\n            RandomBrightnessContrast(),            \n        ], p=0.3),\n    ], p=1)\n</code></pre>\n<p><strong>Model</strong>: All my models are based on <strong>ResNeXt50</strong>, similar to my public kernel, with batch norm in the head replaced with Group-norm. The optimizer, best model selection based on CV, and other things are similar to my public kernel, and I was using 32-48 epochs, depending on the setup. In addition, I tried ResNet34, ResNeXt101, and EfficientNet, while all of them were performing worse. I think ResNet34 may be not capable enough for this task, while ResNeXt101 is too large to do training on my computer with sufficient bs. However, I would say that <strong>the model is the minor thing in this competition, and the main role is played by considering the noise and by optimizing the pipeline: there is no magic model, but there are hard work and solid understanding of the task and the data</strong>.</p>\n<h2>Final ensemble</h2>\n<p>The submission that gave me the 11th place (<strong>0.930 private LB/0.917 public LB</strong>) is based on a majority voting ensemble of 8 models (4 fold) with 6 TTA. They are trained with different train/val splits and other modifications in the training procedure. On average each of the models trained in such manner gave <strong>~0.930 private and ~0.910 public LB</strong> single model 4 fold performance (with the <strong>maximum of 0.938 and 0.916</strong>, respectively). So, I got quite fair score, not good or bad luck (and my LB position almost haven’t changed). However, the large number of models was a way to survive in this storm. My another ensemble of 11 models, not selected as a final one, got 0.934 private LB. And as I mentioned above, more advanced tiling gives about <strong>0.004 boost</strong> on private LB (while similar public score as my main approach based on 128x128x128 tiles), with the average of <strong>~0.934</strong> and the maximum of <strong>0.941 private LB</strong>, but unfortunately, I haven’t built an ensemble based on them for my final submissions.</p>\n<p><strong>the code snippets are available at:</strong> <a href=\"https://github.com/iafoss/PANDA\" target=\"_blank\">https://github.com/iafoss/PANDA</a></p>\n<p>And I would like to congratulate all participants and wish the best luck in the next competitions. I hope some of my tricks would be useful to you.</p>",
      "rawMarkdown": "## Summary\n- Tile extraction is based on my [public pipeline](https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/146855) with **128x128x128** tiles from intermediate resolution layer\n- Label nose removal gives **~0.005 public and 0.01+ private LB boost**\n- tile cutout + tile selection augmentations\n- [kappa loss](https://arxiv.org/pdf/1509.07107v2.pdf)\n- majority voting ensemble of 8 **ResNeXt50** based models (**0.917 public and 0.930 private LB**)\n- more advanced tile selection could give **~0.004 boost** at private LB on average (and the maximum private LB score of **0.941**)\n\n\n## Introduction\n\nTo begin with, I would really like to express my gratitude to organizes and kaggle team for making this competition possible. It was really enjoying working on it and learned many new things. By sharing some of my ideas in this competition, such as [tile pooling base pipeline](https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/146855) used by many participants, gave me 3 kernel gold medals, so I have reached the kernel grand-master rank. And I have received my first solo competition gold medal. Also, I would like to say congratulations to all winners and people who received medals.\n\nHowever, this day is quite sad for many participants, especially ones who worked very hard throughout the entire competition and got down at private LB. The **chose of the metric by organizers could be done more wisely**: 500+500 test set is definitely not enough for QWK. It is not really normal when LB score is changing by 0.005+ when a different seed is used. The things got deteriorated when the third digit became available for LB score: many people got seduced by overfitting LB noise.\n\nBelow I outline the main things that worked for me. I have tried many more, but most of them have never worked, and I couldn't get any further improvement of my LB during the last month.\n\n## Main challenges\n\nThis competition to a large extent was about dealing with noisy data and train/test bias: as reported by organizers, the Redbound train data has only about **0.853** QWK, and I expect that Karolinska train data has 0.95-0.96 QWK. Beyond this, since Redbound data is graded by students, and Karolinska data is graded by only a single expert, while the test data is graded by 3 experts, there could be train/test bias because of the subjective opinion of people performing grading the train set. Therefore, **solely relaying on CV was not really good strategy in this competition**: at some point I saw a consistent decrease (~10 different models) of LB score when I ran training for longer, while CV was increasing. It confirms the hypothesis about the bias, and the trick was to train models only for limited number of epochs (even if CV could be increased), 32-48 depending on the setup, to **prevent learning the bias**.\n\nMeanwhile, LB was also not the best thing to trust because of severe noise, but some ppl tried to fit random seed as a hyperparameter 😄. The right thing, in my opinion, in this competition was to find the balance between CV and LB, and **trust to your intuition and the experience gained in the previous competitions**.\n\n## Noise\n\nIt is the most important part of this competition, in my opinion. After organizers have disclosed that there is a substantial level of noise, especially in Redbound train data, I have explored a number of techniques to deal with the noise: progressive label distillation, JoCoR (Joint Training with Co-Regularization), Co-teaching, negative learning, excluding hard examples from the batch, etc. However, most of them didn’t really work well here. The additional challenge is the bias between train and test and unstable LB. The thing I found to be the best for this data is removal of the uncertain examples from training set based on the out of fold predictions. I excluded ~1400 Redbound and 300 Karalinska data, so my clean training set contains about 8700 items. Relabeling the excluded images didn’t improve the performance. At the end of the competition [some ppl have discovered this trick as well](https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/161909), so I got nervous about my LB position 😬 \n**This trick gave ~0.005 public LB boost and 0.01+ private LB boost.**\n\n## Pipeline\n\nThe method I have used is mainly based on my [tile pooling pipeline](https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/146855) with several additional tricks:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1212661%2Fe6fe32d759a28480343001aa3c661723%2FTILE.png?generation=1588094975239255&amp;alt=media)\n\nBased on my public kernel, one could reach ~0.90 public and 0.91 private LB averaged (over different submissions) using 36x256x256 tile setup and the kappa loss (see below) without any other changes.\n\n[**kappa loss**](https://arxiv.org/pdf/1509.07107v2.pdf): I have used one minus\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1212661%2F99945116e2c9228e352645ee5f0bdfcc%2F2.png?generation=1595472932908419&amp;alt=media)\n\n(both predictions and labels are centered based on the mean value of labels). In my experiments I found that kappa loss &gt; sorted [binning loss](https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/155424) &gt; binning loss &gt; MSE &gt; CE. The only issue with the loss is that bs should be sufficiently large: I needed to pretrain models on low resolution and then continue training on intermediate resolution with bs =6-8 (**progressive resizing**), while at bs 1-2 I couldn't get convergence. The predicted value is limited within [-0.5,5.5] as `yp = 6*sigmoid(p) - 0.5`. In addition, I have CE aux for prediction of the Gleason score with 0.08 weight.\n\n**tile cutout**:  **Instead of using all tile tiles, why not to randomly select part of them** (let's say 96 out of 128). So, I can use large bs and the model is regularized in the same way as if cutout is used. It gave me quite good boost for CV and quite fair boost at LB.\n\n**128x128x128 tiles from intermediate resolution**: It appeared that many smaller tiles work better than 36x256x256. I think that it helps to select the tissue areas more effectively and at the same time prevents overfitting. \n\n**tile selection augmentation** The idea is quite simple: instead of [generating a single tile set](https://www.kaggle.com/iafoss/panda-16x128x128-tiles), I can generate 4 with adding sz/2 padding to x, y, or both before cutting the image into tiles and selecting ones having the most of the tissue. So, each tile in these 4 datasets will be different, but it is important not to mix tiles from them. During training I select the dataset by random, so effectively I have x4 data. It is an approximation of tile selection with random offset each time, which would be even more effective (based on my experience in Severstal competition), though, too slow to be used with intermediate res images. I also tried TTA based on tile selection, (as well as selection of the tile set with the largest tissue area out of 4), but I couldn't get any statistically significant improvement.\n\n**The above tricks gave ~0.005 boost over my baseline if I consider multiple submissions**. Though, score from submission to submission could change quite a bit.\n\n**Advanced tile selection**: In addition to my main pipeline I also tried to use the method proposed by @akensert [here](https://www.kaggle.com/akensert/panda-optimized-tiling-tf-data-dataset) with 128x128 tiles (but didn't use it for my final sub). It gave ~0.934 private LB single model performance on average for 8 different single 4 fold model subs (**and maximum 0.941 private LB**) and ~0.910 average public LB (~0.912 maximum). Too bad that I didn't create an ensemble based on this method. The trick that I have used in the model for training with such tiles is **n-pooling**: at test I apply the pooling only to nonempty tiles (n for a particular image, while the batch may contain some extra empty tiles for padding), and training is done with random selection of 96 tiles with repetitions (so I don't consider white tiles, which can change the mean statistics at pooling).\n\n**High resolution**: I tried to train several models on high res/2 resolution, 128x256x256 tiles. With tile cutout I could use batches of size 4 (and include 64 random tiles). However, the results were slightly worse than ones for 128x128x128 tiles from the intermediate resolution layer. It indicates that **going to higher resolution would likely provide only a minor boost**, even if I try to optimize my pipeline for training with small batches. Some idea I had is based on having two conv parts for intermediate and high res tiles. First pass through the low res model selects tiles having the highest uncertainty. Next, the selected tiles (but in high res) are passed through the second conv part. The produced feature maps are downscaled twice and replace the low res feature maps that had high uncertainty. Finally, pooling and head are applied to produce the final prediction. This method would allow to keep overall statistics of tiles with only correcting ones that model is not confident about. However, too large level of noise in the training set, noisy LB inconsistent with train labeling, and small potential gain, which would likely be overshadowed by the noise, have prevented me from going into this direction. Also, more complicated pipeline is more likely to be broken under such competition, where there is no certain way to evaluate the performance.\n\n**Augmentation**: I have used Albumentations with the following parameters:\n```\nCompose([\n        HorizontalFlip(),\n        VerticalFlip(),\n        RandomRotate90(),\n        ShiftScaleRotate(shift_limit=0.0625, scale_limit=0.3, rotate_limit=15, p=0.9, \n                         border_mode=cv2.BORDER_CONSTANT),\n        OneOf([#off in most cases\n            MotionBlur(blur_limit=3, p=0.1),\n            MedianBlur(blur_limit=3, p=0.1),\n            Blur(blur_limit=3, p=0.1),\n        ], p=0.2),\n        OneOf([#off in most cases\n            OpticalDistortion(p=0.3),\n            GridDistortion(p=.1),\n            IAAPiecewiseAffine(p=0.3),\n        ], p=0.3),\n        OneOf([\n            HueSaturationValue(10,15,10),\n            CLAHE(clip_limit=2),\n            RandomBrightnessContrast(),            \n        ], p=0.3),\n    ], p=1)\n```\n\n**Model**: All my models are based on **ResNeXt50**, similar to my public kernel, with batch norm in the head replaced with Group-norm. The optimizer, best model selection based on CV, and other things are similar to my public kernel, and I was using 32-48 epochs, depending on the setup. In addition, I tried ResNet34, ResNeXt101, and EfficientNet, while all of them were performing worse. I think ResNet34 may be not capable enough for this task, while ResNeXt101 is too large to do training on my computer with sufficient bs. However, I would say that **the model is the minor thing in this competition, and the main role is played by considering the noise and by optimizing the pipeline: there is no magic model, but there are hard work and solid understanding of the task and the data**.\n\n\n## Final ensemble\n\nThe submission that gave me the 11th place (**0.930 private LB/0.917 public LB**) is based on a majority voting ensemble of 8 models (4 fold) with 6 TTA. They are trained with different train/val splits and other modifications in the training procedure. On average each of the models trained in such manner gave **~0.930 private and ~0.910 public LB** single model 4 fold performance (with the **maximum of 0.938 and 0.916**, respectively). So, I got quite fair score, not good or bad luck (and my LB position almost haven’t changed). However, the large number of models was a way to survive in this storm. My another ensemble of 11 models, not selected as a final one, got 0.934 private LB. And as I mentioned above, more advanced tiling gives about **0.004 boost** on private LB (while similar public score as my main approach based on 128x128x128 tiles), with the average of **~0.934** and the maximum of **0.941 private LB**, but unfortunately, I haven’t built an ensemble based on them for my final submissions.\n\n**the code snippets are available at:** https://github.com/iafoss/PANDA\n\nAnd I would like to congratulate all participants and wish the best luck in the next competitions. I hope some of my tricks would be useful to you.\n",
      "votes": 125
    },
    {
      "id": 941428,
      "postDate": "2020-07-23T08:05:10.673Z",
      "content": "<p>Congrats solo gold !!!\nI think you are the player who has contributed the most to this competition. :)\nYour tile method helps almost player.</p>",
      "rawMarkdown": "Congrats solo gold !!!\nI think you are the player who has contributed the most to this competition. :)\nYour tile method helps almost player.",
      "votes": 5,
      "replies": [
        {
          "id": 941719,
          "postDate": "2020-07-23T11:24:50.860Z",
          "content": "<p>Agreed. This is a well deserved solo gold !</p>",
          "rawMarkdown": "Agreed. This is a well deserved solo gold !",
          "votes": 1
        },
        {
          "id": 942375,
          "postDate": "2020-07-23T18:03:10.987Z",
          "content": "<p>Thank you so much, I'm happy about it.</p>",
          "rawMarkdown": "Thank you so much, I'm happy about it.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1217213,
      "postDate": "2021-02-24T21:07:27.383Z",
      "content": "<p>Congratulation! (after 7 months, I know I'm late here lol) I really love your approach!! Thank you so much for sharing those grate notebooks with explanations!</p>",
      "rawMarkdown": "Congratulation! (after 7 months, I know I'm late here lol) I really love your approach!! Thank you so much for sharing those grate notebooks with explanations!",
      "votes": 1,
      "replies": [
        {
          "id": 1217387,
          "postDate": "2021-02-25T02:20:41.553Z",
          "content": "<p>You are very welcome</p>",
          "rawMarkdown": "You are very welcome"
        }
      ]
    },
    {
      "id": 964939,
      "postDate": "2020-08-10T09:27:03.447Z",
      "content": "<p>Amazing  content. Keep sharing </p>",
      "rawMarkdown": "Amazing  content. Keep sharing ",
      "votes": 1,
      "replies": [
        {
          "id": 965415,
          "postDate": "2020-08-10T15:56:16.437Z",
          "content": "<p>Thank you.</p>",
          "rawMarkdown": "Thank you."
        }
      ]
    },
    {
      "id": 953091,
      "postDate": "2020-07-31T14:25:33.923Z",
      "content": "<p>Congratulations <a href=\"/iafoss\">@iafoss</a> for success on this kernal. Keep sharing nice things like this. 💕</p>",
      "rawMarkdown": "Congratulations @iafoss for success on this kernal. Keep sharing nice things like this. 💕",
      "votes": 1,
      "replies": [
        {
          "id": 953161,
          "postDate": "2020-07-31T15:18:59.843Z",
          "content": "<p>Thank you</p>",
          "rawMarkdown": "Thank you"
        }
      ]
    },
    {
      "id": 951890,
      "postDate": "2020-07-30T13:38:25.480Z",
      "content": "<p>Congratulations <a href=\"/iafoss\">@iafoss</a> for this achievement. And thank for sharing the details.👍 </p>",
      "rawMarkdown": "Congratulations @iafoss for this achievement. And thank for sharing the details.👍 ",
      "votes": 1,
      "replies": [
        {
          "id": 952204,
          "postDate": "2020-07-30T17:52:01.397Z",
          "content": "<p>You re very welcome</p>",
          "rawMarkdown": "You re very welcome"
        }
      ]
    },
    {
      "id": 951849,
      "postDate": "2020-07-30T12:58:04.433Z",
      "content": "<p>Congratulations :-)</p>",
      "rawMarkdown": "Congratulations :-)",
      "votes": 1,
      "replies": [
        {
          "id": 952205,
          "postDate": "2020-07-30T17:52:16.567Z",
          "content": "<p>Thanks</p>",
          "rawMarkdown": "Thanks"
        }
      ]
    },
    {
      "id": 951694,
      "postDate": "2020-07-30T10:14:17.373Z",
      "content": "<p>Congratulations! Learned a lot from your kernels!\nThank you!</p>",
      "rawMarkdown": "Congratulations! Learned a lot from your kernels!\nThank you!",
      "votes": 1,
      "replies": [
        {
          "id": 952211,
          "postDate": "2020-07-30T17:54:16.347Z",
          "content": "<p>Thank you so much, you are very welcome.</p>",
          "rawMarkdown": "Thank you so much, you are very welcome."
        }
      ]
    },
    {
      "id": 949602,
      "postDate": "2020-07-28T18:36:46.423Z",
      "content": "<p>Congrats!</p>",
      "rawMarkdown": "Congrats!",
      "votes": 1,
      "replies": [
        {
          "id": 952207,
          "postDate": "2020-07-30T17:52:36.353Z",
          "content": "<p>Thank you</p>",
          "rawMarkdown": "Thank you"
        }
      ]
    },
    {
      "id": 949583,
      "postDate": "2020-07-28T18:17:38.900Z",
      "content": "<p>Congratulations !!!</p>",
      "rawMarkdown": "Congratulations !!!",
      "votes": 1,
      "replies": [
        {
          "id": 952209,
          "postDate": "2020-07-30T17:53:30.817Z",
          "content": "<p>Thanks</p>",
          "rawMarkdown": "Thanks"
        }
      ]
    },
    {
      "id": 948266,
      "postDate": "2020-07-27T19:09:26.563Z",
      "content": "<p>Congrats!</p>",
      "rawMarkdown": "Congrats!",
      "votes": 1,
      "replies": [
        {
          "id": 948292,
          "postDate": "2020-07-27T19:30:31.743Z",
          "content": "<p>Thanks</p>",
          "rawMarkdown": "Thanks"
        }
      ]
    },
    {
      "id": 948010,
      "postDate": "2020-07-27T15:57:53.943Z",
      "content": "<p>Simply brilliant! Well done <a href=\"/iafoss\">@iafoss</a>. I have learned and keep on learning a lot from people like you. Keep it up!✊ </p>",
      "rawMarkdown": "Simply brilliant! Well done @iafoss. I have learned and keep on learning a lot from people like you. Keep it up!✊ ",
      "votes": 1,
      "replies": [
        {
          "id": 948068,
          "postDate": "2020-07-27T16:43:06.193Z",
          "content": "<p>You are welcome.</p>",
          "rawMarkdown": "You are welcome."
        }
      ]
    },
    {
      "id": 946825,
      "postDate": "2020-07-26T21:42:41.840Z",
      "content": "<p><a href=\"/iafoss\">@iafoss</a> \nCongratulation on winning sole gold, and once again for becoming kernel GM. 🔥 </p>",
      "rawMarkdown": "@iafoss \nCongratulation on winning sole gold, and once again for becoming kernel GM. 🔥 ",
      "votes": 1,
      "replies": [
        {
          "id": 946834,
          "postDate": "2020-07-26T21:56:43.907Z",
          "content": "<p>Thank you so much</p>",
          "rawMarkdown": "Thank you so much",
          "votes": 1
        }
      ]
    },
    {
      "id": 943175,
      "postDate": "2020-07-24T07:24:22.947Z",
      "content": "<p>Well done, big thanks for your starter kernel and your help during this competition! you deserve top 3 haha!</p>",
      "rawMarkdown": "Well done, big thanks for your starter kernel and your help during this competition! you deserve top 3 haha!",
      "votes": 1,
      "replies": [
        {
          "id": 943184,
          "postDate": "2020-07-24T07:30:50.977Z",
          "content": "<p>You are very welcome. Thanks, it happens, with such unstable LB and train/test mismatch even staying in gold range was not that simple. </p>",
          "rawMarkdown": "You are very welcome. Thanks, it happens, with such unstable LB and train/test mismatch even staying in gold range was not that simple. ",
          "votes": 1
        },
        {
          "id": 943194,
          "postDate": "2020-07-24T07:41:27.790Z",
          "content": "<p>Agree ! hope to see you in next CV competitions</p>",
          "rawMarkdown": "Agree ! hope to see you in next CV competitions"
        }
      ]
    },
    {
      "id": 942626,
      "postDate": "2020-07-23T22:11:21.367Z",
      "content": "<p>congrats!  I wish I had known about kappa loss.  How much did you get from it compared to binning loss when you compared?</p>",
      "rawMarkdown": "congrats!  I wish I had known about kappa loss.  How much did you get from it compared to binning loss when you compared?",
      "votes": 1,
      "replies": [
        {
          "id": 942648,
          "postDate": "2020-07-23T22:41:28.853Z",
          "content": "<p>For 4 fold CV on low res I have 0.872 for sorted binning loss vs 0.880 for kappa (pay attention that I exclude noisy images). For single fold intermediate res I got 0.950 (0.952 with individual threshold adjustment, but I'd expect it's just overfitting the val during adjustment) vs 0.950. In terms of LB performance I had only a single fold sub for intermediate res images: 0.92922/0.90120 vs. 0.93299/0.91343 (for the same conf but kappa loss). So the difference at private is not that huge, and the results are likely affected a lot by noise, but low public score has demotivated me enough from looking more into binning loss, when it was posted. My expectation is that those two losses are giving nearly the same result, while kappa loss is slightly better.</p>",
          "rawMarkdown": "For 4 fold CV on low res I have 0.872 for sorted binning loss vs 0.880 for kappa (pay attention that I exclude noisy images). For single fold intermediate res I got 0.950 (0.952 with individual threshold adjustment, but I'd expect it's just overfitting the val during adjustment) vs 0.950. In terms of LB performance I had only a single fold sub for intermediate res images: 0.92922/0.90120 vs. 0.93299/0.91343 (for the same conf but kappa loss). So the difference at private is not that huge, and the results are likely affected a lot by noise, but low public score has demotivated me enough from looking more into binning loss, when it was posted. My expectation is that those two losses are giving nearly the same result, while kappa loss is slightly better.",
          "votes": 1
        },
        {
          "id": 949606,
          "postDate": "2020-07-28T18:38:25.340Z",
          "content": "<p>Thanks.  I'll use it next time ;)</p>",
          "rawMarkdown": "Thanks.  I'll use it next time ;)"
        }
      ]
    },
    {
      "id": 942498,
      "postDate": "2020-07-23T19:23:08.833Z",
      "content": "<p>Congrats on making GM! You really did a great job this comp not just as a competitor but also sharing simple and effective ideas!</p>",
      "rawMarkdown": "Congrats on making GM! You really did a great job this comp not just as a competitor but also sharing simple and effective ideas!",
      "votes": 1,
      "replies": [
        {
          "id": 942549,
          "postDate": "2020-07-23T20:05:24.523Z",
          "content": "<p>Thank you. I'm really sorrow that you dropped so much at private LB. Getting competition GM is a little bit far away for me, but I made one of the required steps - solo gold (it was my main objective in this competition). </p>",
          "rawMarkdown": "Thank you. I'm really sorrow that you dropped so much at private LB. Getting competition GM is a little bit far away for me, but I made one of the required steps - solo gold (it was my main objective in this competition). "
        },
        {
          "id": 942563,
          "postDate": "2020-07-23T20:25:23.390Z",
          "content": "<p>Sorry I didn't know there was more requirements, but you definitely deserve it and i hope it comes soon for you! As for my team, it is quite sad we dropped 17 places, but we were just somewhat unlucky, so it can't be helped really.</p>",
          "rawMarkdown": "Sorry I didn't know there was more requirements, but you definitely deserve it and i hope it comes soon for you! As for my team, it is quite sad we dropped 17 places, but we were just somewhat unlucky, so it can't be helped really."
        }
      ]
    },
    {
      "id": 941985,
      "postDate": "2020-07-23T14:28:49.660Z",
      "content": "<p><a href=\"/iafoss\">@iafoss</a> Congratulations for your solo gold. Your notebooks are game changers in this competition.Thank you very much for sharing your knowledge. You are like a invisible team mate in every team.</p>",
      "rawMarkdown": "@iafoss Congratulations for your solo gold. Your notebooks are game changers in this competition.Thank you very much for sharing your knowledge. You are like a invisible team mate in every team.",
      "votes": 1,
      "replies": [
        {
          "id": 942543,
          "postDate": "2020-07-23T19:59:06.630Z",
          "content": "<p>You are welcome, I'm really glad that my idea about tiling worked so great in this competition.</p>",
          "rawMarkdown": "You are welcome, I'm really glad that my idea about tiling worked so great in this competition.",
          "votes": 1
        }
      ]
    },
    {
      "id": 941737,
      "postDate": "2020-07-23T11:36:38.493Z",
      "content": "<p><a href=\"/iafoss\">@iafoss</a> thank you for putting this level of hardwork into the competition. I really learned a lot from your kernels !</p>\n\n<p>If you don't mind, I have some questions for you:</p>\n\n<p>1) if I put y=2 and y_hat=2 in your formula above, I get a loss of one for a perfect prediction. How does that work exactly ?</p>\n\n<p>2) you mention not getting statistically significant improvement from TTA based on tile selection. How do you test for this ? is it gut feeling, or do you run the experiment multiple times and get statistics like mean and std of loss/kappa_metric (which seems rather impractical given the time it takes to train a model)</p>\n\n<p>3) How did you pick which images to remove from the training dataset ?</p>\n\n<p>4) \"I have explored a number of techniques to deal with the noise: progressive label distillation, JoCoR (Joint Training with Co-Regularization), Co-teaching, negative learning, excluding hard examples from the batch, etc. \"\nThat's the hell of a lot of stuff to try. What's your rationale for testing so many things ? Do you give yourself some time (\"I'll try this for three days and if no improvement I quit\"), or put more simply, when/how do you decide to quit a dead end ? </p>\n\n<p>And finally, one comment: if you resize lvl0 slide down by four, you get the exact same dimensions than lvl1. Yet, if you zoom on one part of the image and plot it twice (one for lvl0//4, one for lvl1), you'll see a neat difference in quality! That's because the images are compressed through JPEG inside the tiff format. So lvl0/4 &gt;&gt; lvl1 in terms of quality. Yet training the same pipe on lvl0/4 gets worse results than lvl1 slides. I believe the jpeg compression could act like a first layer of convolution, roughly speaking. Hence the not so better score using highest resolution ? </p>",
      "rawMarkdown": "@iafoss thank you for putting this level of hardwork into the competition. I really learned a lot from your kernels !\n\nIf you don't mind, I have some questions for you:\n\n1) if I put y=2 and y_hat=2 in your formula above, I get a loss of one for a perfect prediction. How does that work exactly ?\n\n2) you mention not getting statistically significant improvement from TTA based on tile selection. How do you test for this ? is it gut feeling, or do you run the experiment multiple times and get statistics like mean and std of loss/kappa_metric (which seems rather impractical given the time it takes to train a model)\n\n3) How did you pick which images to remove from the training dataset ?\n\n4) \"I have explored a number of techniques to deal with the noise: progressive label distillation, JoCoR (Joint Training with Co-Regularization), Co-teaching, negative learning, excluding hard examples from the batch, etc. \"\nThat's the hell of a lot of stuff to try. What's your rationale for testing so many things ? Do you give yourself some time (\"I'll try this for three days and if no improvement I quit\"), or put more simply, when/how do you decide to quit a dead end ? \n\n\nAnd finally, one comment: if you resize lvl0 slide down by four, you get the exact same dimensions than lvl1. Yet, if you zoom on one part of the image and plot it twice (one for lvl0//4, one for lvl1), you'll see a neat difference in quality! That's because the images are compressed through JPEG inside the tiff format. So lvl0/4 &gt;&gt; lvl1 in terms of quality. Yet training the same pipe on lvl0/4 gets worse results than lvl1 slides. I believe the jpeg compression could act like a first layer of convolution, roughly speaking. Hence the not so better score using highest resolution ? ",
      "votes": 1,
      "replies": [
        {
          "id": 942362,
          "postDate": "2020-07-23T17:57:33.560Z",
          "content": "<p>You are welcome.\n1) You are right, I have used 1 - k, I fixed the typo.\n2) When I checked tile selection TTA, I got slightly lower CV for several models (and lower LB as well). It was not very rigorous, but I wouldn't expect such TTA to be important in the competition (I just put more models in the ensemble to make it more stable). Meanwhile, I was using tile selection for training because it looks to be a natural augmentation for this data. \n3) I tried several things. One worked the best for me is based on the following. I ran progressive label distillation d1 and d2, and generated soft adjusted labels as <code>l_a = (4*l_true + l_d1 + l_d2)/6</code> and <code>l_a = (6*l_true + l_d1 + l_d2)/8</code> for Redbound and Karalinska data respectively (the weights are selected to have about ~1200 and ~200 different labels for those datasets). After that I drop images with <code>abs(l_true - l_a) &amp;gt; 0.5</code>. In addition, I dropped Redbound data with <code>abs(l_d2-l_true) &gt; 0.75</code> and images suggested <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/148060\">here</a>. I'm not sure right now about the exact numbers, but my training set was including 8700 images.\nOne I expected to work the best, but in reality it gave lower results for 4 4-fold models I submitted to LB (and private LB is lower on average as well) is the following. About a month before the end of the competition I had 10+ models trained with the noise removal procedure from above. I computed mean and std values for each prediction and dropped images with <code>abs(l_avr - l_true)/std &amp;gt; 10</code>. The idea is that <strong>if the difference is small, but the models are confident about their predictions, I must exclude the image</strong>. However, <strong>if the difference is large while models are not sure about the prediction, I'd expect it to be a hard example rather than an incorrect label</strong>. It excluded 990 and 236 Redbound and Karolinska data point with the same criterion applied to both datasets. Moreover, the CV for models trained with such exclusion was ~0.93/~0.94 for different providers (quite close, which I expected to be good), not like with the above method giving ~0.93/0.96. However, I couldn't get anything good at LB on average, not exactly sure why this method failed.\n4) I started with running experiments on low res layer tiles, so single fold training takes only ~30 min and 4 fold ~2 hours. Therefore, I could run a number of them quite quickly. When let's say, after several trials for a particular method I couldn't get CV above 0.80 (while CV for simple training is 0.84), I just quit the method. It probably took about a day to understand and implement a new method and get preliminary results. If I got CV comparable with the baseline method, I have a try for training on intermediate resolution tiles, but it's more time consuming and takes 1-2 days for 4 folds to train depending on the method.\n5) It's quite interesting observation. I've never tried to downsize large images by 4 times instead of using the intermediate layer. Also, I would rather expect that if more information is provided, the better model performance, unless there are some other things related to bs, loss behavior, etc.</p>",
          "rawMarkdown": "You are welcome.\n1) You are right, I have used 1 - k, I fixed the typo.\n2) When I checked tile selection TTA, I got slightly lower CV for several models (and lower LB as well). It was not very rigorous, but I wouldn't expect such TTA to be important in the competition (I just put more models in the ensemble to make it more stable). Meanwhile, I was using tile selection for training because it looks to be a natural augmentation for this data. \n3) I tried several things. One worked the best for me is based on the following. I ran progressive label distillation d1 and d2, and generated soft adjusted labels as `l_a = (4*l_true + l_d1 + l_d2)/6` and `l_a = (6*l_true + l_d1 + l_d2)/8` for Redbound and Karalinska data respectively (the weights are selected to have about ~1200 and ~200 different labels for those datasets). After that I drop images with `abs(l_true - l_a) &gt; 0.5`. In addition, I dropped Redbound data with `abs(l_d2-l_true) &gt; 0.75` and images suggested [here](https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/148060). I'm not sure right now about the exact numbers, but my training set was including 8700 images.\nOne I expected to work the best, but in reality it gave lower results for 4 4-fold models I submitted to LB (and private LB is lower on average as well) is the following. About a month before the end of the competition I had 10+ models trained with the noise removal procedure from above. I computed mean and std values for each prediction and dropped images with `abs(l_avr - l_true)/std &gt; 10`. The idea is that **if the difference is small, but the models are confident about their predictions, I must exclude the image**. However, **if the difference is large while models are not sure about the prediction, I'd expect it to be a hard example rather than an incorrect label**. It excluded 990 and 236 Redbound and Karolinska data point with the same criterion applied to both datasets. Moreover, the CV for models trained with such exclusion was ~0.93/~0.94 for different providers (quite close, which I expected to be good), not like with the above method giving ~0.93/0.96. However, I couldn't get anything good at LB on average, not exactly sure why this method failed.\n4) I started with running experiments on low res layer tiles, so single fold training takes only ~30 min and 4 fold ~2 hours. Therefore, I could run a number of them quite quickly. When let's say, after several trials for a particular method I couldn't get CV above 0.80 (while CV for simple training is 0.84), I just quit the method. It probably took about a day to understand and implement a new method and get preliminary results. If I got CV comparable with the baseline method, I have a try for training on intermediate resolution tiles, but it's more time consuming and takes 1-2 days for 4 folds to train depending on the method.\n5) It's quite interesting observation. I've never tried to downsize large images by 4 times instead of using the intermediate layer. Also, I would rather expect that if more information is provided, the better model performance, unless there are some other things related to bs, loss behavior, etc.",
          "votes": 2
        },
        {
          "id": 943657,
          "postDate": "2020-07-24T14:00:15.060Z",
          "content": "<p>Thanks for taking the time to answer so thoroughly all the questions. IMHO the meta game (how to make decisions, when to quit an experiment, the time-frame one allows himself for each experiment, etc..) is probably as important as the nitty-gritty code details, and I see I still have a lot to learn. Especially, I think I should have ran more experiments on lower res images to save time... I'll try to remember this !</p>\n\n<p>Congratulations again on this well-deserved solo Gold ! </p>",
          "rawMarkdown": "Thanks for taking the time to answer so thoroughly all the questions. IMHO the meta game (how to make decisions, when to quit an experiment, the time-frame one allows himself for each experiment, etc..) is probably as important as the nitty-gritty code details, and I see I still have a lot to learn. Especially, I think I should have ran more experiments on lower res images to save time... I'll try to remember this !\n\nCongratulations again on this well-deserved solo Gold ! "
        }
      ]
    },
    {
      "id": 941690,
      "postDate": "2020-07-23T10:59:51.277Z",
      "content": "<p>Congratulations! Well deserved!</p>",
      "rawMarkdown": "Congratulations! Well deserved!",
      "votes": 1,
      "replies": [
        {
          "id": 942222,
          "postDate": "2020-07-23T16:39:00.290Z",
          "content": "<p>Thank you.</p>",
          "rawMarkdown": "Thank you."
        }
      ]
    },
    {
      "id": 941535,
      "postDate": "2020-07-23T09:31:50.743Z",
      "content": "<p>Congratulations on your first solo competition gold medal and kernel grand-master rank!!  You contribute so much and explain in a such good way for all here.  Pretty sure everyone in this was helped by you.  Really felt for you when prize rank did not happen, was so sure it would.  Of course, always good luck for next time.     </p>",
      "rawMarkdown": "Congratulations on your first solo competition gold medal and kernel grand-master rank!!  You contribute so much and explain in a such good way for all here.  Pretty sure everyone in this was helped by you.  Really felt for you when prize rank did not happen, was so sure it would.  Of course, always good luck for next time.     ",
      "votes": 1,
      "replies": [
        {
          "id": 942221,
          "postDate": "2020-07-23T16:38:36.400Z",
          "content": "<p>Thank you so much, I really appreciate it.</p>",
          "rawMarkdown": "Thank you so much, I really appreciate it."
        }
      ]
    },
    {
      "id": 2555976,
      "postDate": "2023-12-10T10:58:00.013Z",
      "content": "<p>Great Job!</p>",
      "rawMarkdown": "Great Job!"
    },
    {
      "id": 1084059,
      "postDate": "2020-11-19T17:20:45.530Z",
      "content": "<p>Hi lafoss, thank a lot for sharing, I wanted to comeback and study little bit more your solution with fastai, I was wondering if you can share too and have a link for the submission notebook you used last?</p>\n<p>Thanks again.</p>",
      "rawMarkdown": "Hi lafoss, thank a lot for sharing, I wanted to comeback and study little bit more your solution with fastai, I was wondering if you can share too and have a link for the submission notebook you used last?\n\nThanks again.",
      "replies": [
        {
          "id": 1084147,
          "postDate": "2020-11-19T19:33:02.763Z",
          "content": "<p>You can check the <a href=\"https://github.com/iafoss/PANDA\" target=\"_blank\">my github repo</a> that provides the code I used for training my final models and links to different inference kernels.</p>",
          "rawMarkdown": "You can check the [my github repo](https://github.com/iafoss/PANDA) that provides the code I used for training my final models and links to different inference kernels.",
          "votes": 1
        },
        {
          "id": 1084248,
          "postDate": "2020-11-19T21:29:32.100Z",
          "content": "<p>Hi Amigo, thanks for the replay, I was looking into, but there is only PANDA_create and PANDA_pretrain and Train notebooks. but there is no inference notebooks.</p>",
          "rawMarkdown": "Hi Amigo, thanks for the replay, I was looking into, but there is only PANDA_create and PANDA_pretrain and Train notebooks. but there is no inference notebooks."
        },
        {
          "id": 1084284,
          "postDate": "2020-11-19T22:11:15.167Z",
          "content": "<p>Read the description provided to the repo, it includes the links.</p>",
          "rawMarkdown": "Read the description provided to the repo, it includes the links.",
          "votes": 1
        }
      ]
    },
    {
      "id": 957245,
      "postDate": "2020-08-04T07:12:11.347Z",
      "content": "<p><a href=\"/iafoss\">@iafoss</a>  We couldn't get some of our approaches to yield good performance with kappa loss. We based our approach on <a href=\"https://www.sciencedirect.com/science/article/abs/pii/S0167865517301666\">this </a> paper, which has been recently implemented in Tensorflow addons. </p>\n\n<p>Could you please share some more details of your training e.g. how many epochs it took to converge in the first run at low-res, the progression of qkappa metric at the intermediate resolution and how many epochs it took?</p>\n\n<p>We also observed that batch size of at least 5 was needed for convergence.</p>\n\n<p>But in our case, we couldn't get over 0.7 qkappa on our validation set with various LR/batch-sizes based on ballpark from above paper.\n(The same architecture, even with categorical cross-entropy loss, went upto 0.85 qkappa on validation set.)</p>\n\n<p>Any thoughts/suggestions on this would be helpful, as this is our first outing with kappa loss.</p>\n\n<p>Thanks, and congratulations!</p>",
      "rawMarkdown": "@iafoss  We couldn't get some of our approaches to yield good performance with kappa loss. We based our approach on [this ](https://www.sciencedirect.com/science/article/abs/pii/S0167865517301666) paper, which has been recently implemented in Tensorflow addons. \n\nCould you please share some more details of your training e.g. how many epochs it took to converge in the first run at low-res, the progression of qkappa metric at the intermediate resolution and how many epochs it took?\n\nWe also observed that batch size of at least 5 was needed for convergence.\n\nBut in our case, we couldn't get over 0.7 qkappa on our validation set with various LR/batch-sizes based on ballpark from above paper.\n(The same architecture, even with categorical cross-entropy loss, went upto 0.85 qkappa on validation set.)\n\nAny thoughts/suggestions on this would be helpful, as this is our first outing with kappa loss.\n\nThanks, and congratulations!",
      "replies": [
        {
          "id": 958012,
          "postDate": "2020-08-04T17:33:56.297Z",
          "content": "<p>It seems that u are referring to a <strong>different kappa loss</strong> that I haven't tried.  If u want to try mine, make sure that you are using centered yp and y in the loss based on mean value of labels (see my writeup) and <code>yp = 6*sigmoid(p) - 0.5</code>.\nI expect you are interested in a <strong>run without noisy labels removal</strong>, otherwise CV is different. I ran such training only <strong>with nearly basic setup ~3 month ago</strong>. In my run below the things are nearly identical to my public kernel, but bs = 64, and the loss is kappa + 0.1*aux (Gleason score), [fold 0 in my kernel]:\n<code>\n0   0.619686    0.726757    0.441646    01:10\n1   0.503786    0.420728    0.734977    01:05\n2   0.446397    0.388014    0.749429    01:04\n3   0.421686    0.371897    0.765389    01:05\n4   0.424101    0.440600    0.702009    01:06\n5   0.388839    0.349748    0.775638    01:06\n6   0.387164    0.497251    0.642067    01:06\n7   0.364439    0.363708    0.763292    01:06\n8   0.367793    0.471586    0.694158    01:06\n9   0.365068    0.383838    0.752568    01:06\n10  0.359331    0.413165    0.738522    01:06\n11  0.339751    0.343124    0.783672    01:06\n12  0.325875    0.324556    0.789526    01:06\n13  0.335336    0.349040    0.776613    01:06\n14  0.335342    0.606504    0.568188    01:06\n15  0.333862    0.304566    0.805746    01:06\n16  0.315612    0.302613    0.808751    01:06\n17  0.298632    0.285233    0.826465    01:07\n18  0.293765    0.318087    0.801963    01:06\n19  0.287576    0.295604    0.811252    01:06\n20  0.275854    0.303913    0.806941    01:06\n21  0.266260    0.291326    0.812348    01:07\n22  0.257481    0.347336    0.778186    01:07\n23  0.241969    0.262429    0.836433    01:07\n24  0.237443    0.264275    0.835410    01:07\n25  0.221909    0.263148    0.837879    01:07\n26  0.215041    0.251062    0.846085    01:07\n27  0.206626    0.257341    0.840549    01:07\n28  0.193016    0.248696    0.847726    01:08\n29  0.188344    0.258596    0.844116    01:07\n30  0.178238    0.251107    0.846781    01:08\n31  0.173396    0.246989    0.850495    01:08\n32  0.164201    0.245519    0.848950    01:08\n33  0.161503    0.242258    0.852881    01:08\n34  0.153965    0.240758    0.853131    01:08\n35  0.157219    0.242864    0.853106    01:08\n</code>\nNext I took the produced model and continued training it on 36x256x256 setup with bs = 6 and max lr of (1e-4,1e-3) [It's nearly my first run on intermediate res without any optimization]:\n<code>\nepoch   train_loss  valid_loss  d_kappa_score   kappa_k     kappa_r     time\n0   0.344915    0.283919    0.815302    0.791614    0.782375    09:47\n1   0.335626    0.285945    0.819567    0.763080    0.806630    09:46\n2   0.379465    0.268116    0.829591    0.842391    0.774673    09:48\n3   0.316176    0.233314    0.854984    0.860656    0.807655    09:50\n4   0.292933    0.361592    0.745009    0.664846    0.725402    09:49\n5   0.332429    0.244246    0.846598    0.859488    0.799022    09:54\n6   0.281213    0.257858    0.830744    0.812808    0.802393    09:52\n7   0.283063    0.226978    0.860622    0.841165    0.833361    09:52\n8   0.244150    0.212531    0.871141    0.873798    0.833969    09:55\n9   0.252861    0.231499    0.852559    0.857308    0.810779    09:52\n10  0.303027    0.253684    0.839478    0.855612    0.786137    09:53\n11  0.279593    0.214638    0.868501    0.869391    0.834018    09:55\n12  0.250773    0.203354    0.878194    0.886358    0.837234    09:56\n13  0.239266    0.203964    0.877210    0.887096    0.836048    09:58\n14  0.259900    0.193430    0.881883    0.889351    0.847430    09:58\n15  0.190317    0.215813    0.863405    0.861846    0.830002    10:01\n16  0.268994    0.207127    0.875795    0.881434    0.832446    10:03\n17  0.157652    0.210765    0.866115    0.870783    0.833130    10:05\n18  0.220389    0.195599    0.884750    0.882381    0.855319    10:07\n19  0.207806    0.190755    0.883505    0.888639    0.850455    10:08\n20  0.188866    0.204237    0.874564    0.883008    0.833871    10:10\n21  0.186228    0.202038    0.878626    0.887706    0.838257    10:10\n22  0.226118    0.236399    0.848630    0.874477    0.778278    10:10\n23  0.174489    0.184914    0.888356    0.896993    0.849679    10:09\n24  0.198199    0.202828    0.877724    0.870652    0.852208    10:09\n25  0.174334    0.176265    0.893822    0.906256    0.857446    10:10\n26  0.175599    0.179044    0.892619    0.901621    0.858356    10:02\n27  0.165271    0.177503    0.895425    0.904506    0.860751    09:58\n28  0.145186    0.181107    0.888282    0.895656    0.853604    09:55\n29  0.158865    0.179561    0.892518    0.895663    0.862738    09:57\n30  0.192933    0.182072    0.894250    0.905655    0.857467    09:56\n31  0.166352    0.171575    0.899262    0.900083    0.873222    10:00\n32  0.133393    0.185716    0.886050    0.896071    0.850738    09:59\n33  0.145026    0.190942    0.880710    0.891096    0.844466    10:00\n34  0.133870    0.170243    0.897480    0.907618    0.863825    10:02\n35  0.122654    0.169633    0.901467    0.911328    0.868964    10:04\n36  0.130831    0.178497    0.893576    0.904467    0.856378    10:07\n37  0.135886    0.172254    0.898089    0.908669    0.864243    09:53\n38  0.140824    0.167994    0.900024    0.912449    0.864562    09:55\n39  0.120919    0.168626    0.898917    0.911365    0.864052    09:56\n40  0.111123    0.166505    0.900856    0.908967    0.869388    09:58\n41  0.140773    0.167067    0.899353    0.908505    0.866439    10:03\n42  0.122937    0.168974    0.900271    0.911877    0.866887    10:03\n43  0.127722    0.170843    0.900197    0.909060    0.867776    10:06\n44  0.109837    0.168787    0.901611    0.908684    0.871180    09:51\n45  0.110626    0.165956    0.902472    0.912478    0.869869    16:13\n46  0.117968    0.167239    0.901686    0.910909    0.869625    09:55\n47  0.123976    0.166429    0.902506    0.913288    0.869211    09:56\n</code></p>",
          "rawMarkdown": "It seems that u are referring to a **different kappa loss** that I haven't tried.  If u want to try mine, make sure that you are using centered yp and y in the loss based on mean value of labels (see my writeup) and `yp = 6*sigmoid(p) - 0.5`.\nI expect you are interested in a **run without noisy labels removal**, otherwise CV is different. I ran such training only **with nearly basic setup ~3 month ago**. In my run below the things are nearly identical to my public kernel, but bs = 64, and the loss is kappa + 0.1*aux (Gleason score), [fold 0 in my kernel]:\n```\n0 \t0.619686 \t0.726757 \t0.441646 \t01:10\n1 \t0.503786 \t0.420728 \t0.734977 \t01:05\n2 \t0.446397 \t0.388014 \t0.749429 \t01:04\n3 \t0.421686 \t0.371897 \t0.765389 \t01:05\n4 \t0.424101 \t0.440600 \t0.702009 \t01:06\n5 \t0.388839 \t0.349748 \t0.775638 \t01:06\n6 \t0.387164 \t0.497251 \t0.642067 \t01:06\n7 \t0.364439 \t0.363708 \t0.763292 \t01:06\n8 \t0.367793 \t0.471586 \t0.694158 \t01:06\n9 \t0.365068 \t0.383838 \t0.752568 \t01:06\n10 \t0.359331 \t0.413165 \t0.738522 \t01:06\n11 \t0.339751 \t0.343124 \t0.783672 \t01:06\n12 \t0.325875 \t0.324556 \t0.789526 \t01:06\n13 \t0.335336 \t0.349040 \t0.776613 \t01:06\n14 \t0.335342 \t0.606504 \t0.568188 \t01:06\n15 \t0.333862 \t0.304566 \t0.805746 \t01:06\n16 \t0.315612 \t0.302613 \t0.808751 \t01:06\n17 \t0.298632 \t0.285233 \t0.826465 \t01:07\n18 \t0.293765 \t0.318087 \t0.801963 \t01:06\n19 \t0.287576 \t0.295604 \t0.811252 \t01:06\n20 \t0.275854 \t0.303913 \t0.806941 \t01:06\n21 \t0.266260 \t0.291326 \t0.812348 \t01:07\n22 \t0.257481 \t0.347336 \t0.778186 \t01:07\n23 \t0.241969 \t0.262429 \t0.836433 \t01:07\n24 \t0.237443 \t0.264275 \t0.835410 \t01:07\n25 \t0.221909 \t0.263148 \t0.837879 \t01:07\n26 \t0.215041 \t0.251062 \t0.846085 \t01:07\n27 \t0.206626 \t0.257341 \t0.840549 \t01:07\n28 \t0.193016 \t0.248696 \t0.847726 \t01:08\n29 \t0.188344 \t0.258596 \t0.844116 \t01:07\n30 \t0.178238 \t0.251107 \t0.846781 \t01:08\n31 \t0.173396 \t0.246989 \t0.850495 \t01:08\n32 \t0.164201 \t0.245519 \t0.848950 \t01:08\n33 \t0.161503 \t0.242258 \t0.852881 \t01:08\n34 \t0.153965 \t0.240758 \t0.853131 \t01:08\n35 \t0.157219 \t0.242864 \t0.853106 \t01:08\n```\nNext I took the produced model and continued training it on 36x256x256 setup with bs = 6 and max lr of (1e-4,1e-3) [It's nearly my first run on intermediate res without any optimization]:\n```\nepoch \ttrain_loss \tvalid_loss \td_kappa_score \tkappa_k \tkappa_r \ttime\n0 \t0.344915 \t0.283919 \t0.815302 \t0.791614 \t0.782375 \t09:47\n1 \t0.335626 \t0.285945 \t0.819567 \t0.763080 \t0.806630 \t09:46\n2 \t0.379465 \t0.268116 \t0.829591 \t0.842391 \t0.774673 \t09:48\n3 \t0.316176 \t0.233314 \t0.854984 \t0.860656 \t0.807655 \t09:50\n4 \t0.292933 \t0.361592 \t0.745009 \t0.664846 \t0.725402 \t09:49\n5 \t0.332429 \t0.244246 \t0.846598 \t0.859488 \t0.799022 \t09:54\n6 \t0.281213 \t0.257858 \t0.830744 \t0.812808 \t0.802393 \t09:52\n7 \t0.283063 \t0.226978 \t0.860622 \t0.841165 \t0.833361 \t09:52\n8 \t0.244150 \t0.212531 \t0.871141 \t0.873798 \t0.833969 \t09:55\n9 \t0.252861 \t0.231499 \t0.852559 \t0.857308 \t0.810779 \t09:52\n10 \t0.303027 \t0.253684 \t0.839478 \t0.855612 \t0.786137 \t09:53\n11 \t0.279593 \t0.214638 \t0.868501 \t0.869391 \t0.834018 \t09:55\n12 \t0.250773 \t0.203354 \t0.878194 \t0.886358 \t0.837234 \t09:56\n13 \t0.239266 \t0.203964 \t0.877210 \t0.887096 \t0.836048 \t09:58\n14 \t0.259900 \t0.193430 \t0.881883 \t0.889351 \t0.847430 \t09:58\n15 \t0.190317 \t0.215813 \t0.863405 \t0.861846 \t0.830002 \t10:01\n16 \t0.268994 \t0.207127 \t0.875795 \t0.881434 \t0.832446 \t10:03\n17 \t0.157652 \t0.210765 \t0.866115 \t0.870783 \t0.833130 \t10:05\n18 \t0.220389 \t0.195599 \t0.884750 \t0.882381 \t0.855319 \t10:07\n19 \t0.207806 \t0.190755 \t0.883505 \t0.888639 \t0.850455 \t10:08\n20 \t0.188866 \t0.204237 \t0.874564 \t0.883008 \t0.833871 \t10:10\n21 \t0.186228 \t0.202038 \t0.878626 \t0.887706 \t0.838257 \t10:10\n22 \t0.226118 \t0.236399 \t0.848630 \t0.874477 \t0.778278 \t10:10\n23 \t0.174489 \t0.184914 \t0.888356 \t0.896993 \t0.849679 \t10:09\n24 \t0.198199 \t0.202828 \t0.877724 \t0.870652 \t0.852208 \t10:09\n25 \t0.174334 \t0.176265 \t0.893822 \t0.906256 \t0.857446 \t10:10\n26 \t0.175599 \t0.179044 \t0.892619 \t0.901621 \t0.858356 \t10:02\n27 \t0.165271 \t0.177503 \t0.895425 \t0.904506 \t0.860751 \t09:58\n28 \t0.145186 \t0.181107 \t0.888282 \t0.895656 \t0.853604 \t09:55\n29 \t0.158865 \t0.179561 \t0.892518 \t0.895663 \t0.862738 \t09:57\n30 \t0.192933 \t0.182072 \t0.894250 \t0.905655 \t0.857467 \t09:56\n31 \t0.166352 \t0.171575 \t0.899262 \t0.900083 \t0.873222 \t10:00\n32 \t0.133393 \t0.185716 \t0.886050 \t0.896071 \t0.850738 \t09:59\n33 \t0.145026 \t0.190942 \t0.880710 \t0.891096 \t0.844466 \t10:00\n34 \t0.133870 \t0.170243 \t0.897480 \t0.907618 \t0.863825 \t10:02\n35 \t0.122654 \t0.169633 \t0.901467 \t0.911328 \t0.868964 \t10:04\n36 \t0.130831 \t0.178497 \t0.893576 \t0.904467 \t0.856378 \t10:07\n37 \t0.135886 \t0.172254 \t0.898089 \t0.908669 \t0.864243 \t09:53\n38 \t0.140824 \t0.167994 \t0.900024 \t0.912449 \t0.864562 \t09:55\n39 \t0.120919 \t0.168626 \t0.898917 \t0.911365 \t0.864052 \t09:56\n40 \t0.111123 \t0.166505 \t0.900856 \t0.908967 \t0.869388 \t09:58\n41 \t0.140773 \t0.167067 \t0.899353 \t0.908505 \t0.866439 \t10:03\n42 \t0.122937 \t0.168974 \t0.900271 \t0.911877 \t0.866887 \t10:03\n43 \t0.127722 \t0.170843 \t0.900197 \t0.909060 \t0.867776 \t10:06\n44 \t0.109837 \t0.168787 \t0.901611 \t0.908684 \t0.871180 \t09:51\n45 \t0.110626 \t0.165956 \t0.902472 \t0.912478 \t0.869869 \t16:13\n46 \t0.117968 \t0.167239 \t0.901686 \t0.910909 \t0.869625 \t09:55\n47 \t0.123976 \t0.166429 \t0.902506 \t0.913288 \t0.869211 \t09:56\n```"
        },
        {
          "id": 966285,
          "postDate": "2020-08-11T10:01:05.060Z",
          "content": "<p>Thank you!<br>\n(Yes, I'm referring to a different kappa loss, and this is for initial run without noisy labels removal.)</p>",
          "rawMarkdown": "Thank you!\n(Yes, I'm referring to a different kappa loss, and this is for initial run without noisy labels removal.)"
        }
      ]
    },
    {
      "id": 944025,
      "postDate": "2020-07-24T18:47:43.077Z",
      "content": "<p>First off, congratulations! Well deserved.</p>\n\n<p>I tried to implement kappa loss, like so:</p>\n\n<p>class KappaLoss(nn.Module):</p>\n\n<pre><code>def __init__(self, labels_mean):\n    super().__init__()\n    self.labels_mean = labels_mean\n\ndef forward(self, y_hat, y):\n    y_hat.requires_grad = True\n    y_hat = y_hat - self.labels_mean\n    y = y - self.labels_mean\n    loss = (1 - (2 * torch.dot(y_hat, y) /\n            (torch.dot(y, y) + torch.dot(y_hat, y_hat))))\n\n    return loss\n</code></pre>\n\n<p>But when I run a training for one epoch and plot the loss vs learning rate, I get this:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2275947%2F69c1820e38c980307c6cce0e4f635906%2Flr_finder_16_batch.jpg?generation=1595616365041574&amp;alt=media\" alt=\"\"></p>\n\n<p>So I feel like something is wrong. Would love your input!</p>",
      "rawMarkdown": "First off, congratulations! Well deserved.\n\nI tried to implement kappa loss, like so:\n\nclass KappaLoss(nn.Module):\n\n    def __init__(self, labels_mean):\n        super().__init__()\n        self.labels_mean = labels_mean\n\n    def forward(self, y_hat, y):\n        y_hat.requires_grad = True\n        y_hat = y_hat - self.labels_mean\n        y = y - self.labels_mean\n        loss = (1 - (2 * torch.dot(y_hat, y) /\n                (torch.dot(y, y) + torch.dot(y_hat, y_hat))))\n\n        return loss\n\nBut when I run a training for one epoch and plot the loss vs learning rate, I get this:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2275947%2F69c1820e38c980307c6cce0e4f635906%2Flr_finder_16_batch.jpg?generation=1595616365041574&amp;alt=media)\n\nSo I feel like something is wrong. Would love your input!",
      "replies": [
        {
          "id": 944058,
          "postDate": "2020-07-24T19:28:21.070Z",
          "content": "<p>I have used the following one:</p>\n\n<p><code>\ny_shift = df.isup_grade.mean()\nNg = 6\ndef Kloss(x, target):\n    x = Ng*torch.sigmoid(x.float()).view(-1) - 0.5\n    target = target.float()\n    return 1.0 - (2.0*((x-y_shift)*(target-y_shift)).sum() - 1e-3)/\\\n        (((x-y_shift)**2).sum() + ((target-y_shift)**2).sum() + 1e-3)\n</code></p>\n\n<p>Though, without sigmoid it should also work. What is your bs?</p>",
          "rawMarkdown": "I have used the following one:\n\n```\ny_shift = df.isup_grade.mean()\nNg = 6\ndef Kloss(x, target):\n    x = Ng*torch.sigmoid(x.float()).view(-1) - 0.5\n    target = target.float()\n    return 1.0 - (2.0*((x-y_shift)*(target-y_shift)).sum() - 1e-3)/\\\n        (((x-y_shift)**2).sum() + ((target-y_shift)**2).sum() + 1e-3)\n```\n\nThough, without sigmoid it should also work. What is your bs?",
          "votes": 2
        },
        {
          "id": 944102,
          "postDate": "2020-07-24T20:39:47.153Z",
          "content": "<p>I used a batch size of 16.</p>\n\n<p>What is x in your case?</p>",
          "rawMarkdown": "I used a batch size of 16.\n\nWhat is x in your case?"
        },
        {
          "id": 944127,
          "postDate": "2020-07-24T21:10:35.740Z",
          "content": "<p>x is the output from the model. Try bs 64, in my procedure I first did pretraining on low res and then used intermediate res with bs = 8. Though, even with bs = 16 there should be convergence, not like u are showing. Did u try just to do regular training with lr ~1e-3? Also if u use sigmoid, the predicted labels in the metric evaluation should be computed as lp = (Ng*torch.sigmoid(x.float()).view(-1)).long().</p>",
          "rawMarkdown": "x is the output from the model. Try bs 64, in my procedure I first did pretraining on low res and then used intermediate res with bs = 8. Though, even with bs = 16 there should be convergence, not like u are showing. Did u try just to do regular training with lr ~1e-3? Also if u use sigmoid, the predicted labels in the metric evaluation should be computed as lp = (Ng*torch.sigmoid(x.float()).view(-1)).long()."
        },
        {
          "id": 944157,
          "postDate": "2020-07-24T22:00:32.883Z",
          "content": "<p>Sorry I meant is x the output of a regression model or classification model? Are the targets just the isup_grades?</p>\n\n<p>What are the shapes of both?</p>",
          "rawMarkdown": "Sorry I meant is x the output of a regression model or classification model? Are the targets just the isup_grades?\n\nWhat are the shapes of both?"
        },
        {
          "id": 944174,
          "postDate": "2020-07-24T22:26:42.047Z",
          "content": "<p>The model outputs a single value per image if u want to use this loss (regression), so u get the same number of elements in your output and target. For the code above, the target is just isup_grades.</p>",
          "rawMarkdown": "The model outputs a single value per image if u want to use this loss (regression), so u get the same number of elements in your output and target. For the code above, the target is just isup_grades.",
          "votes": 1
        },
        {
          "id": 944176,
          "postDate": "2020-07-24T22:30:10.643Z",
          "content": "<p>Okay, that was my error, I was getting y_hat by argmaxing a softmax output. I'll retry. Thank you!</p>",
          "rawMarkdown": "Okay, that was my error, I was getting y_hat by argmaxing a softmax output. I'll retry. Thank you!"
        }
      ]
    },
    {
      "id": 1036751,
      "postDate": "2020-10-04T08:46:58.587Z",
      "content": "<p>Congrats! Thank you for sharing.</p>",
      "rawMarkdown": "Congrats! Thank you for sharing.",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 1037230,
          "postDate": "2020-10-04T17:51:31.370Z",
          "content": "<p>You are welcome</p>",
          "rawMarkdown": "You are welcome"
        }
      ]
    },
    {
      "id": 943666,
      "postDate": "2020-07-24T14:07:28.847Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 943838,
          "postDate": "2020-07-24T16:04:10.490Z",
          "content": "<p>Thanks</p>",
          "rawMarkdown": "Thanks"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 941428,
      "author_name": "fam_taro",
      "author_url": "",
      "post_date": "2020-07-23T08:05:10.673000",
      "content": "<p>Congrats solo gold !!!\nI think you are the player who has contributed the most to this competition. :)\nYour tile method helps almost player.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 941719,
          "author_name": "Benjamin Dubreu",
          "author_url": "",
          "post_date": "2020-07-23T11:24:50.860000",
          "content": "<p>Agreed. This is a well deserved solo gold !</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 942375,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-07-23T18:03:10.987000",
          "content": "<p>Thank you so much, I'm happy about it.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1217213,
      "author_name": "Swikwislkdjc",
      "author_url": "",
      "post_date": "2021-02-24T21:07:27.383000",
      "content": "<p>Congratulation! (after 7 months, I know I'm late here lol) I really love your approach!! Thank you so much for sharing those grate notebooks with explanations!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1217387,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2021-02-25T02:20:41.553000",
          "content": "<p>You are very welcome</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 964939,
      "author_name": "__n1kshaN__",
      "author_url": "",
      "post_date": "2020-08-10T09:27:03.447000",
      "content": "<p>Amazing  content. Keep sharing </p>",
      "votes": 1,
      "replies": [
        {
          "id": 965415,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-08-10T15:56:16.437000",
          "content": "<p>Thank you.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 953091,
      "author_name": "PatrickChoDev",
      "author_url": "",
      "post_date": "2020-07-31T14:25:33.923000",
      "content": "<p>Congratulations <a href=\"/iafoss\">@iafoss</a> for success on this kernal. Keep sharing nice things like this. 💕</p>",
      "votes": 1,
      "replies": [
        {
          "id": 953161,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-07-31T15:18:59.843000",
          "content": "<p>Thank you</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 951890,
      "author_name": "Manish Kumar",
      "author_url": "",
      "post_date": "2020-07-30T13:38:25.480000",
      "content": "<p>Congratulations <a href=\"/iafoss\">@iafoss</a> for this achievement. And thank for sharing the details.👍 </p>",
      "votes": 1,
      "replies": [
        {
          "id": 952204,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-07-30T17:52:01.397000",
          "content": "<p>You re very welcome</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 951849,
      "author_name": "Ashish Pokhriyal",
      "author_url": "",
      "post_date": "2020-07-30T12:58:04.433000",
      "content": "<p>Congratulations :-)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 952205,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-07-30T17:52:16.567000",
          "content": "<p>Thanks</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 951694,
      "author_name": "Parth Agrawal",
      "author_url": "",
      "post_date": "2020-07-30T10:14:17.373000",
      "content": "<p>Congratulations! Learned a lot from your kernels!\nThank you!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 952211,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-07-30T17:54:16.347000",
          "content": "<p>Thank you so much, you are very welcome.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 949602,
      "author_name": "Bruno Bento",
      "author_url": "",
      "post_date": "2020-07-28T18:36:46.423000",
      "content": "<p>Congrats!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 952207,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-07-30T17:52:36.353000",
          "content": "<p>Thank you</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 949583,
      "author_name": "raavi soni",
      "author_url": "",
      "post_date": "2020-07-28T18:17:38.900000",
      "content": "<p>Congratulations !!!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 952209,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-07-30T17:53:30.817000",
          "content": "<p>Thanks</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 948266,
      "author_name": "Sanjay K",
      "author_url": "",
      "post_date": "2020-07-27T19:09:26.563000",
      "content": "<p>Congrats!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 948292,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-07-27T19:30:31.743000",
          "content": "<p>Thanks</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 948010,
      "author_name": "Laique M. Djeutchouang",
      "author_url": "",
      "post_date": "2020-07-27T15:57:53.943000",
      "content": "<p>Simply brilliant! Well done <a href=\"/iafoss\">@iafoss</a>. I have learned and keep on learning a lot from people like you. Keep it up!✊ </p>",
      "votes": 1,
      "replies": [
        {
          "id": 948068,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-07-27T16:43:06.193000",
          "content": "<p>You are welcome.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 946825,
      "author_name": "Innat",
      "author_url": "",
      "post_date": "2020-07-26T21:42:41.840000",
      "content": "<p><a href=\"/iafoss\">@iafoss</a> \nCongratulation on winning sole gold, and once again for becoming kernel GM. 🔥 </p>",
      "votes": 1,
      "replies": [
        {
          "id": 946834,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-07-26T21:56:43.907000",
          "content": "<p>Thank you so much</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 943175,
      "author_name": "Alex",
      "author_url": "",
      "post_date": "2020-07-24T07:24:22.947000",
      "content": "<p>Well done, big thanks for your starter kernel and your help during this competition! you deserve top 3 haha!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 943184,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-07-24T07:30:50.977000",
          "content": "<p>You are very welcome. Thanks, it happens, with such unstable LB and train/test mismatch even staying in gold range was not that simple. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 943194,
          "author_name": "Alex",
          "author_url": "",
          "post_date": "2020-07-24T07:41:27.790000",
          "content": "<p>Agree ! hope to see you in next CV competitions</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 942626,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2020-07-23T22:11:21.367000",
      "content": "<p>congrats!  I wish I had known about kappa loss.  How much did you get from it compared to binning loss when you compared?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 942648,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-07-23T22:41:28.853000",
          "content": "<p>For 4 fold CV on low res I have 0.872 for sorted binning loss vs 0.880 for kappa (pay attention that I exclude noisy images). For single fold intermediate res I got 0.950 (0.952 with individual threshold adjustment, but I'd expect it's just overfitting the val during adjustment) vs 0.950. In terms of LB performance I had only a single fold sub for intermediate res images: 0.92922/0.90120 vs. 0.93299/0.91343 (for the same conf but kappa loss). So the difference at private is not that huge, and the results are likely affected a lot by noise, but low public score has demotivated me enough from looking more into binning loss, when it was posted. My expectation is that those two losses are giving nearly the same result, while kappa loss is slightly better.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 949606,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-07-28T18:38:25.340000",
          "content": "<p>Thanks.  I'll use it next time ;)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 942498,
      "author_name": "Shujun",
      "author_url": "",
      "post_date": "2020-07-23T19:23:08.833000",
      "content": "<p>Congrats on making GM! You really did a great job this comp not just as a competitor but also sharing simple and effective ideas!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 942549,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-07-23T20:05:24.523000",
          "content": "<p>Thank you. I'm really sorrow that you dropped so much at private LB. Getting competition GM is a little bit far away for me, but I made one of the required steps - solo gold (it was my main objective in this competition). </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 942563,
          "author_name": "Shujun",
          "author_url": "",
          "post_date": "2020-07-23T20:25:23.390000",
          "content": "<p>Sorry I didn't know there was more requirements, but you definitely deserve it and i hope it comes soon for you! As for my team, it is quite sad we dropped 17 places, but we were just somewhat unlucky, so it can't be helped really.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 941985,
      "author_name": "ratan rohith",
      "author_url": "",
      "post_date": "2020-07-23T14:28:49.660000",
      "content": "<p><a href=\"/iafoss\">@iafoss</a> Congratulations for your solo gold. Your notebooks are game changers in this competition.Thank you very much for sharing your knowledge. You are like a invisible team mate in every team.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 942543,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-07-23T19:59:06.630000",
          "content": "<p>You are welcome, I'm really glad that my idea about tiling worked so great in this competition.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 941737,
      "author_name": "Benjamin Dubreu",
      "author_url": "",
      "post_date": "2020-07-23T11:36:38.493000",
      "content": "<p><a href=\"/iafoss\">@iafoss</a> thank you for putting this level of hardwork into the competition. I really learned a lot from your kernels !</p>\n\n<p>If you don't mind, I have some questions for you:</p>\n\n<p>1) if I put y=2 and y_hat=2 in your formula above, I get a loss of one for a perfect prediction. How does that work exactly ?</p>\n\n<p>2) you mention not getting statistically significant improvement from TTA based on tile selection. How do you test for this ? is it gut feeling, or do you run the experiment multiple times and get statistics like mean and std of loss/kappa_metric (which seems rather impractical given the time it takes to train a model)</p>\n\n<p>3) How did you pick which images to remove from the training dataset ?</p>\n\n<p>4) \"I have explored a number of techniques to deal with the noise: progressive label distillation, JoCoR (Joint Training with Co-Regularization), Co-teaching, negative learning, excluding hard examples from the batch, etc. \"\nThat's the hell of a lot of stuff to try. What's your rationale for testing so many things ? Do you give yourself some time (\"I'll try this for three days and if no improvement I quit\"), or put more simply, when/how do you decide to quit a dead end ? </p>\n\n<p>And finally, one comment: if you resize lvl0 slide down by four, you get the exact same dimensions than lvl1. Yet, if you zoom on one part of the image and plot it twice (one for lvl0//4, one for lvl1), you'll see a neat difference in quality! That's because the images are compressed through JPEG inside the tiff format. So lvl0/4 &gt;&gt; lvl1 in terms of quality. Yet training the same pipe on lvl0/4 gets worse results than lvl1 slides. I believe the jpeg compression could act like a first layer of convolution, roughly speaking. Hence the not so better score using highest resolution ? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 942362,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-07-23T17:57:33.560000",
          "content": "<p>You are welcome.\n1) You are right, I have used 1 - k, I fixed the typo.\n2) When I checked tile selection TTA, I got slightly lower CV for several models (and lower LB as well). It was not very rigorous, but I wouldn't expect such TTA to be important in the competition (I just put more models in the ensemble to make it more stable). Meanwhile, I was using tile selection for training because it looks to be a natural augmentation for this data. \n3) I tried several things. One worked the best for me is based on the following. I ran progressive label distillation d1 and d2, and generated soft adjusted labels as <code>l_a = (4*l_true + l_d1 + l_d2)/6</code> and <code>l_a = (6*l_true + l_d1 + l_d2)/8</code> for Redbound and Karalinska data respectively (the weights are selected to have about ~1200 and ~200 different labels for those datasets). After that I drop images with <code>abs(l_true - l_a) &amp;gt; 0.5</code>. In addition, I dropped Redbound data with <code>abs(l_d2-l_true) &gt; 0.75</code> and images suggested <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/148060\">here</a>. I'm not sure right now about the exact numbers, but my training set was including 8700 images.\nOne I expected to work the best, but in reality it gave lower results for 4 4-fold models I submitted to LB (and private LB is lower on average as well) is the following. About a month before the end of the competition I had 10+ models trained with the noise removal procedure from above. I computed mean and std values for each prediction and dropped images with <code>abs(l_avr - l_true)/std &amp;gt; 10</code>. The idea is that <strong>if the difference is small, but the models are confident about their predictions, I must exclude the image</strong>. However, <strong>if the difference is large while models are not sure about the prediction, I'd expect it to be a hard example rather than an incorrect label</strong>. It excluded 990 and 236 Redbound and Karolinska data point with the same criterion applied to both datasets. Moreover, the CV for models trained with such exclusion was ~0.93/~0.94 for different providers (quite close, which I expected to be good), not like with the above method giving ~0.93/0.96. However, I couldn't get anything good at LB on average, not exactly sure why this method failed.\n4) I started with running experiments on low res layer tiles, so single fold training takes only ~30 min and 4 fold ~2 hours. Therefore, I could run a number of them quite quickly. When let's say, after several trials for a particular method I couldn't get CV above 0.80 (while CV for simple training is 0.84), I just quit the method. It probably took about a day to understand and implement a new method and get preliminary results. If I got CV comparable with the baseline method, I have a try for training on intermediate resolution tiles, but it's more time consuming and takes 1-2 days for 4 folds to train depending on the method.\n5) It's quite interesting observation. I've never tried to downsize large images by 4 times instead of using the intermediate layer. Also, I would rather expect that if more information is provided, the better model performance, unless there are some other things related to bs, loss behavior, etc.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 943657,
          "author_name": "Benjamin Dubreu",
          "author_url": "",
          "post_date": "2020-07-24T14:00:15.060000",
          "content": "<p>Thanks for taking the time to answer so thoroughly all the questions. IMHO the meta game (how to make decisions, when to quit an experiment, the time-frame one allows himself for each experiment, etc..) is probably as important as the nitty-gritty code details, and I see I still have a lot to learn. Especially, I think I should have ran more experiments on lower res images to save time... I'll try to remember this !</p>\n\n<p>Congratulations again on this well-deserved solo Gold ! </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 941690,
      "author_name": "bluetrain",
      "author_url": "",
      "post_date": "2020-07-23T10:59:51.277000",
      "content": "<p>Congratulations! Well deserved!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 942222,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-07-23T16:39:00.290000",
          "content": "<p>Thank you.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 941535,
      "author_name": "something4kag",
      "author_url": "",
      "post_date": "2020-07-23T09:31:50.743000",
      "content": "<p>Congratulations on your first solo competition gold medal and kernel grand-master rank!!  You contribute so much and explain in a such good way for all here.  Pretty sure everyone in this was helped by you.  Really felt for you when prize rank did not happen, was so sure it would.  Of course, always good luck for next time.     </p>",
      "votes": 1,
      "replies": [
        {
          "id": 942221,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-07-23T16:38:36.400000",
          "content": "<p>Thank you so much, I really appreciate it.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2555976,
      "author_name": "Mahmoud Essam",
      "author_url": "",
      "post_date": "2023-12-10T10:58:00.013000",
      "content": "<p>Great Job!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1084059,
      "author_name": "TheStoneMX",
      "author_url": "",
      "post_date": "2020-11-19T17:20:45.530000",
      "content": "<p>Hi lafoss, thank a lot for sharing, I wanted to comeback and study little bit more your solution with fastai, I was wondering if you can share too and have a link for the submission notebook you used last?</p>\n<p>Thanks again.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1084147,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-11-19T19:33:02.763000",
          "content": "<p>You can check the <a href=\"https://github.com/iafoss/PANDA\" target=\"_blank\">my github repo</a> that provides the code I used for training my final models and links to different inference kernels.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1084248,
          "author_name": "TheStoneMX",
          "author_url": "",
          "post_date": "2020-11-19T21:29:32.100000",
          "content": "<p>Hi Amigo, thanks for the replay, I was looking into, but there is only PANDA_create and PANDA_pretrain and Train notebooks. but there is no inference notebooks.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1084284,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-11-19T22:11:15.167000",
          "content": "<p>Read the description provided to the repo, it includes the links.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 957245,
      "author_name": "Srinath K",
      "author_url": "",
      "post_date": "2020-08-04T07:12:11.347000",
      "content": "<p><a href=\"/iafoss\">@iafoss</a>  We couldn't get some of our approaches to yield good performance with kappa loss. We based our approach on <a href=\"https://www.sciencedirect.com/science/article/abs/pii/S0167865517301666\">this </a> paper, which has been recently implemented in Tensorflow addons. </p>\n\n<p>Could you please share some more details of your training e.g. how many epochs it took to converge in the first run at low-res, the progression of qkappa metric at the intermediate resolution and how many epochs it took?</p>\n\n<p>We also observed that batch size of at least 5 was needed for convergence.</p>\n\n<p>But in our case, we couldn't get over 0.7 qkappa on our validation set with various LR/batch-sizes based on ballpark from above paper.\n(The same architecture, even with categorical cross-entropy loss, went upto 0.85 qkappa on validation set.)</p>\n\n<p>Any thoughts/suggestions on this would be helpful, as this is our first outing with kappa loss.</p>\n\n<p>Thanks, and congratulations!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 958012,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-08-04T17:33:56.297000",
          "content": "<p>It seems that u are referring to a <strong>different kappa loss</strong> that I haven't tried.  If u want to try mine, make sure that you are using centered yp and y in the loss based on mean value of labels (see my writeup) and <code>yp = 6*sigmoid(p) - 0.5</code>.\nI expect you are interested in a <strong>run without noisy labels removal</strong>, otherwise CV is different. I ran such training only <strong>with nearly basic setup ~3 month ago</strong>. In my run below the things are nearly identical to my public kernel, but bs = 64, and the loss is kappa + 0.1*aux (Gleason score), [fold 0 in my kernel]:\n<code>\n0   0.619686    0.726757    0.441646    01:10\n1   0.503786    0.420728    0.734977    01:05\n2   0.446397    0.388014    0.749429    01:04\n3   0.421686    0.371897    0.765389    01:05\n4   0.424101    0.440600    0.702009    01:06\n5   0.388839    0.349748    0.775638    01:06\n6   0.387164    0.497251    0.642067    01:06\n7   0.364439    0.363708    0.763292    01:06\n8   0.367793    0.471586    0.694158    01:06\n9   0.365068    0.383838    0.752568    01:06\n10  0.359331    0.413165    0.738522    01:06\n11  0.339751    0.343124    0.783672    01:06\n12  0.325875    0.324556    0.789526    01:06\n13  0.335336    0.349040    0.776613    01:06\n14  0.335342    0.606504    0.568188    01:06\n15  0.333862    0.304566    0.805746    01:06\n16  0.315612    0.302613    0.808751    01:06\n17  0.298632    0.285233    0.826465    01:07\n18  0.293765    0.318087    0.801963    01:06\n19  0.287576    0.295604    0.811252    01:06\n20  0.275854    0.303913    0.806941    01:06\n21  0.266260    0.291326    0.812348    01:07\n22  0.257481    0.347336    0.778186    01:07\n23  0.241969    0.262429    0.836433    01:07\n24  0.237443    0.264275    0.835410    01:07\n25  0.221909    0.263148    0.837879    01:07\n26  0.215041    0.251062    0.846085    01:07\n27  0.206626    0.257341    0.840549    01:07\n28  0.193016    0.248696    0.847726    01:08\n29  0.188344    0.258596    0.844116    01:07\n30  0.178238    0.251107    0.846781    01:08\n31  0.173396    0.246989    0.850495    01:08\n32  0.164201    0.245519    0.848950    01:08\n33  0.161503    0.242258    0.852881    01:08\n34  0.153965    0.240758    0.853131    01:08\n35  0.157219    0.242864    0.853106    01:08\n</code>\nNext I took the produced model and continued training it on 36x256x256 setup with bs = 6 and max lr of (1e-4,1e-3) [It's nearly my first run on intermediate res without any optimization]:\n<code>\nepoch   train_loss  valid_loss  d_kappa_score   kappa_k     kappa_r     time\n0   0.344915    0.283919    0.815302    0.791614    0.782375    09:47\n1   0.335626    0.285945    0.819567    0.763080    0.806630    09:46\n2   0.379465    0.268116    0.829591    0.842391    0.774673    09:48\n3   0.316176    0.233314    0.854984    0.860656    0.807655    09:50\n4   0.292933    0.361592    0.745009    0.664846    0.725402    09:49\n5   0.332429    0.244246    0.846598    0.859488    0.799022    09:54\n6   0.281213    0.257858    0.830744    0.812808    0.802393    09:52\n7   0.283063    0.226978    0.860622    0.841165    0.833361    09:52\n8   0.244150    0.212531    0.871141    0.873798    0.833969    09:55\n9   0.252861    0.231499    0.852559    0.857308    0.810779    09:52\n10  0.303027    0.253684    0.839478    0.855612    0.786137    09:53\n11  0.279593    0.214638    0.868501    0.869391    0.834018    09:55\n12  0.250773    0.203354    0.878194    0.886358    0.837234    09:56\n13  0.239266    0.203964    0.877210    0.887096    0.836048    09:58\n14  0.259900    0.193430    0.881883    0.889351    0.847430    09:58\n15  0.190317    0.215813    0.863405    0.861846    0.830002    10:01\n16  0.268994    0.207127    0.875795    0.881434    0.832446    10:03\n17  0.157652    0.210765    0.866115    0.870783    0.833130    10:05\n18  0.220389    0.195599    0.884750    0.882381    0.855319    10:07\n19  0.207806    0.190755    0.883505    0.888639    0.850455    10:08\n20  0.188866    0.204237    0.874564    0.883008    0.833871    10:10\n21  0.186228    0.202038    0.878626    0.887706    0.838257    10:10\n22  0.226118    0.236399    0.848630    0.874477    0.778278    10:10\n23  0.174489    0.184914    0.888356    0.896993    0.849679    10:09\n24  0.198199    0.202828    0.877724    0.870652    0.852208    10:09\n25  0.174334    0.176265    0.893822    0.906256    0.857446    10:10\n26  0.175599    0.179044    0.892619    0.901621    0.858356    10:02\n27  0.165271    0.177503    0.895425    0.904506    0.860751    09:58\n28  0.145186    0.181107    0.888282    0.895656    0.853604    09:55\n29  0.158865    0.179561    0.892518    0.895663    0.862738    09:57\n30  0.192933    0.182072    0.894250    0.905655    0.857467    09:56\n31  0.166352    0.171575    0.899262    0.900083    0.873222    10:00\n32  0.133393    0.185716    0.886050    0.896071    0.850738    09:59\n33  0.145026    0.190942    0.880710    0.891096    0.844466    10:00\n34  0.133870    0.170243    0.897480    0.907618    0.863825    10:02\n35  0.122654    0.169633    0.901467    0.911328    0.868964    10:04\n36  0.130831    0.178497    0.893576    0.904467    0.856378    10:07\n37  0.135886    0.172254    0.898089    0.908669    0.864243    09:53\n38  0.140824    0.167994    0.900024    0.912449    0.864562    09:55\n39  0.120919    0.168626    0.898917    0.911365    0.864052    09:56\n40  0.111123    0.166505    0.900856    0.908967    0.869388    09:58\n41  0.140773    0.167067    0.899353    0.908505    0.866439    10:03\n42  0.122937    0.168974    0.900271    0.911877    0.866887    10:03\n43  0.127722    0.170843    0.900197    0.909060    0.867776    10:06\n44  0.109837    0.168787    0.901611    0.908684    0.871180    09:51\n45  0.110626    0.165956    0.902472    0.912478    0.869869    16:13\n46  0.117968    0.167239    0.901686    0.910909    0.869625    09:55\n47  0.123976    0.166429    0.902506    0.913288    0.869211    09:56\n</code></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 966285,
          "author_name": "Srinath K",
          "author_url": "",
          "post_date": "2020-08-11T10:01:05.060000",
          "content": "<p>Thank you!<br>\n(Yes, I'm referring to a different kappa loss, and this is for initial run without noisy labels removal.)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 944025,
      "author_name": "Yousef Rabi",
      "author_url": "",
      "post_date": "2020-07-24T18:47:43.077000",
      "content": "<p>First off, congratulations! Well deserved.</p>\n\n<p>I tried to implement kappa loss, like so:</p>\n\n<p>class KappaLoss(nn.Module):</p>\n\n<pre><code>def __init__(self, labels_mean):\n    super().__init__()\n    self.labels_mean = labels_mean\n\ndef forward(self, y_hat, y):\n    y_hat.requires_grad = True\n    y_hat = y_hat - self.labels_mean\n    y = y - self.labels_mean\n    loss = (1 - (2 * torch.dot(y_hat, y) /\n            (torch.dot(y, y) + torch.dot(y_hat, y_hat))))\n\n    return loss\n</code></pre>\n\n<p>But when I run a training for one epoch and plot the loss vs learning rate, I get this:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2275947%2F69c1820e38c980307c6cce0e4f635906%2Flr_finder_16_batch.jpg?generation=1595616365041574&amp;alt=media\" alt=\"\"></p>\n\n<p>So I feel like something is wrong. Would love your input!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 944058,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-07-24T19:28:21.070000",
          "content": "<p>I have used the following one:</p>\n\n<p><code>\ny_shift = df.isup_grade.mean()\nNg = 6\ndef Kloss(x, target):\n    x = Ng*torch.sigmoid(x.float()).view(-1) - 0.5\n    target = target.float()\n    return 1.0 - (2.0*((x-y_shift)*(target-y_shift)).sum() - 1e-3)/\\\n        (((x-y_shift)**2).sum() + ((target-y_shift)**2).sum() + 1e-3)\n</code></p>\n\n<p>Though, without sigmoid it should also work. What is your bs?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 944102,
          "author_name": "Yousef Rabi",
          "author_url": "",
          "post_date": "2020-07-24T20:39:47.153000",
          "content": "<p>I used a batch size of 16.</p>\n\n<p>What is x in your case?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 944127,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-07-24T21:10:35.740000",
          "content": "<p>x is the output from the model. Try bs 64, in my procedure I first did pretraining on low res and then used intermediate res with bs = 8. Though, even with bs = 16 there should be convergence, not like u are showing. Did u try just to do regular training with lr ~1e-3? Also if u use sigmoid, the predicted labels in the metric evaluation should be computed as lp = (Ng*torch.sigmoid(x.float()).view(-1)).long().</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 944157,
          "author_name": "Yousef Rabi",
          "author_url": "",
          "post_date": "2020-07-24T22:00:32.883000",
          "content": "<p>Sorry I meant is x the output of a regression model or classification model? Are the targets just the isup_grades?</p>\n\n<p>What are the shapes of both?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 944174,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-07-24T22:26:42.047000",
          "content": "<p>The model outputs a single value per image if u want to use this loss (regression), so u get the same number of elements in your output and target. For the code above, the target is just isup_grades.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 944176,
          "author_name": "Yousef Rabi",
          "author_url": "",
          "post_date": "2020-07-24T22:30:10.643000",
          "content": "<p>Okay, that was my error, I was getting y_hat by argmaxing a softmax output. I'll retry. Thank you!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1036751,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-04T08:46:58.587000",
      "content": "<p>Congrats! Thank you for sharing.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1037230,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-10-04T17:51:31.370000",
          "content": "<p>You are welcome</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 943666,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-24T14:07:28.847000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 943838,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-07-24T16:04:10.490000",
          "content": "<p>Thanks</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "941402": "## Summary\n- Tile extraction is based on my [public pipeline](https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/146855) with **128x128x128** tiles from intermediate resolution layer\n- Label nose removal gives **~0.005 public and 0.01+ private LB boost**\n- tile cutout + tile selection augmentations\n- [kappa loss](https://arxiv.org/pdf/1509.07107v2.pdf)\n- majority voting ensemble of 8 **ResNeXt50** based models (**0.917 public and 0.930 private LB**)\n- more advanced tile selection could give **~0.004 boost** at private LB on average (and the maximum private LB score of **0.941**)\n\n\n## Introduction\n\nTo begin with, I would really like to express my gratitude to organizes and kaggle team for making this competition possible. It was really enjoying working on it and learned many new things. By sharing some of my ideas in this competition, such as [tile pooling base pipeline](https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/146855) used by many participants, gave me 3 kernel gold medals, so I have reached the kernel grand-master rank. And I have received my first solo competition gold medal. Also, I would like to say congratulations to all winners and people who received medals.\n\nHowever, this day is quite sad for many participants, especially ones who worked very hard throughout the entire competition and got down at private LB. The **chose of the metric by organizers could be done more wisely**: 500+500 test set is definitely not enough for QWK. It is not really normal when LB score is changing by 0.005+ when a different seed is used. The things got deteriorated when the third digit became available for LB score: many people got seduced by overfitting LB noise.\n\nBelow I outline the main things that worked for me. I have tried many more, but most of them have never worked, and I couldn't get any further improvement of my LB during the last month.\n\n## Main challenges\n\nThis competition to a large extent was about dealing with noisy data and train/test bias: as reported by organizers, the Redbound train data has only about **0.853** QWK, and I expect that Karolinska train data has 0.95-0.96 QWK. Beyond this, since Redbound data is graded by students, and Karolinska data is graded by only a single expert, while the test data is graded by 3 experts, there could be train/test bias because of the subjective opinion of people performing grading the train set. Therefore, **solely relaying on CV was not really good strategy in this competition**: at some point I saw a consistent decrease (~10 different models) of LB score when I ran training for longer, while CV was increasing. It confirms the hypothesis about the bias, and the trick was to train models only for limited number of epochs (even if CV could be increased), 32-48 depending on the setup, to **prevent learning the bias**.\n\nMeanwhile, LB was also not the best thing to trust because of severe noise, but some ppl tried to fit random seed as a hyperparameter 😄. The right thing, in my opinion, in this competition was to find the balance between CV and LB, and **trust to your intuition and the experience gained in the previous competitions**.\n\n## Noise\n\nIt is the most important part of this competition, in my opinion. After organizers have disclosed that there is a substantial level of noise, especially in Redbound train data, I have explored a number of techniques to deal with the noise: progressive label distillation, JoCoR (Joint Training with Co-Regularization), Co-teaching, negative learning, excluding hard examples from the batch, etc. However, most of them didn’t really work well here. The additional challenge is the bias between train and test and unstable LB. The thing I found to be the best for this data is removal of the uncertain examples from training set based on the out of fold predictions. I excluded ~1400 Redbound and 300 Karalinska data, so my clean training set contains about 8700 items. Relabeling the excluded images didn’t improve the performance. At the end of the competition [some ppl have discovered this trick as well](https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/161909), so I got nervous about my LB position 😬 \n**This trick gave ~0.005 public LB boost and 0.01+ private LB boost.**\n\n## Pipeline\n\nThe method I have used is mainly based on my [tile pooling pipeline](https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/146855) with several additional tricks:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1212661%2Fe6fe32d759a28480343001aa3c661723%2FTILE.png?generation=1588094975239255&amp;alt=media)\n\nBased on my public kernel, one could reach ~0.90 public and 0.91 private LB averaged (over different submissions) using 36x256x256 tile setup and the kappa loss (see below) without any other changes.\n\n[**kappa loss**](https://arxiv.org/pdf/1509.07107v2.pdf): I have used one minus\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1212661%2F99945116e2c9228e352645ee5f0bdfcc%2F2.png?generation=1595472932908419&amp;alt=media)\n\n(both predictions and labels are centered based on the mean value of labels). In my experiments I found that kappa loss &gt; sorted [binning loss](https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/155424) &gt; binning loss &gt; MSE &gt; CE. The only issue with the loss is that bs should be sufficiently large: I needed to pretrain models on low resolution and then continue training on intermediate resolution with bs =6-8 (**progressive resizing**), while at bs 1-2 I couldn't get convergence. The predicted value is limited within [-0.5,5.5] as `yp = 6*sigmoid(p) - 0.5`. In addition, I have CE aux for prediction of the Gleason score with 0.08 weight.\n\n**tile cutout**:  **Instead of using all tile tiles, why not to randomly select part of them** (let's say 96 out of 128). So, I can use large bs and the model is regularized in the same way as if cutout is used. It gave me quite good boost for CV and quite fair boost at LB.\n\n**128x128x128 tiles from intermediate resolution**: It appeared that many smaller tiles work better than 36x256x256. I think that it helps to select the tissue areas more effectively and at the same time prevents overfitting. \n\n**tile selection augmentation** The idea is quite simple: instead of [generating a single tile set](https://www.kaggle.com/iafoss/panda-16x128x128-tiles), I can generate 4 with adding sz/2 padding to x, y, or both before cutting the image into tiles and selecting ones having the most of the tissue. So, each tile in these 4 datasets will be different, but it is important not to mix tiles from them. During training I select the dataset by random, so effectively I have x4 data. It is an approximation of tile selection with random offset each time, which would be even more effective (based on my experience in Severstal competition), though, too slow to be used with intermediate res images. I also tried TTA based on tile selection, (as well as selection of the tile set with the largest tissue area out of 4), but I couldn't get any statistically significant improvement.\n\n**The above tricks gave ~0.005 boost over my baseline if I consider multiple submissions**. Though, score from submission to submission could change quite a bit.\n\n**Advanced tile selection**: In addition to my main pipeline I also tried to use the method proposed by @akensert [here](https://www.kaggle.com/akensert/panda-optimized-tiling-tf-data-dataset) with 128x128 tiles (but didn't use it for my final sub). It gave ~0.934 private LB single model performance on average for 8 different single 4 fold model subs (**and maximum 0.941 private LB**) and ~0.910 average public LB (~0.912 maximum). Too bad that I didn't create an ensemble based on this method. The trick that I have used in the model for training with such tiles is **n-pooling**: at test I apply the pooling only to nonempty tiles (n for a particular image, while the batch may contain some extra empty tiles for padding), and training is done with random selection of 96 tiles with repetitions (so I don't consider white tiles, which can change the mean statistics at pooling).\n\n**High resolution**: I tried to train several models on high res/2 resolution, 128x256x256 tiles. With tile cutout I could use batches of size 4 (and include 64 random tiles). However, the results were slightly worse than ones for 128x128x128 tiles from the intermediate resolution layer. It indicates that **going to higher resolution would likely provide only a minor boost**, even if I try to optimize my pipeline for training with small batches. Some idea I had is based on having two conv parts for intermediate and high res tiles. First pass through the low res model selects tiles having the highest uncertainty. Next, the selected tiles (but in high res) are passed through the second conv part. The produced feature maps are downscaled twice and replace the low res feature maps that had high uncertainty. Finally, pooling and head are applied to produce the final prediction. This method would allow to keep overall statistics of tiles with only correcting ones that model is not confident about. However, too large level of noise in the training set, noisy LB inconsistent with train labeling, and small potential gain, which would likely be overshadowed by the noise, have prevented me from going into this direction. Also, more complicated pipeline is more likely to be broken under such competition, where there is no certain way to evaluate the performance.\n\n**Augmentation**: I have used Albumentations with the following parameters:\n```\nCompose([\n        HorizontalFlip(),\n        VerticalFlip(),\n        RandomRotate90(),\n        ShiftScaleRotate(shift_limit=0.0625, scale_limit=0.3, rotate_limit=15, p=0.9, \n                         border_mode=cv2.BORDER_CONSTANT),\n        OneOf([#off in most cases\n            MotionBlur(blur_limit=3, p=0.1),\n            MedianBlur(blur_limit=3, p=0.1),\n            Blur(blur_limit=3, p=0.1),\n        ], p=0.2),\n        OneOf([#off in most cases\n            OpticalDistortion(p=0.3),\n            GridDistortion(p=.1),\n            IAAPiecewiseAffine(p=0.3),\n        ], p=0.3),\n        OneOf([\n            HueSaturationValue(10,15,10),\n            CLAHE(clip_limit=2),\n            RandomBrightnessContrast(),            \n        ], p=0.3),\n    ], p=1)\n```\n\n**Model**: All my models are based on **ResNeXt50**, similar to my public kernel, with batch norm in the head replaced with Group-norm. The optimizer, best model selection based on CV, and other things are similar to my public kernel, and I was using 32-48 epochs, depending on the setup. In addition, I tried ResNet34, ResNeXt101, and EfficientNet, while all of them were performing worse. I think ResNet34 may be not capable enough for this task, while ResNeXt101 is too large to do training on my computer with sufficient bs. However, I would say that **the model is the minor thing in this competition, and the main role is played by considering the noise and by optimizing the pipeline: there is no magic model, but there are hard work and solid understanding of the task and the data**.\n\n\n## Final ensemble\n\nThe submission that gave me the 11th place (**0.930 private LB/0.917 public LB**) is based on a majority voting ensemble of 8 models (4 fold) with 6 TTA. They are trained with different train/val splits and other modifications in the training procedure. On average each of the models trained in such manner gave **~0.930 private and ~0.910 public LB** single model 4 fold performance (with the **maximum of 0.938 and 0.916**, respectively). So, I got quite fair score, not good or bad luck (and my LB position almost haven’t changed). However, the large number of models was a way to survive in this storm. My another ensemble of 11 models, not selected as a final one, got 0.934 private LB. And as I mentioned above, more advanced tiling gives about **0.004 boost** on private LB (while similar public score as my main approach based on 128x128x128 tiles), with the average of **~0.934** and the maximum of **0.941 private LB**, but unfortunately, I haven’t built an ensemble based on them for my final submissions.\n\n**the code snippets are available at:** https://github.com/iafoss/PANDA\n\nAnd I would like to congratulate all participants and wish the best luck in the next competitions. I hope some of my tricks would be useful to you.\n",
    "941428": "Congrats solo gold !!!\nI think you are the player who has contributed the most to this competition. :)\nYour tile method helps almost player.",
    "1217213": "Congratulation! (after 7 months, I know I'm late here lol) I really love your approach!! Thank you so much for sharing those grate notebooks with explanations!",
    "964939": "Amazing  content. Keep sharing ",
    "953091": "Congratulations @iafoss for success on this kernal. Keep sharing nice things like this. 💕",
    "951890": "Congratulations @iafoss for this achievement. And thank for sharing the details.👍 ",
    "951849": "Congratulations :-)",
    "951694": "Congratulations! Learned a lot from your kernels!\nThank you!",
    "949602": "Congrats!",
    "949583": "Congratulations !!!",
    "948266": "Congrats!",
    "948010": "Simply brilliant! Well done @iafoss. I have learned and keep on learning a lot from people like you. Keep it up!✊ ",
    "946825": "@iafoss \nCongratulation on winning sole gold, and once again for becoming kernel GM. 🔥 ",
    "943175": "Well done, big thanks for your starter kernel and your help during this competition! you deserve top 3 haha!",
    "942626": "congrats!  I wish I had known about kappa loss.  How much did you get from it compared to binning loss when you compared?",
    "942498": "Congrats on making GM! You really did a great job this comp not just as a competitor but also sharing simple and effective ideas!",
    "941985": "@iafoss Congratulations for your solo gold. Your notebooks are game changers in this competition.Thank you very much for sharing your knowledge. You are like a invisible team mate in every team.",
    "941737": "@iafoss thank you for putting this level of hardwork into the competition. I really learned a lot from your kernels !\n\nIf you don't mind, I have some questions for you:\n\n1) if I put y=2 and y_hat=2 in your formula above, I get a loss of one for a perfect prediction. How does that work exactly ?\n\n2) you mention not getting statistically significant improvement from TTA based on tile selection. How do you test for this ? is it gut feeling, or do you run the experiment multiple times and get statistics like mean and std of loss/kappa_metric (which seems rather impractical given the time it takes to train a model)\n\n3) How did you pick which images to remove from the training dataset ?\n\n4) \"I have explored a number of techniques to deal with the noise: progressive label distillation, JoCoR (Joint Training with Co-Regularization), Co-teaching, negative learning, excluding hard examples from the batch, etc. \"\nThat's the hell of a lot of stuff to try. What's your rationale for testing so many things ? Do you give yourself some time (\"I'll try this for three days and if no improvement I quit\"), or put more simply, when/how do you decide to quit a dead end ? \n\n\nAnd finally, one comment: if you resize lvl0 slide down by four, you get the exact same dimensions than lvl1. Yet, if you zoom on one part of the image and plot it twice (one for lvl0//4, one for lvl1), you'll see a neat difference in quality! That's because the images are compressed through JPEG inside the tiff format. So lvl0/4 &gt;&gt; lvl1 in terms of quality. Yet training the same pipe on lvl0/4 gets worse results than lvl1 slides. I believe the jpeg compression could act like a first layer of convolution, roughly speaking. Hence the not so better score using highest resolution ? ",
    "941690": "Congratulations! Well deserved!",
    "941535": "Congratulations on your first solo competition gold medal and kernel grand-master rank!!  You contribute so much and explain in a such good way for all here.  Pretty sure everyone in this was helped by you.  Really felt for you when prize rank did not happen, was so sure it would.  Of course, always good luck for next time.     ",
    "2555976": "Great Job!",
    "1084059": "Hi lafoss, thank a lot for sharing, I wanted to comeback and study little bit more your solution with fastai, I was wondering if you can share too and have a link for the submission notebook you used last?\n\nThanks again.",
    "957245": "@iafoss  We couldn't get some of our approaches to yield good performance with kappa loss. We based our approach on [this ](https://www.sciencedirect.com/science/article/abs/pii/S0167865517301666) paper, which has been recently implemented in Tensorflow addons. \n\nCould you please share some more details of your training e.g. how many epochs it took to converge in the first run at low-res, the progression of qkappa metric at the intermediate resolution and how many epochs it took?\n\nWe also observed that batch size of at least 5 was needed for convergence.\n\nBut in our case, we couldn't get over 0.7 qkappa on our validation set with various LR/batch-sizes based on ballpark from above paper.\n(The same architecture, even with categorical cross-entropy loss, went upto 0.85 qkappa on validation set.)\n\nAny thoughts/suggestions on this would be helpful, as this is our first outing with kappa loss.\n\nThanks, and congratulations!",
    "944025": "First off, congratulations! Well deserved.\n\nI tried to implement kappa loss, like so:\n\nclass KappaLoss(nn.Module):\n\n    def __init__(self, labels_mean):\n        super().__init__()\n        self.labels_mean = labels_mean\n\n    def forward(self, y_hat, y):\n        y_hat.requires_grad = True\n        y_hat = y_hat - self.labels_mean\n        y = y - self.labels_mean\n        loss = (1 - (2 * torch.dot(y_hat, y) /\n                (torch.dot(y, y) + torch.dot(y_hat, y_hat))))\n\n        return loss\n\nBut when I run a training for one epoch and plot the loss vs learning rate, I get this:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2275947%2F69c1820e38c980307c6cce0e4f635906%2Flr_finder_16_batch.jpg?generation=1595616365041574&amp;alt=media)\n\nSo I feel like something is wrong. Would love your input!",
    "1036751": "Congrats! Thank you for sharing.",
    "943666": ""
  }
}