{
  "id": 169230,
  "title": "6th place solution : noise robust learning [BarelyBears]",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/169230",
  "author_name": "RabotniKuma",
  "post_date": "2020-07-23T09:28:25.175000",
  "votes": 55,
  "comment_count": 9,
  "views": 0,
  "content": "<p>First of all, we would like to thank organizers for such an interesting and realistic problem. \nImplementations are available:\n<a href=\"https://github.com/analokmaus/kaggle-panda-challenge-public\">https://github.com/analokmaus/kaggle-panda-challenge-public</a></p>\n\n<h1>TL;DR</h1>\n\n<p>Label noise is the biggest challenge in this competition.\nWe used <strong>online uncertainty sample mining(OUSM)</strong>  and <strong>mixup</strong> to robustly fit CNN models, and blended 4 models with different settings to stabilize the results. \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1973217%2F7c9a14d7f3d3de7bf65ce0b4c363fe69%2Fkaggle-panda-challenge.001.png?generation=1595831415189859&amp;alt=media&amp;width=500\" alt=\"\"></p>\n\n<h1>Tile-based multi instance learning model</h1>\n\n<p>The first challenge in this comp was how to deal with those extremely large images. Thanks to <a href=\"/iafoss\">@iafoss</a> ’s great notebook, we used almost identical model with various backbones and tile sizes. We modified classifier part to ordinal regression (<a href=\"https://arxiv.org/abs/1901.07884\">CORAL loss</a>).</p>\n\n<h1>Preprocessing and data augmentation</h1>\n\n<p>Data augmentation for tile-based CNN model can be applied in two ways: slide level and tile level. Slide level augmentations are applied to whole slide images before tiles are extracted. Since the point is to create slightly different tile sets, we used shift, scale, rotate.\nTile level augmentations aim to improve feature extractor performance, and we used shift, scale, rotate, flip, and random dropout. The idea of random dropout is to randomly fill a tile with mean pixel value and regularize the model.</p>\n\n<h1>Postprocessing</h1>\n\n<p>We used 4 times TTA during inference and optimized thresholds to maximize QWK values. </p>\n\n<h1>Validation strategy</h1>\n\n<p>As written in task description, the label quality in train data differs a lot from those in test data. So from the very beginning we assumed this part would be critical in this comp. Roughly speaking, \ntrain data: noisy and big\npublic test data: clean but small\nprivate test data: clean but small\nSo our strategy is <strong>IGNORE CV, CARE ABOUT PUBLIC LB, AND TRUST METHODOLOGY.</strong>\nFor us, the results were unstable due to small size of test data, but not so ‘lottery’.</p>\n\n<h1>Handling noisy labels</h1>\n\n<p>We read tens of papers about handling noisy labels, and implemented some of them such as:</p>\n\n<ul>\n<li>loss functions (DMI loss, DAC loss, Symmetric loss, <strong>OUSM loss</strong>, etc..)</li>\n<li>training procedure (CleanNet, Iterative Self-training, <strong>mixup</strong>, etc..)</li>\n</ul>\n\n<p>The common ideas among them are that noisy samples should have different features from  correct samples, thus noisy ones should have bigger loss. \nOUSM(Online Uncertainty Sample Mining) is an approach, in which samples with high loss are excluded from each mini-batch. According to <a href=\"https://arxiv.org/abs/1901.07759\">previous research</a>, this method works with skin lesion classification problem where similar kind of label noise exists. In PANDA competition, it gave us stable boost from around 0.87 to 0.90 on public LB. \nThen we trained models with different random seeds, and collected samples which are often judged as noise(with big loss). We excluded 10% of ‘most likely to be noisy’ samples from each label because due to the imbalance in label distribution, grade &gt;= 2 samples are more likely to be judged as noise. This new datasets should be less noisy than the original one, and models trained on this new dataset achieve 0.91 on public LB.\nApart from OUSM, mixup also showed good performance on public LB. This is consistent with <a href=\"https://arxiv.org/abs/1710.09412\">original paper</a> which reported performance improvement with label corruption.</p>\n\n<h1>Pipeline overview</h1>\n\n<p>Our pipeline is simple average of the following models</p>\n\n<ul>\n<li>5 fold 224x64Tile-based model, se-resnext50 (OUSM)</li>\n<li>5 fold 224x64 Tile-based model, se-resnext50 (OUSM with different params)</li>\n<li>5 fold 224x64 Tile-based model, se-resnext101 (OUSM)</li>\n<li>5 fold 256x36 Tile-based model, efficientnet-b0 (mixup)\nThis model scored 0.903 on public LB, and 0.932 on private LB. \nCompared to <a href=\"/haqishen\">@haqishen</a> 's model with no denoising, our final model showed +0.018 on public and +0.017 on private.</li>\n</ul>",
  "messages": [
    {
      "id": 941533,
      "postDate": "2020-07-23T09:28:25.177Z",
      "content": "<p>First of all, we would like to thank organizers for such an interesting and realistic problem. \nImplementations are available:\n<a href=\"https://github.com/analokmaus/kaggle-panda-challenge-public\">https://github.com/analokmaus/kaggle-panda-challenge-public</a></p>\n\n<h1>TL;DR</h1>\n\n<p>Label noise is the biggest challenge in this competition.\nWe used <strong>online uncertainty sample mining(OUSM)</strong>  and <strong>mixup</strong> to robustly fit CNN models, and blended 4 models with different settings to stabilize the results. \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1973217%2F7c9a14d7f3d3de7bf65ce0b4c363fe69%2Fkaggle-panda-challenge.001.png?generation=1595831415189859&amp;alt=media&amp;width=500\" alt=\"\"></p>\n\n<h1>Tile-based multi instance learning model</h1>\n\n<p>The first challenge in this comp was how to deal with those extremely large images. Thanks to <a href=\"/iafoss\">@iafoss</a> ’s great notebook, we used almost identical model with various backbones and tile sizes. We modified classifier part to ordinal regression (<a href=\"https://arxiv.org/abs/1901.07884\">CORAL loss</a>).</p>\n\n<h1>Preprocessing and data augmentation</h1>\n\n<p>Data augmentation for tile-based CNN model can be applied in two ways: slide level and tile level. Slide level augmentations are applied to whole slide images before tiles are extracted. Since the point is to create slightly different tile sets, we used shift, scale, rotate.\nTile level augmentations aim to improve feature extractor performance, and we used shift, scale, rotate, flip, and random dropout. The idea of random dropout is to randomly fill a tile with mean pixel value and regularize the model.</p>\n\n<h1>Postprocessing</h1>\n\n<p>We used 4 times TTA during inference and optimized thresholds to maximize QWK values. </p>\n\n<h1>Validation strategy</h1>\n\n<p>As written in task description, the label quality in train data differs a lot from those in test data. So from the very beginning we assumed this part would be critical in this comp. Roughly speaking, \ntrain data: noisy and big\npublic test data: clean but small\nprivate test data: clean but small\nSo our strategy is <strong>IGNORE CV, CARE ABOUT PUBLIC LB, AND TRUST METHODOLOGY.</strong>\nFor us, the results were unstable due to small size of test data, but not so ‘lottery’.</p>\n\n<h1>Handling noisy labels</h1>\n\n<p>We read tens of papers about handling noisy labels, and implemented some of them such as:</p>\n\n<ul>\n<li>loss functions (DMI loss, DAC loss, Symmetric loss, <strong>OUSM loss</strong>, etc..)</li>\n<li>training procedure (CleanNet, Iterative Self-training, <strong>mixup</strong>, etc..)</li>\n</ul>\n\n<p>The common ideas among them are that noisy samples should have different features from  correct samples, thus noisy ones should have bigger loss. \nOUSM(Online Uncertainty Sample Mining) is an approach, in which samples with high loss are excluded from each mini-batch. According to <a href=\"https://arxiv.org/abs/1901.07759\">previous research</a>, this method works with skin lesion classification problem where similar kind of label noise exists. In PANDA competition, it gave us stable boost from around 0.87 to 0.90 on public LB. \nThen we trained models with different random seeds, and collected samples which are often judged as noise(with big loss). We excluded 10% of ‘most likely to be noisy’ samples from each label because due to the imbalance in label distribution, grade &gt;= 2 samples are more likely to be judged as noise. This new datasets should be less noisy than the original one, and models trained on this new dataset achieve 0.91 on public LB.\nApart from OUSM, mixup also showed good performance on public LB. This is consistent with <a href=\"https://arxiv.org/abs/1710.09412\">original paper</a> which reported performance improvement with label corruption.</p>\n\n<h1>Pipeline overview</h1>\n\n<p>Our pipeline is simple average of the following models</p>\n\n<ul>\n<li>5 fold 224x64Tile-based model, se-resnext50 (OUSM)</li>\n<li>5 fold 224x64 Tile-based model, se-resnext50 (OUSM with different params)</li>\n<li>5 fold 224x64 Tile-based model, se-resnext101 (OUSM)</li>\n<li>5 fold 256x36 Tile-based model, efficientnet-b0 (mixup)\nThis model scored 0.903 on public LB, and 0.932 on private LB. \nCompared to <a href=\"/haqishen\">@haqishen</a> 's model with no denoising, our final model showed +0.018 on public and +0.017 on private.</li>\n</ul>",
      "rawMarkdown": "First of all, we would like to thank organizers for such an interesting and realistic problem. \nImplementations are available:\nhttps://github.com/analokmaus/kaggle-panda-challenge-public\n\n# TL;DR\nLabel noise is the biggest challenge in this competition.\nWe used **online uncertainty sample mining(OUSM)**  and **mixup** to robustly fit CNN models, and blended 4 models with different settings to stabilize the results. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1973217%2F7c9a14d7f3d3de7bf65ce0b4c363fe69%2Fkaggle-panda-challenge.001.png?generation=1595831415189859&amp;alt=media&amp;width=500)\n\n\n# Tile-based multi instance learning model\nThe first challenge in this comp was how to deal with those extremely large images. Thanks to @iafoss ’s great notebook, we used almost identical model with various backbones and tile sizes. We modified classifier part to ordinal regression ([CORAL loss](https://arxiv.org/abs/1901.07884)).\n\n# Preprocessing and data augmentation\nData augmentation for tile-based CNN model can be applied in two ways: slide level and tile level. Slide level augmentations are applied to whole slide images before tiles are extracted. Since the point is to create slightly different tile sets, we used shift, scale, rotate.\nTile level augmentations aim to improve feature extractor performance, and we used shift, scale, rotate, flip, and random dropout. The idea of random dropout is to randomly fill a tile with mean pixel value and regularize the model.\n\n# Postprocessing\nWe used 4 times TTA during inference and optimized thresholds to maximize QWK values. \n\n# Validation strategy\nAs written in task description, the label quality in train data differs a lot from those in test data. So from the very beginning we assumed this part would be critical in this comp. Roughly speaking, \ntrain data: noisy and big\npublic test data: clean but small\nprivate test data: clean but small\nSo our strategy is **IGNORE CV, CARE ABOUT PUBLIC LB, AND TRUST METHODOLOGY.**\nFor us, the results were unstable due to small size of test data, but not so ‘lottery’.\n\n# Handling noisy labels\nWe read tens of papers about handling noisy labels, and implemented some of them such as:\n\n- loss functions (DMI loss, DAC loss, Symmetric loss, **OUSM loss**, etc..)\n- training procedure (CleanNet, Iterative Self-training, **mixup**, etc..)\n\nThe common ideas among them are that noisy samples should have different features from  correct samples, thus noisy ones should have bigger loss. \nOUSM(Online Uncertainty Sample Mining) is an approach, in which samples with high loss are excluded from each mini-batch. According to [previous research](https://arxiv.org/abs/1901.07759), this method works with skin lesion classification problem where similar kind of label noise exists. In PANDA competition, it gave us stable boost from around 0.87 to 0.90 on public LB. \nThen we trained models with different random seeds, and collected samples which are often judged as noise(with big loss). We excluded 10% of ‘most likely to be noisy’ samples from each label because due to the imbalance in label distribution, grade &gt;= 2 samples are more likely to be judged as noise. This new datasets should be less noisy than the original one, and models trained on this new dataset achieve 0.91 on public LB.\nApart from OUSM, mixup also showed good performance on public LB. This is consistent with [original paper](https://arxiv.org/abs/1710.09412) which reported performance improvement with label corruption.\n\n# Pipeline overview\n\nOur pipeline is simple average of the following models\n\n- 5 fold 224x64Tile-based model, se-resnext50 (OUSM)\n- 5 fold 224x64 Tile-based model, se-resnext50 (OUSM with different params)\n- 5 fold 224x64 Tile-based model, se-resnext101 (OUSM)\n- 5 fold 256x36 Tile-based model, efficientnet-b0 (mixup)\nThis model scored 0.903 on public LB, and 0.932 on private LB. \nCompared to @haqishen 's model with no denoising, our final model showed +0.018 on public and +0.017 on private.",
      "votes": 55
    },
    {
      "id": 1121965,
      "postDate": "2020-12-22T04:12:30.010Z",
      "content": "<p>Thank you for this great knowledge sharing.   </p>",
      "rawMarkdown": "Thank you for this great knowledge sharing.   ",
      "votes": -1
    },
    {
      "id": 1592976,
      "postDate": "2021-11-23T14:12:45.687Z",
      "content": "<p>Thank you for sharing! Does trust LB mean that it should be validated with a clean dataset?<br>\n I'm wondering what to do if the percentage of public LB data is very small.<br>\nShould I remove the label noise (and hard sample) from the training data before trusting the CV?</p>",
      "rawMarkdown": "Thank you for sharing! Does trust LB mean that it should be validated with a clean dataset?\n I'm wondering what to do if the percentage of public LB data is very small.\nShould I remove the label noise (and hard sample) from the training data before trusting the CV?"
    },
    {
      "id": 947507,
      "postDate": "2020-07-27T10:04:42.500Z",
      "content": "<p>Nice</p>",
      "rawMarkdown": "Nice"
    },
    {
      "id": 946442,
      "postDate": "2020-07-26T15:24:12.933Z",
      "content": "<p>Congrats ! I like the way on how you handled the noise\nthanks for your code btw, very informative</p>",
      "rawMarkdown": "Congrats ! I like the way on how you handled the noise\nthanks for your code btw, very informative"
    },
    {
      "id": 941561,
      "postDate": "2020-07-23T09:51:36.950Z",
      "content": "<p>congratulations and  thank you for sharing this awesome solution,,will you please share the  code you used for online uncertainty sample mining(OUSM)? thank  you</p>",
      "rawMarkdown": "congratulations and  thank you for sharing this awesome solution,,will you please share the  code you used for online uncertainty sample mining(OUSM)? thank  you",
      "replies": [
        {
          "id": 941569,
          "postDate": "2020-07-23T09:58:04.670Z",
          "content": "<p>Thank you. I’m currently cleaning our git repo, one moment please :)</p>",
          "rawMarkdown": "Thank you. I’m currently cleaning our git repo, one moment please :)",
          "votes": 1
        },
        {
          "id": 941576,
          "postDate": "2020-07-23T10:00:04.413Z",
          "content": "<p><a href=\"/analokamus\">@analokamus</a>  can't  wait for one moment,i am desperate for awesome OUSM code,not finding it in google,have you implemented it or collected from google? i can't find code for that. great job :)</p>",
          "rawMarkdown": "@analokamus  can't  wait for one moment,i am desperate for awesome OUSM code,not finding it in google,have you implemented it or collected from google? i can't find code for that. great job :)"
        },
        {
          "id": 946236,
          "postDate": "2020-07-26T13:02:44.910Z",
          "content": "<p><a href=\"/mobassir\">@mobassir</a> \nOUSM implementation is here: \n<a href=\"https://github.com/analokmaus/kaggle-panda-challenge-public/blob/208caf4c83a5ab9d181e66eee447cd2e475d58dc/models/noisy_loss.py#L97\">https://github.com/analokmaus/kaggle-panda-challenge-public/blob/208caf4c83a5ab9d181e66eee447cd2e475d58dc/models/noisy_loss.py#L97</a>\nIt is very simple!</p>",
          "rawMarkdown": "@mobassir \nOUSM implementation is here: \nhttps://github.com/analokmaus/kaggle-panda-challenge-public/blob/208caf4c83a5ab9d181e66eee447cd2e475d58dc/models/noisy_loss.py#L97\nIt is very simple!",
          "votes": 2
        },
        {
          "id": 947517,
          "postDate": "2020-07-27T10:12:07.923Z",
          "content": "<p>thank you mate <a href=\"/analokamus\">@analokamus</a> </p>",
          "rawMarkdown": "thank you mate @analokamus "
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1121965,
      "author_name": "AIFahim",
      "author_url": "",
      "post_date": "2020-12-22T04:12:30.010000",
      "content": "<p>Thank you for this great knowledge sharing.   </p>",
      "votes": -1,
      "replies": []
    },
    {
      "id": 1592976,
      "author_name": "patriot",
      "author_url": "",
      "post_date": "2021-11-23T14:12:45.687000",
      "content": "<p>Thank you for sharing! Does trust LB mean that it should be validated with a clean dataset?<br>\n I'm wondering what to do if the percentage of public LB data is very small.<br>\nShould I remove the label noise (and hard sample) from the training data before trusting the CV?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 947507,
      "author_name": "fromavocado",
      "author_url": "",
      "post_date": "2020-07-27T10:04:42.500000",
      "content": "<p>Nice</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 946442,
      "author_name": "Alex",
      "author_url": "",
      "post_date": "2020-07-26T15:24:12.933000",
      "content": "<p>Congrats ! I like the way on how you handled the noise\nthanks for your code btw, very informative</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 941561,
      "author_name": "Mobassir",
      "author_url": "",
      "post_date": "2020-07-23T09:51:36.950000",
      "content": "<p>congratulations and  thank you for sharing this awesome solution,,will you please share the  code you used for online uncertainty sample mining(OUSM)? thank  you</p>",
      "votes": 0,
      "replies": [
        {
          "id": 941569,
          "author_name": "RabotniKuma",
          "author_url": "",
          "post_date": "2020-07-23T09:58:04.670000",
          "content": "<p>Thank you. I’m currently cleaning our git repo, one moment please :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 941576,
          "author_name": "Mobassir",
          "author_url": "",
          "post_date": "2020-07-23T10:00:04.413000",
          "content": "<p><a href=\"/analokamus\">@analokamus</a>  can't  wait for one moment,i am desperate for awesome OUSM code,not finding it in google,have you implemented it or collected from google? i can't find code for that. great job :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 946236,
          "author_name": "RabotniKuma",
          "author_url": "",
          "post_date": "2020-07-26T13:02:44.910000",
          "content": "<p><a href=\"/mobassir\">@mobassir</a> \nOUSM implementation is here: \n<a href=\"https://github.com/analokmaus/kaggle-panda-challenge-public/blob/208caf4c83a5ab9d181e66eee447cd2e475d58dc/models/noisy_loss.py#L97\">https://github.com/analokmaus/kaggle-panda-challenge-public/blob/208caf4c83a5ab9d181e66eee447cd2e475d58dc/models/noisy_loss.py#L97</a>\nIt is very simple!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 947517,
          "author_name": "Mobassir",
          "author_url": "",
          "post_date": "2020-07-27T10:12:07.923000",
          "content": "<p>thank you mate <a href=\"/analokamus\">@analokamus</a> </p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "941533": "First of all, we would like to thank organizers for such an interesting and realistic problem. \nImplementations are available:\nhttps://github.com/analokmaus/kaggle-panda-challenge-public\n\n# TL;DR\nLabel noise is the biggest challenge in this competition.\nWe used **online uncertainty sample mining(OUSM)**  and **mixup** to robustly fit CNN models, and blended 4 models with different settings to stabilize the results. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1973217%2F7c9a14d7f3d3de7bf65ce0b4c363fe69%2Fkaggle-panda-challenge.001.png?generation=1595831415189859&amp;alt=media&amp;width=500)\n\n\n# Tile-based multi instance learning model\nThe first challenge in this comp was how to deal with those extremely large images. Thanks to @iafoss ’s great notebook, we used almost identical model with various backbones and tile sizes. We modified classifier part to ordinal regression ([CORAL loss](https://arxiv.org/abs/1901.07884)).\n\n# Preprocessing and data augmentation\nData augmentation for tile-based CNN model can be applied in two ways: slide level and tile level. Slide level augmentations are applied to whole slide images before tiles are extracted. Since the point is to create slightly different tile sets, we used shift, scale, rotate.\nTile level augmentations aim to improve feature extractor performance, and we used shift, scale, rotate, flip, and random dropout. The idea of random dropout is to randomly fill a tile with mean pixel value and regularize the model.\n\n# Postprocessing\nWe used 4 times TTA during inference and optimized thresholds to maximize QWK values. \n\n# Validation strategy\nAs written in task description, the label quality in train data differs a lot from those in test data. So from the very beginning we assumed this part would be critical in this comp. Roughly speaking, \ntrain data: noisy and big\npublic test data: clean but small\nprivate test data: clean but small\nSo our strategy is **IGNORE CV, CARE ABOUT PUBLIC LB, AND TRUST METHODOLOGY.**\nFor us, the results were unstable due to small size of test data, but not so ‘lottery’.\n\n# Handling noisy labels\nWe read tens of papers about handling noisy labels, and implemented some of them such as:\n\n- loss functions (DMI loss, DAC loss, Symmetric loss, **OUSM loss**, etc..)\n- training procedure (CleanNet, Iterative Self-training, **mixup**, etc..)\n\nThe common ideas among them are that noisy samples should have different features from  correct samples, thus noisy ones should have bigger loss. \nOUSM(Online Uncertainty Sample Mining) is an approach, in which samples with high loss are excluded from each mini-batch. According to [previous research](https://arxiv.org/abs/1901.07759), this method works with skin lesion classification problem where similar kind of label noise exists. In PANDA competition, it gave us stable boost from around 0.87 to 0.90 on public LB. \nThen we trained models with different random seeds, and collected samples which are often judged as noise(with big loss). We excluded 10% of ‘most likely to be noisy’ samples from each label because due to the imbalance in label distribution, grade &gt;= 2 samples are more likely to be judged as noise. This new datasets should be less noisy than the original one, and models trained on this new dataset achieve 0.91 on public LB.\nApart from OUSM, mixup also showed good performance on public LB. This is consistent with [original paper](https://arxiv.org/abs/1710.09412) which reported performance improvement with label corruption.\n\n# Pipeline overview\n\nOur pipeline is simple average of the following models\n\n- 5 fold 224x64Tile-based model, se-resnext50 (OUSM)\n- 5 fold 224x64 Tile-based model, se-resnext50 (OUSM with different params)\n- 5 fold 224x64 Tile-based model, se-resnext101 (OUSM)\n- 5 fold 256x36 Tile-based model, efficientnet-b0 (mixup)\nThis model scored 0.903 on public LB, and 0.932 on private LB. \nCompared to @haqishen 's model with no denoising, our final model showed +0.018 on public and +0.017 on private.",
    "1121965": "Thank you for this great knowledge sharing.   ",
    "1592976": "Thank you for sharing! Does trust LB mean that it should be validated with a clean dataset?\n I'm wondering what to do if the percentage of public LB data is very small.\nShould I remove the label noise (and hard sample) from the training data before trusting the CV?",
    "947507": "Nice",
    "946442": "Congrats ! I like the way on how you handled the noise\nthanks for your code btw, very informative",
    "941561": "congratulations and  thank you for sharing this awesome solution,,will you please share the  code you used for online uncertainty sample mining(OUSM)? thank  you"
  }
}