{
  "id": 169143,
  "title": "1st Place Solution [PND]",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/169143",
  "author_name": "fam_taro",
  "post_date": "2020-07-23T03:20:51.396000",
  "votes": 225,
  "comment_count": 116,
  "views": 0,
  "content": "<p>Congratulations to everyone and thanks for the hosts for preparing this competition!</p>\n<h4>I published <a href=\"https://docs.google.com/presentation/d/1Ies4vnyVtW5U3XNDr_fom43ZJDIodu1SV6DSK8di6fs/edit?usp=sharing\" target=\"_blank\">slide</a>!</h4>\n<h4>Our code is <a href=\"https://github.com/kentaroy47/Kaggle-PANDA-1st-place-solution\" target=\"_blank\">here</a>!</h4>\n<h1>Proposed Denoising Method</h1>\n<p>We're very suprised that we finished 1st, and our simple label-denoising method (suprisingly) boosted up PB.</p>\n<p>The competition was all about handling noisy labels, so we worked hard on finding good ways to denoising.</p>\n<p>Here is our simple denoising method by <a href=\"https://www.kaggle.com/kyoshioka47\" target=\"_blank\">@kyoshioka47</a>:</p>\n<h2>Getting cleaned labels</h2>\n<ul>\n<li>train k-folds with effnet-b1 (Almost identical to Qishen's kernel)  <ul>\n<li>Model specifics in fam_taro( <a href=\"https://www.kaggle.com/yukkyo\" target=\"_blank\">@yukkyo</a> ) part</li></ul></li>\n<li>Predict hold-out sets with the trained model. We get <code>pred</code> with this step.</li>\n<li>Remove the training data which has a high disparity between ground truth and pred. The filtered labels will be called cleaned labels.</li>\n</ul>\n<p>We calculate <code>disparity</code> by the absolute difference of ISUP between GT and pred. Data with disparity larger than 1.6 was simply removed.  </p>\n<p>Here is the psuedo codes. probs_raw is the raw prediciton results (ISUP)</p>\n<pre><code># Base arutema method\ndef remove_noisy(df, thresh):\n    gap = np.abs(df[\"isup_grade\"] - df[\"probs_raw\"])\n    df_removed = df[gap &gt; thresh].reset_index(drop=True)\n    df_keep = df[gap &lt;= thresh].reset_index(drop=True)\n    return df_keep, df_removed\n\ndf_keep, df_remove = remove_noisy(df, thresh=1.6)\nshow_keep_remove(df, df_keep, df_remove)\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1670024%2F6f6bbe0dfdb2bd5ba10057a1ba32f040%2Farutema.png?generation=1595474367096439&amp;alt=media\" alt=\"\"></p>\n<h2>Retraining</h2>\n<p>Retrain model with using the denoised labels. <br>\nWe get CV 0.94 LB 0.90 PB 0.934 with a simple Qishen Eff-b0 model with k-folds.<br>\nEnsambling with different models further boosted to 1st place.</p>\n<p>We tried CleanLab too, but that did not perform well in CV/LB so we sticked with this.</p>\n<h1>1. Our final submission</h1>\n<ul>\n<li>Select 1 (public LB 0.910, private LB 0.922)<ul>\n<li>Resnext50_32x4d(poteman)</li></ul></li>\n<li>Select 2 (public LB 0.904, private LB 0.940)<ul>\n<li>Effnet-B0(arutema47) + Effnet-B1(fam_taro)<ul>\n<li>Simple average (<code>1 : 1</code>)</li></ul></li></ul></li>\n</ul>\n<p>Suprisingly, even with several weight patterns, the PB was 0.940.</p>\n<h1>2. Resnext50_32x4d( <a href=\"https://www.kaggle.com/poteman\" target=\"_blank\">@poteman</a> ), public 0.910, private 0.922</h1>\n<p>This was our best LB model.</p>\n<ul>\n<li>Split kfold: stratified kfold with imghash(threshold 0.90)</li>\n<li>iafoss tile method<ul>\n<li>tile size 256, tile num 64</li></ul></li>\n<li>model:resnext50_32x4d</li>\n<li>head: 3 * reg_head + 1 * softmax head</li>\n</ul>\n<h1>3. Effnet-B1(fam_taro), public 0.901, private 0.932</h1>\n<ul>\n<li>Split kfold<ul>\n<li>stratified 5 kfold with gleason-score and imghash similarity (threshold 0.90)<ul>\n<li>convert <code>negative</code> to <code>0+0</code></li>\n<li>how to grouping by imghash similarity<ul>\n<li>This is based on <a href=\"https://www.kaggle.com/appian\" target=\"_blank\">@appian</a> 's kernel<ul>\n<li><a href=\"https://www.kaggle.com/appian/panda-imagehash-to-detect-duplicate-images\" target=\"_blank\">https://www.kaggle.com/appian/panda-imagehash-to-detect-duplicate-images</a></li></ul></li>\n<li><a href=\"https://www.kaggle.com/yukkyo/imagehash-to-detect-duplicate-images-and-grouping\" target=\"_blank\">https://www.kaggle.com/yukkyo/imagehash-to-detect-duplicate-images-and-grouping</a></li></ul></li></ul></li>\n<li>In my opinion, split method is import point for our denoise method.<ul>\n<li>Because we use prediction of out of fold</li>\n<li>If you put the duplicate images in a different fold, I don't think denoise will work for them</li></ul></li></ul></li>\n<li>Data<ul>\n<li>iafoss tile method</li>\n<li>tile size 192, tile num 64</li></ul></li>\n<li>Model: Effnet-B1 + GeM<ul>\n<li>label: isup-grade and first score of gleason(10 dim bin)</li></ul></li>\n<li>Make final sub by 3 steps<ul>\n<li>Local train &amp; predict</li>\n<li>Remove noisy label<ul>\n<li>extended <a href=\"https://www.kaggle.com/kyoshioka47\" target=\"_blank\">@kyoshioka47</a> method</li>\n<li>Change threshold for each isup-grade and data-provider</li></ul></li>\n<li>Re-train</li></ul></li>\n<li>Not work for me<ul>\n<li>Remove noisy by confident-learning</li>\n<li>Cycle GAN augmentation(karolinska  radboud)</li>\n<li>test with AdaBN &amp; Freezing BN at train</li>\n<li>CutMix, Mixup (before denoising)</li></ul></li>\n</ul>\n<pre><code>def remove_noisy2(df, thresholds):\n    gap = np.abs(df[\"isup_grade\"] - df[\"probs_raw\"])\n\n    df_keeps = list()\n    df_removes = list()\n\n    for label, thresh in enumerate(thresholds):\n        df_tmp = df[df.isup_grade == label].reset_index(drop=True)\n        gap_tmp = gap[df.isup_grade == label].reset_index(drop=True)\n\n        df_remove_tmp = df_tmp[gap_tmp &gt; thresh].reset_index(drop=True)\n        df_keep_tmp = df_tmp[gap_tmp &lt;= thresh].reset_index(drop=True)\n\n        df_removes.append(df_remove_tmp)\n        df_keeps.append(df_keep_tmp)\n\n    df_keep = pd.concat(df_keeps, axis=0)\n    df_removed = pd.concat(df_removes, axis=0)\n    return df_keep, df_removed\n\ndef remove_noisy3(df, thresholds_rad, thresholds_ka):\n    df_r = df[df.data_provider == \"radboud\"].reset_index(drop=True)\n    df_k = df[df.data_provider != \"radboud\"].reset_index(drop=True)\n\n    dfs = [df_r, df_k]\n    thresholds = [thresholds_rad, thresholds_ka]\n    df_keeps = list()\n    df_removes = list()\n\n    for df_tmp, thresholds_tmp in zip(dfs, thresholds):\n        df_keep_tmp, df_remove_tmp = remove_noisy2(df_tmp, thresholds_tmp)\n        df_keeps.append(df_keep_tmp)\n        df_removes.append(df_remove_tmp)\n\n    df_keep = pd.concat(df_keeps, axis=0)\n    df_removed = pd.concat(df_removes, axis=0)\n    return df_keep, df_removed\n\n# Change thresh each label each dataprovider\nthresholds_rad=[1.3, 0.8, 0.8, 0.8, 0.8, 1.3]\nthresholds_ka=[1.5, 1.0, 1.0, 1.0, 1.0, 1.5]\n\ndf_keep, df_removed = remove_noisy3(df, thresholds_rad=thresholds_rad, thresholds_ka=thresholds_ka)\nshow_keep_remove(df, df_keep, df_removed)\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1670024%2F8f930b3a7b3e30fe877ab049d0ed3b13%2F2020-07-23%2017.14.29.png?generation=1595492134245391&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 940582,
      "postDate": "2020-07-23T03:20:51.397Z",
      "content": "<p>Congratulations to everyone and thanks for the hosts for preparing this competition!</p>\n<h4>I published <a href=\"https://docs.google.com/presentation/d/1Ies4vnyVtW5U3XNDr_fom43ZJDIodu1SV6DSK8di6fs/edit?usp=sharing\" target=\"_blank\">slide</a>!</h4>\n<h4>Our code is <a href=\"https://github.com/kentaroy47/Kaggle-PANDA-1st-place-solution\" target=\"_blank\">here</a>!</h4>\n<h1>Proposed Denoising Method</h1>\n<p>We're very suprised that we finished 1st, and our simple label-denoising method (suprisingly) boosted up PB.</p>\n<p>The competition was all about handling noisy labels, so we worked hard on finding good ways to denoising.</p>\n<p>Here is our simple denoising method by <a href=\"https://www.kaggle.com/kyoshioka47\" target=\"_blank\">@kyoshioka47</a>:</p>\n<h2>Getting cleaned labels</h2>\n<ul>\n<li>train k-folds with effnet-b1 (Almost identical to Qishen's kernel)  <ul>\n<li>Model specifics in fam_taro( <a href=\"https://www.kaggle.com/yukkyo\" target=\"_blank\">@yukkyo</a> ) part</li></ul></li>\n<li>Predict hold-out sets with the trained model. We get <code>pred</code> with this step.</li>\n<li>Remove the training data which has a high disparity between ground truth and pred. The filtered labels will be called cleaned labels.</li>\n</ul>\n<p>We calculate <code>disparity</code> by the absolute difference of ISUP between GT and pred. Data with disparity larger than 1.6 was simply removed.  </p>\n<p>Here is the psuedo codes. probs_raw is the raw prediciton results (ISUP)</p>\n<pre><code># Base arutema method\ndef remove_noisy(df, thresh):\n    gap = np.abs(df[\"isup_grade\"] - df[\"probs_raw\"])\n    df_removed = df[gap &gt; thresh].reset_index(drop=True)\n    df_keep = df[gap &lt;= thresh].reset_index(drop=True)\n    return df_keep, df_removed\n\ndf_keep, df_remove = remove_noisy(df, thresh=1.6)\nshow_keep_remove(df, df_keep, df_remove)\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1670024%2F6f6bbe0dfdb2bd5ba10057a1ba32f040%2Farutema.png?generation=1595474367096439&amp;alt=media\" alt=\"\"></p>\n<h2>Retraining</h2>\n<p>Retrain model with using the denoised labels. <br>\nWe get CV 0.94 LB 0.90 PB 0.934 with a simple Qishen Eff-b0 model with k-folds.<br>\nEnsambling with different models further boosted to 1st place.</p>\n<p>We tried CleanLab too, but that did not perform well in CV/LB so we sticked with this.</p>\n<h1>1. Our final submission</h1>\n<ul>\n<li>Select 1 (public LB 0.910, private LB 0.922)<ul>\n<li>Resnext50_32x4d(poteman)</li></ul></li>\n<li>Select 2 (public LB 0.904, private LB 0.940)<ul>\n<li>Effnet-B0(arutema47) + Effnet-B1(fam_taro)<ul>\n<li>Simple average (<code>1 : 1</code>)</li></ul></li></ul></li>\n</ul>\n<p>Suprisingly, even with several weight patterns, the PB was 0.940.</p>\n<h1>2. Resnext50_32x4d( <a href=\"https://www.kaggle.com/poteman\" target=\"_blank\">@poteman</a> ), public 0.910, private 0.922</h1>\n<p>This was our best LB model.</p>\n<ul>\n<li>Split kfold: stratified kfold with imghash(threshold 0.90)</li>\n<li>iafoss tile method<ul>\n<li>tile size 256, tile num 64</li></ul></li>\n<li>model:resnext50_32x4d</li>\n<li>head: 3 * reg_head + 1 * softmax head</li>\n</ul>\n<h1>3. Effnet-B1(fam_taro), public 0.901, private 0.932</h1>\n<ul>\n<li>Split kfold<ul>\n<li>stratified 5 kfold with gleason-score and imghash similarity (threshold 0.90)<ul>\n<li>convert <code>negative</code> to <code>0+0</code></li>\n<li>how to grouping by imghash similarity<ul>\n<li>This is based on <a href=\"https://www.kaggle.com/appian\" target=\"_blank\">@appian</a> 's kernel<ul>\n<li><a href=\"https://www.kaggle.com/appian/panda-imagehash-to-detect-duplicate-images\" target=\"_blank\">https://www.kaggle.com/appian/panda-imagehash-to-detect-duplicate-images</a></li></ul></li>\n<li><a href=\"https://www.kaggle.com/yukkyo/imagehash-to-detect-duplicate-images-and-grouping\" target=\"_blank\">https://www.kaggle.com/yukkyo/imagehash-to-detect-duplicate-images-and-grouping</a></li></ul></li></ul></li>\n<li>In my opinion, split method is import point for our denoise method.<ul>\n<li>Because we use prediction of out of fold</li>\n<li>If you put the duplicate images in a different fold, I don't think denoise will work for them</li></ul></li></ul></li>\n<li>Data<ul>\n<li>iafoss tile method</li>\n<li>tile size 192, tile num 64</li></ul></li>\n<li>Model: Effnet-B1 + GeM<ul>\n<li>label: isup-grade and first score of gleason(10 dim bin)</li></ul></li>\n<li>Make final sub by 3 steps<ul>\n<li>Local train &amp; predict</li>\n<li>Remove noisy label<ul>\n<li>extended <a href=\"https://www.kaggle.com/kyoshioka47\" target=\"_blank\">@kyoshioka47</a> method</li>\n<li>Change threshold for each isup-grade and data-provider</li></ul></li>\n<li>Re-train</li></ul></li>\n<li>Not work for me<ul>\n<li>Remove noisy by confident-learning</li>\n<li>Cycle GAN augmentation(karolinska  radboud)</li>\n<li>test with AdaBN &amp; Freezing BN at train</li>\n<li>CutMix, Mixup (before denoising)</li></ul></li>\n</ul>\n<pre><code>def remove_noisy2(df, thresholds):\n    gap = np.abs(df[\"isup_grade\"] - df[\"probs_raw\"])\n\n    df_keeps = list()\n    df_removes = list()\n\n    for label, thresh in enumerate(thresholds):\n        df_tmp = df[df.isup_grade == label].reset_index(drop=True)\n        gap_tmp = gap[df.isup_grade == label].reset_index(drop=True)\n\n        df_remove_tmp = df_tmp[gap_tmp &gt; thresh].reset_index(drop=True)\n        df_keep_tmp = df_tmp[gap_tmp &lt;= thresh].reset_index(drop=True)\n\n        df_removes.append(df_remove_tmp)\n        df_keeps.append(df_keep_tmp)\n\n    df_keep = pd.concat(df_keeps, axis=0)\n    df_removed = pd.concat(df_removes, axis=0)\n    return df_keep, df_removed\n\ndef remove_noisy3(df, thresholds_rad, thresholds_ka):\n    df_r = df[df.data_provider == \"radboud\"].reset_index(drop=True)\n    df_k = df[df.data_provider != \"radboud\"].reset_index(drop=True)\n\n    dfs = [df_r, df_k]\n    thresholds = [thresholds_rad, thresholds_ka]\n    df_keeps = list()\n    df_removes = list()\n\n    for df_tmp, thresholds_tmp in zip(dfs, thresholds):\n        df_keep_tmp, df_remove_tmp = remove_noisy2(df_tmp, thresholds_tmp)\n        df_keeps.append(df_keep_tmp)\n        df_removes.append(df_remove_tmp)\n\n    df_keep = pd.concat(df_keeps, axis=0)\n    df_removed = pd.concat(df_removes, axis=0)\n    return df_keep, df_removed\n\n# Change thresh each label each dataprovider\nthresholds_rad=[1.3, 0.8, 0.8, 0.8, 0.8, 1.3]\nthresholds_ka=[1.5, 1.0, 1.0, 1.0, 1.0, 1.5]\n\ndf_keep, df_removed = remove_noisy3(df, thresholds_rad=thresholds_rad, thresholds_ka=thresholds_ka)\nshow_keep_remove(df, df_keep, df_removed)\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1670024%2F8f930b3a7b3e30fe877ab049d0ed3b13%2F2020-07-23%2017.14.29.png?generation=1595492134245391&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Congratulations to everyone and thanks for the hosts for preparing this competition!\n\n#### I published [slide](https://docs.google.com/presentation/d/1Ies4vnyVtW5U3XNDr_fom43ZJDIodu1SV6DSK8di6fs/edit?usp=sharing)!\n#### Our code is [here](https://github.com/kentaroy47/Kaggle-PANDA-1st-place-solution)!\n\n# Proposed Denoising Method\nWe're very suprised that we finished 1st, and our simple label-denoising method (suprisingly) boosted up PB.\n\nThe competition was all about handling noisy labels, so we worked hard on finding good ways to denoising.\n\nHere is our simple denoising method by @kyoshioka47:\n\n## Getting cleaned labels\n- train k-folds with effnet-b1 (Almost identical to Qishen's kernel)  \n  - Model specifics in fam_taro( @yukkyo ) part\n- Predict hold-out sets with the trained model. We get `pred` with this step.\n- Remove the training data which has a high disparity between ground truth and pred. The filtered labels will be called cleaned labels.\n\nWe calculate `disparity` by the absolute difference of ISUP between GT and pred. Data with disparity larger than 1.6 was simply removed.  \n\nHere is the psuedo codes. probs_raw is the raw prediciton results (ISUP)\n\n```python\n# Base arutema method\ndef remove_noisy(df, thresh):\n    gap = np.abs(df[\"isup_grade\"] - df[\"probs_raw\"])\n    df_removed = df[gap &gt; thresh].reset_index(drop=True)\n    df_keep = df[gap &lt;= thresh].reset_index(drop=True)\n    return df_keep, df_removed\n\ndf_keep, df_remove = remove_noisy(df, thresh=1.6)\nshow_keep_remove(df, df_keep, df_remove)\n```\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1670024%2F6f6bbe0dfdb2bd5ba10057a1ba32f040%2Farutema.png?generation=1595474367096439&amp;alt=media)\n\n\n## Retraining\nRetrain model with using the denoised labels. \nWe get CV 0.94 LB 0.90 PB 0.934 with a simple Qishen Eff-b0 model with k-folds.\nEnsambling with different models further boosted to 1st place.\n\nWe tried CleanLab too, but that did not perform well in CV/LB so we sticked with this.\n\n# 1. Our final submission\n\n- Select 1 (public LB 0.910, private LB 0.922)\n    - Resnext50_32x4d(poteman)\n- Select 2 (public LB 0.904, private LB 0.940)\n    - Effnet-B0(arutema47) + Effnet-B1(fam_taro)\n        - Simple average (`1 : 1`)\n\nSuprisingly, even with several weight patterns, the PB was 0.940.\n\n# 2. Resnext50_32x4d( @poteman ), public 0.910, private 0.922\nThis was our best LB model.\n\n- Split kfold: stratified kfold with imghash(threshold 0.90)\n- iafoss tile method\n    - tile size 256, tile num 64\n- model:resnext50_32x4d\n- head: 3 * reg_head + 1 * softmax head\n\n# 3. Effnet-B1(fam_taro), public 0.901, private 0.932\n\n- Split kfold\n    - stratified 5 kfold with gleason-score and imghash similarity (threshold 0.90)\n        - convert `negative` to `0+0`\n        - how to grouping by imghash similarity\n            - This is based on @appian 's kernel\n                - https://www.kaggle.com/appian/panda-imagehash-to-detect-duplicate-images\n            - https://www.kaggle.com/yukkyo/imagehash-to-detect-duplicate-images-and-grouping\n    - In my opinion, split method is import point for our denoise method.\n        - Because we use prediction of out of fold\n        - If you put the duplicate images in a different fold, I don't think denoise will work for them\n- Data\n  - iafoss tile method\n    - tile size 192, tile num 64\n- Model: Effnet-B1 + GeM\n    - label: isup-grade and first score of gleason(10 dim bin)\n- Make final sub by 3 steps\n    - Local train &amp; predict\n    - Remove noisy label\n        - extended @kyoshioka47 method\n        - Change threshold for each isup-grade and data-provider\n    - Re-train\n- Not work for me\n    - Remove noisy by confident-learning\n    - Cycle GAN augmentation(karolinska &lt;-&gt; radboud)\n    - test with AdaBN &amp; Freezing BN at train\n    - CutMix, Mixup (before denoising)\n  \n```python\ndef remove_noisy2(df, thresholds):\n    gap = np.abs(df[\"isup_grade\"] - df[\"probs_raw\"])\n    \n    df_keeps = list()\n    df_removes = list()\n    \n    for label, thresh in enumerate(thresholds):\n        df_tmp = df[df.isup_grade == label].reset_index(drop=True)\n        gap_tmp = gap[df.isup_grade == label].reset_index(drop=True)\n        \n        df_remove_tmp = df_tmp[gap_tmp &gt; thresh].reset_index(drop=True)\n        df_keep_tmp = df_tmp[gap_tmp &lt;= thresh].reset_index(drop=True)\n        \n        df_removes.append(df_remove_tmp)\n        df_keeps.append(df_keep_tmp)\n    \n    df_keep = pd.concat(df_keeps, axis=0)\n    df_removed = pd.concat(df_removes, axis=0)\n    return df_keep, df_removed\n\ndef remove_noisy3(df, thresholds_rad, thresholds_ka):\n    df_r = df[df.data_provider == \"radboud\"].reset_index(drop=True)\n    df_k = df[df.data_provider != \"radboud\"].reset_index(drop=True)\n    \n    dfs = [df_r, df_k]\n    thresholds = [thresholds_rad, thresholds_ka]\n    df_keeps = list()\n    df_removes = list()\n    \n    for df_tmp, thresholds_tmp in zip(dfs, thresholds):\n        df_keep_tmp, df_remove_tmp = remove_noisy2(df_tmp, thresholds_tmp)\n        df_keeps.append(df_keep_tmp)\n        df_removes.append(df_remove_tmp)\n    \n    df_keep = pd.concat(df_keeps, axis=0)\n    df_removed = pd.concat(df_removes, axis=0)\n    return df_keep, df_removed\n\n# Change thresh each label each dataprovider\nthresholds_rad=[1.3, 0.8, 0.8, 0.8, 0.8, 1.3]\nthresholds_ka=[1.5, 1.0, 1.0, 1.0, 1.0, 1.5]\n\ndf_keep, df_removed = remove_noisy3(df, thresholds_rad=thresholds_rad, thresholds_ka=thresholds_ka)\nshow_keep_remove(df, df_keep, df_removed)\n```\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1670024%2F8f930b3a7b3e30fe877ab049d0ed3b13%2F2020-07-23%2017.14.29.png?generation=1595492134245391&amp;alt=media)",
      "votes": 225
    },
    {
      "id": 941544,
      "postDate": "2020-07-23T09:36:35.487Z",
      "content": "<p>Our single model <a href=\"/kyoshioka47\">@kyoshioka47</a> can get 3th place.👍 </p>",
      "rawMarkdown": "Our single model @kyoshioka47 can get 3th place.👍 ",
      "votes": 6
    },
    {
      "id": 951178,
      "postDate": "2020-07-30T01:04:38.143Z",
      "content": "<p>After doing some analysis on late subs, I have a feeling that mixup augmentation was effective for generalization. I think this is a famous technique and most top-teams have implemented this too.</p>\n\n<p>For mixup, I simply made an another tile and mixed two images.\nI just implemented mixup from the original paper, but some implementation can be found from Bengali contests.</p>\n\n<p>Also tried cutmix, does not work well. Maybe because tile includes important information partially.</p>\n\n<p>mixup: Beyond Empirical Risk Minimization\n<a href=\"https://arxiv.org/abs/1710.09412\">https://arxiv.org/abs/1710.09412</a></p>\n\n<h2>viz</h2>\n\n<p>You get weird images like this\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3508221%2F4434018a8c402c1a39b687e088a18cff%2F2020-07-30%20100643.png?generation=1596071226907240&amp;alt=media\" alt=\"\"></p>\n\n<h2>some codes</h2>\n\n<p>```python\nrd = np.random.rand()\nif mixup and rd &lt; 0.3:\n    mix_idx = np.random.random_integers(0, len(self.df))\n    row2 = self.df.iloc[mix_idx]\n    img_id2 = row2.image_id\n    file = os.path.join(\"train_{}_{}\".format(self.image_size, self.n_tiles), f'{img_id2}.npz')\n    images2 = np.load(file)[\"arr_0\"]\n    images2 = images2.transpose(2, 0, 1)</p>\n\n<pre><code>if self.transform is not None:\n    images2 = self.transform(image=images2)['image']\n\n# blend image\ngamma = np.random.beta(1,1)\nimages = ((images*gamma + images2*(1-gamma))).astype(np.uint8)\n# blend labels\nlabel2 = np.zeros(5).astype(np.float32)\nlabel2[:row2.isup_grade] = 1.\nlabel = (label*gamma+label2*(1-gamma))       \n</code></pre>\n\n<p>```</p>",
      "rawMarkdown": "After doing some analysis on late subs, I have a feeling that mixup augmentation was effective for generalization. I think this is a famous technique and most top-teams have implemented this too.\n\nFor mixup, I simply made an another tile and mixed two images.\nI just implemented mixup from the original paper, but some implementation can be found from Bengali contests.\n\nAlso tried cutmix, does not work well. Maybe because tile includes important information partially.\n\nmixup: Beyond Empirical Risk Minimization\nhttps://arxiv.org/abs/1710.09412\n\n## viz\nYou get weird images like this\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3508221%2F4434018a8c402c1a39b687e088a18cff%2F2020-07-30%20100643.png?generation=1596071226907240&amp;alt=media)\n\n\n## some codes\n```python\nrd = np.random.rand()\nif mixup and rd &lt; 0.3:\n    mix_idx = np.random.random_integers(0, len(self.df))\n    row2 = self.df.iloc[mix_idx]\n    img_id2 = row2.image_id\n    file = os.path.join(\"train_{}_{}\".format(self.image_size, self.n_tiles), f'{img_id2}.npz')\n    images2 = np.load(file)[\"arr_0\"]\n    images2 = images2.transpose(2, 0, 1)\n\n    if self.transform is not None:\n        images2 = self.transform(image=images2)['image']\n\n    # blend image\n    gamma = np.random.beta(1,1)\n    images = ((images*gamma + images2*(1-gamma))).astype(np.uint8)\n    # blend labels\n    label2 = np.zeros(5).astype(np.float32)\n    label2[:row2.isup_grade] = 1.\n    label = (label*gamma+label2*(1-gamma))       \n```",
      "votes": 4,
      "replies": [
        {
          "id": 1351327,
          "postDate": "2021-06-16T08:14:59.670Z",
          "content": "<p>thanks for sharing!!</p>\n<p>One question is, is there any other way to use mixup to predict scale variables other than bin processing as in the example above?</p>",
          "rawMarkdown": "thanks for sharing!!\n\nOne question is, is there any other way to use mixup to predict scale variables other than bin processing as in the example above?"
        }
      ]
    },
    {
      "id": 940614,
      "postDate": "2020-07-23T03:50:50.700Z",
      "content": "<p>Very nice congrats.\n I did exactly that in last few days although not systematically as you . Got 93 with this strategy ,few more iterations could have got more ,sad I din't believe that solution of self as it scored poor at public </p>",
      "rawMarkdown": "Very nice congrats.\n I did exactly that in last few days although not systematically as you . Got 93 with this strategy ,few more iterations could have got more ,sad I din't believe that solution of self as it scored poor at public ",
      "votes": 4,
      "replies": [
        {
          "id": 940652,
          "postDate": "2020-07-23T04:21:48.243Z",
          "content": "<p>Actually it was <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/165774\">your discussion</a> which lead us to this idea :) big thanks</p>\n\n<p>Currently looking at our submissions, this solution overfits to the PB, still need to investigate if this works well as a machine learning solution for prostitute cancer.</p>",
          "rawMarkdown": "Actually it was [your discussion](https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/165774) which lead us to this idea :) big thanks\n\nCurrently looking at our submissions, this solution overfits to the PB, still need to investigate if this works well as a machine learning solution for prostitute cancer.",
          "votes": 1
        },
        {
          "id": 940665,
          "postDate": "2020-07-23T04:35:36.023Z",
          "content": "<p>Thanks ...it helped youn team where is my share of prize :) kidding . I spent lot in figuring it out after I failed to move up in ladder .</p>",
          "rawMarkdown": "Thanks ...it helped youn team where is my share of prize :) kidding . I spent lot in figuring it out after I failed to move up in ladder .",
          "votes": 1
        },
        {
          "id": 940804,
          "postDate": "2020-07-23T05:43:44.813Z",
          "content": "<p>Truly, lots of share goes to you,and  <a href=\"/iafoss\">@iafoss</a> <a href=\"/haqishen\">@haqishen</a> who built the basis of this competition.</p>",
          "rawMarkdown": "Truly, lots of share goes to you,and  @iafoss @haqishen who built the basis of this competition.",
          "votes": 2
        },
        {
          "id": 940813,
          "postDate": "2020-07-23T05:52:09.043Z",
          "content": "<p><a href=\"/kyoshioka47\">@kyoshioka47</a>  Thanks for acknowledging m happy,if not i some one else got best out of it :) \ni was looking for team up in GWD ,are you in.. we can work it out fast we are ranked 202 currently ?</p>",
          "rawMarkdown": "@kyoshioka47  Thanks for acknowledging m happy,if not i some one else got best out of it :) \ni was looking for team up in GWD ,are you in.. we can work it out fast we are ranked 202 currently ?"
        }
      ]
    },
    {
      "id": 957359,
      "postDate": "2020-08-04T08:40:00.647Z",
      "content": "<p>Congratulation, very good. Thaks  for the information</p>",
      "rawMarkdown": "Congratulation, very good. Thaks  for the information",
      "votes": 1
    },
    {
      "id": 944369,
      "postDate": "2020-07-25T04:07:32.830Z",
      "content": "<p>Wow, a very novel solution and a great writeup, thanks for sharing. Congratulations! </p>",
      "rawMarkdown": "Wow, a very novel solution and a great writeup, thanks for sharing. Congratulations! ",
      "votes": 1
    },
    {
      "id": 944138,
      "postDate": "2020-07-24T21:30:23.017Z",
      "content": "<p>Lots of things to remember for upcoming competitions! How did you come up with the noise removal technique? 🤔 \nAnd finally, congratulations on your win!</p>",
      "rawMarkdown": "Lots of things to remember for upcoming competitions! How did you come up with the noise removal technique? 🤔 \nAnd finally, congratulations on your win!",
      "votes": 1,
      "replies": [
        {
          "id": 944321,
          "postDate": "2020-07-25T02:43:52.687Z",
          "content": "<p><a href=\"/yassinealouini\">@yassinealouini</a> thx!\nThis idea came to me by @kentaroy47.\nHowever, it's simple, but it came out of many experiments.</p>\n\n<p>The point of this idea is that my model may be more accurate than the Original Label.</p>\n\n<p>Specifically, prior to this idea I found that updating all of Radboud's labels to my out of fold predictions and re-training them would raise the public LB.\nThis led me to think that the accuracy of the model might be better than Original Label.</p>\n\n<p>However, this denoising method has many problems (ex. breaking LocalCV, discarding to hard examples). Be careful if you use this method.</p>",
          "rawMarkdown": "@yassinealouini thx!\nThis idea came to me by @kentaroy47.\nHowever, it's simple, but it came out of many experiments.\n\nThe point of this idea is that my model may be more accurate than the Original Label.\n\nSpecifically, prior to this idea I found that updating all of Radboud's labels to my out of fold predictions and re-training them would raise the public LB.\nThis led me to think that the accuracy of the model might be better than Original Label.\n\nHowever, this denoising method has many problems (ex. breaking LocalCV, discarding to hard examples). Be careful if you use this method.",
          "votes": 7
        },
        {
          "id": 944494,
          "postDate": "2020-07-25T06:46:05.697Z",
          "content": "<p>Alright, thanks for the tips!</p>",
          "rawMarkdown": "Alright, thanks for the tips!",
          "votes": 1
        },
        {
          "id": 959151,
          "postDate": "2020-08-05T11:41:48.633Z",
          "content": "<p><a href=\"/yukkyo\">@yukkyo</a> \nI am just curious, if you say it comes from \"many experiments\", how much time is behind those experiments? hours, days, weeks?\nThanks 🙊 </p>",
          "rawMarkdown": "@yukkyo \nI am just curious, if you say it comes from \"many experiments\", how much time is behind those experiments? hours, days, weeks?\nThanks 🙊 "
        }
      ]
    },
    {
      "id": 942758,
      "postDate": "2020-07-24T01:34:55.583Z",
      "content": "<p>Congratulations! I love this write up and I love your your techniques. Boss move / well done.</p>",
      "rawMarkdown": "Congratulations! I love this write up and I love your your techniques. Boss move / well done.",
      "votes": 1
    },
    {
      "id": 942276,
      "postDate": "2020-07-23T17:12:23.050Z",
      "content": "<p>Thanks <a href=\"/yukkyo\">@yukkyo</a> for your outstanding work and share with us</p>",
      "rawMarkdown": "Thanks @yukkyo for your outstanding work and share with us",
      "votes": 1
    },
    {
      "id": 941313,
      "postDate": "2020-07-23T06:59:08.413Z",
      "content": "<p>Congrats</p>",
      "rawMarkdown": "Congrats",
      "votes": 1
    },
    {
      "id": 941245,
      "postDate": "2020-07-23T06:34:34.330Z",
      "content": "<p>Congrats your team!!!\nDoes <code>probs_raw</code> mean regression value?\nFor example, it is possible that the model outputs a large negative value for a sample whose label is 0. Did you do processing such as clipping?</p>",
      "rawMarkdown": "Congrats your team!!!\nDoes `probs_raw` mean regression value?\nFor example, it is possible that the model outputs a large negative value for a sample whose label is 0. Did you do processing such as clipping?",
      "votes": 1,
      "replies": [
        {
          "id": 941281,
          "postDate": "2020-07-23T06:46:04.730Z",
          "content": "<p><a href=\"/tattaka\">@tattaka</a> thx !\nI convert each label value to bin (ex. 2 -&gt; <code>[1, 1, 0, 0, 0]</code>) and using sigmoid for predicting.\nSo there are no negative value on my case.</p>",
          "rawMarkdown": "@tattaka thx !\nI convert each label value to bin (ex. 2 -&gt; `[1, 1, 0, 0, 0]`) and using sigmoid for predicting.\nSo there are no negative value on my case."
        },
        {
          "id": 941300,
          "postDate": "2020-07-23T06:55:06.330Z",
          "content": "<p>Thanks.</p>",
          "rawMarkdown": "Thanks.\n",
          "votes": 1
        },
        {
          "id": 964923,
          "postDate": "2020-08-10T09:13:57.523Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/yukkyo\" target=\"_blank\">@yukkyo</a>!! <br>\nWhat is the usual motivation use bin label?</p>",
          "rawMarkdown": "Hi @yukkyo!! \nWhat is the usual motivation use bin label?"
        }
      ]
    },
    {
      "id": 941231,
      "postDate": "2020-07-23T06:20:57.540Z",
      "content": "<p>Congrats! The method to get clean label looks fancy. I wonder score before/after applying this method.</p>",
      "rawMarkdown": "Congrats! The method to get clean label looks fancy. I wonder score before/after applying this method.",
      "votes": 1,
      "replies": [
        {
          "id": 941264,
          "postDate": "2020-07-23T06:38:53.417Z",
          "content": "<p><a href=\"/songwonho\">@songwonho</a> \nthx !\nOn my single model, before and after removing the noise, the following is what it looks like</p>\n\n<ul>\n<li>before: Public: 0.892, Private: 0.916</li>\n<li>after: Public: 0.901, Private: 0.932</li>\n</ul>",
          "rawMarkdown": "@songwonho \nthx !\nOn my single model, before and after removing the noise, the following is what it looks like\n\n- before: Public: 0.892, Private: 0.916\n- after: Public: 0.901, Private: 0.932",
          "votes": 5
        },
        {
          "id": 941436,
          "postDate": "2020-07-23T08:08:25.850Z",
          "content": "<p>Interesting ! Thanks.</p>",
          "rawMarkdown": "Interesting ! Thanks.",
          "votes": 1
        }
      ]
    },
    {
      "id": 941230,
      "postDate": "2020-07-23T06:20:25.130Z",
      "content": "<p>Congrats ! Gratuluje ! Marek</p>",
      "rawMarkdown": "Congrats ! Gratuluje ! Marek",
      "votes": 1
    },
    {
      "id": 941056,
      "postDate": "2020-07-23T06:13:45.167Z",
      "content": "<p>Congrats ! Thanks for sharing your method to handle noisy labels, very interesting. </p>",
      "rawMarkdown": "Congrats ! Thanks for sharing your method to handle noisy labels, very interesting. ",
      "votes": 1
    },
    {
      "id": 1037502,
      "postDate": "2020-10-05T04:43:54.317Z",
      "content": "<p>We will share our codes!<br>\n<a href=\"https://github.com/kentaroy47/Kaggle-PANDA-1st-place-solution\" target=\"_blank\">https://github.com/kentaroy47/Kaggle-PANDA-1st-place-solution</a></p>",
      "rawMarkdown": "We will share our codes!\nhttps://github.com/kentaroy47/Kaggle-PANDA-1st-place-solution",
      "votes": 2
    },
    {
      "id": 955237,
      "postDate": "2020-08-02T13:30:01.177Z",
      "content": "<p>Thanks for sharing. It was quite insighful.</p>",
      "rawMarkdown": "Thanks for sharing. It was quite insighful.",
      "votes": 2
    },
    {
      "id": 942942,
      "postDate": "2020-07-24T04:31:25.780Z",
      "content": "<p>Thanks for sharing your method for handling noisy labels</p>",
      "rawMarkdown": "Thanks for sharing your method for handling noisy labels",
      "votes": 2
    },
    {
      "id": 941287,
      "postDate": "2020-07-23T06:50:11.867Z",
      "content": "<p>Congratulations, very interesting solution! I also tried cleaning labels online, but CV was unstable and I dropped this idea. Never got to try it the way you did. And our team also tried various stain normalizations, not GAN however. Also failed.</p>",
      "rawMarkdown": "Congratulations, very interesting solution! I also tried cleaning labels online, but CV was unstable and I dropped this idea. Never got to try it the way you did. And our team also tried various stain normalizations, not GAN however. Also failed.",
      "votes": 2,
      "replies": [
        {
          "id": 941307,
          "postDate": "2020-07-23T06:58:04.133Z",
          "content": "<p><a href=\"/cateek\">@cateek</a> thx and congrats 2nd place !\nVery interesting.\nIndeed, my CV was unstable in my case as well.</p>\n\n<p>I also considered stain normalization, but I didn't adopt it because I was afraid of test run time.</p>",
          "rawMarkdown": "@cateek thx and congrats 2nd place !\nVery interesting.\nIndeed, my CV was unstable in my case as well.\n\nI also considered stain normalization, but I didn't adopt it because I was afraid of test run time.",
          "votes": 1
        },
        {
          "id": 946914,
          "postDate": "2020-07-27T00:00:29.637Z",
          "content": "<p><a href=\"/cateek\">@cateek</a> \nI'll add a few things that I remembered.</p>\n\n<p>As you pointed out, the CV was unstable and I was looking at Public LB to make some adjustments. Also at that time I continued to use a light model (EfficientNet-B1) to avoid overfitting to <code>noise that remained in local after denoising</code> and  <code>Public LB</code>.</p>",
          "rawMarkdown": "@cateek \nI'll add a few things that I remembered.\n\nAs you pointed out, the CV was unstable and I was looking at Public LB to make some adjustments. Also at that time I continued to use a light model (EfficientNet-B1) to avoid overfitting to `noise that remained in local after denoising` and  `Public LB`."
        }
      ]
    },
    {
      "id": 941038,
      "postDate": "2020-07-23T06:06:26.013Z",
      "content": "<p>Wow, congrats <a href=\"/kyoshioka47\">@kyoshioka47</a>, <a href=\"/yukkyo\">@yukkyo</a>, <a href=\"/poteman\">@poteman</a>.</p>\n\n<p>I did not participate in this competition. However, it is a pleasure to see the familiar guys I saw at other competitions.</p>\n\n<p>I look forward to your other competitions. 👍 </p>",
      "rawMarkdown": "Wow, congrats @kyoshioka47, @yukkyo, @poteman.\n\nI did not participate in this competition. However, it is a pleasure to see the familiar guys I saw at other competitions.\n\nI look forward to your other competitions. 👍 ",
      "votes": 2,
      "replies": [
        {
          "id": 941047,
          "postDate": "2020-07-23T06:10:36.330Z",
          "content": "<p><a href=\"/piantic\">@piantic</a> thx ! I'll see you at another competition soon!</p>",
          "rawMarkdown": "@piantic thx ! I'll see you at another competition soon!"
        }
      ]
    },
    {
      "id": 954512,
      "postDate": "2020-08-01T19:37:06.050Z",
      "content": "<p>congratulations!!</p>",
      "rawMarkdown": "congratulations!!"
    },
    {
      "id": 954453,
      "postDate": "2020-08-01T18:31:04.340Z",
      "content": "<p>Well done</p>",
      "rawMarkdown": "Well done"
    },
    {
      "id": 954217,
      "postDate": "2020-08-01T14:01:38.237Z",
      "content": "<p>Congratulations</p>",
      "rawMarkdown": "Congratulations"
    },
    {
      "id": 953975,
      "postDate": "2020-08-01T09:31:46.620Z",
      "content": "<p>Great Solution</p>",
      "rawMarkdown": "Great Solution"
    },
    {
      "id": 953706,
      "postDate": "2020-08-01T03:12:39.977Z",
      "content": "<p>Congratulations and thank you. I learned a lot!</p>",
      "rawMarkdown": "Congratulations and thank you. I learned a lot!"
    },
    {
      "id": 953694,
      "postDate": "2020-08-01T02:44:13.053Z",
      "content": "<p>Your work really helped me to learn and understand. Thank you.</p>",
      "rawMarkdown": "Your work really helped me to learn and understand. Thank you."
    },
    {
      "id": 952857,
      "postDate": "2020-07-31T09:15:55.437Z",
      "content": "<p>Congratulation...!!</p>",
      "rawMarkdown": "Congratulation...!!"
    },
    {
      "id": 952828,
      "postDate": "2020-07-31T08:55:34.337Z",
      "content": "<p>Congratulations</p>",
      "rawMarkdown": "Congratulations"
    },
    {
      "id": 952684,
      "postDate": "2020-07-31T06:30:08.077Z",
      "content": "<p>congratulations!!</p>",
      "rawMarkdown": "congratulations!!"
    },
    {
      "id": 951636,
      "postDate": "2020-07-30T09:29:02.370Z",
      "content": "<p>Congrats. </p>",
      "rawMarkdown": "Congrats. "
    },
    {
      "id": 951348,
      "postDate": "2020-07-30T04:45:27.670Z",
      "content": "<p>Congrats and really nice write up!</p>",
      "rawMarkdown": "Congrats and really nice write up!"
    },
    {
      "id": 951177,
      "postDate": "2020-07-30T01:04:29.617Z",
      "content": "<p>Congrats on the win! Beautiful solution!</p>",
      "rawMarkdown": "Congrats on the win! Beautiful solution!"
    },
    {
      "id": 950814,
      "postDate": "2020-07-29T16:34:13.587Z",
      "content": "<p>Congrats</p>",
      "rawMarkdown": "Congrats"
    },
    {
      "id": 950376,
      "postDate": "2020-07-29T11:13:19.093Z",
      "content": "<p>Congratulations and thank you. I learned a lot!</p>",
      "rawMarkdown": "Congratulations and thank you. I learned a lot!"
    },
    {
      "id": 949951,
      "postDate": "2020-07-29T05:03:49.207Z",
      "content": "<p>tHX FOR SHARING THE IDEAS</p>",
      "rawMarkdown": "tHX FOR SHARING THE IDEAS"
    },
    {
      "id": 949831,
      "postDate": "2020-07-29T01:38:29.840Z",
      "content": "<p>Awesome!</p>",
      "rawMarkdown": "Awesome!"
    },
    {
      "id": 949660,
      "postDate": "2020-07-28T19:26:01.597Z",
      "content": "<p>Congratz</p>",
      "rawMarkdown": "Congratz"
    },
    {
      "id": 949530,
      "postDate": "2020-07-28T17:33:07.337Z",
      "content": "<p>Great!</p>",
      "rawMarkdown": "Great!"
    },
    {
      "id": 949449,
      "postDate": "2020-07-28T16:33:01.133Z",
      "content": "<p>great</p>",
      "rawMarkdown": "great"
    },
    {
      "id": 949383,
      "postDate": "2020-07-28T15:54:00.023Z",
      "content": "<p>Great work congrats</p>",
      "rawMarkdown": "Great work congrats"
    },
    {
      "id": 949369,
      "postDate": "2020-07-28T15:38:11.320Z",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!"
    },
    {
      "id": 949160,
      "postDate": "2020-07-28T13:01:36.517Z",
      "content": "<p>Congratulations! </p>",
      "rawMarkdown": "Congratulations! "
    },
    {
      "id": 948828,
      "postDate": "2020-07-28T08:48:26.393Z",
      "content": "<p>Nice</p>",
      "rawMarkdown": "Nice"
    },
    {
      "id": 948745,
      "postDate": "2020-07-28T07:29:20.490Z",
      "content": "<p>Thank you for sharing and congratulations!</p>",
      "rawMarkdown": "Thank you for sharing and congratulations!"
    },
    {
      "id": 948554,
      "postDate": "2020-07-28T04:09:00.630Z",
      "content": "<p>Great Solution, Congratulations!!!</p>",
      "rawMarkdown": "Great Solution, Congratulations!!!"
    },
    {
      "id": 948516,
      "postDate": "2020-07-28T03:08:41.060Z",
      "content": "<p>Excellent novel solution and great visualisation. Congratulations </p>",
      "rawMarkdown": "Excellent novel solution and great visualisation. Congratulations "
    },
    {
      "id": 948496,
      "postDate": "2020-07-28T02:38:19.107Z",
      "content": "<p>very novel solution。 thanks for sharing.</p>",
      "rawMarkdown": "very novel solution。 thanks for sharing."
    },
    {
      "id": 948405,
      "postDate": "2020-07-27T22:47:12.343Z",
      "content": "<p>Congrats. Good job!</p>",
      "rawMarkdown": "Congrats. Good job!"
    },
    {
      "id": 947948,
      "postDate": "2020-07-27T15:09:57.240Z",
      "content": "<p>Congrats</p>",
      "rawMarkdown": "Congrats"
    },
    {
      "id": 947902,
      "postDate": "2020-07-27T14:45:19.357Z",
      "content": "<p>Thanks for sharing this technique. Worth reading. Never seen this anywhere else. Keep innovating 👍 </p>",
      "rawMarkdown": "Thanks for sharing this technique. Worth reading. Never seen this anywhere else. Keep innovating 👍 "
    },
    {
      "id": 947820,
      "postDate": "2020-07-27T13:54:05.277Z",
      "content": "<p>keep going</p>",
      "rawMarkdown": "keep going\n"
    },
    {
      "id": 947393,
      "postDate": "2020-07-27T08:38:50.990Z",
      "content": "<p>Congratulations!!</p>",
      "rawMarkdown": "Congratulations!!"
    },
    {
      "id": 947172,
      "postDate": "2020-07-27T05:41:53.087Z",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!"
    },
    {
      "id": 947077,
      "postDate": "2020-07-27T04:17:29.067Z",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!"
    },
    {
      "id": 946992,
      "postDate": "2020-07-27T02:00:18.707Z",
      "content": "<p>Congrats, nice work!</p>",
      "rawMarkdown": "Congrats, nice work!"
    },
    {
      "id": 946987,
      "postDate": "2020-07-27T01:59:20.017Z",
      "content": "<p>Very cool! Great teamwork!</p>",
      "rawMarkdown": "Very cool! Great teamwork!"
    },
    {
      "id": 946737,
      "postDate": "2020-07-26T19:23:15.440Z",
      "content": "<p>Hi and congrats for your novel solution. I think that label removing has really improved your model, and this idea really surprised me. Thank you for the insight!</p>\n\n<p>May I ask you with which machine you trained your model?</p>",
      "rawMarkdown": "Hi and congrats for your novel solution. I think that label removing has really improved your model, and this idea really surprised me. Thank you for the insight!\n\nMay I ask you with which machine you trained your model?",
      "replies": [
        {
          "id": 946876,
          "postDate": "2020-07-26T23:13:59.497Z",
          "content": "<p><a href=\"/mawanda\">@mawanda</a> thx!\nI only used my own machine (with TitanRTX x 2).\nEach training requires 1 GPU.</p>\n\n<p>However, I don't know about my other teammates' machines.</p>",
          "rawMarkdown": "@mawanda thx!\nI only used my own machine (with TitanRTX x 2).\nEach training requires 1 GPU.\n\nHowever, I don't know about my other teammates' machines."
        },
        {
          "id": 946926,
          "postDate": "2020-07-27T00:28:01.917Z",
          "content": "<p>I used 2080ti to train models. \nSpeed up training by saving tile files to png images prior.\nThe speed was about 15 min/epoch (effb0)</p>",
          "rawMarkdown": "I used 2080ti to train models. \nSpeed up training by saving tile files to png images prior.\nThe speed was about 15 min/epoch (effb0)"
        },
        {
          "id": 949292,
          "postDate": "2020-07-28T14:34:19.563Z",
          "content": "<p>Sure thing. I admit that my machine is nothing compared (2070S). Thank you for the information. I was only able to train a ResNet18 \"siamese\" model, with which I obtained a discrete 0.8966. Wondering about what I could have obtained with more resources.</p>\n\n<p>Congratulations again to your work and your team! You rocked. :)</p>",
          "rawMarkdown": "Sure thing. I admit that my machine is nothing compared (2070S). Thank you for the information. I was only able to train a ResNet18 \"siamese\" model, with which I obtained a discrete 0.8966. Wondering about what I could have obtained with more resources.\n\nCongratulations again to your work and your team! You rocked. :)"
        }
      ]
    },
    {
      "id": 946702,
      "postDate": "2020-07-26T18:56:15.387Z",
      "content": "<p>Congrats!😃</p>",
      "rawMarkdown": "Congrats!😃"
    },
    {
      "id": 946677,
      "postDate": "2020-07-26T18:12:41.230Z",
      "content": "<p>Congrats!</p>",
      "rawMarkdown": "Congrats!"
    },
    {
      "id": 946227,
      "postDate": "2020-07-26T12:55:03.047Z",
      "content": "<p>Congratulations..and Thanks for sharing...</p>",
      "rawMarkdown": "Congratulations..and Thanks for sharing..."
    },
    {
      "id": 946016,
      "postDate": "2020-07-26T09:31:34.613Z",
      "content": "<p>Congrats</p>",
      "rawMarkdown": "Congrats"
    },
    {
      "id": 945453,
      "postDate": "2020-07-25T20:16:49.233Z",
      "content": "<p>Congrats on the win, and thanks for sharing the method</p>",
      "rawMarkdown": "Congrats on the win, and thanks for sharing the method"
    },
    {
      "id": 945313,
      "postDate": "2020-07-25T18:21:08.100Z",
      "content": "<p>Congratulations! Keep up the good work :)</p>",
      "rawMarkdown": "Congratulations! Keep up the good work :)"
    },
    {
      "id": 945310,
      "postDate": "2020-07-25T18:18:29.613Z",
      "content": "<p>elegant!</p>",
      "rawMarkdown": "elegant!"
    },
    {
      "id": 944531,
      "postDate": "2020-07-25T07:11:05.440Z",
      "content": "<p>Niceee</p>",
      "rawMarkdown": "Niceee"
    },
    {
      "id": 942096,
      "postDate": "2020-07-23T15:22:57.337Z",
      "content": "<p><a href=\"/yukkyo\">@yukkyo</a> Congratulations!!! What is the idea of ​​freezing BN during training?</p>",
      "rawMarkdown": "@yukkyo Congratulations!!! What is the idea of ​​freezing BN during training?",
      "replies": [
        {
          "id": 942236,
          "postDate": "2020-07-23T16:47:27.007Z",
          "content": "<p>Freezing bn is to use the pre-train parameters without updating the BN parameters.\nIf you use PyTorch, you can implement Freeze BN by replacing <code>model.train()</code> to <code>model.eval()</code>.\n(However, be careful when using things like <code>nn.Dropout()</code>)</p>",
          "rawMarkdown": "Freezing bn is to use the pre-train parameters without updating the BN parameters.\nIf you use PyTorch, you can implement Freeze BN by replacing `model.train()` to `model.eval()`.\n(However, be careful when using things like `nn.Dropout()`)"
        },
        {
          "id": 942329,
          "postDate": "2020-07-23T17:44:49.213Z",
          "content": "<p>thanks, I got it.</p>",
          "rawMarkdown": "thanks, I got it.",
          "votes": 1
        }
      ]
    },
    {
      "id": 941809,
      "postDate": "2020-07-23T12:28:53.350Z",
      "rawMarkdown": ""
    },
    {
      "id": 941234,
      "postDate": "2020-07-23T06:22:00.870Z",
      "content": "<p>First of all congratulations to winning this competition.\nI tried the same approach of denoising the labels: Train effnet-b0 very similar to Qishen's kernel. Then remove data if the difference between predicted label and the ground truth label is &gt;= 2. This gave me 0.921 on private leaderboard. In comparison using a very similar effnet-b0 with the original labels (no label denoising)  gave private leaderboard of 0.929. Therefore, I would argue that this kind of label denoising is not necessarily a key ingredient for this challenge.</p>",
      "rawMarkdown": "First of all congratulations to winning this competition.\nI tried the same approach of denoising the labels: Train effnet-b0 very similar to Qishen's kernel. Then remove data if the difference between predicted label and the ground truth label is &gt;= 2. This gave me 0.921 on private leaderboard. In comparison using a very similar effnet-b0 with the original labels (no label denoising)  gave private leaderboard of 0.929. Therefore, I would argue that this kind of label denoising is not necessarily a key ingredient for this challenge.",
      "replies": [
        {
          "id": 941275,
          "postDate": "2020-07-23T06:43:33.790Z",
          "content": "<p><a href=\"/jakobw\">@jakobw</a> thx! This point is very useful.\nWhen using this method, I think it's important that how to split Train/Valid.\nHow did you split it?</p>",
          "rawMarkdown": "@jakobw thx! This point is very useful.\nWhen using this method, I think it's important that how to split Train/Valid.\nHow did you split it?",
          "votes": 1
        },
        {
          "id": 941418,
          "postDate": "2020-07-23T08:03:02.147Z",
          "content": "<p>You're point is interesting. \nseems like <a href=\"/iafoss\">@iafoss</a> utilizes <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169205\">similar ideas</a> to get gold, I think base model selection and data split was important too.</p>",
          "rawMarkdown": "You're point is interesting. \nseems like @iafoss utilizes [similar ideas](https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169205) to get gold, I think base model selection and data split was important too."
        },
        {
          "id": 942529,
          "postDate": "2020-07-23T19:46:18.650Z",
          "content": "<p>I started using this trick nearly almost from the beginning of the competition, so I don't have quite good statistics on subs without it to make a solid comparison, but when I checked my subs from 2 month ago I see that my private LB has increased on average from ~0.915 to ~0.925-0.930 (though I didn't check it for different seeds etc. at that time). \nThere is always some variation of LB (and it probably may explain your inconsistency), like one of my models 2 month ago got 0.938 using quite similar pipeline to my final sub, but on average I say it gets ~0.930 (and 0.938 is just private LB noise), because if I look into other subs using similar approach but different seed, or other minor adjustments, I see that it can get as low as 0.924. If you are lucky u select one that gives you the best score, if you are not lucky, you can drop on LB, and if you want to be safe u ensemble many models and get less affected by the noise bringing you up or down.</p>",
          "rawMarkdown": "I started using this trick nearly almost from the beginning of the competition, so I don't have quite good statistics on subs without it to make a solid comparison, but when I checked my subs from 2 month ago I see that my private LB has increased on average from ~0.915 to ~0.925-0.930 (though I didn't check it for different seeds etc. at that time). \nThere is always some variation of LB (and it probably may explain your inconsistency), like one of my models 2 month ago got 0.938 using quite similar pipeline to my final sub, but on average I say it gets ~0.930 (and 0.938 is just private LB noise), because if I look into other subs using similar approach but different seed, or other minor adjustments, I see that it can get as low as 0.924. If you are lucky u select one that gives you the best score, if you are not lucky, you can drop on LB, and if you want to be safe u ensemble many models and get less affected by the noise bringing you up or down.",
          "votes": 3
        },
        {
          "id": 944647,
          "postDate": "2020-07-25T08:21:18.390Z",
          "content": "<p><a href=\"/yukkyo\">@yukkyo</a> I used a stratified 5-fold split.</p>",
          "rawMarkdown": "@yukkyo I used a stratified 5-fold split."
        },
        {
          "id": 944682,
          "postDate": "2020-07-25T08:47:19.957Z",
          "content": "<p><a href=\"/jakobw\">@jakobw</a>\nIf you apply just a stratfied kfold, the duplicate image would go into a different fold.\nIn that case I don't think this denoising method is very good for the duplicate images.\nAnd I think that affects the Score.</p>\n\n<p>We used imghash to put the duplicate images in the same fold.\nHow do you handle duplicate images?</p>",
          "rawMarkdown": "@jakobw\nIf you apply just a stratfied kfold, the duplicate image would go into a different fold.\nIn that case I don't think this denoising method is very good for the duplicate images.\nAnd I think that affects the Score.\n\nWe used imghash to put the duplicate images in the same fold.\nHow do you handle duplicate images?",
          "votes": 1
        },
        {
          "id": 946693,
          "postDate": "2020-07-26T18:39:10.137Z",
          "content": "<p><a href=\"/yukkyo\">@yukkyo</a> , How many duplicates did u find?</p>",
          "rawMarkdown": "@yukkyo , How many duplicates did u find?"
        },
        {
          "id": 946884,
          "postDate": "2020-07-26T23:25:28.403Z",
          "content": "<p><a href=\"/iafoss\">@iafoss</a> \nI treated 2,121 images as <code>duplicates</code> (script is below).\n<a href=\"https://www.kaggle.com/yukkyo/imagehash-to-detect-duplicate-images-and-grouping\">https://www.kaggle.com/yukkyo/imagehash-to-detect-duplicate-images-and-grouping</a></p>\n\n<p>Of course, this includes False Positive (not true duplicate image) and False Negative (true duplicate image, but I've missed it).\nAnd we can change this rate by imghash threshold.</p>\n\n<p>If you look at example of above kernel, you can see that there are many False Positive and this threshold value(0.9) seems a bit low. But I chose this value because I wanted to avoid putting the same image in different folds any more than that.</p>",
          "rawMarkdown": "@iafoss \nI treated 2,121 images as `duplicates` (script is below).\nhttps://www.kaggle.com/yukkyo/imagehash-to-detect-duplicate-images-and-grouping\n\nOf course, this includes False Positive (not true duplicate image) and False Negative (true duplicate image, but I've missed it).\nAnd we can change this rate by imghash threshold.\n\nIf you look at example of above kernel, you can see that there are many False Positive and this threshold value(0.9) seems a bit low. But I chose this value because I wanted to avoid putting the same image in different folds any more than that.",
          "votes": 1
        }
      ]
    },
    {
      "id": 951366,
      "postDate": "2020-07-30T05:12:25.373Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 950532,
      "postDate": "2020-07-29T13:02:48.643Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 947988,
      "postDate": "2020-07-27T15:38:20.187Z",
      "content": "<p>Congrats</p>",
      "rawMarkdown": "Congrats",
      "isDeleted": true
    },
    {
      "id": 944170,
      "postDate": "2020-07-24T22:19:55.033Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 942190,
      "postDate": "2020-07-23T16:23:07.013Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 942900,
      "postDate": "2020-07-24T04:03:17.743Z",
      "content": "<p>Congrats and thanks!</p>",
      "rawMarkdown": "Congrats and thanks!",
      "votes": 1
    },
    {
      "id": 955719,
      "postDate": "2020-08-02T21:10:03.687Z",
      "content": "<p>Thanks.</p>",
      "rawMarkdown": "Thanks."
    },
    {
      "id": 955365,
      "postDate": "2020-08-02T15:21:38.250Z",
      "content": "<p>Thanks for tips</p>",
      "rawMarkdown": "Thanks for tips"
    },
    {
      "id": 955185,
      "postDate": "2020-08-02T12:25:03.110Z",
      "content": "<p>Congrats and thanks for sharing.</p>",
      "rawMarkdown": "Congrats and thanks for sharing."
    },
    {
      "id": 955167,
      "postDate": "2020-08-02T12:11:31.137Z",
      "content": "<p>Awesome read, thanks for sharing</p>",
      "rawMarkdown": "Awesome read, thanks for sharing"
    },
    {
      "id": 954106,
      "postDate": "2020-08-01T11:41:02.597Z",
      "content": "<p>Thanks for sharing.</p>",
      "rawMarkdown": "Thanks for sharing."
    },
    {
      "id": 952772,
      "postDate": "2020-07-31T07:59:39.883Z",
      "content": "<p>Thanks for sharing your notebook!</p>",
      "rawMarkdown": "Thanks for sharing your notebook!"
    },
    {
      "id": 951566,
      "postDate": "2020-07-30T08:09:38.543Z",
      "content": "<p>Thanks for sharing the method.</p>",
      "rawMarkdown": "Thanks for sharing the method."
    },
    {
      "id": 950699,
      "postDate": "2020-07-29T14:57:05.450Z",
      "content": "<p>thanks for sharing.</p>",
      "rawMarkdown": "thanks for sharing."
    },
    {
      "id": 947619,
      "postDate": "2020-07-27T11:31:54.290Z",
      "content": "<p>Congrats and thanks!</p>",
      "rawMarkdown": "Congrats and thanks!"
    },
    {
      "id": 946814,
      "postDate": "2020-07-26T21:29:20.623Z",
      "content": "<p>Amazing ! thanks for the trick :)</p>",
      "rawMarkdown": "Amazing ! thanks for the trick :)"
    },
    {
      "id": 946540,
      "postDate": "2020-07-26T16:30:03.540Z",
      "content": "<p><em>Congrats and thanks!</em></p>",
      "rawMarkdown": "*Congrats and thanks!*"
    },
    {
      "id": 946301,
      "postDate": "2020-07-26T13:52:20.867Z",
      "content": "<p>Nice! Thanks for sharing!</p>",
      "rawMarkdown": "Nice! Thanks for sharing!"
    },
    {
      "id": 946092,
      "postDate": "2020-07-26T11:00:17.277Z",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing"
    },
    {
      "id": 945674,
      "postDate": "2020-07-26T04:38:43.110Z",
      "content": "<p>thanks!</p>",
      "rawMarkdown": "thanks!"
    },
    {
      "id": 945307,
      "postDate": "2020-07-25T18:16:29.463Z",
      "content": "<p>Thanks for the share. Kudos!</p>",
      "rawMarkdown": "Thanks for the share. Kudos!"
    },
    {
      "id": 944792,
      "postDate": "2020-07-25T10:52:52.613Z",
      "content": "<p>Thanks for sharing. Congratulations!!!</p>",
      "rawMarkdown": "Thanks for sharing. Congratulations!!!"
    }
  ],
  "comments": [
    {
      "id": 941544,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-23T09:36:35.487000",
      "content": "<p>Our single model <a href=\"/kyoshioka47\">@kyoshioka47</a> can get 3th place.👍 </p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 951178,
      "author_name": "arutema47",
      "author_url": "",
      "post_date": "2020-07-30T01:04:38.143000",
      "content": "<p>After doing some analysis on late subs, I have a feeling that mixup augmentation was effective for generalization. I think this is a famous technique and most top-teams have implemented this too.</p>\n\n<p>For mixup, I simply made an another tile and mixed two images.\nI just implemented mixup from the original paper, but some implementation can be found from Bengali contests.</p>\n\n<p>Also tried cutmix, does not work well. Maybe because tile includes important information partially.</p>\n\n<p>mixup: Beyond Empirical Risk Minimization\n<a href=\"https://arxiv.org/abs/1710.09412\">https://arxiv.org/abs/1710.09412</a></p>\n\n<h2>viz</h2>\n\n<p>You get weird images like this\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3508221%2F4434018a8c402c1a39b687e088a18cff%2F2020-07-30%20100643.png?generation=1596071226907240&amp;alt=media\" alt=\"\"></p>\n\n<h2>some codes</h2>\n\n<p>```python\nrd = np.random.rand()\nif mixup and rd &lt; 0.3:\n    mix_idx = np.random.random_integers(0, len(self.df))\n    row2 = self.df.iloc[mix_idx]\n    img_id2 = row2.image_id\n    file = os.path.join(\"train_{}_{}\".format(self.image_size, self.n_tiles), f'{img_id2}.npz')\n    images2 = np.load(file)[\"arr_0\"]\n    images2 = images2.transpose(2, 0, 1)</p>\n\n<pre><code>if self.transform is not None:\n    images2 = self.transform(image=images2)['image']\n\n# blend image\ngamma = np.random.beta(1,1)\nimages = ((images*gamma + images2*(1-gamma))).astype(np.uint8)\n# blend labels\nlabel2 = np.zeros(5).astype(np.float32)\nlabel2[:row2.isup_grade] = 1.\nlabel = (label*gamma+label2*(1-gamma))       \n</code></pre>\n\n<p>```</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1351327,
          "author_name": "patriot",
          "author_url": "",
          "post_date": "2021-06-16T08:14:59.670000",
          "content": "<p>thanks for sharing!!</p>\n<p>One question is, is there any other way to use mixup to predict scale variables other than bin processing as in the example above?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 940614,
      "author_name": "Jaideep",
      "author_url": "",
      "post_date": "2020-07-23T03:50:50.700000",
      "content": "<p>Very nice congrats.\n I did exactly that in last few days although not systematically as you . Got 93 with this strategy ,few more iterations could have got more ,sad I din't believe that solution of self as it scored poor at public </p>",
      "votes": 4,
      "replies": [
        {
          "id": 940652,
          "author_name": "arutema47",
          "author_url": "",
          "post_date": "2020-07-23T04:21:48.243000",
          "content": "<p>Actually it was <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/165774\">your discussion</a> which lead us to this idea :) big thanks</p>\n\n<p>Currently looking at our submissions, this solution overfits to the PB, still need to investigate if this works well as a machine learning solution for prostitute cancer.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 940665,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-07-23T04:35:36.023000",
          "content": "<p>Thanks ...it helped youn team where is my share of prize :) kidding . I spent lot in figuring it out after I failed to move up in ladder .</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 940804,
          "author_name": "arutema47",
          "author_url": "",
          "post_date": "2020-07-23T05:43:44.813000",
          "content": "<p>Truly, lots of share goes to you,and  <a href=\"/iafoss\">@iafoss</a> <a href=\"/haqishen\">@haqishen</a> who built the basis of this competition.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 940813,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-07-23T05:52:09.043000",
          "content": "<p><a href=\"/kyoshioka47\">@kyoshioka47</a>  Thanks for acknowledging m happy,if not i some one else got best out of it :) \ni was looking for team up in GWD ,are you in.. we can work it out fast we are ranked 202 currently ?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 957359,
      "author_name": "DhruvGargg02",
      "author_url": "",
      "post_date": "2020-08-04T08:40:00.647000",
      "content": "<p>Congratulation, very good. Thaks  for the information</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 944369,
      "author_name": "Stanley Zheng",
      "author_url": "",
      "post_date": "2020-07-25T04:07:32.830000",
      "content": "<p>Wow, a very novel solution and a great writeup, thanks for sharing. Congratulations! </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 944138,
      "author_name": "Yassine Alouini",
      "author_url": "",
      "post_date": "2020-07-24T21:30:23.017000",
      "content": "<p>Lots of things to remember for upcoming competitions! How did you come up with the noise removal technique? 🤔 \nAnd finally, congratulations on your win!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 944321,
          "author_name": "fam_taro",
          "author_url": "",
          "post_date": "2020-07-25T02:43:52.687000",
          "content": "<p><a href=\"/yassinealouini\">@yassinealouini</a> thx!\nThis idea came to me by @kentaroy47.\nHowever, it's simple, but it came out of many experiments.</p>\n\n<p>The point of this idea is that my model may be more accurate than the Original Label.</p>\n\n<p>Specifically, prior to this idea I found that updating all of Radboud's labels to my out of fold predictions and re-training them would raise the public LB.\nThis led me to think that the accuracy of the model might be better than Original Label.</p>\n\n<p>However, this denoising method has many problems (ex. breaking LocalCV, discarding to hard examples). Be careful if you use this method.</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 944494,
          "author_name": "Yassine Alouini",
          "author_url": "",
          "post_date": "2020-07-25T06:46:05.697000",
          "content": "<p>Alright, thanks for the tips!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 959151,
          "author_name": "Michal Brezak",
          "author_url": "",
          "post_date": "2020-08-05T11:41:48.633000",
          "content": "<p><a href=\"/yukkyo\">@yukkyo</a> \nI am just curious, if you say it comes from \"many experiments\", how much time is behind those experiments? hours, days, weeks?\nThanks 🙊 </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 942758,
      "author_name": "عثمان",
      "author_url": "",
      "post_date": "2020-07-24T01:34:55.583000",
      "content": "<p>Congratulations! I love this write up and I love your your techniques. Boss move / well done.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 942276,
      "author_name": "Mahmud Hasan",
      "author_url": "",
      "post_date": "2020-07-23T17:12:23.050000",
      "content": "<p>Thanks <a href=\"/yukkyo\">@yukkyo</a> for your outstanding work and share with us</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 941313,
      "author_name": "Walia",
      "author_url": "",
      "post_date": "2020-07-23T06:59:08.413000",
      "content": "<p>Congrats</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 941245,
      "author_name": "tattaka",
      "author_url": "",
      "post_date": "2020-07-23T06:34:34.330000",
      "content": "<p>Congrats your team!!!\nDoes <code>probs_raw</code> mean regression value?\nFor example, it is possible that the model outputs a large negative value for a sample whose label is 0. Did you do processing such as clipping?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 941281,
          "author_name": "fam_taro",
          "author_url": "",
          "post_date": "2020-07-23T06:46:04.730000",
          "content": "<p><a href=\"/tattaka\">@tattaka</a> thx !\nI convert each label value to bin (ex. 2 -&gt; <code>[1, 1, 0, 0, 0]</code>) and using sigmoid for predicting.\nSo there are no negative value on my case.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 941300,
          "author_name": "tattaka",
          "author_url": "",
          "post_date": "2020-07-23T06:55:06.330000",
          "content": "<p>Thanks.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 964923,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-10T09:13:57.523000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/yukkyo\" target=\"_blank\">@yukkyo</a>!! <br>\nWhat is the usual motivation use bin label?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 941231,
      "author_name": "Wonho Song",
      "author_url": "",
      "post_date": "2020-07-23T06:20:57.540000",
      "content": "<p>Congrats! The method to get clean label looks fancy. I wonder score before/after applying this method.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 941264,
          "author_name": "fam_taro",
          "author_url": "",
          "post_date": "2020-07-23T06:38:53.417000",
          "content": "<p><a href=\"/songwonho\">@songwonho</a> \nthx !\nOn my single model, before and after removing the noise, the following is what it looks like</p>\n\n<ul>\n<li>before: Public: 0.892, Private: 0.916</li>\n<li>after: Public: 0.901, Private: 0.932</li>\n</ul>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 941436,
          "author_name": "Wonho Song",
          "author_url": "",
          "post_date": "2020-07-23T08:08:25.850000",
          "content": "<p>Interesting ! Thanks.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 941230,
      "author_name": "Marek Nurzynski",
      "author_url": "",
      "post_date": "2020-07-23T06:20:25.130000",
      "content": "<p>Congrats ! Gratuluje ! Marek</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 941056,
      "author_name": "Alex",
      "author_url": "",
      "post_date": "2020-07-23T06:13:45.167000",
      "content": "<p>Congrats ! Thanks for sharing your method to handle noisy labels, very interesting. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1037502,
      "author_name": "arutema47",
      "author_url": "",
      "post_date": "2020-10-05T04:43:54.317000",
      "content": "<p>We will share our codes!<br>\n<a href=\"https://github.com/kentaroy47/Kaggle-PANDA-1st-place-solution\" target=\"_blank\">https://github.com/kentaroy47/Kaggle-PANDA-1st-place-solution</a></p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 955237,
      "author_name": "Amritvir Singh",
      "author_url": "",
      "post_date": "2020-08-02T13:30:01.177000",
      "content": "<p>Thanks for sharing. It was quite insighful.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 942942,
      "author_name": "Vinay Phadnis",
      "author_url": "",
      "post_date": "2020-07-24T04:31:25.780000",
      "content": "<p>Thanks for sharing your method for handling noisy labels</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 941287,
      "author_name": "Eek The Cat",
      "author_url": "",
      "post_date": "2020-07-23T06:50:11.867000",
      "content": "<p>Congratulations, very interesting solution! I also tried cleaning labels online, but CV was unstable and I dropped this idea. Never got to try it the way you did. And our team also tried various stain normalizations, not GAN however. Also failed.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 941307,
          "author_name": "fam_taro",
          "author_url": "",
          "post_date": "2020-07-23T06:58:04.133000",
          "content": "<p><a href=\"/cateek\">@cateek</a> thx and congrats 2nd place !\nVery interesting.\nIndeed, my CV was unstable in my case as well.</p>\n\n<p>I also considered stain normalization, but I didn't adopt it because I was afraid of test run time.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 946914,
          "author_name": "fam_taro",
          "author_url": "",
          "post_date": "2020-07-27T00:00:29.637000",
          "content": "<p><a href=\"/cateek\">@cateek</a> \nI'll add a few things that I remembered.</p>\n\n<p>As you pointed out, the CV was unstable and I was looking at Public LB to make some adjustments. Also at that time I continued to use a light model (EfficientNet-B1) to avoid overfitting to <code>noise that remained in local after denoising</code> and  <code>Public LB</code>.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 941038,
      "author_name": "Heroseo",
      "author_url": "",
      "post_date": "2020-07-23T06:06:26.013000",
      "content": "<p>Wow, congrats <a href=\"/kyoshioka47\">@kyoshioka47</a>, <a href=\"/yukkyo\">@yukkyo</a>, <a href=\"/poteman\">@poteman</a>.</p>\n\n<p>I did not participate in this competition. However, it is a pleasure to see the familiar guys I saw at other competitions.</p>\n\n<p>I look forward to your other competitions. 👍 </p>",
      "votes": 2,
      "replies": [
        {
          "id": 941047,
          "author_name": "fam_taro",
          "author_url": "",
          "post_date": "2020-07-23T06:10:36.330000",
          "content": "<p><a href=\"/piantic\">@piantic</a> thx ! I'll see you at another competition soon!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 954512,
      "author_name": "Erik Larson",
      "author_url": "",
      "post_date": "2020-08-01T19:37:06.050000",
      "content": "<p>congratulations!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 954453,
      "author_name": "Sami.Alashabi",
      "author_url": "",
      "post_date": "2020-08-01T18:31:04.340000",
      "content": "<p>Well done</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 954217,
      "author_name": "Rafael Pompeu",
      "author_url": "",
      "post_date": "2020-08-01T14:01:38.237000",
      "content": "<p>Congratulations</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 953975,
      "author_name": "Saksham Agrawal",
      "author_url": "",
      "post_date": "2020-08-01T09:31:46.620000",
      "content": "<p>Great Solution</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 953706,
      "author_name": "Alex Rain",
      "author_url": "",
      "post_date": "2020-08-01T03:12:39.977000",
      "content": "<p>Congratulations and thank you. I learned a lot!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 953694,
      "author_name": "Adhyan Maji",
      "author_url": "",
      "post_date": "2020-08-01T02:44:13.053000",
      "content": "<p>Your work really helped me to learn and understand. Thank you.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 952857,
      "author_name": "LeeSangHun_JackJack",
      "author_url": "",
      "post_date": "2020-07-31T09:15:55.437000",
      "content": "<p>Congratulation...!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 952828,
      "author_name": "pranav0198",
      "author_url": "",
      "post_date": "2020-07-31T08:55:34.337000",
      "content": "<p>Congratulations</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 952684,
      "author_name": "zahed khan pathan",
      "author_url": "",
      "post_date": "2020-07-31T06:30:08.077000",
      "content": "<p>congratulations!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 951636,
      "author_name": "Ishant",
      "author_url": "",
      "post_date": "2020-07-30T09:29:02.370000",
      "content": "<p>Congrats. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 951348,
      "author_name": "James Sharwin",
      "author_url": "",
      "post_date": "2020-07-30T04:45:27.670000",
      "content": "<p>Congrats and really nice write up!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 951177,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-30T01:04:29.617000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 950814,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-29T16:34:13.587000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 950376,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-29T11:13:19.093000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 949951,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-29T05:03:49.207000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 949831,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-29T01:38:29.840000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 949660,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-28T19:26:01.597000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 949530,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-28T17:33:07.337000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 949449,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-28T16:33:01.133000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 949383,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-28T15:54:00.023000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 949369,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-28T15:38:11.320000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 949160,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-28T13:01:36.517000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 948828,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-28T08:48:26.393000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 948745,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-28T07:29:20.490000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 948554,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-28T04:09:00.630000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 948516,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-28T03:08:41.060000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 948496,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-28T02:38:19.107000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 948405,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-27T22:47:12.343000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 947948,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-27T15:09:57.240000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 947902,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-27T14:45:19.357000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 947820,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-27T13:54:05.277000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 947393,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-27T08:38:50.990000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 947172,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-27T05:41:53.087000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 947077,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-27T04:17:29.067000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 946992,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-27T02:00:18.707000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 946987,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-27T01:59:20.017000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 946737,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-26T19:23:15.440000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 946876,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-26T23:13:59.497000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 946926,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-27T00:28:01.917000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 949292,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-28T14:34:19.563000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 946702,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-26T18:56:15.387000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 946677,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-26T18:12:41.230000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 946227,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-26T12:55:03.047000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 946016,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-26T09:31:34.613000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 945453,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-25T20:16:49.233000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 945313,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-25T18:21:08.100000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 945310,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-25T18:18:29.613000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 944531,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-25T07:11:05.440000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 942096,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-23T15:22:57.337000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 942236,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-23T16:47:27.007000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 942329,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-23T17:44:49.213000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 941809,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-23T12:28:53.350000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 941234,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-23T06:22:00.870000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 941275,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-23T06:43:33.790000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 941418,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-23T08:03:02.147000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 942529,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-23T19:46:18.650000",
          "content": "",
          "votes": 3,
          "replies": []
        },
        {
          "id": 944647,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-25T08:21:18.390000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 944682,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-25T08:47:19.957000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 946693,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-26T18:39:10.137000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 946884,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-26T23:25:28.403000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 951366,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-30T05:12:25.373000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 950532,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-29T13:02:48.643000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 947988,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-27T15:38:20.187000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 944170,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-24T22:19:55.033000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 942190,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-23T16:23:07.013000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 942900,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-24T04:03:17.743000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 955719,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-02T21:10:03.687000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 955365,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-02T15:21:38.250000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 955185,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-02T12:25:03.110000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 955167,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-02T12:11:31.137000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 954106,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-01T11:41:02.597000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 952772,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-31T07:59:39.883000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 951566,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-30T08:09:38.543000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 950699,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-29T14:57:05.450000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 947619,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-27T11:31:54.290000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 946814,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-26T21:29:20.623000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 946540,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-26T16:30:03.540000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 946301,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-26T13:52:20.867000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 946092,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-26T11:00:17.277000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 945674,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-26T04:38:43.110000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 945307,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-25T18:16:29.463000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 944792,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-25T10:52:52.613000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "940582": "Congratulations to everyone and thanks for the hosts for preparing this competition!\n\n#### I published [slide](https://docs.google.com/presentation/d/1Ies4vnyVtW5U3XNDr_fom43ZJDIodu1SV6DSK8di6fs/edit?usp=sharing)!\n#### Our code is [here](https://github.com/kentaroy47/Kaggle-PANDA-1st-place-solution)!\n\n# Proposed Denoising Method\nWe're very suprised that we finished 1st, and our simple label-denoising method (suprisingly) boosted up PB.\n\nThe competition was all about handling noisy labels, so we worked hard on finding good ways to denoising.\n\nHere is our simple denoising method by @kyoshioka47:\n\n## Getting cleaned labels\n- train k-folds with effnet-b1 (Almost identical to Qishen's kernel)  \n  - Model specifics in fam_taro( @yukkyo ) part\n- Predict hold-out sets with the trained model. We get `pred` with this step.\n- Remove the training data which has a high disparity between ground truth and pred. The filtered labels will be called cleaned labels.\n\nWe calculate `disparity` by the absolute difference of ISUP between GT and pred. Data with disparity larger than 1.6 was simply removed.  \n\nHere is the psuedo codes. probs_raw is the raw prediciton results (ISUP)\n\n```python\n# Base arutema method\ndef remove_noisy(df, thresh):\n    gap = np.abs(df[\"isup_grade\"] - df[\"probs_raw\"])\n    df_removed = df[gap &gt; thresh].reset_index(drop=True)\n    df_keep = df[gap &lt;= thresh].reset_index(drop=True)\n    return df_keep, df_removed\n\ndf_keep, df_remove = remove_noisy(df, thresh=1.6)\nshow_keep_remove(df, df_keep, df_remove)\n```\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1670024%2F6f6bbe0dfdb2bd5ba10057a1ba32f040%2Farutema.png?generation=1595474367096439&amp;alt=media)\n\n\n## Retraining\nRetrain model with using the denoised labels. \nWe get CV 0.94 LB 0.90 PB 0.934 with a simple Qishen Eff-b0 model with k-folds.\nEnsambling with different models further boosted to 1st place.\n\nWe tried CleanLab too, but that did not perform well in CV/LB so we sticked with this.\n\n# 1. Our final submission\n\n- Select 1 (public LB 0.910, private LB 0.922)\n    - Resnext50_32x4d(poteman)\n- Select 2 (public LB 0.904, private LB 0.940)\n    - Effnet-B0(arutema47) + Effnet-B1(fam_taro)\n        - Simple average (`1 : 1`)\n\nSuprisingly, even with several weight patterns, the PB was 0.940.\n\n# 2. Resnext50_32x4d( @poteman ), public 0.910, private 0.922\nThis was our best LB model.\n\n- Split kfold: stratified kfold with imghash(threshold 0.90)\n- iafoss tile method\n    - tile size 256, tile num 64\n- model:resnext50_32x4d\n- head: 3 * reg_head + 1 * softmax head\n\n# 3. Effnet-B1(fam_taro), public 0.901, private 0.932\n\n- Split kfold\n    - stratified 5 kfold with gleason-score and imghash similarity (threshold 0.90)\n        - convert `negative` to `0+0`\n        - how to grouping by imghash similarity\n            - This is based on @appian 's kernel\n                - https://www.kaggle.com/appian/panda-imagehash-to-detect-duplicate-images\n            - https://www.kaggle.com/yukkyo/imagehash-to-detect-duplicate-images-and-grouping\n    - In my opinion, split method is import point for our denoise method.\n        - Because we use prediction of out of fold\n        - If you put the duplicate images in a different fold, I don't think denoise will work for them\n- Data\n  - iafoss tile method\n    - tile size 192, tile num 64\n- Model: Effnet-B1 + GeM\n    - label: isup-grade and first score of gleason(10 dim bin)\n- Make final sub by 3 steps\n    - Local train &amp; predict\n    - Remove noisy label\n        - extended @kyoshioka47 method\n        - Change threshold for each isup-grade and data-provider\n    - Re-train\n- Not work for me\n    - Remove noisy by confident-learning\n    - Cycle GAN augmentation(karolinska &lt;-&gt; radboud)\n    - test with AdaBN &amp; Freezing BN at train\n    - CutMix, Mixup (before denoising)\n  \n```python\ndef remove_noisy2(df, thresholds):\n    gap = np.abs(df[\"isup_grade\"] - df[\"probs_raw\"])\n    \n    df_keeps = list()\n    df_removes = list()\n    \n    for label, thresh in enumerate(thresholds):\n        df_tmp = df[df.isup_grade == label].reset_index(drop=True)\n        gap_tmp = gap[df.isup_grade == label].reset_index(drop=True)\n        \n        df_remove_tmp = df_tmp[gap_tmp &gt; thresh].reset_index(drop=True)\n        df_keep_tmp = df_tmp[gap_tmp &lt;= thresh].reset_index(drop=True)\n        \n        df_removes.append(df_remove_tmp)\n        df_keeps.append(df_keep_tmp)\n    \n    df_keep = pd.concat(df_keeps, axis=0)\n    df_removed = pd.concat(df_removes, axis=0)\n    return df_keep, df_removed\n\ndef remove_noisy3(df, thresholds_rad, thresholds_ka):\n    df_r = df[df.data_provider == \"radboud\"].reset_index(drop=True)\n    df_k = df[df.data_provider != \"radboud\"].reset_index(drop=True)\n    \n    dfs = [df_r, df_k]\n    thresholds = [thresholds_rad, thresholds_ka]\n    df_keeps = list()\n    df_removes = list()\n    \n    for df_tmp, thresholds_tmp in zip(dfs, thresholds):\n        df_keep_tmp, df_remove_tmp = remove_noisy2(df_tmp, thresholds_tmp)\n        df_keeps.append(df_keep_tmp)\n        df_removes.append(df_remove_tmp)\n    \n    df_keep = pd.concat(df_keeps, axis=0)\n    df_removed = pd.concat(df_removes, axis=0)\n    return df_keep, df_removed\n\n# Change thresh each label each dataprovider\nthresholds_rad=[1.3, 0.8, 0.8, 0.8, 0.8, 1.3]\nthresholds_ka=[1.5, 1.0, 1.0, 1.0, 1.0, 1.5]\n\ndf_keep, df_removed = remove_noisy3(df, thresholds_rad=thresholds_rad, thresholds_ka=thresholds_ka)\nshow_keep_remove(df, df_keep, df_removed)\n```\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1670024%2F8f930b3a7b3e30fe877ab049d0ed3b13%2F2020-07-23%2017.14.29.png?generation=1595492134245391&amp;alt=media)",
    "941544": "Our single model @kyoshioka47 can get 3th place.👍 ",
    "951178": "After doing some analysis on late subs, I have a feeling that mixup augmentation was effective for generalization. I think this is a famous technique and most top-teams have implemented this too.\n\nFor mixup, I simply made an another tile and mixed two images.\nI just implemented mixup from the original paper, but some implementation can be found from Bengali contests.\n\nAlso tried cutmix, does not work well. Maybe because tile includes important information partially.\n\nmixup: Beyond Empirical Risk Minimization\nhttps://arxiv.org/abs/1710.09412\n\n## viz\nYou get weird images like this\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3508221%2F4434018a8c402c1a39b687e088a18cff%2F2020-07-30%20100643.png?generation=1596071226907240&amp;alt=media)\n\n\n## some codes\n```python\nrd = np.random.rand()\nif mixup and rd &lt; 0.3:\n    mix_idx = np.random.random_integers(0, len(self.df))\n    row2 = self.df.iloc[mix_idx]\n    img_id2 = row2.image_id\n    file = os.path.join(\"train_{}_{}\".format(self.image_size, self.n_tiles), f'{img_id2}.npz')\n    images2 = np.load(file)[\"arr_0\"]\n    images2 = images2.transpose(2, 0, 1)\n\n    if self.transform is not None:\n        images2 = self.transform(image=images2)['image']\n\n    # blend image\n    gamma = np.random.beta(1,1)\n    images = ((images*gamma + images2*(1-gamma))).astype(np.uint8)\n    # blend labels\n    label2 = np.zeros(5).astype(np.float32)\n    label2[:row2.isup_grade] = 1.\n    label = (label*gamma+label2*(1-gamma))       \n```",
    "940614": "Very nice congrats.\n I did exactly that in last few days although not systematically as you . Got 93 with this strategy ,few more iterations could have got more ,sad I din't believe that solution of self as it scored poor at public ",
    "957359": "Congratulation, very good. Thaks  for the information",
    "944369": "Wow, a very novel solution and a great writeup, thanks for sharing. Congratulations! ",
    "944138": "Lots of things to remember for upcoming competitions! How did you come up with the noise removal technique? 🤔 \nAnd finally, congratulations on your win!",
    "942758": "Congratulations! I love this write up and I love your your techniques. Boss move / well done.",
    "942276": "Thanks @yukkyo for your outstanding work and share with us",
    "941313": "Congrats",
    "941245": "Congrats your team!!!\nDoes `probs_raw` mean regression value?\nFor example, it is possible that the model outputs a large negative value for a sample whose label is 0. Did you do processing such as clipping?",
    "941231": "Congrats! The method to get clean label looks fancy. I wonder score before/after applying this method.",
    "941230": "Congrats ! Gratuluje ! Marek",
    "941056": "Congrats ! Thanks for sharing your method to handle noisy labels, very interesting. ",
    "1037502": "We will share our codes!\nhttps://github.com/kentaroy47/Kaggle-PANDA-1st-place-solution",
    "955237": "Thanks for sharing. It was quite insighful.",
    "942942": "Thanks for sharing your method for handling noisy labels",
    "941287": "Congratulations, very interesting solution! I also tried cleaning labels online, but CV was unstable and I dropped this idea. Never got to try it the way you did. And our team also tried various stain normalizations, not GAN however. Also failed.",
    "941038": "Wow, congrats @kyoshioka47, @yukkyo, @poteman.\n\nI did not participate in this competition. However, it is a pleasure to see the familiar guys I saw at other competitions.\n\nI look forward to your other competitions. 👍 ",
    "954512": "congratulations!!",
    "954453": "Well done",
    "954217": "Congratulations",
    "953975": "Great Solution",
    "953706": "Congratulations and thank you. I learned a lot!",
    "953694": "Your work really helped me to learn and understand. Thank you.",
    "952857": "Congratulation...!!",
    "952828": "Congratulations",
    "952684": "congratulations!!",
    "951636": "Congrats. ",
    "951348": "Congrats and really nice write up!",
    "951177": "Congrats on the win! Beautiful solution!",
    "950814": "Congrats",
    "950376": "Congratulations and thank you. I learned a lot!",
    "949951": "tHX FOR SHARING THE IDEAS",
    "949831": "Awesome!",
    "949660": "Congratz",
    "949530": "Great!",
    "949449": "great",
    "949383": "Great work congrats",
    "949369": "Congratulations!",
    "949160": "Congratulations! ",
    "948828": "Nice",
    "948745": "Thank you for sharing and congratulations!",
    "948554": "Great Solution, Congratulations!!!",
    "948516": "Excellent novel solution and great visualisation. Congratulations ",
    "948496": "very novel solution。 thanks for sharing.",
    "948405": "Congrats. Good job!",
    "947948": "Congrats",
    "947902": "Thanks for sharing this technique. Worth reading. Never seen this anywhere else. Keep innovating 👍 ",
    "947820": "keep going\n",
    "947393": "Congratulations!!",
    "947172": "Congratulations!",
    "947077": "Congratulations!",
    "946992": "Congrats, nice work!",
    "946987": "Very cool! Great teamwork!",
    "946737": "Hi and congrats for your novel solution. I think that label removing has really improved your model, and this idea really surprised me. Thank you for the insight!\n\nMay I ask you with which machine you trained your model?",
    "946702": "Congrats!😃",
    "946677": "Congrats!",
    "946227": "Congratulations..and Thanks for sharing...",
    "946016": "Congrats",
    "945453": "Congrats on the win, and thanks for sharing the method",
    "945313": "Congratulations! Keep up the good work :)",
    "945310": "elegant!",
    "944531": "Niceee",
    "942096": "@yukkyo Congratulations!!! What is the idea of ​​freezing BN during training?",
    "941809": "",
    "941234": "First of all congratulations to winning this competition.\nI tried the same approach of denoising the labels: Train effnet-b0 very similar to Qishen's kernel. Then remove data if the difference between predicted label and the ground truth label is &gt;= 2. This gave me 0.921 on private leaderboard. In comparison using a very similar effnet-b0 with the original labels (no label denoising)  gave private leaderboard of 0.929. Therefore, I would argue that this kind of label denoising is not necessarily a key ingredient for this challenge.",
    "951366": "",
    "950532": "",
    "947988": "Congrats",
    "944170": "",
    "942190": "",
    "942900": "Congrats and thanks!",
    "955719": "Thanks.",
    "955365": "Thanks for tips",
    "955185": "Congrats and thanks for sharing.",
    "955167": "Awesome read, thanks for sharing",
    "954106": "Thanks for sharing.",
    "952772": "Thanks for sharing your notebook!",
    "951566": "Thanks for sharing the method.",
    "950699": "thanks for sharing.",
    "947619": "Congrats and thanks!",
    "946814": "Amazing ! thanks for the trick :)",
    "946540": "*Congrats and thanks!*",
    "946301": "Nice! Thanks for sharing!",
    "946092": "Thanks for sharing",
    "945674": "thanks!",
    "945307": "Thanks for the share. Kudos!",
    "944792": "Thanks for sharing. Congratulations!!!"
  }
}