{
  "id": 358029,
  "title": "5th Place Solution",
  "url": "/competitions/mayo-clinic-strip-ai/discussion/358029",
  "author_name": "Theo Viel",
  "post_date": "2022-10-06T11:26:13.159000",
  "votes": 33,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Thanks everyone for the competition, congratz to the winners &amp; kudos to the learners. </p>\n<p>Although this one is a controversial one because most people couldn't even beat the sample submission, I still believe it was an alright competition. The key is to understand that there is not a lot of signal, and that the competition metric is not friendly with overconfident models, more on that later.</p>\n<p>Because of the nature of the competition (not a lot of signal, and basically no public LB), I decided not too invest too much time on it, so I am very happy with the result !</p>\n<h3>Data</h3>\n<ul>\n<li>Color normalization : <ul>\n<li>Detect the background color</li>\n<li>Normalize all the image by  <code>𝑟 = 𝑏𝑎𝑐𝑘𝑔𝑟𝑜𝑢𝑛𝑑_𝑐𝑜𝑙𝑜𝑟/(255,   255,  255)</code></li></ul></li>\n<li>Simple Approach : <ul>\n<li>Remove all white chunks from the image </li>\n<li>Resize to 1024x1024</li></ul></li>\n<li>Advanced Approach (cf figure below): <ul>\n<li>Detect duplicated areas in the image</li>\n<li>Keep only one of them</li>\n<li>Resize to have the longest edge to 1024, preserving aspect ratio</li></ul></li>\n</ul>\n<p><a href=\"https://ibb.co/fpnzYN1\"><img src=\"https://i.ibb.co/HzxfDPF/mayo-data-pipe.png\" alt=\"mayo-data-pipe\"></a></p>\n<p>You can see some examples <a href=\"https://www.kaggle.com/code/theoviel/inference-mayo\" target=\"_blank\">here</a>.</p>\n<h3>Models</h3>\n<ul>\n<li>3 Small Models  -  <strong>CV AUC 0.684</strong><ol>\n<li>Resnet10t – AUC 0.661</li>\n<li>EfficientNet-b0 – AUC 0.671</li>\n<li>EfficientNet-b0 using simple approach images - AUC 0.662</li></ol></li>\n<li>Training<ul>\n<li>Image size : 1024x1024</li>\n<li>Ranger + lr=5e-4 (a, b)   - Adam + lr = 1e-4 (c)</li>\n<li>Class-balanced BCE to mimic the metric + Label smoothing</li>\n<li>10 epochs, bs=16</li>\n<li>Flip &amp; color augmentations</li></ul></li>\n<li>Inference<ul>\n<li>4 flips TTA</li>\n<li>Simple average</li>\n<li>Scaling !</li></ul></li>\n</ul>\n<h3>Adapting to the Competition Metric</h3>\n<p><a href=\"https://ibb.co/C1BYWbF\"><img src=\"https://i.ibb.co/LkScvxf/logloss.png\" alt=\"logloss\"></a></p>\n<ul>\n<li>The logloss heavily penalizes confident mistakes, comparatively to the reward for confident guesses<ul>\n<li>Since our models have low AUCs, we want to avoid the “Heavy Penalty” zone</li></ul></li>\n<li>My ensemble already outputs most of its scores in the “safe” range possibly because  my models were designed to underfit (Label smoothing, small model size, short training with low LRs)<ul>\n<li><strong>Ensemble CV : 0.655</strong></li></ul></li>\n<li>You can force your model to be safer ! (cf <a href=\"https://www.kaggle.com/competitions/mayo-clinic-strip-ai/discussion/357877\" target=\"_blank\">this post</a>)<ul>\n<li>Linearly Rescale to the 0.15 / 0.85 range</li>\n<li>Clip to 0.25 / 0.75</li>\n<li><strong>Final CV : 0.640</strong> - <strong>Public LB 0.733</strong> - <strong>Private LB 0.666</strong></li></ul></li>\n</ul>\n<p>Thanks for reading :)</p>",
  "messages": [
    {
      "id": 1974650,
      "postDate": "2022-10-06T11:26:13.160Z",
      "content": "<p>Thanks everyone for the competition, congratz to the winners &amp; kudos to the learners. </p>\n<p>Although this one is a controversial one because most people couldn't even beat the sample submission, I still believe it was an alright competition. The key is to understand that there is not a lot of signal, and that the competition metric is not friendly with overconfident models, more on that later.</p>\n<p>Because of the nature of the competition (not a lot of signal, and basically no public LB), I decided not too invest too much time on it, so I am very happy with the result !</p>\n<h3>Data</h3>\n<ul>\n<li>Color normalization : <ul>\n<li>Detect the background color</li>\n<li>Normalize all the image by  <code>𝑟 = 𝑏𝑎𝑐𝑘𝑔𝑟𝑜𝑢𝑛𝑑_𝑐𝑜𝑙𝑜𝑟/(255,   255,  255)</code></li></ul></li>\n<li>Simple Approach : <ul>\n<li>Remove all white chunks from the image </li>\n<li>Resize to 1024x1024</li></ul></li>\n<li>Advanced Approach (cf figure below): <ul>\n<li>Detect duplicated areas in the image</li>\n<li>Keep only one of them</li>\n<li>Resize to have the longest edge to 1024, preserving aspect ratio</li></ul></li>\n</ul>\n<p><a href=\"https://ibb.co/fpnzYN1\"><img src=\"https://i.ibb.co/HzxfDPF/mayo-data-pipe.png\" alt=\"mayo-data-pipe\"></a></p>\n<p>You can see some examples <a href=\"https://www.kaggle.com/code/theoviel/inference-mayo\" target=\"_blank\">here</a>.</p>\n<h3>Models</h3>\n<ul>\n<li>3 Small Models  -  <strong>CV AUC 0.684</strong><ol>\n<li>Resnet10t – AUC 0.661</li>\n<li>EfficientNet-b0 – AUC 0.671</li>\n<li>EfficientNet-b0 using simple approach images - AUC 0.662</li></ol></li>\n<li>Training<ul>\n<li>Image size : 1024x1024</li>\n<li>Ranger + lr=5e-4 (a, b)   - Adam + lr = 1e-4 (c)</li>\n<li>Class-balanced BCE to mimic the metric + Label smoothing</li>\n<li>10 epochs, bs=16</li>\n<li>Flip &amp; color augmentations</li></ul></li>\n<li>Inference<ul>\n<li>4 flips TTA</li>\n<li>Simple average</li>\n<li>Scaling !</li></ul></li>\n</ul>\n<h3>Adapting to the Competition Metric</h3>\n<p><a href=\"https://ibb.co/C1BYWbF\"><img src=\"https://i.ibb.co/LkScvxf/logloss.png\" alt=\"logloss\"></a></p>\n<ul>\n<li>The logloss heavily penalizes confident mistakes, comparatively to the reward for confident guesses<ul>\n<li>Since our models have low AUCs, we want to avoid the “Heavy Penalty” zone</li></ul></li>\n<li>My ensemble already outputs most of its scores in the “safe” range possibly because  my models were designed to underfit (Label smoothing, small model size, short training with low LRs)<ul>\n<li><strong>Ensemble CV : 0.655</strong></li></ul></li>\n<li>You can force your model to be safer ! (cf <a href=\"https://www.kaggle.com/competitions/mayo-clinic-strip-ai/discussion/357877\" target=\"_blank\">this post</a>)<ul>\n<li>Linearly Rescale to the 0.15 / 0.85 range</li>\n<li>Clip to 0.25 / 0.75</li>\n<li><strong>Final CV : 0.640</strong> - <strong>Public LB 0.733</strong> - <strong>Private LB 0.666</strong></li></ul></li>\n</ul>\n<p>Thanks for reading :)</p>",
      "rawMarkdown": "Thanks everyone for the competition, congratz to the winners & kudos to the learners. \n\nAlthough this one is a controversial one because most people couldn't even beat the sample submission, I still believe it was an alright competition. The key is to understand that there is not a lot of signal, and that the competition metric is not friendly with overconfident models, more on that later.\n\nBecause of the nature of the competition (not a lot of signal, and basically no public LB), I decided not too invest too much time on it, so I am very happy with the result !\n\n### Data\n\n- Color normalization : \n - Detect the background color\n - Normalize all the image by  `𝑟 = 𝑏𝑎𝑐𝑘𝑔𝑟𝑜𝑢𝑛𝑑_𝑐𝑜𝑙𝑜𝑟/(255,   255,  255)`\n- Simple Approach : \n - Remove all white chunks from the image \n - Resize to 1024x1024\n- Advanced Approach (cf figure below): \n - Detect duplicated areas in the image\n - Keep only one of them\n - Resize to have the longest edge to 1024, preserving aspect ratio\n\n<a href=\"https://ibb.co/fpnzYN1\"><img src=\"https://i.ibb.co/HzxfDPF/mayo-data-pipe.png\" alt=\"mayo-data-pipe\" border=\"0\"></a>\n\nYou can see some examples [here](https://www.kaggle.com/code/theoviel/inference-mayo).\n\n\n### Models\n\n- 3 Small Models  -  **CV AUC 0.684**\n  1. Resnet10t – AUC 0.661\n  2. EfficientNet-b0 – AUC 0.671\n  3. EfficientNet-b0 using simple approach images - AUC 0.662\n- Training\n - Image size : 1024x1024\n - Ranger + lr=5e-4 (a, b)   - Adam + lr = 1e-4 (c)\n - Class-balanced BCE to mimic the metric + Label smoothing\n - 10 epochs, bs=16\n - Flip & color augmentations\n- Inference\n  - 4 flips TTA\n  - Simple average\n  - Scaling !\n\n### Adapting to the Competition Metric\n\n<a href=\"https://ibb.co/C1BYWbF\"><img src=\"https://i.ibb.co/LkScvxf/logloss.png\" alt=\"logloss\" border=\"0\"></a>\n\n- The logloss heavily penalizes confident mistakes, comparatively to the reward for confident guesses\n  - Since our models have low AUCs, we want to avoid the “Heavy Penalty” zone\n- My ensemble already outputs most of its scores in the “safe” range possibly because  my models were designed to underfit (Label smoothing, small model size, short training with low LRs)\n  - **Ensemble CV : 0.655**\n- You can force your model to be safer ! (cf [this post](https://www.kaggle.com/competitions/mayo-clinic-strip-ai/discussion/357877))\n  - Linearly Rescale to the 0.15 / 0.85 range\n  - Clip to 0.25 / 0.75\n  - **Final CV : 0.640** - **Public LB 0.733** - **Private LB 0.666**\n\nThanks for reading :)",
      "votes": 33
    },
    {
      "id": 1974895,
      "postDate": "2022-10-06T14:15:19.917Z",
      "content": "<p>Congrats solo gold!!  The data pipeline is amazing,  did you try some heavy models?</p>",
      "rawMarkdown": "Congrats solo gold!!  The data pipeline is amazing,  did you try some heavy models?",
      "votes": 3,
      "replies": [
        {
          "id": 1974939,
          "postDate": "2022-10-06T14:42:14.760Z",
          "content": "<p>Thanks ! <br>\nI tried effnets up to b5, resnet 18, 24 and 50, but the smallest ones worked better</p>",
          "rawMarkdown": "Thanks ! \nI tried effnets up to b5, resnet 18, 24 and 50, but the smallest ones worked better",
          "votes": 3
        },
        {
          "id": 1974960,
          "postDate": "2022-10-06T14:56:47.837Z",
          "content": "<p>Yeah amazing pre-processing and post-processing. Can you try DenseNet variations with your pipeline? They were my best models in both this competition and RSNA-MICCAI Brain Tumor Radiogenomic Classification for some reason. I think dense connections of low level features could be the reason.</p>",
          "rawMarkdown": "Yeah amazing pre-processing and post-processing. Can you try DenseNet variations with your pipeline? They were my best models in both this competition and RSNA-MICCAI Brain Tumor Radiogenomic Classification for some reason. I think dense connections of low level features could be the reason.",
          "votes": 3
        },
        {
          "id": 1975022,
          "postDate": "2022-10-06T15:25:08.600Z",
          "content": "<p>Densenet scores about 0.01/0.02 lower on AUC so it's not really good on my pipeline </p>\n<p>Other interesting models (but not good enough to help my ensemble) were convnext-pico, convnext-atto, efficientnet-b1 &amp; effnet-v2-tiny.</p>",
          "rawMarkdown": "Densenet scores about 0.01/0.02 lower on AUC so it's not really good on my pipeline \n\nOther interesting models (but not good enough to help my ensemble) were convnext-pico, convnext-atto, efficientnet-b1 & effnet-v2-tiny.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1974891,
      "postDate": "2022-10-06T14:12:52.383Z",
      "content": "<p>Congrats! I like the postprocessing part, it seems like a very good boost. Also it is amazing to have such a good cv with efb0. </p>",
      "rawMarkdown": "Congrats! I like the postprocessing part, it seems like a very good boost. Also it is amazing to have such a good cv with efb0. ",
      "votes": 3,
      "replies": [
        {
          "id": 1974940,
          "postDate": "2022-10-06T14:42:48.747Z",
          "content": "<p>Thanks, and congratulations again on the win ! </p>",
          "rawMarkdown": "Thanks, and congratulations again on the win ! ",
          "votes": 2
        }
      ]
    },
    {
      "id": 1978840,
      "postDate": "2022-10-09T03:36:05.373Z",
      "content": "<p>Congrats on your success!</p>",
      "rawMarkdown": "Congrats on your success!",
      "votes": 1
    },
    {
      "id": 1976544,
      "postDate": "2022-10-07T12:14:21.950Z",
      "content": "<p>This is wonderful and the approach is intuitive and powerful. Well earned!</p>",
      "rawMarkdown": "This is wonderful and the approach is intuitive and powerful. Well earned!",
      "votes": 1
    },
    {
      "id": 1976085,
      "postDate": "2022-10-07T06:41:44.807Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> thanks for putting out your solution here. How did you shrink the 395 GB of image data, did you work with the data on Kaggle only or did you take it on computer to work on it. I am asking this as I am a machine learning beginner and the best resource I have is Kaggle, can you share some initial preprocessing techniques for such large image data.</p>",
      "rawMarkdown": "Congratulations @theoviel thanks for putting out your solution here. How did you shrink the 395 GB of image data, did you work with the data on Kaggle only or did you take it on computer to work on it. I am asking this as I am a machine learning beginner and the best resource I have is Kaggle, can you share some initial preprocessing techniques for such large image data.",
      "votes": 1,
      "replies": [
        {
          "id": 1976195,
          "postDate": "2022-10-07T08:04:47.543Z",
          "content": "<p>I worked on my local machine :)</p>\n<p>My preprocessing code is available <a href=\"https://www.kaggle.com/code/theoviel/inference-mayo\" target=\"_blank\">here</a>, I basically run it once and then work with the 1024 images.</p>",
          "rawMarkdown": "I worked on my local machine :)\n\nMy preprocessing code is available [here](https://www.kaggle.com/code/theoviel/inference-mayo), I basically run it once and then work with the 1024 images.",
          "votes": 1,
          "replies": [
            {
              "id": 2177041,
              "postDate": "2023-03-11T06:35:19.983Z",
              "content": "<p>Thank a lot</p>",
              "rawMarkdown": "Thank a lot"
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1974895,
      "author_name": "KKY",
      "author_url": "",
      "post_date": "2022-10-06T14:15:19.917000",
      "content": "<p>Congrats solo gold!!  The data pipeline is amazing,  did you try some heavy models?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1974939,
          "author_name": "Theo Viel",
          "author_url": "",
          "post_date": "2022-10-06T14:42:14.760000",
          "content": "<p>Thanks ! <br>\nI tried effnets up to b5, resnet 18, 24 and 50, but the smallest ones worked better</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1974960,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2022-10-06T14:56:47.837000",
          "content": "<p>Yeah amazing pre-processing and post-processing. Can you try DenseNet variations with your pipeline? They were my best models in both this competition and RSNA-MICCAI Brain Tumor Radiogenomic Classification for some reason. I think dense connections of low level features could be the reason.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1975022,
          "author_name": "Theo Viel",
          "author_url": "",
          "post_date": "2022-10-06T15:25:08.600000",
          "content": "<p>Densenet scores about 0.01/0.02 lower on AUC so it's not really good on my pipeline </p>\n<p>Other interesting models (but not good enough to help my ensemble) were convnext-pico, convnext-atto, efficientnet-b1 &amp; effnet-v2-tiny.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1974891,
      "author_name": "khyeh",
      "author_url": "",
      "post_date": "2022-10-06T14:12:52.383000",
      "content": "<p>Congrats! I like the postprocessing part, it seems like a very good boost. Also it is amazing to have such a good cv with efb0. </p>",
      "votes": 3,
      "replies": [
        {
          "id": 1974940,
          "author_name": "Theo Viel",
          "author_url": "",
          "post_date": "2022-10-06T14:42:48.747000",
          "content": "<p>Thanks, and congratulations again on the win ! </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1978840,
      "author_name": "Will",
      "author_url": "",
      "post_date": "2022-10-09T03:36:05.373000",
      "content": "<p>Congrats on your success!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1976544,
      "author_name": "Darien Schettler",
      "author_url": "",
      "post_date": "2022-10-07T12:14:21.950000",
      "content": "<p>This is wonderful and the approach is intuitive and powerful. Well earned!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1976085,
      "author_name": "Manish Tripathi",
      "author_url": "",
      "post_date": "2022-10-07T06:41:44.807000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> thanks for putting out your solution here. How did you shrink the 395 GB of image data, did you work with the data on Kaggle only or did you take it on computer to work on it. I am asking this as I am a machine learning beginner and the best resource I have is Kaggle, can you share some initial preprocessing techniques for such large image data.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1976195,
          "author_name": "Theo Viel",
          "author_url": "",
          "post_date": "2022-10-07T08:04:47.543000",
          "content": "<p>I worked on my local machine :)</p>\n<p>My preprocessing code is available <a href=\"https://www.kaggle.com/code/theoviel/inference-mayo\" target=\"_blank\">here</a>, I basically run it once and then work with the 1024 images.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2177041,
              "author_name": "Manish Tripathi",
              "author_url": "",
              "post_date": "2023-03-11T06:35:19.983000",
              "content": "<p>Thank a lot</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1974650": "Thanks everyone for the competition, congratz to the winners & kudos to the learners. \n\nAlthough this one is a controversial one because most people couldn't even beat the sample submission, I still believe it was an alright competition. The key is to understand that there is not a lot of signal, and that the competition metric is not friendly with overconfident models, more on that later.\n\nBecause of the nature of the competition (not a lot of signal, and basically no public LB), I decided not too invest too much time on it, so I am very happy with the result !\n\n### Data\n\n- Color normalization : \n - Detect the background color\n - Normalize all the image by  `𝑟 = 𝑏𝑎𝑐𝑘𝑔𝑟𝑜𝑢𝑛𝑑_𝑐𝑜𝑙𝑜𝑟/(255,   255,  255)`\n- Simple Approach : \n - Remove all white chunks from the image \n - Resize to 1024x1024\n- Advanced Approach (cf figure below): \n - Detect duplicated areas in the image\n - Keep only one of them\n - Resize to have the longest edge to 1024, preserving aspect ratio\n\n<a href=\"https://ibb.co/fpnzYN1\"><img src=\"https://i.ibb.co/HzxfDPF/mayo-data-pipe.png\" alt=\"mayo-data-pipe\" border=\"0\"></a>\n\nYou can see some examples [here](https://www.kaggle.com/code/theoviel/inference-mayo).\n\n\n### Models\n\n- 3 Small Models  -  **CV AUC 0.684**\n  1. Resnet10t – AUC 0.661\n  2. EfficientNet-b0 – AUC 0.671\n  3. EfficientNet-b0 using simple approach images - AUC 0.662\n- Training\n - Image size : 1024x1024\n - Ranger + lr=5e-4 (a, b)   - Adam + lr = 1e-4 (c)\n - Class-balanced BCE to mimic the metric + Label smoothing\n - 10 epochs, bs=16\n - Flip & color augmentations\n- Inference\n  - 4 flips TTA\n  - Simple average\n  - Scaling !\n\n### Adapting to the Competition Metric\n\n<a href=\"https://ibb.co/C1BYWbF\"><img src=\"https://i.ibb.co/LkScvxf/logloss.png\" alt=\"logloss\" border=\"0\"></a>\n\n- The logloss heavily penalizes confident mistakes, comparatively to the reward for confident guesses\n  - Since our models have low AUCs, we want to avoid the “Heavy Penalty” zone\n- My ensemble already outputs most of its scores in the “safe” range possibly because  my models were designed to underfit (Label smoothing, small model size, short training with low LRs)\n  - **Ensemble CV : 0.655**\n- You can force your model to be safer ! (cf [this post](https://www.kaggle.com/competitions/mayo-clinic-strip-ai/discussion/357877))\n  - Linearly Rescale to the 0.15 / 0.85 range\n  - Clip to 0.25 / 0.75\n  - **Final CV : 0.640** - **Public LB 0.733** - **Private LB 0.666**\n\nThanks for reading :)",
    "1974895": "Congrats solo gold!!  The data pipeline is amazing,  did you try some heavy models?",
    "1974891": "Congrats! I like the postprocessing part, it seems like a very good boost. Also it is amazing to have such a good cv with efb0. ",
    "1978840": "Congrats on your success!",
    "1976544": "This is wonderful and the approach is intuitive and powerful. Well earned!",
    "1976085": "Congratulations @theoviel thanks for putting out your solution here. How did you shrink the 395 GB of image data, did you work with the data on Kaggle only or did you take it on computer to work on it. I am asking this as I am a machine learning beginner and the best resource I have is Kaggle, can you share some initial preprocessing techniques for such large image data."
  }
}