{
  "id": 391378,
  "title": "10th Place Solution",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/391378",
  "author_name": "moritake04",
  "post_date": "2023-03-01T09:41:13.958000",
  "votes": 28,
  "comment_count": 10,
  "views": 0,
  "content": "<p>First, I would like to thank the competition hosts for organizing this competition and also the participants.</p>\n<p>I’m very happy to have won my first gold medal at Kaggle!</p>\n<h2>Overview</h2>\n<p>My model is a very simple single input model, with a simple pipeline. (So I didn't expect to win a gold medal.)</p>\n<p>My final model is a simple average ensemble of the following 3 models. (model names are from timm)</p>\n<ul>\n<li>tf_efficientnetv2_s (no aux loss)</li>\n<li>tf_efficientnetv2_s (using aux loss)</li>\n<li>maxvit_tiny_tf_384.in1k (using aux loss)</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9058219%2F6c141d2897b20fac001918828caaef9b%2Fpipeline.png?generation=1677661238257982&amp;alt=media\" alt=\"pipeline\"></p>\n<p>I built the Baseline myself, but it is similar to <a href=\"https://www.kaggle.com/vslaykovsky\" target=\"_blank\">@vslaykovsky</a>'s (<a href=\"https://www.kaggle.com/code/vslaykovsky/train-pytorch-aux-targets-weighted-loss-thres\" target=\"_blank\">https://www.kaggle.com/code/vslaykovsky/train-pytorch-aux-targets-weighted-loss-thres</a>) and I actually used it as a reference.</p>\n<h2>Solution</h2>\n<ul>\n<li>CV Strategy<ul>\n<li>StratifiedGroupKFold (n=4, groups=patient_id)</li></ul></li>\n<li>Data・Preprocessing<ul>\n<li>No external data was used.</li>\n<li>ROI extraction was performed using rule-based method (reference: <a href=\"https://www.kaggle.com/code/vslaykovsky/rsna-cut-off-empty-space-from-images\" target=\"_blank\">https://www.kaggle.com/code/vslaykovsky/rsna-cut-off-empty-space-from-images</a>).</li>\n<li>Fast loading was achieved using nvJPEG2000 (reference: <a href=\"https://www.kaggle.com/code/snaker/easy-load-the-image-with-nvjpeg2000\" target=\"_blank\">https://www.kaggle.com/code/snaker/easy-load-the-image-with-nvjpeg2000</a>).</li>\n<li>The resolution was set to 1536x960 for efficientnet and 1536x768 for maxvit.</li>\n<li>The channel number was set to 3 (to use pretrained models).</li>\n<li>Min-max scaling (-1.0 ~ 1.0) was applied.</li>\n<li>VOI LUT was applied.</li>\n<li>8 bit</li>\n<li>Care was taken to ensure that the preprocessing pipeline remains unchanged between training and inference.</li></ul></li>\n<li>Dealing with Imbalance Data<ul>\n<li>Batch size of 8. Adjusted to have a majority (Not Cancer) to minority (Cancer) ratio of 7:1 for each batch. This is essentially oversampling.</li></ul></li>\n<li>Data Augmentation<ul>\n<li>I used albumentations (a data augmentation library).</li></ul></li>\n</ul>\n<pre><code> albumentations  A\n\nA.HorizontalFlip(p=)\nA.VerticalFlip(p=)\nA.ShiftScaleRotate(shift_limit=, scale_limit=, rotate_limit=, p=)\nA.OneOf([\n    A.RandomGamma(gamma_limit=(, ), p=),\n    A.RandomBrightnessContrast(brightness_limit=, contrast_limit=, p=)\n], p=)\nA.CoarseDropout(max_height=, max_width=, p=)\n</code></pre>\n<ul>\n<li>Model Parameters<ul>\n<li>drop_rate: 0.8</li>\n<li>drop_path_rate: 0.2</li>\n<li>criterion: BCEWithLogitsLoss</li>\n<li>optimizer: Adam (lr: 1.0e-4)</li>\n<li>scheduler: OneCycleLR (pct_start: 0.1, div_factor: 1.0e+3, max_lr: 1.0e-4)</li>\n<li>epoch: 5</li>\n<li>batch_size: 8 (and accumulate_grad_batches=4, so 8*4=32)</li>\n<li>fp16 (training and inference)</li>\n<li>aux loss<ul>\n<li>aux target: site_id, laterality, view, implant, biopsy, invasive, BIRADS, density, difficult_negative_case, age</li>\n<li>aux loss weight: 1.0</li>\n<li>From the discussion (<a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/370341\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/370341</a>), there was information that there were machine_id in the LB that were not in the training data, so the machine_id was not used.</li>\n<li>However, there is a feeling that a model specific to this competition has been created by using site_id. It is questionable whether this model can be applied to mammography images from other facilities…</li></ul></li></ul></li>\n<li>Inference<ul>\n<li>top3 mean aggregation（Average of the top 3 predictions per prediction id）</li>\n<li>To reduce inference time, when doing the ensemble, the 4 models from the 4-fold cross-validation were not used for inference, but a single model trained on all data was used.<br>\n(cross-validation was used to check CV scores and to find thresholds.)</li>\n<li>The ensemble is a simple average; I also tried optimising the weights using the Nelder-Mead method, but as the weights were calculated almost identically across all models, I opted for a simple average.</li></ul></li>\n<li>Results</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>AUROC</th>\n<th>local cv pF1</th>\n<th>public pF1</th>\n<th>private pF1</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1. tf_efficientnetv2_s (no aux loss)</td>\n<td>0.897</td>\n<td>0.424</td>\n<td>0.48※</td>\n<td>0.41※</td>\n</tr>\n<tr>\n<td>2. tf_efficientnetv2_s (using aux loss)</td>\n<td>0.894</td>\n<td>0.435</td>\n<td>0.48※</td>\n<td>0.45※</td>\n</tr>\n<tr>\n<td>3. maxvit_tiny_tf_384.in1k (using aux loss)</td>\n<td>0.896</td>\n<td>0.443</td>\n<td>Not submitted</td>\n<td>Not submitted</td>\n</tr>\n<tr>\n<td>ensemble (1+2+3 prob mean)</td>\n<td>0.917</td>\n<td>0.482</td>\n<td>0.54</td>\n<td>0.49</td>\n</tr>\n</tbody>\n</table>\n<p>※ simple mean aggregation (not top3 mean aggregation)</p>\n<h2>What Worked</h2>\n<ul>\n<li>Large image size</li>\n<li>Strong dropout (dropout rate=0.8)</li>\n<li>Stochastic depth rate (timm's drop_path_rate)</li>\n<li>Strong augmentation<ul>\n<li>By default, the border_mode of ShiftScaleRotate in Albumentations is cv2.BORDER_REFLECT_101, which produced better results than cv2.BORDER_CONSTANT.</li>\n<li>If border_mode = cv2.BORDER_REFLECT_101, the overflowing image is reflected by the affine transformation and a weird image is generated. ↓<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9058219%2F3bc241b5353f842f03bdecab1625886a%2Freflect2.png?generation=1677661658109905&amp;alt=media\" alt=\"reflect2\"></li>\n<li>However, as the results were better when the image was reflected, I considered that the task of this competition was more about texture than about the shape of the image.</li></ul></li>\n<li>Dealing with imbalances in mini-batches.</li>\n<li>ROI extraction</li>\n<li>Aux Loss</li>\n<li>VOI-LUT</li>\n<li>Anomaly image processing<ul>\n<li>Some images were black and full of noise, which were removed. (reference: <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/373208\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/373208</a>)</li></ul></li>\n<li>Ensemble</li>\n</ul>\n<h2>What Didn’t Worked</h2>\n<ul>\n<li>Dealing with imbalances by loss functions<ul>\n<li>Focal Loss, Weighted Cross Entropy</li></ul></li>\n<li>Metric learning</li>\n<li>Center-crop to maintain aspect ratio</li>\n<li>TTA<ul>\n<li>Not adopted as results did not change much and inference time increased</li></ul></li>\n<li>Large parameter models<ul>\n<li>Only tried efficientnetv2_l, but it is only slightly better than s. Not adopted due to increased inference time.</li></ul></li>\n<li>Some data extensions<ul>\n<li>CutMix, MixUp, Grid distortion, elastic transformation, CLAHE</li></ul></li>\n</ul>\n<h2>Code</h2>\n<ul>\n<li>GitHub -&gt; <a href=\"https://github.com/moritake04/rsna-breast-cancer-detection\" target=\"_blank\">https://github.com/moritake04/rsna-breast-cancer-detection</a></li>\n<li>inference notebook -&gt; <a href=\"https://www.kaggle.com/code/moritake04/private10th-final-sub\" target=\"_blank\">https://www.kaggle.com/code/moritake04/private10th-final-sub</a></li>\n</ul>\n<p>(Sorry for the dirty code.) </p>\n<h2>Thanks and Acknowledgements</h2>\n<p>Again, thank you to the organizers and participants of this competition, and I will continue to work hard to become a kaggle master. </p>",
  "messages": [
    {
      "id": 2164086,
      "postDate": "2023-03-01T09:41:13.960Z",
      "content": "<p>First, I would like to thank the competition hosts for organizing this competition and also the participants.</p>\n<p>I’m very happy to have won my first gold medal at Kaggle!</p>\n<h2>Overview</h2>\n<p>My model is a very simple single input model, with a simple pipeline. (So I didn't expect to win a gold medal.)</p>\n<p>My final model is a simple average ensemble of the following 3 models. (model names are from timm)</p>\n<ul>\n<li>tf_efficientnetv2_s (no aux loss)</li>\n<li>tf_efficientnetv2_s (using aux loss)</li>\n<li>maxvit_tiny_tf_384.in1k (using aux loss)</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9058219%2F6c141d2897b20fac001918828caaef9b%2Fpipeline.png?generation=1677661238257982&amp;alt=media\" alt=\"pipeline\"></p>\n<p>I built the Baseline myself, but it is similar to <a href=\"https://www.kaggle.com/vslaykovsky\" target=\"_blank\">@vslaykovsky</a>'s (<a href=\"https://www.kaggle.com/code/vslaykovsky/train-pytorch-aux-targets-weighted-loss-thres\" target=\"_blank\">https://www.kaggle.com/code/vslaykovsky/train-pytorch-aux-targets-weighted-loss-thres</a>) and I actually used it as a reference.</p>\n<h2>Solution</h2>\n<ul>\n<li>CV Strategy<ul>\n<li>StratifiedGroupKFold (n=4, groups=patient_id)</li></ul></li>\n<li>Data・Preprocessing<ul>\n<li>No external data was used.</li>\n<li>ROI extraction was performed using rule-based method (reference: <a href=\"https://www.kaggle.com/code/vslaykovsky/rsna-cut-off-empty-space-from-images\" target=\"_blank\">https://www.kaggle.com/code/vslaykovsky/rsna-cut-off-empty-space-from-images</a>).</li>\n<li>Fast loading was achieved using nvJPEG2000 (reference: <a href=\"https://www.kaggle.com/code/snaker/easy-load-the-image-with-nvjpeg2000\" target=\"_blank\">https://www.kaggle.com/code/snaker/easy-load-the-image-with-nvjpeg2000</a>).</li>\n<li>The resolution was set to 1536x960 for efficientnet and 1536x768 for maxvit.</li>\n<li>The channel number was set to 3 (to use pretrained models).</li>\n<li>Min-max scaling (-1.0 ~ 1.0) was applied.</li>\n<li>VOI LUT was applied.</li>\n<li>8 bit</li>\n<li>Care was taken to ensure that the preprocessing pipeline remains unchanged between training and inference.</li></ul></li>\n<li>Dealing with Imbalance Data<ul>\n<li>Batch size of 8. Adjusted to have a majority (Not Cancer) to minority (Cancer) ratio of 7:1 for each batch. This is essentially oversampling.</li></ul></li>\n<li>Data Augmentation<ul>\n<li>I used albumentations (a data augmentation library).</li></ul></li>\n</ul>\n<pre><code> albumentations  A\n\nA.HorizontalFlip(p=)\nA.VerticalFlip(p=)\nA.ShiftScaleRotate(shift_limit=, scale_limit=, rotate_limit=, p=)\nA.OneOf([\n    A.RandomGamma(gamma_limit=(, ), p=),\n    A.RandomBrightnessContrast(brightness_limit=, contrast_limit=, p=)\n], p=)\nA.CoarseDropout(max_height=, max_width=, p=)\n</code></pre>\n<ul>\n<li>Model Parameters<ul>\n<li>drop_rate: 0.8</li>\n<li>drop_path_rate: 0.2</li>\n<li>criterion: BCEWithLogitsLoss</li>\n<li>optimizer: Adam (lr: 1.0e-4)</li>\n<li>scheduler: OneCycleLR (pct_start: 0.1, div_factor: 1.0e+3, max_lr: 1.0e-4)</li>\n<li>epoch: 5</li>\n<li>batch_size: 8 (and accumulate_grad_batches=4, so 8*4=32)</li>\n<li>fp16 (training and inference)</li>\n<li>aux loss<ul>\n<li>aux target: site_id, laterality, view, implant, biopsy, invasive, BIRADS, density, difficult_negative_case, age</li>\n<li>aux loss weight: 1.0</li>\n<li>From the discussion (<a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/370341\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/370341</a>), there was information that there were machine_id in the LB that were not in the training data, so the machine_id was not used.</li>\n<li>However, there is a feeling that a model specific to this competition has been created by using site_id. It is questionable whether this model can be applied to mammography images from other facilities…</li></ul></li></ul></li>\n<li>Inference<ul>\n<li>top3 mean aggregation（Average of the top 3 predictions per prediction id）</li>\n<li>To reduce inference time, when doing the ensemble, the 4 models from the 4-fold cross-validation were not used for inference, but a single model trained on all data was used.<br>\n(cross-validation was used to check CV scores and to find thresholds.)</li>\n<li>The ensemble is a simple average; I also tried optimising the weights using the Nelder-Mead method, but as the weights were calculated almost identically across all models, I opted for a simple average.</li></ul></li>\n<li>Results</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>AUROC</th>\n<th>local cv pF1</th>\n<th>public pF1</th>\n<th>private pF1</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1. tf_efficientnetv2_s (no aux loss)</td>\n<td>0.897</td>\n<td>0.424</td>\n<td>0.48※</td>\n<td>0.41※</td>\n</tr>\n<tr>\n<td>2. tf_efficientnetv2_s (using aux loss)</td>\n<td>0.894</td>\n<td>0.435</td>\n<td>0.48※</td>\n<td>0.45※</td>\n</tr>\n<tr>\n<td>3. maxvit_tiny_tf_384.in1k (using aux loss)</td>\n<td>0.896</td>\n<td>0.443</td>\n<td>Not submitted</td>\n<td>Not submitted</td>\n</tr>\n<tr>\n<td>ensemble (1+2+3 prob mean)</td>\n<td>0.917</td>\n<td>0.482</td>\n<td>0.54</td>\n<td>0.49</td>\n</tr>\n</tbody>\n</table>\n<p>※ simple mean aggregation (not top3 mean aggregation)</p>\n<h2>What Worked</h2>\n<ul>\n<li>Large image size</li>\n<li>Strong dropout (dropout rate=0.8)</li>\n<li>Stochastic depth rate (timm's drop_path_rate)</li>\n<li>Strong augmentation<ul>\n<li>By default, the border_mode of ShiftScaleRotate in Albumentations is cv2.BORDER_REFLECT_101, which produced better results than cv2.BORDER_CONSTANT.</li>\n<li>If border_mode = cv2.BORDER_REFLECT_101, the overflowing image is reflected by the affine transformation and a weird image is generated. ↓<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9058219%2F3bc241b5353f842f03bdecab1625886a%2Freflect2.png?generation=1677661658109905&amp;alt=media\" alt=\"reflect2\"></li>\n<li>However, as the results were better when the image was reflected, I considered that the task of this competition was more about texture than about the shape of the image.</li></ul></li>\n<li>Dealing with imbalances in mini-batches.</li>\n<li>ROI extraction</li>\n<li>Aux Loss</li>\n<li>VOI-LUT</li>\n<li>Anomaly image processing<ul>\n<li>Some images were black and full of noise, which were removed. (reference: <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/373208\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/373208</a>)</li></ul></li>\n<li>Ensemble</li>\n</ul>\n<h2>What Didn’t Worked</h2>\n<ul>\n<li>Dealing with imbalances by loss functions<ul>\n<li>Focal Loss, Weighted Cross Entropy</li></ul></li>\n<li>Metric learning</li>\n<li>Center-crop to maintain aspect ratio</li>\n<li>TTA<ul>\n<li>Not adopted as results did not change much and inference time increased</li></ul></li>\n<li>Large parameter models<ul>\n<li>Only tried efficientnetv2_l, but it is only slightly better than s. Not adopted due to increased inference time.</li></ul></li>\n<li>Some data extensions<ul>\n<li>CutMix, MixUp, Grid distortion, elastic transformation, CLAHE</li></ul></li>\n</ul>\n<h2>Code</h2>\n<ul>\n<li>GitHub -&gt; <a href=\"https://github.com/moritake04/rsna-breast-cancer-detection\" target=\"_blank\">https://github.com/moritake04/rsna-breast-cancer-detection</a></li>\n<li>inference notebook -&gt; <a href=\"https://www.kaggle.com/code/moritake04/private10th-final-sub\" target=\"_blank\">https://www.kaggle.com/code/moritake04/private10th-final-sub</a></li>\n</ul>\n<p>(Sorry for the dirty code.) </p>\n<h2>Thanks and Acknowledgements</h2>\n<p>Again, thank you to the organizers and participants of this competition, and I will continue to work hard to become a kaggle master. </p>",
      "rawMarkdown": "First, I would like to thank the competition hosts for organizing this competition and also the participants.\n\nI’m very happy to have won my first gold medal at Kaggle!\n\n## Overview\n\nMy model is a very simple single input model, with a simple pipeline. (So I didn't expect to win a gold medal.)\n\nMy final model is a simple average ensemble of the following 3 models. (model names are from timm)\n- tf_efficientnetv2_s (no aux loss)\n- tf_efficientnetv2_s (using aux loss)\n- maxvit_tiny_tf_384.in1k (using aux loss)\n\n![pipeline](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9058219%2F6c141d2897b20fac001918828caaef9b%2Fpipeline.png?generation=1677661238257982&alt=media)\n\nI built the Baseline myself, but it is similar to @vslaykovsky's (https://www.kaggle.com/code/vslaykovsky/train-pytorch-aux-targets-weighted-loss-thres) and I actually used it as a reference.\n\n## Solution\n\n- CV Strategy\n    - StratifiedGroupKFold (n=4, groups=patient_id)\n- Data・Preprocessing\n    - No external data was used.\n    - ROI extraction was performed using rule-based method (reference: [https://www.kaggle.com/code/vslaykovsky/rsna-cut-off-empty-space-from-images](https://www.kaggle.com/code/vslaykovsky/rsna-cut-off-empty-space-from-images)).\n    - Fast loading was achieved using nvJPEG2000 (reference: [https://www.kaggle.com/code/snaker/easy-load-the-image-with-nvjpeg2000](https://www.kaggle.com/code/snaker/easy-load-the-image-with-nvjpeg2000)).\n    - The resolution was set to 1536x960 for efficientnet and 1536x768 for maxvit.\n    - The channel number was set to 3 (to use pretrained models).\n    - Min-max scaling (-1.0 ~ 1.0) was applied.\n    - VOI LUT was applied.\n    - 8 bit\n    - Care was taken to ensure that the preprocessing pipeline remains unchanged between training and inference.\n- Dealing with Imbalance Data\n    - Batch size of 8. Adjusted to have a majority (Not Cancer) to minority (Cancer) ratio of 7:1 for each batch. This is essentially oversampling.\n- Data Augmentation\n    - I used albumentations (a data augmentation library).\n\n```python\nimport albumentations as A\n\nA.HorizontalFlip(p=0.5)\nA.VerticalFlip(p=0.5)\nA.ShiftScaleRotate(shift_limit=0.1, scale_limit=0.2, rotate_limit=45, p=0.8)\nA.OneOf([\n    A.RandomGamma(gamma_limit=(50, 150), p=0.5),\n    A.RandomBrightnessContrast(brightness_limit=0.5, contrast_limit=0.5, p=0.5)\n], p=0.5)\nA.CoarseDropout(max_height=8, max_width=8, p=0.5)\n```\n\n- Model Parameters\n    - drop_rate: 0.8\n    - drop_path_rate: 0.2\n    - criterion: BCEWithLogitsLoss\n    - optimizer: Adam (lr: 1.0e-4)\n    - scheduler: OneCycleLR (pct_start: 0.1, div_factor: 1.0e+3, max_lr: 1.0e-4)\n    - epoch: 5\n    - batch_size: 8 (and accumulate_grad_batches=4, so 8*4=32)\n    - fp16 (training and inference)\n    - aux loss\n        - aux target: site_id, laterality, view, implant, biopsy, invasive, BIRADS, density, difficult_negative_case, age\n        - aux loss weight: 1.0\n        - From the discussion ([https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/370341](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/370341)), there was information that there were machine_id in the LB that were not in the training data, so the machine_id was not used.\n        - However, there is a feeling that a model specific to this competition has been created by using site_id. It is questionable whether this model can be applied to mammography images from other facilities…\n- Inference\n    - top3 mean aggregation（Average of the top 3 predictions per prediction id）\n    - To reduce inference time, when doing the ensemble, the 4 models from the 4-fold cross-validation were not used for inference, but a single model trained on all data was used.\n    (cross-validation was used to check CV scores and to find thresholds.)\n    - The ensemble is a simple average; I also tried optimising the weights using the Nelder-Mead method, but as the weights were calculated almost identically across all models, I opted for a simple average.\n- Results\n\n| Model | AUROC | local cv pF1  | public pF1 | private pF1 |\n| --- | --- | --- | --- | --- |\n| 1. tf_efficientnetv2_s (no aux loss) | 0.897 | 0.424 | 0.48※ | 0.41※ |\n| 2. tf_efficientnetv2_s (using aux loss) | 0.894 | 0.435 | 0.48※ | 0.45※ |\n| 3. maxvit_tiny_tf_384.in1k (using aux loss) | 0.896 | 0.443 | Not submitted | Not submitted |\n| ensemble (1+2+3 prob mean) | 0.917 | 0.482 | 0.54 | 0.49 |\n\n※ simple mean aggregation (not top3 mean aggregation)\n\n## What Worked\n\n- Large image size\n- Strong dropout (dropout rate=0.8)\n- Stochastic depth rate (timm's drop_path_rate)\n- Strong augmentation\n    - By default, the border_mode of ShiftScaleRotate in Albumentations is cv2.BORDER_REFLECT_101, which produced better results than cv2.BORDER_CONSTANT.\n    - If border_mode = cv2.BORDER_REFLECT_101, the overflowing image is reflected by the affine transformation and a weird image is generated. ↓![reflect2](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9058219%2F3bc241b5353f842f03bdecab1625886a%2Freflect2.png?generation=1677661658109905&alt=media)\n    - However, as the results were better when the image was reflected, I considered that the task of this competition was more about texture than about the shape of the image.\n- Dealing with imbalances in mini-batches.\n- ROI extraction\n- Aux Loss\n- VOI-LUT\n- Anomaly image processing\n    - Some images were black and full of noise, which were removed. (reference: [https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/373208](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/373208))\n- Ensemble\n\n## What Didn’t Worked\n\n- Dealing with imbalances by loss functions\n    - Focal Loss, Weighted Cross Entropy\n- Metric learning\n- Center-crop to maintain aspect ratio\n- TTA\n    - Not adopted as results did not change much and inference time increased\n- Large parameter models\n    - Only tried efficientnetv2_l, but it is only slightly better than s. Not adopted due to increased inference time.\n- Some data extensions\n    - CutMix, MixUp, Grid distortion, elastic transformation, CLAHE\n\n## Code\n\n- GitHub -> https://github.com/moritake04/rsna-breast-cancer-detection\n- inference notebook -> https://www.kaggle.com/code/moritake04/private10th-final-sub\n\n(Sorry for the dirty code.) \n\n## Thanks and Acknowledgements\n\nAgain, thank you to the organizers and participants of this competition, and I will continue to work hard to become a kaggle master. ",
      "votes": 28
    },
    {
      "id": 2164212,
      "postDate": "2023-03-01T12:13:33.973Z",
      "content": "<blockquote>\n  <p>However, as the results were better when the image was reflected, I considered that the task of this competition was more about texture than about the shape of the image.</p>\n</blockquote>\n<p>Good intuition. Multi-view/laterality and multi-task might deal more with shape. Congratulations, it was highly interesting to read what you had done. </p>",
      "rawMarkdown": ">However, as the results were better when the image was reflected, I considered that the task of this competition was more about texture than about the shape of the image.\n\nGood intuition. Multi-view/laterality and multi-task might deal more with shape. Congratulations, it was highly interesting to read what you had done. ",
      "votes": 1,
      "replies": [
        {
          "id": 2164303,
          "postDate": "2023-03-01T13:15:58.270Z",
          "content": "<p>Thank you!<br>\nCertainly, shape information may be important for multi-input approaches such as Multi-view/laterality, and shape information may be used to improve results.<br>\nI had an idea for multi-input, but lacked the implementation time and computing resources😅</p>",
          "rawMarkdown": "Thank you!\nCertainly, shape information may be important for multi-input approaches such as Multi-view/laterality, and shape information may be used to improve results.\nI had an idea for multi-input, but lacked the implementation time and computing resources😅",
          "votes": 1
        }
      ]
    },
    {
      "id": 2164095,
      "postDate": "2023-03-01T09:56:47.483Z",
      "content": "<p>Congrats! Wow drop_rate=0.8 😳, do you have the specific improvement in CV between drop_rate=0.8 and let's say drop_rate=0.2 ?</p>",
      "rawMarkdown": "Congrats! Wow drop_rate=0.8 😳, do you have the specific improvement in CV between drop_rate=0.8 and let's say drop_rate=0.2 ?",
      "votes": 1,
      "replies": [
        {
          "id": 2164133,
          "postDate": "2023-03-01T10:25:23.453Z",
          "content": "<p>Thank you!<br>\nI did a grid search with drop_rate of 0.2, 0.5 and 0.8 and the higher the drop_rate the better the auc and pf1.</p>\n<p>Here is a graph of my mid-experiment (yellow: 0.8, grey: 0.5, blue: 0.2).<br>\nThis is the result for only one fold during the 4-fold cross-validation, with a difference of about 0.05 between drop_rate = 0.2 and 0.8 in pf1 (per image).<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9058219%2Fc291d279c3d26c8eda338b0e6d876daf%2F2023-03-01%20192035.png?generation=1677666089756731&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9058219%2F18dbadb9e5900e5d300829e0f8f5e9a2%2F2023-03-01%20192054.png?generation=1677666104255578&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "Thank you!\nI did a grid search with drop_rate of 0.2, 0.5 and 0.8 and the higher the drop_rate the better the auc and pf1.\n\nHere is a graph of my mid-experiment (yellow: 0.8, grey: 0.5, blue: 0.2).\nThis is the result for only one fold during the 4-fold cross-validation, with a difference of about 0.05 between drop_rate = 0.2 and 0.8 in pf1 (per image).\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9058219%2Fc291d279c3d26c8eda338b0e6d876daf%2F2023-03-01%20192035.png?generation=1677666089756731&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9058219%2F18dbadb9e5900e5d300829e0f8f5e9a2%2F2023-03-01%20192054.png?generation=1677666104255578&alt=media)",
          "votes": 7,
          "replies": [
            {
              "id": 2164151,
              "postDate": "2023-03-01T10:43:36.617Z",
              "content": "<p>Amazing thank you!</p>",
              "rawMarkdown": "Amazing thank you!",
              "votes": 1
            },
            {
              "id": 2165305,
              "postDate": "2023-03-02T05:17:48.437Z",
              "content": "<p>Thanks.  Big improvement with 0.8.  I only used 0.2.  Thanks for the write up.  </p>",
              "rawMarkdown": "Thanks.  Big improvement with 0.8.  I only used 0.2.  Thanks for the write up.  ",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3239881,
      "postDate": "2025-07-03T08:46:40.183Z",
      "content": "<p>I don't know if my experiments were too few and Focal Loss hasn't been effective, or if downsampling and fixed positive and negative sample ratios are more effective</p>",
      "rawMarkdown": "I don't know if my experiments were too few and Focal Loss hasn't been effective, or if downsampling and fixed positive and negative sample ratios are more effective"
    },
    {
      "id": 2250858,
      "postDate": "2023-05-08T20:57:20.917Z",
      "content": "<p>I met ERROR: torch-1.12.1+cu113-cp37-cp37m-linux_x86_64.whl is not a supported wheel on this platform in your project.</p>",
      "rawMarkdown": "I met ERROR: torch-1.12.1+cu113-cp37-cp37m-linux_x86_64.whl is not a supported wheel on this platform in your project.\n\n ",
      "replies": [
        {
          "id": 2253628,
          "postDate": "2023-05-10T10:56:19.557Z",
          "content": "<p>Maybe this is because the python version of the kaggle notebook has recently been upgraded from 3.7 to 3.10.<br>\nIf you set the python version of the notebook to 3.7, you should be able to run it. (Sorry, I don't know the specific solution right now…)</p>",
          "rawMarkdown": "Maybe this is because the python version of the kaggle notebook has recently been upgraded from 3.7 to 3.10.\nIf you set the python version of the notebook to 3.7, you should be able to run it. (Sorry, I don't know the specific solution right now...)",
          "votes": 1,
          "replies": [
            {
              "id": 2274031,
              "postDate": "2023-05-25T15:05:20.410Z",
              "content": "<p>Thank you!</p>",
              "rawMarkdown": "Thank you!"
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2164212,
      "author_name": "Antti Isosalo",
      "author_url": "",
      "post_date": "2023-03-01T12:13:33.973000",
      "content": "<blockquote>\n  <p>However, as the results were better when the image was reflected, I considered that the task of this competition was more about texture than about the shape of the image.</p>\n</blockquote>\n<p>Good intuition. Multi-view/laterality and multi-task might deal more with shape. Congratulations, it was highly interesting to read what you had done. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2164303,
          "author_name": "moritake04",
          "author_url": "",
          "post_date": "2023-03-01T13:15:58.270000",
          "content": "<p>Thank you!<br>\nCertainly, shape information may be important for multi-input approaches such as Multi-view/laterality, and shape information may be used to improve results.<br>\nI had an idea for multi-input, but lacked the implementation time and computing resources😅</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2164095,
      "author_name": "Optimo",
      "author_url": "",
      "post_date": "2023-03-01T09:56:47.483000",
      "content": "<p>Congrats! Wow drop_rate=0.8 😳, do you have the specific improvement in CV between drop_rate=0.8 and let's say drop_rate=0.2 ?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2164133,
          "author_name": "moritake04",
          "author_url": "",
          "post_date": "2023-03-01T10:25:23.453000",
          "content": "<p>Thank you!<br>\nI did a grid search with drop_rate of 0.2, 0.5 and 0.8 and the higher the drop_rate the better the auc and pf1.</p>\n<p>Here is a graph of my mid-experiment (yellow: 0.8, grey: 0.5, blue: 0.2).<br>\nThis is the result for only one fold during the 4-fold cross-validation, with a difference of about 0.05 between drop_rate = 0.2 and 0.8 in pf1 (per image).<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9058219%2Fc291d279c3d26c8eda338b0e6d876daf%2F2023-03-01%20192035.png?generation=1677666089756731&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9058219%2F18dbadb9e5900e5d300829e0f8f5e9a2%2F2023-03-01%20192054.png?generation=1677666104255578&amp;alt=media\" alt=\"\"></p>",
          "votes": 7,
          "replies": [
            {
              "id": 2164151,
              "author_name": "Optimo",
              "author_url": "",
              "post_date": "2023-03-01T10:43:36.617000",
              "content": "<p>Amazing thank you!</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2165305,
              "author_name": "joejeo1",
              "author_url": "",
              "post_date": "2023-03-02T05:17:48.437000",
              "content": "<p>Thanks.  Big improvement with 0.8.  I only used 0.2.  Thanks for the write up.  </p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3239881,
      "author_name": "TinyFavor",
      "author_url": "",
      "post_date": "2025-07-03T08:46:40.183000",
      "content": "<p>I don't know if my experiments were too few and Focal Loss hasn't been effective, or if downsampling and fixed positive and negative sample ratios are more effective</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2250858,
      "author_name": "Dabone. Souleymane.",
      "author_url": "",
      "post_date": "2023-05-08T20:57:20.917000",
      "content": "<p>I met ERROR: torch-1.12.1+cu113-cp37-cp37m-linux_x86_64.whl is not a supported wheel on this platform in your project.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2253628,
          "author_name": "moritake04",
          "author_url": "",
          "post_date": "2023-05-10T10:56:19.557000",
          "content": "<p>Maybe this is because the python version of the kaggle notebook has recently been upgraded from 3.7 to 3.10.<br>\nIf you set the python version of the notebook to 3.7, you should be able to run it. (Sorry, I don't know the specific solution right now…)</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2274031,
              "author_name": "Dabone. Souleymane.",
              "author_url": "",
              "post_date": "2023-05-25T15:05:20.410000",
              "content": "<p>Thank you!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2164086": "First, I would like to thank the competition hosts for organizing this competition and also the participants.\n\nI’m very happy to have won my first gold medal at Kaggle!\n\n## Overview\n\nMy model is a very simple single input model, with a simple pipeline. (So I didn't expect to win a gold medal.)\n\nMy final model is a simple average ensemble of the following 3 models. (model names are from timm)\n- tf_efficientnetv2_s (no aux loss)\n- tf_efficientnetv2_s (using aux loss)\n- maxvit_tiny_tf_384.in1k (using aux loss)\n\n![pipeline](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9058219%2F6c141d2897b20fac001918828caaef9b%2Fpipeline.png?generation=1677661238257982&alt=media)\n\nI built the Baseline myself, but it is similar to @vslaykovsky's (https://www.kaggle.com/code/vslaykovsky/train-pytorch-aux-targets-weighted-loss-thres) and I actually used it as a reference.\n\n## Solution\n\n- CV Strategy\n    - StratifiedGroupKFold (n=4, groups=patient_id)\n- Data・Preprocessing\n    - No external data was used.\n    - ROI extraction was performed using rule-based method (reference: [https://www.kaggle.com/code/vslaykovsky/rsna-cut-off-empty-space-from-images](https://www.kaggle.com/code/vslaykovsky/rsna-cut-off-empty-space-from-images)).\n    - Fast loading was achieved using nvJPEG2000 (reference: [https://www.kaggle.com/code/snaker/easy-load-the-image-with-nvjpeg2000](https://www.kaggle.com/code/snaker/easy-load-the-image-with-nvjpeg2000)).\n    - The resolution was set to 1536x960 for efficientnet and 1536x768 for maxvit.\n    - The channel number was set to 3 (to use pretrained models).\n    - Min-max scaling (-1.0 ~ 1.0) was applied.\n    - VOI LUT was applied.\n    - 8 bit\n    - Care was taken to ensure that the preprocessing pipeline remains unchanged between training and inference.\n- Dealing with Imbalance Data\n    - Batch size of 8. Adjusted to have a majority (Not Cancer) to minority (Cancer) ratio of 7:1 for each batch. This is essentially oversampling.\n- Data Augmentation\n    - I used albumentations (a data augmentation library).\n\n```python\nimport albumentations as A\n\nA.HorizontalFlip(p=0.5)\nA.VerticalFlip(p=0.5)\nA.ShiftScaleRotate(shift_limit=0.1, scale_limit=0.2, rotate_limit=45, p=0.8)\nA.OneOf([\n    A.RandomGamma(gamma_limit=(50, 150), p=0.5),\n    A.RandomBrightnessContrast(brightness_limit=0.5, contrast_limit=0.5, p=0.5)\n], p=0.5)\nA.CoarseDropout(max_height=8, max_width=8, p=0.5)\n```\n\n- Model Parameters\n    - drop_rate: 0.8\n    - drop_path_rate: 0.2\n    - criterion: BCEWithLogitsLoss\n    - optimizer: Adam (lr: 1.0e-4)\n    - scheduler: OneCycleLR (pct_start: 0.1, div_factor: 1.0e+3, max_lr: 1.0e-4)\n    - epoch: 5\n    - batch_size: 8 (and accumulate_grad_batches=4, so 8*4=32)\n    - fp16 (training and inference)\n    - aux loss\n        - aux target: site_id, laterality, view, implant, biopsy, invasive, BIRADS, density, difficult_negative_case, age\n        - aux loss weight: 1.0\n        - From the discussion ([https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/370341](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/370341)), there was information that there were machine_id in the LB that were not in the training data, so the machine_id was not used.\n        - However, there is a feeling that a model specific to this competition has been created by using site_id. It is questionable whether this model can be applied to mammography images from other facilities…\n- Inference\n    - top3 mean aggregation（Average of the top 3 predictions per prediction id）\n    - To reduce inference time, when doing the ensemble, the 4 models from the 4-fold cross-validation were not used for inference, but a single model trained on all data was used.\n    (cross-validation was used to check CV scores and to find thresholds.)\n    - The ensemble is a simple average; I also tried optimising the weights using the Nelder-Mead method, but as the weights were calculated almost identically across all models, I opted for a simple average.\n- Results\n\n| Model | AUROC | local cv pF1  | public pF1 | private pF1 |\n| --- | --- | --- | --- | --- |\n| 1. tf_efficientnetv2_s (no aux loss) | 0.897 | 0.424 | 0.48※ | 0.41※ |\n| 2. tf_efficientnetv2_s (using aux loss) | 0.894 | 0.435 | 0.48※ | 0.45※ |\n| 3. maxvit_tiny_tf_384.in1k (using aux loss) | 0.896 | 0.443 | Not submitted | Not submitted |\n| ensemble (1+2+3 prob mean) | 0.917 | 0.482 | 0.54 | 0.49 |\n\n※ simple mean aggregation (not top3 mean aggregation)\n\n## What Worked\n\n- Large image size\n- Strong dropout (dropout rate=0.8)\n- Stochastic depth rate (timm's drop_path_rate)\n- Strong augmentation\n    - By default, the border_mode of ShiftScaleRotate in Albumentations is cv2.BORDER_REFLECT_101, which produced better results than cv2.BORDER_CONSTANT.\n    - If border_mode = cv2.BORDER_REFLECT_101, the overflowing image is reflected by the affine transformation and a weird image is generated. ↓![reflect2](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9058219%2F3bc241b5353f842f03bdecab1625886a%2Freflect2.png?generation=1677661658109905&alt=media)\n    - However, as the results were better when the image was reflected, I considered that the task of this competition was more about texture than about the shape of the image.\n- Dealing with imbalances in mini-batches.\n- ROI extraction\n- Aux Loss\n- VOI-LUT\n- Anomaly image processing\n    - Some images were black and full of noise, which were removed. (reference: [https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/373208](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/373208))\n- Ensemble\n\n## What Didn’t Worked\n\n- Dealing with imbalances by loss functions\n    - Focal Loss, Weighted Cross Entropy\n- Metric learning\n- Center-crop to maintain aspect ratio\n- TTA\n    - Not adopted as results did not change much and inference time increased\n- Large parameter models\n    - Only tried efficientnetv2_l, but it is only slightly better than s. Not adopted due to increased inference time.\n- Some data extensions\n    - CutMix, MixUp, Grid distortion, elastic transformation, CLAHE\n\n## Code\n\n- GitHub -> https://github.com/moritake04/rsna-breast-cancer-detection\n- inference notebook -> https://www.kaggle.com/code/moritake04/private10th-final-sub\n\n(Sorry for the dirty code.) \n\n## Thanks and Acknowledgements\n\nAgain, thank you to the organizers and participants of this competition, and I will continue to work hard to become a kaggle master. ",
    "2164212": ">However, as the results were better when the image was reflected, I considered that the task of this competition was more about texture than about the shape of the image.\n\nGood intuition. Multi-view/laterality and multi-task might deal more with shape. Congratulations, it was highly interesting to read what you had done. ",
    "2164095": "Congrats! Wow drop_rate=0.8 😳, do you have the specific improvement in CV between drop_rate=0.8 and let's say drop_rate=0.2 ?",
    "3239881": "I don't know if my experiments were too few and Focal Loss hasn't been effective, or if downsampling and fixed positive and negative sample ratios are more effective",
    "2250858": "I met ERROR: torch-1.12.1+cu113-cp37-cp37m-linux_x86_64.whl is not a supported wheel on this platform in your project.\n\n "
  }
}