{
  "id": 392449,
  "title": "1st place solution",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/392449",
  "author_name": "Đăng Nguyễn Hồng",
  "post_date": "2023-03-05T10:56:38.190000",
  "votes": 178,
  "comment_count": 51,
  "views": 0,
  "content": "<p><strong><em>Update 10/04/2023</em></strong></p>\n<p>Ablation study: despite of a large improvement on a particular fold (below), soft positive label trick does not show a clearly improvement in performance over the standard label smoothing technique. The use of external datasets improve F1-score about 0.02 on local OOF validation and Private Leaderboard (with extractly same training pipeline + hyper-params).</p>\n<table>\n<thead>\n<tr>\n<th>External data</th>\n<th>Loss</th>\n<th>OOF F1</th>\n<th>LB</th>\n<th>PL</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>x</td>\n<td>label smoothing=0.1</td>\n<td>0.4921</td>\n<td>0.60</td>\n<td>0.53</td>\n</tr>\n<tr>\n<td>x</td>\n<td>soft positive label=0.8</td>\n<td>0.4853</td>\n<td>0.60</td>\n<td>0.53</td>\n</tr>\n<tr>\n<td>✓</td>\n<td>label smoothing = 0.1</td>\n<td>0.5161</td>\n<td>0.58</td>\n<td><strong>0.56</strong></td>\n</tr>\n<tr>\n<td>✓</td>\n<td>soft positive label = 0.9</td>\n<td><strong>0.5182</strong></td>\n<td><strong>0.61</strong></td>\n<td>0.55</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<p>First of all, I would like to thank Kaggle and competition's host for such an amazing challenge, a lofty goal with high data quality. Thank you to all paticipants/kagglers with many active and helpful discussions/codes. My solution was just built up from every pieces of kindly shares from you. I learned a lot and I'm very appreciated for that.<br>\nI'm also very happy and suprised with the 1st place. This is my first gold medal and I'm writing my first writeup. It was such a great journey for me.<br>\nFor the solution, I use a very simple pipeline which can be described in just few lines:</p>\n<ul>\n<li>Use some external datasets: VinDr-Mammo, MiniDDSM, CMMD, CDD-CESM, BMCD.</li>\n<li>4 x Convnextv1-small 2048x1024, validated on 4-folds splits of competition data.</li>\n<li>Soft positive label</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10254700%2F15d5d89bcd9cd18e45b6da133d5f5ad7%2Fpreprocess.jpg?generation=1685726162522472&amp;alt=media\" alt=\"inference pipeline\"></p>\n<p>Now I want to share some experiments and my thought about those. Many of theme could be found in another discussions by excellent kagglers. Many of theme seem obvious. Hope this helps some new comer getting started in the future. Kindly note that it's just my own opinion/thoughts with very limited experiments and knownledge. I'm appreciated for your discussions and feel free to correct me if something was wrong.</p>\n<h1>1. ROI crop</h1>\n<p>ROI cropping was performed since it effectively help keeping more texture/detail given a fixed resolution. I use YOLOX-nano 416x416 for ROI detector. The advantage of DL detector vs rule-based methods is the obtained bbox is smaller, aspect ratio is more stable and focus to the breast region.</p>\n<ol>\n<li>Train a YOLOX on <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> 's <a href=\"https://www.kaggle.com/datasets/remekkinas/rsna-roi-detector-annotations-yolo\" target=\"_blank\">dataset</a> (472 bbox-annotated images)</li>\n<li>Inference on all available training data with low <code>conf_thres</code> and high <code>iou_thres</code>. Only 3 miss-detected images (all contain noise) and over 100 images with 2 boxes (almost overlapped). I manually select and label 99 of those images. Therefore, I have 571 annotated images in total.</li>\n<li>Retrain YOLOX on new images: 521 for train, 50 for val. Note that these 50 val images include all 47 val images of original <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> 's dataset. The new dataset version contain original resolution images (same as original dicoms), preprocessed with simple min-max normalization. Tried various model/image sizes and finally choose YOLOX-nano 416x416 as final model due to the consistent result and small overhead.</li>\n</ol>\n<table>\n<thead>\n<tr>\n<th><strong>model size</strong></th>\n<th><strong>image size</strong></th>\n<th><strong>interpolation</strong></th>\n<th><strong>AP_new_val</strong></th>\n<th><strong>AP_remek_val</strong></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><em>nano (selection)</em></td>\n<td><em>416</em></td>\n<td><em>LINEAR</em></td>\n<td><em>96.26</em></td>\n<td><em>94.21</em></td>\n</tr>\n<tr>\n<td>nano</td>\n<td>416</td>\n<td>AREA</td>\n<td>94.09</td>\n<td>91.60</td>\n</tr>\n<tr>\n<td>nano</td>\n<td>640</td>\n<td>LINEAR</td>\n<td>95.85</td>\n<td>88.40</td>\n</tr>\n<tr>\n<td>nano</td>\n<td>768</td>\n<td>LINEAR</td>\n<td>96.22</td>\n<td>82.09</td>\n</tr>\n<tr>\n<td>nano</td>\n<td>1024</td>\n<td>LINEAR</td>\n<td>94.92</td>\n<td>89.40</td>\n</tr>\n<tr>\n<td>tiny</td>\n<td>416</td>\n<td>LINEAR</td>\n<td>94.23</td>\n<td>90.20</td>\n</tr>\n<tr>\n<td>tiny</td>\n<td>640</td>\n<td>LINEAR</td>\n<td>94.95</td>\n<td>89.84</td>\n</tr>\n<tr>\n<td>tiny</td>\n<td>768</td>\n<td>AREA</td>\n<td>96.21</td>\n<td>68.03</td>\n</tr>\n<tr>\n<td>tiny</td>\n<td>1024</td>\n<td>AREA</td>\n<td>93.69</td>\n<td>73.70</td>\n</tr>\n<tr>\n<td>s</td>\n<td>416</td>\n<td>LINEAR</td>\n<td>95.03</td>\n<td>0.86</td>\n</tr>\n<tr>\n<td>s</td>\n<td>640</td>\n<td>LINEAR</td>\n<td>96.10</td>\n<td>70.80</td>\n</tr>\n<tr>\n<td>s</td>\n<td>768</td>\n<td>LINEAR</td>\n<td>96.79</td>\n<td>78.70</td>\n</tr>\n</tbody>\n</table>\n<p>AP@0.5 is 1.0 in all experiments. We see a large gap between AP@0.5-0.95 between two validation sets. Some reasons for that:</p>\n<ul>\n<li>New version add more 3/50 typical hard cases.</li>\n<li>Inconsistent processing pipeline: val images in Remek's val was resized 2 times (original --&gt; 1024 --&gt; 416) </li>\n<li>Training images is annotated according to personal bias (no standard way/consentration to annotate the breast boxes correctly). So higher AP may not indicate a better model.</li>\n<li>The validation size is also not large enough to judge</li>\n<li>No hyper parameters tuning</li>\n</ul>\n<p>Did these things led to the large gap, particularly with stronger model and larger image size ?<br>\nAll these efforts are just to ensure an \"as good as posible\" ROI detection model. I think <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> 's <a href=\"https://www.kaggle.com/datasets/remekkinas/rsna-roi-detector-annotations-yolo\" target=\"_blank\">dataset</a> is enough to train good YOLOX models and they could perform equally well in hidden test set.</p>\n<p>Simpler Otsu thresholding + findCountours() slightly modified from <a href=\"https://www.kaggle.com/code/snnclsr/roi-extraction-using-opencv\" target=\"_blank\">this notebook</a> is used to find breast bbox as a fall back in case of YOLOX's miss-detection.<br>\nOr, if both miss the breast box, just use the whole image without any cropping.</p>\n<h1>2. The inference pipeline</h1>\n<p>Operations on large array take time, so I try to transfer the computation task to GPU as much as posible.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10254700%2F9de52f41201b4a92774f5fee331cb879%2Fwriteup-Page-1.drawio.png?generation=1678001361222187&amp;alt=media\" alt=\"\"></p>\n<h1>3. Early experiments</h1>\n<p>My final solution use external datasets, but I stitch up with using only competition data for almost of the time (until \"7 days to go\"). Hence most of my experiments are done on competition data only: 5-folds splits with StratifiedGroupKFold based on <code>patient_id</code>. Training hyperparams used for the final solution are almost inherited from these early experiments.</p>\n<h2>3.1. About the metric</h2>\n<p>The competition pF1 score is not stable and hard to track for me. Therefore, I mainly track my experiments based on multiple metrics: <code>{ PR_AUC, ROC_AUC, best_PF1 (binarized), best_threshold }</code> instead of just one.</p>\n<ul>\n<li><strong>PR_AUC</strong>: correlated with but more stable than best_PF1. It focuses on positive cases, and is strongly affected by prior data distribution (% of positive).</li>\n<li><strong>ROC_AUC</strong>: less affected by prior data distribution. Much more stable, but seem to be over optimistic which led to just a small gap between a good model and a bad model.</li>\n<li>To get a high binaried pf1, model should not predict too many positives which usually led to large FP --&gt; dramatically reduce best_PF1. A good scored pf1 model tends to prioritize Precision over Recall. I personaly don't like this behaviour, especially for real life application.</li>\n</ul>\n<h2>3.2. Augmentations</h2>\n<p>I stitch with this augmentation pipeline for all experiments, no tuning at all:</p>\n<pre><code>A.Compose([\n    \n    custom_augs.CustomRandomSizedCropNoResize(scale=(, ), ratio=(, ), p=),\n    \n    A.HorizontalFlip(p=),\n    A.VerticalFlip(p=),\n    \n    A.OneOf([\n        A.Downscale(scale_min=, scale_max=, interpolation=(upscale=cv2.INTER_LINEAR, downscale=cv2.INTER_AREA), p=),\n        A.Downscale(scale_min=, scale_max=, interpolation=(upscale=cv2.INTER_LANCZOS4, downscale=cv2.INTER_AREA), p=),\n        A.Downscale(scale_min=, scale_max=, interpolation=(upscale=cv2.INTER_LINEAR, downscale=cv2.INTER_LINEAR), p=),\n    ], p=),\n    \n    A.OneOf([\n        A.RandomToneCurve(scale=, p=),\n        A.RandomBrightnessContrast(brightness_limit=(-, ), contrast_limit=(-, ), brightness_by_max=, always_apply=, p=)\n    ], p=),\n    \n    A.OneOf(\n        [\n            A.ShiftScaleRotate(shift_limit=, scale_limit=[-, ], rotate_limit=[-, ], interpolation=cv2.INTER_LINEAR,\n                               border_mode=cv2.BORDER_CONSTANT, value=, mask_value=, shift_limit_x=[-, ],\n                               shift_limit_y=[-, ], rotate_method=, p=),\n            A.ElasticTransform(alpha=, sigma=, alpha_affine=, interpolation=cv2.INTER_LINEAR, border_mode=cv2.BORDER_CONSTANT,\n                               value=, mask_value=, approximate=, same_dxdy=, p=),\n            A.GridDistortion(num_steps=, distort_limit=, interpolation=cv2.INTER_LINEAR, border_mode=cv2.BORDER_CONSTANT,\n                             value=, mask_value=, normalized=, p=),\n        ], p=),\n    \n    A.CoarseDropout(max_holes=, max_height=, max_width=, min_holes=, min_height=, min_width=,\n                    fill_value=, mask_fill_value=, p=),\n    ], p=)\n</code></pre>\n<p>For the random crop choice: real breast size/ratio vary largly between images --&gt; popular pipeline of <code>longest resize + padding</code> introduces multi-scales problem. Of course, it would introduce higher risks of wrong positive label.</p>\n<p><em>Example batch</em><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10254700%2F4f149f8617c07c9a8e0d8e571f6a68fa%2Fexample_batch.jpg?generation=1685726233222765&amp;alt=media\" alt=\"example batch\"></p>\n<h2>3.3. Up/down sampling</h2>\n<p>I upsample pos cases in each epoch for all of my experiments.</p>\n<ul>\n<li>Ensuring at least 1 pos in a batch/iteration is pretty important to stablize training. I found training difficult while set 0.5 pos per batch.</li>\n<li>Upsampling ratio can largly affect CV score and prediction distribution. It also vary between backbones and other hyperparams choices. However, I prefer smallest pos/neg ratio as posible (since it's near to the real data distribution) but ensure at least 1 pos/batch.</li>\n<li>Large pos/neg ratio helps training faster in early epochs. I tried linearly increase/decrease the pos/neg ratio between epochs to face with some problems of prediction distribution/threshold (especially EffB4). But in the end, i got no improvement in CV.</li>\n</ul>\n<h2>3.4. Model/backbone</h2>\n<p>I tried <code>Eff-B2</code>, <code>Eff-B4</code>, <code>Effv2-s</code> and <code>Convnextv1-small</code></p>\n<ul>\n<li>Each model has its own characteristic and training phenomenon.</li>\n<li>All models could perform equally well in local CV. Except that <code>Convnextv1-small</code> give higher CV score.</li>\n<li><code>EfficientNet</code> (no model EMA) tends to overfit quickly after fews epochs with high pos/neg ratio: longer training reduce AUC largely and may slightly increase best_pf1 --&gt; model tends to predict less positives. Small pos/neg ratio helps training more stable but reduce CV. Linearly increase pos/neg ratio between epochs (by a sampler) did not help much.</li>\n<li><code>Convnext-small</code> (with/without EMA) shows both stable training and better CV.</li>\n</ul>\n<h2>3.5. drop_rate, drop_path_rate</h2>\n<p>Playing with drop_rate and drop_path_rate:</p>\n<ul>\n<li>We can use large dropout rate of &gt;= 0.5 to regularize training and reduce overfiting.</li>\n<li>With very large drop_rate = 0.9 or drop_path_rate = 0.5, I still can get a \"not bad as expected\" model in CV score. The following results is for Eff-B4, pos/neg = 1/3 on fold 0:</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>drop_rate</th>\n<th>drop_path_rate</th>\n<th>auc</th>\n<th>best_pf1</th>\n<th>best_thres</th>\n<th>epoch</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.9</td>\n<td>0.2</td>\n<td>91.90</td>\n<td>47.73</td>\n<td>0.78</td>\n<td>4</td>\n</tr>\n<tr>\n<td>0.7</td>\n<td>0.2</td>\n<td>90.36</td>\n<td>52.27</td>\n<td>0.55</td>\n<td>3</td>\n</tr>\n<tr>\n<td>0.5</td>\n<td>0.2</td>\n<td>90.02</td>\n<td>50.00</td>\n<td>0.72</td>\n<td>4</td>\n</tr>\n<tr>\n<td>0.3</td>\n<td>0.2</td>\n<td>91.23</td>\n<td>48.24</td>\n<td>0.82</td>\n<td>4</td>\n</tr>\n<tr>\n<td>0.5</td>\n<td>0.5</td>\n<td>90.45</td>\n<td>46.33</td>\n<td>0.95</td>\n<td>4</td>\n</tr>\n</tbody>\n</table>\n<p>However, I set drop_rate = 0.5 and drop_path_rate = 0.2 for most of my experiments, including the final ones.</p>\n<h2>3.6. Global pooling</h2>\n<p>I stitch with max pooling for almost my experiments as my inductive bias:</p>\n<ul>\n<li>Max pooling is suitable and seem to be effective for anomalies detection or \"needle in the haystack\" tasks in literature. In this case, cancer may appear in a very small region and the rest are all normal.</li>\n<li><code>max()</code> provides stronger learning signal, but less stable than mean() in term of gradient.</li>\n<li>My guess: <code>gem</code> &gt; <code>max</code> &gt; <code>mean</code> when enough data provided.</li>\n</ul>\n<h2>3.7. Soft positive label/ Positive label smoothing</h2>\n<p>Convnextv1-small look good in CV scores with stable AUC, PR_AUC, best PF1 across epochs. But there're differences in behaviour between Effv2-s and Convnextv1-small especially in best threshold for image/breast level.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10254700%2F6771d0c6e87067873e53021742e7bf6b%2Fwriteup_smoothing.png?generation=1678006236367845&amp;alt=media\" alt=\"\"><br>\n<em>The two lower green ones belong to <code>Effv2-s</code> and the other belong to <code>Convnextv1-small</code></em></p>\n<p>Some discussions suggest smaller best threshold (&lt;0.55) may indicate a better model. For single-image, Convnext show a very high threshold of &gt; 0.92, which could indicate the problem of over-confident. Stronger models with larger number of parameters is easier to be over-confident or overfitted, especialy in this highly imballanced dataset scenario. About the above figure, <code>label_smoothing = 0.1</code> was used but seem like it was not enough.<br>\nSo, just add harder label smoothing to regularize training. Or use positive weight &lt; 1.0 to reduce the priority of positive samples.</p>\n<table>\n<thead>\n<tr>\n<th>loss</th>\n<th>num_logits</th>\n<th>target {neg, pos}</th>\n<th>pr_auc</th>\n<th>roc_auc</th>\n<th>best_pf1</th>\n<th>best_thres</th>\n<th>epoch</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><em>(baseline) bce_smooth 0.1</em></td>\n<td>2</td>\n<td>{ [0.95, 0.05], [0.05, 0.95] }</td>\n<td>0.4755</td>\n<td>0.9278</td>\n<td>0.497</td>\n<td>0.66</td>\n<td>14</td>\n</tr>\n<tr>\n<td>bce_smooth 0.4</td>\n<td>2</td>\n<td>{ [0.8, 0.2], [0.2, 0.8] }</td>\n<td>0.4749</td>\n<td>0.9248</td>\n<td>0.5</td>\n<td>0.6</td>\n<td>25</td>\n</tr>\n<tr>\n<td>bce_pos_smooth 0.4</td>\n<td>2</td>\n<td>{ [1.0, 0.0], [0.2, 0.8] }</td>\n<td>0.5191</td>\n<td>0.9153</td>\n<td>0.5488</td>\n<td>0.53</td>\n<td>13.5</td>\n</tr>\n<tr>\n<td><em>(best) bce_pos_smooth 0.2</em></td>\n<td>1</td>\n<td>{ 0.0, 0.8 }</td>\n<td><strong>0.5401</strong></td>\n<td>0.9281</td>\n<td><strong>0.5714</strong></td>\n<td>0.49</td>\n<td>20</td>\n</tr>\n<tr>\n<td>bce_pos_smooth 0.3</td>\n<td>1</td>\n<td>{ 0.0, 0.7 }</td>\n<td>0.522</td>\n<td><strong>0.933</strong></td>\n<td>0.517</td>\n<td>0.5</td>\n<td>17</td>\n</tr>\n<tr>\n<td>bce_smooth 0.1 + pos_weight 0.4</td>\n<td>1</td>\n<td>{ 0.05, 0.95 }</td>\n<td>0.4946</td>\n<td>0.9146</td>\n<td>0.5393</td>\n<td>0.39</td>\n<td>19</td>\n</tr>\n</tbody>\n</table>\n<p><em>Note:</em></p>\n<ul>\n<li><em>Table above is fold 0 CV results</em></li>\n<li><em><code>num_logits = 2</code> means using <code>sigmoid</code> (BCEWithLogitsLoss) for training and <code>softmax</code> for inference. Refer <a href=\"https://github.com/ultralytics/yolov5/issues/5401\" target=\"_blank\">here</a>.</em></li>\n</ul>\n<p>Soft positive labeling look reasonable: we have per-breast label and not per-image label. For some images belong to same patient, cancer signal may not appears clearly in some images, or even all images (MG is not enough to judge for cancer/non-cancer) --&gt; the positive label should not be the maximum bound value of 1.0, but less confident.</p>\n<p>Soft postive label trick improve CV and helps threshold looks much better.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10254700%2Fc36afde0a06afbeb3dbd38a76753276c%2Fpos_smooth_thres.png?generation=1678007271357249&amp;alt=media\" alt=\"\"></p>\n<p>As some discussions, very sharp prediction distribution may indicate worse result/generalization.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10254700%2F7caf77b112791cd16eb20de6f6d443bd%2Fdistribution.png?generation=1678007391776163&amp;alt=media\" alt=\"\"></p>\n<hr>\n<h1>4. Final experiments</h1>\n<h2>4.1. External datasets</h2>\n<p>A week left to the competition deadline, I was thinking about training final experiments for the final submission and should not make any mistakes or missing something. I read some discussions again and relized I was missing a big part: external data. In particular, external data contains a large number of positive cases which are valuable.<br>\nThese external datasets summary:</p>\n<table>\n<thead>\n<tr>\n<th>Dataset</th>\n<th>num_patients*</th>\n<th>num_samples*</th>\n<th>num_pos_samples*</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><a href=\"https://physionet.org/content/vindr-mammo/1.0.0/\" target=\"_blank\">VinDr-Mammo</a></td>\n<td>5000</td>\n<td>20000</td>\n<td>226 (1.13 %)</td>\n</tr>\n<tr>\n<td><a href=\"https://www.kaggle.com/datasets/cheddad/miniddsm2\" target=\"_blank\">MiniDDSM</a></td>\n<td>1952</td>\n<td>7808</td>\n<td>1480 (18.95 %)</td>\n</tr>\n<tr>\n<td><a href=\"https://wiki.cancerimagingarchive.net/pages/viewpage.action?pageId=70230508\" target=\"_blank\">CMMD</a></td>\n<td>1775</td>\n<td>5202</td>\n<td>2632 (50.6%)</td>\n</tr>\n<tr>\n<td><a href=\"https://wiki.cancerimagingarchive.net/pages/viewpage.action?pageId=109379611\" target=\"_blank\">CDD-CESM</a></td>\n<td>326</td>\n<td>1003</td>\n<td>331 (33 %)</td>\n</tr>\n<tr>\n<td><a href=\"https://zenodo.org/record/5036062\" target=\"_blank\">BMCD</a></td>\n<td>82</td>\n<td>328</td>\n<td>22 (6.71 %)</td>\n</tr>\n<tr>\n<td>All</td>\n<td>9135</td>\n<td>34341</td>\n<td>4691 (13.66 %)</td>\n</tr>\n</tbody>\n</table>\n<p><strong>*</strong> The number may not indicate original dataset characteristics, but processed data I used for this competition.</p>\n<p><strong>Some details:</strong></p>\n<ol>\n<li><p><strong>VinDr-Mammo</strong>: contains BIRADS scores for each image of 0-5. I treated BIRADS 5 as cancer (1) and all other as normal (0). With only Digital Mammograms and BIRADS categories, one can't confirm 100% if a case is cancer or not . BIRADS 4 indicate 30% chance of cancer, then I was treated it as normal. My decision only happen in just a few seconds as i could remember. Reading other posts, i'm feeling my decision is not as good, except that it helps reduce sensitivity and can \"improve\" the pF1 (I don't want to see it that way). Maybe I make a huge mistake here. Better solution is to use soft/uncertain label or pseudo labeling for these ambigous (BIRADS-4) cases instead. Some images has LUTDescriptor. The image look over-exposured when apply VOILUT (voi + windowing), so I just apply windowing on this dataset, equivalent to pydicom's <code>apply_voi_lut(prefer_lut = False)</code></p></li>\n<li><p><strong>MiniDDSM</strong>: I used all 7808 samples. Found that the status label of <code>{Cancer, Benign, Normal}</code> is per patient_id, not per laterality. So I treat a patient-laterality as cancer if and only if status == 'Cancer' and at least 1 supicious region annotation (segmentation map) for that laterality is available. End up in 1480 positives and the remaining 6318 negatives. I use 16-bits png part for the less information loss. No windowing parameters as presented. Since there're watermarks noise with very high pixel intensity in ROI crop, percentile min-max scaled  was performed instead of min-max scaled for normalization.</p></li>\n<li><p><strong>CMMD</strong>: Total of 5202 breast images belong to 1872 patient id. Note that some patient ids start with 'D2' are almost malignant and usually had label for one laterality only. For those cases, I treated the other laterality (no laterality-level label specified in csv file, but still have image) as normal (EDA from the competition data show that cancer only appears in one laterality). Original dicom images are in 8-bits depth with windowing parameters available.</p></li>\n<li><p><strong>CDD-CESM</strong>: consisting Contrast-enhanced spectral mammography (CESM) images. This dataset contain label of <code>{Normal, Malignant, Benign}</code>. I treated Malignant as cancer, Normal or Benign as normal and only use the <strong>low-energy images</strong> as it is comparable to digital mammograms (MG), or at least they look pretty similar for me. Low-energy images is in 8-bits jpeg, no windowing information.</p></li>\n<li><p><strong>BMCD</strong>: contains 100 patients (50 normal + 50 suspicious cases) with 82 biopsy-confirmed cases of <code>{'NORMAL', 'BENIGN', 'DCIS', 'MALIGNANT'}</code> and mammogram images of them at the time of screening and  avg 2.2 year before. I treat 'DCIS' or 'MALIGNANT' patient's last screening images as cancer and all the remaining as normal. Original dicom images is in 16-bits depth and windowing parameters are available.</p></li>\n</ol>\n<h2>4.2. Validation strategy</h2>\n<p>I found inconsistence in CV between 5-folds splits of competition data, probably because the number of positive is not sufficient. Although, hidden test should has distribution/property closer to the competition data, so I change validation strategy to 4-splits as:</p>\n<ul>\n<li>Do 4-folds splitting on competition data</li>\n<li>Use 1 fold for validation, the rest 3 folds + all external data for training.<br>\nThen, for each split, training data contain about 5560/75400 positive cases (~7.38 %).</li>\n</ul>\n<h2>4.3. Training</h2>\n<p>I managed to get 4 x Convnextv1-small corresponding to the above 4 splits. Some unexpected results were founded during training, so the training stages was changed and in short consist of:</p>\n<ol>\n<li>Train 2 models on fold 0 and fold 1 with <code>soft_pos_label = 0.8</code></li>\n<li>Train 2 models on fold 2 and fold 3 with <code>soft_pos_label = 0.9</code></li>\n<li>Finetune 2 models obtained from stage 1 on fold 0 and fold 1 with <code>soft_pos_label = 0.9</code></li>\n</ol>\n<p>In details, I start training on first two folds: fold 0 and fold 1 with the following config:</p>\n<ul>\n<li>Model: timm's convnext_small.fb_in22k_ft_in1k_384</li>\n<li>Input size: 2048x1024</li>\n<li>Loss: vanila BCE (no class weight)</li>\n<li>Sampler: upsampling pos samples per epoch to pos/neg = 1/7, ensure each batch contains at least 1 pos sample.</li>\n<li>Batchsize: 8</li>\n<li>Automatic Mixed Precision (AMP): enable</li>\n<li>Model EMA: enable</li>\n<li>Global pooling: max</li>\n<li>Soft positive label = 0.8</li>\n<li>Optimizer: SGD with momemtum=0.9</li>\n<li>Scheduler: Cosine lr decay(epoch = 24, lr = 1e-3, min_lr = 1e-5) + linear warmup(warmup_lr = 1e-5, warmup_epoch = 4)</li>\n<li>Drop_rate = 0.5, drop_path_rate = 0.2</li>\n</ul>\n<p>Once training finished, results are not as my expectation on fold 0:</p>\n<ul>\n<li>CV results are not good</li>\n<li>Threshold is much smaller. I expected it to be in the range [0.35, 0.5], but it's just around 0.25+-0.02</li>\n<li>Training is not converged yet. I guess that CV can be improved with more additional training epochs.</li>\n</ul>\n<p>So I start train fold 2 and fold 3 with few changes: longer training with larger learning rate and reduce soft positive label.</p>\n<ul>\n<li>Scheduler: Cosine lr decay(epoch = 30, lr = 3e-3, min_lr = 5e-5) + linear warmup(warmup_lr = 3e-5, warmup_epoch = 4)</li>\n<li>Soft positive label: 0.9</li>\n</ul>\n<p>Results on fold 2 and fold 3 seem to be better. So I decided to finetune fold 0 and fold 1 with the same value of soft_positive_label = 0.9 from the previous last checkpoints.</p>\n<h2>4.4. Checkpoints selection</h2>\n<ol>\n<li>For each fold, I manually select 3-7 best checkpoints by looking at multiple metrics.</li>\n<li>Grid search over all posible combinations of each fold's checkpoints (3x5x7x5 = 525 combinations for my case), compute Out-Of-Fold (OOF) results for each combination.</li>\n<li>Select the best combination with highest binarized pF1. The best combination has local OOF pf1 = 0.5187, while the worst with 0.4951.</li>\n</ol>\n<p>Final results:</p>\n<table>\n<thead>\n<tr>\n<th>name</th>\n<th>soft_positive_label</th>\n<th>pr_auc</th>\n<th>roc_auc</th>\n<th>best_pf1</th>\n<th>best_thres</th>\n<th>epoch</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>fold 0</td>\n<td>0.8</td>\n<td>0.3983</td>\n<td>0.9142</td>\n<td>0.4716</td>\n<td>0.25</td>\n<td>24</td>\n</tr>\n<tr>\n<td><strong>(selected)</strong> fold 0 + fine-tune</td>\n<td>0.9</td>\n<td>0.4363</td>\n<td>0.9119</td>\n<td>0.4785</td>\n<td>0.35</td>\n<td>11 (24 + 11)</td>\n</tr>\n<tr>\n<td>fold 1</td>\n<td>0.8</td>\n<td>0.5151</td>\n<td>0.9202</td>\n<td>0.5291</td>\n<td>0.34</td>\n<td>18</td>\n</tr>\n<tr>\n<td><strong>(selected)</strong> fold 1 + fine-tune</td>\n<td>0.9</td>\n<td>0.5381</td>\n<td>0.9149</td>\n<td>0.5381</td>\n<td>0.34</td>\n<td>8 (24 + 8)</td>\n</tr>\n<tr>\n<td><strong>(selected)</strong> fold 2</td>\n<td>0.9</td>\n<td>0.4946</td>\n<td>0.9234</td>\n<td>0.5185</td>\n<td>0.34</td>\n<td>26</td>\n</tr>\n<tr>\n<td><strong>(selected)</strong> fold 3</td>\n<td>0.9</td>\n<td>0.5088</td>\n<td>0.9401</td>\n<td>0.5455</td>\n<td>0.31</td>\n<td>19</td>\n</tr>\n</tbody>\n</table>\n<p>Some thoughts:</p>\n<ul>\n<li>Fold 0 result looks weird? Maybe I need more time inspecting it.</li>\n<li>With much more data especially positive ones, guesses for the best hyperparams became obsolete. Maybe I could forget the soft positive label trick and get better result?</li>\n</ul>\n<p>OOF validation was done to determine the best threshold value of 0.34</p>\n<pre><code>                        auc      @th     f1      |  prec    recall  |   sens    spec \nsingle image     [0]    0.87296 0.40000 0.41907 |   0.48365 0.37047 |   0.37047 0.99145\ngrouby mean()    [0]    0.92043 0.34000 0.51820 |   0.60989 0.45122 |   0.45122 0.99391\ngrouby max()     [0]    0.91939 0.61000 0.50913 |   0.57545 0.45732 |   0.45732 0.99289\n--------------\n\nsingle image     [1]    0.84866 0.40000 0.33649 |   0.38241 0.30120 |   0.30120 0.98881\ngrouby mean()    [1]    0.89225 0.34000 0.39587 |   0.47027 0.34252 |   0.34252 0.99139\ngrouby max()     [1]    0.88917 0.61000 0.39424 |   0.44554 0.35433 |   0.35433 0.99016\n--------------\n\nsingle image     [2]    0.89611 0.40000 0.53331 |   0.62912 0.46356 |   0.46356 0.99453\ngrouby mean()    [2]    0.94288 0.34000 0.64699 |   0.75419 0.56722 |   0.56723 0.99632\ngrouby max()     [2]    0.94329 0.61000 0.63182 |   0.71428 0.56722 |   0.56723 0.99548\n--------------\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10254700%2Feeb0cc23fad1497f249613e0d4f87307%2Ffinal_sub_plot.png?generation=1678010502203196&amp;alt=media\" alt=\"\"></p>\n<p><em>The results was generated by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> ’s <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/378521\" target=\"_blank\">script</a></em></p>\n<h2>4.5. Submission</h2>\n<p>My training progress was done in the last day of the competition.<br>\n5 final submissions include:</p>\n<ul>\n<li>1 submission of models in earlier epochs while waiting training to be done. The main purpose is to ensure the pipeline is correct and no exception occurs, which give me 0.56 LB.</li>\n<li>4 submission of same final (best) model with different threshold: 0.31, 0.34 (best oof), 0.37, 0.40 . I guessed the LB's best threshold &gt;  CV's best threshold, but I'm totally wrong (I did not probe the LB). But it was fortunate that threshold=0.31 is the one on the peak :D</li>\n</ul>\n<p>Those submissions bring me from LB 600th to LB 22nd in one day. The PL 0.55 submission was successfully finished when ~30 mins left to the deadline.</p>\n<table>\n<thead>\n<tr>\n<th>Threshold</th>\n<th>OOF</th>\n<th>LB</th>\n<th>PL</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.27 (late sub)</td>\n<td>0.4877</td>\n<td>0.60</td>\n<td>0.53</td>\n</tr>\n<tr>\n<td>0.28 (late sub)</td>\n<td>0.4917</td>\n<td>0.60</td>\n<td>0.54</td>\n</tr>\n<tr>\n<td>0.29 (late sub)</td>\n<td>0.4973</td>\n<td>0.60</td>\n<td>0.54</td>\n</tr>\n<tr>\n<td>0.30 (late sub)</td>\n<td>0.5027</td>\n<td>0.60</td>\n<td>0.55</td>\n</tr>\n<tr>\n<td><strong>(selection) 0.31</strong></td>\n<td><strong>0.5049</strong></td>\n<td><strong>0.61</strong></td>\n<td><strong>0.55</strong></td>\n</tr>\n<tr>\n<td><strong>(selection) 0.34</strong></td>\n<td><strong>0.5187</strong></td>\n<td><strong>0.58</strong></td>\n<td><strong>0.53</strong></td>\n</tr>\n<tr>\n<td>0.37</td>\n<td>0.5000</td>\n<td>0.55</td>\n<td>0.52</td>\n</tr>\n<tr>\n<td>0.40</td>\n<td>0.4896</td>\n<td>0.54</td>\n<td>0.50</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h1>5. Code</h1>\n<ul>\n<li>Submission notebook: <a href=\"https://www.kaggle.com/dangnh0611/1st-place-submission-code\" target=\"_blank\">https://www.kaggle.com/dangnh0611/1st-place-submission-code</a> </li>\n<li>Training code: <a href=\"https://github.com/dangnh0611/kaggle_rsna_breast_cancer\" target=\"_blank\">https://github.com/dangnh0611/kaggle_rsna_breast_cancer</a></li>\n</ul>\n<hr>\n<p>I just got luck with simple pipeline and simple decision. Many teams had much better models but did not select it in final, as I could see.<br>\nThanks for your attention.</p>",
  "messages": [
    {
      "id": 2169701,
      "postDate": "2023-03-05T10:56:38.190Z",
      "content": "<p><strong><em>Update 10/04/2023</em></strong></p>\n<p>Ablation study: despite of a large improvement on a particular fold (below), soft positive label trick does not show a clearly improvement in performance over the standard label smoothing technique. The use of external datasets improve F1-score about 0.02 on local OOF validation and Private Leaderboard (with extractly same training pipeline + hyper-params).</p>\n<table>\n<thead>\n<tr>\n<th>External data</th>\n<th>Loss</th>\n<th>OOF F1</th>\n<th>LB</th>\n<th>PL</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>x</td>\n<td>label smoothing=0.1</td>\n<td>0.4921</td>\n<td>0.60</td>\n<td>0.53</td>\n</tr>\n<tr>\n<td>x</td>\n<td>soft positive label=0.8</td>\n<td>0.4853</td>\n<td>0.60</td>\n<td>0.53</td>\n</tr>\n<tr>\n<td>✓</td>\n<td>label smoothing = 0.1</td>\n<td>0.5161</td>\n<td>0.58</td>\n<td><strong>0.56</strong></td>\n</tr>\n<tr>\n<td>✓</td>\n<td>soft positive label = 0.9</td>\n<td><strong>0.5182</strong></td>\n<td><strong>0.61</strong></td>\n<td>0.55</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<p>First of all, I would like to thank Kaggle and competition's host for such an amazing challenge, a lofty goal with high data quality. Thank you to all paticipants/kagglers with many active and helpful discussions/codes. My solution was just built up from every pieces of kindly shares from you. I learned a lot and I'm very appreciated for that.<br>\nI'm also very happy and suprised with the 1st place. This is my first gold medal and I'm writing my first writeup. It was such a great journey for me.<br>\nFor the solution, I use a very simple pipeline which can be described in just few lines:</p>\n<ul>\n<li>Use some external datasets: VinDr-Mammo, MiniDDSM, CMMD, CDD-CESM, BMCD.</li>\n<li>4 x Convnextv1-small 2048x1024, validated on 4-folds splits of competition data.</li>\n<li>Soft positive label</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10254700%2F15d5d89bcd9cd18e45b6da133d5f5ad7%2Fpreprocess.jpg?generation=1685726162522472&amp;alt=media\" alt=\"inference pipeline\"></p>\n<p>Now I want to share some experiments and my thought about those. Many of theme could be found in another discussions by excellent kagglers. Many of theme seem obvious. Hope this helps some new comer getting started in the future. Kindly note that it's just my own opinion/thoughts with very limited experiments and knownledge. I'm appreciated for your discussions and feel free to correct me if something was wrong.</p>\n<h1>1. ROI crop</h1>\n<p>ROI cropping was performed since it effectively help keeping more texture/detail given a fixed resolution. I use YOLOX-nano 416x416 for ROI detector. The advantage of DL detector vs rule-based methods is the obtained bbox is smaller, aspect ratio is more stable and focus to the breast region.</p>\n<ol>\n<li>Train a YOLOX on <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> 's <a href=\"https://www.kaggle.com/datasets/remekkinas/rsna-roi-detector-annotations-yolo\" target=\"_blank\">dataset</a> (472 bbox-annotated images)</li>\n<li>Inference on all available training data with low <code>conf_thres</code> and high <code>iou_thres</code>. Only 3 miss-detected images (all contain noise) and over 100 images with 2 boxes (almost overlapped). I manually select and label 99 of those images. Therefore, I have 571 annotated images in total.</li>\n<li>Retrain YOLOX on new images: 521 for train, 50 for val. Note that these 50 val images include all 47 val images of original <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> 's dataset. The new dataset version contain original resolution images (same as original dicoms), preprocessed with simple min-max normalization. Tried various model/image sizes and finally choose YOLOX-nano 416x416 as final model due to the consistent result and small overhead.</li>\n</ol>\n<table>\n<thead>\n<tr>\n<th><strong>model size</strong></th>\n<th><strong>image size</strong></th>\n<th><strong>interpolation</strong></th>\n<th><strong>AP_new_val</strong></th>\n<th><strong>AP_remek_val</strong></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><em>nano (selection)</em></td>\n<td><em>416</em></td>\n<td><em>LINEAR</em></td>\n<td><em>96.26</em></td>\n<td><em>94.21</em></td>\n</tr>\n<tr>\n<td>nano</td>\n<td>416</td>\n<td>AREA</td>\n<td>94.09</td>\n<td>91.60</td>\n</tr>\n<tr>\n<td>nano</td>\n<td>640</td>\n<td>LINEAR</td>\n<td>95.85</td>\n<td>88.40</td>\n</tr>\n<tr>\n<td>nano</td>\n<td>768</td>\n<td>LINEAR</td>\n<td>96.22</td>\n<td>82.09</td>\n</tr>\n<tr>\n<td>nano</td>\n<td>1024</td>\n<td>LINEAR</td>\n<td>94.92</td>\n<td>89.40</td>\n</tr>\n<tr>\n<td>tiny</td>\n<td>416</td>\n<td>LINEAR</td>\n<td>94.23</td>\n<td>90.20</td>\n</tr>\n<tr>\n<td>tiny</td>\n<td>640</td>\n<td>LINEAR</td>\n<td>94.95</td>\n<td>89.84</td>\n</tr>\n<tr>\n<td>tiny</td>\n<td>768</td>\n<td>AREA</td>\n<td>96.21</td>\n<td>68.03</td>\n</tr>\n<tr>\n<td>tiny</td>\n<td>1024</td>\n<td>AREA</td>\n<td>93.69</td>\n<td>73.70</td>\n</tr>\n<tr>\n<td>s</td>\n<td>416</td>\n<td>LINEAR</td>\n<td>95.03</td>\n<td>0.86</td>\n</tr>\n<tr>\n<td>s</td>\n<td>640</td>\n<td>LINEAR</td>\n<td>96.10</td>\n<td>70.80</td>\n</tr>\n<tr>\n<td>s</td>\n<td>768</td>\n<td>LINEAR</td>\n<td>96.79</td>\n<td>78.70</td>\n</tr>\n</tbody>\n</table>\n<p>AP@0.5 is 1.0 in all experiments. We see a large gap between AP@0.5-0.95 between two validation sets. Some reasons for that:</p>\n<ul>\n<li>New version add more 3/50 typical hard cases.</li>\n<li>Inconsistent processing pipeline: val images in Remek's val was resized 2 times (original --&gt; 1024 --&gt; 416) </li>\n<li>Training images is annotated according to personal bias (no standard way/consentration to annotate the breast boxes correctly). So higher AP may not indicate a better model.</li>\n<li>The validation size is also not large enough to judge</li>\n<li>No hyper parameters tuning</li>\n</ul>\n<p>Did these things led to the large gap, particularly with stronger model and larger image size ?<br>\nAll these efforts are just to ensure an \"as good as posible\" ROI detection model. I think <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> 's <a href=\"https://www.kaggle.com/datasets/remekkinas/rsna-roi-detector-annotations-yolo\" target=\"_blank\">dataset</a> is enough to train good YOLOX models and they could perform equally well in hidden test set.</p>\n<p>Simpler Otsu thresholding + findCountours() slightly modified from <a href=\"https://www.kaggle.com/code/snnclsr/roi-extraction-using-opencv\" target=\"_blank\">this notebook</a> is used to find breast bbox as a fall back in case of YOLOX's miss-detection.<br>\nOr, if both miss the breast box, just use the whole image without any cropping.</p>\n<h1>2. The inference pipeline</h1>\n<p>Operations on large array take time, so I try to transfer the computation task to GPU as much as posible.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10254700%2F9de52f41201b4a92774f5fee331cb879%2Fwriteup-Page-1.drawio.png?generation=1678001361222187&amp;alt=media\" alt=\"\"></p>\n<h1>3. Early experiments</h1>\n<p>My final solution use external datasets, but I stitch up with using only competition data for almost of the time (until \"7 days to go\"). Hence most of my experiments are done on competition data only: 5-folds splits with StratifiedGroupKFold based on <code>patient_id</code>. Training hyperparams used for the final solution are almost inherited from these early experiments.</p>\n<h2>3.1. About the metric</h2>\n<p>The competition pF1 score is not stable and hard to track for me. Therefore, I mainly track my experiments based on multiple metrics: <code>{ PR_AUC, ROC_AUC, best_PF1 (binarized), best_threshold }</code> instead of just one.</p>\n<ul>\n<li><strong>PR_AUC</strong>: correlated with but more stable than best_PF1. It focuses on positive cases, and is strongly affected by prior data distribution (% of positive).</li>\n<li><strong>ROC_AUC</strong>: less affected by prior data distribution. Much more stable, but seem to be over optimistic which led to just a small gap between a good model and a bad model.</li>\n<li>To get a high binaried pf1, model should not predict too many positives which usually led to large FP --&gt; dramatically reduce best_PF1. A good scored pf1 model tends to prioritize Precision over Recall. I personaly don't like this behaviour, especially for real life application.</li>\n</ul>\n<h2>3.2. Augmentations</h2>\n<p>I stitch with this augmentation pipeline for all experiments, no tuning at all:</p>\n<pre><code>A.Compose([\n    \n    custom_augs.CustomRandomSizedCropNoResize(scale=(, ), ratio=(, ), p=),\n    \n    A.HorizontalFlip(p=),\n    A.VerticalFlip(p=),\n    \n    A.OneOf([\n        A.Downscale(scale_min=, scale_max=, interpolation=(upscale=cv2.INTER_LINEAR, downscale=cv2.INTER_AREA), p=),\n        A.Downscale(scale_min=, scale_max=, interpolation=(upscale=cv2.INTER_LANCZOS4, downscale=cv2.INTER_AREA), p=),\n        A.Downscale(scale_min=, scale_max=, interpolation=(upscale=cv2.INTER_LINEAR, downscale=cv2.INTER_LINEAR), p=),\n    ], p=),\n    \n    A.OneOf([\n        A.RandomToneCurve(scale=, p=),\n        A.RandomBrightnessContrast(brightness_limit=(-, ), contrast_limit=(-, ), brightness_by_max=, always_apply=, p=)\n    ], p=),\n    \n    A.OneOf(\n        [\n            A.ShiftScaleRotate(shift_limit=, scale_limit=[-, ], rotate_limit=[-, ], interpolation=cv2.INTER_LINEAR,\n                               border_mode=cv2.BORDER_CONSTANT, value=, mask_value=, shift_limit_x=[-, ],\n                               shift_limit_y=[-, ], rotate_method=, p=),\n            A.ElasticTransform(alpha=, sigma=, alpha_affine=, interpolation=cv2.INTER_LINEAR, border_mode=cv2.BORDER_CONSTANT,\n                               value=, mask_value=, approximate=, same_dxdy=, p=),\n            A.GridDistortion(num_steps=, distort_limit=, interpolation=cv2.INTER_LINEAR, border_mode=cv2.BORDER_CONSTANT,\n                             value=, mask_value=, normalized=, p=),\n        ], p=),\n    \n    A.CoarseDropout(max_holes=, max_height=, max_width=, min_holes=, min_height=, min_width=,\n                    fill_value=, mask_fill_value=, p=),\n    ], p=)\n</code></pre>\n<p>For the random crop choice: real breast size/ratio vary largly between images --&gt; popular pipeline of <code>longest resize + padding</code> introduces multi-scales problem. Of course, it would introduce higher risks of wrong positive label.</p>\n<p><em>Example batch</em><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10254700%2F4f149f8617c07c9a8e0d8e571f6a68fa%2Fexample_batch.jpg?generation=1685726233222765&amp;alt=media\" alt=\"example batch\"></p>\n<h2>3.3. Up/down sampling</h2>\n<p>I upsample pos cases in each epoch for all of my experiments.</p>\n<ul>\n<li>Ensuring at least 1 pos in a batch/iteration is pretty important to stablize training. I found training difficult while set 0.5 pos per batch.</li>\n<li>Upsampling ratio can largly affect CV score and prediction distribution. It also vary between backbones and other hyperparams choices. However, I prefer smallest pos/neg ratio as posible (since it's near to the real data distribution) but ensure at least 1 pos/batch.</li>\n<li>Large pos/neg ratio helps training faster in early epochs. I tried linearly increase/decrease the pos/neg ratio between epochs to face with some problems of prediction distribution/threshold (especially EffB4). But in the end, i got no improvement in CV.</li>\n</ul>\n<h2>3.4. Model/backbone</h2>\n<p>I tried <code>Eff-B2</code>, <code>Eff-B4</code>, <code>Effv2-s</code> and <code>Convnextv1-small</code></p>\n<ul>\n<li>Each model has its own characteristic and training phenomenon.</li>\n<li>All models could perform equally well in local CV. Except that <code>Convnextv1-small</code> give higher CV score.</li>\n<li><code>EfficientNet</code> (no model EMA) tends to overfit quickly after fews epochs with high pos/neg ratio: longer training reduce AUC largely and may slightly increase best_pf1 --&gt; model tends to predict less positives. Small pos/neg ratio helps training more stable but reduce CV. Linearly increase pos/neg ratio between epochs (by a sampler) did not help much.</li>\n<li><code>Convnext-small</code> (with/without EMA) shows both stable training and better CV.</li>\n</ul>\n<h2>3.5. drop_rate, drop_path_rate</h2>\n<p>Playing with drop_rate and drop_path_rate:</p>\n<ul>\n<li>We can use large dropout rate of &gt;= 0.5 to regularize training and reduce overfiting.</li>\n<li>With very large drop_rate = 0.9 or drop_path_rate = 0.5, I still can get a \"not bad as expected\" model in CV score. The following results is for Eff-B4, pos/neg = 1/3 on fold 0:</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>drop_rate</th>\n<th>drop_path_rate</th>\n<th>auc</th>\n<th>best_pf1</th>\n<th>best_thres</th>\n<th>epoch</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.9</td>\n<td>0.2</td>\n<td>91.90</td>\n<td>47.73</td>\n<td>0.78</td>\n<td>4</td>\n</tr>\n<tr>\n<td>0.7</td>\n<td>0.2</td>\n<td>90.36</td>\n<td>52.27</td>\n<td>0.55</td>\n<td>3</td>\n</tr>\n<tr>\n<td>0.5</td>\n<td>0.2</td>\n<td>90.02</td>\n<td>50.00</td>\n<td>0.72</td>\n<td>4</td>\n</tr>\n<tr>\n<td>0.3</td>\n<td>0.2</td>\n<td>91.23</td>\n<td>48.24</td>\n<td>0.82</td>\n<td>4</td>\n</tr>\n<tr>\n<td>0.5</td>\n<td>0.5</td>\n<td>90.45</td>\n<td>46.33</td>\n<td>0.95</td>\n<td>4</td>\n</tr>\n</tbody>\n</table>\n<p>However, I set drop_rate = 0.5 and drop_path_rate = 0.2 for most of my experiments, including the final ones.</p>\n<h2>3.6. Global pooling</h2>\n<p>I stitch with max pooling for almost my experiments as my inductive bias:</p>\n<ul>\n<li>Max pooling is suitable and seem to be effective for anomalies detection or \"needle in the haystack\" tasks in literature. In this case, cancer may appear in a very small region and the rest are all normal.</li>\n<li><code>max()</code> provides stronger learning signal, but less stable than mean() in term of gradient.</li>\n<li>My guess: <code>gem</code> &gt; <code>max</code> &gt; <code>mean</code> when enough data provided.</li>\n</ul>\n<h2>3.7. Soft positive label/ Positive label smoothing</h2>\n<p>Convnextv1-small look good in CV scores with stable AUC, PR_AUC, best PF1 across epochs. But there're differences in behaviour between Effv2-s and Convnextv1-small especially in best threshold for image/breast level.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10254700%2F6771d0c6e87067873e53021742e7bf6b%2Fwriteup_smoothing.png?generation=1678006236367845&amp;alt=media\" alt=\"\"><br>\n<em>The two lower green ones belong to <code>Effv2-s</code> and the other belong to <code>Convnextv1-small</code></em></p>\n<p>Some discussions suggest smaller best threshold (&lt;0.55) may indicate a better model. For single-image, Convnext show a very high threshold of &gt; 0.92, which could indicate the problem of over-confident. Stronger models with larger number of parameters is easier to be over-confident or overfitted, especialy in this highly imballanced dataset scenario. About the above figure, <code>label_smoothing = 0.1</code> was used but seem like it was not enough.<br>\nSo, just add harder label smoothing to regularize training. Or use positive weight &lt; 1.0 to reduce the priority of positive samples.</p>\n<table>\n<thead>\n<tr>\n<th>loss</th>\n<th>num_logits</th>\n<th>target {neg, pos}</th>\n<th>pr_auc</th>\n<th>roc_auc</th>\n<th>best_pf1</th>\n<th>best_thres</th>\n<th>epoch</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><em>(baseline) bce_smooth 0.1</em></td>\n<td>2</td>\n<td>{ [0.95, 0.05], [0.05, 0.95] }</td>\n<td>0.4755</td>\n<td>0.9278</td>\n<td>0.497</td>\n<td>0.66</td>\n<td>14</td>\n</tr>\n<tr>\n<td>bce_smooth 0.4</td>\n<td>2</td>\n<td>{ [0.8, 0.2], [0.2, 0.8] }</td>\n<td>0.4749</td>\n<td>0.9248</td>\n<td>0.5</td>\n<td>0.6</td>\n<td>25</td>\n</tr>\n<tr>\n<td>bce_pos_smooth 0.4</td>\n<td>2</td>\n<td>{ [1.0, 0.0], [0.2, 0.8] }</td>\n<td>0.5191</td>\n<td>0.9153</td>\n<td>0.5488</td>\n<td>0.53</td>\n<td>13.5</td>\n</tr>\n<tr>\n<td><em>(best) bce_pos_smooth 0.2</em></td>\n<td>1</td>\n<td>{ 0.0, 0.8 }</td>\n<td><strong>0.5401</strong></td>\n<td>0.9281</td>\n<td><strong>0.5714</strong></td>\n<td>0.49</td>\n<td>20</td>\n</tr>\n<tr>\n<td>bce_pos_smooth 0.3</td>\n<td>1</td>\n<td>{ 0.0, 0.7 }</td>\n<td>0.522</td>\n<td><strong>0.933</strong></td>\n<td>0.517</td>\n<td>0.5</td>\n<td>17</td>\n</tr>\n<tr>\n<td>bce_smooth 0.1 + pos_weight 0.4</td>\n<td>1</td>\n<td>{ 0.05, 0.95 }</td>\n<td>0.4946</td>\n<td>0.9146</td>\n<td>0.5393</td>\n<td>0.39</td>\n<td>19</td>\n</tr>\n</tbody>\n</table>\n<p><em>Note:</em></p>\n<ul>\n<li><em>Table above is fold 0 CV results</em></li>\n<li><em><code>num_logits = 2</code> means using <code>sigmoid</code> (BCEWithLogitsLoss) for training and <code>softmax</code> for inference. Refer <a href=\"https://github.com/ultralytics/yolov5/issues/5401\" target=\"_blank\">here</a>.</em></li>\n</ul>\n<p>Soft positive labeling look reasonable: we have per-breast label and not per-image label. For some images belong to same patient, cancer signal may not appears clearly in some images, or even all images (MG is not enough to judge for cancer/non-cancer) --&gt; the positive label should not be the maximum bound value of 1.0, but less confident.</p>\n<p>Soft postive label trick improve CV and helps threshold looks much better.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10254700%2Fc36afde0a06afbeb3dbd38a76753276c%2Fpos_smooth_thres.png?generation=1678007271357249&amp;alt=media\" alt=\"\"></p>\n<p>As some discussions, very sharp prediction distribution may indicate worse result/generalization.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10254700%2F7caf77b112791cd16eb20de6f6d443bd%2Fdistribution.png?generation=1678007391776163&amp;alt=media\" alt=\"\"></p>\n<hr>\n<h1>4. Final experiments</h1>\n<h2>4.1. External datasets</h2>\n<p>A week left to the competition deadline, I was thinking about training final experiments for the final submission and should not make any mistakes or missing something. I read some discussions again and relized I was missing a big part: external data. In particular, external data contains a large number of positive cases which are valuable.<br>\nThese external datasets summary:</p>\n<table>\n<thead>\n<tr>\n<th>Dataset</th>\n<th>num_patients*</th>\n<th>num_samples*</th>\n<th>num_pos_samples*</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><a href=\"https://physionet.org/content/vindr-mammo/1.0.0/\" target=\"_blank\">VinDr-Mammo</a></td>\n<td>5000</td>\n<td>20000</td>\n<td>226 (1.13 %)</td>\n</tr>\n<tr>\n<td><a href=\"https://www.kaggle.com/datasets/cheddad/miniddsm2\" target=\"_blank\">MiniDDSM</a></td>\n<td>1952</td>\n<td>7808</td>\n<td>1480 (18.95 %)</td>\n</tr>\n<tr>\n<td><a href=\"https://wiki.cancerimagingarchive.net/pages/viewpage.action?pageId=70230508\" target=\"_blank\">CMMD</a></td>\n<td>1775</td>\n<td>5202</td>\n<td>2632 (50.6%)</td>\n</tr>\n<tr>\n<td><a href=\"https://wiki.cancerimagingarchive.net/pages/viewpage.action?pageId=109379611\" target=\"_blank\">CDD-CESM</a></td>\n<td>326</td>\n<td>1003</td>\n<td>331 (33 %)</td>\n</tr>\n<tr>\n<td><a href=\"https://zenodo.org/record/5036062\" target=\"_blank\">BMCD</a></td>\n<td>82</td>\n<td>328</td>\n<td>22 (6.71 %)</td>\n</tr>\n<tr>\n<td>All</td>\n<td>9135</td>\n<td>34341</td>\n<td>4691 (13.66 %)</td>\n</tr>\n</tbody>\n</table>\n<p><strong>*</strong> The number may not indicate original dataset characteristics, but processed data I used for this competition.</p>\n<p><strong>Some details:</strong></p>\n<ol>\n<li><p><strong>VinDr-Mammo</strong>: contains BIRADS scores for each image of 0-5. I treated BIRADS 5 as cancer (1) and all other as normal (0). With only Digital Mammograms and BIRADS categories, one can't confirm 100% if a case is cancer or not . BIRADS 4 indicate 30% chance of cancer, then I was treated it as normal. My decision only happen in just a few seconds as i could remember. Reading other posts, i'm feeling my decision is not as good, except that it helps reduce sensitivity and can \"improve\" the pF1 (I don't want to see it that way). Maybe I make a huge mistake here. Better solution is to use soft/uncertain label or pseudo labeling for these ambigous (BIRADS-4) cases instead. Some images has LUTDescriptor. The image look over-exposured when apply VOILUT (voi + windowing), so I just apply windowing on this dataset, equivalent to pydicom's <code>apply_voi_lut(prefer_lut = False)</code></p></li>\n<li><p><strong>MiniDDSM</strong>: I used all 7808 samples. Found that the status label of <code>{Cancer, Benign, Normal}</code> is per patient_id, not per laterality. So I treat a patient-laterality as cancer if and only if status == 'Cancer' and at least 1 supicious region annotation (segmentation map) for that laterality is available. End up in 1480 positives and the remaining 6318 negatives. I use 16-bits png part for the less information loss. No windowing parameters as presented. Since there're watermarks noise with very high pixel intensity in ROI crop, percentile min-max scaled  was performed instead of min-max scaled for normalization.</p></li>\n<li><p><strong>CMMD</strong>: Total of 5202 breast images belong to 1872 patient id. Note that some patient ids start with 'D2' are almost malignant and usually had label for one laterality only. For those cases, I treated the other laterality (no laterality-level label specified in csv file, but still have image) as normal (EDA from the competition data show that cancer only appears in one laterality). Original dicom images are in 8-bits depth with windowing parameters available.</p></li>\n<li><p><strong>CDD-CESM</strong>: consisting Contrast-enhanced spectral mammography (CESM) images. This dataset contain label of <code>{Normal, Malignant, Benign}</code>. I treated Malignant as cancer, Normal or Benign as normal and only use the <strong>low-energy images</strong> as it is comparable to digital mammograms (MG), or at least they look pretty similar for me. Low-energy images is in 8-bits jpeg, no windowing information.</p></li>\n<li><p><strong>BMCD</strong>: contains 100 patients (50 normal + 50 suspicious cases) with 82 biopsy-confirmed cases of <code>{'NORMAL', 'BENIGN', 'DCIS', 'MALIGNANT'}</code> and mammogram images of them at the time of screening and  avg 2.2 year before. I treat 'DCIS' or 'MALIGNANT' patient's last screening images as cancer and all the remaining as normal. Original dicom images is in 16-bits depth and windowing parameters are available.</p></li>\n</ol>\n<h2>4.2. Validation strategy</h2>\n<p>I found inconsistence in CV between 5-folds splits of competition data, probably because the number of positive is not sufficient. Although, hidden test should has distribution/property closer to the competition data, so I change validation strategy to 4-splits as:</p>\n<ul>\n<li>Do 4-folds splitting on competition data</li>\n<li>Use 1 fold for validation, the rest 3 folds + all external data for training.<br>\nThen, for each split, training data contain about 5560/75400 positive cases (~7.38 %).</li>\n</ul>\n<h2>4.3. Training</h2>\n<p>I managed to get 4 x Convnextv1-small corresponding to the above 4 splits. Some unexpected results were founded during training, so the training stages was changed and in short consist of:</p>\n<ol>\n<li>Train 2 models on fold 0 and fold 1 with <code>soft_pos_label = 0.8</code></li>\n<li>Train 2 models on fold 2 and fold 3 with <code>soft_pos_label = 0.9</code></li>\n<li>Finetune 2 models obtained from stage 1 on fold 0 and fold 1 with <code>soft_pos_label = 0.9</code></li>\n</ol>\n<p>In details, I start training on first two folds: fold 0 and fold 1 with the following config:</p>\n<ul>\n<li>Model: timm's convnext_small.fb_in22k_ft_in1k_384</li>\n<li>Input size: 2048x1024</li>\n<li>Loss: vanila BCE (no class weight)</li>\n<li>Sampler: upsampling pos samples per epoch to pos/neg = 1/7, ensure each batch contains at least 1 pos sample.</li>\n<li>Batchsize: 8</li>\n<li>Automatic Mixed Precision (AMP): enable</li>\n<li>Model EMA: enable</li>\n<li>Global pooling: max</li>\n<li>Soft positive label = 0.8</li>\n<li>Optimizer: SGD with momemtum=0.9</li>\n<li>Scheduler: Cosine lr decay(epoch = 24, lr = 1e-3, min_lr = 1e-5) + linear warmup(warmup_lr = 1e-5, warmup_epoch = 4)</li>\n<li>Drop_rate = 0.5, drop_path_rate = 0.2</li>\n</ul>\n<p>Once training finished, results are not as my expectation on fold 0:</p>\n<ul>\n<li>CV results are not good</li>\n<li>Threshold is much smaller. I expected it to be in the range [0.35, 0.5], but it's just around 0.25+-0.02</li>\n<li>Training is not converged yet. I guess that CV can be improved with more additional training epochs.</li>\n</ul>\n<p>So I start train fold 2 and fold 3 with few changes: longer training with larger learning rate and reduce soft positive label.</p>\n<ul>\n<li>Scheduler: Cosine lr decay(epoch = 30, lr = 3e-3, min_lr = 5e-5) + linear warmup(warmup_lr = 3e-5, warmup_epoch = 4)</li>\n<li>Soft positive label: 0.9</li>\n</ul>\n<p>Results on fold 2 and fold 3 seem to be better. So I decided to finetune fold 0 and fold 1 with the same value of soft_positive_label = 0.9 from the previous last checkpoints.</p>\n<h2>4.4. Checkpoints selection</h2>\n<ol>\n<li>For each fold, I manually select 3-7 best checkpoints by looking at multiple metrics.</li>\n<li>Grid search over all posible combinations of each fold's checkpoints (3x5x7x5 = 525 combinations for my case), compute Out-Of-Fold (OOF) results for each combination.</li>\n<li>Select the best combination with highest binarized pF1. The best combination has local OOF pf1 = 0.5187, while the worst with 0.4951.</li>\n</ol>\n<p>Final results:</p>\n<table>\n<thead>\n<tr>\n<th>name</th>\n<th>soft_positive_label</th>\n<th>pr_auc</th>\n<th>roc_auc</th>\n<th>best_pf1</th>\n<th>best_thres</th>\n<th>epoch</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>fold 0</td>\n<td>0.8</td>\n<td>0.3983</td>\n<td>0.9142</td>\n<td>0.4716</td>\n<td>0.25</td>\n<td>24</td>\n</tr>\n<tr>\n<td><strong>(selected)</strong> fold 0 + fine-tune</td>\n<td>0.9</td>\n<td>0.4363</td>\n<td>0.9119</td>\n<td>0.4785</td>\n<td>0.35</td>\n<td>11 (24 + 11)</td>\n</tr>\n<tr>\n<td>fold 1</td>\n<td>0.8</td>\n<td>0.5151</td>\n<td>0.9202</td>\n<td>0.5291</td>\n<td>0.34</td>\n<td>18</td>\n</tr>\n<tr>\n<td><strong>(selected)</strong> fold 1 + fine-tune</td>\n<td>0.9</td>\n<td>0.5381</td>\n<td>0.9149</td>\n<td>0.5381</td>\n<td>0.34</td>\n<td>8 (24 + 8)</td>\n</tr>\n<tr>\n<td><strong>(selected)</strong> fold 2</td>\n<td>0.9</td>\n<td>0.4946</td>\n<td>0.9234</td>\n<td>0.5185</td>\n<td>0.34</td>\n<td>26</td>\n</tr>\n<tr>\n<td><strong>(selected)</strong> fold 3</td>\n<td>0.9</td>\n<td>0.5088</td>\n<td>0.9401</td>\n<td>0.5455</td>\n<td>0.31</td>\n<td>19</td>\n</tr>\n</tbody>\n</table>\n<p>Some thoughts:</p>\n<ul>\n<li>Fold 0 result looks weird? Maybe I need more time inspecting it.</li>\n<li>With much more data especially positive ones, guesses for the best hyperparams became obsolete. Maybe I could forget the soft positive label trick and get better result?</li>\n</ul>\n<p>OOF validation was done to determine the best threshold value of 0.34</p>\n<pre><code>                        auc      @th     f1      |  prec    recall  |   sens    spec \nsingle image     [0]    0.87296 0.40000 0.41907 |   0.48365 0.37047 |   0.37047 0.99145\ngrouby mean()    [0]    0.92043 0.34000 0.51820 |   0.60989 0.45122 |   0.45122 0.99391\ngrouby max()     [0]    0.91939 0.61000 0.50913 |   0.57545 0.45732 |   0.45732 0.99289\n--------------\n\nsingle image     [1]    0.84866 0.40000 0.33649 |   0.38241 0.30120 |   0.30120 0.98881\ngrouby mean()    [1]    0.89225 0.34000 0.39587 |   0.47027 0.34252 |   0.34252 0.99139\ngrouby max()     [1]    0.88917 0.61000 0.39424 |   0.44554 0.35433 |   0.35433 0.99016\n--------------\n\nsingle image     [2]    0.89611 0.40000 0.53331 |   0.62912 0.46356 |   0.46356 0.99453\ngrouby mean()    [2]    0.94288 0.34000 0.64699 |   0.75419 0.56722 |   0.56723 0.99632\ngrouby max()     [2]    0.94329 0.61000 0.63182 |   0.71428 0.56722 |   0.56723 0.99548\n--------------\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10254700%2Feeb0cc23fad1497f249613e0d4f87307%2Ffinal_sub_plot.png?generation=1678010502203196&amp;alt=media\" alt=\"\"></p>\n<p><em>The results was generated by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> ’s <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/378521\" target=\"_blank\">script</a></em></p>\n<h2>4.5. Submission</h2>\n<p>My training progress was done in the last day of the competition.<br>\n5 final submissions include:</p>\n<ul>\n<li>1 submission of models in earlier epochs while waiting training to be done. The main purpose is to ensure the pipeline is correct and no exception occurs, which give me 0.56 LB.</li>\n<li>4 submission of same final (best) model with different threshold: 0.31, 0.34 (best oof), 0.37, 0.40 . I guessed the LB's best threshold &gt;  CV's best threshold, but I'm totally wrong (I did not probe the LB). But it was fortunate that threshold=0.31 is the one on the peak :D</li>\n</ul>\n<p>Those submissions bring me from LB 600th to LB 22nd in one day. The PL 0.55 submission was successfully finished when ~30 mins left to the deadline.</p>\n<table>\n<thead>\n<tr>\n<th>Threshold</th>\n<th>OOF</th>\n<th>LB</th>\n<th>PL</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.27 (late sub)</td>\n<td>0.4877</td>\n<td>0.60</td>\n<td>0.53</td>\n</tr>\n<tr>\n<td>0.28 (late sub)</td>\n<td>0.4917</td>\n<td>0.60</td>\n<td>0.54</td>\n</tr>\n<tr>\n<td>0.29 (late sub)</td>\n<td>0.4973</td>\n<td>0.60</td>\n<td>0.54</td>\n</tr>\n<tr>\n<td>0.30 (late sub)</td>\n<td>0.5027</td>\n<td>0.60</td>\n<td>0.55</td>\n</tr>\n<tr>\n<td><strong>(selection) 0.31</strong></td>\n<td><strong>0.5049</strong></td>\n<td><strong>0.61</strong></td>\n<td><strong>0.55</strong></td>\n</tr>\n<tr>\n<td><strong>(selection) 0.34</strong></td>\n<td><strong>0.5187</strong></td>\n<td><strong>0.58</strong></td>\n<td><strong>0.53</strong></td>\n</tr>\n<tr>\n<td>0.37</td>\n<td>0.5000</td>\n<td>0.55</td>\n<td>0.52</td>\n</tr>\n<tr>\n<td>0.40</td>\n<td>0.4896</td>\n<td>0.54</td>\n<td>0.50</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h1>5. Code</h1>\n<ul>\n<li>Submission notebook: <a href=\"https://www.kaggle.com/dangnh0611/1st-place-submission-code\" target=\"_blank\">https://www.kaggle.com/dangnh0611/1st-place-submission-code</a> </li>\n<li>Training code: <a href=\"https://github.com/dangnh0611/kaggle_rsna_breast_cancer\" target=\"_blank\">https://github.com/dangnh0611/kaggle_rsna_breast_cancer</a></li>\n</ul>\n<hr>\n<p>I just got luck with simple pipeline and simple decision. Many teams had much better models but did not select it in final, as I could see.<br>\nThanks for your attention.</p>",
      "rawMarkdown": "***Update 10/04/2023***\n\nAblation study: despite of a large improvement on a particular fold (below), soft positive label trick does not show a clearly improvement in performance over the standard label smoothing technique. The use of external datasets improve F1-score about 0.02 on local OOF validation and Private Leaderboard (with extractly same training pipeline + hyper-params).\n\n| External data |            Loss           |   OOF F1   |    LB    |    PL    |\n|:-------------:|:-------------------------:|:----------:|:--------:|:--------:|\n|       x       |    label smoothing=0.1    |  0.4921    |   0.60   |   0.53   |\n|       x       |  soft positive label=0.8  |   0.4853   |   0.60   |   0.53   |\n|       ✓       |   label smoothing = 0.1   |   0.5161   |   0.58   | **0.56** |\n|       ✓       | soft positive label = 0.9 | **0.5182** | **0.61** |   0.55   |\n\n\n-----\n\nFirst of all, I would like to thank Kaggle and competition's host for such an amazing challenge, a lofty goal with high data quality. Thank you to all paticipants/kagglers with many active and helpful discussions/codes. My solution was just built up from every pieces of kindly shares from you. I learned a lot and I'm very appreciated for that.\nI'm also very happy and suprised with the 1st place. This is my first gold medal and I'm writing my first writeup. It was such a great journey for me.\nFor the solution, I use a very simple pipeline which can be described in just few lines:\n- Use some external datasets: VinDr-Mammo, MiniDDSM, CMMD, CDD-CESM, BMCD.\n- 4 x Convnextv1-small 2048x1024, validated on 4-folds splits of competition data.\n- Soft positive label\n\n![inference pipeline](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10254700%2F15d5d89bcd9cd18e45b6da133d5f5ad7%2Fpreprocess.jpg?generation=1685726162522472&alt=media)\n\nNow I want to share some experiments and my thought about those. Many of theme could be found in another discussions by excellent kagglers. Many of theme seem obvious. Hope this helps some new comer getting started in the future. Kindly note that it's just my own opinion/thoughts with very limited experiments and knownledge. I'm appreciated for your discussions and feel free to correct me if something was wrong.\n\n# 1. ROI crop\nROI cropping was performed since it effectively help keeping more texture/detail given a fixed resolution. I use YOLOX-nano 416x416 for ROI detector. The advantage of DL detector vs rule-based methods is the obtained bbox is smaller, aspect ratio is more stable and focus to the breast region.\n1. Train a YOLOX on @remekkinas 's [dataset](https://www.kaggle.com/datasets/remekkinas/rsna-roi-detector-annotations-yolo) (472 bbox-annotated images)\n2. Inference on all available training data with low `conf_thres` and high `iou_thres`. Only 3 miss-detected images (all contain noise) and over 100 images with 2 boxes (almost overlapped). I manually select and label 99 of those images. Therefore, I have 571 annotated images in total.\n3. Retrain YOLOX on new images: 521 for train, 50 for val. Note that these 50 val images include all 47 val images of original @remekkinas 's dataset. The new dataset version contain original resolution images (same as original dicoms), preprocessed with simple min-max normalization. Tried various model/image sizes and finally choose YOLOX-nano 416x416 as final model due to the consistent result and small overhead.\n\n|   **model size**   | **image size** | **interpolation** | **AP_new_val** | **AP_remek_val** |\n|:------------------:|:--------------:|:-----------------:|:----------:|:--------------:|\n| _nano (selection)_ |      _416_     |      _LINEAR_     |   _96.26_  |     _94.21_    |\n|        nano        |       416      |        AREA       |    94.09   |      91.60      |\n|        nano        |       640      |       LINEAR      |    95.85   |      88.40      |\n|        nano        |       768      |       LINEAR      |    96.22   |      82.09     |\n|        nano        |      1024      |       LINEAR      |    94.92   |      89.40      |\n|        tiny        |       416      |       LINEAR      |    94.23   |      90.20      |\n|        tiny        |       640      |       LINEAR      |    94.95   |      89.84     |\n|        tiny        |       768      |        AREA       |    96.21   |      68.03     |\n|        tiny        |      1024      |        AREA       |    93.69   |      73.70      |\n|          s         |       416      |       LINEAR      |    95.03   |      0.86      |\n|          s         |       640      |       LINEAR      |    96.10    |      70.80      |\n|          s         |       768      |       LINEAR      |    96.79   |      78.70      |\n\n\nAP@0.5 is 1.0 in all experiments. We see a large gap between AP@0.5-0.95 between two validation sets. Some reasons for that:\n- New version add more 3/50 typical hard cases.\n- Inconsistent processing pipeline: val images in Remek's val was resized 2 times (original --> 1024 --> 416) \n- Training images is annotated according to personal bias (no standard way/consentration to annotate the breast boxes correctly). So higher AP may not indicate a better model.\n- The validation size is also not large enough to judge\n- No hyper parameters tuning\n\nDid these things led to the large gap, particularly with stronger model and larger image size ?\nAll these efforts are just to ensure an \"as good as posible\" ROI detection model. I think @remekkinas 's [dataset](https://www.kaggle.com/datasets/remekkinas/rsna-roi-detector-annotations-yolo) is enough to train good YOLOX models and they could perform equally well in hidden test set.\n\nSimpler Otsu thresholding + findCountours() slightly modified from [this notebook](https://www.kaggle.com/code/snnclsr/roi-extraction-using-opencv) is used to find breast bbox as a fall back in case of YOLOX's miss-detection.\nOr, if both miss the breast box, just use the whole image without any cropping.\n\n# 2. The inference pipeline\nOperations on large array take time, so I try to transfer the computation task to GPU as much as posible.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10254700%2F9de52f41201b4a92774f5fee331cb879%2Fwriteup-Page-1.drawio.png?generation=1678001361222187&alt=media)\n\n# 3. Early experiments\nMy final solution use external datasets, but I stitch up with using only competition data for almost of the time (until \"7 days to go\"). Hence most of my experiments are done on competition data only: 5-folds splits with StratifiedGroupKFold based on `patient_id`. Training hyperparams used for the final solution are almost inherited from these early experiments.\n\n## 3.1. About the metric\nThe competition pF1 score is not stable and hard to track for me. Therefore, I mainly track my experiments based on multiple metrics: `{ PR_AUC, ROC_AUC, best_PF1 (binarized), best_threshold }` instead of just one.\n- **PR_AUC**: correlated with but more stable than best_PF1. It focuses on positive cases, and is strongly affected by prior data distribution (% of positive).\n- **ROC_AUC**: less affected by prior data distribution. Much more stable, but seem to be over optimistic which led to just a small gap between a good model and a bad model.\n- To get a high binaried pf1, model should not predict too many positives which usually led to large FP --> dramatically reduce best_PF1. A good scored pf1 model tends to prioritize Precision over Recall. I personaly don't like this behaviour, especially for real life application.\n\n## 3.2. Augmentations\nI stitch with this augmentation pipeline for all experiments, no tuning at all:\n\n```python\nA.Compose([\n    # crop, tweak from A.RandomSizedCrop()\n    custom_augs.CustomRandomSizedCropNoResize(scale=(0.5, 1.0), ratio=(0.5, 0.8), p=0.4),\n    # flip\n    A.HorizontalFlip(p=0.5),\n    A.VerticalFlip(p=0.5),\n    # downscale\n    A.OneOf([\n        A.Downscale(scale_min=0.75, scale_max=0.95, interpolation=dict(upscale=cv2.INTER_LINEAR, downscale=cv2.INTER_AREA), p=0.1),\n        A.Downscale(scale_min=0.75, scale_max=0.95, interpolation=dict(upscale=cv2.INTER_LANCZOS4, downscale=cv2.INTER_AREA), p=0.1),\n        A.Downscale(scale_min=0.75, scale_max=0.95, interpolation=dict(upscale=cv2.INTER_LINEAR, downscale=cv2.INTER_LINEAR), p=0.8),\n    ], p=0.125),\n    # contrast\n    A.OneOf([\n        A.RandomToneCurve(scale=0.3, p=0.5),\n        A.RandomBrightnessContrast(brightness_limit=(-0.1, 0.2), contrast_limit=(-0.4, 0.5), brightness_by_max=True, always_apply=False, p=0.5)\n    ], p=0.5),\n    # geometric\n    A.OneOf(\n        [\n            A.ShiftScaleRotate(shift_limit=None, scale_limit=[-0.15, 0.15], rotate_limit=[-30, 30], interpolation=cv2.INTER_LINEAR,\n                               border_mode=cv2.BORDER_CONSTANT, value=0, mask_value=None, shift_limit_x=[-0.1, 0.1],\n                               shift_limit_y=[-0.2, 0.2], rotate_method='largest_box', p=0.6),\n            A.ElasticTransform(alpha=1, sigma=20, alpha_affine=10, interpolation=cv2.INTER_LINEAR, border_mode=cv2.BORDER_CONSTANT,\n                               value=0, mask_value=None, approximate=False, same_dxdy=False, p=0.2),\n            A.GridDistortion(num_steps=5, distort_limit=0.3, interpolation=cv2.INTER_LINEAR, border_mode=cv2.BORDER_CONSTANT,\n                             value=0, mask_value=None, normalized=True, p=0.2),\n        ], p=0.5),\n    # random erase\n    A.CoarseDropout(max_holes=6, max_height=0.15, max_width=0.25, min_holes=1, min_height=0.05, min_width=0.1,\n                    fill_value=0, mask_fill_value=None, p=0.25),\n    ], p=0.9)\n```\nFor the random crop choice: real breast size/ratio vary largly between images --> popular pipeline of `longest resize + padding` introduces multi-scales problem. Of course, it would introduce higher risks of wrong positive label.\n\n*Example batch*\n![example batch](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10254700%2F4f149f8617c07c9a8e0d8e571f6a68fa%2Fexample_batch.jpg?generation=1685726233222765&alt=media)\n\n\n## 3.3. Up/down sampling\nI upsample pos cases in each epoch for all of my experiments.\n- Ensuring at least 1 pos in a batch/iteration is pretty important to stablize training. I found training difficult while set 0.5 pos per batch.\n- Upsampling ratio can largly affect CV score and prediction distribution. It also vary between backbones and other hyperparams choices. However, I prefer smallest pos/neg ratio as posible (since it's near to the real data distribution) but ensure at least 1 pos/batch.\n- Large pos/neg ratio helps training faster in early epochs. I tried linearly increase/decrease the pos/neg ratio between epochs to face with some problems of prediction distribution/threshold (especially EffB4). But in the end, i got no improvement in CV.\n\n## 3.4. Model/backbone\nI tried `Eff-B2`, `Eff-B4`, `Effv2-s` and `Convnextv1-small`\n- Each model has its own characteristic and training phenomenon.\n- All models could perform equally well in local CV. Except that `Convnextv1-small` give higher CV score.\n- `EfficientNet` (no model EMA) tends to overfit quickly after fews epochs with high pos/neg ratio: longer training reduce AUC largely and may slightly increase best_pf1 --> model tends to predict less positives. Small pos/neg ratio helps training more stable but reduce CV. Linearly increase pos/neg ratio between epochs (by a sampler) did not help much.\n- `Convnext-small` (with/without EMA) shows both stable training and better CV.\n\n## 3.5. drop_rate, drop_path_rate\nPlaying with drop_rate and drop_path_rate:\n- We can use large dropout rate of >= 0.5 to regularize training and reduce overfiting.\n- With very large drop_rate = 0.9 or drop_path_rate = 0.5, I still can get a \"not bad as expected\" model in CV score. The following results is for Eff-B4, pos/neg = 1/3 on fold 0:\n\n| drop_rate | drop_path_rate |  auc  | best_pf1 | best_thres | epoch |\n|:---------:|:--------------:|:-----:|:--------:|:----------:|:-----:|\n|    0.9    |       0.2      | 91.90 |   47.73  |    0.78    |   4   |\n|    0.7    |       0.2      | 90.36 |   52.27  |    0.55    |   3   |\n|    0.5    |       0.2      | 90.02 |   50.00  |    0.72    |   4   |\n|    0.3    |       0.2      | 91.23 |   48.24  |    0.82    |   4   |\n|    0.5    |       0.5      | 90.45 |   46.33  |    0.95    |   4   |\n\nHowever, I set drop_rate = 0.5 and drop_path_rate = 0.2 for most of my experiments, including the final ones.\n\n## 3.6. Global pooling\nI stitch with max pooling for almost my experiments as my inductive bias:\n\n- Max pooling is suitable and seem to be effective for anomalies detection or \"needle in the haystack\" tasks in literature. In this case, cancer may appear in a very small region and the rest are all normal.\n- `max()` provides stronger learning signal, but less stable than mean() in term of gradient.\n- My guess: `gem` > `max` > `mean` when enough data provided.\n\n## 3.7. Soft positive label/ Positive label smoothing\nConvnextv1-small look good in CV scores with stable AUC, PR_AUC, best PF1 across epochs. But there're differences in behaviour between Effv2-s and Convnextv1-small especially in best threshold for image/breast level.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10254700%2F6771d0c6e87067873e53021742e7bf6b%2Fwriteup_smoothing.png?generation=1678006236367845&alt=media)\n*The two lower green ones belong to `Effv2-s` and the other belong to `Convnextv1-small`*\n\nSome discussions suggest smaller best threshold (<0.55) may indicate a better model. For single-image, Convnext show a very high threshold of > 0.92, which could indicate the problem of over-confident. Stronger models with larger number of parameters is easier to be over-confident or overfitted, especialy in this highly imballanced dataset scenario. About the above figure, `label_smoothing = 0.1` was used but seem like it was not enough.\nSo, just add harder label smoothing to regularize training. Or use positive weight < 1.0 to reduce the priority of positive samples.\n\n\n| loss                            | num_logits | target {neg, pos}              | pr_auc     | roc_auc       | best_pf1   | best_thres | epoch |\n|---------------------------------|------------|--------------------------------|------------|-----------|------------|------------|-------|\n| _(baseline) bce_smooth 0.1_     | 2          | { [0.95, 0.05], [0.05, 0.95] } | 0.4755     | 0.9278    | 0.497      | 0.66       | 14    |\n| bce_smooth 0.4                  | 2          | { [0.8, 0.2], [0.2, 0.8] }     | 0.4749     | 0.9248    | 0.5        | 0.6        | 25    |\n| bce_pos_smooth 0.4              | 2          | { [1.0, 0.0], [0.2, 0.8] }     | 0.5191     | 0.9153    | 0.5488     | 0.53       | 13.5  |\n| _(best) bce_pos_smooth 0.2_     | 1          | { 0.0, 0.8 }                   | **0.5401** | 0.9281    | **0.5714** | 0.49       | 20    |\n| bce_pos_smooth 0.3              | 1          | { 0.0, 0.7 }                   | 0.522      | **0.933** | 0.517      | 0.5        | 17    |\n| bce_smooth 0.1 + pos_weight 0.4 | 1          | { 0.05, 0.95 }                 | 0.4946     | 0.9146    | 0.5393     | 0.39       | 19    |\n\n*Note:*\n- *Table above is fold 0 CV results*\n- *`num_logits = 2` means using `sigmoid` (BCEWithLogitsLoss) for training and `softmax` for inference. Refer [here](https://github.com/ultralytics/yolov5/issues/5401).*\n\nSoft positive labeling look reasonable: we have per-breast label and not per-image label. For some images belong to same patient, cancer signal may not appears clearly in some images, or even all images (MG is not enough to judge for cancer/non-cancer) --> the positive label should not be the maximum bound value of 1.0, but less confident.\n\nSoft postive label trick improve CV and helps threshold looks much better.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10254700%2Fc36afde0a06afbeb3dbd38a76753276c%2Fpos_smooth_thres.png?generation=1678007271357249&alt=media)\n\nAs some discussions, very sharp prediction distribution may indicate worse result/generalization.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10254700%2F7caf77b112791cd16eb20de6f6d443bd%2Fdistribution.png?generation=1678007391776163&alt=media)\n\n\n---------------------\n# 4. Final experiments\n## 4.1. External datasets\nA week left to the competition deadline, I was thinking about training final experiments for the final submission and should not make any mistakes or missing something. I read some discussions again and relized I was missing a big part: external data. In particular, external data contains a large number of positive cases which are valuable.\nThese external datasets summary:\n\n| Dataset     | num_patients* | num_samples* | num_pos_samples* | \n|-------------|---------------|--------------|------------------|\n| [VinDr-Mammo](https://physionet.org/content/vindr-mammo/1.0.0/) | 5000          | 20000        | 226 (1.13 %)     | \n| [MiniDDSM](https://www.kaggle.com/datasets/cheddad/miniddsm2)   | 1952          | 7808         | 1480 (18.95 %)   |\n| [CMMD](https://wiki.cancerimagingarchive.net/pages/viewpage.action?pageId=70230508)        | 1775          | 5202         | 2632 (50.6%)     |\n| [CDD-CESM](https://wiki.cancerimagingarchive.net/pages/viewpage.action?pageId=109379611)    | 326           | 1003         | 331 (33 %)       | \n| [BMCD](https://zenodo.org/record/5036062)        | 82            | 328          | 22 (6.71 %)      | [1, 2]    |\n| All         | 9135          | 34341        | 4691 (13.66 %)   |\n\n***** The number may not indicate original dataset characteristics, but processed data I used for this competition.\n\n**Some details:**\n1. **VinDr-Mammo**: contains BIRADS scores for each image of 0-5. I treated BIRADS 5 as cancer (1) and all other as normal (0). With only Digital Mammograms and BIRADS categories, one can't confirm 100% if a case is cancer or not . BIRADS 4 indicate 30% chance of cancer, then I was treated it as normal. My decision only happen in just a few seconds as i could remember. Reading other posts, i'm feeling my decision is not as good, except that it helps reduce sensitivity and can \"improve\" the pF1 (I don't want to see it that way). Maybe I make a huge mistake here. Better solution is to use soft/uncertain label or pseudo labeling for these ambigous (BIRADS-4) cases instead. Some images has LUTDescriptor. The image look over-exposured when apply VOILUT (voi + windowing), so I just apply windowing on this dataset, equivalent to pydicom's `apply_voi_lut(prefer_lut = False)`\n\n2. **MiniDDSM**: I used all 7808 samples. Found that the status label of `{Cancer, Benign, Normal}` is per patient_id, not per laterality. So I treat a patient-laterality as cancer if and only if status == 'Cancer' and at least 1 supicious region annotation (segmentation map) for that laterality is available. End up in 1480 positives and the remaining 6318 negatives. I use 16-bits png part for the less information loss. No windowing parameters as presented. Since there're watermarks noise with very high pixel intensity in ROI crop, percentile min-max scaled  was performed instead of min-max scaled for normalization.\n\n3. **CMMD**: Total of 5202 breast images belong to 1872 patient id. Note that some patient ids start with 'D2' are almost malignant and usually had label for one laterality only. For those cases, I treated the other laterality (no laterality-level label specified in csv file, but still have image) as normal (EDA from the competition data show that cancer only appears in one laterality). Original dicom images are in 8-bits depth with windowing parameters available.\n\n4. **CDD-CESM**: consisting Contrast-enhanced spectral mammography (CESM) images. This dataset contain label of `{Normal, Malignant, Benign}`. I treated Malignant as cancer, Normal or Benign as normal and only use the **low-energy images** as it is comparable to digital mammograms (MG), or at least they look pretty similar for me. Low-energy images is in 8-bits jpeg, no windowing information.\n\n5. **BMCD**: contains 100 patients (50 normal + 50 suspicious cases) with 82 biopsy-confirmed cases of `{'NORMAL', 'BENIGN', 'DCIS', 'MALIGNANT'}` and mammogram images of them at the time of screening and  avg 2.2 year before. I treat 'DCIS' or 'MALIGNANT' patient's last screening images as cancer and all the remaining as normal. Original dicom images is in 16-bits depth and windowing parameters are available.\n\n## 4.2. Validation strategy\nI found inconsistence in CV between 5-folds splits of competition data, probably because the number of positive is not sufficient. Although, hidden test should has distribution/property closer to the competition data, so I change validation strategy to 4-splits as:\n- Do 4-folds splitting on competition data\n- Use 1 fold for validation, the rest 3 folds + all external data for training.\nThen, for each split, training data contain about 5560/75400 positive cases (~7.38 %).\n\n## 4.3. Training\nI managed to get 4 x Convnextv1-small corresponding to the above 4 splits. Some unexpected results were founded during training, so the training stages was changed and in short consist of:\n1. Train 2 models on fold 0 and fold 1 with `soft_pos_label = 0.8`\n2. Train 2 models on fold 2 and fold 3 with `soft_pos_label = 0.9`\n3. Finetune 2 models obtained from stage 1 on fold 0 and fold 1 with `soft_pos_label = 0.9`\n\nIn details, I start training on first two folds: fold 0 and fold 1 with the following config:\n- Model: timm's convnext_small.fb_in22k_ft_in1k_384\n- Input size: 2048x1024\n- Loss: vanila BCE (no class weight)\n- Sampler: upsampling pos samples per epoch to pos/neg = 1/7, ensure each batch contains at least 1 pos sample.\n- Batchsize: 8\n- Automatic Mixed Precision (AMP): enable\n- Model EMA: enable\n- Global pooling: max\n- Soft positive label = 0.8\n- Optimizer: SGD with momemtum=0.9\n- Scheduler: Cosine lr decay(epoch = 24, lr = 1e-3, min_lr = 1e-5) + linear warmup(warmup_lr = 1e-5, warmup_epoch = 4)\n- Drop_rate = 0.5, drop_path_rate = 0.2\n\nOnce training finished, results are not as my expectation on fold 0:\n- CV results are not good\n- Threshold is much smaller. I expected it to be in the range [0.35, 0.5], but it's just around 0.25+-0.02\n- Training is not converged yet. I guess that CV can be improved with more additional training epochs.\n\nSo I start train fold 2 and fold 3 with few changes: longer training with larger learning rate and reduce soft positive label.\n- Scheduler: Cosine lr decay(epoch = 30, lr = 3e-3, min_lr = 5e-5) + linear warmup(warmup_lr = 3e-5, warmup_epoch = 4)\n- Soft positive label: 0.9\n\nResults on fold 2 and fold 3 seem to be better. So I decided to finetune fold 0 and fold 1 with the same value of soft_positive_label = 0.9 from the previous last checkpoints.\n\n## 4.4. Checkpoints selection\n1. For each fold, I manually select 3-7 best checkpoints by looking at multiple metrics.\n2. Grid search over all posible combinations of each fold's checkpoints (3x5x7x5 = 525 combinations for my case), compute Out-Of-Fold (OOF) results for each combination.\n3. Select the best combination with highest binarized pF1. The best combination has local OOF pf1 = 0.5187, while the worst with 0.4951.\n\nFinal results:\n\n| name                          | soft_positive_label | pr_auc | roc_auc | best_pf1 | best_thres | epoch        |\n|-------------------------------|---------------------|--------|---------|----------|------------|--------------|\n| fold 0                        | 0.8                 | 0.3983 | 0.9142  | 0.4716   | 0.25       | 24           |\n| **(selected)** fold 0 + fine-tune | 0.9                 | 0.4363 | 0.9119  | 0.4785   | 0.35       | 11 (24 + 11) |\n| fold 1                        | 0.8                 | 0.5151 | 0.9202  | 0.5291   | 0.34       | 18           |\n| **(selected)** fold 1 + fine-tune | 0.9                 | 0.5381 | 0.9149  | 0.5381   | 0.34       | 8 (24 + 8)   |\n| **(selected)** fold 2             | 0.9                 | 0.4946 | 0.9234  | 0.5185   | 0.34       | 26           |\n| **(selected)** fold 3             | 0.9                 | 0.5088 | 0.9401  | 0.5455   | 0.31       | 19           |\n\nSome thoughts:\n- Fold 0 result looks weird? Maybe I need more time inspecting it.\n- With much more data especially positive ones, guesses for the best hyperparams became obsolete. Maybe I could forget the soft positive label trick and get better result?\n\nOOF validation was done to determine the best threshold value of 0.34\n\n```\n                    \tauc      @th     f1      | \tprec    recall  | \tsens    spec \nsingle image     [0]\t0.87296\t0.40000\t0.41907 | \t0.48365\t0.37047 | \t0.37047\t0.99145\ngrouby mean()    [0]\t0.92043\t0.34000\t0.51820 | \t0.60989\t0.45122 | \t0.45122\t0.99391\ngrouby max()     [0]\t0.91939\t0.61000\t0.50913 | \t0.57545\t0.45732 | \t0.45732\t0.99289\n--------------\n\nsingle image     [1]\t0.84866\t0.40000\t0.33649 | \t0.38241\t0.30120 | \t0.30120\t0.98881\ngrouby mean()    [1]\t0.89225\t0.34000\t0.39587 | \t0.47027\t0.34252 | \t0.34252\t0.99139\ngrouby max()     [1]\t0.88917\t0.61000\t0.39424 | \t0.44554\t0.35433 | \t0.35433\t0.99016\n--------------\n\nsingle image     [2]\t0.89611\t0.40000\t0.53331 | \t0.62912\t0.46356 | \t0.46356\t0.99453\ngrouby mean()    [2]\t0.94288\t0.34000\t0.64699 | \t0.75419\t0.56722 | \t0.56723\t0.99632\ngrouby max()     [2]\t0.94329\t0.61000\t0.63182 | \t0.71428\t0.56722 | \t0.56723\t0.99548\n--------------\n```\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10254700%2Feeb0cc23fad1497f249613e0d4f87307%2Ffinal_sub_plot.png?generation=1678010502203196&alt=media)\n\n*The results was generated by @hengck23 ’s [script](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/378521)*\n## 4.5. Submission\nMy training progress was done in the last day of the competition.\n5 final submissions include:\n- 1 submission of models in earlier epochs while waiting training to be done. The main purpose is to ensure the pipeline is correct and no exception occurs, which give me 0.56 LB.\n- 4 submission of same final (best) model with different threshold: 0.31, 0.34 (best oof), 0.37, 0.40 . I guessed the LB's best threshold >  CV's best threshold, but I'm totally wrong (I did not probe the LB). But it was fortunate that threshold=0.31 is the one on the peak :D\n\nThose submissions bring me from LB 600th to LB 22nd in one day. The PL 0.55 submission was successfully finished when ~30 mins left to the deadline.\n\n| Threshold            | OOF        | LB       | PL       |\n|----------------------|------------|----------|----------|\n| 0.27 (late sub)      | 0.4877     | 0.60     | 0.53     |\n| 0.28 (late sub)      | 0.4917     | 0.60     | 0.54     |\n| 0.29 (late sub)      | 0.4973     | 0.60     | 0.54     |\n| 0.30 (late sub)      | 0.5027     | 0.60     | 0.55     |\n| **(selection) 0.31** | **0.5049** | **0.61** | **0.55** |\n| **(selection) 0.34** | **0.5187** | **0.58** | **0.53** |\n| 0.37                 | 0.5000     | 0.55     | 0.52     |\n| 0.40                 | 0.4896     | 0.54     | 0.50     |\n\n\n ---\n# 5. Code\n- Submission notebook: https://www.kaggle.com/dangnh0611/1st-place-submission-code \n- Training code: https://github.com/dangnh0611/kaggle_rsna_breast_cancer\n\n\n---\n\nI just got luck with simple pipeline and simple decision. Many teams had much better models but did not select it in final, as I could see.\nThanks for your attention.",
      "votes": 178
    },
    {
      "id": 2169786,
      "postDate": "2023-03-05T12:35:48.053Z",
      "content": "<p>Thanks a lot for sharing such a detailed write up. Congratulations on your Gold… looking at the way you broke down the problem, its a well deserved gold. In my experiments also found that Maxpooling worked the best, while average pooling tended to wash away the required signals</p>",
      "rawMarkdown": "Thanks a lot for sharing such a detailed write up. Congratulations on your Gold... looking at the way you broke down the problem, its a well deserved gold. In my experiments also found that Maxpooling worked the best, while average pooling tended to wash away the required signals",
      "votes": 9,
      "replies": [
        {
          "id": 2169862,
          "postDate": "2023-03-05T13:59:42.743Z",
          "content": "<p>Thanks for your kind words !</p>",
          "rawMarkdown": "Thanks for your kind words !"
        }
      ]
    },
    {
      "id": 2191737,
      "postDate": "2023-03-22T07:14:32.980Z",
      "content": "<p>Could you share more training parameters of ROI detection model? I can't get the same AP@0.5-0.95 as you mentioned above. I only train on <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a>'s dataset. thank you!</p>",
      "rawMarkdown": "Could you share more training parameters of ROI detection model? I can't get the same AP@0.5-0.95 as you mentioned above. I only train on @remekkinas's dataset. thank you!",
      "votes": 1,
      "replies": [
        {
          "id": 2192430,
          "postDate": "2023-03-22T16:44:04.817Z",
          "content": "<p>Hi, insufficient number of val samples and personal labeling bias can make the AP@0.5-0.95 unstable. <br>\nFor YOLOX, I use <a href=\"https://github.com/dangnh0611/kaggle_rsna_breast_cancer/blob/reproduce/src/roi_det/YOLOX/exps/projects/rsna/yolox_nano_bre_416.py\" target=\"_blank\">this hyperparams config</a></p>",
          "rawMarkdown": "Hi, insufficient number of val samples and personal labeling bias can make the AP@0.5-0.95 unstable. \nFor YOLOX, I use [this hyperparams config](https://github.com/dangnh0611/kaggle_rsna_breast_cancer/blob/reproduce/src/roi_det/YOLOX/exps/projects/rsna/yolox_nano_bre_416.py)"
        }
      ]
    },
    {
      "id": 2174211,
      "postDate": "2023-03-09T00:18:06.917Z",
      "content": "<p>Really great work. </p>",
      "rawMarkdown": "Really great work. ",
      "votes": 1
    },
    {
      "id": 2172703,
      "postDate": "2023-03-07T18:05:37.537Z",
      "content": "<p>Hi, you mentioned that you used image size 2048x1024 with Convnextv1-small model, but when you processed the images yolox produce as output image of 416?.</p>\n<p>Thank in advance!</p>",
      "rawMarkdown": "Hi, you mentioned that you used image size 2048x1024 with Convnextv1-small model, but when you processed the images yolox produce as output image of 416?.\n\nThank in advance!",
      "votes": 1,
      "replies": [
        {
          "id": 2180975,
          "postDate": "2023-03-14T08:19:48.853Z",
          "content": "<p>Hi, i get the bounding box output from YOLOX, then the coordinates will be used to crop the original resolution image, as described <a href=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10254700%2F9de52f41201b4a92774f5fee331cb879%2Fwriteup-Page-1.drawio.png?generation=1678001361222187&amp;alt=media\" target=\"_blank\">here</a></p>",
          "rawMarkdown": "Hi, i get the bounding box output from YOLOX, then the coordinates will be used to crop the original resolution image, as described [here](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10254700%2F9de52f41201b4a92774f5fee331cb879%2Fwriteup-Page-1.drawio.png?generation=1678001361222187&alt=media)",
          "votes": 1
        }
      ]
    },
    {
      "id": 2171012,
      "postDate": "2023-03-06T13:11:48.213Z",
      "content": "<p>great work, and thanks to share the solution </p>",
      "rawMarkdown": "great work, and thanks to share the solution ",
      "votes": 1
    },
    {
      "id": 2170281,
      "postDate": "2023-03-05T20:11:53.087Z",
      "content": "<p>Absolutely great work! And a overwhelming amount of experiments you have done starting from pre-processing and all the way to the inference. Also you seem to have picked up many things during this short time that took me couple of years to get a grasp over while doing PhD on this very topic. In my opinion publishing these results and answering those remaining questions that you have listed would get you far for a PhD, unless of course you already have one. 😃</p>",
      "rawMarkdown": "Absolutely great work! And a overwhelming amount of experiments you have done starting from pre-processing and all the way to the inference. Also you seem to have picked up many things during this short time that took me couple of years to get a grasp over while doing PhD on this very topic. In my opinion publishing these results and answering those remaining questions that you have listed would get you far for a PhD, unless of course you already have one. 😃",
      "votes": 1,
      "replies": [
        {
          "id": 2180985,
          "postDate": "2023-03-14T08:23:21.857Z",
          "content": "<p>Thanks for you kind words and suggestions 😄 </p>\n<blockquote>\n  <p>unless of course you already have one</p>\n</blockquote>\n<p>I did not even have a Master's degree 😂</p>",
          "rawMarkdown": "Thanks for you kind words and suggestions 😄 \n>unless of course you already have one\n\nI did not even have a Master's degree 😂",
          "votes": 3
        }
      ]
    },
    {
      "id": 2169955,
      "postDate": "2023-03-05T15:17:35.710Z",
      "content": "<p>Thank you for your kind message! I'm looking forward to your training code! I have learned a lot from reading your sharing.</p>",
      "rawMarkdown": "Thank you for your kind message! I'm looking forward to your training code! I have learned a lot from reading your sharing.",
      "votes": 1,
      "replies": [
        {
          "id": 2170164,
          "postDate": "2023-03-05T18:19:09.730Z",
          "content": "<p>Thank you! I guess that the training code should be available in next 2 days.</p>",
          "rawMarkdown": "Thank you! I guess that the training code should be available in next 2 days.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2169869,
      "postDate": "2023-03-05T14:09:39.483Z",
      "content": "<p>Really great work. I like the augmentations, especially the different downscales. </p>",
      "rawMarkdown": "Really great work. I like the augmentations, especially the different downscales. ",
      "votes": 1,
      "replies": [
        {
          "id": 2169905,
          "postDate": "2023-03-05T14:37:13.920Z",
          "content": "<p>Thank you!</p>",
          "rawMarkdown": "Thank you!"
        }
      ]
    },
    {
      "id": 2388905,
      "postDate": "2023-08-13T17:07:32.820Z",
      "content": "<p>Interesting!!!</p>",
      "rawMarkdown": "Interesting!!!"
    },
    {
      "id": 2353153,
      "postDate": "2023-07-21T13:38:16.540Z",
      "content": "<p>Sorry to bother you. I want to reproduce this project, but I can't install nvidia-dali-nightly-cuda. Is this package necessary</p>",
      "rawMarkdown": "Sorry to bother you. I want to reproduce this project, but I can't install nvidia-dali-nightly-cuda. Is this package necessary"
    },
    {
      "id": 2313320,
      "postDate": "2023-06-22T15:27:29.573Z",
      "content": "<p>Great work:), congrats. I just read the solution and review a little bit your code, I would like to know what is the image type input for the windowing process? (I know it must be an array but like, an array from the original dicom image…?)<br>\nThx!</p>",
      "rawMarkdown": "Great work:), congrats. I just read the solution and review a little bit your code, I would like to know what is the image type input for the windowing process? (I know it must be an array but like, an array from the original dicom image...?)\nThx!",
      "replies": [
        {
          "id": 2315725,
          "postDate": "2023-06-24T11:03:16.300Z",
          "content": "<p>Yes, it's raw pixel array from the original dicom image, e.g <code>pydicom.dcmread(dcm_path).pixel_array</code></p>",
          "rawMarkdown": "Yes, it's raw pixel array from the original dicom image, e.g `pydicom.dcmread(dcm_path).pixel_array`",
          "votes": 1,
          "replies": [
            {
              "id": 2316028,
              "postDate": "2023-06-24T15:47:57.880Z",
              "content": "<p>Ohh I see, thank u very much. Considering working with the raw dicom images as arrays I just wonder how did you manage memory and computation resources?:c hahaha</p>",
              "rawMarkdown": "Ohh I see, thank u very much. Considering working with the raw dicom images as arrays I just wonder how did you manage memory and computation resources?:c hahaha"
            },
            {
              "id": 2316041,
              "postDate": "2023-06-24T15:58:07.553Z",
              "content": "<p>ah sorry, just a small note that during inference, windowing is applied to the ROI-cropped patch only, which then resized to smaller fixed size, e.g 2048x1024. But, this only saves a little computation 👀</p>",
              "rawMarkdown": "ah sorry, just a small note that during inference, windowing is applied to the ROI-cropped patch only, which then resized to smaller fixed size, e.g 2048x1024. But, this only saves a little computation 👀"
            }
          ]
        }
      ]
    },
    {
      "id": 2305243,
      "postDate": "2023-06-16T14:26:40.557Z",
      "content": "<p>Helpful information, thank you for sharing.</p>",
      "rawMarkdown": "Helpful information, thank you for sharing."
    },
    {
      "id": 2301571,
      "postDate": "2023-06-14T03:17:06.527Z",
      "content": "<p>Hi, thanks for the detailed writeup. I learnt many new things going through your codebase. I have a question. Winning model is an ensemble of four ConvNext models, did using an ensemble model show significantly better performance than a single model? Thanks</p>",
      "rawMarkdown": "Hi, thanks for the detailed writeup. I learnt many new things going through your codebase. I have a question. Winning model is an ensemble of four ConvNext models, did using an ensemble model show significantly better performance than a single model? Thanks",
      "replies": [
        {
          "id": 2301667,
          "postDate": "2023-06-14T04:57:53.917Z",
          "content": "<p>In my quick test without threshold tunning, single model on single fold could archive LB 0.59 and PB 0.54</p>",
          "rawMarkdown": "In my quick test without threshold tunning, single model on single fold could archive LB 0.59 and PB 0.54",
          "replies": [
            {
              "id": 2301697,
              "postDate": "2023-06-14T05:32:15.057Z",
              "content": "<p>Thanks. If I get this correctly, based on the pb scores single model trained on single fold performs just as well as the ensemble right? <br>\nI'm just curious as to why four convnext models were used</p>",
              "rawMarkdown": "Thanks. If I get this correctly, based on the pb scores single model trained on single fold performs just as well as the ensemble right? \nI'm just curious as to why four convnext models were used"
            },
            {
              "id": 2301713,
              "postDate": "2023-06-14T05:56:39.603Z",
              "content": "<p>Because that's all I had at that time :D. Winning submission was created and submitted in the last day of this competition and I have no other stronger models (e.g with different architectures or training strategies) at that time --&gt; no more complex ensemble should be tried.<br>\n4 is always better stabilization/generalization/score than 1 in almost every cases. In this particular case, 0.02 LB and 0.01 PB difference is not a negligible improvement.</p>",
              "rawMarkdown": "Because that's all I had at that time :D. Winning submission was created and submitted in the last day of this competition and I have no other stronger models (e.g with different architectures or training strategies) at that time --> no more complex ensemble should be tried.\n4 is always better stabilization/generalization/score than 1 in almost every cases. In this particular case, 0.02 LB and 0.01 PB difference is not a negligible improvement."
            },
            {
              "id": 2301722,
              "postDate": "2023-06-14T06:08:04.180Z",
              "content": "<p>Thank you very much for explaining it. </p>",
              "rawMarkdown": "Thank you very much for explaining it. "
            }
          ]
        }
      ]
    },
    {
      "id": 2261305,
      "postDate": "2023-05-16T08:32:56.707Z",
      "content": "<p>num_logits = 2 means using sigmoid (BCEWithLogitsLoss) for training and softmax .Can you explain why this strategy is used? Is there any basis for it? Usually the same activation function is used for training and inference.Thanks in advance.</p>",
      "rawMarkdown": "num_logits = 2 means using sigmoid (BCEWithLogitsLoss) for training and softmax .Can you explain why this strategy is used? Is there any basis for it? Usually the same activation function is used for training and inference.Thanks in advance.",
      "replies": [
        {
          "id": 2261316,
          "postDate": "2023-05-16T08:50:40.447Z",
          "content": "<p>Hi,<br>\nIt's just a baseline when I'm start with this competition. I experienced slightly better results with BCE instead of CCE for binary classification tasks in the past. Another reason is to easily integrating custom auxiliary losses (usually BCEs too) and balancing these loss weights, e.g all losses have same scales. Also, 2 logits --&gt; double the number of params in the linear head </p>",
          "rawMarkdown": "Hi,\nIt's just a baseline when I'm start with this competition. I experienced slightly better results with BCE instead of CCE for binary classification tasks in the past. Another reason is to easily integrating custom auxiliary losses (usually BCEs too) and balancing these loss weights, e.g all losses have same scales. Also, 2 logits --> double the number of params in the linear head ",
          "replies": [
            {
              "id": 2261636,
              "postDate": "2023-05-16T13:26:52.977Z",
              "content": "<p>Thank you for your reply. My question is why sigmoid is used in the training phase and softmax is used in the inference phase. This usage is very uncommon. Does it help the results?</p>",
              "rawMarkdown": "Thank you for your reply. My question is why sigmoid is used in the training phase and softmax is used in the inference phase. This usage is very uncommon. Does it help the results?"
            },
            {
              "id": 2261793,
              "postDate": "2023-05-16T14:51:04.733Z",
              "content": "<p>As you see the result (table above), sigmoid with num_logits=1 was used in my final submission since “sigmoid for training+softmax for inference” worsen the result (did not help)</p>",
              "rawMarkdown": "As you see the result (table above), sigmoid with num_logits=1 was used in my final submission since “sigmoid for training+softmax for inference” worsen the result (did not help)"
            },
            {
              "id": 2283399,
              "postDate": "2023-06-01T08:21:48.583Z",
              "content": "<p>Here you mentioned pos cases upsample, I would like to ask whether your upsample is simply copying samples or other methods so that each batch has at least one pos case.thank you!</p>",
              "rawMarkdown": "Here you mentioned pos cases upsample, I would like to ask whether your upsample is simply copying samples or other methods so that each batch has at least one pos case.thank you!"
            },
            {
              "id": 2285394,
              "postDate": "2023-06-02T17:13:03.367Z",
              "content": "<p>Yes, it's simply copying samples using a custom batch sampler</p>",
              "rawMarkdown": "Yes, it's simply copying samples using a custom batch sampler"
            }
          ]
        }
      ]
    },
    {
      "id": 2173792,
      "postDate": "2023-03-08T16:39:01.987Z",
      "content": "<p>Congrats for the win! Thanks for the detailed write up, I have learned a lot.</p>",
      "rawMarkdown": "Congrats for the win! Thanks for the detailed write up, I have learned a lot.",
      "replies": [
        {
          "id": 2180989,
          "postDate": "2023-03-14T08:25:21.330Z",
          "content": "<p>Thank you. Hope it helps</p>",
          "rawMarkdown": "Thank you. Hope it helps"
        }
      ]
    },
    {
      "id": 2172491,
      "postDate": "2023-03-07T14:58:01.233Z",
      "content": "<p>You're a living genius. Bravo</p>",
      "rawMarkdown": "You're a living genius. Bravo"
    },
    {
      "id": 2172374,
      "postDate": "2023-03-07T13:40:11.260Z",
      "content": "<p>Congratz on the win ! <br>\nAnd thanks for the really nice write-up :)</p>",
      "rawMarkdown": "Congratz on the win ! \nAnd thanks for the really nice write-up :)",
      "replies": [
        {
          "id": 2180992,
          "postDate": "2023-03-14T08:26:00.457Z",
          "content": "<p>Thanks, I also learned a lot from you 💯</p>",
          "rawMarkdown": "Thanks, I also learned a lot from you 💯"
        }
      ]
    },
    {
      "id": 2170596,
      "postDate": "2023-03-06T05:52:18.743Z",
      "content": "<p>Gem was not work for me.<br>\nI think Gem is very tricky.In the same competition , some people say it works, but others say it doesn't work</p>",
      "rawMarkdown": "Gem was not work for me.\nI think Gem is very tricky.In the same competition , some people say it works, but others say it doesn't work",
      "replies": [
        {
          "id": 2173793,
          "postDate": "2023-03-08T16:39:42.787Z",
          "content": "<p>What is GEM?</p>",
          "rawMarkdown": "What is GEM?",
          "replies": [
            {
              "id": 2180999,
              "postDate": "2023-03-14T08:28:41.263Z",
              "content": "<blockquote>\n  <p>What is GEM?</p>\n</blockquote>\n<p>Generalized Mean Pooling</p>",
              "rawMarkdown": "> What is GEM?\n\nGeneralized Mean Pooling"
            }
          ]
        }
      ]
    },
    {
      "id": 2170123,
      "postDate": "2023-03-05T17:52:40.727Z",
      "content": "<p>Many thanks for the elaborate write-up and for sharing your reasonings throughout it!</p>\n<p>As you comment in <a href=\"https://github.com/dangnh0611/kaggle_rsna_breast_cancer/blob/dev/notebooks/roi_yolov5.ipynb\" target=\"_blank\">https://github.com/dangnh0611/kaggle_rsna_breast_cancer/blob/dev/notebooks/roi_yolov5.ipynb</a> can you share the annotation dataset and the yolov5 training notebook?</p>\n<p>Looking forward to study training code in the dev branch even before you publish the refactored one, if possible.</p>",
      "rawMarkdown": "Many thanks for the elaborate write-up and for sharing your reasonings throughout it!\n\nAs you comment in https://github.com/dangnh0611/kaggle_rsna_breast_cancer/blob/dev/notebooks/roi_yolov5.ipynb can you share the annotation dataset and the yolov5 training notebook?\n\nLooking forward to study training code in the dev branch even before you publish the refactored one, if possible.",
      "replies": [
        {
          "id": 2170160,
          "postDate": "2023-03-05T18:16:24.280Z",
          "content": "<p>I'm sorry, but the notebook you mentioned should be move to <code>notebooks/3rd/</code> directory instead since i just download <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> 's <a href=\"https://www.kaggle.com/code/remekkinas/breast-cancer-roi-brest-extractor\" target=\"_blank\">notebook</a>. Giving credit to his awesome discussions (<a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/371630\" target=\"_blank\">here</a> and <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369749\" target=\"_blank\">here</a>), <a href=\"https://www.kaggle.com/code/remekkinas/roi-detector-yolov5-training-and-annotations/notebook\" target=\"_blank\">yolov5 training notebook</a>, <a href=\"https://www.kaggle.com/code/remekkinas/breast-cancer-roi-brest-extractor\" target=\"_blank\">yolov5 inference notebook</a> and <a href=\"https://www.kaggle.com/datasets/remekkinas/rsna-roi-detector-annotations-yolo\" target=\"_blank\">annotated dataset</a>. </p>",
          "rawMarkdown": "I'm sorry, but the notebook you mentioned should be move to `notebooks/3rd/` directory instead since i just download @remekkinas 's [notebook](https://www.kaggle.com/code/remekkinas/breast-cancer-roi-brest-extractor). Giving credit to his awesome discussions ([here](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/371630) and [here](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369749)), [yolov5 training notebook](https://www.kaggle.com/code/remekkinas/roi-detector-yolov5-training-and-annotations/notebook), [yolov5 inference notebook](https://www.kaggle.com/code/remekkinas/breast-cancer-roi-brest-extractor) and [annotated dataset](https://www.kaggle.com/datasets/remekkinas/rsna-roi-detector-annotations-yolo). ",
          "votes": 1,
          "replies": [
            {
              "id": 2170342,
              "postDate": "2023-03-05T22:09:35.930Z",
              "content": "<p>Thank you.</p>\n<p>Training code of the Convnextv1-small classification models is available at <code>dev/src/pytorch-image-models/projects/rsna</code>, correct? Thanks again.</p>",
              "rawMarkdown": "Thank you.\n\nTraining code of the Convnextv1-small classification models is available at `dev/src/pytorch-image-models/projects/rsna`, correct? Thanks again."
            },
            {
              "id": 2180995,
              "postDate": "2023-03-14T08:27:29.507Z",
              "content": "<p>Yes. And I have updated <a href=\"https://github.com/dangnh0611/kaggle_rsna_breast_cancer\" target=\"_blank\">the code</a>. Thanks for your attention 😄</p>",
              "rawMarkdown": "Yes. And I have updated [the code](https://github.com/dangnh0611/kaggle_rsna_breast_cancer). Thanks for your attention 😄"
            }
          ]
        }
      ]
    },
    {
      "id": 2169724,
      "postDate": "2023-03-05T11:32:39.940Z",
      "content": "<p>Thank you for your kind share! Do you train single-view models or dual-view models to achieve such an impressive performance?</p>",
      "rawMarkdown": "Thank you for your kind share! Do you train single-view models or dual-view models to achieve such an impressive performance?",
      "replies": [
        {
          "id": 2169730,
          "postDate": "2023-03-05T11:39:21.357Z",
          "content": "<p>Hi, I just tried single-view models. Multi-views model is potential so I'll give it a try later 😊</p>",
          "rawMarkdown": "Hi, I just tried single-view models. Multi-views model is potential so I'll give it a try later 😊"
        }
      ]
    },
    {
      "id": 2893530,
      "postDate": "2024-06-27T21:09:34.420Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2283395,
      "postDate": "2023-06-01T08:20:58.220Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2261301,
      "postDate": "2023-05-16T08:29:20.033Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2175611,
      "postDate": "2023-03-10T02:25:04.217Z",
      "rawMarkdown": "",
      "votes": -1,
      "isDeleted": true,
      "replies": [
        {
          "id": 2180987,
          "postDate": "2023-03-14T08:24:59.253Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2169786,
      "author_name": "sandy1112",
      "author_url": "",
      "post_date": "2023-03-05T12:35:48.053000",
      "content": "<p>Thanks a lot for sharing such a detailed write up. Congratulations on your Gold… looking at the way you broke down the problem, its a well deserved gold. In my experiments also found that Maxpooling worked the best, while average pooling tended to wash away the required signals</p>",
      "votes": 9,
      "replies": [
        {
          "id": 2169862,
          "author_name": "Đăng Nguyễn Hồng",
          "author_url": "",
          "post_date": "2023-03-05T13:59:42.743000",
          "content": "<p>Thanks for your kind words !</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2191737,
      "author_name": "Reyihi",
      "author_url": "",
      "post_date": "2023-03-22T07:14:32.980000",
      "content": "<p>Could you share more training parameters of ROI detection model? I can't get the same AP@0.5-0.95 as you mentioned above. I only train on <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a>'s dataset. thank you!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2192430,
          "author_name": "Đăng Nguyễn Hồng",
          "author_url": "",
          "post_date": "2023-03-22T16:44:04.817000",
          "content": "<p>Hi, insufficient number of val samples and personal labeling bias can make the AP@0.5-0.95 unstable. <br>\nFor YOLOX, I use <a href=\"https://github.com/dangnh0611/kaggle_rsna_breast_cancer/blob/reproduce/src/roi_det/YOLOX/exps/projects/rsna/yolox_nano_bre_416.py\" target=\"_blank\">this hyperparams config</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2174211,
      "author_name": "Jhye O'Meley",
      "author_url": "",
      "post_date": "2023-03-09T00:18:06.917000",
      "content": "<p>Really great work. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2172703,
      "author_name": "Pablo Larrosa",
      "author_url": "",
      "post_date": "2023-03-07T18:05:37.537000",
      "content": "<p>Hi, you mentioned that you used image size 2048x1024 with Convnextv1-small model, but when you processed the images yolox produce as output image of 416?.</p>\n<p>Thank in advance!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2180975,
          "author_name": "Đăng Nguyễn Hồng",
          "author_url": "",
          "post_date": "2023-03-14T08:19:48.853000",
          "content": "<p>Hi, i get the bounding box output from YOLOX, then the coordinates will be used to crop the original resolution image, as described <a href=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10254700%2F9de52f41201b4a92774f5fee331cb879%2Fwriteup-Page-1.drawio.png?generation=1678001361222187&amp;alt=media\" target=\"_blank\">here</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2171012,
      "author_name": "Abdallah Ragab223",
      "author_url": "",
      "post_date": "2023-03-06T13:11:48.213000",
      "content": "<p>great work, and thanks to share the solution </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2170281,
      "author_name": "Antti Isosalo",
      "author_url": "",
      "post_date": "2023-03-05T20:11:53.087000",
      "content": "<p>Absolutely great work! And a overwhelming amount of experiments you have done starting from pre-processing and all the way to the inference. Also you seem to have picked up many things during this short time that took me couple of years to get a grasp over while doing PhD on this very topic. In my opinion publishing these results and answering those remaining questions that you have listed would get you far for a PhD, unless of course you already have one. 😃</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2180985,
          "author_name": "Đăng Nguyễn Hồng",
          "author_url": "",
          "post_date": "2023-03-14T08:23:21.857000",
          "content": "<p>Thanks for you kind words and suggestions 😄 </p>\n<blockquote>\n  <p>unless of course you already have one</p>\n</blockquote>\n<p>I did not even have a Master's degree 😂</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2169955,
      "author_name": "Mr.G",
      "author_url": "",
      "post_date": "2023-03-05T15:17:35.710000",
      "content": "<p>Thank you for your kind message! I'm looking forward to your training code! I have learned a lot from reading your sharing.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2170164,
          "author_name": "Đăng Nguyễn Hồng",
          "author_url": "",
          "post_date": "2023-03-05T18:19:09.730000",
          "content": "<p>Thank you! I guess that the training code should be available in next 2 days.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2169869,
      "author_name": "Darragh",
      "author_url": "",
      "post_date": "2023-03-05T14:09:39.483000",
      "content": "<p>Really great work. I like the augmentations, especially the different downscales. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2169905,
          "author_name": "Đăng Nguyễn Hồng",
          "author_url": "",
          "post_date": "2023-03-05T14:37:13.920000",
          "content": "<p>Thank you!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2388905,
      "author_name": "继续战斗继续前进",
      "author_url": "",
      "post_date": "2023-08-13T17:07:32.820000",
      "content": "<p>Interesting!!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2353153,
      "author_name": "HANF Wang",
      "author_url": "",
      "post_date": "2023-07-21T13:38:16.540000",
      "content": "<p>Sorry to bother you. I want to reproduce this project, but I can't install nvidia-dali-nightly-cuda. Is this package necessary</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2313320,
      "author_name": "Pedro Martínez Barrón",
      "author_url": "",
      "post_date": "2023-06-22T15:27:29.573000",
      "content": "<p>Great work:), congrats. I just read the solution and review a little bit your code, I would like to know what is the image type input for the windowing process? (I know it must be an array but like, an array from the original dicom image…?)<br>\nThx!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2315725,
          "author_name": "Đăng Nguyễn Hồng",
          "author_url": "",
          "post_date": "2023-06-24T11:03:16.300000",
          "content": "<p>Yes, it's raw pixel array from the original dicom image, e.g <code>pydicom.dcmread(dcm_path).pixel_array</code></p>",
          "votes": 1,
          "replies": [
            {
              "id": 2316028,
              "author_name": "Pedro Martínez Barrón",
              "author_url": "",
              "post_date": "2023-06-24T15:47:57.880000",
              "content": "<p>Ohh I see, thank u very much. Considering working with the raw dicom images as arrays I just wonder how did you manage memory and computation resources?:c hahaha</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2316041,
              "author_name": "Đăng Nguyễn Hồng",
              "author_url": "",
              "post_date": "2023-06-24T15:58:07.553000",
              "content": "<p>ah sorry, just a small note that during inference, windowing is applied to the ROI-cropped patch only, which then resized to smaller fixed size, e.g 2048x1024. But, this only saves a little computation 👀</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2305243,
      "author_name": "Lidiia Kachmarska",
      "author_url": "",
      "post_date": "2023-06-16T14:26:40.557000",
      "content": "<p>Helpful information, thank you for sharing.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2301571,
      "author_name": "Bharys J",
      "author_url": "",
      "post_date": "2023-06-14T03:17:06.527000",
      "content": "<p>Hi, thanks for the detailed writeup. I learnt many new things going through your codebase. I have a question. Winning model is an ensemble of four ConvNext models, did using an ensemble model show significantly better performance than a single model? Thanks</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2301667,
          "author_name": "Đăng Nguyễn Hồng",
          "author_url": "",
          "post_date": "2023-06-14T04:57:53.917000",
          "content": "<p>In my quick test without threshold tunning, single model on single fold could archive LB 0.59 and PB 0.54</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2301697,
              "author_name": "Bharys J",
              "author_url": "",
              "post_date": "2023-06-14T05:32:15.057000",
              "content": "<p>Thanks. If I get this correctly, based on the pb scores single model trained on single fold performs just as well as the ensemble right? <br>\nI'm just curious as to why four convnext models were used</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2301713,
              "author_name": "Đăng Nguyễn Hồng",
              "author_url": "",
              "post_date": "2023-06-14T05:56:39.603000",
              "content": "<p>Because that's all I had at that time :D. Winning submission was created and submitted in the last day of this competition and I have no other stronger models (e.g with different architectures or training strategies) at that time --&gt; no more complex ensemble should be tried.<br>\n4 is always better stabilization/generalization/score than 1 in almost every cases. In this particular case, 0.02 LB and 0.01 PB difference is not a negligible improvement.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2301722,
              "author_name": "Bharys J",
              "author_url": "",
              "post_date": "2023-06-14T06:08:04.180000",
              "content": "<p>Thank you very much for explaining it. </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2261305,
      "author_name": "CJLUer",
      "author_url": "",
      "post_date": "2023-05-16T08:32:56.707000",
      "content": "<p>num_logits = 2 means using sigmoid (BCEWithLogitsLoss) for training and softmax .Can you explain why this strategy is used? Is there any basis for it? Usually the same activation function is used for training and inference.Thanks in advance.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2261316,
          "author_name": "Đăng Nguyễn Hồng",
          "author_url": "",
          "post_date": "2023-05-16T08:50:40.447000",
          "content": "<p>Hi,<br>\nIt's just a baseline when I'm start with this competition. I experienced slightly better results with BCE instead of CCE for binary classification tasks in the past. Another reason is to easily integrating custom auxiliary losses (usually BCEs too) and balancing these loss weights, e.g all losses have same scales. Also, 2 logits --&gt; double the number of params in the linear head </p>",
          "votes": 0,
          "replies": [
            {
              "id": 2261636,
              "author_name": "CJLUer",
              "author_url": "",
              "post_date": "2023-05-16T13:26:52.977000",
              "content": "<p>Thank you for your reply. My question is why sigmoid is used in the training phase and softmax is used in the inference phase. This usage is very uncommon. Does it help the results?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2261793,
              "author_name": "Đăng Nguyễn Hồng",
              "author_url": "",
              "post_date": "2023-05-16T14:51:04.733000",
              "content": "<p>As you see the result (table above), sigmoid with num_logits=1 was used in my final submission since “sigmoid for training+softmax for inference” worsen the result (did not help)</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2283399,
              "author_name": "CJLUer",
              "author_url": "",
              "post_date": "2023-06-01T08:21:48.583000",
              "content": "<p>Here you mentioned pos cases upsample, I would like to ask whether your upsample is simply copying samples or other methods so that each batch has at least one pos case.thank you!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2285394,
              "author_name": "Đăng Nguyễn Hồng",
              "author_url": "",
              "post_date": "2023-06-02T17:13:03.367000",
              "content": "<p>Yes, it's simply copying samples using a custom batch sampler</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2173792,
      "author_name": "Fernando Cossio",
      "author_url": "",
      "post_date": "2023-03-08T16:39:01.987000",
      "content": "<p>Congrats for the win! Thanks for the detailed write up, I have learned a lot.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2180989,
          "author_name": "Đăng Nguyễn Hồng",
          "author_url": "",
          "post_date": "2023-03-14T08:25:21.330000",
          "content": "<p>Thank you. Hope it helps</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2172491,
      "author_name": "Ifeanyichukwu Nwobodo",
      "author_url": "",
      "post_date": "2023-03-07T14:58:01.233000",
      "content": "<p>You're a living genius. Bravo</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2172374,
      "author_name": "Theo Viel",
      "author_url": "",
      "post_date": "2023-03-07T13:40:11.260000",
      "content": "<p>Congratz on the win ! <br>\nAnd thanks for the really nice write-up :)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2180992,
          "author_name": "Đăng Nguyễn Hồng",
          "author_url": "",
          "post_date": "2023-03-14T08:26:00.457000",
          "content": "<p>Thanks, I also learned a lot from you 💯</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2170596,
      "author_name": "fate",
      "author_url": "",
      "post_date": "2023-03-06T05:52:18.743000",
      "content": "<p>Gem was not work for me.<br>\nI think Gem is very tricky.In the same competition , some people say it works, but others say it doesn't work</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2173793,
          "author_name": "Fernando Cossio",
          "author_url": "",
          "post_date": "2023-03-08T16:39:42.787000",
          "content": "<p>What is GEM?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2180999,
              "author_name": "Đăng Nguyễn Hồng",
              "author_url": "",
              "post_date": "2023-03-14T08:28:41.263000",
              "content": "<blockquote>\n  <p>What is GEM?</p>\n</blockquote>\n<p>Generalized Mean Pooling</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2170123,
      "author_name": "Pablo Rios",
      "author_url": "",
      "post_date": "2023-03-05T17:52:40.727000",
      "content": "<p>Many thanks for the elaborate write-up and for sharing your reasonings throughout it!</p>\n<p>As you comment in <a href=\"https://github.com/dangnh0611/kaggle_rsna_breast_cancer/blob/dev/notebooks/roi_yolov5.ipynb\" target=\"_blank\">https://github.com/dangnh0611/kaggle_rsna_breast_cancer/blob/dev/notebooks/roi_yolov5.ipynb</a> can you share the annotation dataset and the yolov5 training notebook?</p>\n<p>Looking forward to study training code in the dev branch even before you publish the refactored one, if possible.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2170160,
          "author_name": "Đăng Nguyễn Hồng",
          "author_url": "",
          "post_date": "2023-03-05T18:16:24.280000",
          "content": "<p>I'm sorry, but the notebook you mentioned should be move to <code>notebooks/3rd/</code> directory instead since i just download <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> 's <a href=\"https://www.kaggle.com/code/remekkinas/breast-cancer-roi-brest-extractor\" target=\"_blank\">notebook</a>. Giving credit to his awesome discussions (<a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/371630\" target=\"_blank\">here</a> and <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369749\" target=\"_blank\">here</a>), <a href=\"https://www.kaggle.com/code/remekkinas/roi-detector-yolov5-training-and-annotations/notebook\" target=\"_blank\">yolov5 training notebook</a>, <a href=\"https://www.kaggle.com/code/remekkinas/breast-cancer-roi-brest-extractor\" target=\"_blank\">yolov5 inference notebook</a> and <a href=\"https://www.kaggle.com/datasets/remekkinas/rsna-roi-detector-annotations-yolo\" target=\"_blank\">annotated dataset</a>. </p>",
          "votes": 1,
          "replies": [
            {
              "id": 2170342,
              "author_name": "Pablo Rios",
              "author_url": "",
              "post_date": "2023-03-05T22:09:35.930000",
              "content": "<p>Thank you.</p>\n<p>Training code of the Convnextv1-small classification models is available at <code>dev/src/pytorch-image-models/projects/rsna</code>, correct? Thanks again.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2180995,
              "author_name": "Đăng Nguyễn Hồng",
              "author_url": "",
              "post_date": "2023-03-14T08:27:29.507000",
              "content": "<p>Yes. And I have updated <a href=\"https://github.com/dangnh0611/kaggle_rsna_breast_cancer\" target=\"_blank\">the code</a>. Thanks for your attention 😄</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2169724,
      "author_name": "kingjohnson",
      "author_url": "",
      "post_date": "2023-03-05T11:32:39.940000",
      "content": "<p>Thank you for your kind share! Do you train single-view models or dual-view models to achieve such an impressive performance?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2169730,
          "author_name": "Đăng Nguyễn Hồng",
          "author_url": "",
          "post_date": "2023-03-05T11:39:21.357000",
          "content": "<p>Hi, I just tried single-view models. Multi-views model is potential so I'll give it a try later 😊</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2893530,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-06-27T21:09:34.420000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2283395,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-06-01T08:20:58.220000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2261301,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-05-16T08:29:20.033000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2175611,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-03-10T02:25:04.217000",
      "content": "",
      "votes": -1,
      "replies": [
        {
          "id": 2180987,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-03-14T08:24:59.253000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2169701": "***Update 10/04/2023***\n\nAblation study: despite of a large improvement on a particular fold (below), soft positive label trick does not show a clearly improvement in performance over the standard label smoothing technique. The use of external datasets improve F1-score about 0.02 on local OOF validation and Private Leaderboard (with extractly same training pipeline + hyper-params).\n\n| External data |            Loss           |   OOF F1   |    LB    |    PL    |\n|:-------------:|:-------------------------:|:----------:|:--------:|:--------:|\n|       x       |    label smoothing=0.1    |  0.4921    |   0.60   |   0.53   |\n|       x       |  soft positive label=0.8  |   0.4853   |   0.60   |   0.53   |\n|       ✓       |   label smoothing = 0.1   |   0.5161   |   0.58   | **0.56** |\n|       ✓       | soft positive label = 0.9 | **0.5182** | **0.61** |   0.55   |\n\n\n-----\n\nFirst of all, I would like to thank Kaggle and competition's host for such an amazing challenge, a lofty goal with high data quality. Thank you to all paticipants/kagglers with many active and helpful discussions/codes. My solution was just built up from every pieces of kindly shares from you. I learned a lot and I'm very appreciated for that.\nI'm also very happy and suprised with the 1st place. This is my first gold medal and I'm writing my first writeup. It was such a great journey for me.\nFor the solution, I use a very simple pipeline which can be described in just few lines:\n- Use some external datasets: VinDr-Mammo, MiniDDSM, CMMD, CDD-CESM, BMCD.\n- 4 x Convnextv1-small 2048x1024, validated on 4-folds splits of competition data.\n- Soft positive label\n\n![inference pipeline](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10254700%2F15d5d89bcd9cd18e45b6da133d5f5ad7%2Fpreprocess.jpg?generation=1685726162522472&alt=media)\n\nNow I want to share some experiments and my thought about those. Many of theme could be found in another discussions by excellent kagglers. Many of theme seem obvious. Hope this helps some new comer getting started in the future. Kindly note that it's just my own opinion/thoughts with very limited experiments and knownledge. I'm appreciated for your discussions and feel free to correct me if something was wrong.\n\n# 1. ROI crop\nROI cropping was performed since it effectively help keeping more texture/detail given a fixed resolution. I use YOLOX-nano 416x416 for ROI detector. The advantage of DL detector vs rule-based methods is the obtained bbox is smaller, aspect ratio is more stable and focus to the breast region.\n1. Train a YOLOX on @remekkinas 's [dataset](https://www.kaggle.com/datasets/remekkinas/rsna-roi-detector-annotations-yolo) (472 bbox-annotated images)\n2. Inference on all available training data with low `conf_thres` and high `iou_thres`. Only 3 miss-detected images (all contain noise) and over 100 images with 2 boxes (almost overlapped). I manually select and label 99 of those images. Therefore, I have 571 annotated images in total.\n3. Retrain YOLOX on new images: 521 for train, 50 for val. Note that these 50 val images include all 47 val images of original @remekkinas 's dataset. The new dataset version contain original resolution images (same as original dicoms), preprocessed with simple min-max normalization. Tried various model/image sizes and finally choose YOLOX-nano 416x416 as final model due to the consistent result and small overhead.\n\n|   **model size**   | **image size** | **interpolation** | **AP_new_val** | **AP_remek_val** |\n|:------------------:|:--------------:|:-----------------:|:----------:|:--------------:|\n| _nano (selection)_ |      _416_     |      _LINEAR_     |   _96.26_  |     _94.21_    |\n|        nano        |       416      |        AREA       |    94.09   |      91.60      |\n|        nano        |       640      |       LINEAR      |    95.85   |      88.40      |\n|        nano        |       768      |       LINEAR      |    96.22   |      82.09     |\n|        nano        |      1024      |       LINEAR      |    94.92   |      89.40      |\n|        tiny        |       416      |       LINEAR      |    94.23   |      90.20      |\n|        tiny        |       640      |       LINEAR      |    94.95   |      89.84     |\n|        tiny        |       768      |        AREA       |    96.21   |      68.03     |\n|        tiny        |      1024      |        AREA       |    93.69   |      73.70      |\n|          s         |       416      |       LINEAR      |    95.03   |      0.86      |\n|          s         |       640      |       LINEAR      |    96.10    |      70.80      |\n|          s         |       768      |       LINEAR      |    96.79   |      78.70      |\n\n\nAP@0.5 is 1.0 in all experiments. We see a large gap between AP@0.5-0.95 between two validation sets. Some reasons for that:\n- New version add more 3/50 typical hard cases.\n- Inconsistent processing pipeline: val images in Remek's val was resized 2 times (original --> 1024 --> 416) \n- Training images is annotated according to personal bias (no standard way/consentration to annotate the breast boxes correctly). So higher AP may not indicate a better model.\n- The validation size is also not large enough to judge\n- No hyper parameters tuning\n\nDid these things led to the large gap, particularly with stronger model and larger image size ?\nAll these efforts are just to ensure an \"as good as posible\" ROI detection model. I think @remekkinas 's [dataset](https://www.kaggle.com/datasets/remekkinas/rsna-roi-detector-annotations-yolo) is enough to train good YOLOX models and they could perform equally well in hidden test set.\n\nSimpler Otsu thresholding + findCountours() slightly modified from [this notebook](https://www.kaggle.com/code/snnclsr/roi-extraction-using-opencv) is used to find breast bbox as a fall back in case of YOLOX's miss-detection.\nOr, if both miss the breast box, just use the whole image without any cropping.\n\n# 2. The inference pipeline\nOperations on large array take time, so I try to transfer the computation task to GPU as much as posible.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10254700%2F9de52f41201b4a92774f5fee331cb879%2Fwriteup-Page-1.drawio.png?generation=1678001361222187&alt=media)\n\n# 3. Early experiments\nMy final solution use external datasets, but I stitch up with using only competition data for almost of the time (until \"7 days to go\"). Hence most of my experiments are done on competition data only: 5-folds splits with StratifiedGroupKFold based on `patient_id`. Training hyperparams used for the final solution are almost inherited from these early experiments.\n\n## 3.1. About the metric\nThe competition pF1 score is not stable and hard to track for me. Therefore, I mainly track my experiments based on multiple metrics: `{ PR_AUC, ROC_AUC, best_PF1 (binarized), best_threshold }` instead of just one.\n- **PR_AUC**: correlated with but more stable than best_PF1. It focuses on positive cases, and is strongly affected by prior data distribution (% of positive).\n- **ROC_AUC**: less affected by prior data distribution. Much more stable, but seem to be over optimistic which led to just a small gap between a good model and a bad model.\n- To get a high binaried pf1, model should not predict too many positives which usually led to large FP --> dramatically reduce best_PF1. A good scored pf1 model tends to prioritize Precision over Recall. I personaly don't like this behaviour, especially for real life application.\n\n## 3.2. Augmentations\nI stitch with this augmentation pipeline for all experiments, no tuning at all:\n\n```python\nA.Compose([\n    # crop, tweak from A.RandomSizedCrop()\n    custom_augs.CustomRandomSizedCropNoResize(scale=(0.5, 1.0), ratio=(0.5, 0.8), p=0.4),\n    # flip\n    A.HorizontalFlip(p=0.5),\n    A.VerticalFlip(p=0.5),\n    # downscale\n    A.OneOf([\n        A.Downscale(scale_min=0.75, scale_max=0.95, interpolation=dict(upscale=cv2.INTER_LINEAR, downscale=cv2.INTER_AREA), p=0.1),\n        A.Downscale(scale_min=0.75, scale_max=0.95, interpolation=dict(upscale=cv2.INTER_LANCZOS4, downscale=cv2.INTER_AREA), p=0.1),\n        A.Downscale(scale_min=0.75, scale_max=0.95, interpolation=dict(upscale=cv2.INTER_LINEAR, downscale=cv2.INTER_LINEAR), p=0.8),\n    ], p=0.125),\n    # contrast\n    A.OneOf([\n        A.RandomToneCurve(scale=0.3, p=0.5),\n        A.RandomBrightnessContrast(brightness_limit=(-0.1, 0.2), contrast_limit=(-0.4, 0.5), brightness_by_max=True, always_apply=False, p=0.5)\n    ], p=0.5),\n    # geometric\n    A.OneOf(\n        [\n            A.ShiftScaleRotate(shift_limit=None, scale_limit=[-0.15, 0.15], rotate_limit=[-30, 30], interpolation=cv2.INTER_LINEAR,\n                               border_mode=cv2.BORDER_CONSTANT, value=0, mask_value=None, shift_limit_x=[-0.1, 0.1],\n                               shift_limit_y=[-0.2, 0.2], rotate_method='largest_box', p=0.6),\n            A.ElasticTransform(alpha=1, sigma=20, alpha_affine=10, interpolation=cv2.INTER_LINEAR, border_mode=cv2.BORDER_CONSTANT,\n                               value=0, mask_value=None, approximate=False, same_dxdy=False, p=0.2),\n            A.GridDistortion(num_steps=5, distort_limit=0.3, interpolation=cv2.INTER_LINEAR, border_mode=cv2.BORDER_CONSTANT,\n                             value=0, mask_value=None, normalized=True, p=0.2),\n        ], p=0.5),\n    # random erase\n    A.CoarseDropout(max_holes=6, max_height=0.15, max_width=0.25, min_holes=1, min_height=0.05, min_width=0.1,\n                    fill_value=0, mask_fill_value=None, p=0.25),\n    ], p=0.9)\n```\nFor the random crop choice: real breast size/ratio vary largly between images --> popular pipeline of `longest resize + padding` introduces multi-scales problem. Of course, it would introduce higher risks of wrong positive label.\n\n*Example batch*\n![example batch](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10254700%2F4f149f8617c07c9a8e0d8e571f6a68fa%2Fexample_batch.jpg?generation=1685726233222765&alt=media)\n\n\n## 3.3. Up/down sampling\nI upsample pos cases in each epoch for all of my experiments.\n- Ensuring at least 1 pos in a batch/iteration is pretty important to stablize training. I found training difficult while set 0.5 pos per batch.\n- Upsampling ratio can largly affect CV score and prediction distribution. It also vary between backbones and other hyperparams choices. However, I prefer smallest pos/neg ratio as posible (since it's near to the real data distribution) but ensure at least 1 pos/batch.\n- Large pos/neg ratio helps training faster in early epochs. I tried linearly increase/decrease the pos/neg ratio between epochs to face with some problems of prediction distribution/threshold (especially EffB4). But in the end, i got no improvement in CV.\n\n## 3.4. Model/backbone\nI tried `Eff-B2`, `Eff-B4`, `Effv2-s` and `Convnextv1-small`\n- Each model has its own characteristic and training phenomenon.\n- All models could perform equally well in local CV. Except that `Convnextv1-small` give higher CV score.\n- `EfficientNet` (no model EMA) tends to overfit quickly after fews epochs with high pos/neg ratio: longer training reduce AUC largely and may slightly increase best_pf1 --> model tends to predict less positives. Small pos/neg ratio helps training more stable but reduce CV. Linearly increase pos/neg ratio between epochs (by a sampler) did not help much.\n- `Convnext-small` (with/without EMA) shows both stable training and better CV.\n\n## 3.5. drop_rate, drop_path_rate\nPlaying with drop_rate and drop_path_rate:\n- We can use large dropout rate of >= 0.5 to regularize training and reduce overfiting.\n- With very large drop_rate = 0.9 or drop_path_rate = 0.5, I still can get a \"not bad as expected\" model in CV score. The following results is for Eff-B4, pos/neg = 1/3 on fold 0:\n\n| drop_rate | drop_path_rate |  auc  | best_pf1 | best_thres | epoch |\n|:---------:|:--------------:|:-----:|:--------:|:----------:|:-----:|\n|    0.9    |       0.2      | 91.90 |   47.73  |    0.78    |   4   |\n|    0.7    |       0.2      | 90.36 |   52.27  |    0.55    |   3   |\n|    0.5    |       0.2      | 90.02 |   50.00  |    0.72    |   4   |\n|    0.3    |       0.2      | 91.23 |   48.24  |    0.82    |   4   |\n|    0.5    |       0.5      | 90.45 |   46.33  |    0.95    |   4   |\n\nHowever, I set drop_rate = 0.5 and drop_path_rate = 0.2 for most of my experiments, including the final ones.\n\n## 3.6. Global pooling\nI stitch with max pooling for almost my experiments as my inductive bias:\n\n- Max pooling is suitable and seem to be effective for anomalies detection or \"needle in the haystack\" tasks in literature. In this case, cancer may appear in a very small region and the rest are all normal.\n- `max()` provides stronger learning signal, but less stable than mean() in term of gradient.\n- My guess: `gem` > `max` > `mean` when enough data provided.\n\n## 3.7. Soft positive label/ Positive label smoothing\nConvnextv1-small look good in CV scores with stable AUC, PR_AUC, best PF1 across epochs. But there're differences in behaviour between Effv2-s and Convnextv1-small especially in best threshold for image/breast level.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10254700%2F6771d0c6e87067873e53021742e7bf6b%2Fwriteup_smoothing.png?generation=1678006236367845&alt=media)\n*The two lower green ones belong to `Effv2-s` and the other belong to `Convnextv1-small`*\n\nSome discussions suggest smaller best threshold (<0.55) may indicate a better model. For single-image, Convnext show a very high threshold of > 0.92, which could indicate the problem of over-confident. Stronger models with larger number of parameters is easier to be over-confident or overfitted, especialy in this highly imballanced dataset scenario. About the above figure, `label_smoothing = 0.1` was used but seem like it was not enough.\nSo, just add harder label smoothing to regularize training. Or use positive weight < 1.0 to reduce the priority of positive samples.\n\n\n| loss                            | num_logits | target {neg, pos}              | pr_auc     | roc_auc       | best_pf1   | best_thres | epoch |\n|---------------------------------|------------|--------------------------------|------------|-----------|------------|------------|-------|\n| _(baseline) bce_smooth 0.1_     | 2          | { [0.95, 0.05], [0.05, 0.95] } | 0.4755     | 0.9278    | 0.497      | 0.66       | 14    |\n| bce_smooth 0.4                  | 2          | { [0.8, 0.2], [0.2, 0.8] }     | 0.4749     | 0.9248    | 0.5        | 0.6        | 25    |\n| bce_pos_smooth 0.4              | 2          | { [1.0, 0.0], [0.2, 0.8] }     | 0.5191     | 0.9153    | 0.5488     | 0.53       | 13.5  |\n| _(best) bce_pos_smooth 0.2_     | 1          | { 0.0, 0.8 }                   | **0.5401** | 0.9281    | **0.5714** | 0.49       | 20    |\n| bce_pos_smooth 0.3              | 1          | { 0.0, 0.7 }                   | 0.522      | **0.933** | 0.517      | 0.5        | 17    |\n| bce_smooth 0.1 + pos_weight 0.4 | 1          | { 0.05, 0.95 }                 | 0.4946     | 0.9146    | 0.5393     | 0.39       | 19    |\n\n*Note:*\n- *Table above is fold 0 CV results*\n- *`num_logits = 2` means using `sigmoid` (BCEWithLogitsLoss) for training and `softmax` for inference. Refer [here](https://github.com/ultralytics/yolov5/issues/5401).*\n\nSoft positive labeling look reasonable: we have per-breast label and not per-image label. For some images belong to same patient, cancer signal may not appears clearly in some images, or even all images (MG is not enough to judge for cancer/non-cancer) --> the positive label should not be the maximum bound value of 1.0, but less confident.\n\nSoft postive label trick improve CV and helps threshold looks much better.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10254700%2Fc36afde0a06afbeb3dbd38a76753276c%2Fpos_smooth_thres.png?generation=1678007271357249&alt=media)\n\nAs some discussions, very sharp prediction distribution may indicate worse result/generalization.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10254700%2F7caf77b112791cd16eb20de6f6d443bd%2Fdistribution.png?generation=1678007391776163&alt=media)\n\n\n---------------------\n# 4. Final experiments\n## 4.1. External datasets\nA week left to the competition deadline, I was thinking about training final experiments for the final submission and should not make any mistakes or missing something. I read some discussions again and relized I was missing a big part: external data. In particular, external data contains a large number of positive cases which are valuable.\nThese external datasets summary:\n\n| Dataset     | num_patients* | num_samples* | num_pos_samples* | \n|-------------|---------------|--------------|------------------|\n| [VinDr-Mammo](https://physionet.org/content/vindr-mammo/1.0.0/) | 5000          | 20000        | 226 (1.13 %)     | \n| [MiniDDSM](https://www.kaggle.com/datasets/cheddad/miniddsm2)   | 1952          | 7808         | 1480 (18.95 %)   |\n| [CMMD](https://wiki.cancerimagingarchive.net/pages/viewpage.action?pageId=70230508)        | 1775          | 5202         | 2632 (50.6%)     |\n| [CDD-CESM](https://wiki.cancerimagingarchive.net/pages/viewpage.action?pageId=109379611)    | 326           | 1003         | 331 (33 %)       | \n| [BMCD](https://zenodo.org/record/5036062)        | 82            | 328          | 22 (6.71 %)      | [1, 2]    |\n| All         | 9135          | 34341        | 4691 (13.66 %)   |\n\n***** The number may not indicate original dataset characteristics, but processed data I used for this competition.\n\n**Some details:**\n1. **VinDr-Mammo**: contains BIRADS scores for each image of 0-5. I treated BIRADS 5 as cancer (1) and all other as normal (0). With only Digital Mammograms and BIRADS categories, one can't confirm 100% if a case is cancer or not . BIRADS 4 indicate 30% chance of cancer, then I was treated it as normal. My decision only happen in just a few seconds as i could remember. Reading other posts, i'm feeling my decision is not as good, except that it helps reduce sensitivity and can \"improve\" the pF1 (I don't want to see it that way). Maybe I make a huge mistake here. Better solution is to use soft/uncertain label or pseudo labeling for these ambigous (BIRADS-4) cases instead. Some images has LUTDescriptor. The image look over-exposured when apply VOILUT (voi + windowing), so I just apply windowing on this dataset, equivalent to pydicom's `apply_voi_lut(prefer_lut = False)`\n\n2. **MiniDDSM**: I used all 7808 samples. Found that the status label of `{Cancer, Benign, Normal}` is per patient_id, not per laterality. So I treat a patient-laterality as cancer if and only if status == 'Cancer' and at least 1 supicious region annotation (segmentation map) for that laterality is available. End up in 1480 positives and the remaining 6318 negatives. I use 16-bits png part for the less information loss. No windowing parameters as presented. Since there're watermarks noise with very high pixel intensity in ROI crop, percentile min-max scaled  was performed instead of min-max scaled for normalization.\n\n3. **CMMD**: Total of 5202 breast images belong to 1872 patient id. Note that some patient ids start with 'D2' are almost malignant and usually had label for one laterality only. For those cases, I treated the other laterality (no laterality-level label specified in csv file, but still have image) as normal (EDA from the competition data show that cancer only appears in one laterality). Original dicom images are in 8-bits depth with windowing parameters available.\n\n4. **CDD-CESM**: consisting Contrast-enhanced spectral mammography (CESM) images. This dataset contain label of `{Normal, Malignant, Benign}`. I treated Malignant as cancer, Normal or Benign as normal and only use the **low-energy images** as it is comparable to digital mammograms (MG), or at least they look pretty similar for me. Low-energy images is in 8-bits jpeg, no windowing information.\n\n5. **BMCD**: contains 100 patients (50 normal + 50 suspicious cases) with 82 biopsy-confirmed cases of `{'NORMAL', 'BENIGN', 'DCIS', 'MALIGNANT'}` and mammogram images of them at the time of screening and  avg 2.2 year before. I treat 'DCIS' or 'MALIGNANT' patient's last screening images as cancer and all the remaining as normal. Original dicom images is in 16-bits depth and windowing parameters are available.\n\n## 4.2. Validation strategy\nI found inconsistence in CV between 5-folds splits of competition data, probably because the number of positive is not sufficient. Although, hidden test should has distribution/property closer to the competition data, so I change validation strategy to 4-splits as:\n- Do 4-folds splitting on competition data\n- Use 1 fold for validation, the rest 3 folds + all external data for training.\nThen, for each split, training data contain about 5560/75400 positive cases (~7.38 %).\n\n## 4.3. Training\nI managed to get 4 x Convnextv1-small corresponding to the above 4 splits. Some unexpected results were founded during training, so the training stages was changed and in short consist of:\n1. Train 2 models on fold 0 and fold 1 with `soft_pos_label = 0.8`\n2. Train 2 models on fold 2 and fold 3 with `soft_pos_label = 0.9`\n3. Finetune 2 models obtained from stage 1 on fold 0 and fold 1 with `soft_pos_label = 0.9`\n\nIn details, I start training on first two folds: fold 0 and fold 1 with the following config:\n- Model: timm's convnext_small.fb_in22k_ft_in1k_384\n- Input size: 2048x1024\n- Loss: vanila BCE (no class weight)\n- Sampler: upsampling pos samples per epoch to pos/neg = 1/7, ensure each batch contains at least 1 pos sample.\n- Batchsize: 8\n- Automatic Mixed Precision (AMP): enable\n- Model EMA: enable\n- Global pooling: max\n- Soft positive label = 0.8\n- Optimizer: SGD with momemtum=0.9\n- Scheduler: Cosine lr decay(epoch = 24, lr = 1e-3, min_lr = 1e-5) + linear warmup(warmup_lr = 1e-5, warmup_epoch = 4)\n- Drop_rate = 0.5, drop_path_rate = 0.2\n\nOnce training finished, results are not as my expectation on fold 0:\n- CV results are not good\n- Threshold is much smaller. I expected it to be in the range [0.35, 0.5], but it's just around 0.25+-0.02\n- Training is not converged yet. I guess that CV can be improved with more additional training epochs.\n\nSo I start train fold 2 and fold 3 with few changes: longer training with larger learning rate and reduce soft positive label.\n- Scheduler: Cosine lr decay(epoch = 30, lr = 3e-3, min_lr = 5e-5) + linear warmup(warmup_lr = 3e-5, warmup_epoch = 4)\n- Soft positive label: 0.9\n\nResults on fold 2 and fold 3 seem to be better. So I decided to finetune fold 0 and fold 1 with the same value of soft_positive_label = 0.9 from the previous last checkpoints.\n\n## 4.4. Checkpoints selection\n1. For each fold, I manually select 3-7 best checkpoints by looking at multiple metrics.\n2. Grid search over all posible combinations of each fold's checkpoints (3x5x7x5 = 525 combinations for my case), compute Out-Of-Fold (OOF) results for each combination.\n3. Select the best combination with highest binarized pF1. The best combination has local OOF pf1 = 0.5187, while the worst with 0.4951.\n\nFinal results:\n\n| name                          | soft_positive_label | pr_auc | roc_auc | best_pf1 | best_thres | epoch        |\n|-------------------------------|---------------------|--------|---------|----------|------------|--------------|\n| fold 0                        | 0.8                 | 0.3983 | 0.9142  | 0.4716   | 0.25       | 24           |\n| **(selected)** fold 0 + fine-tune | 0.9                 | 0.4363 | 0.9119  | 0.4785   | 0.35       | 11 (24 + 11) |\n| fold 1                        | 0.8                 | 0.5151 | 0.9202  | 0.5291   | 0.34       | 18           |\n| **(selected)** fold 1 + fine-tune | 0.9                 | 0.5381 | 0.9149  | 0.5381   | 0.34       | 8 (24 + 8)   |\n| **(selected)** fold 2             | 0.9                 | 0.4946 | 0.9234  | 0.5185   | 0.34       | 26           |\n| **(selected)** fold 3             | 0.9                 | 0.5088 | 0.9401  | 0.5455   | 0.31       | 19           |\n\nSome thoughts:\n- Fold 0 result looks weird? Maybe I need more time inspecting it.\n- With much more data especially positive ones, guesses for the best hyperparams became obsolete. Maybe I could forget the soft positive label trick and get better result?\n\nOOF validation was done to determine the best threshold value of 0.34\n\n```\n                    \tauc      @th     f1      | \tprec    recall  | \tsens    spec \nsingle image     [0]\t0.87296\t0.40000\t0.41907 | \t0.48365\t0.37047 | \t0.37047\t0.99145\ngrouby mean()    [0]\t0.92043\t0.34000\t0.51820 | \t0.60989\t0.45122 | \t0.45122\t0.99391\ngrouby max()     [0]\t0.91939\t0.61000\t0.50913 | \t0.57545\t0.45732 | \t0.45732\t0.99289\n--------------\n\nsingle image     [1]\t0.84866\t0.40000\t0.33649 | \t0.38241\t0.30120 | \t0.30120\t0.98881\ngrouby mean()    [1]\t0.89225\t0.34000\t0.39587 | \t0.47027\t0.34252 | \t0.34252\t0.99139\ngrouby max()     [1]\t0.88917\t0.61000\t0.39424 | \t0.44554\t0.35433 | \t0.35433\t0.99016\n--------------\n\nsingle image     [2]\t0.89611\t0.40000\t0.53331 | \t0.62912\t0.46356 | \t0.46356\t0.99453\ngrouby mean()    [2]\t0.94288\t0.34000\t0.64699 | \t0.75419\t0.56722 | \t0.56723\t0.99632\ngrouby max()     [2]\t0.94329\t0.61000\t0.63182 | \t0.71428\t0.56722 | \t0.56723\t0.99548\n--------------\n```\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10254700%2Feeb0cc23fad1497f249613e0d4f87307%2Ffinal_sub_plot.png?generation=1678010502203196&alt=media)\n\n*The results was generated by @hengck23 ’s [script](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/378521)*\n## 4.5. Submission\nMy training progress was done in the last day of the competition.\n5 final submissions include:\n- 1 submission of models in earlier epochs while waiting training to be done. The main purpose is to ensure the pipeline is correct and no exception occurs, which give me 0.56 LB.\n- 4 submission of same final (best) model with different threshold: 0.31, 0.34 (best oof), 0.37, 0.40 . I guessed the LB's best threshold >  CV's best threshold, but I'm totally wrong (I did not probe the LB). But it was fortunate that threshold=0.31 is the one on the peak :D\n\nThose submissions bring me from LB 600th to LB 22nd in one day. The PL 0.55 submission was successfully finished when ~30 mins left to the deadline.\n\n| Threshold            | OOF        | LB       | PL       |\n|----------------------|------------|----------|----------|\n| 0.27 (late sub)      | 0.4877     | 0.60     | 0.53     |\n| 0.28 (late sub)      | 0.4917     | 0.60     | 0.54     |\n| 0.29 (late sub)      | 0.4973     | 0.60     | 0.54     |\n| 0.30 (late sub)      | 0.5027     | 0.60     | 0.55     |\n| **(selection) 0.31** | **0.5049** | **0.61** | **0.55** |\n| **(selection) 0.34** | **0.5187** | **0.58** | **0.53** |\n| 0.37                 | 0.5000     | 0.55     | 0.52     |\n| 0.40                 | 0.4896     | 0.54     | 0.50     |\n\n\n ---\n# 5. Code\n- Submission notebook: https://www.kaggle.com/dangnh0611/1st-place-submission-code \n- Training code: https://github.com/dangnh0611/kaggle_rsna_breast_cancer\n\n\n---\n\nI just got luck with simple pipeline and simple decision. Many teams had much better models but did not select it in final, as I could see.\nThanks for your attention.",
    "2169786": "Thanks a lot for sharing such a detailed write up. Congratulations on your Gold... looking at the way you broke down the problem, its a well deserved gold. In my experiments also found that Maxpooling worked the best, while average pooling tended to wash away the required signals",
    "2191737": "Could you share more training parameters of ROI detection model? I can't get the same AP@0.5-0.95 as you mentioned above. I only train on @remekkinas's dataset. thank you!",
    "2174211": "Really great work. ",
    "2172703": "Hi, you mentioned that you used image size 2048x1024 with Convnextv1-small model, but when you processed the images yolox produce as output image of 416?.\n\nThank in advance!",
    "2171012": "great work, and thanks to share the solution ",
    "2170281": "Absolutely great work! And a overwhelming amount of experiments you have done starting from pre-processing and all the way to the inference. Also you seem to have picked up many things during this short time that took me couple of years to get a grasp over while doing PhD on this very topic. In my opinion publishing these results and answering those remaining questions that you have listed would get you far for a PhD, unless of course you already have one. 😃",
    "2169955": "Thank you for your kind message! I'm looking forward to your training code! I have learned a lot from reading your sharing.",
    "2169869": "Really great work. I like the augmentations, especially the different downscales. ",
    "2388905": "Interesting!!!",
    "2353153": "Sorry to bother you. I want to reproduce this project, but I can't install nvidia-dali-nightly-cuda. Is this package necessary",
    "2313320": "Great work:), congrats. I just read the solution and review a little bit your code, I would like to know what is the image type input for the windowing process? (I know it must be an array but like, an array from the original dicom image...?)\nThx!",
    "2305243": "Helpful information, thank you for sharing.",
    "2301571": "Hi, thanks for the detailed writeup. I learnt many new things going through your codebase. I have a question. Winning model is an ensemble of four ConvNext models, did using an ensemble model show significantly better performance than a single model? Thanks",
    "2261305": "num_logits = 2 means using sigmoid (BCEWithLogitsLoss) for training and softmax .Can you explain why this strategy is used? Is there any basis for it? Usually the same activation function is used for training and inference.Thanks in advance.",
    "2173792": "Congrats for the win! Thanks for the detailed write up, I have learned a lot.",
    "2172491": "You're a living genius. Bravo",
    "2172374": "Congratz on the win ! \nAnd thanks for the really nice write-up :)",
    "2170596": "Gem was not work for me.\nI think Gem is very tricky.In the same competition , some people say it works, but others say it doesn't work",
    "2170123": "Many thanks for the elaborate write-up and for sharing your reasonings throughout it!\n\nAs you comment in https://github.com/dangnh0611/kaggle_rsna_breast_cancer/blob/dev/notebooks/roi_yolov5.ipynb can you share the annotation dataset and the yolov5 training notebook?\n\nLooking forward to study training code in the dev branch even before you publish the refactored one, if possible.",
    "2169724": "Thank you for your kind share! Do you train single-view models or dual-view models to achieve such an impressive performance?",
    "2893530": "",
    "2283395": "",
    "2261301": "",
    "2175611": ""
  }
}