{
  "id": 391725,
  "title": "3rd Place Solution (Part of data processing and Image-level model)",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/391725",
  "author_name": "ForcewithMe",
  "post_date": "2023-03-02T13:14:42.871000",
  "votes": 31,
  "comment_count": 3,
  "views": 0,
  "content": "<h2>Introduction</h2>\n<p>Thank you to all the participants for your hard work in the competition.We are honored to have achieved a good result, coming in third place in this competition. We also want to express our deepest gratitude to the organizers for putting together such a fantastic event.Thank you very much.<br>\nFinally, I want to thank my excellent teammates <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> , <a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a> , <a href=\"https://www.kaggle.com/kevin1742064161\" target=\"_blank\">@kevin1742064161</a> . On behalf of my teammates, I would like to introduce part of our solution, and another part is presented by <a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a> <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/391779\" target=\"_blank\">in another thread</a>.</p>\n<h2>1. Overview of the pipeline</h2>\n<ul>\n<li>Extract ROI with a fixed aspect ratio(1.6:1) using YOLOX</li>\n<li>Feed the ROI into different classification models</li>\n<li>Average weighting fusion of the results from the classification models</li>\n</ul>\n<h2>2. External Data</h2>\n<p>We use 4 external data in total. Not all models used all the external data. Some models only used CBIS-DDSM + CMMD, while the remaining models used all four external data. Although these external data appear to be different from the competition data, they can improve the CV and significantly enhance the stability of the training.</p>\n<p>1) <a href=\"https://wiki.cancerimagingarchive.net/pages/viewpage.action?pageId=22516629\" target=\"_blank\">CBIS-DDSM</a><br>\nThe classification labels of CBIS-DDSM are: MALIGNANT, BENIGN WITHOUT CALLBACK, BENIGN. We consider MALIGNANT as positive and the others as negative, resulting in 1,350 positive and 1,753 negative samples.</p>\n<p>2) <a href=\"https://wiki.cancerimagingarchive.net/pages/viewpage.action?pageId=70230508\" target=\"_blank\">CMMD</a><br>\nThe classification labels of CMMD are: MALIGNANT, BENIGN. We consider MALIGNANT as positive and BENIGN as negative, resulting in 4,094 positive and 1,108 negative samples.</p>\n<p>3) <a href=\"https://physionet.org/content/vindr-mammo/1.0.0/\" target=\"_blank\">Vindr</a><br>\nVindr did not provide classification labels, but provided BI-RADS indices. The detailed explanation of the indices can be found <a href=\"https://radiopaedia.org/articles/breast-imaging-reporting-and-data-system-bi-rads\" target=\"_blank\">here</a>: <br>\nWe finally chose to consider BI-RADS-4 and BI-RADS-5 as positive,  BI-RADS-2 and BI-RADS-3 as negative, and discard other categories.  Finally, we obtained 988 positive and 5,606 negative samples.</p>\n<p>4) <a href=\"https://www.kaggle.com/datasets/cheddad/miniddsm2\" target=\"_blank\">Mini-DDSM</a><br>\n  The classification labels of Mini-DDSM are: Benign, Cancer, Normal. Since we found that the detection model had a large number of wrong bboxes in the Normal category, we finally chose to discard all Normal images, consider Cancer as positive, and Benign as negative. Finally, we obtained 2,716 positive and 2,684 negative samples.</p>\n<h2>3. Preprocessing</h2>\n<p>1) Firstly, we use 1,000 annotated images, 500 of which are labeled using the open-source annotations provided by <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> and 500 of which are manually annotated by us (referred to as BCD_1k), to train YOLOX-m. No validation set is used, and all data is used for both training and validation. <br>\n2) Using the model from step 1, we predict the whole dataset using YOLOX-m, which will be used as the ROI crop for the official dataset. Meanwhile, the detection boxes will be saved and a pseudo-label dataset (referred to as BCD_all) will be created. <br>\n3) Using the BCD_all dataset obtained in step 2, YOLOX-nano and YOLOX-x are trained, with no validation set. YOLOX-nano will be used for final online submissions, while YOLOX-x will be used for ROI crop on external datasets.<br>\n4) Images were cropped according to the detection boxes and resized to a 1.6:1 aspect ratio (1536<em>960 or 1280</em>800) with padding.</p>\n<p>The link to the BCD_1k and BCD_all datasets:<br>\n<a href=\"https://www.kaggle.com/datasets/kevin1742064161/bcd-dataset\" target=\"_blank\">https://www.kaggle.com/datasets/kevin1742064161/bcd-dataset</a><br>\nThe link to the YOLOX code:<br>\n<a href=\"https://www.kaggle.com/datasets/kevin1742064161/yolo-x\" target=\"_blank\">https://www.kaggle.com/datasets/kevin1742064161/yolo-x</a><br>\nThe link to the bboxes of official datasets and external datasets:<br>\n<a href=\"https://www.kaggle.com/datasets/kevin1742064161/yolox-bboxes\" target=\"_blank\">https://www.kaggle.com/datasets/kevin1742064161/yolox-bboxes</a></p>\n<h2>4. Data augmentation</h2>\n<p>HorizontalFlip, VerticalFlip, RandomBrightnessContrast, ShiftScaleRotate, MedianBlur, GaussianBlur, GaussNoise, ElasticTransform, GridDistortion, OpticalDistortion, CoarseDropout, Mixup</p>\n<h2>5. Model</h2>\n<p>We used two types of classification models. <br>\nOne type is a CNN ( EfficientNet or Convnext) that integrates metadata at the image-level and uses Mean to generate prediction scores. <br>\nThe other is a multi-view model based on CNN+LSTM, which applies the idea of multiple instance learning. (We will call this CNN+LSTM models <code>LSTM</code> for convenience in the following). <a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a> has provided a <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/391779\" target=\"_blank\">more detailed explanation of this model</a>. </p>\n<h3>5.1 Meta data</h3>\n<h4>Motivation of meta data</h4>\n<p>There is a significant score difference between the two sites, which is essentially caused by different machines.</p>\n<h4>How to encode and insert meta data</h4>\n<ul>\n<li>Considering that there are machines in the test set that have not appeared in the training set, <a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a> suggests to use one-hot encoding for machines. Since there are 10 machines in the training set, the encoding is a 1x10 vector. If it is external data or a machine that has not appeared in the training set, it is a vector of all zeros.</li>\n<li>Site and view are encoded in the same way, as a 1x2 vector.</li>\n<li>The machine or machine+site+view vectors are sent to an MLP for encoding, and the encoded features are concatenated with the CNN backbone features and sent to another MLP for classification.</li>\n</ul>\n<h2>6. Ensemble</h2>\n<ul>\n<li>Considering the significant risk of shake in this competition, we adopted a relatively conservative strategy for our final submissions. One submission focused on the Mean model and included 7 image-level models and 4 LSTM models. The other submission focused on LSTM and included 7 LSTM models and 4 image-level models. The threshold for combining the two types of models was optimized using the best cross-validation (CV) threshold obtained from the out-of-fold (OOF) data. </li>\n<li>The image-level models were trained using a 4-fold split, while the LSTM models were trained using a 5-fold split. When searching for the threshold, the entire OOF was used. If a model utilized <code>n</code> folds in the submitted code, it was given a weight of <code>n</code>. Avoiding overemphasis on the weight of individual models was a key strategy for avoiding overfitting.</li>\n<li>In the end, the LSTM-focused model achieved a higher score in both submissions, which was consistent with our CV and simulated private scores. For more details about our simulated private score, please check <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/390958\" target=\"_blank\">here</a>. We are not sure that we simulated in the right way, just a funny story.</li>\n</ul>\n<h3>6.1 First submission: 7 Image-level(Mean) + 4 LSTM</h3>\n<p>Private LB: 0.51 <br>\nPublic LB: 0.67 <br>\nCV: 0.549<br>\nLocal Private score: 0.521</p>\n<p><a href=\"https://postimg.cc/Y4BNyqKk\" target=\"_blank\"><img src=\"https://i.postimg.cc/nLCdrC3s/sub1.png\" alt=\"sub1.png\"></a></p>\n<h2>6.2 Second submission: 7 LSTM + 4 Image-level(Mean)</h2>\n<p>Private LB: 0.53<br>\nPublic LB: 0.67<br>\nCV: 0.554<br>\nLocal private Score: 0.524</p>\n<p><a href=\"https://postimg.cc/9RMDLvMC\" target=\"_blank\"><img src=\"https://i.postimg.cc/63VCpKx2/sub2.png\" alt=\"sub2.png\"></a></p>\n<h2>7. Some interesting thoughts</h2>\n<p>1) LSTM has a much more stable threshold in CV, but performs slightly worse than Mean in LB. For the Mean model, a threshold change of ±0.1 can cause a fluctuation of about 0.05 in pF1. For example, if the optimal threshold is 0.5, the worst pF1 can be around 0.45 when the threshold is 0.4~0.6. However, for the LSTM model, a threshold change of 0.3 may only cause a fluctuation of 0.05.</p>\n<p>2) We had 14 submissions that scored above 0.55 in private leaderboard, including two submissions that ranked highest in the Public Leaderboard (0.68). However, their CV and simulated Private scores were not as good as our final two submissions, so we didn't choose them. </p>\n<p>3) We believe that the LSTM model is the better model (than Image-level models). Although the local pF1 and AUC are similar, using LSTM is obviously a better choice from a theoretical standpoint. Moreover, the LSTM model has a more stable threshold.</p>\n<p><strong>Inference Code is public!</strong><br>\nThe code of the two submissions has been publicly available in <a href=\"https://www.kaggle.com/code/forcewithme/0226-yoloxnano-yoloxs-mean2\" target=\"_blank\">here</a>  and <a href=\"https://www.kaggle.com/code/forcewithme/final-lstm2\" target=\"_blank\">here</a></p>",
  "messages": [
    {
      "id": 2165801,
      "postDate": "2023-03-02T13:14:42.870Z",
      "content": "<h2>Introduction</h2>\n<p>Thank you to all the participants for your hard work in the competition.We are honored to have achieved a good result, coming in third place in this competition. We also want to express our deepest gratitude to the organizers for putting together such a fantastic event.Thank you very much.<br>\nFinally, I want to thank my excellent teammates <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> , <a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a> , <a href=\"https://www.kaggle.com/kevin1742064161\" target=\"_blank\">@kevin1742064161</a> . On behalf of my teammates, I would like to introduce part of our solution, and another part is presented by <a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a> <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/391779\" target=\"_blank\">in another thread</a>.</p>\n<h2>1. Overview of the pipeline</h2>\n<ul>\n<li>Extract ROI with a fixed aspect ratio(1.6:1) using YOLOX</li>\n<li>Feed the ROI into different classification models</li>\n<li>Average weighting fusion of the results from the classification models</li>\n</ul>\n<h2>2. External Data</h2>\n<p>We use 4 external data in total. Not all models used all the external data. Some models only used CBIS-DDSM + CMMD, while the remaining models used all four external data. Although these external data appear to be different from the competition data, they can improve the CV and significantly enhance the stability of the training.</p>\n<p>1) <a href=\"https://wiki.cancerimagingarchive.net/pages/viewpage.action?pageId=22516629\" target=\"_blank\">CBIS-DDSM</a><br>\nThe classification labels of CBIS-DDSM are: MALIGNANT, BENIGN WITHOUT CALLBACK, BENIGN. We consider MALIGNANT as positive and the others as negative, resulting in 1,350 positive and 1,753 negative samples.</p>\n<p>2) <a href=\"https://wiki.cancerimagingarchive.net/pages/viewpage.action?pageId=70230508\" target=\"_blank\">CMMD</a><br>\nThe classification labels of CMMD are: MALIGNANT, BENIGN. We consider MALIGNANT as positive and BENIGN as negative, resulting in 4,094 positive and 1,108 negative samples.</p>\n<p>3) <a href=\"https://physionet.org/content/vindr-mammo/1.0.0/\" target=\"_blank\">Vindr</a><br>\nVindr did not provide classification labels, but provided BI-RADS indices. The detailed explanation of the indices can be found <a href=\"https://radiopaedia.org/articles/breast-imaging-reporting-and-data-system-bi-rads\" target=\"_blank\">here</a>: <br>\nWe finally chose to consider BI-RADS-4 and BI-RADS-5 as positive,  BI-RADS-2 and BI-RADS-3 as negative, and discard other categories.  Finally, we obtained 988 positive and 5,606 negative samples.</p>\n<p>4) <a href=\"https://www.kaggle.com/datasets/cheddad/miniddsm2\" target=\"_blank\">Mini-DDSM</a><br>\n  The classification labels of Mini-DDSM are: Benign, Cancer, Normal. Since we found that the detection model had a large number of wrong bboxes in the Normal category, we finally chose to discard all Normal images, consider Cancer as positive, and Benign as negative. Finally, we obtained 2,716 positive and 2,684 negative samples.</p>\n<h2>3. Preprocessing</h2>\n<p>1) Firstly, we use 1,000 annotated images, 500 of which are labeled using the open-source annotations provided by <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> and 500 of which are manually annotated by us (referred to as BCD_1k), to train YOLOX-m. No validation set is used, and all data is used for both training and validation. <br>\n2) Using the model from step 1, we predict the whole dataset using YOLOX-m, which will be used as the ROI crop for the official dataset. Meanwhile, the detection boxes will be saved and a pseudo-label dataset (referred to as BCD_all) will be created. <br>\n3) Using the BCD_all dataset obtained in step 2, YOLOX-nano and YOLOX-x are trained, with no validation set. YOLOX-nano will be used for final online submissions, while YOLOX-x will be used for ROI crop on external datasets.<br>\n4) Images were cropped according to the detection boxes and resized to a 1.6:1 aspect ratio (1536<em>960 or 1280</em>800) with padding.</p>\n<p>The link to the BCD_1k and BCD_all datasets:<br>\n<a href=\"https://www.kaggle.com/datasets/kevin1742064161/bcd-dataset\" target=\"_blank\">https://www.kaggle.com/datasets/kevin1742064161/bcd-dataset</a><br>\nThe link to the YOLOX code:<br>\n<a href=\"https://www.kaggle.com/datasets/kevin1742064161/yolo-x\" target=\"_blank\">https://www.kaggle.com/datasets/kevin1742064161/yolo-x</a><br>\nThe link to the bboxes of official datasets and external datasets:<br>\n<a href=\"https://www.kaggle.com/datasets/kevin1742064161/yolox-bboxes\" target=\"_blank\">https://www.kaggle.com/datasets/kevin1742064161/yolox-bboxes</a></p>\n<h2>4. Data augmentation</h2>\n<p>HorizontalFlip, VerticalFlip, RandomBrightnessContrast, ShiftScaleRotate, MedianBlur, GaussianBlur, GaussNoise, ElasticTransform, GridDistortion, OpticalDistortion, CoarseDropout, Mixup</p>\n<h2>5. Model</h2>\n<p>We used two types of classification models. <br>\nOne type is a CNN ( EfficientNet or Convnext) that integrates metadata at the image-level and uses Mean to generate prediction scores. <br>\nThe other is a multi-view model based on CNN+LSTM, which applies the idea of multiple instance learning. (We will call this CNN+LSTM models <code>LSTM</code> for convenience in the following). <a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a> has provided a <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/391779\" target=\"_blank\">more detailed explanation of this model</a>. </p>\n<h3>5.1 Meta data</h3>\n<h4>Motivation of meta data</h4>\n<p>There is a significant score difference between the two sites, which is essentially caused by different machines.</p>\n<h4>How to encode and insert meta data</h4>\n<ul>\n<li>Considering that there are machines in the test set that have not appeared in the training set, <a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a> suggests to use one-hot encoding for machines. Since there are 10 machines in the training set, the encoding is a 1x10 vector. If it is external data or a machine that has not appeared in the training set, it is a vector of all zeros.</li>\n<li>Site and view are encoded in the same way, as a 1x2 vector.</li>\n<li>The machine or machine+site+view vectors are sent to an MLP for encoding, and the encoded features are concatenated with the CNN backbone features and sent to another MLP for classification.</li>\n</ul>\n<h2>6. Ensemble</h2>\n<ul>\n<li>Considering the significant risk of shake in this competition, we adopted a relatively conservative strategy for our final submissions. One submission focused on the Mean model and included 7 image-level models and 4 LSTM models. The other submission focused on LSTM and included 7 LSTM models and 4 image-level models. The threshold for combining the two types of models was optimized using the best cross-validation (CV) threshold obtained from the out-of-fold (OOF) data. </li>\n<li>The image-level models were trained using a 4-fold split, while the LSTM models were trained using a 5-fold split. When searching for the threshold, the entire OOF was used. If a model utilized <code>n</code> folds in the submitted code, it was given a weight of <code>n</code>. Avoiding overemphasis on the weight of individual models was a key strategy for avoiding overfitting.</li>\n<li>In the end, the LSTM-focused model achieved a higher score in both submissions, which was consistent with our CV and simulated private scores. For more details about our simulated private score, please check <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/390958\" target=\"_blank\">here</a>. We are not sure that we simulated in the right way, just a funny story.</li>\n</ul>\n<h3>6.1 First submission: 7 Image-level(Mean) + 4 LSTM</h3>\n<p>Private LB: 0.51 <br>\nPublic LB: 0.67 <br>\nCV: 0.549<br>\nLocal Private score: 0.521</p>\n<p><a href=\"https://postimg.cc/Y4BNyqKk\" target=\"_blank\"><img src=\"https://i.postimg.cc/nLCdrC3s/sub1.png\" alt=\"sub1.png\"></a></p>\n<h2>6.2 Second submission: 7 LSTM + 4 Image-level(Mean)</h2>\n<p>Private LB: 0.53<br>\nPublic LB: 0.67<br>\nCV: 0.554<br>\nLocal private Score: 0.524</p>\n<p><a href=\"https://postimg.cc/9RMDLvMC\" target=\"_blank\"><img src=\"https://i.postimg.cc/63VCpKx2/sub2.png\" alt=\"sub2.png\"></a></p>\n<h2>7. Some interesting thoughts</h2>\n<p>1) LSTM has a much more stable threshold in CV, but performs slightly worse than Mean in LB. For the Mean model, a threshold change of ±0.1 can cause a fluctuation of about 0.05 in pF1. For example, if the optimal threshold is 0.5, the worst pF1 can be around 0.45 when the threshold is 0.4~0.6. However, for the LSTM model, a threshold change of 0.3 may only cause a fluctuation of 0.05.</p>\n<p>2) We had 14 submissions that scored above 0.55 in private leaderboard, including two submissions that ranked highest in the Public Leaderboard (0.68). However, their CV and simulated Private scores were not as good as our final two submissions, so we didn't choose them. </p>\n<p>3) We believe that the LSTM model is the better model (than Image-level models). Although the local pF1 and AUC are similar, using LSTM is obviously a better choice from a theoretical standpoint. Moreover, the LSTM model has a more stable threshold.</p>\n<p><strong>Inference Code is public!</strong><br>\nThe code of the two submissions has been publicly available in <a href=\"https://www.kaggle.com/code/forcewithme/0226-yoloxnano-yoloxs-mean2\" target=\"_blank\">here</a>  and <a href=\"https://www.kaggle.com/code/forcewithme/final-lstm2\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "## Introduction\nThank you to all the participants for your hard work in the competition.We are honored to have achieved a good result, coming in third place in this competition. We also want to express our deepest gratitude to the organizers for putting together such a fantastic event.Thank you very much.\nFinally, I want to thank my excellent teammates @haqishen , @boliu0 , @kevin1742064161 . On behalf of my teammates, I would like to introduce part of our solution, and another part is presented by @boliu0 [in another thread](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/391779).\n\n## 1. Overview of the pipeline\n- Extract ROI with a fixed aspect ratio(1.6:1) using YOLOX\n- Feed the ROI into different classification models\n- Average weighting fusion of the results from the classification models\n\n## 2. External Data\nWe use 4 external data in total. Not all models used all the external data. Some models only used CBIS-DDSM + CMMD, while the remaining models used all four external data. Although these external data appear to be different from the competition data, they can improve the CV and significantly enhance the stability of the training.\n\n1) [CBIS-DDSM](https://wiki.cancerimagingarchive.net/pages/viewpage.action?pageId=22516629)\nThe classification labels of CBIS-DDSM are: MALIGNANT, BENIGN WITHOUT CALLBACK, BENIGN. We consider MALIGNANT as positive and the others as negative, resulting in 1,350 positive and 1,753 negative samples.\n\n2) [CMMD](https://wiki.cancerimagingarchive.net/pages/viewpage.action?pageId=70230508)\nThe classification labels of CMMD are: MALIGNANT, BENIGN. We consider MALIGNANT as positive and BENIGN as negative, resulting in 4,094 positive and 1,108 negative samples.\n\n3) [Vindr](https://physionet.org/content/vindr-mammo/1.0.0/)\nVindr did not provide classification labels, but provided BI-RADS indices. The detailed explanation of the indices can be found [here](https://radiopaedia.org/articles/breast-imaging-reporting-and-data-system-bi-rads): \nWe finally chose to consider BI-RADS-4 and BI-RADS-5 as positive,  BI-RADS-2 and BI-RADS-3 as negative, and discard other categories.  Finally, we obtained 988 positive and 5,606 negative samples.\n\n4) [Mini-DDSM](https://www.kaggle.com/datasets/cheddad/miniddsm2)\n  The classification labels of Mini-DDSM are: Benign, Cancer, Normal. Since we found that the detection model had a large number of wrong bboxes in the Normal category, we finally chose to discard all Normal images, consider Cancer as positive, and Benign as negative. Finally, we obtained 2,716 positive and 2,684 negative samples.\n\n## 3. Preprocessing\n1) Firstly, we use 1,000 annotated images, 500 of which are labeled using the open-source annotations provided by @remekkinas and 500 of which are manually annotated by us (referred to as BCD_1k), to train YOLOX-m. No validation set is used, and all data is used for both training and validation. \n2) Using the model from step 1, we predict the whole dataset using YOLOX-m, which will be used as the ROI crop for the official dataset. Meanwhile, the detection boxes will be saved and a pseudo-label dataset (referred to as BCD_all) will be created. \n3) Using the BCD_all dataset obtained in step 2, YOLOX-nano and YOLOX-x are trained, with no validation set. YOLOX-nano will be used for final online submissions, while YOLOX-x will be used for ROI crop on external datasets.\n4) Images were cropped according to the detection boxes and resized to a 1.6:1 aspect ratio (1536*960 or 1280*800) with padding.\n\nThe link to the BCD_1k and BCD_all datasets:\n[https://www.kaggle.com/datasets/kevin1742064161/bcd-dataset](https://www.kaggle.com/datasets/kevin1742064161/bcd-dataset)\nThe link to the YOLOX code:\n[https://www.kaggle.com/datasets/kevin1742064161/yolo-x](https://www.kaggle.com/datasets/kevin1742064161/yolo-x)\nThe link to the bboxes of official datasets and external datasets:\n[https://www.kaggle.com/datasets/kevin1742064161/yolox-bboxes](https://www.kaggle.com/datasets/kevin1742064161/yolox-bboxes)\n\n## 4. Data augmentation\nHorizontalFlip, VerticalFlip, RandomBrightnessContrast, ShiftScaleRotate, MedianBlur, GaussianBlur, GaussNoise, ElasticTransform, GridDistortion, OpticalDistortion, CoarseDropout, Mixup\n\n## 5. Model\nWe used two types of classification models. \nOne type is a CNN ( EfficientNet or Convnext) that integrates metadata at the image-level and uses Mean to generate prediction scores. \nThe other is a multi-view model based on CNN+LSTM, which applies the idea of multiple instance learning. (We will call this CNN+LSTM models `LSTM` for convenience in the following). @boliu0 has provided a [more detailed explanation of this model](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/391779). \n\n### 5.1 Meta data\n#### Motivation of meta data\nThere is a significant score difference between the two sites, which is essentially caused by different machines.\n\n#### How to encode and insert meta data\n- Considering that there are machines in the test set that have not appeared in the training set, @boliu0 suggests to use one-hot encoding for machines. Since there are 10 machines in the training set, the encoding is a 1x10 vector. If it is external data or a machine that has not appeared in the training set, it is a vector of all zeros.\n- Site and view are encoded in the same way, as a 1x2 vector.\n- The machine or machine+site+view vectors are sent to an MLP for encoding, and the encoded features are concatenated with the CNN backbone features and sent to another MLP for classification.\n\n## 6. Ensemble\n- Considering the significant risk of shake in this competition, we adopted a relatively conservative strategy for our final submissions. One submission focused on the Mean model and included 7 image-level models and 4 LSTM models. The other submission focused on LSTM and included 7 LSTM models and 4 image-level models. The threshold for combining the two types of models was optimized using the best cross-validation (CV) threshold obtained from the out-of-fold (OOF) data. \n- The image-level models were trained using a 4-fold split, while the LSTM models were trained using a 5-fold split. When searching for the threshold, the entire OOF was used. If a model utilized `n` folds in the submitted code, it was given a weight of `n`. Avoiding overemphasis on the weight of individual models was a key strategy for avoiding overfitting.\n- In the end, the LSTM-focused model achieved a higher score in both submissions, which was consistent with our CV and simulated private scores. For more details about our simulated private score, please check [here](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/390958). We are not sure that we simulated in the right way, just a funny story.\n\n### 6.1 First submission: 7 Image-level(Mean) + 4 LSTM\nPrivate LB: 0.51 \nPublic LB: 0.67 \nCV: 0.549\nLocal Private score: 0.521\n\n[![sub1.png](https://i.postimg.cc/nLCdrC3s/sub1.png)](https://postimg.cc/Y4BNyqKk)\n\n## 6.2 Second submission: 7 LSTM + 4 Image-level(Mean)\nPrivate LB: 0.53\nPublic LB: 0.67\nCV: 0.554\nLocal private Score: 0.524\n\n[![sub2.png](https://i.postimg.cc/63VCpKx2/sub2.png)](https://postimg.cc/9RMDLvMC)\n\n## 7. Some interesting thoughts\n\n1) LSTM has a much more stable threshold in CV, but performs slightly worse than Mean in LB. For the Mean model, a threshold change of ±0.1 can cause a fluctuation of about 0.05 in pF1. For example, if the optimal threshold is 0.5, the worst pF1 can be around 0.45 when the threshold is 0.4~0.6. However, for the LSTM model, a threshold change of 0.3 may only cause a fluctuation of 0.05.\n\n2) We had 14 submissions that scored above 0.55 in private leaderboard, including two submissions that ranked highest in the Public Leaderboard (0.68). However, their CV and simulated Private scores were not as good as our final two submissions, so we didn't choose them. \n\n3) We believe that the LSTM model is the better model (than Image-level models). Although the local pF1 and AUC are similar, using LSTM is obviously a better choice from a theoretical standpoint. Moreover, the LSTM model has a more stable threshold.\n\n**Inference Code is public!**\nThe code of the two submissions has been publicly available in [here](https://www.kaggle.com/code/forcewithme/0226-yoloxnano-yoloxs-mean2)  and [here](https://www.kaggle.com/code/forcewithme/final-lstm2)\n",
      "votes": 31
    },
    {
      "id": 2165938,
      "postDate": "2023-03-02T14:47:27.910Z",
      "content": "<blockquote>\n  <p>We believe that the LSTM model is the better model (than Image-level models). Although the local pF1 and AUC are similar, using LSTM is obviously a better choice from a theoretical standpoint. Moreover, the LSTM model has a more stable threshold.</p>\n</blockquote>\n<p>Hope you get to investigate this further during the follow-up analysis. Congrats and thanks for sharing, your deliverables/datasets are very comprehensive.</p>",
      "rawMarkdown": ">We believe that the LSTM model is the better model (than Image-level models). Although the local pF1 and AUC are similar, using LSTM is obviously a better choice from a theoretical standpoint. Moreover, the LSTM model has a more stable threshold.\n\nHope you get to investigate this further during the follow-up analysis. Congrats and thanks for sharing, your deliverables/datasets are very comprehensive.",
      "votes": 1,
      "replies": [
        {
          "id": 2165987,
          "postDate": "2023-03-02T15:10:56.087Z",
          "content": "<p>Thank you for your suggestions and congratulations.</p>",
          "rawMarkdown": "Thank you for your suggestions and congratulations.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2173131,
      "postDate": "2023-03-08T05:49:03.253Z",
      "content": "<p>I don't know yolox, you can help me?</p>",
      "rawMarkdown": "I don't know yolox, you can help me?"
    }
  ],
  "comments": [
    {
      "id": 2165938,
      "author_name": "Antti Isosalo",
      "author_url": "",
      "post_date": "2023-03-02T14:47:27.910000",
      "content": "<blockquote>\n  <p>We believe that the LSTM model is the better model (than Image-level models). Although the local pF1 and AUC are similar, using LSTM is obviously a better choice from a theoretical standpoint. Moreover, the LSTM model has a more stable threshold.</p>\n</blockquote>\n<p>Hope you get to investigate this further during the follow-up analysis. Congrats and thanks for sharing, your deliverables/datasets are very comprehensive.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2165987,
          "author_name": "ForcewithMe",
          "author_url": "",
          "post_date": "2023-03-02T15:10:56.087000",
          "content": "<p>Thank you for your suggestions and congratulations.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2173131,
      "author_name": "Qx Nam",
      "author_url": "",
      "post_date": "2023-03-08T05:49:03.253000",
      "content": "<p>I don't know yolox, you can help me?</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2165801": "## Introduction\nThank you to all the participants for your hard work in the competition.We are honored to have achieved a good result, coming in third place in this competition. We also want to express our deepest gratitude to the organizers for putting together such a fantastic event.Thank you very much.\nFinally, I want to thank my excellent teammates @haqishen , @boliu0 , @kevin1742064161 . On behalf of my teammates, I would like to introduce part of our solution, and another part is presented by @boliu0 [in another thread](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/391779).\n\n## 1. Overview of the pipeline\n- Extract ROI with a fixed aspect ratio(1.6:1) using YOLOX\n- Feed the ROI into different classification models\n- Average weighting fusion of the results from the classification models\n\n## 2. External Data\nWe use 4 external data in total. Not all models used all the external data. Some models only used CBIS-DDSM + CMMD, while the remaining models used all four external data. Although these external data appear to be different from the competition data, they can improve the CV and significantly enhance the stability of the training.\n\n1) [CBIS-DDSM](https://wiki.cancerimagingarchive.net/pages/viewpage.action?pageId=22516629)\nThe classification labels of CBIS-DDSM are: MALIGNANT, BENIGN WITHOUT CALLBACK, BENIGN. We consider MALIGNANT as positive and the others as negative, resulting in 1,350 positive and 1,753 negative samples.\n\n2) [CMMD](https://wiki.cancerimagingarchive.net/pages/viewpage.action?pageId=70230508)\nThe classification labels of CMMD are: MALIGNANT, BENIGN. We consider MALIGNANT as positive and BENIGN as negative, resulting in 4,094 positive and 1,108 negative samples.\n\n3) [Vindr](https://physionet.org/content/vindr-mammo/1.0.0/)\nVindr did not provide classification labels, but provided BI-RADS indices. The detailed explanation of the indices can be found [here](https://radiopaedia.org/articles/breast-imaging-reporting-and-data-system-bi-rads): \nWe finally chose to consider BI-RADS-4 and BI-RADS-5 as positive,  BI-RADS-2 and BI-RADS-3 as negative, and discard other categories.  Finally, we obtained 988 positive and 5,606 negative samples.\n\n4) [Mini-DDSM](https://www.kaggle.com/datasets/cheddad/miniddsm2)\n  The classification labels of Mini-DDSM are: Benign, Cancer, Normal. Since we found that the detection model had a large number of wrong bboxes in the Normal category, we finally chose to discard all Normal images, consider Cancer as positive, and Benign as negative. Finally, we obtained 2,716 positive and 2,684 negative samples.\n\n## 3. Preprocessing\n1) Firstly, we use 1,000 annotated images, 500 of which are labeled using the open-source annotations provided by @remekkinas and 500 of which are manually annotated by us (referred to as BCD_1k), to train YOLOX-m. No validation set is used, and all data is used for both training and validation. \n2) Using the model from step 1, we predict the whole dataset using YOLOX-m, which will be used as the ROI crop for the official dataset. Meanwhile, the detection boxes will be saved and a pseudo-label dataset (referred to as BCD_all) will be created. \n3) Using the BCD_all dataset obtained in step 2, YOLOX-nano and YOLOX-x are trained, with no validation set. YOLOX-nano will be used for final online submissions, while YOLOX-x will be used for ROI crop on external datasets.\n4) Images were cropped according to the detection boxes and resized to a 1.6:1 aspect ratio (1536*960 or 1280*800) with padding.\n\nThe link to the BCD_1k and BCD_all datasets:\n[https://www.kaggle.com/datasets/kevin1742064161/bcd-dataset](https://www.kaggle.com/datasets/kevin1742064161/bcd-dataset)\nThe link to the YOLOX code:\n[https://www.kaggle.com/datasets/kevin1742064161/yolo-x](https://www.kaggle.com/datasets/kevin1742064161/yolo-x)\nThe link to the bboxes of official datasets and external datasets:\n[https://www.kaggle.com/datasets/kevin1742064161/yolox-bboxes](https://www.kaggle.com/datasets/kevin1742064161/yolox-bboxes)\n\n## 4. Data augmentation\nHorizontalFlip, VerticalFlip, RandomBrightnessContrast, ShiftScaleRotate, MedianBlur, GaussianBlur, GaussNoise, ElasticTransform, GridDistortion, OpticalDistortion, CoarseDropout, Mixup\n\n## 5. Model\nWe used two types of classification models. \nOne type is a CNN ( EfficientNet or Convnext) that integrates metadata at the image-level and uses Mean to generate prediction scores. \nThe other is a multi-view model based on CNN+LSTM, which applies the idea of multiple instance learning. (We will call this CNN+LSTM models `LSTM` for convenience in the following). @boliu0 has provided a [more detailed explanation of this model](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/391779). \n\n### 5.1 Meta data\n#### Motivation of meta data\nThere is a significant score difference between the two sites, which is essentially caused by different machines.\n\n#### How to encode and insert meta data\n- Considering that there are machines in the test set that have not appeared in the training set, @boliu0 suggests to use one-hot encoding for machines. Since there are 10 machines in the training set, the encoding is a 1x10 vector. If it is external data or a machine that has not appeared in the training set, it is a vector of all zeros.\n- Site and view are encoded in the same way, as a 1x2 vector.\n- The machine or machine+site+view vectors are sent to an MLP for encoding, and the encoded features are concatenated with the CNN backbone features and sent to another MLP for classification.\n\n## 6. Ensemble\n- Considering the significant risk of shake in this competition, we adopted a relatively conservative strategy for our final submissions. One submission focused on the Mean model and included 7 image-level models and 4 LSTM models. The other submission focused on LSTM and included 7 LSTM models and 4 image-level models. The threshold for combining the two types of models was optimized using the best cross-validation (CV) threshold obtained from the out-of-fold (OOF) data. \n- The image-level models were trained using a 4-fold split, while the LSTM models were trained using a 5-fold split. When searching for the threshold, the entire OOF was used. If a model utilized `n` folds in the submitted code, it was given a weight of `n`. Avoiding overemphasis on the weight of individual models was a key strategy for avoiding overfitting.\n- In the end, the LSTM-focused model achieved a higher score in both submissions, which was consistent with our CV and simulated private scores. For more details about our simulated private score, please check [here](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/390958). We are not sure that we simulated in the right way, just a funny story.\n\n### 6.1 First submission: 7 Image-level(Mean) + 4 LSTM\nPrivate LB: 0.51 \nPublic LB: 0.67 \nCV: 0.549\nLocal Private score: 0.521\n\n[![sub1.png](https://i.postimg.cc/nLCdrC3s/sub1.png)](https://postimg.cc/Y4BNyqKk)\n\n## 6.2 Second submission: 7 LSTM + 4 Image-level(Mean)\nPrivate LB: 0.53\nPublic LB: 0.67\nCV: 0.554\nLocal private Score: 0.524\n\n[![sub2.png](https://i.postimg.cc/63VCpKx2/sub2.png)](https://postimg.cc/9RMDLvMC)\n\n## 7. Some interesting thoughts\n\n1) LSTM has a much more stable threshold in CV, but performs slightly worse than Mean in LB. For the Mean model, a threshold change of ±0.1 can cause a fluctuation of about 0.05 in pF1. For example, if the optimal threshold is 0.5, the worst pF1 can be around 0.45 when the threshold is 0.4~0.6. However, for the LSTM model, a threshold change of 0.3 may only cause a fluctuation of 0.05.\n\n2) We had 14 submissions that scored above 0.55 in private leaderboard, including two submissions that ranked highest in the Public Leaderboard (0.68). However, their CV and simulated Private scores were not as good as our final two submissions, so we didn't choose them. \n\n3) We believe that the LSTM model is the better model (than Image-level models). Although the local pF1 and AUC are similar, using LSTM is obviously a better choice from a theoretical standpoint. Moreover, the LSTM model has a more stable threshold.\n\n**Inference Code is public!**\nThe code of the two submissions has been publicly available in [here](https://www.kaggle.com/code/forcewithme/0226-yoloxnano-yoloxs-mean2)  and [here](https://www.kaggle.com/code/forcewithme/final-lstm2)\n",
    "2165938": ">We believe that the LSTM model is the better model (than Image-level models). Although the local pF1 and AUC are similar, using LSTM is obviously a better choice from a theoretical standpoint. Moreover, the LSTM model has a more stable threshold.\n\nHope you get to investigate this further during the follow-up analysis. Congrats and thanks for sharing, your deliverables/datasets are very comprehensive.",
    "2173131": "I don't know yolox, you can help me?"
  }
}