{
  "id": 391125,
  "title": "7th Solution",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/391125",
  "author_name": "Mt.Panda",
  "post_date": "2023-02-28T13:07:27.491000",
  "votes": 32,
  "comment_count": 26,
  "views": 0,
  "content": "<p>First of all, we would like to express our respect to all the participants and thank the organizers for making this competition possible. This competition was tough for us due to the volatility of the metrics and the fact that we had to complete our inference in less than 9 hours. In the end, by choosing our best local cv model (CV0.534), we were able to win 7th place.<br>\nHere we will describe how we increased our CV score and how we made our entire inference faster.</p>\n<h2>Summary</h2>\n<ul>\n<li>Images: Kaggle train data and VinDr-Mammo as the external data</li>\n<li>Preprocess: ROI cropping in a rule-based way and sigmoid windowing</li>\n<li>Resolution: 1520x912</li>\n<li>Model: EfficientNetV2S (and EfficientNet B5) with GeM pooling (p=3)</li>\n<li>CV: 4-fold, grouped by patient and stratified by cancer, BIRADS, density, age, biopsy, implant, and machine_id</li>\n<li>Train:<ul>\n<li>Augmentation: V/H Flip, Geometric transformation (Affine and Elastic)</li>\n<li>Loss function: BCE</li>\n<li>Optimizer: Adam</li>\n<li>Scheduler: Cosine decay (starting from 5e-5)</li></ul></li>\n<li>Inference: 2xTTA (vertical flip)</li>\n<li>Ensemble: seed averaging (2 seeds) and 2 level ensemble (breast- and laterality-level)</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3706526%2Ff93bb633bb5bcfeffde33cb6931f5bd3%2Fsolution.png?generation=1677589173238328&amp;alt=media\" alt=\"\"></p>\n<h2>Preprocessing</h2>\n<p>We obtained the ROI breast in a rule-based way. In brief, after converting values below 40 to 0, uniform columns and rows were removed because they were assumed to be background. This process was so simple yet effective and fast. After this preprocessing, we obtained images with an aspect ratio of 1:1.6~2 on average, and which were then resized to 1520x912. This resolution was determined after comparing four options: ①640x640, ②1024x1024, ③1520x912, and ④2689x1569. The order of CV score was ① &lt; ② &lt; ③ &gt; ④. Additionally, we generated images with sigmoid windowing applied as well, which did not have a significant effect on improving the score, and even on ensemble. However, we trained some models with these images and included them in the ensemble to make our prediction more robust.<br>\nThe most important thing was to generate 2 types of input images, namely breast-level and laterality-level. The breast-level consisted of one breast per image, while the laterality-level had two or more breasts per image by simply concatenating breasts in columns. The models trained with the laterality-level was much weaker (-0.04) than those with the breast-level, but the ensemble of the two levels was very effective.</p>\n<h2>Model Architectures</h2>\n<p>We found that larger models did not necessarily score better (EfficientNetB2 &lt; EfficientNetB5 &gt; EfficientNetB7). Out of several models, including EfficientNetV2, EfficientNetB5, SeResNeXt50, ConvNeXt tiny/small, and NextViT base, EfficientNetB5 performed the best (CV0.47 on 1520x912), and EfficientNetV2S came in second-best (CV0.45 on 1520x912). Additionally, we replaced the pooling layer from 'average' to 'generalized mean' (GeM), which resulted in a slight improvement in score (0.005~0.01). Note that we ultimately used EfficientNetV2S in inference because EfficientNetB5 takes longer to infer than EfficientNetV2S. However, the difference in metrics between the two was negligible (CV0.49, single model) by using B5's predictions on external data when training V2S, as described later.</p>\n<h2>External Data</h2>\n<p>We used VinDr-Mammo dataset as external data, whose labels were defined by breast-level predictions after aggregating them into laterality-level. We found a significant improvement in the CV score by 0.02 using this dataset.</p>\n<h2>Inference Speed Up</h2>\n<p>The inference time limit was very tight, so we made efforts to ensure inference completed in time. First, we used DALI to decode images and preprocessed most of them on GPU. This significantly increased processing speed. However, as mentioned <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/384231\" target=\"_blank\">here</a>, some images could not be processed using DALI, so we used dicomSDL and cupy as a fallback method for decoding and preprocessing images. Additionally, we implemented a 2-stage method as shown in the following image. Through data analysis, we observed that the data contained many explicit negative samples which did not require prediction by a strong model or ensemble. By filtering out those data with the threshold of 0.01, we were able to reduce the number of data (to 25%) which need to be saved to disk and predicted without losing accuracy.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3706526%2Fcfd0afe964fcafd12ca8460914bd6816%2Fsolution1.png?generation=1677589474115470&amp;alt=media\" alt=\"\"><br>\nTo speed up prediction, we compiled our models using tensorRT. Through our experiments, we observed that the fp16 model was 1.3 times faster than fp32, but its numerical error is not negligible. In the end, we used models compiled in fp32.</p>\n<h2>Code</h2>\n<p>train: <a href=\"https://github.com/Masaaaato/RSNABreast7thPlace\" target=\"_blank\">https://github.com/Masaaaato/RSNABreast7thPlace</a><br>\ninference: <a href=\"https://www.kaggle.com/code/masato114/2stage-ensemble/notebook\" target=\"_blank\">https://www.kaggle.com/code/masato114/2stage-ensemble/notebook</a></p>",
  "messages": [
    {
      "id": 2162872,
      "postDate": "2023-02-28T13:07:27.490Z",
      "content": "<p>First of all, we would like to express our respect to all the participants and thank the organizers for making this competition possible. This competition was tough for us due to the volatility of the metrics and the fact that we had to complete our inference in less than 9 hours. In the end, by choosing our best local cv model (CV0.534), we were able to win 7th place.<br>\nHere we will describe how we increased our CV score and how we made our entire inference faster.</p>\n<h2>Summary</h2>\n<ul>\n<li>Images: Kaggle train data and VinDr-Mammo as the external data</li>\n<li>Preprocess: ROI cropping in a rule-based way and sigmoid windowing</li>\n<li>Resolution: 1520x912</li>\n<li>Model: EfficientNetV2S (and EfficientNet B5) with GeM pooling (p=3)</li>\n<li>CV: 4-fold, grouped by patient and stratified by cancer, BIRADS, density, age, biopsy, implant, and machine_id</li>\n<li>Train:<ul>\n<li>Augmentation: V/H Flip, Geometric transformation (Affine and Elastic)</li>\n<li>Loss function: BCE</li>\n<li>Optimizer: Adam</li>\n<li>Scheduler: Cosine decay (starting from 5e-5)</li></ul></li>\n<li>Inference: 2xTTA (vertical flip)</li>\n<li>Ensemble: seed averaging (2 seeds) and 2 level ensemble (breast- and laterality-level)</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3706526%2Ff93bb633bb5bcfeffde33cb6931f5bd3%2Fsolution.png?generation=1677589173238328&amp;alt=media\" alt=\"\"></p>\n<h2>Preprocessing</h2>\n<p>We obtained the ROI breast in a rule-based way. In brief, after converting values below 40 to 0, uniform columns and rows were removed because they were assumed to be background. This process was so simple yet effective and fast. After this preprocessing, we obtained images with an aspect ratio of 1:1.6~2 on average, and which were then resized to 1520x912. This resolution was determined after comparing four options: ①640x640, ②1024x1024, ③1520x912, and ④2689x1569. The order of CV score was ① &lt; ② &lt; ③ &gt; ④. Additionally, we generated images with sigmoid windowing applied as well, which did not have a significant effect on improving the score, and even on ensemble. However, we trained some models with these images and included them in the ensemble to make our prediction more robust.<br>\nThe most important thing was to generate 2 types of input images, namely breast-level and laterality-level. The breast-level consisted of one breast per image, while the laterality-level had two or more breasts per image by simply concatenating breasts in columns. The models trained with the laterality-level was much weaker (-0.04) than those with the breast-level, but the ensemble of the two levels was very effective.</p>\n<h2>Model Architectures</h2>\n<p>We found that larger models did not necessarily score better (EfficientNetB2 &lt; EfficientNetB5 &gt; EfficientNetB7). Out of several models, including EfficientNetV2, EfficientNetB5, SeResNeXt50, ConvNeXt tiny/small, and NextViT base, EfficientNetB5 performed the best (CV0.47 on 1520x912), and EfficientNetV2S came in second-best (CV0.45 on 1520x912). Additionally, we replaced the pooling layer from 'average' to 'generalized mean' (GeM), which resulted in a slight improvement in score (0.005~0.01). Note that we ultimately used EfficientNetV2S in inference because EfficientNetB5 takes longer to infer than EfficientNetV2S. However, the difference in metrics between the two was negligible (CV0.49, single model) by using B5's predictions on external data when training V2S, as described later.</p>\n<h2>External Data</h2>\n<p>We used VinDr-Mammo dataset as external data, whose labels were defined by breast-level predictions after aggregating them into laterality-level. We found a significant improvement in the CV score by 0.02 using this dataset.</p>\n<h2>Inference Speed Up</h2>\n<p>The inference time limit was very tight, so we made efforts to ensure inference completed in time. First, we used DALI to decode images and preprocessed most of them on GPU. This significantly increased processing speed. However, as mentioned <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/384231\" target=\"_blank\">here</a>, some images could not be processed using DALI, so we used dicomSDL and cupy as a fallback method for decoding and preprocessing images. Additionally, we implemented a 2-stage method as shown in the following image. Through data analysis, we observed that the data contained many explicit negative samples which did not require prediction by a strong model or ensemble. By filtering out those data with the threshold of 0.01, we were able to reduce the number of data (to 25%) which need to be saved to disk and predicted without losing accuracy.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3706526%2Fcfd0afe964fcafd12ca8460914bd6816%2Fsolution1.png?generation=1677589474115470&amp;alt=media\" alt=\"\"><br>\nTo speed up prediction, we compiled our models using tensorRT. Through our experiments, we observed that the fp16 model was 1.3 times faster than fp32, but its numerical error is not negligible. In the end, we used models compiled in fp32.</p>\n<h2>Code</h2>\n<p>train: <a href=\"https://github.com/Masaaaato/RSNABreast7thPlace\" target=\"_blank\">https://github.com/Masaaaato/RSNABreast7thPlace</a><br>\ninference: <a href=\"https://www.kaggle.com/code/masato114/2stage-ensemble/notebook\" target=\"_blank\">https://www.kaggle.com/code/masato114/2stage-ensemble/notebook</a></p>",
      "rawMarkdown": "First of all, we would like to express our respect to all the participants and thank the organizers for making this competition possible. This competition was tough for us due to the volatility of the metrics and the fact that we had to complete our inference in less than 9 hours. In the end, by choosing our best local cv model (CV0.534), we were able to win 7th place.\nHere we will describe how we increased our CV score and how we made our entire inference faster.\n\n## Summary\n- Images: Kaggle train data and VinDr-Mammo as the external data\n- Preprocess: ROI cropping in a rule-based way and sigmoid windowing\n- Resolution: 1520x912\n- Model: EfficientNetV2S (and EfficientNet B5) with GeM pooling (p=3)\n- CV: 4-fold, grouped by patient and stratified by cancer, BIRADS, density, age, biopsy, implant, and machine_id\n- Train:\n    - Augmentation: V/H Flip, Geometric transformation (Affine and Elastic)\n    - Loss function: BCE\n    - Optimizer: Adam\n    - Scheduler: Cosine decay (starting from 5e-5)\n- Inference: 2xTTA (vertical flip)\n- Ensemble: seed averaging (2 seeds) and 2 level ensemble (breast- and laterality-level)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3706526%2Ff93bb633bb5bcfeffde33cb6931f5bd3%2Fsolution.png?generation=1677589173238328&alt=media)\n\n## Preprocessing\nWe obtained the ROI breast in a rule-based way. In brief, after converting values below 40 to 0, uniform columns and rows were removed because they were assumed to be background. This process was so simple yet effective and fast. After this preprocessing, we obtained images with an aspect ratio of 1:1.6~2 on average, and which were then resized to 1520x912. This resolution was determined after comparing four options: ①640x640, ②1024x1024, ③1520x912, and ④2689x1569. The order of CV score was ① < ② < ③ > ④. Additionally, we generated images with sigmoid windowing applied as well, which did not have a significant effect on improving the score, and even on ensemble. However, we trained some models with these images and included them in the ensemble to make our prediction more robust.\nThe most important thing was to generate 2 types of input images, namely breast-level and laterality-level. The breast-level consisted of one breast per image, while the laterality-level had two or more breasts per image by simply concatenating breasts in columns. The models trained with the laterality-level was much weaker (-0.04) than those with the breast-level, but the ensemble of the two levels was very effective.\n\n## Model Architectures\nWe found that larger models did not necessarily score better (EfficientNetB2 < EfficientNetB5 > EfficientNetB7). Out of several models, including EfficientNetV2, EfficientNetB5, SeResNeXt50, ConvNeXt tiny/small, and NextViT base, EfficientNetB5 performed the best (CV0.47 on 1520x912), and EfficientNetV2S came in second-best (CV0.45 on 1520x912). Additionally, we replaced the pooling layer from 'average' to 'generalized mean' (GeM), which resulted in a slight improvement in score (0.005~0.01). Note that we ultimately used EfficientNetV2S in inference because EfficientNetB5 takes longer to infer than EfficientNetV2S. However, the difference in metrics between the two was negligible (CV0.49, single model) by using B5's predictions on external data when training V2S, as described later.\n\n## External Data\nWe used VinDr-Mammo dataset as external data, whose labels were defined by breast-level predictions after aggregating them into laterality-level. We found a significant improvement in the CV score by 0.02 using this dataset.\n\n## Inference Speed Up\nThe inference time limit was very tight, so we made efforts to ensure inference completed in time. First, we used DALI to decode images and preprocessed most of them on GPU. This significantly increased processing speed. However, as mentioned [here](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/384231), some images could not be processed using DALI, so we used dicomSDL and cupy as a fallback method for decoding and preprocessing images. Additionally, we implemented a 2-stage method as shown in the following image. Through data analysis, we observed that the data contained many explicit negative samples which did not require prediction by a strong model or ensemble. By filtering out those data with the threshold of 0.01, we were able to reduce the number of data (to 25%) which need to be saved to disk and predicted without losing accuracy.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3706526%2Fcfd0afe964fcafd12ca8460914bd6816%2Fsolution1.png?generation=1677589474115470&alt=media)\nTo speed up prediction, we compiled our models using tensorRT. Through our experiments, we observed that the fp16 model was 1.3 times faster than fp32, but its numerical error is not negligible. In the end, we used models compiled in fp32.\n\n## Code\ntrain: [https://github.com/Masaaaato/RSNABreast7thPlace](https://github.com/Masaaaato/RSNABreast7thPlace)\ninference: [https://www.kaggle.com/code/masato114/2stage-ensemble/notebook](https://www.kaggle.com/code/masato114/2stage-ensemble/notebook)",
      "votes": 32
    },
    {
      "id": 2164003,
      "postDate": "2023-03-01T08:43:37.773Z",
      "content": "<p>Great work and congrats!</p>\n<p>\"The models trained with the laterality-level was much weaker (-0.04) than those with the breast-level, but the ensemble of the two levels was very effective.\"</p>\n<p>I have the first observation too but I just stopped there… 😑</p>",
      "rawMarkdown": "Great work and congrats!\n\n\"The models trained with the laterality-level was much weaker (-0.04) than those with the breast-level, but the ensemble of the two levels was very effective.\"\n\nI have the first observation too but I just stopped there... 😑",
      "votes": 1,
      "replies": [
        {
          "id": 2165561,
          "postDate": "2023-03-02T09:23:04.330Z",
          "content": "<p>Of course I considered discarding that weaker models, but I thought surely the lower covariance between breast- and laterality-level would be beneficial to the ensemble, and I was reassured by the fact that the scores were up in LB😏</p>",
          "rawMarkdown": "Of course I considered discarding that weaker models, but I thought surely the lower covariance between breast- and laterality-level would be beneficial to the ensemble, and I was reassured by the fact that the scores were up in LB😏",
          "votes": 1
        }
      ]
    },
    {
      "id": 2263351,
      "postDate": "2023-05-17T14:21:00.723Z",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/masato114\" target=\"_blank\">@masato114</a> i checked your code. So in the instruction after training with rsna images (step 1), creating pseudo labels (step 2), you instructed the below step 3 instruction: <br>\nChange the config file as follows according to your purpose.<br>\nFor breast-level/external dataset/no windowing: config1.yaml<br>\nFor laterality-level/external dataest/no windowing: config2.yaml<br>\nFor breast-level/external dataset/windowing: config3.yaml<br>\nFor laterality-level/external dataset/windowing: config4.yaml</p>\n<p>So my question is: for step 3 i will use train.py with the above configs right? Also do i need to finetune with lateral with vindr and then breast with vindr as shown in the figure or i can finetune them in any order and ensemble later? </p>",
      "rawMarkdown": "Hello @masato114 i checked your code. So in the instruction after training with rsna images (step 1), creating pseudo labels (step 2), you instructed the below step 3 instruction: \nChange the config file as follows according to your purpose.\nFor breast-level/external dataset/no windowing: config1.yaml\nFor laterality-level/external dataest/no windowing: config2.yaml\nFor breast-level/external dataset/windowing: config3.yaml\nFor laterality-level/external dataset/windowing: config4.yaml\n\nSo my question is: for step 3 i will use train.py with the above configs right? Also do i need to finetune with lateral with vindr and then breast with vindr as shown in the figure or i can finetune them in any order and ensemble later? ",
      "replies": [
        {
          "id": 2265147,
          "postDate": "2023-05-19T02:49:34.330Z",
          "content": "<p>Hi, thanks for your comment.</p>\n<p>Yes, you can use train.py with the configs you indicated for step 3. Once you successfully generated pseudolabels on vindr, which are labeled for each breast of each image, you can make lateral-level labels by aggregating them on both patient_id and laterality. This should be done before proceeding to step 3.</p>\n<p>As you are saying, you can finetune them in any order and ensemble later.</p>",
          "rawMarkdown": "Hi, thanks for your comment.\n\nYes, you can use train.py with the configs you indicated for step 3. Once you successfully generated pseudolabels on vindr, which are labeled for each breast of each image, you can make lateral-level labels by aggregating them on both patient_id and laterality. This should be done before proceeding to step 3.\n\nAs you are saying, you can finetune them in any order and ensemble later.",
          "replies": [
            {
              "id": 2265182,
              "postDate": "2023-05-19T03:50:09.807Z",
              "content": "<p><a href=\"https://www.kaggle.com/masato114\" target=\"_blank\">@masato114</a> Thanks for the reply. My further question is: <br>\n--<strong>This should be done before proceeding to step 3.</strong><br>\nSo is this the pseudo label generation step as mentioned in your github as: <br>\nConduct pseudolabeling on external dataset.<br>\npython -u src/external_pseudolabeling.py configs/config0.yaml</p>\n<p>These are the steps from your github readme.md:</p>\n<ol>\n<li>Train only using kaggle dataset like below.<br>\npython -u src/train.py configs/config0.yaml</li>\n<li>Conduct pseudolabeling on external dataset.<br>\npython -u src/external_pseudolabeling.py configs/config0.yaml</li>\n<li>Change the config file as follows according to your purpose.<br>\n(<strong>This was my question and i understand that i have to use train.py</strong>)<br>\nFor breast-level/external dataset/no windowing: config1.yaml<br>\nFor laterality-level/external dataest/no windowing: config2.yaml<br>\nFor breast-level/external dataset/windowing: config3.yaml<br>\nFor laterality-level/external dataset/windowing: config4.yaml</li>\n</ol>\n<p>In short, are the 3 steps specified in the github readme.md file are sufficient right? or u want any intermediate step as for generating lateral-level labels by aggregating them on both patient_id and laterality? Can you please clarify where to put any extra steps in the above-mentioned three steps as i am seeing<br>\nyou have generated the lateral labels in the external_pseudolabeling.py file (in VinDrMammo_breast-level_annotations.csv)? Can u please confirm?</p>",
              "rawMarkdown": "@masato114 Thanks for the reply. My further question is: \n--**This should be done before proceeding to step 3.**\nSo is this the pseudo label generation step as mentioned in your github as: \nConduct pseudolabeling on external dataset.\npython -u src/external_pseudolabeling.py configs/config0.yaml\n\nThese are the steps from your github readme.md:\n1. Train only using kaggle dataset like below.\npython -u src/train.py configs/config0.yaml\n2. Conduct pseudolabeling on external dataset.\npython -u src/external_pseudolabeling.py configs/config0.yaml\n3. Change the config file as follows according to your purpose.\n(**This was my question and i understand that i have to use train.py**)\nFor breast-level/external dataset/no windowing: config1.yaml\nFor laterality-level/external dataest/no windowing: config2.yaml\nFor breast-level/external dataset/windowing: config3.yaml\nFor laterality-level/external dataset/windowing: config4.yaml\n\nIn short, are the 3 steps specified in the github readme.md file are sufficient right? or u want any intermediate step as for generating lateral-level labels by aggregating them on both patient_id and laterality? Can you please clarify where to put any extra steps in the above-mentioned three steps as i am seeing\nyou have generated the lateral labels in the external_pseudolabeling.py file (in VinDrMammo_breast-level_annotations.csv)? Can u please confirm?"
            },
            {
              "id": 2265211,
              "postDate": "2023-05-19T04:44:58.087Z",
              "content": "<p>Hi, I apologize for making you confused.<br>\nThere are no extra steps between step 2 and 3. But perhaps you need to rename columns of vindr annotation file so that they are the same as of train.csv.  I may have overlooked something about vindr whose patient column is named as 'study_id' not 'patient_id', I guess.<br>\nIf you have hit errors, please kindly share the concrete problems.</p>",
              "rawMarkdown": "Hi, I apologize for making you confused.\nThere are no extra steps between step 2 and 3. But perhaps you need to rename columns of vindr annotation file so that they are the same as of train.csv.  I may have overlooked something about vindr whose patient column is named as 'study_id' not 'patient_id', I guess.\nIf you have hit errors, please kindly share the concrete problems."
            },
            {
              "id": 2265256,
              "postDate": "2023-05-19T05:24:37.293Z",
              "content": "<p>Thanks a lot..</p>",
              "rawMarkdown": "Thanks a lot.."
            },
            {
              "id": 2294255,
              "postDate": "2023-06-09T23:12:39.713Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2305764,
              "postDate": "2023-06-16T22:39:44.253Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2305778,
              "postDate": "2023-06-16T22:48:08.277Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2305791,
              "postDate": "2023-06-16T23:05:07.107Z",
              "content": "<p>Yes, I did the same process as what is applied to test images in inference notebook.</p>",
              "rawMarkdown": "Yes, I did the same process as what is applied to test images in inference notebook."
            },
            {
              "id": 2305800,
              "postDate": "2023-06-16T23:22:50.380Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/masato114\" target=\"_blank\">@masato114</a> , did u perform the same preprocessing for vindr dataset as u did for kaggle? can u please share the preprocessing for vindr dataset?</p>",
              "rawMarkdown": "Hi @masato114 , did u perform the same preprocessing for vindr dataset as u did for kaggle? can u please share the preprocessing for vindr dataset?"
            },
            {
              "id": 2305811,
              "postDate": "2023-06-16T23:30:55.130Z",
              "content": "<p>You can find it in <a href=\"https://www.kaggle.com/code/masato114/rsna-generate-train-images/notebook\" target=\"_blank\">https://www.kaggle.com/code/masato114/rsna-generate-train-images/notebook</a>.</p>",
              "rawMarkdown": "You can find it in https://www.kaggle.com/code/masato114/rsna-generate-train-images/notebook."
            },
            {
              "id": 2305832,
              "postDate": "2023-06-16T23:57:39.507Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2305854,
              "postDate": "2023-06-17T00:33:03.157Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2305968,
              "postDate": "2023-06-17T03:18:22.597Z",
              "content": "<p><a href=\"https://www.kaggle.com/masato114\" target=\"_blank\">@masato114</a>  the TransferSyntaxUID is 1.2.840.10008.1.2.1, it is different from the RSNA dataset. So, if condition of convert_dicom_to_j2k is not executed for vindr and format is also .dicom for vindr instead of .dcm for rsna. So either u have commented the if check for PhotometricInterpretation and TransferSyntaxUID. Can u please check?</p>\n<p>Also because Vindr images have different formats, the dali is throwing an error jp2:<br>\n[/opt/dali/dali/imgcodec/image_decoder.cc:448] Cannot parse the image: /ocean/projects/asc170022p/shg121/PhD/RSNA_Breast_Imaging/Dataset/External/Vindr/vindr-mammo-a-large-scale-benchmark-dataset-for-computer-aided-detection-and-diagnosis-in-full-field-digital-mammography-1.0.0/tmp/j2k/42ec6b7233c3e984cec02a1dd702dcb0_33c3590bb1738c0b9512630d9741bf65.jp2</p>\n<p>The temporary jp2 image is actually not forming up correctly.</p>\n<p>So, i am doing like this:</p>\n<p>import cv2</p>\n<p>SIZE = (912, 1520)<br>\npixel_data = ds.pixel_array<br>\nnormalized_data = ((pixel_data - np.min(pixel_data)) / np.ptp(pixel_data) * 255).astype(np.uint8)<br>\nimg = cv2.resize(normalized_data, SIZE, interpolation=cv2.INTER_AREA)<br>\nplt.imshow(img, cmap=plt.cm.bone)  <br>\nplt.show()</p>",
              "rawMarkdown": "@masato114  the TransferSyntaxUID is 1.2.840.10008.1.2.1, it is different from the RSNA dataset. So, if condition of convert_dicom_to_j2k is not executed for vindr and format is also .dicom for vindr instead of .dcm for rsna. So either u have commented the if check for PhotometricInterpretation and TransferSyntaxUID. Can u please check?\n\nAlso because Vindr images have different formats, the dali is throwing an error jp2:\n[/opt/dali/dali/imgcodec/image_decoder.cc:448] Cannot parse the image: /ocean/projects/asc170022p/shg121/PhD/RSNA_Breast_Imaging/Dataset/External/Vindr/vindr-mammo-a-large-scale-benchmark-dataset-for-computer-aided-detection-and-diagnosis-in-full-field-digital-mammography-1.0.0/tmp/j2k/42ec6b7233c3e984cec02a1dd702dcb0_33c3590bb1738c0b9512630d9741bf65.jp2\n\nThe temporary jp2 image is actually not forming up correctly.\n\nSo, i am doing like this:\n\nimport cv2\n\nSIZE = (912, 1520)\npixel_data = ds.pixel_array\nnormalized_data = ((pixel_data - np.min(pixel_data)) / np.ptp(pixel_data) * 255).astype(np.uint8)\nimg = cv2.resize(normalized_data, SIZE, interpolation=cv2.INTER_AREA)\nplt.imshow(img, cmap=plt.cm.bone)  \nplt.show()\n "
            },
            {
              "id": 2306019,
              "postDate": "2023-06-17T04:05:51.200Z",
              "content": "<p>Try the following.</p>\n<pre><code> ():\n    img_copy = img.copy()\n    img = np.where(img &lt;= , , img) \n    height, _ = img.shape\n\n    \n    y_a = height //  + (height*)\n    y_b = height //  - (height*)\n    b_arr = img[y_b:y_a].std(axis=) != \n    continuing_ones = CountUpContinuingOnes(b_arr)\n    \n    col_ind = np.where(continuing_ones == continuing_ones.())[]\n    img = img[:, col_ind]\n\n    \n    _, width = img.shape\n    x_a = width //  + (width*)\n    x_b = width //  - (width*)\n    b_arr = img[:,x_b:x_a].std(axis=) != \n    continuing_ones = CountUpContinuingOnes(b_arr)\n    \n    row_ind = np.where(continuing_ones == continuing_ones.())[]\n\n     img_copy[row_ind][:, col_ind]\n\n ():\n    dicom = dicomsdl.(in_path)\n    data = dicom.pixelData()\n    data = data[:-, :-]\n     dicom.getPixelDataInfo()[] == :\n        data = np.amax(data) - data\n\n    data = data - np.(data)\n    data = data / np.(data)\n    data = (data * ).astype(np.uint8)\n\n    img = ExtractBreast(data)\n    img = cv2.resize(img, SIZE, interpolation = cv2.INTER_AREA)\n    cv2.imwrite(out_path, img)\n</code></pre>",
              "rawMarkdown": "Try the following.\n```python\ndef ExtractBreast(img):\n    img_copy = img.copy()\n    img = np.where(img <= 40, 0, img) # To detect backgrounds easily\n    height, _ = img.shape\n\n    # whether each col is non-constant or not\n    y_a = height // 2 + int(height*0.4)\n    y_b = height // 2 - int(height*0.4)\n    b_arr = img[y_b:y_a].std(axis=0) != 0\n    continuing_ones = CountUpContinuingOnes(b_arr)\n    # longest should be the breast\n    col_ind = np.where(continuing_ones == continuing_ones.max())[0]\n    img = img[:, col_ind]\n\n    # whether each row is non-constant or not\n    _, width = img.shape\n    x_a = width // 2 + int(width*0.4)\n    x_b = width // 2 - int(width*0.4)\n    b_arr = img[:,x_b:x_a].std(axis=1) != 0\n    continuing_ones = CountUpContinuingOnes(b_arr)\n    # longest should be the breast\n    row_ind = np.where(continuing_ones == continuing_ones.max())[0]\n\n    return img_copy[row_ind][:, col_ind]\n\ndef func(in_path, out_path, SIZE=(912, 1520)):\n    dicom = dicomsdl.open(in_path)\n    data = dicom.pixelData()\n    data = data[5:-5, 5:-5]\n    if dicom.getPixelDataInfo()['PhotometricInterpretation'] == \"MONOCHROME1\":\n        data = np.amax(data) - data\n\n    data = data - np.min(data)\n    data = data / np.max(data)\n    data = (data * 255).astype(np.uint8)\n\n    img = ExtractBreast(data)\n    img = cv2.resize(img, SIZE, interpolation = cv2.INTER_AREA)\n    cv2.imwrite(out_path, img)\n```"
            },
            {
              "id": 2306140,
              "postDate": "2023-06-17T05:56:00.767Z",
              "content": "<p>Great thanks <a href=\"https://www.kaggle.com/masato114\" target=\"_blank\">@masato114</a> </p>",
              "rawMarkdown": "Great thanks @masato114 "
            },
            {
              "id": 2333448,
              "postDate": "2023-07-07T01:44:45.600Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/masato114\" target=\"_blank\">@masato114</a>, thanks for all the clarification for the training. I really appreciate your help and i am able to<br>\nsuccessfully replicate your training scripts for all the four cases.</p>\n<p>Now after training, following is the version number of all the models and their corresponding architectures:</p>\n<table>\n<thead>\n<tr>\n<th>Configs</th>\n<th>Breast/Laterity</th>\n<th>Dataset type</th>\n<th>Model type</th>\n<th>Model version</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>config0.yaml</td>\n<td>Breast Level</td>\n<td>Kaggle Data</td>\n<td>Efficientnet_b5</td>\n<td>model_ver_084</td>\n</tr>\n<tr>\n<td>config1.yaml</td>\n<td>Breast Level</td>\n<td>Kaggle + External Data</td>\n<td>Efficientnet_b2</td>\n<td>model_ver_084</td>\n</tr>\n<tr>\n<td>config2.yaml</td>\n<td>Lateral Level</td>\n<td>Kaggle + External Data</td>\n<td>Efficientnet_b2</td>\n<td>model_ver_085</td>\n</tr>\n<tr>\n<td>config3.yaml</td>\n<td>Breast Level</td>\n<td>Kaggle + External Data (with Sigmoid windowing)</td>\n<td>Efficientnet_b2</td>\n<td>model_ver_099</td>\n</tr>\n<tr>\n<td>config4.yaml</td>\n<td>Lateral Level</td>\n<td>Kaggle + External Data (with Sigmoid windowing)</td>\n<td>Efficientnet_b2</td>\n<td>model_ver_100</td>\n</tr>\n</tbody>\n</table>\n<p>Now when I am trying to replicate the code for ensemble as per <a href=\"https://www.kaggle.com/code/masato114/2stage-ensemble/notebook\" target=\"_blank\">https://www.kaggle.com/code/masato114/2stage-ensemble/notebook</a>, in the 20^th notebook, you are using the model kaggle/input/compile-ver49-83-83/ver49-{fold}.ts' to fill df['prediction'] data.</p>\n<p>Then in 23rd notebook for breast level, /kaggle/input/efficientnetv2s-exp084-tensorrt/efficientnetv2s_ver084-seed10-{fold}-swa.ts, /kaggle/input/efficientnetv2s-exp084-tensorrt/efficientnetv2s_ver084-seed660-{fold}-swa.ts and /kaggle/input/rsna-efficientnetv2s-pseudo-sig-tensorrt/efficientnetv2s_ver099-seed10-{fold}.ts are used to predict test_preds_1, test_preds_2 and test_preds_3 respectively. These 3 values are ensembled to calculate second_df['prediction_ver84'].</p>\n<p>Similary in 29^th notebook for laterality Level, /kaggle/input/efficientnetv2s-exp085-tensorrt/efficientnetv2s_ver085-{fold}.ts, kaggle/input/efficientnetv2s-exp085-tensorrt/efficientnetv2s_ver085-seed660-{fold}.ts, and /kaggle/input/rsna-efficientnetv2s-pseudo-sig-tensorrt/efficientnetv2s_ver100-seed10-{fold}.ts are used predict test_preds_1, test_preds_2 and test_preds_3 respectively. These 3 values are ensembled to calculate second_df['prediction_ver85'].</p>\n<p>Now my questions are as follows:</p>\n<ol>\n<li>How did you get model with version 049?</li>\n<li>To get the models with seed660 did you again train on breast and lateral level with config1 and config2? Is there any reason to use seed660?</li>\n<li>It means you only use efficientnetb5 to get the pseudo labels? And, you did not use efficientnetb5 for ensembling?</li>\n</ol>\n<p>It will be nice if you make these points and ensembling a bit more clear.</p>",
              "rawMarkdown": "Hi @masato114, thanks for all the clarification for the training. I really appreciate your help and i am able to\nsuccessfully replicate your training scripts for all the four cases.\n\nNow after training, following is the version number of all the models and their corresponding architectures:\n\n| Configs      | Breast/Laterity | Dataset type                                    | Model type      | Model version |\n|--------------|-----------------|-------------------------------------------------|-----------------|---------------|\n| config0.yaml | Breast Level    | Kaggle Data                                     | Efficientnet_b5 | model_ver_084 |\n| config1.yaml | Breast Level    | Kaggle + External Data                          | Efficientnet_b2 | model_ver_084 |\n| config2.yaml | Lateral Level   | Kaggle + External Data                          | Efficientnet_b2 | model_ver_085 |\n| config3.yaml | Breast Level    | Kaggle + External Data (with Sigmoid windowing) | Efficientnet_b2 | model_ver_099 |\n| config4.yaml | Lateral Level   | Kaggle + External Data (with Sigmoid windowing) | Efficientnet_b2 | model_ver_100 | \n\nNow when I am trying to replicate the code for ensemble as per https://www.kaggle.com/code/masato114/2stage-ensemble/notebook, in the 20^th notebook, you are using the model kaggle/input/compile-ver49-83-83/ver49-{fold}.ts' to fill df['prediction'] data.\n\nThen in 23rd notebook for breast level, /kaggle/input/efficientnetv2s-exp084-tensorrt/efficientnetv2s_ver084-seed10-{fold}-swa.ts, /kaggle/input/efficientnetv2s-exp084-tensorrt/efficientnetv2s_ver084-seed660-{fold}-swa.ts and /kaggle/input/rsna-efficientnetv2s-pseudo-sig-tensorrt/efficientnetv2s_ver099-seed10-{fold}.ts are used to predict test_preds_1, test_preds_2 and test_preds_3 respectively. These 3 values are ensembled to calculate second_df['prediction_ver84'].\n\nSimilary in 29^th notebook for laterality Level, /kaggle/input/efficientnetv2s-exp085-tensorrt/efficientnetv2s_ver085-{fold}.ts, kaggle/input/efficientnetv2s-exp085-tensorrt/efficientnetv2s_ver085-seed660-{fold}.ts, and /kaggle/input/rsna-efficientnetv2s-pseudo-sig-tensorrt/efficientnetv2s_ver100-seed10-{fold}.ts are used predict test_preds_1, test_preds_2 and test_preds_3 respectively. These 3 values are ensembled to calculate second_df['prediction_ver85'].\n\nNow my questions are as follows:\n1. How did you get model with version 049?\n2. To get the models with seed660 did you again train on breast and lateral level with config1 and config2? Is there any reason to use seed660?\n3. It means you only use efficientnetb5 to get the pseudo labels? And, you did not use efficientnetb5 for ensembling?\n\nIt will be nice if you make these points and ensembling a bit more clear."
            },
            {
              "id": 2333483,
              "postDate": "2023-07-07T03:00:54.890Z",
              "content": "<p>In short,</p>\n<ol>\n<li>Low resolution images were used for ver049. Since this is not special as to both the model architecture and the training procedure, we were not specific on that. They should be generated with reference to the inference code. However, if you simply want to keep the inference time to 9 hours, use the pre-trained models provided.</li>\n<li>That is because it was better out of several initial conditions.</li>\n<li>To shorten the inference time. Effb5 is heavier than Effv2s.</li>\n</ol>",
              "rawMarkdown": "In short,\n1. Low resolution images were used for ver049. Since this is not special as to both the model architecture and the training procedure, we were not specific on that. They should be generated with reference to the inference code. However, if you simply want to keep the inference time to 9 hours, use the pre-trained models provided.\n2. That is because it was better out of several initial conditions.\n3. To shorten the inference time. Effb5 is heavier than Effv2s."
            },
            {
              "id": 2333510,
              "postDate": "2023-07-07T03:17:16.300Z",
              "content": "<p>Great. Thanks for your help. Happy to collaborate.</p>",
              "rawMarkdown": "Great. Thanks for your help. Happy to collaborate."
            }
          ]
        }
      ]
    },
    {
      "id": 2164074,
      "postDate": "2023-03-01T09:30:50.307Z",
      "content": "<blockquote>\n  <p>Through data analysis, we observed that the data contained many explicit negative samples which did not require prediction by a strong model or ensemble. By filtering out those data with the threshold of 0.01, we were able to reduce the number of data (to 25%) which need to be saved to disk and predicted without losing accuracy.</p>\n</blockquote>\n<p>Congratulations and thank you for sharing your work! May I know what data analysis you applied to filter the test (hidden) images?</p>",
      "rawMarkdown": ">Through data analysis, we observed that the data contained many explicit negative samples which did not require prediction by a strong model or ensemble. By filtering out those data with the threshold of 0.01, we were able to reduce the number of data (to 25%) which need to be saved to disk and predicted without losing accuracy.\n\nCongratulations and thank you for sharing your work! May I know what data analysis you applied to filter the test (hidden) images?",
      "replies": [
        {
          "id": 2164931,
          "postDate": "2023-03-01T21:41:33.947Z",
          "content": "<p>Thank you for your question.<br>\nI did an analysis against oof predictions like this.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9268451%2F980fe355177fd722e585bd5d7e5dcae2%2Fanasysis.png?generation=1677707481367684&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "Thank you for your question.\nI did an analysis against oof predictions like this.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9268451%2F980fe355177fd722e585bd5d7e5dcae2%2Fanasysis.png?generation=1677707481367684&alt=media)",
          "votes": 1
        }
      ]
    },
    {
      "id": 2163816,
      "postDate": "2023-03-01T05:16:45.963Z",
      "content": "<blockquote>\n  <p>Through data analysis, we observed that the data contained many explicit negative samples which did not require prediction by a strong model or ensemble. By filtering out those data with the threshold of 0.01, we were able to reduce the number of data (to 25%) which need to be saved to disk and predicted without losing accuracy.</p>\n</blockquote>\n<p>That is very inspiring👍</p>",
      "rawMarkdown": ">Through data analysis, we observed that the data contained many explicit negative samples which did not require prediction by a strong model or ensemble. By filtering out those data with the threshold of 0.01, we were able to reduce the number of data (to 25%) which need to be saved to disk and predicted without losing accuracy.\n\nThat is very inspiring👍"
    },
    {
      "id": 2163005,
      "postDate": "2023-02-28T14:27:02.817Z",
      "content": "<p>Thanks for the thorough explanation! 👍 Was DALI used for the augmentation also? Elastic deformation can be quite resource consuming, did you have your own implementation for that one?</p>",
      "rawMarkdown": "Thanks for the thorough explanation! 👍 Was DALI used for the augmentation also? Elastic deformation can be quite resource consuming, did you have your own implementation for that one?",
      "replies": [
        {
          "id": 2163071,
          "postDate": "2023-02-28T15:11:12.577Z",
          "content": "<p>Thanks. Once images were loaded with DALI and saved in storage, DALI was no longer used, even for augmentation. Elastic deformation was implemented by 　<code>albumentations.augmentations.geometric.transforms.ElasticTransform</code> with alpha=10 and sigma=15.</p>",
          "rawMarkdown": "Thanks. Once images were loaded with DALI and saved in storage, DALI was no longer used, even for augmentation. Elastic deformation was implemented by 　`albumentations.augmentations.geometric.transforms.ElasticTransform` with alpha=10 and sigma=15.",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2164003,
      "author_name": "Chenglu",
      "author_url": "",
      "post_date": "2023-03-01T08:43:37.773000",
      "content": "<p>Great work and congrats!</p>\n<p>\"The models trained with the laterality-level was much weaker (-0.04) than those with the breast-level, but the ensemble of the two levels was very effective.\"</p>\n<p>I have the first observation too but I just stopped there… 😑</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2165561,
          "author_name": "Mt.Panda",
          "author_url": "",
          "post_date": "2023-03-02T09:23:04.330000",
          "content": "<p>Of course I considered discarding that weaker models, but I thought surely the lower covariance between breast- and laterality-level would be beneficial to the ensemble, and I was reassured by the fact that the scores were up in LB😏</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2263351,
      "author_name": "Shantanu Ghosh",
      "author_url": "",
      "post_date": "2023-05-17T14:21:00.723000",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/masato114\" target=\"_blank\">@masato114</a> i checked your code. So in the instruction after training with rsna images (step 1), creating pseudo labels (step 2), you instructed the below step 3 instruction: <br>\nChange the config file as follows according to your purpose.<br>\nFor breast-level/external dataset/no windowing: config1.yaml<br>\nFor laterality-level/external dataest/no windowing: config2.yaml<br>\nFor breast-level/external dataset/windowing: config3.yaml<br>\nFor laterality-level/external dataset/windowing: config4.yaml</p>\n<p>So my question is: for step 3 i will use train.py with the above configs right? Also do i need to finetune with lateral with vindr and then breast with vindr as shown in the figure or i can finetune them in any order and ensemble later? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2265147,
          "author_name": "Mt.Panda",
          "author_url": "",
          "post_date": "2023-05-19T02:49:34.330000",
          "content": "<p>Hi, thanks for your comment.</p>\n<p>Yes, you can use train.py with the configs you indicated for step 3. Once you successfully generated pseudolabels on vindr, which are labeled for each breast of each image, you can make lateral-level labels by aggregating them on both patient_id and laterality. This should be done before proceeding to step 3.</p>\n<p>As you are saying, you can finetune them in any order and ensemble later.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2265182,
              "author_name": "Shantanu Ghosh",
              "author_url": "",
              "post_date": "2023-05-19T03:50:09.807000",
              "content": "<p><a href=\"https://www.kaggle.com/masato114\" target=\"_blank\">@masato114</a> Thanks for the reply. My further question is: <br>\n--<strong>This should be done before proceeding to step 3.</strong><br>\nSo is this the pseudo label generation step as mentioned in your github as: <br>\nConduct pseudolabeling on external dataset.<br>\npython -u src/external_pseudolabeling.py configs/config0.yaml</p>\n<p>These are the steps from your github readme.md:</p>\n<ol>\n<li>Train only using kaggle dataset like below.<br>\npython -u src/train.py configs/config0.yaml</li>\n<li>Conduct pseudolabeling on external dataset.<br>\npython -u src/external_pseudolabeling.py configs/config0.yaml</li>\n<li>Change the config file as follows according to your purpose.<br>\n(<strong>This was my question and i understand that i have to use train.py</strong>)<br>\nFor breast-level/external dataset/no windowing: config1.yaml<br>\nFor laterality-level/external dataest/no windowing: config2.yaml<br>\nFor breast-level/external dataset/windowing: config3.yaml<br>\nFor laterality-level/external dataset/windowing: config4.yaml</li>\n</ol>\n<p>In short, are the 3 steps specified in the github readme.md file are sufficient right? or u want any intermediate step as for generating lateral-level labels by aggregating them on both patient_id and laterality? Can you please clarify where to put any extra steps in the above-mentioned three steps as i am seeing<br>\nyou have generated the lateral labels in the external_pseudolabeling.py file (in VinDrMammo_breast-level_annotations.csv)? Can u please confirm?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2265211,
              "author_name": "Mt.Panda",
              "author_url": "",
              "post_date": "2023-05-19T04:44:58.087000",
              "content": "<p>Hi, I apologize for making you confused.<br>\nThere are no extra steps between step 2 and 3. But perhaps you need to rename columns of vindr annotation file so that they are the same as of train.csv.  I may have overlooked something about vindr whose patient column is named as 'study_id' not 'patient_id', I guess.<br>\nIf you have hit errors, please kindly share the concrete problems.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2265256,
              "author_name": "Shantanu Ghosh",
              "author_url": "",
              "post_date": "2023-05-19T05:24:37.293000",
              "content": "<p>Thanks a lot..</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2294255,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-06-09T23:12:39.713000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2305764,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-06-16T22:39:44.253000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2305778,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-06-16T22:48:08.277000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2305791,
              "author_name": "Mt.Panda",
              "author_url": "",
              "post_date": "2023-06-16T23:05:07.107000",
              "content": "<p>Yes, I did the same process as what is applied to test images in inference notebook.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2305800,
              "author_name": "Shantanu Ghosh",
              "author_url": "",
              "post_date": "2023-06-16T23:22:50.380000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/masato114\" target=\"_blank\">@masato114</a> , did u perform the same preprocessing for vindr dataset as u did for kaggle? can u please share the preprocessing for vindr dataset?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2305811,
              "author_name": "Mt.Panda",
              "author_url": "",
              "post_date": "2023-06-16T23:30:55.130000",
              "content": "<p>You can find it in <a href=\"https://www.kaggle.com/code/masato114/rsna-generate-train-images/notebook\" target=\"_blank\">https://www.kaggle.com/code/masato114/rsna-generate-train-images/notebook</a>.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2305832,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-06-16T23:57:39.507000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2305854,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-06-17T00:33:03.157000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2305968,
              "author_name": "Shantanu Ghosh",
              "author_url": "",
              "post_date": "2023-06-17T03:18:22.597000",
              "content": "<p><a href=\"https://www.kaggle.com/masato114\" target=\"_blank\">@masato114</a>  the TransferSyntaxUID is 1.2.840.10008.1.2.1, it is different from the RSNA dataset. So, if condition of convert_dicom_to_j2k is not executed for vindr and format is also .dicom for vindr instead of .dcm for rsna. So either u have commented the if check for PhotometricInterpretation and TransferSyntaxUID. Can u please check?</p>\n<p>Also because Vindr images have different formats, the dali is throwing an error jp2:<br>\n[/opt/dali/dali/imgcodec/image_decoder.cc:448] Cannot parse the image: /ocean/projects/asc170022p/shg121/PhD/RSNA_Breast_Imaging/Dataset/External/Vindr/vindr-mammo-a-large-scale-benchmark-dataset-for-computer-aided-detection-and-diagnosis-in-full-field-digital-mammography-1.0.0/tmp/j2k/42ec6b7233c3e984cec02a1dd702dcb0_33c3590bb1738c0b9512630d9741bf65.jp2</p>\n<p>The temporary jp2 image is actually not forming up correctly.</p>\n<p>So, i am doing like this:</p>\n<p>import cv2</p>\n<p>SIZE = (912, 1520)<br>\npixel_data = ds.pixel_array<br>\nnormalized_data = ((pixel_data - np.min(pixel_data)) / np.ptp(pixel_data) * 255).astype(np.uint8)<br>\nimg = cv2.resize(normalized_data, SIZE, interpolation=cv2.INTER_AREA)<br>\nplt.imshow(img, cmap=plt.cm.bone)  <br>\nplt.show()</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2306019,
              "author_name": "Mt.Panda",
              "author_url": "",
              "post_date": "2023-06-17T04:05:51.200000",
              "content": "<p>Try the following.</p>\n<pre><code> ():\n    img_copy = img.copy()\n    img = np.where(img &lt;= , , img) \n    height, _ = img.shape\n\n    \n    y_a = height //  + (height*)\n    y_b = height //  - (height*)\n    b_arr = img[y_b:y_a].std(axis=) != \n    continuing_ones = CountUpContinuingOnes(b_arr)\n    \n    col_ind = np.where(continuing_ones == continuing_ones.())[]\n    img = img[:, col_ind]\n\n    \n    _, width = img.shape\n    x_a = width //  + (width*)\n    x_b = width //  - (width*)\n    b_arr = img[:,x_b:x_a].std(axis=) != \n    continuing_ones = CountUpContinuingOnes(b_arr)\n    \n    row_ind = np.where(continuing_ones == continuing_ones.())[]\n\n     img_copy[row_ind][:, col_ind]\n\n ():\n    dicom = dicomsdl.(in_path)\n    data = dicom.pixelData()\n    data = data[:-, :-]\n     dicom.getPixelDataInfo()[] == :\n        data = np.amax(data) - data\n\n    data = data - np.(data)\n    data = data / np.(data)\n    data = (data * ).astype(np.uint8)\n\n    img = ExtractBreast(data)\n    img = cv2.resize(img, SIZE, interpolation = cv2.INTER_AREA)\n    cv2.imwrite(out_path, img)\n</code></pre>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2306140,
              "author_name": "Shantanu Ghosh",
              "author_url": "",
              "post_date": "2023-06-17T05:56:00.767000",
              "content": "<p>Great thanks <a href=\"https://www.kaggle.com/masato114\" target=\"_blank\">@masato114</a> </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2333448,
              "author_name": "Shantanu Ghosh",
              "author_url": "",
              "post_date": "2023-07-07T01:44:45.600000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/masato114\" target=\"_blank\">@masato114</a>, thanks for all the clarification for the training. I really appreciate your help and i am able to<br>\nsuccessfully replicate your training scripts for all the four cases.</p>\n<p>Now after training, following is the version number of all the models and their corresponding architectures:</p>\n<table>\n<thead>\n<tr>\n<th>Configs</th>\n<th>Breast/Laterity</th>\n<th>Dataset type</th>\n<th>Model type</th>\n<th>Model version</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>config0.yaml</td>\n<td>Breast Level</td>\n<td>Kaggle Data</td>\n<td>Efficientnet_b5</td>\n<td>model_ver_084</td>\n</tr>\n<tr>\n<td>config1.yaml</td>\n<td>Breast Level</td>\n<td>Kaggle + External Data</td>\n<td>Efficientnet_b2</td>\n<td>model_ver_084</td>\n</tr>\n<tr>\n<td>config2.yaml</td>\n<td>Lateral Level</td>\n<td>Kaggle + External Data</td>\n<td>Efficientnet_b2</td>\n<td>model_ver_085</td>\n</tr>\n<tr>\n<td>config3.yaml</td>\n<td>Breast Level</td>\n<td>Kaggle + External Data (with Sigmoid windowing)</td>\n<td>Efficientnet_b2</td>\n<td>model_ver_099</td>\n</tr>\n<tr>\n<td>config4.yaml</td>\n<td>Lateral Level</td>\n<td>Kaggle + External Data (with Sigmoid windowing)</td>\n<td>Efficientnet_b2</td>\n<td>model_ver_100</td>\n</tr>\n</tbody>\n</table>\n<p>Now when I am trying to replicate the code for ensemble as per <a href=\"https://www.kaggle.com/code/masato114/2stage-ensemble/notebook\" target=\"_blank\">https://www.kaggle.com/code/masato114/2stage-ensemble/notebook</a>, in the 20^th notebook, you are using the model kaggle/input/compile-ver49-83-83/ver49-{fold}.ts' to fill df['prediction'] data.</p>\n<p>Then in 23rd notebook for breast level, /kaggle/input/efficientnetv2s-exp084-tensorrt/efficientnetv2s_ver084-seed10-{fold}-swa.ts, /kaggle/input/efficientnetv2s-exp084-tensorrt/efficientnetv2s_ver084-seed660-{fold}-swa.ts and /kaggle/input/rsna-efficientnetv2s-pseudo-sig-tensorrt/efficientnetv2s_ver099-seed10-{fold}.ts are used to predict test_preds_1, test_preds_2 and test_preds_3 respectively. These 3 values are ensembled to calculate second_df['prediction_ver84'].</p>\n<p>Similary in 29^th notebook for laterality Level, /kaggle/input/efficientnetv2s-exp085-tensorrt/efficientnetv2s_ver085-{fold}.ts, kaggle/input/efficientnetv2s-exp085-tensorrt/efficientnetv2s_ver085-seed660-{fold}.ts, and /kaggle/input/rsna-efficientnetv2s-pseudo-sig-tensorrt/efficientnetv2s_ver100-seed10-{fold}.ts are used predict test_preds_1, test_preds_2 and test_preds_3 respectively. These 3 values are ensembled to calculate second_df['prediction_ver85'].</p>\n<p>Now my questions are as follows:</p>\n<ol>\n<li>How did you get model with version 049?</li>\n<li>To get the models with seed660 did you again train on breast and lateral level with config1 and config2? Is there any reason to use seed660?</li>\n<li>It means you only use efficientnetb5 to get the pseudo labels? And, you did not use efficientnetb5 for ensembling?</li>\n</ol>\n<p>It will be nice if you make these points and ensembling a bit more clear.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2333483,
              "author_name": "Mt.Panda",
              "author_url": "",
              "post_date": "2023-07-07T03:00:54.890000",
              "content": "<p>In short,</p>\n<ol>\n<li>Low resolution images were used for ver049. Since this is not special as to both the model architecture and the training procedure, we were not specific on that. They should be generated with reference to the inference code. However, if you simply want to keep the inference time to 9 hours, use the pre-trained models provided.</li>\n<li>That is because it was better out of several initial conditions.</li>\n<li>To shorten the inference time. Effb5 is heavier than Effv2s.</li>\n</ol>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2333510,
              "author_name": "Shantanu Ghosh",
              "author_url": "",
              "post_date": "2023-07-07T03:17:16.300000",
              "content": "<p>Great. Thanks for your help. Happy to collaborate.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2164074,
      "author_name": "Chua Jia Qing Isaiah",
      "author_url": "",
      "post_date": "2023-03-01T09:30:50.307000",
      "content": "<blockquote>\n  <p>Through data analysis, we observed that the data contained many explicit negative samples which did not require prediction by a strong model or ensemble. By filtering out those data with the threshold of 0.01, we were able to reduce the number of data (to 25%) which need to be saved to disk and predicted without losing accuracy.</p>\n</blockquote>\n<p>Congratulations and thank you for sharing your work! May I know what data analysis you applied to filter the test (hidden) images?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2164931,
          "author_name": "luddite^",
          "author_url": "",
          "post_date": "2023-03-01T21:41:33.947000",
          "content": "<p>Thank you for your question.<br>\nI did an analysis against oof predictions like this.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9268451%2F980fe355177fd722e585bd5d7e5dcae2%2Fanasysis.png?generation=1677707481367684&amp;alt=media\" alt=\"\"></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2163816,
      "author_name": "RickyLu",
      "author_url": "",
      "post_date": "2023-03-01T05:16:45.963000",
      "content": "<blockquote>\n  <p>Through data analysis, we observed that the data contained many explicit negative samples which did not require prediction by a strong model or ensemble. By filtering out those data with the threshold of 0.01, we were able to reduce the number of data (to 25%) which need to be saved to disk and predicted without losing accuracy.</p>\n</blockquote>\n<p>That is very inspiring👍</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2163005,
      "author_name": "Antti Isosalo",
      "author_url": "",
      "post_date": "2023-02-28T14:27:02.817000",
      "content": "<p>Thanks for the thorough explanation! 👍 Was DALI used for the augmentation also? Elastic deformation can be quite resource consuming, did you have your own implementation for that one?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2163071,
          "author_name": "Mt.Panda",
          "author_url": "",
          "post_date": "2023-02-28T15:11:12.577000",
          "content": "<p>Thanks. Once images were loaded with DALI and saved in storage, DALI was no longer used, even for augmentation. Elastic deformation was implemented by 　<code>albumentations.augmentations.geometric.transforms.ElasticTransform</code> with alpha=10 and sigma=15.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2162872": "First of all, we would like to express our respect to all the participants and thank the organizers for making this competition possible. This competition was tough for us due to the volatility of the metrics and the fact that we had to complete our inference in less than 9 hours. In the end, by choosing our best local cv model (CV0.534), we were able to win 7th place.\nHere we will describe how we increased our CV score and how we made our entire inference faster.\n\n## Summary\n- Images: Kaggle train data and VinDr-Mammo as the external data\n- Preprocess: ROI cropping in a rule-based way and sigmoid windowing\n- Resolution: 1520x912\n- Model: EfficientNetV2S (and EfficientNet B5) with GeM pooling (p=3)\n- CV: 4-fold, grouped by patient and stratified by cancer, BIRADS, density, age, biopsy, implant, and machine_id\n- Train:\n    - Augmentation: V/H Flip, Geometric transformation (Affine and Elastic)\n    - Loss function: BCE\n    - Optimizer: Adam\n    - Scheduler: Cosine decay (starting from 5e-5)\n- Inference: 2xTTA (vertical flip)\n- Ensemble: seed averaging (2 seeds) and 2 level ensemble (breast- and laterality-level)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3706526%2Ff93bb633bb5bcfeffde33cb6931f5bd3%2Fsolution.png?generation=1677589173238328&alt=media)\n\n## Preprocessing\nWe obtained the ROI breast in a rule-based way. In brief, after converting values below 40 to 0, uniform columns and rows were removed because they were assumed to be background. This process was so simple yet effective and fast. After this preprocessing, we obtained images with an aspect ratio of 1:1.6~2 on average, and which were then resized to 1520x912. This resolution was determined after comparing four options: ①640x640, ②1024x1024, ③1520x912, and ④2689x1569. The order of CV score was ① < ② < ③ > ④. Additionally, we generated images with sigmoid windowing applied as well, which did not have a significant effect on improving the score, and even on ensemble. However, we trained some models with these images and included them in the ensemble to make our prediction more robust.\nThe most important thing was to generate 2 types of input images, namely breast-level and laterality-level. The breast-level consisted of one breast per image, while the laterality-level had two or more breasts per image by simply concatenating breasts in columns. The models trained with the laterality-level was much weaker (-0.04) than those with the breast-level, but the ensemble of the two levels was very effective.\n\n## Model Architectures\nWe found that larger models did not necessarily score better (EfficientNetB2 < EfficientNetB5 > EfficientNetB7). Out of several models, including EfficientNetV2, EfficientNetB5, SeResNeXt50, ConvNeXt tiny/small, and NextViT base, EfficientNetB5 performed the best (CV0.47 on 1520x912), and EfficientNetV2S came in second-best (CV0.45 on 1520x912). Additionally, we replaced the pooling layer from 'average' to 'generalized mean' (GeM), which resulted in a slight improvement in score (0.005~0.01). Note that we ultimately used EfficientNetV2S in inference because EfficientNetB5 takes longer to infer than EfficientNetV2S. However, the difference in metrics between the two was negligible (CV0.49, single model) by using B5's predictions on external data when training V2S, as described later.\n\n## External Data\nWe used VinDr-Mammo dataset as external data, whose labels were defined by breast-level predictions after aggregating them into laterality-level. We found a significant improvement in the CV score by 0.02 using this dataset.\n\n## Inference Speed Up\nThe inference time limit was very tight, so we made efforts to ensure inference completed in time. First, we used DALI to decode images and preprocessed most of them on GPU. This significantly increased processing speed. However, as mentioned [here](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/384231), some images could not be processed using DALI, so we used dicomSDL and cupy as a fallback method for decoding and preprocessing images. Additionally, we implemented a 2-stage method as shown in the following image. Through data analysis, we observed that the data contained many explicit negative samples which did not require prediction by a strong model or ensemble. By filtering out those data with the threshold of 0.01, we were able to reduce the number of data (to 25%) which need to be saved to disk and predicted without losing accuracy.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3706526%2Fcfd0afe964fcafd12ca8460914bd6816%2Fsolution1.png?generation=1677589474115470&alt=media)\nTo speed up prediction, we compiled our models using tensorRT. Through our experiments, we observed that the fp16 model was 1.3 times faster than fp32, but its numerical error is not negligible. In the end, we used models compiled in fp32.\n\n## Code\ntrain: [https://github.com/Masaaaato/RSNABreast7thPlace](https://github.com/Masaaaato/RSNABreast7thPlace)\ninference: [https://www.kaggle.com/code/masato114/2stage-ensemble/notebook](https://www.kaggle.com/code/masato114/2stage-ensemble/notebook)",
    "2164003": "Great work and congrats!\n\n\"The models trained with the laterality-level was much weaker (-0.04) than those with the breast-level, but the ensemble of the two levels was very effective.\"\n\nI have the first observation too but I just stopped there... 😑",
    "2263351": "Hello @masato114 i checked your code. So in the instruction after training with rsna images (step 1), creating pseudo labels (step 2), you instructed the below step 3 instruction: \nChange the config file as follows according to your purpose.\nFor breast-level/external dataset/no windowing: config1.yaml\nFor laterality-level/external dataest/no windowing: config2.yaml\nFor breast-level/external dataset/windowing: config3.yaml\nFor laterality-level/external dataset/windowing: config4.yaml\n\nSo my question is: for step 3 i will use train.py with the above configs right? Also do i need to finetune with lateral with vindr and then breast with vindr as shown in the figure or i can finetune them in any order and ensemble later? ",
    "2164074": ">Through data analysis, we observed that the data contained many explicit negative samples which did not require prediction by a strong model or ensemble. By filtering out those data with the threshold of 0.01, we were able to reduce the number of data (to 25%) which need to be saved to disk and predicted without losing accuracy.\n\nCongratulations and thank you for sharing your work! May I know what data analysis you applied to filter the test (hidden) images?",
    "2163816": ">Through data analysis, we observed that the data contained many explicit negative samples which did not require prediction by a strong model or ensemble. By filtering out those data with the threshold of 0.01, we were able to reduce the number of data (to 25%) which need to be saved to disk and predicted without losing accuracy.\n\nThat is very inspiring👍",
    "2163005": "Thanks for the thorough explanation! 👍 Was DALI used for the augmentation also? Elastic deformation can be quite resource consuming, did you have your own implementation for that one?"
  }
}