{
  "id": 391022,
  "title": "26th solution",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/391022",
  "author_name": "Chenglu",
  "post_date": "2023-02-28T06:04:23.739000",
  "votes": 28,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Training source code: <a href=\"https://github.com/louis-she/rsna-2022-public\" target=\"_blank\">https://github.com/louis-she/rsna-2022-public</a><br>\nInfernece notebook: <a href=\"https://www.kaggle.com/code/snaker/rsna-infer-t4-threshold?scriptVersionId=120550686\" target=\"_blank\">https://www.kaggle.com/code/snaker/rsna-infer-t4-threshold?scriptVersionId=120550686</a></p>\n<p>Thank everyone who contributed to the competition and thanks Kaggle for hosting.</p>\n<p>My solution didn't involve complicated pipeline, just as most does, single image classification. Here are some settings you may be interested in:</p>\n<ul>\n<li><strong>Augmentation</strong> RandomRotate90, HorizontalFlip, VerticalFlip</li>\n<li><strong>Models and Resolutions</strong> <ol>\n<li>ConvnextV2_nano @ 1536</li>\n<li>ConvnextV2_nano_another_seed @ 1536</li>\n<li>ConvnextV2_nano @ 2048</li>\n<li>EfficientnetV2_s @ 1536</li></ol></li>\n<li><strong>Optimizer and Scheduler</strong> Adam with 3e-5 + OneCycleLR</li>\n<li><strong>Epochs</strong> 33</li>\n<li><strong>Agg</strong> by maximum</li>\n<li>Aux Losses: no</li>\n<li>Tablet Features: no</li>\n<li>SWA: no</li>\n<li>External Dataset: no</li>\n</ul>\n<h2>The only trick: drop false positives</h2>\n<p>I realized that all the label tagged for <code>laterality + patient_id</code> is the same. That means the label is actually for <code>patient + laterality</code>, not image. So the easiest way to deal with this, is to train OOF and drop False Negatives with high confidence( e.g. the label is 1 but model predict probability is like 0.1 or 0.01). The CV score boost huge after this is done( about 0.03 to 0.04 ).</p>\n<h2>Things I have learned</h2>\n<p>The performance matters a lot in the competition. I'm happy that I take almost one month to optimize the inference pipeline, and I really learned something, like open-sourced the <a href=\"https://github.com/louis-she/nvjpeg2k-python\" target=\"_blank\">nvjpeg2k-python</a> ( first time writing some cuda code, it feels good ). To utilze the maximum of the hardware, I use the 2 x T4 kernel and some multiprocessing tricks. You can see that in the inference notebook.</p>\n<p>I use the percentile threshold trick I mentioned in the thread, cause it boost my LB score about ~0.015 . I know that in this way I will have a very high chance to overfit LB but I still chose it as one of my last selection sub, because it is just too good on LB… The result of manual threshold without percentile is 0.61 , the 0.63 score of LB is so attractive that I left everything behind.</p>\n<hr>\n<p>Congrats to all the winners! And thanks for the great community, let's meet in another competition.</p>",
  "messages": [
    {
      "id": 2162323,
      "postDate": "2023-02-28T06:04:23.740Z",
      "content": "<p>Training source code: <a href=\"https://github.com/louis-she/rsna-2022-public\" target=\"_blank\">https://github.com/louis-she/rsna-2022-public</a><br>\nInfernece notebook: <a href=\"https://www.kaggle.com/code/snaker/rsna-infer-t4-threshold?scriptVersionId=120550686\" target=\"_blank\">https://www.kaggle.com/code/snaker/rsna-infer-t4-threshold?scriptVersionId=120550686</a></p>\n<p>Thank everyone who contributed to the competition and thanks Kaggle for hosting.</p>\n<p>My solution didn't involve complicated pipeline, just as most does, single image classification. Here are some settings you may be interested in:</p>\n<ul>\n<li><strong>Augmentation</strong> RandomRotate90, HorizontalFlip, VerticalFlip</li>\n<li><strong>Models and Resolutions</strong> <ol>\n<li>ConvnextV2_nano @ 1536</li>\n<li>ConvnextV2_nano_another_seed @ 1536</li>\n<li>ConvnextV2_nano @ 2048</li>\n<li>EfficientnetV2_s @ 1536</li></ol></li>\n<li><strong>Optimizer and Scheduler</strong> Adam with 3e-5 + OneCycleLR</li>\n<li><strong>Epochs</strong> 33</li>\n<li><strong>Agg</strong> by maximum</li>\n<li>Aux Losses: no</li>\n<li>Tablet Features: no</li>\n<li>SWA: no</li>\n<li>External Dataset: no</li>\n</ul>\n<h2>The only trick: drop false positives</h2>\n<p>I realized that all the label tagged for <code>laterality + patient_id</code> is the same. That means the label is actually for <code>patient + laterality</code>, not image. So the easiest way to deal with this, is to train OOF and drop False Negatives with high confidence( e.g. the label is 1 but model predict probability is like 0.1 or 0.01). The CV score boost huge after this is done( about 0.03 to 0.04 ).</p>\n<h2>Things I have learned</h2>\n<p>The performance matters a lot in the competition. I'm happy that I take almost one month to optimize the inference pipeline, and I really learned something, like open-sourced the <a href=\"https://github.com/louis-she/nvjpeg2k-python\" target=\"_blank\">nvjpeg2k-python</a> ( first time writing some cuda code, it feels good ). To utilze the maximum of the hardware, I use the 2 x T4 kernel and some multiprocessing tricks. You can see that in the inference notebook.</p>\n<p>I use the percentile threshold trick I mentioned in the thread, cause it boost my LB score about ~0.015 . I know that in this way I will have a very high chance to overfit LB but I still chose it as one of my last selection sub, because it is just too good on LB… The result of manual threshold without percentile is 0.61 , the 0.63 score of LB is so attractive that I left everything behind.</p>\n<hr>\n<p>Congrats to all the winners! And thanks for the great community, let's meet in another competition.</p>",
      "rawMarkdown": "Training source code: https://github.com/louis-she/rsna-2022-public\nInfernece notebook: https://www.kaggle.com/code/snaker/rsna-infer-t4-threshold?scriptVersionId=120550686\n\nThank everyone who contributed to the competition and thanks Kaggle for hosting.\n\nMy solution didn't involve complicated pipeline, just as most does, single image classification. Here are some settings you may be interested in:\n\n* **Augmentation** RandomRotate90, HorizontalFlip, VerticalFlip\n* **Models and Resolutions** \n  1. ConvnextV2_nano @ 1536\n  2. ConvnextV2_nano_another_seed @ 1536\n  3. ConvnextV2_nano @ 2048\n  4. EfficientnetV2_s @ 1536\n* **Optimizer and Scheduler** Adam with 3e-5 + OneCycleLR\n* **Epochs** 33\n* **Agg** by maximum\n* Aux Losses: no\n* Tablet Features: no\n* SWA: no\n* External Dataset: no\n\n## The only trick: drop false positives\n\nI realized that all the label tagged for `laterality + patient_id` is the same. That means the label is actually for `patient + laterality`, not image. So the easiest way to deal with this, is to train OOF and drop False Negatives with high confidence( e.g. the label is 1 but model predict probability is like 0.1 or 0.01). The CV score boost huge after this is done( about 0.03 to 0.04 ).\n\n## Things I have learned\n\nThe performance matters a lot in the competition. I'm happy that I take almost one month to optimize the inference pipeline, and I really learned something, like open-sourced the [nvjpeg2k-python](https://github.com/louis-she/nvjpeg2k-python) ( first time writing some cuda code, it feels good ). To utilze the maximum of the hardware, I use the 2 x T4 kernel and some multiprocessing tricks. You can see that in the inference notebook.\n\nI use the percentile threshold trick I mentioned in the thread, cause it boost my LB score about ~0.015 . I know that in this way I will have a very high chance to overfit LB but I still chose it as one of my last selection sub, because it is just too good on LB... The result of manual threshold without percentile is 0.61 , the 0.63 score of LB is so attractive that I left everything behind.\n\n------------------------------------------\n\nCongrats to all the winners! And thanks for the great community, let's meet in another competition.\n",
      "votes": 28
    },
    {
      "id": 2162597,
      "postDate": "2023-02-28T10:01:01.740Z",
      "content": "<p>I was looking for the trick, <strong>The only trick: drop false positives</strong> to remove incorrect positive labels, for that I tried <em>Multiple Instance Learning</em>, <em>Triplet Loss</em>, <em>Prototype Net</em>, <em>Saliency Map Inconsistency</em>, but never thought it could be as simple as what you did! 🙂<br>\nThanks for sharing the trick! It is another take away from this competition for me.</p>",
      "rawMarkdown": "I was looking for the trick, **The only trick: drop false positives** to remove incorrect positive labels, for that I tried *Multiple Instance Learning*, *Triplet Loss*, *Prototype Net*, *Saliency Map Inconsistency*, but never thought it could be as simple as what you did! 🙂\nThanks for sharing the trick! It is another take away from this competition for me.",
      "votes": 1,
      "replies": [
        {
          "id": 2162667,
          "postDate": "2023-02-28T10:59:19.933Z",
          "content": "<p>Glad I can help… Some words lists by you are unfamiliar to me, I'll have a look at these new techs, thanks for sharing.</p>",
          "rawMarkdown": "Glad I can help... Some words lists by you are unfamiliar to me, I'll have a look at these new techs, thanks for sharing.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2162373,
      "postDate": "2023-02-28T06:51:51.230Z",
      "content": "<p>Congrats! And thanks for open-sourcing the nvjpeg2k-python!</p>",
      "rawMarkdown": "Congrats! And thanks for open-sourcing the nvjpeg2k-python!",
      "votes": 1
    },
    {
      "id": 2163596,
      "postDate": "2023-02-28T23:51:00.653Z",
      "content": "<p>Thank you for sharing a great solution!</p>\n<p>I don't understand 'drop false positive'. Do you mean you drop false positive in train data, and train another model from scratch with train data without false positive? </p>\n<p>I would appreciate it if you answer the question.</p>",
      "rawMarkdown": "Thank you for sharing a great solution!\n\nI don't understand 'drop false positive'. Do you mean you drop false positive in train data, and train another model from scratch with train data without false positive? \n\nI would appreciate it if you answer the question.",
      "replies": [
        {
          "id": 2163633,
          "postDate": "2023-03-01T00:57:48.120Z",
          "content": "<p>\"Do you mean you drop false negative in train data, and train another model from scratch without train data without false negative\"</p>\n<p>that's correct, train model without the FPs, validation should use all the data.</p>",
          "rawMarkdown": "\"Do you mean you drop false negative in train data, and train another model from scratch without train data without false negative\"\n\nthat's correct, train model without the FPs, validation should use all the data.",
          "votes": 1,
          "replies": [
            {
              "id": 2163714,
              "postDate": "2023-03-01T02:43:17.400Z",
              "content": "<p>Thank you for answering! Would I ask one more question?</p>\n<p>I would like you to explain why you decided to remove false positives in train data.<br>\nBecause I was trying to make the model's prediction of false positive data close to zero, I couldn't think of removing false positive in train data.<br>\nI was wondering what's the effect of removing false positive in train data.</p>",
              "rawMarkdown": "Thank you for answering! Would I ask one more question?\n\nI would like you to explain why you decided to remove false positives in train data.\nBecause I was trying to make the model's prediction of false positive data close to zero, I couldn't think of removing false positive in train data.\nI was wondering what's the effect of removing false positive in train data.\n\n"
            },
            {
              "id": 2163854,
              "postDate": "2023-03-01T05:56:00.483Z",
              "content": "<p>remove is based on a fact: label is tagged on <code>laterality + patient_id</code>, not on every image. If one of the images which has the same <code>laterality + patient_id</code> is positive, then all the image belongs to the same <code>laterality + patient_id</code> are marked as positive. So there must be many FPs in the dataset. </p>",
              "rawMarkdown": "remove is based on a fact: label is tagged on `laterality + patient_id`, not on every image. If one of the images which has the same `laterality + patient_id` is positive, then all the image belongs to the same `laterality + patient_id` are marked as positive. So there must be many FPs in the dataset. ",
              "votes": 2
            },
            {
              "id": 2164206,
              "postDate": "2023-03-01T12:07:51.383Z",
              "content": "<p>Thank you very much!!</p>",
              "rawMarkdown": "Thank you very much!!"
            }
          ]
        }
      ]
    },
    {
      "id": 2162850,
      "postDate": "2023-02-28T12:53:39.033Z",
      "content": "<p>Thanks for sharing, I have one question, how did you not overfit with 33 epochs? My models would overfit near the 6th-7th epoch usually. I find it interesting since you also applied very few augmentations, perhaps is the architecture?</p>",
      "rawMarkdown": "Thanks for sharing, I have one question, how did you not overfit with 33 epochs? My models would overfit near the 6th-7th epoch usually. I find it interesting since you also applied very few augmentations, perhaps is the architecture?",
      "replies": [
        {
          "id": 2162856,
          "postDate": "2023-02-28T12:58:09.830Z",
          "content": "<p>To my understanding he didn't use pretrained weights (as they are not allowed for commercial use).</p>",
          "rawMarkdown": "To my understanding he didn't use pretrained weights (as they are not allowed for commercial use)."
        },
        {
          "id": 2162999,
          "postDate": "2023-02-28T14:16:43.430Z",
          "content": "<p>huh i missed something here.. one epoch for me does not mean the whole dataset, instead, it is <code>number_of_positives * (negtive_ratios + 1)</code>, I normally chose negative_ratios as 4 to 6. So here one epochs means all the positive samples and <strong>some</strong> of the negative samples, <strong>some</strong> = <code>number_of_positives * negative_ratios</code></p>",
          "rawMarkdown": "huh i missed something here.. one epoch for me does not mean the whole dataset, instead, it is `number_of_positives * (negtive_ratios + 1)`, I normally chose negative_ratios as 4 to 6. So here one epochs means all the positive samples and **some** of the negative samples, **some** = `number_of_positives * negative_ratios`",
          "votes": 1
        }
      ]
    },
    {
      "id": 2162757,
      "postDate": "2023-02-28T11:57:45.490Z",
      "content": "<p>Thanks for overview your solution. Did you use original data for training stage or some cropped/preprocessed version?</p>",
      "rawMarkdown": "Thanks for overview your solution. Did you use original data for training stage or some cropped/preprocessed version?",
      "replies": [
        {
          "id": 2162776,
          "postDate": "2023-02-28T12:05:24.937Z",
          "content": "<pre><code>image = utils.crop_roi(image)\nlong_edge = (image.shape[:])\npad_fn = A.PadIfNeeded(long_edge, long_edge, border_mode=cv2.BORDER_CONSTANT, position=  self.training  , value=, always_apply=, p=)\n</code></pre>\n<p>for <code>utils.crop_roi</code>, you can find the source here: <a href=\"https://github.com/louis-she/rsna-2022-public/blob/03e9c542b756cb16b8ee6de2bb79d1b2bd346e1f/utils.py#L24\" target=\"_blank\">https://github.com/louis-she/rsna-2022-public/blob/03e9c542b756cb16b8ee6de2bb79d1b2bd346e1f/utils.py#L24</a></p>",
          "rawMarkdown": "```python\nimage = utils.crop_roi(image)\nlong_edge = max(image.shape[:2])\npad_fn = A.PadIfNeeded(long_edge, long_edge, border_mode=cv2.BORDER_CONSTANT, position=\"random\" if self.training else \"center\", value=0, always_apply=True, p=1.0)\n```\n\nfor `utils.crop_roi`, you can find the source here: https://github.com/louis-she/rsna-2022-public/blob/03e9c542b756cb16b8ee6de2bb79d1b2bd346e1f/utils.py#L24",
          "votes": 2
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2162597,
      "author_name": "Rasoul Mojtahedzadeh",
      "author_url": "",
      "post_date": "2023-02-28T10:01:01.740000",
      "content": "<p>I was looking for the trick, <strong>The only trick: drop false positives</strong> to remove incorrect positive labels, for that I tried <em>Multiple Instance Learning</em>, <em>Triplet Loss</em>, <em>Prototype Net</em>, <em>Saliency Map Inconsistency</em>, but never thought it could be as simple as what you did! 🙂<br>\nThanks for sharing the trick! It is another take away from this competition for me.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2162667,
          "author_name": "Chenglu",
          "author_url": "",
          "post_date": "2023-02-28T10:59:19.933000",
          "content": "<p>Glad I can help… Some words lists by you are unfamiliar to me, I'll have a look at these new techs, thanks for sharing.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2162373,
      "author_name": "Antti Isosalo",
      "author_url": "",
      "post_date": "2023-02-28T06:51:51.230000",
      "content": "<p>Congrats! And thanks for open-sourcing the nvjpeg2k-python!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2163596,
      "author_name": "Ju7on9",
      "author_url": "",
      "post_date": "2023-02-28T23:51:00.653000",
      "content": "<p>Thank you for sharing a great solution!</p>\n<p>I don't understand 'drop false positive'. Do you mean you drop false positive in train data, and train another model from scratch with train data without false positive? </p>\n<p>I would appreciate it if you answer the question.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2163633,
          "author_name": "Chenglu",
          "author_url": "",
          "post_date": "2023-03-01T00:57:48.120000",
          "content": "<p>\"Do you mean you drop false negative in train data, and train another model from scratch without train data without false negative\"</p>\n<p>that's correct, train model without the FPs, validation should use all the data.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2163714,
              "author_name": "Ju7on9",
              "author_url": "",
              "post_date": "2023-03-01T02:43:17.400000",
              "content": "<p>Thank you for answering! Would I ask one more question?</p>\n<p>I would like you to explain why you decided to remove false positives in train data.<br>\nBecause I was trying to make the model's prediction of false positive data close to zero, I couldn't think of removing false positive in train data.<br>\nI was wondering what's the effect of removing false positive in train data.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2163854,
              "author_name": "Chenglu",
              "author_url": "",
              "post_date": "2023-03-01T05:56:00.483000",
              "content": "<p>remove is based on a fact: label is tagged on <code>laterality + patient_id</code>, not on every image. If one of the images which has the same <code>laterality + patient_id</code> is positive, then all the image belongs to the same <code>laterality + patient_id</code> are marked as positive. So there must be many FPs in the dataset. </p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2164206,
              "author_name": "Ju7on9",
              "author_url": "",
              "post_date": "2023-03-01T12:07:51.383000",
              "content": "<p>Thank you very much!!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2162850,
      "author_name": "moth",
      "author_url": "",
      "post_date": "2023-02-28T12:53:39.033000",
      "content": "<p>Thanks for sharing, I have one question, how did you not overfit with 33 epochs? My models would overfit near the 6th-7th epoch usually. I find it interesting since you also applied very few augmentations, perhaps is the architecture?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2162856,
          "author_name": "Rasoul Mojtahedzadeh",
          "author_url": "",
          "post_date": "2023-02-28T12:58:09.830000",
          "content": "<p>To my understanding he didn't use pretrained weights (as they are not allowed for commercial use).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2162999,
          "author_name": "Chenglu",
          "author_url": "",
          "post_date": "2023-02-28T14:16:43.430000",
          "content": "<p>huh i missed something here.. one epoch for me does not mean the whole dataset, instead, it is <code>number_of_positives * (negtive_ratios + 1)</code>, I normally chose negative_ratios as 4 to 6. So here one epochs means all the positive samples and <strong>some</strong> of the negative samples, <strong>some</strong> = <code>number_of_positives * negative_ratios</code></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2162757,
      "author_name": "Sergey Bryansky",
      "author_url": "",
      "post_date": "2023-02-28T11:57:45.490000",
      "content": "<p>Thanks for overview your solution. Did you use original data for training stage or some cropped/preprocessed version?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2162776,
          "author_name": "Chenglu",
          "author_url": "",
          "post_date": "2023-02-28T12:05:24.937000",
          "content": "<pre><code>image = utils.crop_roi(image)\nlong_edge = (image.shape[:])\npad_fn = A.PadIfNeeded(long_edge, long_edge, border_mode=cv2.BORDER_CONSTANT, position=  self.training  , value=, always_apply=, p=)\n</code></pre>\n<p>for <code>utils.crop_roi</code>, you can find the source here: <a href=\"https://github.com/louis-she/rsna-2022-public/blob/03e9c542b756cb16b8ee6de2bb79d1b2bd346e1f/utils.py#L24\" target=\"_blank\">https://github.com/louis-she/rsna-2022-public/blob/03e9c542b756cb16b8ee6de2bb79d1b2bd346e1f/utils.py#L24</a></p>",
          "votes": 2,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2162323": "Training source code: https://github.com/louis-she/rsna-2022-public\nInfernece notebook: https://www.kaggle.com/code/snaker/rsna-infer-t4-threshold?scriptVersionId=120550686\n\nThank everyone who contributed to the competition and thanks Kaggle for hosting.\n\nMy solution didn't involve complicated pipeline, just as most does, single image classification. Here are some settings you may be interested in:\n\n* **Augmentation** RandomRotate90, HorizontalFlip, VerticalFlip\n* **Models and Resolutions** \n  1. ConvnextV2_nano @ 1536\n  2. ConvnextV2_nano_another_seed @ 1536\n  3. ConvnextV2_nano @ 2048\n  4. EfficientnetV2_s @ 1536\n* **Optimizer and Scheduler** Adam with 3e-5 + OneCycleLR\n* **Epochs** 33\n* **Agg** by maximum\n* Aux Losses: no\n* Tablet Features: no\n* SWA: no\n* External Dataset: no\n\n## The only trick: drop false positives\n\nI realized that all the label tagged for `laterality + patient_id` is the same. That means the label is actually for `patient + laterality`, not image. So the easiest way to deal with this, is to train OOF and drop False Negatives with high confidence( e.g. the label is 1 but model predict probability is like 0.1 or 0.01). The CV score boost huge after this is done( about 0.03 to 0.04 ).\n\n## Things I have learned\n\nThe performance matters a lot in the competition. I'm happy that I take almost one month to optimize the inference pipeline, and I really learned something, like open-sourced the [nvjpeg2k-python](https://github.com/louis-she/nvjpeg2k-python) ( first time writing some cuda code, it feels good ). To utilze the maximum of the hardware, I use the 2 x T4 kernel and some multiprocessing tricks. You can see that in the inference notebook.\n\nI use the percentile threshold trick I mentioned in the thread, cause it boost my LB score about ~0.015 . I know that in this way I will have a very high chance to overfit LB but I still chose it as one of my last selection sub, because it is just too good on LB... The result of manual threshold without percentile is 0.61 , the 0.63 score of LB is so attractive that I left everything behind.\n\n------------------------------------------\n\nCongrats to all the winners! And thanks for the great community, let's meet in another competition.\n",
    "2162597": "I was looking for the trick, **The only trick: drop false positives** to remove incorrect positive labels, for that I tried *Multiple Instance Learning*, *Triplet Loss*, *Prototype Net*, *Saliency Map Inconsistency*, but never thought it could be as simple as what you did! 🙂\nThanks for sharing the trick! It is another take away from this competition for me.",
    "2162373": "Congrats! And thanks for open-sourcing the nvjpeg2k-python!",
    "2163596": "Thank you for sharing a great solution!\n\nI don't understand 'drop false positive'. Do you mean you drop false positive in train data, and train another model from scratch with train data without false positive? \n\nI would appreciate it if you answer the question.",
    "2162850": "Thanks for sharing, I have one question, how did you not overfit with 33 epochs? My models would overfit near the 6th-7th epoch usually. I find it interesting since you also applied very few augmentations, perhaps is the architecture?",
    "2162757": "Thanks for overview your solution. Did you use original data for training stage or some cropped/preprocessed version?"
  }
}