{
  "id": 372797,
  "title": "Question for the experts:  Would a Yolov5 trained on DDSM GMIC like solution work for this comp?",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/372797",
  "author_name": "@kaggleqrdl",
  "post_date": "2022-12-18T05:48:23.051000",
  "votes": 3,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I'll be the first to admit I'm a novice with CV.   However, after playing around with <a href=\"https://github.com/nyukat/GMIC\" target=\"_blank\">GMIC</a> and looking at some of the yolov5 notebooks, it seems like we could do something similar but perhaps better.</p>\n<p>Basically train yolov5 on DDSM (something like this: <a href=\"https://www.kaggle.com/code/sentrankim/yolov5-ddsm/notebook\" target=\"_blank\">https://www.kaggle.com/code/sentrankim/yolov5-ddsm/notebook</a> ?)  and then train on the extracted ROI patches.  </p>\n<p>The one issue I see in this approach is that some extracted patches will be from breasts which are labeled as cancer, but not all of the extracted patches from that breast will contain malignant tissue and the training will be a bit confused about that.   However, afaict, it already has that problem, just at a greater scale without ROI.   Perhaps this could be a fine tuning issue?</p>\n<p>Another issue is standardizing the ROI patch size in the same way GMIC does.  Resizing wouldn't necessarily be very ideal.    Perhaps you could have seperate models for difference sizes, like micro, small, medium,large, etc,  and then grow the bounding box of the ROI patch until it matches the smallest sized model.</p>\n<p>edit:</p>\n<p>For reference, here is the GMIC model arch<br>\n<a href=\"https://arxiv.org/pdf/2002.07613.pdf\" target=\"_blank\">https://arxiv.org/pdf/2002.07613.pdf</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9052057%2Fdc50fb7c0eb10e0568b17df1160d6032%2FScreenshot%202022-12-18%201.36.01%20AM.png?generation=1671356183767261&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 2068574,
      "postDate": "2022-12-18T05:48:23.053Z",
      "content": "<p>I'll be the first to admit I'm a novice with CV.   However, after playing around with <a href=\"https://github.com/nyukat/GMIC\" target=\"_blank\">GMIC</a> and looking at some of the yolov5 notebooks, it seems like we could do something similar but perhaps better.</p>\n<p>Basically train yolov5 on DDSM (something like this: <a href=\"https://www.kaggle.com/code/sentrankim/yolov5-ddsm/notebook\" target=\"_blank\">https://www.kaggle.com/code/sentrankim/yolov5-ddsm/notebook</a> ?)  and then train on the extracted ROI patches.  </p>\n<p>The one issue I see in this approach is that some extracted patches will be from breasts which are labeled as cancer, but not all of the extracted patches from that breast will contain malignant tissue and the training will be a bit confused about that.   However, afaict, it already has that problem, just at a greater scale without ROI.   Perhaps this could be a fine tuning issue?</p>\n<p>Another issue is standardizing the ROI patch size in the same way GMIC does.  Resizing wouldn't necessarily be very ideal.    Perhaps you could have seperate models for difference sizes, like micro, small, medium,large, etc,  and then grow the bounding box of the ROI patch until it matches the smallest sized model.</p>\n<p>edit:</p>\n<p>For reference, here is the GMIC model arch<br>\n<a href=\"https://arxiv.org/pdf/2002.07613.pdf\" target=\"_blank\">https://arxiv.org/pdf/2002.07613.pdf</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9052057%2Fdc50fb7c0eb10e0568b17df1160d6032%2FScreenshot%202022-12-18%201.36.01%20AM.png?generation=1671356183767261&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I'll be the first to admit I'm a novice with CV.   However, after playing around with [GMIC](https://github.com/nyukat/GMIC) and looking at some of the yolov5 notebooks, it seems like we could do something similar but perhaps better.\n\nBasically train yolov5 on DDSM (something like this: https://www.kaggle.com/code/sentrankim/yolov5-ddsm/notebook ?)  and then train on the extracted ROI patches.  \n\nThe one issue I see in this approach is that some extracted patches will be from breasts which are labeled as cancer, but not all of the extracted patches from that breast will contain malignant tissue and the training will be a bit confused about that.   However, afaict, it already has that problem, just at a greater scale without ROI.   Perhaps this could be a fine tuning issue?\n\nAnother issue is standardizing the ROI patch size in the same way GMIC does.  Resizing wouldn't necessarily be very ideal.    Perhaps you could have seperate models for difference sizes, like micro, small, medium,large, etc,  and then grow the bounding box of the ROI patch until it matches the smallest sized model.\n\nedit:\n\nFor reference, here is the GMIC model arch\nhttps://arxiv.org/pdf/2002.07613.pdf\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9052057%2Fdc50fb7c0eb10e0568b17df1160d6032%2FScreenshot%202022-12-18%201.36.01%20AM.png?generation=1671356183767261&alt=media)\n\n\n\n",
      "votes": 3
    },
    {
      "id": 2069022,
      "postDate": "2022-12-18T14:03:31.277Z",
      "content": "<p>Keep in mind, the DDSM dataset is older digitized film-screen based mammography. The competition dataset is FFDM, or Full Field Digital Mammography. Even the best film based images aren't nearly as clear and detailed as digital ones. There is quite a bit of lost data during film processing and during the digitization process.</p>",
      "rawMarkdown": "Keep in mind, the DDSM dataset is older digitized film-screen based mammography. The competition dataset is FFDM, or Full Field Digital Mammography. Even the best film based images aren't nearly as clear and detailed as digital ones. There is quite a bit of lost data during film processing and during the digitization process.",
      "votes": 2,
      "replies": [
        {
          "id": 2069412,
          "postDate": "2022-12-18T23:56:21.930Z",
          "content": "<p>Yes, this is true, however it has been annotated by trained radiologists.   We want to leverage that insight.</p>",
          "rawMarkdown": "Yes, this is true, however it has been annotated by trained radiologists.   We want to leverage that insight."
        }
      ]
    },
    {
      "id": 2070225,
      "postDate": "2022-12-19T18:01:52.717Z",
      "content": "<p><a href=\"https://physionet.org/content/vindr-mammo/1.0.0/\" target=\"_blank\">https://physionet.org/content/vindr-mammo/1.0.0/</a></p>\n<p><a href=\"https://www.kaggle.com/vbookshelf\" target=\"_blank\">@vbookshelf</a> also created a dataset for it 19 days ago, but if he mentioned it here, I didn't see it.  </p>\n<p>The below-the-radar approach is quite intriguing.</p>\n<p><a href=\"https://www.kaggle.com/datasets/vbookshelf/mammogram-mass-analyzer-v00\" target=\"_blank\">https://www.kaggle.com/datasets/vbookshelf/mammogram-mass-analyzer-v00</a></p>\n<p>The yolo models are a little hit or miss.  But as a proof of concept it seems to work</p>\n<p>poc here - <a href=\"https://www.kaggle.com/code/kaggleqrdl/yolov5-roi-patch-poc\" target=\"_blank\">https://www.kaggle.com/code/kaggleqrdl/yolov5-roi-patch-poc</a></p>\n<p>Looks like it'll work.  yolov5 models likely need more training, some boundary box managment will be needed.</p>",
      "rawMarkdown": "https://physionet.org/content/vindr-mammo/1.0.0/\n\n@vbookshelf also created a dataset for it 19 days ago, but if he mentioned it here, I didn't see it.  \n\nThe below-the-radar approach is quite intriguing.\n\nhttps://www.kaggle.com/datasets/vbookshelf/mammogram-mass-analyzer-v00\n\nThe yolo models are a little hit or miss.  But as a proof of concept it seems to work\n\npoc here - https://www.kaggle.com/code/kaggleqrdl/yolov5-roi-patch-poc\n\nLooks like it'll work.  yolov5 models likely need more training, some boundary box managment will be needed.\n"
    },
    {
      "id": 2069853,
      "postDate": "2022-12-19T11:39:09.067Z",
      "content": "<p>Well, yolo is huge in breast cancer research.  Shame nobody seems to want to share their models.</p>\n<p><a href=\"https://scholar.google.com/scholar?as_ylo=2022&amp;q=breast+cancer+yolo&amp;hl=en&amp;as_sdt=0,5\" target=\"_blank\">https://scholar.google.com/scholar?as_ylo=2022&amp;q=breast+cancer+yolo&amp;hl=en&amp;as_sdt=0,5</a></p>",
      "rawMarkdown": "Well, yolo is huge in breast cancer research.  Shame nobody seems to want to share their models.\n\nhttps://scholar.google.com/scholar?as_ylo=2022&q=breast+cancer+yolo&hl=en&as_sdt=0,5"
    },
    {
      "id": 2068687,
      "postDate": "2022-12-18T08:36:53.140Z",
      "content": "<p>No, if yolov5 is used as a single solution<br>\nYes, if yolov5 is part of the solution (e.g. as post verifier, as pre roi extraction, as ensemble, as feature extraction, etc) </p>",
      "rawMarkdown": "No, if yolov5 is used as a single solution\nYes, if yolov5 is part of the solution (e.g. as post verifier, as pre roi extraction, as ensemble, as feature extraction, etc) ",
      "replies": [
        {
          "id": 2068741,
          "postDate": "2022-12-18T09:22:40.677Z",
          "content": "<p>\"No, if yolov5 is used as a single solution\"  …  <a href=\"https://www.kaggle.com/henk\" target=\"_blank\">@henk</a>, honest question -  what did I write to make you think that was an option? :)</p>",
          "rawMarkdown": "\"No, if yolov5 is used as a single solution\"  ...  @henk, honest question -  what did I write to make you think that was an option? :)",
          "replies": [
            {
              "id": 2068745,
              "postDate": "2022-12-18T09:24:30.340Z",
              "content": "<p>novice with CV</p>",
              "rawMarkdown": "novice with CV",
              "votes": 1
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2069022,
      "author_name": "David Roberts",
      "author_url": "",
      "post_date": "2022-12-18T14:03:31.277000",
      "content": "<p>Keep in mind, the DDSM dataset is older digitized film-screen based mammography. The competition dataset is FFDM, or Full Field Digital Mammography. Even the best film based images aren't nearly as clear and detailed as digital ones. There is quite a bit of lost data during film processing and during the digitization process.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2069412,
          "author_name": "@kaggleqrdl",
          "author_url": "",
          "post_date": "2022-12-18T23:56:21.930000",
          "content": "<p>Yes, this is true, however it has been annotated by trained radiologists.   We want to leverage that insight.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2070225,
      "author_name": "@kaggleqrdl",
      "author_url": "",
      "post_date": "2022-12-19T18:01:52.717000",
      "content": "<p><a href=\"https://physionet.org/content/vindr-mammo/1.0.0/\" target=\"_blank\">https://physionet.org/content/vindr-mammo/1.0.0/</a></p>\n<p><a href=\"https://www.kaggle.com/vbookshelf\" target=\"_blank\">@vbookshelf</a> also created a dataset for it 19 days ago, but if he mentioned it here, I didn't see it.  </p>\n<p>The below-the-radar approach is quite intriguing.</p>\n<p><a href=\"https://www.kaggle.com/datasets/vbookshelf/mammogram-mass-analyzer-v00\" target=\"_blank\">https://www.kaggle.com/datasets/vbookshelf/mammogram-mass-analyzer-v00</a></p>\n<p>The yolo models are a little hit or miss.  But as a proof of concept it seems to work</p>\n<p>poc here - <a href=\"https://www.kaggle.com/code/kaggleqrdl/yolov5-roi-patch-poc\" target=\"_blank\">https://www.kaggle.com/code/kaggleqrdl/yolov5-roi-patch-poc</a></p>\n<p>Looks like it'll work.  yolov5 models likely need more training, some boundary box managment will be needed.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2069853,
      "author_name": "@kaggleqrdl",
      "author_url": "",
      "post_date": "2022-12-19T11:39:09.067000",
      "content": "<p>Well, yolo is huge in breast cancer research.  Shame nobody seems to want to share their models.</p>\n<p><a href=\"https://scholar.google.com/scholar?as_ylo=2022&amp;q=breast+cancer+yolo&amp;hl=en&amp;as_sdt=0,5\" target=\"_blank\">https://scholar.google.com/scholar?as_ylo=2022&amp;q=breast+cancer+yolo&amp;hl=en&amp;as_sdt=0,5</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2068687,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-12-18T08:36:53.140000",
      "content": "<p>No, if yolov5 is used as a single solution<br>\nYes, if yolov5 is part of the solution (e.g. as post verifier, as pre roi extraction, as ensemble, as feature extraction, etc) </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2068741,
          "author_name": "@kaggleqrdl",
          "author_url": "",
          "post_date": "2022-12-18T09:22:40.677000",
          "content": "<p>\"No, if yolov5 is used as a single solution\"  …  <a href=\"https://www.kaggle.com/henk\" target=\"_blank\">@henk</a>, honest question -  what did I write to make you think that was an option? :)</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2068745,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2022-12-18T09:24:30.340000",
              "content": "<p>novice with CV</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2068574": "I'll be the first to admit I'm a novice with CV.   However, after playing around with [GMIC](https://github.com/nyukat/GMIC) and looking at some of the yolov5 notebooks, it seems like we could do something similar but perhaps better.\n\nBasically train yolov5 on DDSM (something like this: https://www.kaggle.com/code/sentrankim/yolov5-ddsm/notebook ?)  and then train on the extracted ROI patches.  \n\nThe one issue I see in this approach is that some extracted patches will be from breasts which are labeled as cancer, but not all of the extracted patches from that breast will contain malignant tissue and the training will be a bit confused about that.   However, afaict, it already has that problem, just at a greater scale without ROI.   Perhaps this could be a fine tuning issue?\n\nAnother issue is standardizing the ROI patch size in the same way GMIC does.  Resizing wouldn't necessarily be very ideal.    Perhaps you could have seperate models for difference sizes, like micro, small, medium,large, etc,  and then grow the bounding box of the ROI patch until it matches the smallest sized model.\n\nedit:\n\nFor reference, here is the GMIC model arch\nhttps://arxiv.org/pdf/2002.07613.pdf\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9052057%2Fdc50fb7c0eb10e0568b17df1160d6032%2FScreenshot%202022-12-18%201.36.01%20AM.png?generation=1671356183767261&alt=media)\n\n\n\n",
    "2069022": "Keep in mind, the DDSM dataset is older digitized film-screen based mammography. The competition dataset is FFDM, or Full Field Digital Mammography. Even the best film based images aren't nearly as clear and detailed as digital ones. There is quite a bit of lost data during film processing and during the digitization process.",
    "2070225": "https://physionet.org/content/vindr-mammo/1.0.0/\n\n@vbookshelf also created a dataset for it 19 days ago, but if he mentioned it here, I didn't see it.  \n\nThe below-the-radar approach is quite intriguing.\n\nhttps://www.kaggle.com/datasets/vbookshelf/mammogram-mass-analyzer-v00\n\nThe yolo models are a little hit or miss.  But as a proof of concept it seems to work\n\npoc here - https://www.kaggle.com/code/kaggleqrdl/yolov5-roi-patch-poc\n\nLooks like it'll work.  yolov5 models likely need more training, some boundary box managment will be needed.\n",
    "2069853": "Well, yolo is huge in breast cancer research.  Shame nobody seems to want to share their models.\n\nhttps://scholar.google.com/scholar?as_ylo=2022&q=breast+cancer+yolo&hl=en&as_sdt=0,5",
    "2068687": "No, if yolov5 is used as a single solution\nYes, if yolov5 is part of the solution (e.g. as post verifier, as pre roi extraction, as ensemble, as feature extraction, etc) "
  }
}