{
  "id": 373052,
  "title": "Venting on breast cancer detection research",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/373052",
  "author_name": "@kaggleqrdl",
  "post_date": "2022-12-19T11:28:27.768000",
  "votes": 7,
  "comment_count": 9,
  "views": 0,
  "content": "<p>After reading the umpteenth paper about some model or collection of models (and a lot of yolo models) that can detect breast cancer - but with no available models for download and repro, I really have to vent.</p>\n<p>This seems like a deeply tragic state of affairs.  So much work and training is being done, which when thoughtfully ensembled together, might actually do some amazing things for a field which could really benefit from it.   Through ensembling, we'd be able to quickly see which research is doing something original and fresh versus which research is just correlating highly with everything else.</p>\n<p>There really needs to be a movement in ML academia that if you're going to publish research, you need to at the very least push your pytorch models to a collective zoo.   <em>Especially</em> any research which is performed under government grants.</p>",
  "messages": [
    {
      "id": 2069839,
      "postDate": "2022-12-19T11:28:27.770Z",
      "content": "<p>After reading the umpteenth paper about some model or collection of models (and a lot of yolo models) that can detect breast cancer - but with no available models for download and repro, I really have to vent.</p>\n<p>This seems like a deeply tragic state of affairs.  So much work and training is being done, which when thoughtfully ensembled together, might actually do some amazing things for a field which could really benefit from it.   Through ensembling, we'd be able to quickly see which research is doing something original and fresh versus which research is just correlating highly with everything else.</p>\n<p>There really needs to be a movement in ML academia that if you're going to publish research, you need to at the very least push your pytorch models to a collective zoo.   <em>Especially</em> any research which is performed under government grants.</p>",
      "rawMarkdown": "After reading the umpteenth paper about some model or collection of models (and a lot of yolo models) that can detect breast cancer - but with no available models for download and repro, I really have to vent.\n\nThis seems like a deeply tragic state of affairs.  So much work and training is being done, which when thoughtfully ensembled together, might actually do some amazing things for a field which could really benefit from it.   Through ensembling, we'd be able to quickly see which research is doing something original and fresh versus which research is just correlating highly with everything else.\n\nThere really needs to be a movement in ML academia that if you're going to publish research, you need to at the very least push your pytorch models to a collective zoo.   *Especially* any research which is performed under government grants.\n",
      "votes": 6
    },
    {
      "id": 2069975,
      "postDate": "2022-12-19T13:45:16.640Z",
      "content": "<p>I think this is common in the medical field. Especially in the US where medical data is hard to come by. Commercial entities who pay for data aren't going to open-source their software. The challenge with radiologic studies is that they're expensive to produce and require specific licenses to order and perform. Add the legal issues of PHI (Patient Health Information), and it's nearly impossible to find anything open-source.</p>",
      "rawMarkdown": "I think this is common in the medical field. Especially in the US where medical data is hard to come by. Commercial entities who pay for data aren't going to open-source their software. The challenge with radiologic studies is that they're expensive to produce and require specific licenses to order and perform. Add the legal issues of PHI (Patient Health Information), and it's nearly impossible to find anything open-source.",
      "votes": 1,
      "replies": [
        {
          "id": 2070065,
          "postDate": "2022-12-19T15:13:34.103Z",
          "content": "<p>The tragic thing is that individually, it's unlikely they can beat SOTA numbers, but ensembling a lot of these models (there are so many), I think it might be possible.  And yolo is quite light and fast and its standardized format means it can be re-used quite easily.  Also, pytorch models shouldn't have any phi issues around them.</p>",
          "rawMarkdown": "The tragic thing is that individually, it's unlikely they can beat SOTA numbers, but ensembling a lot of these models (there are so many), I think it might be possible.  And yolo is quite light and fast and its standardized format means it can be re-used quite easily.  Also, pytorch models shouldn't have any phi issues around them.",
          "replies": [
            {
              "id": 2070082,
              "postDate": "2022-12-19T15:27:29.987Z",
              "content": "<p>Right, the models themselves wouldn't have any PHI, but the legal issues of PHI make it hard (expensive) to find any data to train on. No commercial entity is going to pay for data, then give away models. Likewise, nobody is going to share medical data freely. </p>\n<p>The challenge with medical imaging research is that colleges don't normally have SOTA imaging equipment (a decent mammography machine costs almost a million USD these days), so their work doesn't translate well to real-world data from modern machinery.</p>\n<p>I would love to see an open source 'data-as-a-service' model for medical image AI.</p>",
              "rawMarkdown": "Right, the models themselves wouldn't have any PHI, but the legal issues of PHI make it hard (expensive) to find any data to train on. No commercial entity is going to pay for data, then give away models. Likewise, nobody is going to share medical data freely. \n\nThe challenge with medical imaging research is that colleges don't normally have SOTA imaging equipment (a decent mammography machine costs almost a million USD these days), so their work doesn't translate well to real-world data from modern machinery.\n\nI would love to see an open source 'data-as-a-service' model for medical image AI."
            },
            {
              "id": 2070178,
              "postDate": "2022-12-19T17:17:16.743Z",
              "content": "<p>Makes sense.  <a href=\"https://www.kaggle.com/davidbroberts\" target=\"_blank\">@davidbroberts</a> what are your thoughts on CCMD? </p>\n<p><a href=\"https://wiki.cancerimagingarchive.net/pages/viewpage.action?pageId=70230508\" target=\"_blank\">https://wiki.cancerimagingarchive.net/pages/viewpage.action?pageId=70230508</a></p>\n<p>It was produced with this I believe:</p>\n<p>It looks to be digital  <a href=\"https://www.htig.com/catalog/mammography-equipment/ge-senographe-ds/\" target=\"_blank\">https://www.htig.com/catalog/mammography-equipment/ge-senographe-ds/</a></p>\n<p>Here's your data as a service, I think <a href=\"https://portal.imaging.datacommons.cancer.gov/explore/filters/?collection_id=cmmd\" target=\"_blank\">https://portal.imaging.datacommons.cancer.gov/explore/filters/?collection_id=cmmd</a></p>\n<p>Also this, </p>\n<p><a href=\"https://physionet.org/content/vindr-mammo/1.0.0/\" target=\"_blank\">https://physionet.org/content/vindr-mammo/1.0.0/</a></p>\n<p><a href=\"https://www.kaggle.com/vbookshelf\" target=\"_blank\">@vbookshelf</a> also created a dataset for it 19 days ago, but if he mentioned it here, I didn't see it.</p>\n<p>The below-the-radar approach is quite intriguing.</p>\n<p><a href=\"https://www.kaggle.com/datasets/vbookshelf/mammogram-mass-analyzer-v00\" target=\"_blank\">https://www.kaggle.com/datasets/vbookshelf/mammogram-mass-analyzer-v00</a></p>",
              "rawMarkdown": "Makes sense.  @davidbroberts what are your thoughts on CCMD? \n\nhttps://wiki.cancerimagingarchive.net/pages/viewpage.action?pageId=70230508\n\nIt was produced with this I believe:\n\n It looks to be digital  https://www.htig.com/catalog/mammography-equipment/ge-senographe-ds/\n\nHere's your data as a service, I think https://portal.imaging.datacommons.cancer.gov/explore/filters/?collection_id=cmmd\n\nAlso this, \n\nhttps://physionet.org/content/vindr-mammo/1.0.0/\n\n@vbookshelf also created a dataset for it 19 days ago, but if he mentioned it here, I didn't see it.\n\nThe below-the-radar approach is quite intriguing.\n\nhttps://www.kaggle.com/datasets/vbookshelf/mammogram-mass-analyzer-v00\n"
            },
            {
              "id": 2070517,
              "postDate": "2022-12-20T04:58:41.110Z",
              "content": "<p>I haven't seen the CCMD dataset, but I have worked with the Vindr mammo set. Vindr is very nice quality FFDM, but the labels don't include benign/malignant if I remember correctly. It may be good for pretraining or pseudo-labeling I guess.</p>\n<p>Interesting link to the datacommons site. Thanks!</p>",
              "rawMarkdown": "I haven't seen the CCMD dataset, but I have worked with the Vindr mammo set. Vindr is very nice quality FFDM, but the labels don't include benign/malignant if I remember correctly. It may be good for pretraining or pseudo-labeling I guess.\n\nInteresting link to the datacommons site. Thanks!"
            },
            {
              "id": 2070547,
              "postDate": "2022-12-20T06:17:12.860Z",
              "content": "<p>Well, it has the birads level 4 and 5, which I imagine isn't that far off from what 'cancer' means (?) in this dataset.   It seems a bit ambigous, tbh.  </p>\n<p>But that's not really the point, I think, but rather the goal is ROI patch extraction.   Unless the theory is that all tumours are somehow breast wide and this can be detected by the networks (is that a theory?), our networks are going to be learning a lot of strange things unless they focus in on the right spots.</p>\n<p>If you survey the literature, a great deal of it takes this two stage approach.  ROI patch -&gt; resnet/efnet to tell if the patch contains malignant tissue.   </p>\n<p>This will probably still be problematic as some patches won't have malignant tissue even though they  are extracted from breasts which are diagnosed with cancer.   But, this can be mitigated somewhat by leveraging the Vindr DS and the birads ratings.   Note that  BI-RADS 5 is a category of breast lesions that have at least a 95% probability of malignancy</p>",
              "rawMarkdown": "Well, it has the birads level 4 and 5, which I imagine isn't that far off from what 'cancer' means (?) in this dataset.   It seems a bit ambigous, tbh.  \n\nBut that's not really the point, I think, but rather the goal is ROI patch extraction.   Unless the theory is that all tumours are somehow breast wide and this can be detected by the networks (is that a theory?), our networks are going to be learning a lot of strange things unless they focus in on the right spots.\n\nIf you survey the literature, a great deal of it takes this two stage approach.  ROI patch -> resnet/efnet to tell if the patch contains malignant tissue.   \n\nThis will probably still be problematic as some patches won't have malignant tissue even though they  are extracted from breasts which are diagnosed with cancer.   But, this can be mitigated somewhat by leveraging the Vindr DS and the birads ratings.   Note that  BI-RADS 5 is a category of breast lesions that have at least a 95% probability of malignancy"
            }
          ]
        }
      ]
    },
    {
      "id": 2069952,
      "postDate": "2022-12-19T13:01:21.073Z",
      "content": "<p>Ok, well, I decided to try something.  I emailed 18 different researchers who contributed to a paper in 2022 on using yolo models for detecting breast cancer (there are a lot more) to see if they would share their yolo models for an opensource effort in this area.  I will create a public dataset on Kaggle with any responses.  I am hopeful, but not particularly optimistic.</p>\n<p><a href=\"https://scholar.google.com/scholar?as_ylo=2022&amp;q=breast+cancer+yolo&amp;hl=en&amp;as_sdt=0,5\" target=\"_blank\">https://scholar.google.com/scholar?as_ylo=2022&amp;q=breast+cancer+yolo&amp;hl=en&amp;as_sdt=0,5</a></p>\n<p>888 results, just for 2022.  </p>",
      "rawMarkdown": "Ok, well, I decided to try something.  I emailed 18 different researchers who contributed to a paper in 2022 on using yolo models for detecting breast cancer (there are a lot more) to see if they would share their yolo models for an opensource effort in this area.  I will create a public dataset on Kaggle with any responses.  I am hopeful, but not particularly optimistic.\n\nhttps://scholar.google.com/scholar?as_ylo=2022&q=breast+cancer+yolo&hl=en&as_sdt=0,5\n\n888 results, just for 2022.  ",
      "votes": 2,
      "replies": [
        {
          "id": 2069971,
          "postDate": "2022-12-19T13:29:13.947Z",
          "content": "<p>As to yolo … you can use it. Problem is with annotation I think. I am (and probably most of participants) correct annotate data. Certainly you can use external annotated dataset but I am not sure if such exist or we can use it in this competition.</p>\n<p>I thnink that classifier is not a final solution (I want to see a. good classifier trained on such imbalanced dataset b. I think that there is too little positive cases to generalize well). It can be used as a part of solution. I know that currently it looks like this is the best option but soon you will see different solution (I am sure). </p>",
          "rawMarkdown": "As to yolo ... you can use it. Problem is with annotation I think. I am (and probably most of participants) correct annotate data. Certainly you can use external annotated dataset but I am not sure if such exist or we can use it in this competition.\n\nI thnink that classifier is not a final solution (I want to see a. good classifier trained on such imbalanced dataset b. I think that there is too little positive cases to generalize well). It can be used as a part of solution. I know that currently it looks like this is the best option but soon you will see different solution (I am sure). ",
          "votes": 1,
          "replies": [
            {
              "id": 2070066,
              "postDate": "2022-12-19T15:14:45.537Z",
              "content": "<p>Yeah, I think Yolo is best for finding ROI patches which are then sent through closer looking models.  That seems to be the general solution to most of the papers that I read.</p>",
              "rawMarkdown": "Yeah, I think Yolo is best for finding ROI patches which are then sent through closer looking models.  That seems to be the general solution to most of the papers that I read."
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2069975,
      "author_name": "David Roberts",
      "author_url": "",
      "post_date": "2022-12-19T13:45:16.640000",
      "content": "<p>I think this is common in the medical field. Especially in the US where medical data is hard to come by. Commercial entities who pay for data aren't going to open-source their software. The challenge with radiologic studies is that they're expensive to produce and require specific licenses to order and perform. Add the legal issues of PHI (Patient Health Information), and it's nearly impossible to find anything open-source.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2070065,
          "author_name": "@kaggleqrdl",
          "author_url": "",
          "post_date": "2022-12-19T15:13:34.103000",
          "content": "<p>The tragic thing is that individually, it's unlikely they can beat SOTA numbers, but ensembling a lot of these models (there are so many), I think it might be possible.  And yolo is quite light and fast and its standardized format means it can be re-used quite easily.  Also, pytorch models shouldn't have any phi issues around them.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2070082,
              "author_name": "David Roberts",
              "author_url": "",
              "post_date": "2022-12-19T15:27:29.987000",
              "content": "<p>Right, the models themselves wouldn't have any PHI, but the legal issues of PHI make it hard (expensive) to find any data to train on. No commercial entity is going to pay for data, then give away models. Likewise, nobody is going to share medical data freely. </p>\n<p>The challenge with medical imaging research is that colleges don't normally have SOTA imaging equipment (a decent mammography machine costs almost a million USD these days), so their work doesn't translate well to real-world data from modern machinery.</p>\n<p>I would love to see an open source 'data-as-a-service' model for medical image AI.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2070178,
              "author_name": "@kaggleqrdl",
              "author_url": "",
              "post_date": "2022-12-19T17:17:16.743000",
              "content": "<p>Makes sense.  <a href=\"https://www.kaggle.com/davidbroberts\" target=\"_blank\">@davidbroberts</a> what are your thoughts on CCMD? </p>\n<p><a href=\"https://wiki.cancerimagingarchive.net/pages/viewpage.action?pageId=70230508\" target=\"_blank\">https://wiki.cancerimagingarchive.net/pages/viewpage.action?pageId=70230508</a></p>\n<p>It was produced with this I believe:</p>\n<p>It looks to be digital  <a href=\"https://www.htig.com/catalog/mammography-equipment/ge-senographe-ds/\" target=\"_blank\">https://www.htig.com/catalog/mammography-equipment/ge-senographe-ds/</a></p>\n<p>Here's your data as a service, I think <a href=\"https://portal.imaging.datacommons.cancer.gov/explore/filters/?collection_id=cmmd\" target=\"_blank\">https://portal.imaging.datacommons.cancer.gov/explore/filters/?collection_id=cmmd</a></p>\n<p>Also this, </p>\n<p><a href=\"https://physionet.org/content/vindr-mammo/1.0.0/\" target=\"_blank\">https://physionet.org/content/vindr-mammo/1.0.0/</a></p>\n<p><a href=\"https://www.kaggle.com/vbookshelf\" target=\"_blank\">@vbookshelf</a> also created a dataset for it 19 days ago, but if he mentioned it here, I didn't see it.</p>\n<p>The below-the-radar approach is quite intriguing.</p>\n<p><a href=\"https://www.kaggle.com/datasets/vbookshelf/mammogram-mass-analyzer-v00\" target=\"_blank\">https://www.kaggle.com/datasets/vbookshelf/mammogram-mass-analyzer-v00</a></p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2070517,
              "author_name": "David Roberts",
              "author_url": "",
              "post_date": "2022-12-20T04:58:41.110000",
              "content": "<p>I haven't seen the CCMD dataset, but I have worked with the Vindr mammo set. Vindr is very nice quality FFDM, but the labels don't include benign/malignant if I remember correctly. It may be good for pretraining or pseudo-labeling I guess.</p>\n<p>Interesting link to the datacommons site. Thanks!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2070547,
              "author_name": "@kaggleqrdl",
              "author_url": "",
              "post_date": "2022-12-20T06:17:12.860000",
              "content": "<p>Well, it has the birads level 4 and 5, which I imagine isn't that far off from what 'cancer' means (?) in this dataset.   It seems a bit ambigous, tbh.  </p>\n<p>But that's not really the point, I think, but rather the goal is ROI patch extraction.   Unless the theory is that all tumours are somehow breast wide and this can be detected by the networks (is that a theory?), our networks are going to be learning a lot of strange things unless they focus in on the right spots.</p>\n<p>If you survey the literature, a great deal of it takes this two stage approach.  ROI patch -&gt; resnet/efnet to tell if the patch contains malignant tissue.   </p>\n<p>This will probably still be problematic as some patches won't have malignant tissue even though they  are extracted from breasts which are diagnosed with cancer.   But, this can be mitigated somewhat by leveraging the Vindr DS and the birads ratings.   Note that  BI-RADS 5 is a category of breast lesions that have at least a 95% probability of malignancy</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2069952,
      "author_name": "@kaggleqrdl",
      "author_url": "",
      "post_date": "2022-12-19T13:01:21.073000",
      "content": "<p>Ok, well, I decided to try something.  I emailed 18 different researchers who contributed to a paper in 2022 on using yolo models for detecting breast cancer (there are a lot more) to see if they would share their yolo models for an opensource effort in this area.  I will create a public dataset on Kaggle with any responses.  I am hopeful, but not particularly optimistic.</p>\n<p><a href=\"https://scholar.google.com/scholar?as_ylo=2022&amp;q=breast+cancer+yolo&amp;hl=en&amp;as_sdt=0,5\" target=\"_blank\">https://scholar.google.com/scholar?as_ylo=2022&amp;q=breast+cancer+yolo&amp;hl=en&amp;as_sdt=0,5</a></p>\n<p>888 results, just for 2022.  </p>",
      "votes": 2,
      "replies": [
        {
          "id": 2069971,
          "author_name": "Remek Kinas",
          "author_url": "",
          "post_date": "2022-12-19T13:29:13.947000",
          "content": "<p>As to yolo … you can use it. Problem is with annotation I think. I am (and probably most of participants) correct annotate data. Certainly you can use external annotated dataset but I am not sure if such exist or we can use it in this competition.</p>\n<p>I thnink that classifier is not a final solution (I want to see a. good classifier trained on such imbalanced dataset b. I think that there is too little positive cases to generalize well). It can be used as a part of solution. I know that currently it looks like this is the best option but soon you will see different solution (I am sure). </p>",
          "votes": 1,
          "replies": [
            {
              "id": 2070066,
              "author_name": "@kaggleqrdl",
              "author_url": "",
              "post_date": "2022-12-19T15:14:45.537000",
              "content": "<p>Yeah, I think Yolo is best for finding ROI patches which are then sent through closer looking models.  That seems to be the general solution to most of the papers that I read.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2069839": "After reading the umpteenth paper about some model or collection of models (and a lot of yolo models) that can detect breast cancer - but with no available models for download and repro, I really have to vent.\n\nThis seems like a deeply tragic state of affairs.  So much work and training is being done, which when thoughtfully ensembled together, might actually do some amazing things for a field which could really benefit from it.   Through ensembling, we'd be able to quickly see which research is doing something original and fresh versus which research is just correlating highly with everything else.\n\nThere really needs to be a movement in ML academia that if you're going to publish research, you need to at the very least push your pytorch models to a collective zoo.   *Especially* any research which is performed under government grants.\n",
    "2069975": "I think this is common in the medical field. Especially in the US where medical data is hard to come by. Commercial entities who pay for data aren't going to open-source their software. The challenge with radiologic studies is that they're expensive to produce and require specific licenses to order and perform. Add the legal issues of PHI (Patient Health Information), and it's nearly impossible to find anything open-source.",
    "2069952": "Ok, well, I decided to try something.  I emailed 18 different researchers who contributed to a paper in 2022 on using yolo models for detecting breast cancer (there are a lot more) to see if they would share their yolo models for an opensource effort in this area.  I will create a public dataset on Kaggle with any responses.  I am hopeful, but not particularly optimistic.\n\nhttps://scholar.google.com/scholar?as_ylo=2022&q=breast+cancer+yolo&hl=en&as_sdt=0,5\n\n888 results, just for 2022.  "
  }
}