{
  "id": 374946,
  "title": "Part 2 - Breast Density - Everything you wanted to know about mammography: A radiologist’s guide",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/374946",
  "author_name": "JAbrantes",
  "post_date": "2022-12-29T16:05:19.912000",
  "votes": 31,
  "comment_count": 18,
  "views": 0,
  "content": "<p>Hello 👋</p>\n<p>I’m João Abrantes, MD, Radiologist from Portugal.</p>\n<p>Welcome to part 2 ✌️of this informal guide to mammography analysis, where I aim to offer some insight into the mammography analysis and some tips and tricks that can (hopefully) be helpful for the development of better algorithms.</p>\n<p>✔️You can check part 1 here: <a href=\"url\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/374288</a></p>\n<p>Breast density is an important factor in breast cancer screening, as it can affect the visibility of abnormalities on mammograms. The American College of Radiology (ACR) has developed a classification system to categorize breast density based on the percentage and distribution of glandular tissue and fat in the breasts. </p>\n<p>Breast density is not something that can be visually assessed, but rather it is determined by a radiologist (or algorithm) after the analysis of a mammogram. Breast tissue is made up of both glandular tissue (“white” or “dense” in the mammogram) and fat (“darker” in the mammogram), and the relative proportions of these two types of tissue can affect the visibility of abnormalities on a mammogram.</p>\n<p>The ACR BI-RADS atlas classifies these patterns of density into four categories: A,B,C,D; with A being the most fatty and D being the most dense.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8926747%2F7336026abf23e122078e59cb13a5e2f1%2FACRbreastdensity.png?generation=1672327547984176&amp;alt=media\" alt=\"\"><br>\n<em>Source: ACR BI-RADS Atlas</em></p>\n<h3>BREAST DENSITY AND RISK</h3>\n<p>Breast density affects mammographic screening in two ways:</p>\n<ul>\n<li>Masking Effect – overlapping dense breast tissue on an underlying cancer</li>\n<li>Breast density itself is an independent risk factor for breast cancer (ACR C or D – Relative Risk 1.3 [1], ACR D- Relative Risk 2.1 [2])</li>\n</ul>\n<h3>DENSITY AND AGE</h3>\n<p>A trend exists for a lower percentage of higher breast density with advancing age, as published in data from several screening programs.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8926747%2F335a8947d321c63896d9afeacea2af31%2Fimage2.png?generation=1672327629969925&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8926747%2F914410738ba06b7486727b8900cce97c%2FImagem21.png?generation=1672327957188757&amp;alt=media\" alt=\"\"></p>\n<p>Dutch Screening Data</p>\n<p><em>Source: Carla van Gils - University Medical Center, UMC Utrecht</em></p>\n<h3>DENSITY AND MASKING EFFECT</h3>\n<p>As previously mentioned, one of the ways breast density affects mammographic analysis is by the masking effect of overlapping dense breast tissue on an underlying cancer.<br>\nThis can, at least partially, account for a higher number of interval cancers. </p>\n<blockquote>\n  <p>Side note: An interval cancer is a breast cancer that is diagnosed between regularly scheduled screening mammograms. It is called \"interval\" cancer because it was discovered in the interval between two mammograms. They can be more aggressive and have a worse prognosis as they are often larger and may have spread beyond the breast by the time they are diagnosed.</p>\n</blockquote>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8926747%2F6df95ec016ea1f85fb5c32c98a60263a%2FImagem3.png?generation=1672327712856623&amp;alt=media\" alt=\"\"></p>\n<p><em>Source [3]</em></p>\n<h3>🔥My take:</h3>\n<p>As we’ve seen, breast density is a major factor that can influence breast cancer detection and prognosis.</p>\n<p>How can we leverage this information for the specific task?</p>\n<blockquote>\n  <p>🤓Some ideas:</p>\n  <ul>\n  <li>Develop a breast density classification algorithm to define the category of the mammogram.</li>\n  <li>Train different models for each specific breast density</li>\n  <li>Apply a specific model for the defined breast density category – either in isolation, or with different weights according to the first breast density classification algorithm.<br>\n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8926747%2Fd6738def94bcbb9b3dda2fae25ffb087%2Fimage4.png?generation=1672327799779577&amp;alt=media\" alt=\"\"></li>\n  </ul>\n</blockquote>\n<p>Stay tuned for the next topics which I’ll be covering:</p>\n<ul>\n<li>Left/right and CC/MLO comparison of images;</li>\n<li>Additional Mammographic Views;</li>\n<li>and much more.</li>\n</ul>\n<p>👍I’ll be watching the comments and try to do my best to positively contribute to the discussion!</p>\n<h3>References</h3>\n<p>[1]     M. J., \" Collaborative Modeling of U.S. Breast Cancer Screening Strategies,\" Agency for Healthcare Research and Quality , 2015. <br>\n[2]     P. Freer, \" Mammographic Breast Density: Impact on Breast Cancer Risk and Implications for Screening.,\" RadioGraphics , 2015. <br>\n[3]     J. Wanders, \"Volumetric breast density affects performance of digital screening mammography.,\" Breast Cancer Res Treat , 2017. </p>",
  "messages": [
    {
      "id": 2079792,
      "postDate": "2022-12-29T16:05:19.913Z",
      "content": "<p>Hello 👋</p>\n<p>I’m João Abrantes, MD, Radiologist from Portugal.</p>\n<p>Welcome to part 2 ✌️of this informal guide to mammography analysis, where I aim to offer some insight into the mammography analysis and some tips and tricks that can (hopefully) be helpful for the development of better algorithms.</p>\n<p>✔️You can check part 1 here: <a href=\"url\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/374288</a></p>\n<p>Breast density is an important factor in breast cancer screening, as it can affect the visibility of abnormalities on mammograms. The American College of Radiology (ACR) has developed a classification system to categorize breast density based on the percentage and distribution of glandular tissue and fat in the breasts. </p>\n<p>Breast density is not something that can be visually assessed, but rather it is determined by a radiologist (or algorithm) after the analysis of a mammogram. Breast tissue is made up of both glandular tissue (“white” or “dense” in the mammogram) and fat (“darker” in the mammogram), and the relative proportions of these two types of tissue can affect the visibility of abnormalities on a mammogram.</p>\n<p>The ACR BI-RADS atlas classifies these patterns of density into four categories: A,B,C,D; with A being the most fatty and D being the most dense.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8926747%2F7336026abf23e122078e59cb13a5e2f1%2FACRbreastdensity.png?generation=1672327547984176&amp;alt=media\" alt=\"\"><br>\n<em>Source: ACR BI-RADS Atlas</em></p>\n<h3>BREAST DENSITY AND RISK</h3>\n<p>Breast density affects mammographic screening in two ways:</p>\n<ul>\n<li>Masking Effect – overlapping dense breast tissue on an underlying cancer</li>\n<li>Breast density itself is an independent risk factor for breast cancer (ACR C or D – Relative Risk 1.3 [1], ACR D- Relative Risk 2.1 [2])</li>\n</ul>\n<h3>DENSITY AND AGE</h3>\n<p>A trend exists for a lower percentage of higher breast density with advancing age, as published in data from several screening programs.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8926747%2F335a8947d321c63896d9afeacea2af31%2Fimage2.png?generation=1672327629969925&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8926747%2F914410738ba06b7486727b8900cce97c%2FImagem21.png?generation=1672327957188757&amp;alt=media\" alt=\"\"></p>\n<p>Dutch Screening Data</p>\n<p><em>Source: Carla van Gils - University Medical Center, UMC Utrecht</em></p>\n<h3>DENSITY AND MASKING EFFECT</h3>\n<p>As previously mentioned, one of the ways breast density affects mammographic analysis is by the masking effect of overlapping dense breast tissue on an underlying cancer.<br>\nThis can, at least partially, account for a higher number of interval cancers. </p>\n<blockquote>\n  <p>Side note: An interval cancer is a breast cancer that is diagnosed between regularly scheduled screening mammograms. It is called \"interval\" cancer because it was discovered in the interval between two mammograms. They can be more aggressive and have a worse prognosis as they are often larger and may have spread beyond the breast by the time they are diagnosed.</p>\n</blockquote>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8926747%2F6df95ec016ea1f85fb5c32c98a60263a%2FImagem3.png?generation=1672327712856623&amp;alt=media\" alt=\"\"></p>\n<p><em>Source [3]</em></p>\n<h3>🔥My take:</h3>\n<p>As we’ve seen, breast density is a major factor that can influence breast cancer detection and prognosis.</p>\n<p>How can we leverage this information for the specific task?</p>\n<blockquote>\n  <p>🤓Some ideas:</p>\n  <ul>\n  <li>Develop a breast density classification algorithm to define the category of the mammogram.</li>\n  <li>Train different models for each specific breast density</li>\n  <li>Apply a specific model for the defined breast density category – either in isolation, or with different weights according to the first breast density classification algorithm.<br>\n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8926747%2Fd6738def94bcbb9b3dda2fae25ffb087%2Fimage4.png?generation=1672327799779577&amp;alt=media\" alt=\"\"></li>\n  </ul>\n</blockquote>\n<p>Stay tuned for the next topics which I’ll be covering:</p>\n<ul>\n<li>Left/right and CC/MLO comparison of images;</li>\n<li>Additional Mammographic Views;</li>\n<li>and much more.</li>\n</ul>\n<p>👍I’ll be watching the comments and try to do my best to positively contribute to the discussion!</p>\n<h3>References</h3>\n<p>[1]     M. J., \" Collaborative Modeling of U.S. Breast Cancer Screening Strategies,\" Agency for Healthcare Research and Quality , 2015. <br>\n[2]     P. Freer, \" Mammographic Breast Density: Impact on Breast Cancer Risk and Implications for Screening.,\" RadioGraphics , 2015. <br>\n[3]     J. Wanders, \"Volumetric breast density affects performance of digital screening mammography.,\" Breast Cancer Res Treat , 2017. </p>",
      "rawMarkdown": "Hello 👋\n\nI’m João Abrantes, MD, Radiologist from Portugal.\n\nWelcome to part 2 ✌️of this informal guide to mammography analysis, where I aim to offer some insight into the mammography analysis and some tips and tricks that can (hopefully) be helpful for the development of better algorithms.\n\n✔️You can check part 1 here: [https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/374288](url)\n \nBreast density is an important factor in breast cancer screening, as it can affect the visibility of abnormalities on mammograms. The American College of Radiology (ACR) has developed a classification system to categorize breast density based on the percentage and distribution of glandular tissue and fat in the breasts. \n\nBreast density is not something that can be visually assessed, but rather it is determined by a radiologist (or algorithm) after the analysis of a mammogram. Breast tissue is made up of both glandular tissue (“white” or “dense” in the mammogram) and fat (“darker” in the mammogram), and the relative proportions of these two types of tissue can affect the visibility of abnormalities on a mammogram.\n\nThe ACR BI-RADS atlas classifies these patterns of density into four categories: A,B,C,D; with A being the most fatty and D being the most dense.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8926747%2F7336026abf23e122078e59cb13a5e2f1%2FACRbreastdensity.png?generation=1672327547984176&alt=media)\n*Source: ACR BI-RADS Atlas*\n\n### BREAST DENSITY AND RISK\n\nBreast density affects mammographic screening in two ways:\n- Masking Effect – overlapping dense breast tissue on an underlying cancer\n- Breast density itself is an independent risk factor for breast cancer (ACR C or D – Relative Risk 1.3 [1], ACR D- Relative Risk 2.1 [2])\n\n### DENSITY AND AGE\nA trend exists for a lower percentage of higher breast density with advancing age, as published in data from several screening programs.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8926747%2F335a8947d321c63896d9afeacea2af31%2Fimage2.png?generation=1672327629969925&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8926747%2F914410738ba06b7486727b8900cce97c%2FImagem21.png?generation=1672327957188757&alt=media)\n\n\nDutch Screening Data\n\n*Source: Carla van Gils - University Medical Center, UMC Utrecht*\n\n### DENSITY AND MASKING EFFECT\n\nAs previously mentioned, one of the ways breast density affects mammographic analysis is by the masking effect of overlapping dense breast tissue on an underlying cancer.\nThis can, at least partially, account for a higher number of interval cancers. \n\n>Side note: An interval cancer is a breast cancer that is diagnosed between regularly scheduled screening mammograms. It is called \"interval\" cancer because it was discovered in the interval between two mammograms. They can be more aggressive and have a worse prognosis as they are often larger and may have spread beyond the breast by the time they are diagnosed.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8926747%2F6df95ec016ea1f85fb5c32c98a60263a%2FImagem3.png?generation=1672327712856623&alt=media)\n\n*Source [3]*\n\n### 🔥My take:\n\nAs we’ve seen, breast density is a major factor that can influence breast cancer detection and prognosis.\n\nHow can we leverage this information for the specific task?\n\n>🤓Some ideas:\n- Develop a breast density classification algorithm to define the category of the mammogram.\n- Train different models for each specific breast density\n- Apply a specific model for the defined breast density category – either in isolation, or with different weights according to the first breast density classification algorithm.\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8926747%2Fd6738def94bcbb9b3dda2fae25ffb087%2Fimage4.png?generation=1672327799779577&alt=media)\n\nStay tuned for the next topics which I’ll be covering:\n\n- Left/right and CC/MLO comparison of images;\n- Additional Mammographic Views;\n- and much more.\n \n\n👍I’ll be watching the comments and try to do my best to positively contribute to the discussion!\n\n### References\n\n[1] \tM. J., \" Collaborative Modeling of U.S. Breast Cancer Screening Strategies,\" Agency for Healthcare Research and Quality , 2015. \n[2] \tP. Freer, \" Mammographic Breast Density: Impact on Breast Cancer Risk and Implications for Screening.,\" RadioGraphics , 2015. \n[3] \tJ. Wanders, \"Volumetric breast density affects performance of digital screening mammography.,\" Breast Cancer Res Treat , 2017. \n",
      "votes": 29
    },
    {
      "id": 2079931,
      "postDate": "2022-12-29T18:01:48.617Z",
      "content": "<p>Thank you so much for your insights. I don't know if it was sent in another discussion but I found a paper related to this topic: <a href=\"https://arxiv.org/abs/2209.09809\" target=\"_blank\">High-resolution synthesis of high-density breast mammograms: Application to improved fairness in deep learning based mass detection\n</a>. It is a good read and they developed gan to do breast density transfer with some results in improving metrics. Here is a chart they showed for the detection performance based on breast density.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5918909%2F91c7d4e33bd88f1b6f615ea05894b5c7%2Fchat.PNG?generation=1672335673594753&amp;alt=media\" alt=\"Chart Result\"><br>\n(<em><a href=\"https://arxiv.org/abs/2209.09809\" target=\"_blank\">High-resolution synthesis of high-density breast mammograms: Application to improved fairness in deep learning based mass detection</a> Figure 1</em></p>\n<p>For the RSNA competition dataset, we have a lot of missing data in the density column (maybe because the dataset is a compilation of different private datasets as discussed <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/370333#2076917\" target=\"_blank\">here</a>) but I am guessing the NA values will have a similar distribution.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5918909%2Fedfa96904306f4d7e31eb812c445ea2b%2F__results___3_0%20(1).png?generation=1672336295946289&amp;alt=media\" alt=\"RSNA Competition Data Results\"></p>\n<p>Another idea is someone can train a model to predict densities for missing densities values, and re-train as an aux target or use it in some other compacity (maybe as a weight calculation for each density).</p>\n<p>There is a lot to do and model generalization is very important here because of the many types of images we have as well as the densities each can have.</p>\n<p>It is implemented and made public here if you're interested: <a href=\"https://github.com/RichardObi/medigan\" target=\"_blank\">https://github.com/RichardObi/medigan</a></p>",
      "rawMarkdown": "Thank you so much for your insights. I don't know if it was sent in another discussion but I found a paper related to this topic: [High-resolution synthesis of high-density breast mammograms: Application to improved fairness in deep learning based mass detection\n](https://arxiv.org/abs/2209.09809). It is a good read and they developed gan to do breast density transfer with some results in improving metrics. Here is a chart they showed for the detection performance based on breast density.\n\n![Chart Result](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5918909%2F91c7d4e33bd88f1b6f615ea05894b5c7%2Fchat.PNG?generation=1672335673594753&alt=media)\n(*[High-resolution synthesis of high-density breast mammograms: Application to improved fairness in deep learning based mass detection](https://arxiv.org/abs/2209.09809) Figure 1*\n\nFor the RSNA competition dataset, we have a lot of missing data in the density column (maybe because the dataset is a compilation of different private datasets as discussed [here](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/370333#2076917)) but I am guessing the NA values will have a similar distribution.\n\n![RSNA Competition Data Results](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5918909%2Fedfa96904306f4d7e31eb812c445ea2b%2F__results___3_0%20(1).png?generation=1672336295946289&alt=media)\n\nAnother idea is someone can train a model to predict densities for missing densities values, and re-train as an aux target or use it in some other compacity (maybe as a weight calculation for each density).\n\nThere is a lot to do and model generalization is very important here because of the many types of images we have as well as the densities each can have.\n\nIt is implemented and made public here if you're interested: https://github.com/RichardObi/medigan",
      "votes": 3,
      "replies": [
        {
          "id": 2080083,
          "postDate": "2022-12-29T20:49:17.407Z",
          "content": "<p>Thank you for the input! Yes, I agree that the distribution of the missing data should mimic the distribution of the know data ( and data already known from screening programs). Thank you for your sources👌</p>",
          "rawMarkdown": "Thank you for the input! Yes, I agree that the distribution of the missing data should mimic the distribution of the know data ( and data already known from screening programs). Thank you for your sources👌",
          "votes": 1
        }
      ]
    },
    {
      "id": 2079917,
      "postDate": "2022-12-29T17:52:51.387Z",
      "content": "<p>Good information <a href=\"https://www.kaggle.com/jabrantes\" target=\"_blank\">@jabrantes</a> .. I think the challenge will be in windowing each image properly. There isn't much time during infer to apply much 'per-image' windowing. </p>\n<p>Standardizing/normalizing to 8bit seems counter intuitive without some windowing/leveling first. Likewise, normalizing in 16 bit doesn't make sense either. </p>",
      "rawMarkdown": "Good information @jabrantes .. I think the challenge will be in windowing each image properly. There isn't much time during infer to apply much 'per-image' windowing. \n\nStandardizing/normalizing to 8bit seems counter intuitive without some windowing/leveling first. Likewise, normalizing in 16 bit doesn't make sense either. ",
      "votes": 3
    },
    {
      "id": 2079899,
      "postDate": "2022-12-29T17:32:43.417Z",
      "content": "<p>do you have pfbeta or AUC scores of cancer vs non-cancers for each breast density?<br>\nyou have to prove how big is the problem before the solution.</p>\n<p>e.g. just start with a public notebook or some simple solution and train a single model for all. Then measure the error for each density. </p>\n<p>Search prior art/paper for the same idea. sometimes the results are not what you think (machine and human interpret images differently, etc)</p>",
      "rawMarkdown": "do you have pfbeta or AUC scores of cancer vs non-cancers for each breast density?\nyou have to prove how big is the problem before the solution.\n\ne.g. just start with a public notebook or some simple solution and train a single model for all. Then measure the error for each density. \n\nSearch prior art/paper for the same idea. sometimes the results are not what you think (machine and human interpret images differently, etc)",
      "votes": 3,
      "replies": [
        {
          "id": 2080085,
          "postDate": "2022-12-29T20:50:40.050Z",
          "content": "<p>I’m going to join a team to work on the data and infer if this is a viable solution/ big problem to tackle!</p>",
          "rawMarkdown": "I’m going to join a team to work on the data and infer if this is a viable solution/ big problem to tackle!",
          "votes": 1,
          "replies": [
            {
              "id": 2080220,
              "postDate": "2022-12-30T00:56:51.077Z",
              "content": "<p>the more correct way to use breast density<br>\n<a href=\"https://www.frontiersin.org/articles/10.3389/fradi.2021.796078/full\" target=\"_blank\">https://www.frontiersin.org/articles/10.3389/fradi.2021.796078/full</a><br>\n(figure.1 and table.4)</p>",
              "rawMarkdown": "the more correct way to use breast density\nhttps://www.frontiersin.org/articles/10.3389/fradi.2021.796078/full\n(figure.1 and table.4)",
              "votes": 3
            },
            {
              "id": 2080505,
              "postDate": "2022-12-30T08:17:11.437Z",
              "content": "<p>Great source! Also addresses the multiple view evaluation of mammography (CC/MLO)</p>",
              "rawMarkdown": "Great source! Also addresses the multiple view evaluation of mammography (CC/MLO)"
            },
            {
              "id": 2080589,
              "postDate": "2022-12-30T09:33:34.540Z",
              "content": "<p>Hopefully it won't discourage you from continuing to openly and generously providing info like the above :)  Still, competing is another way to contribute for sure.</p>\n<p>One question I have for you.. this comp has a weird constraint in that we have to infer on 32K images in a 9 hour window.  I'm wondering what the practical utility is of such a solution optimized for those constraints.  Any thoughts?</p>",
              "rawMarkdown": "Hopefully it won't discourage you from continuing to openly and generously providing info like the above :)  Still, competing is another way to contribute for sure.\n\nOne question I have for you.. this comp has a weird constraint in that we have to infer on 32K images in a 9 hour window.  I'm wondering what the practical utility is of such a solution optimized for those constraints.  Any thoughts?",
              "votes": 2
            },
            {
              "id": 2080597,
              "postDate": "2022-12-30T09:48:39.850Z",
              "content": "<p>I’m also in the path to learn about ML/DL in medicine, so this kind of interactions are so valuable and interesting that I’m even more eager to contribute how I can.<br>\nI can only speculate about the time constraints for inference, but I can identify some factors that would benefit a “optimized” solution, such as:</p>\n<ul>\n<li>The clinical use of these models can be as a “computed aided diagnosis”, and as such, if you have a large inference time, the radiologist may not wait around for the results of the AI before interpreting the exam.</li>\n<li>The IT infrastructure in healthcare is very heterogeneous ( specifically if we think about worldwide implementation),  so I can see how a more optimized solution would be more valuable.</li>\n</ul>",
              "rawMarkdown": "I’m also in the path to learn about ML/DL in medicine, so this kind of interactions are so valuable and interesting that I’m even more eager to contribute how I can.\nI can only speculate about the time constraints for inference, but I can identify some factors that would benefit a “optimized” solution, such as:\n- The clinical use of these models can be as a “computed aided diagnosis”, and as such, if you have a large inference time, the radiologist may not wait around for the results of the AI before interpreting the exam.\n- The IT infrastructure in healthcare is very heterogeneous ( specifically if we think about worldwide implementation),  so I can see how a more optimized solution would be more valuable.",
              "votes": 2
            },
            {
              "id": 2080612,
              "postDate": "2022-12-30T10:08:38.510Z",
              "content": "<p><a href=\"https://www.kaggle.com/jabrantes\" target=\"_blank\">@jabrantes</a> Good context.  While I have you, what are you seeing in general in SOTA commercial CAD used in industry at the moment?  </p>\n<p>I see a lot of amazing SOTA results in papers, but I assume some of that stuff is still theoretical.  There was one paper - <a href=\"https://hajim.rochester.edu/ece/sites/parker/assets/pdf/243-diagnostic-performance-of-an-artificial-intelligence-system-in-breast-ultrasound.pdf\" target=\"_blank\">https://hajim.rochester.edu/ece/sites/parker/assets/pdf/243-diagnostic-performance-of-an-artificial-intelligence-system-in-breast-ultrasound.pdf</a> which was interesting.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9052057%2F632e2b9985a30e8f63b6ebb727bd2a1b%2Fresults.png?generation=1672394874249785&amp;alt=media\" alt=\"\"></p>\n<p>It's using a commercial product called S-Detect.  It doesn't have an AUC metric because the company refuses to give confidence levels I think.  It's either suspicious or not suspicious - dichotomous.</p>\n<p>I found it fascinating how the junior radiologists have higher specificity / accuracy than the the senior radiologists, but much lower sensitivity .  Sensitivity is probably the most important, and likely senior folks sacrifice and err on the side of false positives for that, given the high rate of false negatives in breast cancer detection. </p>\n<p>Getting FP/FN down would be great I think in BC detection, a lot of women are resistant to getting scanned because the FP/FN is so high.  Nobody likes being frightened :(  </p>\n<p>Maybe with 3d mammography.</p>",
              "rawMarkdown": "@jabrantes Good context.  While I have you, what are you seeing in general in SOTA commercial CAD used in industry at the moment?  \n\nI see a lot of amazing SOTA results in papers, but I assume some of that stuff is still theoretical.  There was one paper - https://hajim.rochester.edu/ece/sites/parker/assets/pdf/243-diagnostic-performance-of-an-artificial-intelligence-system-in-breast-ultrasound.pdf which was interesting.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9052057%2F632e2b9985a30e8f63b6ebb727bd2a1b%2Fresults.png?generation=1672394874249785&alt=media)\n\nIt's using a commercial product called S-Detect.  It doesn't have an AUC metric because the company refuses to give confidence levels I think.  It's either suspicious or not suspicious - dichotomous.\n\nI found it fascinating how the junior radiologists have higher specificity / accuracy than the the senior radiologists, but much lower sensitivity .  Sensitivity is probably the most important, and likely senior folks sacrifice and err on the side of false positives for that, given the high rate of false negatives in breast cancer detection. \n\nGetting FP/FN down would be great I think in BC detection, a lot of women are resistant to getting scanned because the FP/FN is so high.  Nobody likes being frightened :(  \n\nMaybe with 3d mammography."
            },
            {
              "id": 2080621,
              "postDate": "2022-12-30T10:17:50.567Z",
              "content": "<p>Here's what chatgpt says:</p>\n<p>There are several reasons why senior radiologists may be less accurate or specific in detecting breast cancer while maintaining high sensitivity. Some possible reasons include:</p>\n<p>Bias: Like all humans, radiologists may be subject to biases that can affect their decision-making and judgment. For example, a radiologist who has seen many cases of breast cancer may be more likely to overestimate the likelihood of breast cancer in a given case, leading to lower specificity but higher sensitivity.</p>\n<p>Experience: While experience is generally considered to be a positive factor in many fields, it can also lead to overconfidence and a tendency to rely on past experiences rather than carefully evaluating all of the available data. This can lead to a decrease in accuracy and specificity, but may also result in a higher sensitivity to detecting breast cancer.</p>\n<p>Workload: Senior radiologists may be more likely to be working with heavier workloads, which can lead to fatigue and a decreased ability to focus and pay attention to detail. This can result in lower accuracy and specificity in detecting breast cancer, but may also result in a higher sensitivity to detecting breast cancer.</p>\n<p>Technological limitations: Some senior radiologists may be less familiar with newer technologies and techniques that can improve accuracy and specificity in detecting breast cancer. For example, newer AI algorithms may be able to incorporate more data sources and modalities than are available to human radiologists, which can improve accuracy and specificity. This may result in senior radiologists having lower accuracy and specificity, but higher sensitivity.</p>",
              "rawMarkdown": "Here's what chatgpt says:\n\nThere are several reasons why senior radiologists may be less accurate or specific in detecting breast cancer while maintaining high sensitivity. Some possible reasons include:\n\nBias: Like all humans, radiologists may be subject to biases that can affect their decision-making and judgment. For example, a radiologist who has seen many cases of breast cancer may be more likely to overestimate the likelihood of breast cancer in a given case, leading to lower specificity but higher sensitivity.\n\nExperience: While experience is generally considered to be a positive factor in many fields, it can also lead to overconfidence and a tendency to rely on past experiences rather than carefully evaluating all of the available data. This can lead to a decrease in accuracy and specificity, but may also result in a higher sensitivity to detecting breast cancer.\n\nWorkload: Senior radiologists may be more likely to be working with heavier workloads, which can lead to fatigue and a decreased ability to focus and pay attention to detail. This can result in lower accuracy and specificity in detecting breast cancer, but may also result in a higher sensitivity to detecting breast cancer.\n\nTechnological limitations: Some senior radiologists may be less familiar with newer technologies and techniques that can improve accuracy and specificity in detecting breast cancer. For example, newer AI algorithms may be able to incorporate more data sources and modalities than are available to human radiologists, which can improve accuracy and specificity. This may result in senior radiologists having lower accuracy and specificity, but higher sensitivity."
            },
            {
              "id": 2085103,
              "postDate": "2023-01-03T23:49:05.093Z",
              "content": "<p>I guess, you have to make an inference model that fuses some steps of computations.</p>",
              "rawMarkdown": "I guess, you have to make an inference model that fuses some steps of computations."
            }
          ]
        }
      ]
    },
    {
      "id": 2080548,
      "postDate": "2022-12-30T09:08:15.523Z",
      "content": "<p><a href=\"https://www.kaggle.com/jabrantes\" target=\"_blank\">@jabrantes</a> Awesome information!  It's always fantastic to see domain experts contribute like this.  Makes one very optimistic and quite grateful that kaggle exists.  </p>\n<p><a href=\"https://www.kaggle.com/argv\" target=\"_blank\">@argv</a>  <a href=\"https://www.kaggle.com/mylesoneill\" target=\"_blank\">@mylesoneill</a> It'd be great to surface this sort of thing better on kaggle and find ways to reward/highlight/leverage.  We saw this over in the novo enzymes comp as well and it's amazing getting such domain experts openly and very generously collaborating alongside ml experts.  It really elevates Kaggle, imho.</p>",
      "rawMarkdown": "@jabrantes Awesome information!  It's always fantastic to see domain experts contribute like this.  Makes one very optimistic and quite grateful that kaggle exists.  \n\n@argv  @mylesoneill It'd be great to surface this sort of thing better on kaggle and find ways to reward/highlight/leverage.  We saw this over in the novo enzymes comp as well and it's amazing getting such domain experts openly and very generously collaborating alongside ml experts.  It really elevates Kaggle, imho.\n\n",
      "votes": 2
    },
    {
      "id": 2147643,
      "postDate": "2023-02-16T18:57:19.447Z",
      "content": "<p>maybe any auto-clustering technique could do the job of density classification based on image histograms.. just an idea of lazy person </p>",
      "rawMarkdown": "maybe any auto-clustering technique could do the job of density classification based on image histograms.. just an idea of lazy person "
    },
    {
      "id": 2088697,
      "postDate": "2023-01-06T15:11:23.350Z",
      "content": "<p>Good Information</p>",
      "rawMarkdown": "Good Information"
    },
    {
      "id": 2086180,
      "postDate": "2023-01-04T16:12:49.230Z",
      "content": "<p>Gracias por tu información. Obrigado pela informação. Thanks for the information</p>",
      "rawMarkdown": "Gracias por tu información. Obrigado pela informação. Thanks for the information"
    },
    {
      "id": 2083824,
      "postDate": "2023-01-02T22:04:07.137Z",
      "content": "<p><a href=\"https://www.kaggle.com/jabrantes\" target=\"_blank\">@jabrantes</a> qué Genio!!! Muchas gracias, ingresé hace poco a Kaggle y tu información ha sido realmente muy productiva!!!</p>",
      "rawMarkdown": "@jabrantes qué Genio!!! Muchas gracias, ingresé hace poco a Kaggle y tu información ha sido realmente muy productiva!!!"
    },
    {
      "id": 2082852,
      "postDate": "2023-01-02T02:58:58.163Z",
      "content": "<p>Thanks so much for this. Great intro for us</p>",
      "rawMarkdown": "Thanks so much for this. Great intro for us"
    }
  ],
  "comments": [
    {
      "id": 2079931,
      "author_name": "outwrest",
      "author_url": "",
      "post_date": "2022-12-29T18:01:48.617000",
      "content": "<p>Thank you so much for your insights. I don't know if it was sent in another discussion but I found a paper related to this topic: <a href=\"https://arxiv.org/abs/2209.09809\" target=\"_blank\">High-resolution synthesis of high-density breast mammograms: Application to improved fairness in deep learning based mass detection\n</a>. It is a good read and they developed gan to do breast density transfer with some results in improving metrics. Here is a chart they showed for the detection performance based on breast density.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5918909%2F91c7d4e33bd88f1b6f615ea05894b5c7%2Fchat.PNG?generation=1672335673594753&amp;alt=media\" alt=\"Chart Result\"><br>\n(<em><a href=\"https://arxiv.org/abs/2209.09809\" target=\"_blank\">High-resolution synthesis of high-density breast mammograms: Application to improved fairness in deep learning based mass detection</a> Figure 1</em></p>\n<p>For the RSNA competition dataset, we have a lot of missing data in the density column (maybe because the dataset is a compilation of different private datasets as discussed <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/370333#2076917\" target=\"_blank\">here</a>) but I am guessing the NA values will have a similar distribution.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5918909%2Fedfa96904306f4d7e31eb812c445ea2b%2F__results___3_0%20(1).png?generation=1672336295946289&amp;alt=media\" alt=\"RSNA Competition Data Results\"></p>\n<p>Another idea is someone can train a model to predict densities for missing densities values, and re-train as an aux target or use it in some other compacity (maybe as a weight calculation for each density).</p>\n<p>There is a lot to do and model generalization is very important here because of the many types of images we have as well as the densities each can have.</p>\n<p>It is implemented and made public here if you're interested: <a href=\"https://github.com/RichardObi/medigan\" target=\"_blank\">https://github.com/RichardObi/medigan</a></p>",
      "votes": 3,
      "replies": [
        {
          "id": 2080083,
          "author_name": "JAbrantes",
          "author_url": "",
          "post_date": "2022-12-29T20:49:17.407000",
          "content": "<p>Thank you for the input! Yes, I agree that the distribution of the missing data should mimic the distribution of the know data ( and data already known from screening programs). Thank you for your sources👌</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2079917,
      "author_name": "David Roberts",
      "author_url": "",
      "post_date": "2022-12-29T17:52:51.387000",
      "content": "<p>Good information <a href=\"https://www.kaggle.com/jabrantes\" target=\"_blank\">@jabrantes</a> .. I think the challenge will be in windowing each image properly. There isn't much time during infer to apply much 'per-image' windowing. </p>\n<p>Standardizing/normalizing to 8bit seems counter intuitive without some windowing/leveling first. Likewise, normalizing in 16 bit doesn't make sense either. </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2079899,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-12-29T17:32:43.417000",
      "content": "<p>do you have pfbeta or AUC scores of cancer vs non-cancers for each breast density?<br>\nyou have to prove how big is the problem before the solution.</p>\n<p>e.g. just start with a public notebook or some simple solution and train a single model for all. Then measure the error for each density. </p>\n<p>Search prior art/paper for the same idea. sometimes the results are not what you think (machine and human interpret images differently, etc)</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2080085,
          "author_name": "JAbrantes",
          "author_url": "",
          "post_date": "2022-12-29T20:50:40.050000",
          "content": "<p>I’m going to join a team to work on the data and infer if this is a viable solution/ big problem to tackle!</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2080220,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2022-12-30T00:56:51.077000",
              "content": "<p>the more correct way to use breast density<br>\n<a href=\"https://www.frontiersin.org/articles/10.3389/fradi.2021.796078/full\" target=\"_blank\">https://www.frontiersin.org/articles/10.3389/fradi.2021.796078/full</a><br>\n(figure.1 and table.4)</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2080505,
              "author_name": "JAbrantes",
              "author_url": "",
              "post_date": "2022-12-30T08:17:11.437000",
              "content": "<p>Great source! Also addresses the multiple view evaluation of mammography (CC/MLO)</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2080589,
              "author_name": "@kaggleqrdl",
              "author_url": "",
              "post_date": "2022-12-30T09:33:34.540000",
              "content": "<p>Hopefully it won't discourage you from continuing to openly and generously providing info like the above :)  Still, competing is another way to contribute for sure.</p>\n<p>One question I have for you.. this comp has a weird constraint in that we have to infer on 32K images in a 9 hour window.  I'm wondering what the practical utility is of such a solution optimized for those constraints.  Any thoughts?</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2080597,
              "author_name": "JAbrantes",
              "author_url": "",
              "post_date": "2022-12-30T09:48:39.850000",
              "content": "<p>I’m also in the path to learn about ML/DL in medicine, so this kind of interactions are so valuable and interesting that I’m even more eager to contribute how I can.<br>\nI can only speculate about the time constraints for inference, but I can identify some factors that would benefit a “optimized” solution, such as:</p>\n<ul>\n<li>The clinical use of these models can be as a “computed aided diagnosis”, and as such, if you have a large inference time, the radiologist may not wait around for the results of the AI before interpreting the exam.</li>\n<li>The IT infrastructure in healthcare is very heterogeneous ( specifically if we think about worldwide implementation),  so I can see how a more optimized solution would be more valuable.</li>\n</ul>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2080612,
              "author_name": "@kaggleqrdl",
              "author_url": "",
              "post_date": "2022-12-30T10:08:38.510000",
              "content": "<p><a href=\"https://www.kaggle.com/jabrantes\" target=\"_blank\">@jabrantes</a> Good context.  While I have you, what are you seeing in general in SOTA commercial CAD used in industry at the moment?  </p>\n<p>I see a lot of amazing SOTA results in papers, but I assume some of that stuff is still theoretical.  There was one paper - <a href=\"https://hajim.rochester.edu/ece/sites/parker/assets/pdf/243-diagnostic-performance-of-an-artificial-intelligence-system-in-breast-ultrasound.pdf\" target=\"_blank\">https://hajim.rochester.edu/ece/sites/parker/assets/pdf/243-diagnostic-performance-of-an-artificial-intelligence-system-in-breast-ultrasound.pdf</a> which was interesting.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9052057%2F632e2b9985a30e8f63b6ebb727bd2a1b%2Fresults.png?generation=1672394874249785&amp;alt=media\" alt=\"\"></p>\n<p>It's using a commercial product called S-Detect.  It doesn't have an AUC metric because the company refuses to give confidence levels I think.  It's either suspicious or not suspicious - dichotomous.</p>\n<p>I found it fascinating how the junior radiologists have higher specificity / accuracy than the the senior radiologists, but much lower sensitivity .  Sensitivity is probably the most important, and likely senior folks sacrifice and err on the side of false positives for that, given the high rate of false negatives in breast cancer detection. </p>\n<p>Getting FP/FN down would be great I think in BC detection, a lot of women are resistant to getting scanned because the FP/FN is so high.  Nobody likes being frightened :(  </p>\n<p>Maybe with 3d mammography.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2080621,
              "author_name": "@kaggleqrdl",
              "author_url": "",
              "post_date": "2022-12-30T10:17:50.567000",
              "content": "<p>Here's what chatgpt says:</p>\n<p>There are several reasons why senior radiologists may be less accurate or specific in detecting breast cancer while maintaining high sensitivity. Some possible reasons include:</p>\n<p>Bias: Like all humans, radiologists may be subject to biases that can affect their decision-making and judgment. For example, a radiologist who has seen many cases of breast cancer may be more likely to overestimate the likelihood of breast cancer in a given case, leading to lower specificity but higher sensitivity.</p>\n<p>Experience: While experience is generally considered to be a positive factor in many fields, it can also lead to overconfidence and a tendency to rely on past experiences rather than carefully evaluating all of the available data. This can lead to a decrease in accuracy and specificity, but may also result in a higher sensitivity to detecting breast cancer.</p>\n<p>Workload: Senior radiologists may be more likely to be working with heavier workloads, which can lead to fatigue and a decreased ability to focus and pay attention to detail. This can result in lower accuracy and specificity in detecting breast cancer, but may also result in a higher sensitivity to detecting breast cancer.</p>\n<p>Technological limitations: Some senior radiologists may be less familiar with newer technologies and techniques that can improve accuracy and specificity in detecting breast cancer. For example, newer AI algorithms may be able to incorporate more data sources and modalities than are available to human radiologists, which can improve accuracy and specificity. This may result in senior radiologists having lower accuracy and specificity, but higher sensitivity.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2085103,
              "author_name": "joejeo1",
              "author_url": "",
              "post_date": "2023-01-03T23:49:05.093000",
              "content": "<p>I guess, you have to make an inference model that fuses some steps of computations.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2080548,
      "author_name": "@kaggleqrdl",
      "author_url": "",
      "post_date": "2022-12-30T09:08:15.523000",
      "content": "<p><a href=\"https://www.kaggle.com/jabrantes\" target=\"_blank\">@jabrantes</a> Awesome information!  It's always fantastic to see domain experts contribute like this.  Makes one very optimistic and quite grateful that kaggle exists.  </p>\n<p><a href=\"https://www.kaggle.com/argv\" target=\"_blank\">@argv</a>  <a href=\"https://www.kaggle.com/mylesoneill\" target=\"_blank\">@mylesoneill</a> It'd be great to surface this sort of thing better on kaggle and find ways to reward/highlight/leverage.  We saw this over in the novo enzymes comp as well and it's amazing getting such domain experts openly and very generously collaborating alongside ml experts.  It really elevates Kaggle, imho.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2147643,
      "author_name": "tomassa",
      "author_url": "",
      "post_date": "2023-02-16T18:57:19.447000",
      "content": "<p>maybe any auto-clustering technique could do the job of density classification based on image histograms.. just an idea of lazy person </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2088697,
      "author_name": "tau__tsm1",
      "author_url": "",
      "post_date": "2023-01-06T15:11:23.350000",
      "content": "<p>Good Information</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2086180,
      "author_name": "Egilda Perez",
      "author_url": "",
      "post_date": "2023-01-04T16:12:49.230000",
      "content": "<p>Gracias por tu información. Obrigado pela informação. Thanks for the information</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2083824,
      "author_name": "Cecil Alvarez",
      "author_url": "",
      "post_date": "2023-01-02T22:04:07.137000",
      "content": "<p><a href=\"https://www.kaggle.com/jabrantes\" target=\"_blank\">@jabrantes</a> qué Genio!!! Muchas gracias, ingresé hace poco a Kaggle y tu información ha sido realmente muy productiva!!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2082852,
      "author_name": "Justin Hodges",
      "author_url": "",
      "post_date": "2023-01-02T02:58:58.163000",
      "content": "<p>Thanks so much for this. Great intro for us</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2079792": "Hello 👋\n\nI’m João Abrantes, MD, Radiologist from Portugal.\n\nWelcome to part 2 ✌️of this informal guide to mammography analysis, where I aim to offer some insight into the mammography analysis and some tips and tricks that can (hopefully) be helpful for the development of better algorithms.\n\n✔️You can check part 1 here: [https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/374288](url)\n \nBreast density is an important factor in breast cancer screening, as it can affect the visibility of abnormalities on mammograms. The American College of Radiology (ACR) has developed a classification system to categorize breast density based on the percentage and distribution of glandular tissue and fat in the breasts. \n\nBreast density is not something that can be visually assessed, but rather it is determined by a radiologist (or algorithm) after the analysis of a mammogram. Breast tissue is made up of both glandular tissue (“white” or “dense” in the mammogram) and fat (“darker” in the mammogram), and the relative proportions of these two types of tissue can affect the visibility of abnormalities on a mammogram.\n\nThe ACR BI-RADS atlas classifies these patterns of density into four categories: A,B,C,D; with A being the most fatty and D being the most dense.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8926747%2F7336026abf23e122078e59cb13a5e2f1%2FACRbreastdensity.png?generation=1672327547984176&alt=media)\n*Source: ACR BI-RADS Atlas*\n\n### BREAST DENSITY AND RISK\n\nBreast density affects mammographic screening in two ways:\n- Masking Effect – overlapping dense breast tissue on an underlying cancer\n- Breast density itself is an independent risk factor for breast cancer (ACR C or D – Relative Risk 1.3 [1], ACR D- Relative Risk 2.1 [2])\n\n### DENSITY AND AGE\nA trend exists for a lower percentage of higher breast density with advancing age, as published in data from several screening programs.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8926747%2F335a8947d321c63896d9afeacea2af31%2Fimage2.png?generation=1672327629969925&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8926747%2F914410738ba06b7486727b8900cce97c%2FImagem21.png?generation=1672327957188757&alt=media)\n\n\nDutch Screening Data\n\n*Source: Carla van Gils - University Medical Center, UMC Utrecht*\n\n### DENSITY AND MASKING EFFECT\n\nAs previously mentioned, one of the ways breast density affects mammographic analysis is by the masking effect of overlapping dense breast tissue on an underlying cancer.\nThis can, at least partially, account for a higher number of interval cancers. \n\n>Side note: An interval cancer is a breast cancer that is diagnosed between regularly scheduled screening mammograms. It is called \"interval\" cancer because it was discovered in the interval between two mammograms. They can be more aggressive and have a worse prognosis as they are often larger and may have spread beyond the breast by the time they are diagnosed.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8926747%2F6df95ec016ea1f85fb5c32c98a60263a%2FImagem3.png?generation=1672327712856623&alt=media)\n\n*Source [3]*\n\n### 🔥My take:\n\nAs we’ve seen, breast density is a major factor that can influence breast cancer detection and prognosis.\n\nHow can we leverage this information for the specific task?\n\n>🤓Some ideas:\n- Develop a breast density classification algorithm to define the category of the mammogram.\n- Train different models for each specific breast density\n- Apply a specific model for the defined breast density category – either in isolation, or with different weights according to the first breast density classification algorithm.\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8926747%2Fd6738def94bcbb9b3dda2fae25ffb087%2Fimage4.png?generation=1672327799779577&alt=media)\n\nStay tuned for the next topics which I’ll be covering:\n\n- Left/right and CC/MLO comparison of images;\n- Additional Mammographic Views;\n- and much more.\n \n\n👍I’ll be watching the comments and try to do my best to positively contribute to the discussion!\n\n### References\n\n[1] \tM. J., \" Collaborative Modeling of U.S. Breast Cancer Screening Strategies,\" Agency for Healthcare Research and Quality , 2015. \n[2] \tP. Freer, \" Mammographic Breast Density: Impact on Breast Cancer Risk and Implications for Screening.,\" RadioGraphics , 2015. \n[3] \tJ. Wanders, \"Volumetric breast density affects performance of digital screening mammography.,\" Breast Cancer Res Treat , 2017. \n",
    "2079931": "Thank you so much for your insights. I don't know if it was sent in another discussion but I found a paper related to this topic: [High-resolution synthesis of high-density breast mammograms: Application to improved fairness in deep learning based mass detection\n](https://arxiv.org/abs/2209.09809). It is a good read and they developed gan to do breast density transfer with some results in improving metrics. Here is a chart they showed for the detection performance based on breast density.\n\n![Chart Result](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5918909%2F91c7d4e33bd88f1b6f615ea05894b5c7%2Fchat.PNG?generation=1672335673594753&alt=media)\n(*[High-resolution synthesis of high-density breast mammograms: Application to improved fairness in deep learning based mass detection](https://arxiv.org/abs/2209.09809) Figure 1*\n\nFor the RSNA competition dataset, we have a lot of missing data in the density column (maybe because the dataset is a compilation of different private datasets as discussed [here](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/370333#2076917)) but I am guessing the NA values will have a similar distribution.\n\n![RSNA Competition Data Results](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5918909%2Fedfa96904306f4d7e31eb812c445ea2b%2F__results___3_0%20(1).png?generation=1672336295946289&alt=media)\n\nAnother idea is someone can train a model to predict densities for missing densities values, and re-train as an aux target or use it in some other compacity (maybe as a weight calculation for each density).\n\nThere is a lot to do and model generalization is very important here because of the many types of images we have as well as the densities each can have.\n\nIt is implemented and made public here if you're interested: https://github.com/RichardObi/medigan",
    "2079917": "Good information @jabrantes .. I think the challenge will be in windowing each image properly. There isn't much time during infer to apply much 'per-image' windowing. \n\nStandardizing/normalizing to 8bit seems counter intuitive without some windowing/leveling first. Likewise, normalizing in 16 bit doesn't make sense either. ",
    "2079899": "do you have pfbeta or AUC scores of cancer vs non-cancers for each breast density?\nyou have to prove how big is the problem before the solution.\n\ne.g. just start with a public notebook or some simple solution and train a single model for all. Then measure the error for each density. \n\nSearch prior art/paper for the same idea. sometimes the results are not what you think (machine and human interpret images differently, etc)",
    "2080548": "@jabrantes Awesome information!  It's always fantastic to see domain experts contribute like this.  Makes one very optimistic and quite grateful that kaggle exists.  \n\n@argv  @mylesoneill It'd be great to surface this sort of thing better on kaggle and find ways to reward/highlight/leverage.  We saw this over in the novo enzymes comp as well and it's amazing getting such domain experts openly and very generously collaborating alongside ml experts.  It really elevates Kaggle, imho.\n\n",
    "2147643": "maybe any auto-clustering technique could do the job of density classification based on image histograms.. just an idea of lazy person ",
    "2088697": "Good Information",
    "2086180": "Gracias por tu información. Obrigado pela informação. Thanks for the information",
    "2083824": "@jabrantes qué Genio!!! Muchas gracias, ingresé hace poco a Kaggle y tu información ha sido realmente muy productiva!!!",
    "2082852": "Thanks so much for this. Great intro for us"
  }
}