{
  "id": 375267,
  "title": "need some advice from doctors or radiologists ...",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/375267",
  "author_name": "hengck23",
  "post_date": "2022-12-31T10:11:24.501000",
  "votes": 29,
  "comment_count": 28,
  "views": 0,
  "content": "<p>results on validation:<br>\n<img src=\"https://i.ibb.co/vXx9MX3/Selection-401.png\" alt=\"https://i.ibb.co/vXx9MX3/Selection-401.png\"><br>\n<img src=\"https://i.ibb.co/L00DzmW/Selection-400.png\" alt=\"https://i.ibb.co/L00DzmW/Selection-400.png\"></p>\n<p>Question : How is very small lesions detected? i can find many similar regions that are false positives. How do you look at the whole image to decide if it is cancerous or not? Or do you really can decide just based on local context region?</p>\n<p>(basically i am trying to decide the context for good detection but it seems impossible. It is even difficult on training set)</p>",
  "messages": [
    {
      "id": 2081540,
      "postDate": "2022-12-31T10:11:24.500Z",
      "content": "<p>results on validation:<br>\n<img src=\"https://i.ibb.co/vXx9MX3/Selection-401.png\" alt=\"https://i.ibb.co/vXx9MX3/Selection-401.png\"><br>\n<img src=\"https://i.ibb.co/L00DzmW/Selection-400.png\" alt=\"https://i.ibb.co/L00DzmW/Selection-400.png\"></p>\n<p>Question : How is very small lesions detected? i can find many similar regions that are false positives. How do you look at the whole image to decide if it is cancerous or not? Or do you really can decide just based on local context region?</p>\n<p>(basically i am trying to decide the context for good detection but it seems impossible. It is even difficult on training set)</p>",
      "rawMarkdown": "results on validation:\n![https://i.ibb.co/vXx9MX3/Selection-401.png](https://i.ibb.co/vXx9MX3/Selection-401.png)\n![https://i.ibb.co/L00DzmW/Selection-400.png](https://i.ibb.co/L00DzmW/Selection-400.png)\n\nQuestion : How is very small lesions detected? i can find many similar regions that are false positives. How do you look at the whole image to decide if it is cancerous or not? Or do you really can decide just based on local context region?\n\n(basically i am trying to decide the context for good detection but it seems impossible. It is even difficult on training set)",
      "votes": 29
    },
    {
      "id": 2081791,
      "postDate": "2022-12-31T16:18:52.997Z",
      "content": "<p>I am not a radiologist, but I have worked on medical imaging software for 30 years and am very familiar with mammography. I don't think it's a reasonable assumption that a radiologist would be able to diagnose cancer from a tiny density like the ones in these images.</p>\n<p>Normally, CAD will pick up the small densities and direct the radiologists attention to the area. Then, the radiologist will \"window\" the image manually to try and find the \"sweet spot\" where the density in question is separated from the rest of the tissue as much as possible in order to determine the location, shape, symmetry etc. Most of the time, a follow on ultrasound (and a biopsy) is needed to verify what the dense tissue actually is.</p>\n<p>Since these are screening mammograms (not follow ups), we don't expect to see many cases of large tumors. It seems the goal of this competition is to do something that humans cannot do .. which is to diagnose cancer from imaging alone.</p>",
      "rawMarkdown": "I am not a radiologist, but I have worked on medical imaging software for 30 years and am very familiar with mammography. I don't think it's a reasonable assumption that a radiologist would be able to diagnose cancer from a tiny density like the ones in these images.\n\nNormally, CAD will pick up the small densities and direct the radiologists attention to the area. Then, the radiologist will \"window\" the image manually to try and find the \"sweet spot\" where the density in question is separated from the rest of the tissue as much as possible in order to determine the location, shape, symmetry etc. Most of the time, a follow on ultrasound (and a biopsy) is needed to verify what the dense tissue actually is.\n\nSince these are screening mammograms (not follow ups), we don't expect to see many cases of large tumors. It seems the goal of this competition is to do something that humans cannot do .. which is to diagnose cancer from imaging alone.",
      "votes": 20,
      "replies": [
        {
          "id": 2081918,
          "postDate": "2022-12-31T20:15:56.937Z",
          "content": "<p>thx for this answer <a href=\"https://www.kaggle.com/davidbroberts\" target=\"_blank\">@davidbroberts</a>, so much good info here 🙏 I think I learned more from your reply here on how all of this works, how the work gets done vs many of the much longer posts, really cool 🙂</p>\n<p>thank you!</p>",
          "rawMarkdown": "thx for this answer @davidbroberts, so much good info here 🙏 I think I learned more from your reply here on how all of this works, how the work gets done vs many of the much longer posts, really cool 🙂\n\nthank you!",
          "votes": 1
        },
        {
          "id": 2082005,
          "postDate": "2023-01-01T01:18:40.297Z",
          "content": "<p><a href=\"https://www.kaggle.com/davidbroberts\" target=\"_blank\">@davidbroberts</a> <br>\nthanks a lot for the reply.</p>\n<p>so i can assume:<br>\n[1]  the annotation process is likely to follows what you have described. even in open dataset like vindr, the process will be CAD+verification, i.e. there may be miss detection by CAD.</p>\n<p>[2]  hence one can give out of trying to perfect a segmentation model for mask annotation. instead one should devise a mask model to include all positive mask, but reduce the num of negative mask. it is ok that for one cancer image to have at least one +ve mask and several neg mask.</p>\n<p>or even in aux loss, we cannot assume to have perfect annotation for mask.<br>\nnow i understand why they use weak models like resnet22 in heatmap prediction (and not a strong one like renset50)</p>\n<p>[3]  Now i understand whay the two-stage approach: </p>\n<ul>\n<li>train an high sensitivity (maybe poor specificity) heatmap detector at high resolution.</li>\n<li>then feed the heatmap to a better specificity image model.</li>\n</ul>\n<p>the loss function is important to ensure the model is high sensitivity or specificity in the two stages.</p>\n<p>actually i also have another observation at external data:<br>\nalthough the sore for mask detection is poor (e.g. for small region), the image level prediction (e.g. BIRADS) is generally good.<br>\ni use a joint loss end-to-end model (heatmap prediction + image level prediction)</p>",
          "rawMarkdown": "@davidbroberts \nthanks a lot for the reply.\n\nso i can assume:\n[1]  the annotation process is likely to follows what you have described. even in open dataset like vindr, the process will be CAD+verification, i.e. there may be miss detection by CAD.\n\n[2]  hence one can give out of trying to perfect a segmentation model for mask annotation. instead one should devise a mask model to include all positive mask, but reduce the num of negative mask. it is ok that for one cancer image to have at least one +ve mask and several neg mask.\n\nor even in aux loss, we cannot assume to have perfect annotation for mask.\nnow i understand why they use weak models like resnet22 in heatmap prediction (and not a strong one like renset50)\n\n\n[3]  Now i understand whay the two-stage approach: \n- train an high sensitivity (maybe poor specificity) heatmap detector at high resolution.\n- then feed the heatmap to a better specificity image model.\n\nthe loss function is important to ensure the model is high sensitivity or specificity in the two stages.\n\nactually i also have another observation at external data:\nalthough the sore for mask detection is poor (e.g. for small region), the image level prediction (e.g. BIRADS) is generally good.\ni use a joint loss end-to-end model (heatmap prediction + image level prediction)",
          "votes": 1,
          "replies": [
            {
              "id": 2082041,
              "postDate": "2023-01-01T03:16:59.447Z",
              "content": "<p>check this !!!</p>\n<p><a href=\"https://www.youtube.com/watch?v=FJl_yf5Au68&amp;t=138s\" target=\"_blank\">https://www.youtube.com/watch?v=FJl_yf5Au68&amp;t=138s</a><br>\nHow to read a screening mammogram</p>\n<p>hence, it is not important to consider use of simple model like resnet22,34 efficientnet b2 or complicated large ones like efficientnet b7.</p>\n<p>a good workflow is more important. e.g. models for different scales, different parts, models for single image, models for multiple images of same laterality, multiple images of different laterality, models for ROI level,  models for  image levels ….</p>\n<p>an ensemble of different type of models seems to be the key.<br>\n<a href=\"https://ibb.co/Nr7J5n6\"><img src=\"https://i.ibb.co/1MbP1Jf/Selection-421.png\" alt=\"Selection-421\"></a><br>\n<a href=\"https://ibb.co/7gz7V1c\"><img src=\"https://i.ibb.co/pWPCZLD/Selection-420.png\" alt=\"Selection-420\"></a><br>\n<a href=\"https://ibb.co/g7g4hXY\"><img src=\"https://i.ibb.co/4fN13Dk/Selection-419.png\" alt=\"Selection-419\"></a></p>",
              "rawMarkdown": "check this !!!\n\n\nhttps://www.youtube.com/watch?v=FJl_yf5Au68&t=138s\nHow to read a screening mammogram\n\n\nhence, it is not important to consider use of simple model like resnet22,34 efficientnet b2 or complicated large ones like efficientnet b7.\n\na good workflow is more important. e.g. models for different scales, different parts, models for single image, models for multiple images of same laterality, multiple images of different laterality, models for ROI level,  models for  image levels ....\n\nan ensemble of different type of models seems to be the key.\n<a href=\"https://ibb.co/Nr7J5n6\"><img src=\"https://i.ibb.co/1MbP1Jf/Selection-421.png\" alt=\"Selection-421\" border=\"0\"></a>\n<a href=\"https://ibb.co/7gz7V1c\"><img src=\"https://i.ibb.co/pWPCZLD/Selection-420.png\" alt=\"Selection-420\" border=\"0\"></a>\n<a href=\"https://ibb.co/g7g4hXY\"><img src=\"https://i.ibb.co/4fN13Dk/Selection-419.png\" alt=\"Selection-419\" border=\"0\"></a>\n\n\n",
              "votes": 4
            },
            {
              "id": 2082044,
              "postDate": "2023-01-01T03:28:17.150Z",
              "content": "<p><img src=\"https://i.ibb.co/Z2bW7MY/Selection-425.png\" alt=\"https://i.ibb.co/Z2bW7MY/Selection-425.png\"><br>\n<img src=\"https://i.ibb.co/LRy2GjL/Selection-432.png\" alt=\"https://i.ibb.co/LRy2GjL/Selection-432.png\"></p>",
              "rawMarkdown": "![https://i.ibb.co/Z2bW7MY/Selection-425.png](https://i.ibb.co/Z2bW7MY/Selection-425.png)\n![https://i.ibb.co/LRy2GjL/Selection-432.png](https://i.ibb.co/LRy2GjL/Selection-432.png)",
              "votes": 3
            },
            {
              "id": 2083683,
              "postDate": "2023-01-02T18:27:42.363Z",
              "content": "<p>That's pretty interesting.  Looks like laterality alignment (left ontop of right) is a critical pre-processing activity.</p>",
              "rawMarkdown": "That's pretty interesting.  Looks like laterality alignment (left ontop of right) is a critical pre-processing activity."
            },
            {
              "id": 2090891,
              "postDate": "2023-01-07T18:36:26.867Z",
              "content": "<blockquote>\n  <p>a good workflow is more important. e.g. models for different scales, different parts, models for single image, models for multiple images of same laterality, multiple images of different laterality, models for ROI level, models for image levels ….</p>\n</blockquote>\n<p>and just how would such multiplicity of models work in a notebook in under 9 hours on 43K high resolution images?</p>",
              "rawMarkdown": ">a good workflow is more important. e.g. models for different scales, different parts, models for single image, models for multiple images of same laterality, multiple images of different laterality, models for ROI level, models for image levels ….\n\nand just how would such multiplicity of models work in a notebook in under 9 hours on 43K high resolution images?"
            },
            {
              "id": 2091004,
              "postDate": "2023-01-07T22:29:54.830Z",
              "content": "<p>Yep, that's the issue.  TPU / tensorflow might be the way to go as it performs generally better than pytorch, but harder to use.</p>",
              "rawMarkdown": "Yep, that's the issue.  TPU / tensorflow might be the way to go as it performs generally better than pytorch, but harder to use.",
              "votes": 1
            },
            {
              "id": 2091010,
              "postDate": "2023-01-07T22:45:57.960Z",
              "content": "<p>I agree…<br>\nmy point is, we need to be extremely stingy in our approach… simple and effective trickery is worth more than here than complex models…<br>\nwork in progress…</p>",
              "rawMarkdown": "I agree...\nmy point is, we need to be extremely stingy in our approach... simple and effective trickery is worth more than here than complex models...\nwork in progress..."
            },
            {
              "id": 2091021,
              "postDate": "2023-01-07T23:08:09.170Z",
              "content": "<p><a href=\"https://www.kaggle.com/houssamassila\" target=\"_blank\">@houssamassila</a> </p>\n<p>Check this out - <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/376630\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/376630</a></p>",
              "rawMarkdown": "@houssamassila \n\nCheck this out - https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/376630",
              "votes": 1
            },
            {
              "id": 2091031,
              "postDate": "2023-01-08T00:03:33.193Z",
              "content": "<p>yes I did check that earlier today. thanks.<br>\neven further than just enlarging the usable ram, trickery regarding the dataset is also useful, like using smaller files when possible/reducing the dataset (even if it seems counter-intuitive at first). To me, this seems more the way to go than spending time trying ever more extravaggant models that won't run in 9 hours.</p>",
              "rawMarkdown": "yes I did check that earlier today. thanks.\neven further than just enlarging the usable ram, trickery regarding the dataset is also useful, like using smaller files when possible/reducing the dataset (even if it seems counter-intuitive at first). To me, this seems more the way to go than spending time trying ever more extravaggant models that won't run in 9 hours."
            },
            {
              "id": 2091061,
              "postDate": "2023-01-08T01:13:37.720Z",
              "content": "<p>Hmm, someone is saying we can't use TPU for this comp.  Which is weird, given that someone created a TPU notebook.   But it might be useful for training.</p>\n<p>edit:  reading up on it, I think it's because the TPUs require internet access which is not allowed in this comp</p>",
              "rawMarkdown": "Hmm, someone is saying we can't use TPU for this comp.  Which is weird, given that someone created a TPU notebook.   But it might be useful for training.\n\nedit:  reading up on it, I think it's because the TPUs require internet access which is not allowed in this comp",
              "votes": 1
            },
            {
              "id": 2091353,
              "postDate": "2023-01-08T10:19:10.467Z",
              "content": "<p>it's as you said, TPU will be useful only to train models that can later be loaded into the submission notebook.</p>",
              "rawMarkdown": "it's as you said, TPU will be useful only to train models that can later be loaded into the submission notebook.\n"
            },
            {
              "id": 2094507,
              "postDate": "2023-01-10T19:39:39.123Z",
              "content": "<p><a href=\"https://www.kaggle.com/houssamassila\" target=\"_blank\">@houssamassila</a> <br>\n\"just how would such multiplicity of models work in a notebook in under 9 hours on 43K high resolution images?\"</p>\n<p>\"early rejection\" is a common trick in computer vision. 10x speed up is not problem</p>\n<p><a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/377008#2094477\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/377008#2094477</a></p>",
              "rawMarkdown": "@houssamassila \n\"just how would such multiplicity of models work in a notebook in under 9 hours on 43K high resolution images?\"\n\n\n\"early rejection\" is a common trick in computer vision. 10x speed up is not problem\n\nhttps://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/377008#2094477"
            }
          ]
        }
      ]
    },
    {
      "id": 2125004,
      "postDate": "2023-02-01T11:12:57.127Z",
      "content": "<p>I see you are looking for a radiologist to team up <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>. I worked with <a href=\"https://www.kaggle.com/sandorkonya\" target=\"_blank\">@sandorkonya</a> before and he is the best one I've ever worked with. Tagged both of you for potential collab.</p>",
      "rawMarkdown": "I see you are looking for a radiologist to team up @hengck23. I worked with @sandorkonya before and he is the best one I've ever worked with. Tagged both of you for potential collab.",
      "votes": 2,
      "replies": [
        {
          "id": 2125090,
          "postDate": "2023-02-01T12:14:31.927Z",
          "content": "<p>The idea is simple.<br>\nWith my top models, i will generate CAM heatmap.<br>\nThen the radiology correct it. (Delete very sure wrong heat map for activated case, mark possible one for missed cases)</p>\n<p>i do not need very accurate annotation. one can simply just annotate a point.<br>\nI do not need all annotation, just at least one per pos image.</p>\n<p>only 2% of train  data are pos. this is very few images. annotation can done within hours (because i do not need good annotation).</p>\n<p>Basically it is manual active learning. we will go through 2~3 rounds.</p>\n<p>Let's see if this will fight imbalance data and gives better generalization, rather than relying on unstable LB/CV  </p>\n<hr>\n<p><a href=\"https://ojs.aaai.org/index.php/AAAI/article/download/16153/15960\" target=\"_blank\">https://ojs.aaai.org/index.php/AAAI/article/download/16153/15960</a><br>\nfigure.2 is example of point annotation</p>",
          "rawMarkdown": "The idea is simple.\nWith my top models, i will generate CAM heatmap.\nThen the radiology correct it. (Delete very sure wrong heat map for activated case, mark possible one for missed cases)\n\ni do not need very accurate annotation. one can simply just annotate a point.\nI do not need all annotation, just at least one per pos image.\n\nonly 2% of train  data are pos. this is very few images. annotation can done within hours (because i do not need good annotation).\n\nBasically it is manual active learning. we will go through 2~3 rounds.\n\nLet's see if this will fight imbalance data and gives better generalization, rather than relying on unstable LB/CV  \n\n---\nhttps://ojs.aaai.org/index.php/AAAI/article/download/16153/15960\nfigure.2 is example of point annotation\n",
          "votes": 2,
          "replies": [
            {
              "id": 2132386,
              "postDate": "2023-02-06T19:23:12Z",
              "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, you probably know this, but the <a href=\"https://www.physionet.org/content/vindr-mammo/1.0.0/\" target=\"_blank\">vindr-mammo</a> dataset already has the bounding box annotations that you are looking for. It would be interesting to see how much juice you can squeeze out of those annotations - gives you a quick way to assess if it's worth getting the RSNA dataset annotated by a radiologist…</p>\n<p>My intuition is that it would help a lot. If you can show a model the ROI that causes a mammogram to be tagged as positive, the model is likely to learn the important features better. I explored this a bit, but haven't found a good way to incorporate the bounding box data. One possible method is to divide the images into tiles, label each tile and then train on the tiles and their labels. The hard part, at least for me is aggregating the tile-level predictions into a single prediction for the whole image. A global maxpooling operation doesn't seem to work well.</p>",
              "rawMarkdown": "@hengck23, you probably know this, but the [vindr-mammo](https://www.physionet.org/content/vindr-mammo/1.0.0/) dataset already has the bounding box annotations that you are looking for. It would be interesting to see how much juice you can squeeze out of those annotations - gives you a quick way to assess if it's worth getting the RSNA dataset annotated by a radiologist...\n\nMy intuition is that it would help a lot. If you can show a model the ROI that causes a mammogram to be tagged as positive, the model is likely to learn the important features better. I explored this a bit, but haven't found a good way to incorporate the bounding box data. One possible method is to divide the images into tiles, label each tile and then train on the tiles and their labels. The hard part, at least for me is aggregating the tile-level predictions into a single prediction for the whole image. A global maxpooling operation doesn't seem to work well."
            }
          ]
        },
        {
          "id": 2125687,
          "postDate": "2023-02-01T19:56:41.807Z",
          "content": "<p><a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> - thank you for the generous words i barely deserve it though.<br>\nI read through <em>some</em> of the discussions of the contest.</p>\n<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> are you using the dataset from <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/377790\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/377790</a> already? Those images have a region annotation and you can do a point annotation from that pretty easily. The mammography technique is quite standardised (in contrary to other projectionsradiographic material). I see that there are labels on the above set which could misguide the classification… so i would mask out the mamma on all of the images (of the official and this one above) and just add black bg to all as a preprocessing step. That area does not contain any valuable information regarding the tumor anyway. The fine-grained structure of the organ on the image is very similar!<br>\nI hope you are already working on full resolution (patches, sliding windows ) because you have only chance to detect the suspect microcalcifications on (near?) full res.</p>\n<p>Dense (fibrous and glandular) breast tissue looks \"white\" on a mammogram. Breast masses and cancers can also look \"white\", so the classification in this category is the hardest task. These breasts have usually an uncertain BIRAD score! This is why you probably see some patients having more than 4 images, especially those with fibrous tissue. The extra images are focused, zoomed images or repeated images due to false compression.</p>\n<p>The problem with annotating mammography is that the annotation itself doesn't take long (one or 4 points doesn't matter) … but fiding the lesion! =) Even annotating 1.1k images (2% positive of 55k) would take a substantial time. <br>\nWe could try some to see how it affects the generalisation… i have a reporting monitor <a href=\"https://www.kaggle.com/home\" target=\"_blank\">@home</a> for good contrast.</p>",
          "rawMarkdown": "@gunesevitan - thank you for the generous words i barely deserve it though.\nI read through *some* of the discussions of the contest.\n\n@hengck23 are you using the dataset from https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/377790 already? Those images have a region annotation and you can do a point annotation from that pretty easily. The mammography technique is quite standardised (in contrary to other projectionsradiographic material). I see that there are labels on the above set which could misguide the classification... so i would mask out the mamma on all of the images (of the official and this one above) and just add black bg to all as a preprocessing step. That area does not contain any valuable information regarding the tumor anyway. The fine-grained structure of the organ on the image is very similar!\nI hope you are already working on full resolution (patches, sliding windows ) because you have only chance to detect the suspect microcalcifications on (near?) full res.\n\nDense (fibrous and glandular) breast tissue looks \"white\" on a mammogram. Breast masses and cancers can also look \"white\", so the classification in this category is the hardest task. These breasts have usually an uncertain BIRAD score! This is why you probably see some patients having more than 4 images, especially those with fibrous tissue. The extra images are focused, zoomed images or repeated images due to false compression.\n\nThe problem with annotating mammography is that the annotation itself doesn't take long (one or 4 points doesn't matter) ... but fiding the lesion! =) Even annotating 1.1k images (2% positive of 55k) would take a substantial time. \nWe could try some to see how it affects the generalisation... i have a reporting monitor @home for good contrast.\n",
          "votes": 4,
          "replies": [
            {
              "id": 2125703,
              "postDate": "2023-02-01T20:24:05.240Z",
              "content": "<p>Thank you very much for explanation. Is any chace you show us the example of dense (fibrous and glandular) breast tissue,  breast masses and cancers? </p>\n<p>a. It is hard for me do see these issues on images.<br>\nb. second - I made prototype for synthetic breast cancer generator (but I am not able to confirm if it works - it was made by me for fun - to see if it is possible to generate such changes on images. I am aware that it is far from \"good\" issue we should classify) </p>\n<p><img src=\"http://zapodaj.net/images/6608d09e85fec.jpg\" alt=\"image_1\"></p>\n<p><img src=\"http://zapodaj.net/images/a896deacf2e13.jpg\" alt=\"image_2\"></p>\n<p><img src=\"http://zapodaj.net/images/9131c16081453.jpg\" alt=\"image_3\"></p>\n<p><img src=\"http://zapodaj.net/images/002e5863d493a.jpg\" alt=\"image_4\"></p>",
              "rawMarkdown": "Thank you very much for explanation. Is any chace you show us the example of dense (fibrous and glandular) breast tissue,  breast masses and cancers? \n\na. It is hard for me do see these issues on images.\nb. second - I made prototype for synthetic breast cancer generator (but I am not able to confirm if it works - it was made by me for fun - to see if it is possible to generate such changes on images. I am aware that it is far from \"good\" issue we should classify) \n\n![image_1](http://zapodaj.net/images/6608d09e85fec.jpg)\n\n![image_2](http://zapodaj.net/images/a896deacf2e13.jpg)\n\n![image_3](http://zapodaj.net/images/9131c16081453.jpg)\n\n![image_4](http://zapodaj.net/images/002e5863d493a.jpg)",
              "votes": 3
            },
            {
              "id": 2125805,
              "postDate": "2023-02-01T22:45:37.867Z",
              "content": "<p>Hi there,</p>\n<p>for densities take a look at this pic, you will see what the differences are:<br>\n<a href=\"https://www.cancer.gov/sites/g/files/xnrzdm211/files/styles/cgov_article/public/cgov_image/media_image/2022-04/BIRADS_Updated%20Image.png?itok=sCglGBKJ\" target=\"_blank\">https://www.cancer.gov/sites/g/files/xnrzdm211/files/styles/cgov_article/public/cgov_image/media_image/2022-04/BIRADS_Updated%20Image.png?itok=sCglGBKJ</a></p>\n<p>i really like synthetic data, and i am pretty sure that it would help alot. but. For human reporters, a mass must be seen on at least two different mammographic projections to be suspect (this does not apply for ml models though).<br>\nThe first 3 images show a rather indistinct \"blurry\" lesion where none of the circumference is well defined - this is usually consiered as suspicious. this is only one of the characteristics though a mammographeur is looking for. Suspicious masses can have other characteristics, like they can be spiculated ( like the last image of yours, however the density there is a bit high ). Other features might be seen:</p>\n<ul>\n<li><p>skin retraction</p></li>\n<li><p>nipple retraction</p></li>\n<li><p>skin thickening</p></li>\n<li><p>trabecular thickening</p></li>\n<li><p>axillary adenopathy</p></li>\n<li><p>architectural distortion  &lt;---**</p></li>\n<li><p>calcifications &lt;---***</p>\n<p>** and *** are important ones. <br>\nHowever there are benign and suspicious calcifications… diffuse ones (like speckles) are usually benign characteristic, clustered, ergional and linearly arranged ones are suspicious.<br>\nsee this vid here: <a href=\"https://youtu.be/jWogsnqwb6I?t=70\" target=\"_blank\">https://youtu.be/jWogsnqwb6I?t=70</a></p></li>\n</ul>",
              "rawMarkdown": "Hi there,\n\nfor densities take a look at this pic, you will see what the differences are:\nhttps://www.cancer.gov/sites/g/files/xnrzdm211/files/styles/cgov_article/public/cgov_image/media_image/2022-04/BIRADS_Updated%20Image.png?itok=sCglGBKJ\n\ni really like synthetic data, and i am pretty sure that it would help alot. but. For human reporters, a mass must be seen on at least two different mammographic projections to be suspect (this does not apply for ml models though).\nThe first 3 images show a rather indistinct \"blurry\" lesion where none of the circumference is well defined - this is usually consiered as suspicious. this is only one of the characteristics though a mammographeur is looking for. Suspicious masses can have other characteristics, like they can be spiculated ( like the last image of yours, however the density there is a bit high ). Other features might be seen:\n- skin retraction\n- nipple retraction\n- skin thickening\n- trabecular thickening\n- axillary adenopathy\n- architectural distortion  <---**\n- calcifications <---***\n\n  ** and *** are important ones. \nHowever there are benign and suspicious calcifications... diffuse ones (like speckles) are usually benign characteristic, clustered, ergional and linearly arranged ones are suspicious.\nsee this vid here: https://youtu.be/jWogsnqwb6I?t=70\n",
              "votes": 6
            },
            {
              "id": 2126298,
              "postDate": "2023-02-02T06:50:50.307Z",
              "content": "<p>Thank you very much <a href=\"https://www.kaggle.com/sandorkonya\" target=\"_blank\">@sandorkonya</a> for your help and explanation. I see that synthetic data is posible to create as a separate long term project :) Certainly it requires test and trial where radiologist will play the main role in this process. BTW: this video from youtube is great! 👍👍👍</p>\n<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> - look here - full res images and this paper: <a href=\"https://arxiv.org/abs/2301.13817\" target=\"_blank\">https://arxiv.org/abs/2301.13817</a></p>",
              "rawMarkdown": "Thank you very much @sandorkonya for your help and explanation. I see that synthetic data is posible to create as a separate long term project :) Certainly it requires test and trial where radiologist will play the main role in this process. BTW: this video from youtube is great! 👍👍👍\n\n@hengck23 - look here - full res images and this paper: https://arxiv.org/abs/2301.13817"
            },
            {
              "id": 2132021,
              "postDate": "2023-02-06T14:45:50.130Z",
              "content": "<p>chatGPT is trained via human feedback (like RL critic).<br>\nbut it uses something like ranking.<br>\nI am thinking that this can also be used in image recognition.<br>\n(this may be something new)</p>\n<p>we need papers that apply human in the loop for large scale data. then it would be automatic online training when people work with Ai assistant.</p>",
              "rawMarkdown": "chatGPT is trained via human feedback (like RL critic).\nbut it uses something like ranking.\nI am thinking that this can also be used in image recognition.\n(this may be something new)\n\nwe need papers that apply human in the loop for large scale data. then it would be automatic online training when people work with Ai assistant."
            },
            {
              "id": 2132033,
              "postDate": "2023-02-06T14:55:28.740Z",
              "content": "<p>Do you have a concrete proposal for this in practice?</p>",
              "rawMarkdown": "Do you have a concrete proposal for this in practice?"
            },
            {
              "id": 2132043,
              "postDate": "2023-02-06T15:02:19.677Z",
              "content": "<p>not yet. but medical AI is mainly CAD (assistant diagnostic  tool). so it can get human feedback easily. e.g. for breast mammography, CAD gives some candidates and radiologists verify them or detect new lesion.</p>\n<p>this is already human feedback.</p>\n<p>how to use this feedback to train like chatGPT is really interesting</p>",
              "rawMarkdown": "not yet. but medical AI is mainly CAD (assistant diagnostic  tool). so it can get human feedback easily. e.g. for breast mammography, CAD gives some candidates and radiologists verify them or detect new lesion.\n\nthis is already human feedback.\n\nhow to use this feedback to train like chatGPT is really interesting"
            }
          ]
        }
      ]
    },
    {
      "id": 2083939,
      "postDate": "2023-01-03T01:43:21.670Z",
      "content": "<p>one thing to note that radiologist make decision by comparing  mammography of different years apart, e.g. if there new emerging calcification, has the mass changes, etc. </p>",
      "rawMarkdown": "one thing to note that radiologist make decision by comparing  mammography of different years apart, e.g. if there new emerging calcification, has the mass changes, etc. \n",
      "votes": 2
    },
    {
      "id": 2081701,
      "postDate": "2022-12-31T14:28:48.233Z",
      "content": "<p>i would recommend reading <a href=\"https://radiologyassistant.nl/breast/bi-rads/bi-rads-for-mammography-and-ultrasound-2013\" target=\"_blank\">https://radiologyassistant.nl/breast/bi-rads/bi-rads-for-mammography-and-ultrasound-2013</a><br>\nthe part on distribution of calcifications might be of assistance</p>",
      "rawMarkdown": "i would recommend reading https://radiologyassistant.nl/breast/bi-rads/bi-rads-for-mammography-and-ultrasound-2013\nthe part on distribution of calcifications might be of assistance",
      "votes": 2
    },
    {
      "id": 2081781,
      "postDate": "2022-12-31T16:07:24.373Z",
      "content": "<p>\"Doctors or radiologists\" - I love the distinction, <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> ! 😏</p>",
      "rawMarkdown": "\"Doctors or radiologists\" - I love the distinction, @vaillant ! 😏"
    },
    {
      "id": 2130469,
      "postDate": "2023-02-05T13:24:30.720Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2081791,
      "author_name": "David Roberts",
      "author_url": "",
      "post_date": "2022-12-31T16:18:52.997000",
      "content": "<p>I am not a radiologist, but I have worked on medical imaging software for 30 years and am very familiar with mammography. I don't think it's a reasonable assumption that a radiologist would be able to diagnose cancer from a tiny density like the ones in these images.</p>\n<p>Normally, CAD will pick up the small densities and direct the radiologists attention to the area. Then, the radiologist will \"window\" the image manually to try and find the \"sweet spot\" where the density in question is separated from the rest of the tissue as much as possible in order to determine the location, shape, symmetry etc. Most of the time, a follow on ultrasound (and a biopsy) is needed to verify what the dense tissue actually is.</p>\n<p>Since these are screening mammograms (not follow ups), we don't expect to see many cases of large tumors. It seems the goal of this competition is to do something that humans cannot do .. which is to diagnose cancer from imaging alone.</p>",
      "votes": 20,
      "replies": [
        {
          "id": 2081918,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-12-31T20:15:56.937000",
          "content": "<p>thx for this answer <a href=\"https://www.kaggle.com/davidbroberts\" target=\"_blank\">@davidbroberts</a>, so much good info here 🙏 I think I learned more from your reply here on how all of this works, how the work gets done vs many of the much longer posts, really cool 🙂</p>\n<p>thank you!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2082005,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2023-01-01T01:18:40.297000",
          "content": "<p><a href=\"https://www.kaggle.com/davidbroberts\" target=\"_blank\">@davidbroberts</a> <br>\nthanks a lot for the reply.</p>\n<p>so i can assume:<br>\n[1]  the annotation process is likely to follows what you have described. even in open dataset like vindr, the process will be CAD+verification, i.e. there may be miss detection by CAD.</p>\n<p>[2]  hence one can give out of trying to perfect a segmentation model for mask annotation. instead one should devise a mask model to include all positive mask, but reduce the num of negative mask. it is ok that for one cancer image to have at least one +ve mask and several neg mask.</p>\n<p>or even in aux loss, we cannot assume to have perfect annotation for mask.<br>\nnow i understand why they use weak models like resnet22 in heatmap prediction (and not a strong one like renset50)</p>\n<p>[3]  Now i understand whay the two-stage approach: </p>\n<ul>\n<li>train an high sensitivity (maybe poor specificity) heatmap detector at high resolution.</li>\n<li>then feed the heatmap to a better specificity image model.</li>\n</ul>\n<p>the loss function is important to ensure the model is high sensitivity or specificity in the two stages.</p>\n<p>actually i also have another observation at external data:<br>\nalthough the sore for mask detection is poor (e.g. for small region), the image level prediction (e.g. BIRADS) is generally good.<br>\ni use a joint loss end-to-end model (heatmap prediction + image level prediction)</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2082041,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-01-01T03:16:59.447000",
              "content": "<p>check this !!!</p>\n<p><a href=\"https://www.youtube.com/watch?v=FJl_yf5Au68&amp;t=138s\" target=\"_blank\">https://www.youtube.com/watch?v=FJl_yf5Au68&amp;t=138s</a><br>\nHow to read a screening mammogram</p>\n<p>hence, it is not important to consider use of simple model like resnet22,34 efficientnet b2 or complicated large ones like efficientnet b7.</p>\n<p>a good workflow is more important. e.g. models for different scales, different parts, models for single image, models for multiple images of same laterality, multiple images of different laterality, models for ROI level,  models for  image levels ….</p>\n<p>an ensemble of different type of models seems to be the key.<br>\n<a href=\"https://ibb.co/Nr7J5n6\"><img src=\"https://i.ibb.co/1MbP1Jf/Selection-421.png\" alt=\"Selection-421\"></a><br>\n<a href=\"https://ibb.co/7gz7V1c\"><img src=\"https://i.ibb.co/pWPCZLD/Selection-420.png\" alt=\"Selection-420\"></a><br>\n<a href=\"https://ibb.co/g7g4hXY\"><img src=\"https://i.ibb.co/4fN13Dk/Selection-419.png\" alt=\"Selection-419\"></a></p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 2082044,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-01-01T03:28:17.150000",
              "content": "<p><img src=\"https://i.ibb.co/Z2bW7MY/Selection-425.png\" alt=\"https://i.ibb.co/Z2bW7MY/Selection-425.png\"><br>\n<img src=\"https://i.ibb.co/LRy2GjL/Selection-432.png\" alt=\"https://i.ibb.co/LRy2GjL/Selection-432.png\"></p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2083683,
              "author_name": "@kaggleqrdl",
              "author_url": "",
              "post_date": "2023-01-02T18:27:42.363000",
              "content": "<p>That's pretty interesting.  Looks like laterality alignment (left ontop of right) is a critical pre-processing activity.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2090891,
              "author_name": "TheStrugglingEngineer",
              "author_url": "",
              "post_date": "2023-01-07T18:36:26.867000",
              "content": "<blockquote>\n  <p>a good workflow is more important. e.g. models for different scales, different parts, models for single image, models for multiple images of same laterality, multiple images of different laterality, models for ROI level, models for image levels ….</p>\n</blockquote>\n<p>and just how would such multiplicity of models work in a notebook in under 9 hours on 43K high resolution images?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2091004,
              "author_name": "@kaggleqrdl",
              "author_url": "",
              "post_date": "2023-01-07T22:29:54.830000",
              "content": "<p>Yep, that's the issue.  TPU / tensorflow might be the way to go as it performs generally better than pytorch, but harder to use.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2091010,
              "author_name": "TheStrugglingEngineer",
              "author_url": "",
              "post_date": "2023-01-07T22:45:57.960000",
              "content": "<p>I agree…<br>\nmy point is, we need to be extremely stingy in our approach… simple and effective trickery is worth more than here than complex models…<br>\nwork in progress…</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2091021,
              "author_name": "@kaggleqrdl",
              "author_url": "",
              "post_date": "2023-01-07T23:08:09.170000",
              "content": "<p><a href=\"https://www.kaggle.com/houssamassila\" target=\"_blank\">@houssamassila</a> </p>\n<p>Check this out - <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/376630\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/376630</a></p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2091031,
              "author_name": "TheStrugglingEngineer",
              "author_url": "",
              "post_date": "2023-01-08T00:03:33.193000",
              "content": "<p>yes I did check that earlier today. thanks.<br>\neven further than just enlarging the usable ram, trickery regarding the dataset is also useful, like using smaller files when possible/reducing the dataset (even if it seems counter-intuitive at first). To me, this seems more the way to go than spending time trying ever more extravaggant models that won't run in 9 hours.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2091061,
              "author_name": "@kaggleqrdl",
              "author_url": "",
              "post_date": "2023-01-08T01:13:37.720000",
              "content": "<p>Hmm, someone is saying we can't use TPU for this comp.  Which is weird, given that someone created a TPU notebook.   But it might be useful for training.</p>\n<p>edit:  reading up on it, I think it's because the TPUs require internet access which is not allowed in this comp</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2091353,
              "author_name": "TheStrugglingEngineer",
              "author_url": "",
              "post_date": "2023-01-08T10:19:10.467000",
              "content": "<p>it's as you said, TPU will be useful only to train models that can later be loaded into the submission notebook.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2094507,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-01-10T19:39:39.123000",
              "content": "<p><a href=\"https://www.kaggle.com/houssamassila\" target=\"_blank\">@houssamassila</a> <br>\n\"just how would such multiplicity of models work in a notebook in under 9 hours on 43K high resolution images?\"</p>\n<p>\"early rejection\" is a common trick in computer vision. 10x speed up is not problem</p>\n<p><a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/377008#2094477\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/377008#2094477</a></p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2125004,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2023-02-01T11:12:57.127000",
      "content": "<p>I see you are looking for a radiologist to team up <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>. I worked with <a href=\"https://www.kaggle.com/sandorkonya\" target=\"_blank\">@sandorkonya</a> before and he is the best one I've ever worked with. Tagged both of you for potential collab.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2125090,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2023-02-01T12:14:31.927000",
          "content": "<p>The idea is simple.<br>\nWith my top models, i will generate CAM heatmap.<br>\nThen the radiology correct it. (Delete very sure wrong heat map for activated case, mark possible one for missed cases)</p>\n<p>i do not need very accurate annotation. one can simply just annotate a point.<br>\nI do not need all annotation, just at least one per pos image.</p>\n<p>only 2% of train  data are pos. this is very few images. annotation can done within hours (because i do not need good annotation).</p>\n<p>Basically it is manual active learning. we will go through 2~3 rounds.</p>\n<p>Let's see if this will fight imbalance data and gives better generalization, rather than relying on unstable LB/CV  </p>\n<hr>\n<p><a href=\"https://ojs.aaai.org/index.php/AAAI/article/download/16153/15960\" target=\"_blank\">https://ojs.aaai.org/index.php/AAAI/article/download/16153/15960</a><br>\nfigure.2 is example of point annotation</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2132386,
              "author_name": "Anil Thomas",
              "author_url": "",
              "post_date": "2023-02-06T19:23:12",
              "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, you probably know this, but the <a href=\"https://www.physionet.org/content/vindr-mammo/1.0.0/\" target=\"_blank\">vindr-mammo</a> dataset already has the bounding box annotations that you are looking for. It would be interesting to see how much juice you can squeeze out of those annotations - gives you a quick way to assess if it's worth getting the RSNA dataset annotated by a radiologist…</p>\n<p>My intuition is that it would help a lot. If you can show a model the ROI that causes a mammogram to be tagged as positive, the model is likely to learn the important features better. I explored this a bit, but haven't found a good way to incorporate the bounding box data. One possible method is to divide the images into tiles, label each tile and then train on the tiles and their labels. The hard part, at least for me is aggregating the tile-level predictions into a single prediction for the whole image. A global maxpooling operation doesn't seem to work well.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2125687,
          "author_name": "dr. Konya",
          "author_url": "",
          "post_date": "2023-02-01T19:56:41.807000",
          "content": "<p><a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> - thank you for the generous words i barely deserve it though.<br>\nI read through <em>some</em> of the discussions of the contest.</p>\n<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> are you using the dataset from <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/377790\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/377790</a> already? Those images have a region annotation and you can do a point annotation from that pretty easily. The mammography technique is quite standardised (in contrary to other projectionsradiographic material). I see that there are labels on the above set which could misguide the classification… so i would mask out the mamma on all of the images (of the official and this one above) and just add black bg to all as a preprocessing step. That area does not contain any valuable information regarding the tumor anyway. The fine-grained structure of the organ on the image is very similar!<br>\nI hope you are already working on full resolution (patches, sliding windows ) because you have only chance to detect the suspect microcalcifications on (near?) full res.</p>\n<p>Dense (fibrous and glandular) breast tissue looks \"white\" on a mammogram. Breast masses and cancers can also look \"white\", so the classification in this category is the hardest task. These breasts have usually an uncertain BIRAD score! This is why you probably see some patients having more than 4 images, especially those with fibrous tissue. The extra images are focused, zoomed images or repeated images due to false compression.</p>\n<p>The problem with annotating mammography is that the annotation itself doesn't take long (one or 4 points doesn't matter) … but fiding the lesion! =) Even annotating 1.1k images (2% positive of 55k) would take a substantial time. <br>\nWe could try some to see how it affects the generalisation… i have a reporting monitor <a href=\"https://www.kaggle.com/home\" target=\"_blank\">@home</a> for good contrast.</p>",
          "votes": 4,
          "replies": [
            {
              "id": 2125703,
              "author_name": "Remek Kinas",
              "author_url": "",
              "post_date": "2023-02-01T20:24:05.240000",
              "content": "<p>Thank you very much for explanation. Is any chace you show us the example of dense (fibrous and glandular) breast tissue,  breast masses and cancers? </p>\n<p>a. It is hard for me do see these issues on images.<br>\nb. second - I made prototype for synthetic breast cancer generator (but I am not able to confirm if it works - it was made by me for fun - to see if it is possible to generate such changes on images. I am aware that it is far from \"good\" issue we should classify) </p>\n<p><img src=\"http://zapodaj.net/images/6608d09e85fec.jpg\" alt=\"image_1\"></p>\n<p><img src=\"http://zapodaj.net/images/a896deacf2e13.jpg\" alt=\"image_2\"></p>\n<p><img src=\"http://zapodaj.net/images/9131c16081453.jpg\" alt=\"image_3\"></p>\n<p><img src=\"http://zapodaj.net/images/002e5863d493a.jpg\" alt=\"image_4\"></p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2125805,
              "author_name": "dr. Konya",
              "author_url": "",
              "post_date": "2023-02-01T22:45:37.867000",
              "content": "<p>Hi there,</p>\n<p>for densities take a look at this pic, you will see what the differences are:<br>\n<a href=\"https://www.cancer.gov/sites/g/files/xnrzdm211/files/styles/cgov_article/public/cgov_image/media_image/2022-04/BIRADS_Updated%20Image.png?itok=sCglGBKJ\" target=\"_blank\">https://www.cancer.gov/sites/g/files/xnrzdm211/files/styles/cgov_article/public/cgov_image/media_image/2022-04/BIRADS_Updated%20Image.png?itok=sCglGBKJ</a></p>\n<p>i really like synthetic data, and i am pretty sure that it would help alot. but. For human reporters, a mass must be seen on at least two different mammographic projections to be suspect (this does not apply for ml models though).<br>\nThe first 3 images show a rather indistinct \"blurry\" lesion where none of the circumference is well defined - this is usually consiered as suspicious. this is only one of the characteristics though a mammographeur is looking for. Suspicious masses can have other characteristics, like they can be spiculated ( like the last image of yours, however the density there is a bit high ). Other features might be seen:</p>\n<ul>\n<li><p>skin retraction</p></li>\n<li><p>nipple retraction</p></li>\n<li><p>skin thickening</p></li>\n<li><p>trabecular thickening</p></li>\n<li><p>axillary adenopathy</p></li>\n<li><p>architectural distortion  &lt;---**</p></li>\n<li><p>calcifications &lt;---***</p>\n<p>** and *** are important ones. <br>\nHowever there are benign and suspicious calcifications… diffuse ones (like speckles) are usually benign characteristic, clustered, ergional and linearly arranged ones are suspicious.<br>\nsee this vid here: <a href=\"https://youtu.be/jWogsnqwb6I?t=70\" target=\"_blank\">https://youtu.be/jWogsnqwb6I?t=70</a></p></li>\n</ul>",
              "votes": 6,
              "replies": []
            },
            {
              "id": 2126298,
              "author_name": "Remek Kinas",
              "author_url": "",
              "post_date": "2023-02-02T06:50:50.307000",
              "content": "<p>Thank you very much <a href=\"https://www.kaggle.com/sandorkonya\" target=\"_blank\">@sandorkonya</a> for your help and explanation. I see that synthetic data is posible to create as a separate long term project :) Certainly it requires test and trial where radiologist will play the main role in this process. BTW: this video from youtube is great! 👍👍👍</p>\n<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> - look here - full res images and this paper: <a href=\"https://arxiv.org/abs/2301.13817\" target=\"_blank\">https://arxiv.org/abs/2301.13817</a></p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2132021,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-02-06T14:45:50.130000",
              "content": "<p>chatGPT is trained via human feedback (like RL critic).<br>\nbut it uses something like ranking.<br>\nI am thinking that this can also be used in image recognition.<br>\n(this may be something new)</p>\n<p>we need papers that apply human in the loop for large scale data. then it would be automatic online training when people work with Ai assistant.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2132033,
              "author_name": "dr. Konya",
              "author_url": "",
              "post_date": "2023-02-06T14:55:28.740000",
              "content": "<p>Do you have a concrete proposal for this in practice?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2132043,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-02-06T15:02:19.677000",
              "content": "<p>not yet. but medical AI is mainly CAD (assistant diagnostic  tool). so it can get human feedback easily. e.g. for breast mammography, CAD gives some candidates and radiologists verify them or detect new lesion.</p>\n<p>this is already human feedback.</p>\n<p>how to use this feedback to train like chatGPT is really interesting</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2083939,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-01-03T01:43:21.670000",
      "content": "<p>one thing to note that radiologist make decision by comparing  mammography of different years apart, e.g. if there new emerging calcification, has the mass changes, etc. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2081701,
      "author_name": "Redicee",
      "author_url": "",
      "post_date": "2022-12-31T14:28:48.233000",
      "content": "<p>i would recommend reading <a href=\"https://radiologyassistant.nl/breast/bi-rads/bi-rads-for-mammography-and-ultrasound-2013\" target=\"_blank\">https://radiologyassistant.nl/breast/bi-rads/bi-rads-for-mammography-and-ultrasound-2013</a><br>\nthe part on distribution of calcifications might be of assistance</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2081781,
      "author_name": "James Howard",
      "author_url": "",
      "post_date": "2022-12-31T16:07:24.373000",
      "content": "<p>\"Doctors or radiologists\" - I love the distinction, <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> ! 😏</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2130469,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-02-05T13:24:30.720000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2081540": "results on validation:\n![https://i.ibb.co/vXx9MX3/Selection-401.png](https://i.ibb.co/vXx9MX3/Selection-401.png)\n![https://i.ibb.co/L00DzmW/Selection-400.png](https://i.ibb.co/L00DzmW/Selection-400.png)\n\nQuestion : How is very small lesions detected? i can find many similar regions that are false positives. How do you look at the whole image to decide if it is cancerous or not? Or do you really can decide just based on local context region?\n\n(basically i am trying to decide the context for good detection but it seems impossible. It is even difficult on training set)",
    "2081791": "I am not a radiologist, but I have worked on medical imaging software for 30 years and am very familiar with mammography. I don't think it's a reasonable assumption that a radiologist would be able to diagnose cancer from a tiny density like the ones in these images.\n\nNormally, CAD will pick up the small densities and direct the radiologists attention to the area. Then, the radiologist will \"window\" the image manually to try and find the \"sweet spot\" where the density in question is separated from the rest of the tissue as much as possible in order to determine the location, shape, symmetry etc. Most of the time, a follow on ultrasound (and a biopsy) is needed to verify what the dense tissue actually is.\n\nSince these are screening mammograms (not follow ups), we don't expect to see many cases of large tumors. It seems the goal of this competition is to do something that humans cannot do .. which is to diagnose cancer from imaging alone.",
    "2125004": "I see you are looking for a radiologist to team up @hengck23. I worked with @sandorkonya before and he is the best one I've ever worked with. Tagged both of you for potential collab.",
    "2083939": "one thing to note that radiologist make decision by comparing  mammography of different years apart, e.g. if there new emerging calcification, has the mass changes, etc. \n",
    "2081701": "i would recommend reading https://radiologyassistant.nl/breast/bi-rads/bi-rads-for-mammography-and-ultrasound-2013\nthe part on distribution of calcifications might be of assistance",
    "2081781": "\"Doctors or radiologists\" - I love the distinction, @vaillant ! 😏",
    "2130469": ""
  }
}