{
  "id": 375106,
  "title": "Leaderboard results so far somewhat off from literature / SOTA",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/375106",
  "author_name": "@kaggleqrdl",
  "post_date": "2022-12-30T09:52:36.666000",
  "votes": 2,
  "comment_count": 4,
  "views": 0,
  "content": "<p>chatgpt:  \"Some AI models have been able to achieve very high levels of accuracy in detecting breast cancer, with AUC scores approaching or exceeding 0.9.\"</p>\n<p>This compares to what I've been seeing in the 2022 papers I've been reviewing as well.  </p>\n<p>We have a lot of images (~32K) to infer on in a short time window, which may partly be the reason for lower numbers.</p>\n<p>A question I've been pondering, what will the winning solution look like?  One possibility is that it will be a Dali driven solution such that everything, or at least at much as possible is done on the GPU, including augmentation.  This will allow for as many models / and augmentation / analysis as possible.</p>\n<p>An example is wavelet analysis which is quite slow on the CPU, but blazingly fast on the GPU.  There is quite a lot of recent papers talking about using wavelets when detecting breast cancer.</p>\n<p>Also curious if the multi T4x2 option will outperform the single P100.</p>\n<p>Decoding will unfortunately still be a bottleneck, and fastest I've been able to <a href=\"https://www.kaggle.com/code/kaggleqrdl/baseline-32000-nvjpeg2000-dicomsdl\" target=\"_blank\">prove out</a> is 4400 seconds for 16K jpeg2000 and 16K lossless.  It'd be a bit faster if the jpeg2000 was done on the GPU via dali, but the lossless via dicomsdl will still be a bottleneck.</p>",
  "messages": [
    {
      "id": 2080671,
      "postDate": "2022-12-30T11:17:24.590Z",
      "content": "<p>There are many AI Mammography products.<br>\nI wonder if anyone wants to do this:</p>\n<ol>\n<li>write to some of the vendors and request for demo</li>\n<li>label the kaggle datset with the demo software (segmentation)</li>\n<li>train a model with the label and submit the mode for kaggle test</li>\n</ol>\n<p>while this submission cannot be granted for the prize but it is a benchmarking for the product.<br>\nit can be good (or bad) publicity for the vendor if the results got top rank.</p>\n<hr>\n<p>i think commercial products are trained with millions of images. it is easy to get a large dataset if you are working with a hospital. I do believe that the accuracy of commercial system are not too bad</p>",
      "rawMarkdown": "There are many AI Mammography products.\nI wonder if anyone wants to do this:\n1. write to some of the vendors and request for demo\n2. label the kaggle datset with the demo software (segmentation)\n3. train a model with the label and submit the mode for kaggle test\n\nwhile this submission cannot be granted for the prize but it is a benchmarking for the product.\nit can be good (or bad) publicity for the vendor if the results got top rank.\n\n---\n\ni think commercial products are trained with millions of images. it is easy to get a large dataset if you are working with a hospital. I do believe that the accuracy of commercial system are not too bad",
      "votes": 2,
      "replies": [
        {
          "id": 2080672,
          "postDate": "2022-12-30T11:20:17.393Z",
          "content": "<p>Here's one paper - <a href=\"https://hajim.rochester.edu/ece/sites/parker/assets/pdf/243-diagnostic-performance-of-an-artificial-intelligence-system-in-breast-ultrasound.pdf\" target=\"_blank\">https://hajim.rochester.edu/ece/sites/parker/assets/pdf/243-diagnostic-performance-of-an-artificial-intelligence-system-in-breast-ultrasound.pdf</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9052057%2F056a5deaf7c79b377c685273a33a1bcd%2Fresults.png?generation=1672399122104026&amp;alt=media\" alt=\"\"></p>\n<p>S-Detect doesn't threshold its results, it's 'dichotomous', which I think explains the lack of AUC #.  Note this study has specific population constraints, which don't necessarily generalize immediately to this contest.</p>\n<p>It's interesting to see how junior radiologists outperform senior ones in accuracy and specificity but not sensitivity.  As I mentioned on another thread, I believe it's because senior ones are trying to prioritize / overcome the false negative problem, but there might be other explanations.</p>\n<p>In terms of getting a commercial product to submit, I'm skeptical that they would make the 1-2s time window required for this comp, and that's with lossless decoding required.  That time constraint really doesn't make sense when detecting cancer for a commercial product, especially given the high rate of FP/FN in BC scanning.   This is assuming you can get past the inevitable and very likely insurmountable IP protection issues.</p>",
          "rawMarkdown": "Here's one paper - https://hajim.rochester.edu/ece/sites/parker/assets/pdf/243-diagnostic-performance-of-an-artificial-intelligence-system-in-breast-ultrasound.pdf\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9052057%2F056a5deaf7c79b377c685273a33a1bcd%2Fresults.png?generation=1672399122104026&alt=media)\n\nS-Detect doesn't threshold its results, it's 'dichotomous', which I think explains the lack of AUC #.  Note this study has specific population constraints, which don't necessarily generalize immediately to this contest.\n\nIt's interesting to see how junior radiologists outperform senior ones in accuracy and specificity but not sensitivity.  As I mentioned on another thread, I believe it's because senior ones are trying to prioritize / overcome the false negative problem, but there might be other explanations.\n\nIn terms of getting a commercial product to submit, I'm skeptical that they would make the 1-2s time window required for this comp, and that's with lossless decoding required.  That time constraint really doesn't make sense when detecting cancer for a commercial product, especially given the high rate of FP/FN in BC scanning.   This is assuming you can get past the inevitable and very likely insurmountable IP protection issues."
        }
      ]
    },
    {
      "id": 2080600,
      "postDate": "2022-12-30T09:52:36.667Z",
      "content": "<p>chatgpt:  \"Some AI models have been able to achieve very high levels of accuracy in detecting breast cancer, with AUC scores approaching or exceeding 0.9.\"</p>\n<p>This compares to what I've been seeing in the 2022 papers I've been reviewing as well.  </p>\n<p>We have a lot of images (~32K) to infer on in a short time window, which may partly be the reason for lower numbers.</p>\n<p>A question I've been pondering, what will the winning solution look like?  One possibility is that it will be a Dali driven solution such that everything, or at least at much as possible is done on the GPU, including augmentation.  This will allow for as many models / and augmentation / analysis as possible.</p>\n<p>An example is wavelet analysis which is quite slow on the CPU, but blazingly fast on the GPU.  There is quite a lot of recent papers talking about using wavelets when detecting breast cancer.</p>\n<p>Also curious if the multi T4x2 option will outperform the single P100.</p>\n<p>Decoding will unfortunately still be a bottleneck, and fastest I've been able to <a href=\"https://www.kaggle.com/code/kaggleqrdl/baseline-32000-nvjpeg2000-dicomsdl\" target=\"_blank\">prove out</a> is 4400 seconds for 16K jpeg2000 and 16K lossless.  It'd be a bit faster if the jpeg2000 was done on the GPU via dali, but the lossless via dicomsdl will still be a bottleneck.</p>",
      "rawMarkdown": "chatgpt:  \"Some AI models have been able to achieve very high levels of accuracy in detecting breast cancer, with AUC scores approaching or exceeding 0.9.\"\n\nThis compares to what I've been seeing in the 2022 papers I've been reviewing as well.  \n\nWe have a lot of images (~32K) to infer on in a short time window, which may partly be the reason for lower numbers.\n\nA question I've been pondering, what will the winning solution look like?  One possibility is that it will be a Dali driven solution such that everything, or at least at much as possible is done on the GPU, including augmentation.  This will allow for as many models / and augmentation / analysis as possible.\n\nAn example is wavelet analysis which is quite slow on the CPU, but blazingly fast on the GPU.  There is quite a lot of recent papers talking about using wavelets when detecting breast cancer.\n\nAlso curious if the multi T4x2 option will outperform the single P100.\n\nDecoding will unfortunately still be a bottleneck, and fastest I've been able to [prove out](https://www.kaggle.com/code/kaggleqrdl/baseline-32000-nvjpeg2000-dicomsdl) is 4400 seconds for 16K jpeg2000 and 16K lossless.  It'd be a bit faster if the jpeg2000 was done on the GPU via dali, but the lossless via dicomsdl will still be a bottleneck.\n\n",
      "votes": 2
    },
    {
      "id": 2080644,
      "postDate": "2022-12-30T10:42:11.423Z",
      "content": "<p>Decoding is no problem at all now. It can be bottleneck…. in future. Now I think problem is to find way to learn model or … find best solution (blended classifier for this task is decent solution). </p>",
      "rawMarkdown": "Decoding is no problem at all now. It can be bottleneck…. in future. Now I think problem is to find way to learn model or … find best solution (blended classifier for this task is decent solution). ",
      "replies": [
        {
          "id": 2080653,
          "postDate": "2022-12-30T10:49:10.633Z",
          "content": "<p>Maybe.  Most of the approaches I've been prototyping go over time and don't make sense submitting.  Having more time would help.</p>\n<p>What is the best time for decoding that you're seeing?  As I mentioned above, the best I can do and publicly shared is 4400 seconds for 16K lossless + 16K jpeg2000. </p>\n<p>It all has to go to CPU though in the intermediate, as the techniques I'm interested in / aware of require CPU manipulation before inference on GPU.</p>\n<p>Regardless, I suspect most folks are going to be hitting the ceiling at some point due to the large # of images and squeezing out extra time somewhere, either in decoding, augmentation, or in inference is going to be required.   For me, lossless decoding is a fixed bottleneck as I don't see anyway possible to reduce that time, besides optimizing decoding libraries such as turbojpeg.  The underlying huffman encoding on lossless makes GPU decoding impossible, I believe.</p>",
          "rawMarkdown": "Maybe.  Most of the approaches I've been prototyping go over time and don't make sense submitting.  Having more time would help.\n\nWhat is the best time for decoding that you're seeing?  As I mentioned above, the best I can do and publicly shared is 4400 seconds for 16K lossless + 16K jpeg2000. \n\nIt all has to go to CPU though in the intermediate, as the techniques I'm interested in / aware of require CPU manipulation before inference on GPU.\n\nRegardless, I suspect most folks are going to be hitting the ceiling at some point due to the large # of images and squeezing out extra time somewhere, either in decoding, augmentation, or in inference is going to be required.   For me, lossless decoding is a fixed bottleneck as I don't see anyway possible to reduce that time, besides optimizing decoding libraries such as turbojpeg.  The underlying huffman encoding on lossless makes GPU decoding impossible, I believe."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2080671,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-12-30T11:17:24.590000",
      "content": "<p>There are many AI Mammography products.<br>\nI wonder if anyone wants to do this:</p>\n<ol>\n<li>write to some of the vendors and request for demo</li>\n<li>label the kaggle datset with the demo software (segmentation)</li>\n<li>train a model with the label and submit the mode for kaggle test</li>\n</ol>\n<p>while this submission cannot be granted for the prize but it is a benchmarking for the product.<br>\nit can be good (or bad) publicity for the vendor if the results got top rank.</p>\n<hr>\n<p>i think commercial products are trained with millions of images. it is easy to get a large dataset if you are working with a hospital. I do believe that the accuracy of commercial system are not too bad</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2080672,
          "author_name": "@kaggleqrdl",
          "author_url": "",
          "post_date": "2022-12-30T11:20:17.393000",
          "content": "<p>Here's one paper - <a href=\"https://hajim.rochester.edu/ece/sites/parker/assets/pdf/243-diagnostic-performance-of-an-artificial-intelligence-system-in-breast-ultrasound.pdf\" target=\"_blank\">https://hajim.rochester.edu/ece/sites/parker/assets/pdf/243-diagnostic-performance-of-an-artificial-intelligence-system-in-breast-ultrasound.pdf</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9052057%2F056a5deaf7c79b377c685273a33a1bcd%2Fresults.png?generation=1672399122104026&amp;alt=media\" alt=\"\"></p>\n<p>S-Detect doesn't threshold its results, it's 'dichotomous', which I think explains the lack of AUC #.  Note this study has specific population constraints, which don't necessarily generalize immediately to this contest.</p>\n<p>It's interesting to see how junior radiologists outperform senior ones in accuracy and specificity but not sensitivity.  As I mentioned on another thread, I believe it's because senior ones are trying to prioritize / overcome the false negative problem, but there might be other explanations.</p>\n<p>In terms of getting a commercial product to submit, I'm skeptical that they would make the 1-2s time window required for this comp, and that's with lossless decoding required.  That time constraint really doesn't make sense when detecting cancer for a commercial product, especially given the high rate of FP/FN in BC scanning.   This is assuming you can get past the inevitable and very likely insurmountable IP protection issues.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2080644,
      "author_name": "Remek Kinas",
      "author_url": "",
      "post_date": "2022-12-30T10:42:11.423000",
      "content": "<p>Decoding is no problem at all now. It can be bottleneck…. in future. Now I think problem is to find way to learn model or … find best solution (blended classifier for this task is decent solution). </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2080653,
          "author_name": "@kaggleqrdl",
          "author_url": "",
          "post_date": "2022-12-30T10:49:10.633000",
          "content": "<p>Maybe.  Most of the approaches I've been prototyping go over time and don't make sense submitting.  Having more time would help.</p>\n<p>What is the best time for decoding that you're seeing?  As I mentioned above, the best I can do and publicly shared is 4400 seconds for 16K lossless + 16K jpeg2000. </p>\n<p>It all has to go to CPU though in the intermediate, as the techniques I'm interested in / aware of require CPU manipulation before inference on GPU.</p>\n<p>Regardless, I suspect most folks are going to be hitting the ceiling at some point due to the large # of images and squeezing out extra time somewhere, either in decoding, augmentation, or in inference is going to be required.   For me, lossless decoding is a fixed bottleneck as I don't see anyway possible to reduce that time, besides optimizing decoding libraries such as turbojpeg.  The underlying huffman encoding on lossless makes GPU decoding impossible, I believe.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2080671": "There are many AI Mammography products.\nI wonder if anyone wants to do this:\n1. write to some of the vendors and request for demo\n2. label the kaggle datset with the demo software (segmentation)\n3. train a model with the label and submit the mode for kaggle test\n\nwhile this submission cannot be granted for the prize but it is a benchmarking for the product.\nit can be good (or bad) publicity for the vendor if the results got top rank.\n\n---\n\ni think commercial products are trained with millions of images. it is easy to get a large dataset if you are working with a hospital. I do believe that the accuracy of commercial system are not too bad",
    "2080600": "chatgpt:  \"Some AI models have been able to achieve very high levels of accuracy in detecting breast cancer, with AUC scores approaching or exceeding 0.9.\"\n\nThis compares to what I've been seeing in the 2022 papers I've been reviewing as well.  \n\nWe have a lot of images (~32K) to infer on in a short time window, which may partly be the reason for lower numbers.\n\nA question I've been pondering, what will the winning solution look like?  One possibility is that it will be a Dali driven solution such that everything, or at least at much as possible is done on the GPU, including augmentation.  This will allow for as many models / and augmentation / analysis as possible.\n\nAn example is wavelet analysis which is quite slow on the CPU, but blazingly fast on the GPU.  There is quite a lot of recent papers talking about using wavelets when detecting breast cancer.\n\nAlso curious if the multi T4x2 option will outperform the single P100.\n\nDecoding will unfortunately still be a bottleneck, and fastest I've been able to [prove out](https://www.kaggle.com/code/kaggleqrdl/baseline-32000-nvjpeg2000-dicomsdl) is 4400 seconds for 16K jpeg2000 and 16K lossless.  It'd be a bit faster if the jpeg2000 was done on the GPU via dali, but the lossless via dicomsdl will still be a bottleneck.\n\n",
    "2080644": "Decoding is no problem at all now. It can be bottleneck…. in future. Now I think problem is to find way to learn model or … find best solution (blended classifier for this task is decent solution). "
  }
}