{
  "id": 369282,
  "title": "📸 DICOM files converted to PNGs [314.72 GB -> 921 MB] ",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/369282",
  "author_name": "Radek Osmulski",
  "post_date": "2022-11-29T15:03:57.618000",
  "votes": 118,
  "comment_count": 34,
  "views": 0,
  "content": "<p>Hey!</p>\n<p>Please find the DICOM files converted to PNG here: <a href=\"https://www.kaggle.com/datasets/radek1/rsna-mammography-images-as-pngs/settings\" target=\"_blank\">RSNA Mammography -- images as PNGs (256px x 256px)</a></p>\n<p><strong>EDIT: Based on a request by <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> I created a new version of the dataset that also contains 512px x 512px images!</strong></p>\n<p>I tried converting the files on Kaggle, but the VM was just not enough 🙂 Even though the processing completed after more than ~3 hrs using the 4 available cores, I still didn't manage to commit the output.</p>\n<p>Opted to go for a GCP VM to get the job done 🙂</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2Fdbc2480e7f8863133c72cc37cc9f6b81%2Fgcp_vm_at_work.png?generation=1669731617602415&amp;alt=media\" alt=\"\"></p>\n<p>On GCP the conversion took just under 17 minutes using 64 cores and multiprocessing. Please note, I converted the images in a way that accounts for <code>dicom.PhotometricInterpretation == \"MONOCHROME1\"</code> where the intensities need to be inverted.</p>\n<p>I converted DICOM files to PNG images 256px by 256px. The folder structure is the same as of the DICOM files.</p>\n<p>Here are a couple of examples:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F4fc6e1f8837c3be387dd45077bf7aa7f%2F2057295788.png?generation=1669735207946955&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2Ff0bd8462a9d27d8049185f6039d3536e%2F2046475482.png?generation=1669735227410985&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2Ff155f4f11b2e5d8e836423bf5f30c7a5%2F349510516.png?generation=1669735239725229&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2Ff1d5dd2efa0bbbd68d061f9aad6cd868%2F1365269360.png?generation=1669735252065645&amp;alt=media\" alt=\"\"></p>\n<p>Happy Kaggling! 🙂</p>\n<h3>Other resources you might find useful:</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/how-to-process-dicom-images-to-pngs\" target=\"_blank\">💡 how to process DICOM images to PNGs</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/eda-training-a-fast-ai-model-submission\" target=\"_blank\">📊 EDA + training a fast.ai model + submission 🚀</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/fast-ai-starter-pack-train-inference\" target=\"_blank\">🤖 [fast.ai starter pack] train + inference 🚀</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/371268\" target=\"_blank\">📸 Over 56GB of processed data, 5 different methods 🥳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369706\" target=\"_blank\">3 resources to get started with Computer Vision in this competition 🚀🚀🚀</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369155\" target=\"_blank\">💡 6 Computer Vision tricks for faster training and better models 🚀🚀🚀</a></li>\n</ul>",
  "messages": [
    {
      "id": 2048609,
      "postDate": "2022-11-29T15:03:57.620Z",
      "content": "<p>Hey!</p>\n<p>Please find the DICOM files converted to PNG here: <a href=\"https://www.kaggle.com/datasets/radek1/rsna-mammography-images-as-pngs/settings\" target=\"_blank\">RSNA Mammography -- images as PNGs (256px x 256px)</a></p>\n<p><strong>EDIT: Based on a request by <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> I created a new version of the dataset that also contains 512px x 512px images!</strong></p>\n<p>I tried converting the files on Kaggle, but the VM was just not enough 🙂 Even though the processing completed after more than ~3 hrs using the 4 available cores, I still didn't manage to commit the output.</p>\n<p>Opted to go for a GCP VM to get the job done 🙂</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2Fdbc2480e7f8863133c72cc37cc9f6b81%2Fgcp_vm_at_work.png?generation=1669731617602415&amp;alt=media\" alt=\"\"></p>\n<p>On GCP the conversion took just under 17 minutes using 64 cores and multiprocessing. Please note, I converted the images in a way that accounts for <code>dicom.PhotometricInterpretation == \"MONOCHROME1\"</code> where the intensities need to be inverted.</p>\n<p>I converted DICOM files to PNG images 256px by 256px. The folder structure is the same as of the DICOM files.</p>\n<p>Here are a couple of examples:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F4fc6e1f8837c3be387dd45077bf7aa7f%2F2057295788.png?generation=1669735207946955&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2Ff0bd8462a9d27d8049185f6039d3536e%2F2046475482.png?generation=1669735227410985&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2Ff155f4f11b2e5d8e836423bf5f30c7a5%2F349510516.png?generation=1669735239725229&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2Ff1d5dd2efa0bbbd68d061f9aad6cd868%2F1365269360.png?generation=1669735252065645&amp;alt=media\" alt=\"\"></p>\n<p>Happy Kaggling! 🙂</p>\n<h3>Other resources you might find useful:</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/how-to-process-dicom-images-to-pngs\" target=\"_blank\">💡 how to process DICOM images to PNGs</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/eda-training-a-fast-ai-model-submission\" target=\"_blank\">📊 EDA + training a fast.ai model + submission 🚀</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/fast-ai-starter-pack-train-inference\" target=\"_blank\">🤖 [fast.ai starter pack] train + inference 🚀</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/371268\" target=\"_blank\">📸 Over 56GB of processed data, 5 different methods 🥳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369706\" target=\"_blank\">3 resources to get started with Computer Vision in this competition 🚀🚀🚀</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369155\" target=\"_blank\">💡 6 Computer Vision tricks for faster training and better models 🚀🚀🚀</a></li>\n</ul>",
      "rawMarkdown": "Hey!\n\nPlease find the DICOM files converted to PNG here: [RSNA Mammography -- images as PNGs (256px x 256px)](https://www.kaggle.com/datasets/radek1/rsna-mammography-images-as-pngs/settings)\n\n**EDIT: Based on a request by @remekkinas I created a new version of the dataset that also contains 512px x 512px images!**\n\n\nI tried converting the files on Kaggle, but the VM was just not enough 🙂 Even though the processing completed after more than ~3 hrs using the 4 available cores, I still didn't manage to commit the output.\n\nOpted to go for a GCP VM to get the job done 🙂\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2Fdbc2480e7f8863133c72cc37cc9f6b81%2Fgcp_vm_at_work.png?generation=1669731617602415&alt=media)\n\nOn GCP the conversion took just under 17 minutes using 64 cores and multiprocessing. Please note, I converted the images in a way that accounts for `dicom.PhotometricInterpretation == \"MONOCHROME1\"` where the intensities need to be inverted.\n\nI converted DICOM files to PNG images 256px by 256px. The folder structure is the same as of the DICOM files.\n\nHere are a couple of examples:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F4fc6e1f8837c3be387dd45077bf7aa7f%2F2057295788.png?generation=1669735207946955&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2Ff0bd8462a9d27d8049185f6039d3536e%2F2046475482.png?generation=1669735227410985&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2Ff155f4f11b2e5d8e836423bf5f30c7a5%2F349510516.png?generation=1669735239725229&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2Ff1d5dd2efa0bbbd68d061f9aad6cd868%2F1365269360.png?generation=1669735252065645&alt=media)\n\nHappy Kaggling! 🙂\n\n### Other resources you might find useful:\n\n* [💡 how to process DICOM images to PNGs](https://www.kaggle.com/code/radek1/how-to-process-dicom-images-to-pngs)\n* [📊 EDA + training a fast.ai model + submission 🚀](https://www.kaggle.com/code/radek1/eda-training-a-fast-ai-model-submission)\n* [🤖 [fast.ai starter pack] train + inference 🚀](https://www.kaggle.com/code/radek1/fast-ai-starter-pack-train-inference)\n* [📸 Over 56GB of processed data, 5 different methods 🥳](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/371268)\n* [3 resources to get started with Computer Vision in this competition 🚀🚀🚀](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369706)\n* [💡 6 Computer Vision tricks for faster training and better models 🚀🚀🚀](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369155)\n\n",
      "votes": 117
    },
    {
      "id": 2056098,
      "postDate": "2022-12-05T18:20:02.083Z",
      "content": "<p>Exercise caution when converting from DICOM to PNG. Mammograms are very high resolution (12MP on average) and downsampling heavily may render small findings invisible. 8-bit is sufficient for most cases but 12 and 16-bit should be considered as well. Run small experiments to understand the importance.</p>\n<p>Cropping the mamogram to remove background tissue is a good strategy but consider the implications of padding to achieve the same image size rather than distorting the aspect ratio. This may change the appearance of fine calcifications.</p>\n<p>Good luck!</p>",
      "rawMarkdown": "Exercise caution when converting from DICOM to PNG. Mammograms are very high resolution (12MP on average) and downsampling heavily may render small findings invisible. 8-bit is sufficient for most cases but 12 and 16-bit should be considered as well. Run small experiments to understand the importance.\n\nCropping the mamogram to remove background tissue is a good strategy but consider the implications of padding to achieve the same image size rather than distorting the aspect ratio. This may change the appearance of fine calcifications.\n\nGood luck!",
      "votes": 18
    },
    {
      "id": 2056988,
      "postDate": "2022-12-06T16:33:12.897Z",
      "content": "<p>As this information might get unnoticed, if you have been using this dataset, please be aware of:</p>\n<p><a href=\"https://www.kaggle.com/code/radek1/how-to-process-dicom-images-to-pngs/comments#2056978\" target=\"_blank\">https://www.kaggle.com/code/radek1/how-to-process-dicom-images-to-pngs/comments#2056978</a></p>",
      "rawMarkdown": "As this information might get unnoticed, if you have been using this dataset, please be aware of:\n\nhttps://www.kaggle.com/code/radek1/how-to-process-dicom-images-to-pngs/comments#2056978",
      "votes": 3
    },
    {
      "id": 2048805,
      "postDate": "2022-11-29T17:28:04.683Z",
      "content": "<p>Nic work Radek! I will use your work. <br>\nI thnink that in case of 256x256 … we should create more informative dataset - remove background (additional pixels). Look on second image and last one - we can crop it a little bit. Anyway … it could be done in augumentation pipeline as well so do not need any work here.</p>\n<p>Radek it would be cool if you provide 512x512 datset. Is any chance you create such dataset?</p>",
      "rawMarkdown": "Nic work Radek! I will use your work. \nI thnink that in case of 256x256 ... we should create more informative dataset - remove background (additional pixels). Look on second image and last one - we can crop it a little bit. Anyway ... it could be done in augumentation pipeline as well so do not need any work here.\n\nRadek it would be cool if you provide 512x512 datset. Is any chance you create such dataset?",
      "votes": 3,
      "replies": [
        {
          "id": 2049028,
          "postDate": "2022-11-29T21:56:27.950Z",
          "content": "<p><a href=\"https://www.kaggle.com/datasets/theoviel/rsna-breast-cancer-512-pngs\" target=\"_blank\">https://www.kaggle.com/datasets/theoviel/rsna-breast-cancer-512-pngs</a>  =)</p>",
          "rawMarkdown": "https://www.kaggle.com/datasets/theoviel/rsna-breast-cancer-512-pngs  =)",
          "votes": 6
        },
        {
          "id": 2049084,
          "postDate": "2022-11-29T23:37:13.553Z",
          "content": "<p>Cool <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a>, thx for sharing!</p>\n<p><a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a>, super glad you are finding this useful! 🙌 Just to make sure you have both sizes of images processed in the same way, I have now added images scaled to 512 x 512 to the dataset! 🙂</p>",
          "rawMarkdown": "Cool @theoviel, thx for sharing!\n\n@remekkinas, super glad you are finding this useful! 🙌 Just to make sure you have both sizes of images processed in the same way, I have now added images scaled to 512 x 512 to the dataset! 🙂"
        },
        {
          "id": 2049882,
          "postDate": "2022-11-30T12:02:19.787Z",
          "content": "<p><a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> thank you! You helped a lot. </p>\n<p><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> cool! Working on solution :)</p>",
          "rawMarkdown": "@theoviel thank you! You helped a lot. \n\n@radek1 cool! Working on solution :)",
          "votes": 1
        },
        {
          "id": 2050574,
          "postDate": "2022-11-30T21:03:17.587Z",
          "content": "<p>Awesome, <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a>! 🙂 Crossing my fingers and curious about what you'll be able to cook up! Best of luck!!! 🙂 </p>",
          "rawMarkdown": "Awesome, @remekkinas! 🙂 Crossing my fingers and curious about what you'll be able to cook up! Best of luck!!! 🙂 ",
          "votes": 1
        },
        {
          "id": 2051560,
          "postDate": "2022-12-01T13:54:43.037Z",
          "content": "<p><a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> I was thinking about the augumentation as well.  However, I just thought doing it by rotation / reflection, but i wonder whether the model was supposed to fit only the regular Mammography image.  The same also apply to removing the background.  Just curious for discussion.  </p>\n<p>Looking forward to your work.</p>",
          "rawMarkdown": "@remekkinas I was thinking about the augumentation as well.  However, I just thought doing it by rotation / reflection, but i wonder whether the model was supposed to fit only the regular Mammography image.  The same also apply to removing the background.  Just curious for discussion.  \n\nLooking forward to your work."
        },
        {
          "id": 2051573,
          "postDate": "2022-12-01T13:58:26.443Z",
          "content": "<p>Augumentation will be important here (as usual in CV task). Removing background is not important in my opinion (I will use ROI cropped images). We will see … we have many possibilities in this competition.</p>",
          "rawMarkdown": "Augumentation will be important here (as usual in CV task). Removing background is not important in my opinion (I will use ROI cropped images). We will see ... we have many possibilities in this competition.",
          "votes": 3
        },
        {
          "id": 2051578,
          "postDate": "2022-12-01T14:01:18.290Z",
          "content": "<p>The big component here is being able to train on more information with less compute… that is what the awesome work that you did <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> helps a lot!</p>\n<p>If we go from 512x512 but removing the regions that carry no signal to 256x256 we decrease the input size not by 2x… but by 4x! And this gets compounded down the road. </p>",
          "rawMarkdown": "The big component here is being able to train on more information with less compute... that is what the awesome work that you did @remekkinas helps a lot!\n\nIf we go from 512x512 but removing the regions that carry no signal to 256x256 we decrease the input size not by 2x... but by 4x! And this gets compounded down the road. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 2054262,
      "postDate": "2022-12-04T01:51:29.180Z",
      "content": "<p>thanks so much bro!</p>",
      "rawMarkdown": "thanks so much bro!",
      "votes": 1,
      "replies": [
        {
          "id": 2054290,
          "postDate": "2022-12-04T02:19:57.093Z",
          "content": "<p>pleasure, <a href=\"https://www.kaggle.com/ayaan\" target=\"_blank\">@ayaan</a>! 🙂</p>",
          "rawMarkdown": "pleasure, @ayaan! 🙂"
        }
      ]
    },
    {
      "id": 2049845,
      "postDate": "2022-11-30T11:26:30.683Z",
      "content": "<p>Thanks radek, <br>\nCan you pls share the code you've used to do this like any preprocessing is done by you or not.</p>",
      "rawMarkdown": "Thanks radek, \nCan you pls share the code you've used to do this like any preprocessing is done by you or not.",
      "votes": 1,
      "replies": [
        {
          "id": 2050573,
          "postDate": "2022-11-30T21:02:33.850Z",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/level14taken\" target=\"_blank\">@level14taken</a>! Sure thing, no problem!</p>\n<p>Pleas find the code here:  <a href=\"https://www.kaggle.com/code/radek1/how-i-processed-dicom-images-to-pngs?scriptVersionId=112592164\" target=\"_blank\">💡 how I processed DICOM images to PNGs</a></p>\n<p>All the best!</p>",
          "rawMarkdown": "Hey @level14taken! Sure thing, no problem!\n\nPleas find the code here:  [💡 how I processed DICOM images to PNGs](https://www.kaggle.com/code/radek1/how-i-processed-dicom-images-to-pngs?scriptVersionId=112592164)\n\nAll the best!",
          "votes": 2
        },
        {
          "id": 2051558,
          "postDate": "2022-12-01T13:51:11.070Z",
          "content": "<p><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> Thank you for your sharing.  Sorry but the link is not working.</p>",
          "rawMarkdown": "@radek1 Thank you for your sharing.  Sorry but the link is not working.",
          "votes": 1
        },
        {
          "id": 2051564,
          "postDate": "2022-12-01T13:56:40.103Z",
          "content": "<p>oh no, I didn't make it public…</p>\n<p>Thank you <a href=\"https://www.kaggle.com/jackysywk\" target=\"_blank\">@jackysywk</a> for letting me know 🙏 Should now be fixed, can you confirm, please?</p>",
          "rawMarkdown": "oh no, I didn't make it public...\n\nThank you @jackysywk for letting me know 🙏 Should now be fixed, can you confirm, please?",
          "votes": 1
        },
        {
          "id": 2051592,
          "postDate": "2022-12-01T14:08:12.780Z",
          "content": "<p>It now works perfectly.  Thanks 💯😁 <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> </p>",
          "rawMarkdown": "It now works perfectly.  Thanks 💯😁 @radek1 ",
          "votes": 1
        },
        {
          "id": 2051599,
          "postDate": "2022-12-01T14:11:35.990Z",
          "content": "<p>Wonderful to hear! Thank you for checking and letting me know,  <a href=\"https://www.kaggle.com/jackysywk\" target=\"_blank\">@jackysywk</a>! 🙌 </p>",
          "rawMarkdown": "Wonderful to hear! Thank you for checking and letting me know,  @jackysywk! 🙌 ",
          "votes": 1
        }
      ]
    },
    {
      "id": 2049359,
      "postDate": "2022-11-30T04:38:03.310Z",
      "content": "<p>Nice work~~~ Thanks!</p>",
      "rawMarkdown": "Nice work~~~ Thanks!",
      "votes": 1,
      "replies": [
        {
          "id": 2051567,
          "postDate": "2022-12-01T13:57:04.653Z",
          "content": "<p>np <a href=\"https://www.kaggle.com/gavinxuai\" target=\"_blank\">@gavinxuai</a>, my pleasure! 🙌</p>",
          "rawMarkdown": "np @gavinxuai, my pleasure! 🙌"
        }
      ]
    },
    {
      "id": 2049287,
      "postDate": "2022-11-30T03:16:51.507Z",
      "content": "<p>Would really help if you can make 768x768 and 1024x1024 as well. They can be used with certain transformer models</p>",
      "rawMarkdown": "Would really help if you can make 768x768 and 1024x1024 as well. They can be used with certain transformer models",
      "votes": 1,
      "replies": [
        {
          "id": 2051562,
          "postDate": "2022-12-01T13:56:09.793Z",
          "content": "<p><a href=\"https://www.kaggle.com/nishantbhansali\" target=\"_blank\">@nishantbhansali</a> Wonderful idea<br>\nGrateful if you can check the link above and i can try doing other pixels, if this is not hardware demanding.</p>",
          "rawMarkdown": "@nishantbhansali Wonderful idea\nGrateful if you can check the link above and i can try doing other pixels, if this is not hardware demanding.",
          "votes": 1
        },
        {
          "id": 2051570,
          "postDate": "2022-12-01T13:58:02.120Z",
          "content": "<p>Hey! these have been added to the dataset! 🙂 happy scaling up to bigger inputs! 😄</p>",
          "rawMarkdown": "Hey! these have been added to the dataset! 🙂 happy scaling up to bigger inputs! 😄",
          "votes": 1
        }
      ]
    },
    {
      "id": 2159958,
      "postDate": "2023-02-26T07:31:31.203Z",
      "content": "<p>Three days ago, I could run image.dicom.pixel_array to decompress a Dicom image. At this point, I get the error \"The pixel data with transfer syntax JPEG 2000 Image Compression (Lossless Only), cannot be read because Pillow lacks the JPEG 2000 plugin\". On my computer, it has been resolved by installing 'pip install -U pylibjpeg pylibjpeg-openjpeg pylibjpeg-libjpeg'. <br>\nIn Kaggle, it seems to install it in 'base environment' but it doesn't work, nor can I create another environment. Could someone help me please?</p>",
      "rawMarkdown": "Three days ago, I could run image.dicom.pixel_array to decompress a Dicom image. At this point, I get the error \"The pixel data with transfer syntax JPEG 2000 Image Compression (Lossless Only), cannot be read because Pillow lacks the JPEG 2000 plugin\". On my computer, it has been resolved by installing 'pip install -U pylibjpeg pylibjpeg-openjpeg pylibjpeg-libjpeg'. \nIn Kaggle, it seems to install it in 'base environment' but it doesn't work, nor can I create another environment. Could someone help me please?"
    },
    {
      "id": 2112596,
      "postDate": "2023-01-23T17:46:20.133Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> and all, </p>\n<p>I have one naive question: Since we want to have the same distribution of both test and training data, don't we need to use the same data transformation pipeline during inference time? DICOM to PNGs, crop, etc. ?</p>",
      "rawMarkdown": "Hi @radek1 and all, \n\nI have one naive question: Since we want to have the same distribution of both test and training data, don't we need to use the same data transformation pipeline during inference time? DICOM to PNGs, crop, etc. ?"
    },
    {
      "id": 2102074,
      "postDate": "2023-01-16T11:22:52.837Z",
      "content": "<p>Wow! It will be very helpful, thanks bro :)</p>",
      "rawMarkdown": "Wow! It will be very helpful, thanks bro :)"
    },
    {
      "id": 2057235,
      "postDate": "2022-12-06T21:52:43.297Z",
      "content": "<p>This will not only help us to learn but to save space locally..😄  </p>",
      "rawMarkdown": "This will not only help us to learn but to save space locally..😄  "
    },
    {
      "id": 2052541,
      "postDate": "2022-12-02T09:37:55.550Z",
      "content": "<p>Hi there, don't get me wrong. I'm just new to these kind of competitions. is information loss when we convert DICOM images to PNGs ?</p>",
      "rawMarkdown": "Hi there, don't get me wrong. I'm just new to these kind of competitions. is information loss when we convert DICOM images to PNGs ?",
      "replies": [
        {
          "id": 2054300,
          "postDate": "2022-12-04T02:26:36.843Z",
          "content": "<p>hey! yes, it is always a question how to go from DICOM images to other formats 🙂 There has been some discussion on the forums, for a really good analysis please take a look <a href=\"https://www.kaggle.com/competitions/rsna-intracranial-hemorrhage-detection/discussion/114214?utm_source=pocket_saves\" target=\"_blank\">here</a></p>",
          "rawMarkdown": "hey! yes, it is always a question how to go from DICOM images to other formats 🙂 There has been some discussion on the forums, for a really good analysis please take a look [here](https://www.kaggle.com/competitions/rsna-intracranial-hemorrhage-detection/discussion/114214?utm_source=pocket_saves)"
        },
        {
          "id": 2054303,
          "postDate": "2022-12-04T02:32:17.147Z",
          "content": "<p>okay bro, i will take a look. thank you very much.</p>",
          "rawMarkdown": "okay bro, i will take a look. thank you very much.\n"
        }
      ]
    },
    {
      "id": 2048727,
      "postDate": "2022-11-29T16:29:50.213Z",
      "content": "<p>Great work, one question is, what should we do when we need to submit prediction results?</p>",
      "rawMarkdown": "Great work, one question is, what should we do when we need to submit prediction results?",
      "replies": [
        {
          "id": 2048801,
          "postDate": "2022-11-29T17:25:48.103Z",
          "content": "<p>Nothing with images. You need to submit probability of cancer on particular photo.<br>\n256x256 is great for experimenting. Then we probably need more pixels and more GPU power.</p>",
          "rawMarkdown": "Nothing with images. You need to submit probability of cancer on particular photo.\n256x256 is great for experimenting. Then we probably need more pixels and more GPU power.",
          "votes": 1
        },
        {
          "id": 2056117,
          "postDate": "2022-12-05T18:39:10.313Z",
          "content": "<p>Well, blending with lower resolution models might work.  Need to experiment</p>",
          "rawMarkdown": "Well, blending with lower resolution models might work.  Need to experiment"
        }
      ]
    },
    {
      "id": 2070798,
      "postDate": "2022-12-20T10:53:29.970Z",
      "content": "<p>thanks so much bro!</p>",
      "rawMarkdown": "thanks so much bro!"
    }
  ],
  "comments": [
    {
      "id": 2056098,
      "author_name": "Hari T.",
      "author_url": "",
      "post_date": "2022-12-05T18:20:02.083000",
      "content": "<p>Exercise caution when converting from DICOM to PNG. Mammograms are very high resolution (12MP on average) and downsampling heavily may render small findings invisible. 8-bit is sufficient for most cases but 12 and 16-bit should be considered as well. Run small experiments to understand the importance.</p>\n<p>Cropping the mamogram to remove background tissue is a good strategy but consider the implications of padding to achieve the same image size rather than distorting the aspect ratio. This may change the appearance of fine calcifications.</p>\n<p>Good luck!</p>",
      "votes": 18,
      "replies": []
    },
    {
      "id": 2056988,
      "author_name": "raddar",
      "author_url": "",
      "post_date": "2022-12-06T16:33:12.897000",
      "content": "<p>As this information might get unnoticed, if you have been using this dataset, please be aware of:</p>\n<p><a href=\"https://www.kaggle.com/code/radek1/how-to-process-dicom-images-to-pngs/comments#2056978\" target=\"_blank\">https://www.kaggle.com/code/radek1/how-to-process-dicom-images-to-pngs/comments#2056978</a></p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2048805,
      "author_name": "Remek Kinas",
      "author_url": "",
      "post_date": "2022-11-29T17:28:04.683000",
      "content": "<p>Nic work Radek! I will use your work. <br>\nI thnink that in case of 256x256 … we should create more informative dataset - remove background (additional pixels). Look on second image and last one - we can crop it a little bit. Anyway … it could be done in augumentation pipeline as well so do not need any work here.</p>\n<p>Radek it would be cool if you provide 512x512 datset. Is any chance you create such dataset?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2049028,
          "author_name": "Theo Viel",
          "author_url": "",
          "post_date": "2022-11-29T21:56:27.950000",
          "content": "<p><a href=\"https://www.kaggle.com/datasets/theoviel/rsna-breast-cancer-512-pngs\" target=\"_blank\">https://www.kaggle.com/datasets/theoviel/rsna-breast-cancer-512-pngs</a>  =)</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 2049084,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-11-29T23:37:13.553000",
          "content": "<p>Cool <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a>, thx for sharing!</p>\n<p><a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a>, super glad you are finding this useful! 🙌 Just to make sure you have both sizes of images processed in the same way, I have now added images scaled to 512 x 512 to the dataset! 🙂</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2049882,
          "author_name": "Remek Kinas",
          "author_url": "",
          "post_date": "2022-11-30T12:02:19.787000",
          "content": "<p><a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> thank you! You helped a lot. </p>\n<p><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> cool! Working on solution :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2050574,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-11-30T21:03:17.587000",
          "content": "<p>Awesome, <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a>! 🙂 Crossing my fingers and curious about what you'll be able to cook up! Best of luck!!! 🙂 </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2051560,
          "author_name": "Jacky S",
          "author_url": "",
          "post_date": "2022-12-01T13:54:43.037000",
          "content": "<p><a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> I was thinking about the augumentation as well.  However, I just thought doing it by rotation / reflection, but i wonder whether the model was supposed to fit only the regular Mammography image.  The same also apply to removing the background.  Just curious for discussion.  </p>\n<p>Looking forward to your work.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2051573,
          "author_name": "Remek Kinas",
          "author_url": "",
          "post_date": "2022-12-01T13:58:26.443000",
          "content": "<p>Augumentation will be important here (as usual in CV task). Removing background is not important in my opinion (I will use ROI cropped images). We will see … we have many possibilities in this competition.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 2051578,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-12-01T14:01:18.290000",
          "content": "<p>The big component here is being able to train on more information with less compute… that is what the awesome work that you did <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> helps a lot!</p>\n<p>If we go from 512x512 but removing the regions that carry no signal to 256x256 we decrease the input size not by 2x… but by 4x! And this gets compounded down the road. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2054262,
      "author_name": "ayaan",
      "author_url": "",
      "post_date": "2022-12-04T01:51:29.180000",
      "content": "<p>thanks so much bro!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2054290,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-12-04T02:19:57.093000",
          "content": "<p>pleasure, <a href=\"https://www.kaggle.com/ayaan\" target=\"_blank\">@ayaan</a>! 🙂</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2049845,
      "author_name": "losingself",
      "author_url": "",
      "post_date": "2022-11-30T11:26:30.683000",
      "content": "<p>Thanks radek, <br>\nCan you pls share the code you've used to do this like any preprocessing is done by you or not.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2050573,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-11-30T21:02:33.850000",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/level14taken\" target=\"_blank\">@level14taken</a>! Sure thing, no problem!</p>\n<p>Pleas find the code here:  <a href=\"https://www.kaggle.com/code/radek1/how-i-processed-dicom-images-to-pngs?scriptVersionId=112592164\" target=\"_blank\">💡 how I processed DICOM images to PNGs</a></p>\n<p>All the best!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2051558,
          "author_name": "Jacky S",
          "author_url": "",
          "post_date": "2022-12-01T13:51:11.070000",
          "content": "<p><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> Thank you for your sharing.  Sorry but the link is not working.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2051564,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-12-01T13:56:40.103000",
          "content": "<p>oh no, I didn't make it public…</p>\n<p>Thank you <a href=\"https://www.kaggle.com/jackysywk\" target=\"_blank\">@jackysywk</a> for letting me know 🙏 Should now be fixed, can you confirm, please?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2051592,
          "author_name": "Jacky S",
          "author_url": "",
          "post_date": "2022-12-01T14:08:12.780000",
          "content": "<p>It now works perfectly.  Thanks 💯😁 <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2051599,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-12-01T14:11:35.990000",
          "content": "<p>Wonderful to hear! Thank you for checking and letting me know,  <a href=\"https://www.kaggle.com/jackysywk\" target=\"_blank\">@jackysywk</a>! 🙌 </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2049359,
      "author_name": "zhenghuoer",
      "author_url": "",
      "post_date": "2022-11-30T04:38:03.310000",
      "content": "<p>Nice work~~~ Thanks!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2051567,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-12-01T13:57:04.653000",
          "content": "<p>np <a href=\"https://www.kaggle.com/gavinxuai\" target=\"_blank\">@gavinxuai</a>, my pleasure! 🙌</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2049287,
      "author_name": "Nishant Bhansali",
      "author_url": "",
      "post_date": "2022-11-30T03:16:51.507000",
      "content": "<p>Would really help if you can make 768x768 and 1024x1024 as well. They can be used with certain transformer models</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2051562,
          "author_name": "Jacky S",
          "author_url": "",
          "post_date": "2022-12-01T13:56:09.793000",
          "content": "<p><a href=\"https://www.kaggle.com/nishantbhansali\" target=\"_blank\">@nishantbhansali</a> Wonderful idea<br>\nGrateful if you can check the link above and i can try doing other pixels, if this is not hardware demanding.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2051570,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-12-01T13:58:02.120000",
          "content": "<p>Hey! these have been added to the dataset! 🙂 happy scaling up to bigger inputs! 😄</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2159958,
      "author_name": "Eduardo Gasca Alvarez",
      "author_url": "",
      "post_date": "2023-02-26T07:31:31.203000",
      "content": "<p>Three days ago, I could run image.dicom.pixel_array to decompress a Dicom image. At this point, I get the error \"The pixel data with transfer syntax JPEG 2000 Image Compression (Lossless Only), cannot be read because Pillow lacks the JPEG 2000 plugin\". On my computer, it has been resolved by installing 'pip install -U pylibjpeg pylibjpeg-openjpeg pylibjpeg-libjpeg'. <br>\nIn Kaggle, it seems to install it in 'base environment' but it doesn't work, nor can I create another environment. Could someone help me please?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2112596,
      "author_name": "Ranjeet",
      "author_url": "",
      "post_date": "2023-01-23T17:46:20.133000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> and all, </p>\n<p>I have one naive question: Since we want to have the same distribution of both test and training data, don't we need to use the same data transformation pipeline during inference time? DICOM to PNGs, crop, etc. ?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2102074,
      "author_name": "Ncolmann",
      "author_url": "",
      "post_date": "2023-01-16T11:22:52.837000",
      "content": "<p>Wow! It will be very helpful, thanks bro :)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2057235,
      "author_name": "Saurav Solanki",
      "author_url": "",
      "post_date": "2022-12-06T21:52:43.297000",
      "content": "<p>This will not only help us to learn but to save space locally..😄  </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2052541,
      "author_name": "nadhir hasan",
      "author_url": "",
      "post_date": "2022-12-02T09:37:55.550000",
      "content": "<p>Hi there, don't get me wrong. I'm just new to these kind of competitions. is information loss when we convert DICOM images to PNGs ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2054300,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-12-04T02:26:36.843000",
          "content": "<p>hey! yes, it is always a question how to go from DICOM images to other formats 🙂 There has been some discussion on the forums, for a really good analysis please take a look <a href=\"https://www.kaggle.com/competitions/rsna-intracranial-hemorrhage-detection/discussion/114214?utm_source=pocket_saves\" target=\"_blank\">here</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2054303,
          "author_name": "nadhir hasan",
          "author_url": "",
          "post_date": "2022-12-04T02:32:17.147000",
          "content": "<p>okay bro, i will take a look. thank you very much.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2048727,
      "author_name": "yanqiangmiffy",
      "author_url": "",
      "post_date": "2022-11-29T16:29:50.213000",
      "content": "<p>Great work, one question is, what should we do when we need to submit prediction results?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2048801,
          "author_name": "Remek Kinas",
          "author_url": "",
          "post_date": "2022-11-29T17:25:48.103000",
          "content": "<p>Nothing with images. You need to submit probability of cancer on particular photo.<br>\n256x256 is great for experimenting. Then we probably need more pixels and more GPU power.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2056117,
          "author_name": "@kaggleqrdl",
          "author_url": "",
          "post_date": "2022-12-05T18:39:10.313000",
          "content": "<p>Well, blending with lower resolution models might work.  Need to experiment</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2070798,
      "author_name": "1! 5!",
      "author_url": "",
      "post_date": "2022-12-20T10:53:29.970000",
      "content": "<p>thanks so much bro!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2048609": "Hey!\n\nPlease find the DICOM files converted to PNG here: [RSNA Mammography -- images as PNGs (256px x 256px)](https://www.kaggle.com/datasets/radek1/rsna-mammography-images-as-pngs/settings)\n\n**EDIT: Based on a request by @remekkinas I created a new version of the dataset that also contains 512px x 512px images!**\n\n\nI tried converting the files on Kaggle, but the VM was just not enough 🙂 Even though the processing completed after more than ~3 hrs using the 4 available cores, I still didn't manage to commit the output.\n\nOpted to go for a GCP VM to get the job done 🙂\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2Fdbc2480e7f8863133c72cc37cc9f6b81%2Fgcp_vm_at_work.png?generation=1669731617602415&alt=media)\n\nOn GCP the conversion took just under 17 minutes using 64 cores and multiprocessing. Please note, I converted the images in a way that accounts for `dicom.PhotometricInterpretation == \"MONOCHROME1\"` where the intensities need to be inverted.\n\nI converted DICOM files to PNG images 256px by 256px. The folder structure is the same as of the DICOM files.\n\nHere are a couple of examples:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F4fc6e1f8837c3be387dd45077bf7aa7f%2F2057295788.png?generation=1669735207946955&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2Ff0bd8462a9d27d8049185f6039d3536e%2F2046475482.png?generation=1669735227410985&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2Ff155f4f11b2e5d8e836423bf5f30c7a5%2F349510516.png?generation=1669735239725229&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2Ff1d5dd2efa0bbbd68d061f9aad6cd868%2F1365269360.png?generation=1669735252065645&alt=media)\n\nHappy Kaggling! 🙂\n\n### Other resources you might find useful:\n\n* [💡 how to process DICOM images to PNGs](https://www.kaggle.com/code/radek1/how-to-process-dicom-images-to-pngs)\n* [📊 EDA + training a fast.ai model + submission 🚀](https://www.kaggle.com/code/radek1/eda-training-a-fast-ai-model-submission)\n* [🤖 [fast.ai starter pack] train + inference 🚀](https://www.kaggle.com/code/radek1/fast-ai-starter-pack-train-inference)\n* [📸 Over 56GB of processed data, 5 different methods 🥳](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/371268)\n* [3 resources to get started with Computer Vision in this competition 🚀🚀🚀](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369706)\n* [💡 6 Computer Vision tricks for faster training and better models 🚀🚀🚀](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369155)\n\n",
    "2056098": "Exercise caution when converting from DICOM to PNG. Mammograms are very high resolution (12MP on average) and downsampling heavily may render small findings invisible. 8-bit is sufficient for most cases but 12 and 16-bit should be considered as well. Run small experiments to understand the importance.\n\nCropping the mamogram to remove background tissue is a good strategy but consider the implications of padding to achieve the same image size rather than distorting the aspect ratio. This may change the appearance of fine calcifications.\n\nGood luck!",
    "2056988": "As this information might get unnoticed, if you have been using this dataset, please be aware of:\n\nhttps://www.kaggle.com/code/radek1/how-to-process-dicom-images-to-pngs/comments#2056978",
    "2048805": "Nic work Radek! I will use your work. \nI thnink that in case of 256x256 ... we should create more informative dataset - remove background (additional pixels). Look on second image and last one - we can crop it a little bit. Anyway ... it could be done in augumentation pipeline as well so do not need any work here.\n\nRadek it would be cool if you provide 512x512 datset. Is any chance you create such dataset?",
    "2054262": "thanks so much bro!",
    "2049845": "Thanks radek, \nCan you pls share the code you've used to do this like any preprocessing is done by you or not.",
    "2049359": "Nice work~~~ Thanks!",
    "2049287": "Would really help if you can make 768x768 and 1024x1024 as well. They can be used with certain transformer models",
    "2159958": "Three days ago, I could run image.dicom.pixel_array to decompress a Dicom image. At this point, I get the error \"The pixel data with transfer syntax JPEG 2000 Image Compression (Lossless Only), cannot be read because Pillow lacks the JPEG 2000 plugin\". On my computer, it has been resolved by installing 'pip install -U pylibjpeg pylibjpeg-openjpeg pylibjpeg-libjpeg'. \nIn Kaggle, it seems to install it in 'base environment' but it doesn't work, nor can I create another environment. Could someone help me please?",
    "2112596": "Hi @radek1 and all, \n\nI have one naive question: Since we want to have the same distribution of both test and training data, don't we need to use the same data transformation pipeline during inference time? DICOM to PNGs, crop, etc. ?",
    "2102074": "Wow! It will be very helpful, thanks bro :)",
    "2057235": "This will not only help us to learn but to save space locally..😄  ",
    "2052541": "Hi there, don't get me wrong. I'm just new to these kind of competitions. is information loss when we convert DICOM images to PNGs ?",
    "2048727": "Great work, one question is, what should we do when we need to submit prediction results?",
    "2070798": "thanks so much bro!"
  }
}