{
  "id": 370525,
  "title": "[UPDATE] 📸 DICOM files converted to PNGs -> 256px, 384px, 512px, 768px and 1024px",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/370525",
  "author_name": "Radek Osmulski",
  "post_date": "2022-12-05T03:42:51.828000",
  "votes": 6,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hey,</p>\n<p>I can see that quite a few people have been using a dataset that I shared: <a href=\"https://www.kaggle.com/datasets/radek1/rsna-mammography-images-as-pngs\" target=\"_blank\">RSNA Mammography PNGs 256px, 512px, 768px &amp; 1024px</a> so just would like to give a quick update:</p>\n<h3>Processing using <code>opencv</code> is 10%-20% faster than with <code>PIL</code></h3>\n<p>Due to limited VM time (and us wanting to squeeze as much inference time into the pipeline as possible) we have to be very careful about how long it takes to process the images to PNGs.</p>\n<p>Based on limited experiments <em>I believe</em> that processing using <code>opencv</code> can be 10%-20% faster! As such, I uploaded files processed using <code>opencv</code> to the dataset for resolutions 256px and 512px. Currently also working on adding 384px (halfway between 256px and 512px).</p>\n<p>You can find the code I used to process the images here: <a href=\"https://www.kaggle.com/code/radek1/how-to-process-dicom-images-to-pngs\" target=\"_blank\">💡 how to process DICOM images to PNGs</a></p>\n<h3>I have not found any issues with any versions of the dataset 🥳</h3>\n<p>No one has used the dataset for successful training &amp; submission to the LB yet as far as I know (there are multiple training notebooks that train fine but the inference part has not been made public).</p>\n<p>However, in working quite intensely on <a href=\"https://www.kaggle.com/code/radek1/fast-ai-starter-pack-train-inference\" target=\"_blank\">🤖 [fast.ai starter pack] train + inference 🚀</a> as part of the verification of why my model is not training well, I have been switching across different public datasets, looking very closely at what my model is seeing in training. And I have not observed any issues with the data. The data seems 👌and safe to train on. </p>\n<p>(In fact, I got the best performance on the smallest size 256px x 256px with the data processed using <code>PIL</code> but I wouldn't read too much into those differences, my model training is unstable for now hence it is not a good idea too read too much into differences that don't seem very large)</p>\n<p>But just wanted to let you know that extensive testing has been done on the dataset and so far everything looks okay 🙂</p>",
  "messages": [
    {
      "id": 2055378,
      "postDate": "2022-12-05T03:42:51.830Z",
      "content": "<p>Hey,</p>\n<p>I can see that quite a few people have been using a dataset that I shared: <a href=\"https://www.kaggle.com/datasets/radek1/rsna-mammography-images-as-pngs\" target=\"_blank\">RSNA Mammography PNGs 256px, 512px, 768px &amp; 1024px</a> so just would like to give a quick update:</p>\n<h3>Processing using <code>opencv</code> is 10%-20% faster than with <code>PIL</code></h3>\n<p>Due to limited VM time (and us wanting to squeeze as much inference time into the pipeline as possible) we have to be very careful about how long it takes to process the images to PNGs.</p>\n<p>Based on limited experiments <em>I believe</em> that processing using <code>opencv</code> can be 10%-20% faster! As such, I uploaded files processed using <code>opencv</code> to the dataset for resolutions 256px and 512px. Currently also working on adding 384px (halfway between 256px and 512px).</p>\n<p>You can find the code I used to process the images here: <a href=\"https://www.kaggle.com/code/radek1/how-to-process-dicom-images-to-pngs\" target=\"_blank\">💡 how to process DICOM images to PNGs</a></p>\n<h3>I have not found any issues with any versions of the dataset 🥳</h3>\n<p>No one has used the dataset for successful training &amp; submission to the LB yet as far as I know (there are multiple training notebooks that train fine but the inference part has not been made public).</p>\n<p>However, in working quite intensely on <a href=\"https://www.kaggle.com/code/radek1/fast-ai-starter-pack-train-inference\" target=\"_blank\">🤖 [fast.ai starter pack] train + inference 🚀</a> as part of the verification of why my model is not training well, I have been switching across different public datasets, looking very closely at what my model is seeing in training. And I have not observed any issues with the data. The data seems 👌and safe to train on. </p>\n<p>(In fact, I got the best performance on the smallest size 256px x 256px with the data processed using <code>PIL</code> but I wouldn't read too much into those differences, my model training is unstable for now hence it is not a good idea too read too much into differences that don't seem very large)</p>\n<p>But just wanted to let you know that extensive testing has been done on the dataset and so far everything looks okay 🙂</p>",
      "rawMarkdown": "Hey,\n\nI can see that quite a few people have been using a dataset that I shared: [RSNA Mammography PNGs 256px, 512px, 768px & 1024px](https://www.kaggle.com/datasets/radek1/rsna-mammography-images-as-pngs) so just would like to give a quick update:\n\n### Processing using `opencv` is 10%-20% faster than with `PIL`\n\nDue to limited VM time (and us wanting to squeeze as much inference time into the pipeline as possible) we have to be very careful about how long it takes to process the images to PNGs.\n\nBased on limited experiments *I believe* that processing using `opencv` can be 10%-20% faster! As such, I uploaded files processed using `opencv` to the dataset for resolutions 256px and 512px. Currently also working on adding 384px (halfway between 256px and 512px).\n\nYou can find the code I used to process the images here: [💡 how to process DICOM images to PNGs](https://www.kaggle.com/code/radek1/how-to-process-dicom-images-to-pngs)\n\n### I have not found any issues with any versions of the dataset 🥳\n\nNo one has used the dataset for successful training & submission to the LB yet as far as I know (there are multiple training notebooks that train fine but the inference part has not been made public).\n\nHowever, in working quite intensely on [🤖 [fast.ai starter pack] train + inference 🚀](https://www.kaggle.com/code/radek1/fast-ai-starter-pack-train-inference) as part of the verification of why my model is not training well, I have been switching across different public datasets, looking very closely at what my model is seeing in training. And I have not observed any issues with the data. The data seems 👌and safe to train on. \n\n(In fact, I got the best performance on the smallest size 256px x 256px with the data processed using `PIL` but I wouldn't read too much into those differences, my model training is unstable for now hence it is not a good idea too read too much into differences that don't seem very large)\n\nBut just wanted to let you know that extensive testing has been done on the dataset and so far everything looks okay 🙂",
      "votes": 6
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2055378": "Hey,\n\nI can see that quite a few people have been using a dataset that I shared: [RSNA Mammography PNGs 256px, 512px, 768px & 1024px](https://www.kaggle.com/datasets/radek1/rsna-mammography-images-as-pngs) so just would like to give a quick update:\n\n### Processing using `opencv` is 10%-20% faster than with `PIL`\n\nDue to limited VM time (and us wanting to squeeze as much inference time into the pipeline as possible) we have to be very careful about how long it takes to process the images to PNGs.\n\nBased on limited experiments *I believe* that processing using `opencv` can be 10%-20% faster! As such, I uploaded files processed using `opencv` to the dataset for resolutions 256px and 512px. Currently also working on adding 384px (halfway between 256px and 512px).\n\nYou can find the code I used to process the images here: [💡 how to process DICOM images to PNGs](https://www.kaggle.com/code/radek1/how-to-process-dicom-images-to-pngs)\n\n### I have not found any issues with any versions of the dataset 🥳\n\nNo one has used the dataset for successful training & submission to the LB yet as far as I know (there are multiple training notebooks that train fine but the inference part has not been made public).\n\nHowever, in working quite intensely on [🤖 [fast.ai starter pack] train + inference 🚀](https://www.kaggle.com/code/radek1/fast-ai-starter-pack-train-inference) as part of the verification of why my model is not training well, I have been switching across different public datasets, looking very closely at what my model is seeing in training. And I have not observed any issues with the data. The data seems 👌and safe to train on. \n\n(In fact, I got the best performance on the smallest size 256px x 256px with the data processed using `PIL` but I wouldn't read too much into those differences, my model training is unstable for now hence it is not a good idea too read too much into differences that don't seem very large)\n\nBut just wanted to let you know that extensive testing has been done on the dataset and so far everything looks okay 🙂"
  }
}