{
  "id": 371268,
  "title": "📸 Over 56GB of processed data, 5 different methods 🥳",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/371268",
  "author_name": "Radek Osmulski",
  "post_date": "2022-12-09T01:49:59.065000",
  "votes": 21,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hey!</p>\n<p>I wanted to give you an update on the <a href=\"https://www.kaggle.com/datasets/radek1/rsna-mammography-images-as-pngs\" target=\"_blank\">RSNA Mammo PNGs 256px, 384px, 512px, 768px, 1024px</a> that I know many people have been using!</p>\n<p>(can also be useful to people who just joined the competition)</p>\n<p>I kept adding to the dataset so it now has a decent collection of sizes processed in various ways 🙂</p>\n<p>All processing is documented in <a href=\"https://www.kaggle.com/code/radek1/how-to-process-dicom-images-to-pngs\" target=\"_blank\">💡 how to process DICOM images to PNGs</a> so you can grab it and straight away implement it in your inference notebook!</p>\n<p>The four processing methods are as follows:</p>\n<ul>\n<li>processing using <code>PIL</code> - <code>PIL</code> is a great choice but turned out for me to be 10%-20% slower than <code>opencv</code>, so switched to that</li>\n<li>processing using <code>opencv</code> - like above, but faster, using <code>opencv</code> (these methods normalize images to between 0 and 1 and account fo the <code>MONOCHROME1</code> status)</li>\n<li>processing using <code>opencv</code> with VOI LUT transformation (thank you <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> 🙏)</li>\n<li>processing using <code>opencv</code> without VOI LUT transformation AND with reading the data in by <code>dicomsdl</code> (2x faster then <code>pydicom</code>!)  (thank you <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a>! 🙏) 🔥🔥🔥 </li>\n<li>processing using <code>opencv</code> with keeping the aspect ratio, resizing the longer side to 1024 (please see discussion in this thread on this <br>\nfor further information)</li>\n</ul>\n<p>Hope you find this useful in working on the solution! 🙂</p>\n<h2>Other resources you might find useful:</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/how-to-process-dicom-images-to-pngs\" target=\"_blank\">💡 how to process DICOM images to PNGs</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/eda-training-a-fast-ai-model-submission\" target=\"_blank\">📊 EDA + training a fast.ai model + submission 🚀</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/fast-ai-starter-pack-train-inference\" target=\"_blank\">🤖 [fast.ai starter pack] train + inference 🚀</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/371268\" target=\"_blank\">📸 Over 56GB of processed data, 5 different methods 🥳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369706\" target=\"_blank\">3 resources to get started with Computer Vision in this competition 🚀🚀🚀</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369155\" target=\"_blank\">💡 6 Computer Vision tricks for faster training and better models 🚀🚀🚀</a></li>\n</ul>",
  "messages": [
    {
      "id": 2059535,
      "postDate": "2022-12-09T01:49:59.067Z",
      "content": "<p>Hey!</p>\n<p>I wanted to give you an update on the <a href=\"https://www.kaggle.com/datasets/radek1/rsna-mammography-images-as-pngs\" target=\"_blank\">RSNA Mammo PNGs 256px, 384px, 512px, 768px, 1024px</a> that I know many people have been using!</p>\n<p>(can also be useful to people who just joined the competition)</p>\n<p>I kept adding to the dataset so it now has a decent collection of sizes processed in various ways 🙂</p>\n<p>All processing is documented in <a href=\"https://www.kaggle.com/code/radek1/how-to-process-dicom-images-to-pngs\" target=\"_blank\">💡 how to process DICOM images to PNGs</a> so you can grab it and straight away implement it in your inference notebook!</p>\n<p>The four processing methods are as follows:</p>\n<ul>\n<li>processing using <code>PIL</code> - <code>PIL</code> is a great choice but turned out for me to be 10%-20% slower than <code>opencv</code>, so switched to that</li>\n<li>processing using <code>opencv</code> - like above, but faster, using <code>opencv</code> (these methods normalize images to between 0 and 1 and account fo the <code>MONOCHROME1</code> status)</li>\n<li>processing using <code>opencv</code> with VOI LUT transformation (thank you <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> 🙏)</li>\n<li>processing using <code>opencv</code> without VOI LUT transformation AND with reading the data in by <code>dicomsdl</code> (2x faster then <code>pydicom</code>!)  (thank you <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a>! 🙏) 🔥🔥🔥 </li>\n<li>processing using <code>opencv</code> with keeping the aspect ratio, resizing the longer side to 1024 (please see discussion in this thread on this <br>\nfor further information)</li>\n</ul>\n<p>Hope you find this useful in working on the solution! 🙂</p>\n<h2>Other resources you might find useful:</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/how-to-process-dicom-images-to-pngs\" target=\"_blank\">💡 how to process DICOM images to PNGs</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/eda-training-a-fast-ai-model-submission\" target=\"_blank\">📊 EDA + training a fast.ai model + submission 🚀</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/fast-ai-starter-pack-train-inference\" target=\"_blank\">🤖 [fast.ai starter pack] train + inference 🚀</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/371268\" target=\"_blank\">📸 Over 56GB of processed data, 5 different methods 🥳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369706\" target=\"_blank\">3 resources to get started with Computer Vision in this competition 🚀🚀🚀</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369155\" target=\"_blank\">💡 6 Computer Vision tricks for faster training and better models 🚀🚀🚀</a></li>\n</ul>",
      "rawMarkdown": "Hey!\n\nI wanted to give you an update on the [RSNA Mammo PNGs 256px, 384px, 512px, 768px, 1024px](https://www.kaggle.com/datasets/radek1/rsna-mammography-images-as-pngs) that I know many people have been using!\n\n(can also be useful to people who just joined the competition)\n\nI kept adding to the dataset so it now has a decent collection of sizes processed in various ways 🙂\n\nAll processing is documented in [💡 how to process DICOM images to PNGs](https://www.kaggle.com/code/radek1/how-to-process-dicom-images-to-pngs) so you can grab it and straight away implement it in your inference notebook!\n\nThe four processing methods are as follows:\n* processing using `PIL` - `PIL` is a great choice but turned out for me to be 10%-20% slower than `opencv`, so switched to that\n* processing using `opencv` - like above, but faster, using `opencv` (these methods normalize images to between 0 and 1 and account fo the `MONOCHROME1` status)\n* processing using `opencv` with VOI LUT transformation (thank you @raddar 🙏)\n* processing using `opencv` without VOI LUT transformation AND with reading the data in by `dicomsdl` (2x faster then `pydicom`!)  (thank you @remekkinas! 🙏) 🔥🔥🔥 \n* processing using `opencv` with keeping the aspect ratio, resizing the longer side to 1024 (please see discussion in this thread on this \nfor further information)\n\nHope you find this useful in working on the solution! 🙂\n\n## Other resources you might find useful:\n\n* [💡 how to process DICOM images to PNGs](https://www.kaggle.com/code/radek1/how-to-process-dicom-images-to-pngs)\n* [📊 EDA + training a fast.ai model + submission 🚀](https://www.kaggle.com/code/radek1/eda-training-a-fast-ai-model-submission)\n* [🤖 [fast.ai starter pack] train + inference 🚀](https://www.kaggle.com/code/radek1/fast-ai-starter-pack-train-inference)\n* [📸 Over 56GB of processed data, 5 different methods 🥳](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/371268)\n* [3 resources to get started with Computer Vision in this competition 🚀🚀🚀](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369706)\n* [💡 6 Computer Vision tricks for faster training and better models 🚀🚀🚀](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369155)",
      "votes": 21
    },
    {
      "id": 2059728,
      "postDate": "2022-12-09T06:59:46.113Z",
      "content": "<p>Thank you very much for mentioning. Great work!<br>\nI thnink that one more dataset is required 😁 - images with preserved aspect ratio.</p>\n<pre><code>def image_resize(image, width = None, height = None, inter = cv2.INTER_LINEAR):\n\n    dim = None\n    (h, w) = image.shape[:2]\n\n    if width is None and height is None:\n        return image\n\n    if width is None:\n        r = height / float(h)\n        dim = (int(w * r), height)\n    else:\n        r = width / float(w)\n        dim = (width, int(h * r))\n    resized = cv2.resize(image, dim, interpolation = inter)\n\n    return resized\n\nh, w = tmp_im.shape\n\nif w &gt; h:\n     img = image_resize(tmp_im, width = size)\nelse:\n     img = image_resize(tmp_im, height = size)\n</code></pre>\n<p>What do you think?</p>",
      "rawMarkdown": "Thank you very much for mentioning. Great work!\nI thnink that one more dataset is required 😁 - images with preserved aspect ratio.\n\n```\ndef image_resize(image, width = None, height = None, inter = cv2.INTER_LINEAR):\n\n    dim = None\n    (h, w) = image.shape[:2]\n\n    if width is None and height is None:\n        return image\n    \n    if width is None:\n        r = height / float(h)\n        dim = (int(w * r), height)\n    else:\n        r = width / float(w)\n        dim = (width, int(h * r))\n    resized = cv2.resize(image, dim, interpolation = inter)\n    \n    return resized\n\nh, w = tmp_im.shape\n\nif w > h:\n     img = image_resize(tmp_im, width = size)\nelse:\n     img = image_resize(tmp_im, height = size)\n```\nWhat do you think?",
      "votes": 1,
      "replies": [
        {
          "id": 2059768,
          "postDate": "2022-12-09T07:51:23.997Z",
          "content": "<p>Working on it! Especially for you, <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a>! 🙂 </p>\n<p>Was about to delete that instance, but guess will squeeze in processing the dataset this way as well really quickly 🙂</p>\n<p>It will have the full thing (VOI LUT processing as well) and will be resized to the longer side being 1024 -- this way if someone needs the images smaller, they should be able to resize again without too much loss of quality.</p>\n<p>Should be ready in probably a couple of hours given how slow on my end 🙂</p>",
          "rawMarkdown": "Working on it! Especially for you, @remekkinas! 🙂 \n\nWas about to delete that instance, but guess will squeeze in processing the dataset this way as well really quickly 🙂\n\nIt will have the full thing (VOI LUT processing as well) and will be resized to the longer side being 1024 -- this way if someone needs the images smaller, they should be able to resize again without too much loss of quality.\n\nShould be ready in probably a couple of hours given how slow on my end 🙂",
          "votes": 2
        },
        {
          "id": 2060036,
          "postDate": "2022-12-09T13:53:49.957Z",
          "content": "<p>Files added, processing now 🙂 Updated the notebook as well to show the processing done 🙂 Hope this can be of use!</p>\n<p>Will update the original post I guess once the processing finishes, the dataset grew again! 😄</p>",
          "rawMarkdown": "Files added, processing now 🙂 Updated the notebook as well to show the processing done 🙂 Hope this can be of use!\n\nWill update the original post I guess once the processing finishes, the dataset grew again! 😄",
          "votes": 1
        },
        {
          "id": 2060041,
          "postDate": "2022-12-09T14:03:56.350Z",
          "content": "<p>Great work and dataset! 👍👍👍</p>",
          "rawMarkdown": "Great work and dataset! 👍👍👍",
          "votes": 1
        },
        {
          "id": 2060046,
          "postDate": "2022-12-09T14:10:46.293Z",
          "content": "<p>Thanks, <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a>! 😊 Team effort! Appreciate your suggestions and the awesome things you've found! (like <code>dicomsdl</code>!).</p>\n<p>Here is the fun bit on <code>dicomsdl</code> BTW. On a two core VM it is impossible to process the dataset with the VOI LUT transform AFAICT! (and this is the only thing we really need <code>pydicom</code> for). So I am thinking that quite likely people will start shifting to <code>dicomsdl</code>. Because of the speed.</p>\n<p>You can possibly ensemble a good amount of models if you just use <code>pydicom</code> and <code>opencv</code> (without VOI LUT) but <code>dicomsdl</code> opens up a completely new sea of possibility 🙂 Just gotta verify the images are really good, but glancing over at the results they seem okay.</p>\n<p>Anyhow, those are just some random musings, but maybe someone will find them useful! 🙂</p>\n<p>Thank you again, <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a>, for your kind words and the awesome things you shared! 🙏</p>",
          "rawMarkdown": "Thanks, @remekkinas! 😊 Team effort! Appreciate your suggestions and the awesome things you've found! (like `dicomsdl`!).\n\nHere is the fun bit on `dicomsdl` BTW. On a two core VM it is impossible to process the dataset with the VOI LUT transform AFAICT! (and this is the only thing we really need `pydicom` for). So I am thinking that quite likely people will start shifting to `dicomsdl`. Because of the speed.\n\nYou can possibly ensemble a good amount of models if you just use `pydicom` and `opencv` (without VOI LUT) but `dicomsdl` opens up a completely new sea of possibility 🙂 Just gotta verify the images are really good, but glancing over at the results they seem okay.\n\nAnyhow, those are just some random musings, but maybe someone will find them useful! 🙂\n\nThank you again, @remekkinas, for your kind words and the awesome things you shared! 🙏",
          "votes": 3
        },
        {
          "id": 2060113,
          "postDate": "2022-12-09T14:57:23.827Z",
          "content": "<p>Thanks!<br>\nThis is not my dicovery - <a href=\"https://www.kaggle.com/kaggleqrdl\" target=\"_blank\">@kaggleqrdl</a> described ( <a href=\"https://www.kaggle.com/alenic\" target=\"_blank\">@alenic</a> idea ) but I implemented it only in notebook 🙂 Speed is here important. We have more room for models/blending/TTA … etc.<br>\nOK. We have first stage (data processing) more or less under controll. We can do next step :)</p>",
          "rawMarkdown": "Thanks!\nThis is not my dicovery - @kaggleqrdl described ( @alenic idea ) but I implemented it only in notebook 🙂 Speed is here important. We have more room for models/blending/TTA ... etc.\nOK. We have first stage (data processing) more or less under controll. We can do next step :)",
          "votes": 2
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2059728,
      "author_name": "Remek Kinas",
      "author_url": "",
      "post_date": "2022-12-09T06:59:46.113000",
      "content": "<p>Thank you very much for mentioning. Great work!<br>\nI thnink that one more dataset is required 😁 - images with preserved aspect ratio.</p>\n<pre><code>def image_resize(image, width = None, height = None, inter = cv2.INTER_LINEAR):\n\n    dim = None\n    (h, w) = image.shape[:2]\n\n    if width is None and height is None:\n        return image\n\n    if width is None:\n        r = height / float(h)\n        dim = (int(w * r), height)\n    else:\n        r = width / float(w)\n        dim = (width, int(h * r))\n    resized = cv2.resize(image, dim, interpolation = inter)\n\n    return resized\n\nh, w = tmp_im.shape\n\nif w &gt; h:\n     img = image_resize(tmp_im, width = size)\nelse:\n     img = image_resize(tmp_im, height = size)\n</code></pre>\n<p>What do you think?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2059768,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-12-09T07:51:23.997000",
          "content": "<p>Working on it! Especially for you, <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a>! 🙂 </p>\n<p>Was about to delete that instance, but guess will squeeze in processing the dataset this way as well really quickly 🙂</p>\n<p>It will have the full thing (VOI LUT processing as well) and will be resized to the longer side being 1024 -- this way if someone needs the images smaller, they should be able to resize again without too much loss of quality.</p>\n<p>Should be ready in probably a couple of hours given how slow on my end 🙂</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2060036,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-12-09T13:53:49.957000",
          "content": "<p>Files added, processing now 🙂 Updated the notebook as well to show the processing done 🙂 Hope this can be of use!</p>\n<p>Will update the original post I guess once the processing finishes, the dataset grew again! 😄</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2060041,
          "author_name": "Remek Kinas",
          "author_url": "",
          "post_date": "2022-12-09T14:03:56.350000",
          "content": "<p>Great work and dataset! 👍👍👍</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2060046,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-12-09T14:10:46.293000",
          "content": "<p>Thanks, <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a>! 😊 Team effort! Appreciate your suggestions and the awesome things you've found! (like <code>dicomsdl</code>!).</p>\n<p>Here is the fun bit on <code>dicomsdl</code> BTW. On a two core VM it is impossible to process the dataset with the VOI LUT transform AFAICT! (and this is the only thing we really need <code>pydicom</code> for). So I am thinking that quite likely people will start shifting to <code>dicomsdl</code>. Because of the speed.</p>\n<p>You can possibly ensemble a good amount of models if you just use <code>pydicom</code> and <code>opencv</code> (without VOI LUT) but <code>dicomsdl</code> opens up a completely new sea of possibility 🙂 Just gotta verify the images are really good, but glancing over at the results they seem okay.</p>\n<p>Anyhow, those are just some random musings, but maybe someone will find them useful! 🙂</p>\n<p>Thank you again, <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a>, for your kind words and the awesome things you shared! 🙏</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 2060113,
          "author_name": "Remek Kinas",
          "author_url": "",
          "post_date": "2022-12-09T14:57:23.827000",
          "content": "<p>Thanks!<br>\nThis is not my dicovery - <a href=\"https://www.kaggle.com/kaggleqrdl\" target=\"_blank\">@kaggleqrdl</a> described ( <a href=\"https://www.kaggle.com/alenic\" target=\"_blank\">@alenic</a> idea ) but I implemented it only in notebook 🙂 Speed is here important. We have more room for models/blending/TTA … etc.<br>\nOK. We have first stage (data processing) more or less under controll. We can do next step :)</p>",
          "votes": 2,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2059535": "Hey!\n\nI wanted to give you an update on the [RSNA Mammo PNGs 256px, 384px, 512px, 768px, 1024px](https://www.kaggle.com/datasets/radek1/rsna-mammography-images-as-pngs) that I know many people have been using!\n\n(can also be useful to people who just joined the competition)\n\nI kept adding to the dataset so it now has a decent collection of sizes processed in various ways 🙂\n\nAll processing is documented in [💡 how to process DICOM images to PNGs](https://www.kaggle.com/code/radek1/how-to-process-dicom-images-to-pngs) so you can grab it and straight away implement it in your inference notebook!\n\nThe four processing methods are as follows:\n* processing using `PIL` - `PIL` is a great choice but turned out for me to be 10%-20% slower than `opencv`, so switched to that\n* processing using `opencv` - like above, but faster, using `opencv` (these methods normalize images to between 0 and 1 and account fo the `MONOCHROME1` status)\n* processing using `opencv` with VOI LUT transformation (thank you @raddar 🙏)\n* processing using `opencv` without VOI LUT transformation AND with reading the data in by `dicomsdl` (2x faster then `pydicom`!)  (thank you @remekkinas! 🙏) 🔥🔥🔥 \n* processing using `opencv` with keeping the aspect ratio, resizing the longer side to 1024 (please see discussion in this thread on this \nfor further information)\n\nHope you find this useful in working on the solution! 🙂\n\n## Other resources you might find useful:\n\n* [💡 how to process DICOM images to PNGs](https://www.kaggle.com/code/radek1/how-to-process-dicom-images-to-pngs)\n* [📊 EDA + training a fast.ai model + submission 🚀](https://www.kaggle.com/code/radek1/eda-training-a-fast-ai-model-submission)\n* [🤖 [fast.ai starter pack] train + inference 🚀](https://www.kaggle.com/code/radek1/fast-ai-starter-pack-train-inference)\n* [📸 Over 56GB of processed data, 5 different methods 🥳](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/371268)\n* [3 resources to get started with Computer Vision in this competition 🚀🚀🚀](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369706)\n* [💡 6 Computer Vision tricks for faster training and better models 🚀🚀🚀](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369155)",
    "2059728": "Thank you very much for mentioning. Great work!\nI thnink that one more dataset is required 😁 - images with preserved aspect ratio.\n\n```\ndef image_resize(image, width = None, height = None, inter = cv2.INTER_LINEAR):\n\n    dim = None\n    (h, w) = image.shape[:2]\n\n    if width is None and height is None:\n        return image\n    \n    if width is None:\n        r = height / float(h)\n        dim = (int(w * r), height)\n    else:\n        r = width / float(w)\n        dim = (width, int(h * r))\n    resized = cv2.resize(image, dim, interpolation = inter)\n    \n    return resized\n\nh, w = tmp_im.shape\n\nif w > h:\n     img = image_resize(tmp_im, width = size)\nelse:\n     img = image_resize(tmp_im, height = size)\n```\nWhat do you think?"
  }
}