{
  "id": 376698,
  "title": "Custom Preprocessor for RSNA Mammography Competition",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/376698",
  "author_name": "Paul Bacher",
  "post_date": "2023-01-07T21:59:02.516000",
  "votes": 5,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I want to share with you the custom class I created to facilitate the preprocessing of the competition data. This is my first time participating in a competition and making it helped me to understand the basics of preprocessing, as well as how to work with DICOM files and mammograms.</p>\n<p>For now, it permits to load and preprocess mammography images, including the following steps:</p>\n<ul>\n<li>Applying windowing</li>\n<li>Rescaling and normalizing pixel values</li>\n<li>Flipping the breast side</li>\n<li>Cropping</li>\n<li>Resizing</li>\n</ul>\n<p>You can save your preprocessed images as PNG files in the original folder structure.<br>\nThe images are preprocessed from a list of the DICOM filepaths. I also wrote a function that make sthis list from the CSV file.</p>\n<p><a href=\"https://www.kaggle.com/code/paulbacher/custom-class-preprocessing-rsna-breast-cancer#Custom-class:-MammographyPreprocessor\" target=\"_blank\"><strong>You can access my notebook here</strong></a>.</p>\n<p>This is a preview of a few images and there associated preprocessed version:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11265932%2Fa6963228f9d84657be7ca97e0cda9f3f%2Fbefore_preprocessing.png?generation=1673123435246287&amp;alt=media\" alt=\"Before preprocessing\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11265932%2Ff114ca568d46d669068d0de42041212a%2Fafter_preprocessing.png?generation=1673123473463955&amp;alt=media\" alt=\"After preprocessing\"></p>\n<p>I will make this custom class available from my GitHub in the future.</p>\n<p>Do you have maybe any suggestion regarding the preprocessing steps that could improve the preprocessing?</p>\n<p><strong>Also:</strong><br>\nI noticed that when reshaping the images to squares (256x256 or 512x512), the compression is not even along the horizontal and vertical direction (the cropped images average ratio is around 2). I discussed about it in the notebook and choosing another ratio could improve the model (?). However, I wonder which pre-trained CNN-like model I can use that does not require a square input shape.<br>\nThis might be pointless but I wanted to have your opinion.</p>\n<p>I look forward for your feedbacks! I want to improve my skills.</p>\n<p>Thanks! 👍 </p>",
  "messages": [
    {
      "id": 2090988,
      "postDate": "2023-01-07T22:13:11.103Z",
      "content": "<p>Great work! <br>\nYes, I have one idea to improve it - do not scale it to square. If you scale such way images are distorted. Better is to create dataset using oryginal aspect ratio. Then people can decide if rescale images to square or rectangle or create letterbox (padd them).</p>\n<pre><code>def image_resize(image, width = None, height = None, inter = cv2.INTER_LINEAR):\n\n    dim = None\n    (h, w) = image.shape[:2]\n\n    if width is None and height is None:\n        return image\n\n    if width is None:\n        r = height / float(h)\n        dim = (int(w * r), height)\n    else:\n        r = width / float(w)\n        dim = (width, int(h * r))\n    resized = cv2.resize(image, dim, interpolation = inter)\n\n    return resized\n\nsize = 1024\n\nimg = cv2.imread(\"test.png\")\nh,w, _ = img.shape\n\nif h &gt; w:\n   out_img = image_resize(img, height = size)\nelse:\n   out_img = image_resize(img, width = size)\n</code></pre>",
      "rawMarkdown": "Great work! \nYes, I have one idea to improve it - do not scale it to square. If you scale such way images are distorted. Better is to create dataset using oryginal aspect ratio. Then people can decide if rescale images to square or rectangle or create letterbox (padd them).\n\n```\ndef image_resize(image, width = None, height = None, inter = cv2.INTER_LINEAR):\n\n    dim = None\n    (h, w) = image.shape[:2]\n\n    if width is None and height is None:\n        return image\n    \n    if width is None:\n        r = height / float(h)\n        dim = (int(w * r), height)\n    else:\n        r = width / float(w)\n        dim = (width, int(h * r))\n    resized = cv2.resize(image, dim, interpolation = inter)\n    \n    return resized\n\nsize = 1024\n\nimg = cv2.imread(\"test.png\")\nh,w, _ = img.shape\n\nif h > w:\n   out_img = image_resize(img, height = size)\nelse:\n   out_img = image_resize(img, width = size)\n```",
      "votes": 3,
      "replies": [
        {
          "id": 2091012,
          "postDate": "2023-01-07T22:51:29.707Z",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a>,<br>\nI did the maths! After cropping, the images have an aspect ratio that varies between 1:1,2 and 1:4!!<br>\nWhat I did in the notebook was rescaling everything to 256x512 (so 1:2) because it was the median, but I did not thought about padding them so I can reduce even more the distorsion. Gonna add this to my V2 soon 👍</p>",
          "rawMarkdown": "Thank you @remekkinas,\nI did the maths! After cropping, the images have an aspect ratio that varies between 1:1,2 and 1:4!!\nWhat I did in the notebook was rescaling everything to 256x512 (so 1:2) because it was the median, but I did not thought about padding them so I can reduce even more the distorsion. Gonna add this to my V2 soon 👍",
          "replies": [
            {
              "id": 2091416,
              "postDate": "2023-01-08T11:38:00.327Z",
              "content": "<p>yes, one case could be done better - letterbox (if you want to preserve original aspect ratio). But anyway your dataset is great! 👍</p>",
              "rawMarkdown": "yes, one case could be done better - letterbox (if you want to preserve original aspect ratio). But anyway your dataset is great! 👍",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2090978,
      "postDate": "2023-01-07T21:59:02.517Z",
      "content": "<p>I want to share with you the custom class I created to facilitate the preprocessing of the competition data. This is my first time participating in a competition and making it helped me to understand the basics of preprocessing, as well as how to work with DICOM files and mammograms.</p>\n<p>For now, it permits to load and preprocess mammography images, including the following steps:</p>\n<ul>\n<li>Applying windowing</li>\n<li>Rescaling and normalizing pixel values</li>\n<li>Flipping the breast side</li>\n<li>Cropping</li>\n<li>Resizing</li>\n</ul>\n<p>You can save your preprocessed images as PNG files in the original folder structure.<br>\nThe images are preprocessed from a list of the DICOM filepaths. I also wrote a function that make sthis list from the CSV file.</p>\n<p><a href=\"https://www.kaggle.com/code/paulbacher/custom-class-preprocessing-rsna-breast-cancer#Custom-class:-MammographyPreprocessor\" target=\"_blank\"><strong>You can access my notebook here</strong></a>.</p>\n<p>This is a preview of a few images and there associated preprocessed version:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11265932%2Fa6963228f9d84657be7ca97e0cda9f3f%2Fbefore_preprocessing.png?generation=1673123435246287&amp;alt=media\" alt=\"Before preprocessing\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11265932%2Ff114ca568d46d669068d0de42041212a%2Fafter_preprocessing.png?generation=1673123473463955&amp;alt=media\" alt=\"After preprocessing\"></p>\n<p>I will make this custom class available from my GitHub in the future.</p>\n<p>Do you have maybe any suggestion regarding the preprocessing steps that could improve the preprocessing?</p>\n<p><strong>Also:</strong><br>\nI noticed that when reshaping the images to squares (256x256 or 512x512), the compression is not even along the horizontal and vertical direction (the cropped images average ratio is around 2). I discussed about it in the notebook and choosing another ratio could improve the model (?). However, I wonder which pre-trained CNN-like model I can use that does not require a square input shape.<br>\nThis might be pointless but I wanted to have your opinion.</p>\n<p>I look forward for your feedbacks! I want to improve my skills.</p>\n<p>Thanks! 👍 </p>",
      "rawMarkdown": "I want to share with you the custom class I created to facilitate the preprocessing of the competition data. This is my first time participating in a competition and making it helped me to understand the basics of preprocessing, as well as how to work with DICOM files and mammograms.\n\nFor now, it permits to load and preprocess mammography images, including the following steps:\n- Applying windowing\n- Rescaling and normalizing pixel values\n- Flipping the breast side\n- Cropping\n- Resizing\n\nYou can save your preprocessed images as PNG files in the original folder structure.\nThe images are preprocessed from a list of the DICOM filepaths. I also wrote a function that make sthis list from the CSV file.\n\n[**You can access my notebook here**](https://www.kaggle.com/code/paulbacher/custom-class-preprocessing-rsna-breast-cancer#Custom-class:-MammographyPreprocessor).\n\nThis is a preview of a few images and there associated preprocessed version:\n![Before preprocessing](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11265932%2Fa6963228f9d84657be7ca97e0cda9f3f%2Fbefore_preprocessing.png?generation=1673123435246287&alt=media)\n![After preprocessing](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11265932%2Ff114ca568d46d669068d0de42041212a%2Fafter_preprocessing.png?generation=1673123473463955&alt=media)\n\nI will make this custom class available from my GitHub in the future.\n\nDo you have maybe any suggestion regarding the preprocessing steps that could improve the preprocessing?\n\n**Also:**\nI noticed that when reshaping the images to squares (256x256 or 512x512), the compression is not even along the horizontal and vertical direction (the cropped images average ratio is around 2). I discussed about it in the notebook and choosing another ratio could improve the model (?). However, I wonder which pre-trained CNN-like model I can use that does not require a square input shape.\nThis might be pointless but I wanted to have your opinion.\n\nI look forward for your feedbacks! I want to improve my skills.\n\nThanks! 👍 ",
      "votes": 4
    }
  ],
  "comments": [
    {
      "id": 2090988,
      "author_name": "Remek Kinas",
      "author_url": "",
      "post_date": "2023-01-07T22:13:11.103000",
      "content": "<p>Great work! <br>\nYes, I have one idea to improve it - do not scale it to square. If you scale such way images are distorted. Better is to create dataset using oryginal aspect ratio. Then people can decide if rescale images to square or rectangle or create letterbox (padd them).</p>\n<pre><code>def image_resize(image, width = None, height = None, inter = cv2.INTER_LINEAR):\n\n    dim = None\n    (h, w) = image.shape[:2]\n\n    if width is None and height is None:\n        return image\n\n    if width is None:\n        r = height / float(h)\n        dim = (int(w * r), height)\n    else:\n        r = width / float(w)\n        dim = (width, int(h * r))\n    resized = cv2.resize(image, dim, interpolation = inter)\n\n    return resized\n\nsize = 1024\n\nimg = cv2.imread(\"test.png\")\nh,w, _ = img.shape\n\nif h &gt; w:\n   out_img = image_resize(img, height = size)\nelse:\n   out_img = image_resize(img, width = size)\n</code></pre>",
      "votes": 3,
      "replies": [
        {
          "id": 2091012,
          "author_name": "Paul Bacher",
          "author_url": "",
          "post_date": "2023-01-07T22:51:29.707000",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a>,<br>\nI did the maths! After cropping, the images have an aspect ratio that varies between 1:1,2 and 1:4!!<br>\nWhat I did in the notebook was rescaling everything to 256x512 (so 1:2) because it was the median, but I did not thought about padding them so I can reduce even more the distorsion. Gonna add this to my V2 soon 👍</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2091416,
              "author_name": "Remek Kinas",
              "author_url": "",
              "post_date": "2023-01-08T11:38:00.327000",
              "content": "<p>yes, one case could be done better - letterbox (if you want to preserve original aspect ratio). But anyway your dataset is great! 👍</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2090988": "Great work! \nYes, I have one idea to improve it - do not scale it to square. If you scale such way images are distorted. Better is to create dataset using oryginal aspect ratio. Then people can decide if rescale images to square or rectangle or create letterbox (padd them).\n\n```\ndef image_resize(image, width = None, height = None, inter = cv2.INTER_LINEAR):\n\n    dim = None\n    (h, w) = image.shape[:2]\n\n    if width is None and height is None:\n        return image\n    \n    if width is None:\n        r = height / float(h)\n        dim = (int(w * r), height)\n    else:\n        r = width / float(w)\n        dim = (width, int(h * r))\n    resized = cv2.resize(image, dim, interpolation = inter)\n    \n    return resized\n\nsize = 1024\n\nimg = cv2.imread(\"test.png\")\nh,w, _ = img.shape\n\nif h > w:\n   out_img = image_resize(img, height = size)\nelse:\n   out_img = image_resize(img, width = size)\n```",
    "2090978": "I want to share with you the custom class I created to facilitate the preprocessing of the competition data. This is my first time participating in a competition and making it helped me to understand the basics of preprocessing, as well as how to work with DICOM files and mammograms.\n\nFor now, it permits to load and preprocess mammography images, including the following steps:\n- Applying windowing\n- Rescaling and normalizing pixel values\n- Flipping the breast side\n- Cropping\n- Resizing\n\nYou can save your preprocessed images as PNG files in the original folder structure.\nThe images are preprocessed from a list of the DICOM filepaths. I also wrote a function that make sthis list from the CSV file.\n\n[**You can access my notebook here**](https://www.kaggle.com/code/paulbacher/custom-class-preprocessing-rsna-breast-cancer#Custom-class:-MammographyPreprocessor).\n\nThis is a preview of a few images and there associated preprocessed version:\n![Before preprocessing](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11265932%2Fa6963228f9d84657be7ca97e0cda9f3f%2Fbefore_preprocessing.png?generation=1673123435246287&alt=media)\n![After preprocessing](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11265932%2Ff114ca568d46d669068d0de42041212a%2Fafter_preprocessing.png?generation=1673123473463955&alt=media)\n\nI will make this custom class available from my GitHub in the future.\n\nDo you have maybe any suggestion regarding the preprocessing steps that could improve the preprocessing?\n\n**Also:**\nI noticed that when reshaping the images to squares (256x256 or 512x512), the compression is not even along the horizontal and vertical direction (the cropped images average ratio is around 2). I discussed about it in the notebook and choosing another ratio could improve the model (?). However, I wonder which pre-trained CNN-like model I can use that does not require a square input shape.\nThis might be pointless but I wanted to have your opinion.\n\nI look forward for your feedbacks! I want to improve my skills.\n\nThanks! 👍 "
  }
}