{
  "id": 336194,
  "title": "Detect background patches with OpenCV",
  "url": "/competitions/mayo-clinic-strip-ai/discussion/336194",
  "author_name": "RaghavPrabhakar",
  "post_date": "2022-07-09T21:06:32.456000",
  "votes": 6,
  "comment_count": 2,
  "views": 0,
  "content": "<p>As the images are very large, the current sensible approach seems to be divide the image into patches and perform classification on those patches. Thanks to <a href=\"https://www.kaggle.com/robikscube\" target=\"_blank\">@robikscube</a> we have a dataset of images divided into patched of 1024 X 1024 <a href=\"https://www.kaggle.com/competitions/mayo-clinic-strip-ai/discussion/335755\" target=\"_blank\">here</a>. But this leads to another problem that most patches are white backgrounds which needs to be removed as discussed here. One solution could be train a neural network to identify between background and non-background images and thanks to <a href=\"https://www.kaggle.com/moth\" target=\"_blank\">@moth</a> we have a labelled <a href=\"https://www.kaggle.com/datasets/alejopaullier/strip-ai-background-clot\" target=\"_blank\">dataset</a> of 20k images for that.</p>\n<p>Another solution could be to convert image into grayscale and calculate the ratio of white pixels over whole image and if that ratio is below certain threshold then we discard that image.</p>\n<pre><code>def check_background(filepath, threshold=5, display=False, ax=None):\n    image = cv2.imread(filepath)\n    h, w, _ = image.shape\n    gray = cv2.cvtColor(image, cv2.COLOR_BGR2GRAY)\n    thresh = cv2.threshold(gray, 0, 255, cv2.THRESH_OTSU + cv2.THRESH_BINARY_INV)[1]\n\n    pixels = cv2.countNonZero(thresh)\n    ratio = (pixels/(h * w)) * 100\n    #print('Pixel ratio: {:.2f}%'.format(ratio))\n    roi = 0\n    if ratio &gt;= threshold:\n        roi = 1\n\n    if display and ax is not None:\n        ax.imshow(thresh)\n        ax.set_title('Mostly Background' if not roi else 'Contains region of interest')\n        ax.axis('off')\n\n    return roi\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5710134%2F7b454bf5c175944c99bb8b336ab50d6e%2FScreenshot%20from%202022-07-10%2002-26-47.png?generation=1657400257781731&amp;alt=media\" alt=\"\"></p>\n<blockquote>\n  <p>With Threshold=5%, we do get Accuracy of 95% as seen in the notebook <a href=\"https://www.kaggle.com/code/raghavprabhakar66/stroke-blood-clot-classification\" target=\"_blank\">here</a>.</p>\n</blockquote>",
  "messages": [
    {
      "id": 1849837,
      "postDate": "2022-07-09T21:06:32.457Z",
      "content": "<p>As the images are very large, the current sensible approach seems to be divide the image into patches and perform classification on those patches. Thanks to <a href=\"https://www.kaggle.com/robikscube\" target=\"_blank\">@robikscube</a> we have a dataset of images divided into patched of 1024 X 1024 <a href=\"https://www.kaggle.com/competitions/mayo-clinic-strip-ai/discussion/335755\" target=\"_blank\">here</a>. But this leads to another problem that most patches are white backgrounds which needs to be removed as discussed here. One solution could be train a neural network to identify between background and non-background images and thanks to <a href=\"https://www.kaggle.com/moth\" target=\"_blank\">@moth</a> we have a labelled <a href=\"https://www.kaggle.com/datasets/alejopaullier/strip-ai-background-clot\" target=\"_blank\">dataset</a> of 20k images for that.</p>\n<p>Another solution could be to convert image into grayscale and calculate the ratio of white pixels over whole image and if that ratio is below certain threshold then we discard that image.</p>\n<pre><code>def check_background(filepath, threshold=5, display=False, ax=None):\n    image = cv2.imread(filepath)\n    h, w, _ = image.shape\n    gray = cv2.cvtColor(image, cv2.COLOR_BGR2GRAY)\n    thresh = cv2.threshold(gray, 0, 255, cv2.THRESH_OTSU + cv2.THRESH_BINARY_INV)[1]\n\n    pixels = cv2.countNonZero(thresh)\n    ratio = (pixels/(h * w)) * 100\n    #print('Pixel ratio: {:.2f}%'.format(ratio))\n    roi = 0\n    if ratio &gt;= threshold:\n        roi = 1\n\n    if display and ax is not None:\n        ax.imshow(thresh)\n        ax.set_title('Mostly Background' if not roi else 'Contains region of interest')\n        ax.axis('off')\n\n    return roi\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5710134%2F7b454bf5c175944c99bb8b336ab50d6e%2FScreenshot%20from%202022-07-10%2002-26-47.png?generation=1657400257781731&amp;alt=media\" alt=\"\"></p>\n<blockquote>\n  <p>With Threshold=5%, we do get Accuracy of 95% as seen in the notebook <a href=\"https://www.kaggle.com/code/raghavprabhakar66/stroke-blood-clot-classification\" target=\"_blank\">here</a>.</p>\n</blockquote>",
      "rawMarkdown": "As the images are very large, the current sensible approach seems to be divide the image into patches and perform classification on those patches. Thanks to @robikscube we have a dataset of images divided into patched of 1024 X 1024 [here](https://www.kaggle.com/competitions/mayo-clinic-strip-ai/discussion/335755). But this leads to another problem that most patches are white backgrounds which needs to be removed as discussed here. One solution could be train a neural network to identify between background and non-background images and thanks to @moth we have a labelled [dataset](https://www.kaggle.com/datasets/alejopaullier/strip-ai-background-clot) of 20k images for that.\n\nAnother solution could be to convert image into grayscale and calculate the ratio of white pixels over whole image and if that ratio is below certain threshold then we discard that image.\n\n```python\ndef check_background(filepath, threshold=5, display=False, ax=None):\n    image = cv2.imread(filepath)\n    h, w, _ = image.shape\n    gray = cv2.cvtColor(image, cv2.COLOR_BGR2GRAY)\n    thresh = cv2.threshold(gray, 0, 255, cv2.THRESH_OTSU + cv2.THRESH_BINARY_INV)[1]\n\n    pixels = cv2.countNonZero(thresh)\n    ratio = (pixels/(h * w)) * 100\n    #print('Pixel ratio: {:.2f}%'.format(ratio))\n    roi = 0\n    if ratio >= threshold:\n        roi = 1\n    \n    if display and ax is not None:\n        ax.imshow(thresh)\n        ax.set_title('Mostly Background' if not roi else 'Contains region of interest')\n        ax.axis('off')\n    \n    return roi\n```\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5710134%2F7b454bf5c175944c99bb8b336ab50d6e%2FScreenshot%20from%202022-07-10%2002-26-47.png?generation=1657400257781731&alt=media)\n\n> With Threshold=5%, we do get Accuracy of 95% as seen in the notebook [here](https://www.kaggle.com/code/raghavprabhakar66/stroke-blood-clot-classification).",
      "votes": 6
    },
    {
      "id": 1849891,
      "postDate": "2022-07-09T23:15:23.667Z",
      "content": "<p>Nice work!</p>",
      "rawMarkdown": "Nice work!"
    },
    {
      "id": 1850108,
      "postDate": "2022-07-10T05:11:56.027Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1849891,
      "author_name": "moth",
      "author_url": "",
      "post_date": "2022-07-09T23:15:23.667000",
      "content": "<p>Nice work!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1850108,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-10T05:11:56.027000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1849837": "As the images are very large, the current sensible approach seems to be divide the image into patches and perform classification on those patches. Thanks to @robikscube we have a dataset of images divided into patched of 1024 X 1024 [here](https://www.kaggle.com/competitions/mayo-clinic-strip-ai/discussion/335755). But this leads to another problem that most patches are white backgrounds which needs to be removed as discussed here. One solution could be train a neural network to identify between background and non-background images and thanks to @moth we have a labelled [dataset](https://www.kaggle.com/datasets/alejopaullier/strip-ai-background-clot) of 20k images for that.\n\nAnother solution could be to convert image into grayscale and calculate the ratio of white pixels over whole image and if that ratio is below certain threshold then we discard that image.\n\n```python\ndef check_background(filepath, threshold=5, display=False, ax=None):\n    image = cv2.imread(filepath)\n    h, w, _ = image.shape\n    gray = cv2.cvtColor(image, cv2.COLOR_BGR2GRAY)\n    thresh = cv2.threshold(gray, 0, 255, cv2.THRESH_OTSU + cv2.THRESH_BINARY_INV)[1]\n\n    pixels = cv2.countNonZero(thresh)\n    ratio = (pixels/(h * w)) * 100\n    #print('Pixel ratio: {:.2f}%'.format(ratio))\n    roi = 0\n    if ratio >= threshold:\n        roi = 1\n    \n    if display and ax is not None:\n        ax.imshow(thresh)\n        ax.set_title('Mostly Background' if not roi else 'Contains region of interest')\n        ax.axis('off')\n    \n    return roi\n```\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5710134%2F7b454bf5c175944c99bb8b336ab50d6e%2FScreenshot%20from%202022-07-10%2002-26-47.png?generation=1657400257781731&alt=media)\n\n> With Threshold=5%, we do get Accuracy of 95% as seen in the notebook [here](https://www.kaggle.com/code/raghavprabhakar66/stroke-blood-clot-classification).",
    "1849891": "Nice work!",
    "1850108": ""
  }
}