{
  "id": 336067,
  "title": "CE images are twice the size of LAA images (on average)",
  "url": "/competitions/mayo-clinic-strip-ai/discussion/336067",
  "author_name": "Nikita Kuzmenkov",
  "post_date": "2022-07-09T07:43:59.287000",
  "votes": 15,
  "comment_count": 2,
  "views": 0,
  "content": "<p>As you know, LAA class is nearly 2.6 times less frequent than CE class (547 CE cases vs only 207 LAA cases).</p>\n<p>Furthermore, if you count the number of 1024x1024 tiles in CE and LAA images, you'll get 417,431 and 211,215 tiles, respectively.</p>\n<p>Thus you get only one LAA sample per 5.22 CE samples in your training set, which is a way more severe imbalance than one you get by just counting the whole images. </p>\n<p>Taking this into consideration your real <strong>accuracy baseline is 0.81</strong> and not 0.62.</p>\n<h2>Update</h2>\n<p>It turned out that CE images have more whitespace on average, so relative difference between the number of samples between LAA and CE classes almost wanes to the original (1 to 2.6) after filtering out blank tiles. I find it best to use <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> technique:</p>\n<pre><code>import cv2\n\nIMG_SIZE = 256\nMIN_SATURATION = 40\nMIN_PIXELS = 1000 * (IMG_SIZE // 256) ** 2\n\ntile = ...  # tile to be checked\n\nhsv = cv2.cvtColor(tile, cv2.COLOR_RGB2HSV)\n_, s, _ = cv2.split(hsv)\n\nif ((s &gt; MIN_SATURATION).sum() &lt; MIN_PIXELS) or (tile.sum() &lt; MIN_PIXELS):\n    print(\"Tile is blank!\")\n</code></pre>\n<p>So make sure you have filtered out blank samples to prevent your model from overfitting.</p>\n<p>I applied this technique in my <strong><a href=\"https://www.kaggle.com/code/nickuzmenkov/strip-ai-eda-data-preparation\" target=\"_blank\">STRIP AI - EDA &amp; Data Preparation</a></strong> notebook and published filtered tiles to a <strong><a href=\"https://www.kaggle.com/datasets/nickuzmenkov/strip-ai-256x256-png-tiles\" target=\"_blank\">STRIP AI - 256x256 PNG Tiles</a></strong> dataset.</p>",
  "messages": [
    {
      "id": 1849090,
      "postDate": "2022-07-09T07:43:59.287Z",
      "content": "<p>As you know, LAA class is nearly 2.6 times less frequent than CE class (547 CE cases vs only 207 LAA cases).</p>\n<p>Furthermore, if you count the number of 1024x1024 tiles in CE and LAA images, you'll get 417,431 and 211,215 tiles, respectively.</p>\n<p>Thus you get only one LAA sample per 5.22 CE samples in your training set, which is a way more severe imbalance than one you get by just counting the whole images. </p>\n<p>Taking this into consideration your real <strong>accuracy baseline is 0.81</strong> and not 0.62.</p>\n<h2>Update</h2>\n<p>It turned out that CE images have more whitespace on average, so relative difference between the number of samples between LAA and CE classes almost wanes to the original (1 to 2.6) after filtering out blank tiles. I find it best to use <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> technique:</p>\n<pre><code>import cv2\n\nIMG_SIZE = 256\nMIN_SATURATION = 40\nMIN_PIXELS = 1000 * (IMG_SIZE // 256) ** 2\n\ntile = ...  # tile to be checked\n\nhsv = cv2.cvtColor(tile, cv2.COLOR_RGB2HSV)\n_, s, _ = cv2.split(hsv)\n\nif ((s &gt; MIN_SATURATION).sum() &lt; MIN_PIXELS) or (tile.sum() &lt; MIN_PIXELS):\n    print(\"Tile is blank!\")\n</code></pre>\n<p>So make sure you have filtered out blank samples to prevent your model from overfitting.</p>\n<p>I applied this technique in my <strong><a href=\"https://www.kaggle.com/code/nickuzmenkov/strip-ai-eda-data-preparation\" target=\"_blank\">STRIP AI - EDA &amp; Data Preparation</a></strong> notebook and published filtered tiles to a <strong><a href=\"https://www.kaggle.com/datasets/nickuzmenkov/strip-ai-256x256-png-tiles\" target=\"_blank\">STRIP AI - 256x256 PNG Tiles</a></strong> dataset.</p>",
      "rawMarkdown": "As you know, LAA class is nearly 2.6 times less frequent than CE class (547 CE cases vs only 207 LAA cases).\n\nFurthermore, if you count the number of 1024x1024 tiles in CE and LAA images, you'll get 417,431 and 211,215 tiles, respectively.\n\nThus you get only one LAA sample per 5.22 CE samples in your training set, which is a way more severe imbalance than one you get by just counting the whole images. \n\nTaking this into consideration your real **accuracy baseline is 0.81** and not 0.62.\n\n## Update\n\nIt turned out that CE images have more whitespace on average, so relative difference between the number of samples between LAA and CE classes almost wanes to the original (1 to 2.6) after filtering out blank tiles. I find it best to use @iafoss technique:\n\n```python\nimport cv2\n\nIMG_SIZE = 256\nMIN_SATURATION = 40\nMIN_PIXELS = 1000 * (IMG_SIZE // 256) ** 2\n\ntile = ...  # tile to be checked\n\nhsv = cv2.cvtColor(tile, cv2.COLOR_RGB2HSV)\n_, s, _ = cv2.split(hsv)\n    \nif ((s > MIN_SATURATION).sum() < MIN_PIXELS) or (tile.sum() < MIN_PIXELS):\n    print(\"Tile is blank!\")\n```\nSo make sure you have filtered out blank samples to prevent your model from overfitting.\n\nI applied this technique in my **[STRIP AI - EDA & Data Preparation][1]** notebook and published filtered tiles to a **[STRIP AI - 256x256 PNG Tiles][2]** dataset.\n\n[1]: https://www.kaggle.com/code/nickuzmenkov/strip-ai-eda-data-preparation\n[2]: https://www.kaggle.com/datasets/nickuzmenkov/strip-ai-256x256-png-tiles",
      "votes": 15
    },
    {
      "id": 1856834,
      "postDate": "2022-07-15T16:24:47.513Z",
      "content": "<p>good finding</p>",
      "rawMarkdown": "good finding",
      "votes": 2
    },
    {
      "id": 1849934,
      "postDate": "2022-07-10T00:15:07.737Z",
      "content": "<p>This is a really interesting observation.<br>\nI think you're right that the imbalance in the data is one of the factors that makes this a challenging problem.</p>",
      "rawMarkdown": "This is a really interesting observation.\nI think you're right that the imbalance in the data is one of the factors that makes this a challenging problem.\n"
    }
  ],
  "comments": [
    {
      "id": 1856834,
      "author_name": "makakinho",
      "author_url": "",
      "post_date": "2022-07-15T16:24:47.513000",
      "content": "<p>good finding</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1849934,
      "author_name": "The Devastator",
      "author_url": "",
      "post_date": "2022-07-10T00:15:07.737000",
      "content": "<p>This is a really interesting observation.<br>\nI think you're right that the imbalance in the data is one of the factors that makes this a challenging problem.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1849090": "As you know, LAA class is nearly 2.6 times less frequent than CE class (547 CE cases vs only 207 LAA cases).\n\nFurthermore, if you count the number of 1024x1024 tiles in CE and LAA images, you'll get 417,431 and 211,215 tiles, respectively.\n\nThus you get only one LAA sample per 5.22 CE samples in your training set, which is a way more severe imbalance than one you get by just counting the whole images. \n\nTaking this into consideration your real **accuracy baseline is 0.81** and not 0.62.\n\n## Update\n\nIt turned out that CE images have more whitespace on average, so relative difference between the number of samples between LAA and CE classes almost wanes to the original (1 to 2.6) after filtering out blank tiles. I find it best to use @iafoss technique:\n\n```python\nimport cv2\n\nIMG_SIZE = 256\nMIN_SATURATION = 40\nMIN_PIXELS = 1000 * (IMG_SIZE // 256) ** 2\n\ntile = ...  # tile to be checked\n\nhsv = cv2.cvtColor(tile, cv2.COLOR_RGB2HSV)\n_, s, _ = cv2.split(hsv)\n    \nif ((s > MIN_SATURATION).sum() < MIN_PIXELS) or (tile.sum() < MIN_PIXELS):\n    print(\"Tile is blank!\")\n```\nSo make sure you have filtered out blank samples to prevent your model from overfitting.\n\nI applied this technique in my **[STRIP AI - EDA & Data Preparation][1]** notebook and published filtered tiles to a **[STRIP AI - 256x256 PNG Tiles][2]** dataset.\n\n[1]: https://www.kaggle.com/code/nickuzmenkov/strip-ai-eda-data-preparation\n[2]: https://www.kaggle.com/datasets/nickuzmenkov/strip-ai-256x256-png-tiles",
    "1856834": "good finding",
    "1849934": "This is a really interesting observation.\nI think you're right that the imbalance in the data is one of the factors that makes this a challenging problem.\n"
  }
}