{
  "id": 461049,
  "title": "Question about  enhance the speed of WSI (Whole Slide Image) patch extraction",
  "url": "/competitions/UBC-OCEAN/discussion/461049",
  "author_name": "Seeing Times",
  "post_date": "2023-12-12T11:37:45.445000",
  "votes": 3,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Extraction patches for WSI consumed the majority of time, Is any solution to reduce run time？<br>\n`def extract_image_tiles(p_img, tumor_model, wsi_size: int = 512, tma_size: int = 1024, scale: float = 0.5, <br>\n                        drop_thr: float = 0.7, white_drop_thr: float = 0.7, white_thr: int = 230, <br>\n                        wsi_thr: int = 5500, max_samples: int = 2500) -&gt; list:<br>\n    im = pyvips.Image.new_from_file(p_img)<br>\n    img_name = os.path.splitext(os.path.basename(p_img))[0]  </p>\n<pre><code> im. &lt; wsi_thr  im. &lt; wsi_thr:\n    size = tma_size\n     =   \n:\n    size = wsi_size \n\nim = im.resize()\nw = h = int(size*)\n\nidxs = [(y, y + h, x, x + w)  y  (, im., h)  x  (, im., w)]\n max_samples &lt; len(idxs):\n    idxs = .sample(idxs, max_samples)\ntiles = []\n\n k, (y, y_, x, x_)  enumerate(idxs):\n    tile = im.crop(x, y, (w, im. - x), (h, im. - y)).numpy()[..., :]\n\n     tile.shape[:] != (h, w):\n        continue\n\n     size == tma_size:\n        white_bg = .all(tile &gt;= white_thr, axis=)\n         .(white_bg) &gt;= (.prod(white_bg.shape) * white_drop_thr):\n            continue\n\n\n     size == wsi_size:\n        mask_bg = (.(tile, axis=) &lt;= ) | (.(tile, axis=) &gt;= )\n         .(mask_bg) &gt;= (.prod(mask_bg.shape) * drop_thr):\n            continue\n\n    tile_pil = Image.fromarray(tile.astype(.uint8))\n    tile_tensor_transformed = data_transforms()(tile_pil).unsqueeze().to(CFG.device)\n\n\n img_name, tiles`\n</code></pre>",
  "messages": [
    {
      "id": 2558834,
      "postDate": "2023-12-12T11:37:45.447Z",
      "content": "<p>Extraction patches for WSI consumed the majority of time, Is any solution to reduce run time？<br>\n`def extract_image_tiles(p_img, tumor_model, wsi_size: int = 512, tma_size: int = 1024, scale: float = 0.5, <br>\n                        drop_thr: float = 0.7, white_drop_thr: float = 0.7, white_thr: int = 230, <br>\n                        wsi_thr: int = 5500, max_samples: int = 2500) -&gt; list:<br>\n    im = pyvips.Image.new_from_file(p_img)<br>\n    img_name = os.path.splitext(os.path.basename(p_img))[0]  </p>\n<pre><code> im. &lt; wsi_thr  im. &lt; wsi_thr:\n    size = tma_size\n     =   \n:\n    size = wsi_size \n\nim = im.resize()\nw = h = int(size*)\n\nidxs = [(y, y + h, x, x + w)  y  (, im., h)  x  (, im., w)]\n max_samples &lt; len(idxs):\n    idxs = .sample(idxs, max_samples)\ntiles = []\n\n k, (y, y_, x, x_)  enumerate(idxs):\n    tile = im.crop(x, y, (w, im. - x), (h, im. - y)).numpy()[..., :]\n\n     tile.shape[:] != (h, w):\n        continue\n\n     size == tma_size:\n        white_bg = .all(tile &gt;= white_thr, axis=)\n         .(white_bg) &gt;= (.prod(white_bg.shape) * white_drop_thr):\n            continue\n\n\n     size == wsi_size:\n        mask_bg = (.(tile, axis=) &lt;= ) | (.(tile, axis=) &gt;= )\n         .(mask_bg) &gt;= (.prod(mask_bg.shape) * drop_thr):\n            continue\n\n    tile_pil = Image.fromarray(tile.astype(.uint8))\n    tile_tensor_transformed = data_transforms()(tile_pil).unsqueeze().to(CFG.device)\n\n\n img_name, tiles`\n</code></pre>",
      "rawMarkdown": "Extraction patches for WSI consumed the majority of time, Is any solution to reduce run time？\n`def extract_image_tiles(p_img, tumor_model, wsi_size: int = 512, tma_size: int = 1024, scale: float = 0.5, \n                        drop_thr: float = 0.7, white_drop_thr: float = 0.7, white_thr: int = 230, \n                        wsi_thr: int = 5500, max_samples: int = 2500) -> list:\n    im = pyvips.Image.new_from_file(p_img)\n    img_name = os.path.splitext(os.path.basename(p_img))[0]  \n\n    if im.width < wsi_thr and im.height < wsi_thr:\n        size = tma_size\n        scale = 0.25  \n    else:\n        size = wsi_size \n    \n    im = im.resize(scale)\n    w = h = int(size*scale)\n    \n    idxs = [(y, y + h, x, x + w) for y in range(0, im.height, h) for x in range(0, im.width, w)]\n    if max_samples < len(idxs):\n        idxs = random.sample(idxs, max_samples)\n    tiles = []\n\n    for k, (y, y_, x, x_) in enumerate(idxs):\n        tile = im.crop(x, y, min(w, im.width - x), min(h, im.height - y)).numpy()[..., :3]\n        \n        if tile.shape[:2] != (h, w):\n            continue\n        \n        if size == tma_size:\n            white_bg = np.all(tile >= white_thr, axis=2)\n            if np.sum(white_bg) >= (np.prod(white_bg.shape) * white_drop_thr):\n                continue\n\n    \n        if size == wsi_size:\n            mask_bg = (np.sum(tile, axis=2) <= 10) | (np.max(tile, axis=2) >= 230)\n            if np.sum(mask_bg) >= (np.prod(mask_bg.shape) * drop_thr):\n                continue\n\n        tile_pil = Image.fromarray(tile.astype(np.uint8))\n        tile_tensor_transformed = data_transforms()(tile_pil).unsqueeze(0).to(CFG.device)\n\n            \n    return img_name, tiles`",
      "votes": 3
    },
    {
      "id": 2561316,
      "postDate": "2023-12-14T12:09:57.140Z",
      "content": "<p>first identify which line of your code takes too much time, I assume it is converting pyvips tiles to numpy array, <br>\nyou may convert your for loop into a concurrent function (naive concurrent package from python),</p>\n<p>since now notebooks give 4 core cpu, so you can process a tile on each core</p>",
      "rawMarkdown": "first identify which line of your code takes too much time, I assume it is converting pyvips tiles to numpy array, \nyou may convert your for loop into a concurrent function (naive concurrent package from python),\n\nsince now notebooks give 4 core cpu, so you can process a tile on each core",
      "votes": 2,
      "replies": [
        {
          "id": 2561387,
          "postDate": "2023-12-14T12:53:39.733Z",
          "content": "<p>Thank you for your reply！I‘ll try it later</p>",
          "rawMarkdown": "Thank you for your reply！I‘ll try it later"
        }
      ]
    },
    {
      "id": 2584067,
      "postDate": "2024-01-02T16:47:43.700Z",
      "content": "<p>Look into python-openslide!  I got my code to tile every patch within an image by using a two step process.  One step to get tile coordinates and the next step to do the extraction.  I will be releasing my code shortly on how to do this.  I explain my solution here: <a href=\"https://www.kaggle.com/competitions/UBC-OCEAN/discussion/465030\" target=\"_blank\">https://www.kaggle.com/competitions/UBC-OCEAN/discussion/465030</a></p>",
      "rawMarkdown": "Look into python-openslide!  I got my code to tile every patch within an image by using a two step process.  One step to get tile coordinates and the next step to do the extraction.  I will be releasing my code shortly on how to do this.  I explain my solution here: https://www.kaggle.com/competitions/UBC-OCEAN/discussion/465030"
    }
  ],
  "comments": [
    {
      "id": 2561316,
      "author_name": "Ali",
      "author_url": "",
      "post_date": "2023-12-14T12:09:57.140000",
      "content": "<p>first identify which line of your code takes too much time, I assume it is converting pyvips tiles to numpy array, <br>\nyou may convert your for loop into a concurrent function (naive concurrent package from python),</p>\n<p>since now notebooks give 4 core cpu, so you can process a tile on each core</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2561387,
          "author_name": "Seeing Times",
          "author_url": "",
          "post_date": "2023-12-14T12:53:39.733000",
          "content": "<p>Thank you for your reply！I‘ll try it later</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2584067,
      "author_name": "Connor",
      "author_url": "",
      "post_date": "2024-01-02T16:47:43.700000",
      "content": "<p>Look into python-openslide!  I got my code to tile every patch within an image by using a two step process.  One step to get tile coordinates and the next step to do the extraction.  I will be releasing my code shortly on how to do this.  I explain my solution here: <a href=\"https://www.kaggle.com/competitions/UBC-OCEAN/discussion/465030\" target=\"_blank\">https://www.kaggle.com/competitions/UBC-OCEAN/discussion/465030</a></p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2558834": "Extraction patches for WSI consumed the majority of time, Is any solution to reduce run time？\n`def extract_image_tiles(p_img, tumor_model, wsi_size: int = 512, tma_size: int = 1024, scale: float = 0.5, \n                        drop_thr: float = 0.7, white_drop_thr: float = 0.7, white_thr: int = 230, \n                        wsi_thr: int = 5500, max_samples: int = 2500) -> list:\n    im = pyvips.Image.new_from_file(p_img)\n    img_name = os.path.splitext(os.path.basename(p_img))[0]  \n\n    if im.width < wsi_thr and im.height < wsi_thr:\n        size = tma_size\n        scale = 0.25  \n    else:\n        size = wsi_size \n    \n    im = im.resize(scale)\n    w = h = int(size*scale)\n    \n    idxs = [(y, y + h, x, x + w) for y in range(0, im.height, h) for x in range(0, im.width, w)]\n    if max_samples < len(idxs):\n        idxs = random.sample(idxs, max_samples)\n    tiles = []\n\n    for k, (y, y_, x, x_) in enumerate(idxs):\n        tile = im.crop(x, y, min(w, im.width - x), min(h, im.height - y)).numpy()[..., :3]\n        \n        if tile.shape[:2] != (h, w):\n            continue\n        \n        if size == tma_size:\n            white_bg = np.all(tile >= white_thr, axis=2)\n            if np.sum(white_bg) >= (np.prod(white_bg.shape) * white_drop_thr):\n                continue\n\n    \n        if size == wsi_size:\n            mask_bg = (np.sum(tile, axis=2) <= 10) | (np.max(tile, axis=2) >= 230)\n            if np.sum(mask_bg) >= (np.prod(mask_bg.shape) * drop_thr):\n                continue\n\n        tile_pil = Image.fromarray(tile.astype(np.uint8))\n        tile_tensor_transformed = data_transforms()(tile_pil).unsqueeze(0).to(CFG.device)\n\n            \n    return img_name, tiles`",
    "2561316": "first identify which line of your code takes too much time, I assume it is converting pyvips tiles to numpy array, \nyou may convert your for loop into a concurrent function (naive concurrent package from python),\n\nsince now notebooks give 4 core cpu, so you can process a tile on each core",
    "2584067": "Look into python-openslide!  I got my code to tile every patch within an image by using a two step process.  One step to get tile coordinates and the next step to do the extraction.  I will be releasing my code shortly on how to do this.  I explain my solution here: https://www.kaggle.com/competitions/UBC-OCEAN/discussion/465030"
  }
}