{
  "id": 341752,
  "title": "Build a MONAI based pipeline for this challenge",
  "url": "/competitions/mayo-clinic-strip-ai/discussion/341752",
  "author_name": "Yiheng Wang",
  "post_date": "2022-08-04T07:10:54.769000",
  "votes": 11,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I'm working on building a MONAI based pipeline for this challenge, and any progresses will be updated in this topic.<br>\nI don't have any experience in pathology image, thus I expect to learn a lot during the following serval weeks : )</p>\n<p>update 1:<br>\nI created a notebook that shows how to use MONAI transforms to load whole slide images:<br>\n<a href=\"https://www.kaggle.com/code/yiheng/load-wsi-and-extract-patches-with-monai/notebook\" target=\"_blank\">https://www.kaggle.com/code/yiheng/load-wsi-and-extract-patches-with-monai/notebook</a><br>\na whole slide image will be divided into multiple patches.<br>\nIt can be used as the input of a Multiple Instance Learning (MIL) model, which expected the shape of:<br>\n<code>[Batch, Number of patches, Channel, H, W]</code></p>\n<p>Update 2:<br>\n<br>\nI met several issues, such as data loading and poor model performance.<br>\nNow, I resized the training set into <code>n * 4096</code> (the longest dimension is 4096), and the dataset has been uploaded:<br>\n<a href=\"https://www.kaggle.com/datasets/yiheng/mayo-longest-4096\" target=\"_blank\">https://www.kaggle.com/datasets/yiheng/mayo-longest-4096</a></p>\n<p>For reference, I refered to: <a href=\"https://www.kaggle.com/code/analokamus/how-to-use-pyvips-offline\" target=\"_blank\">https://www.kaggle.com/code/analokamus/how-to-use-pyvips-offline</a> and use <code>pyvips</code> to load the image first, and then resized into <code>n * 4096</code>. The function I used is:</p>\n<pre><code>def save_resized_img(image_id):\n    image_pth = os.path.join(test_img_dir, image_id + \".tif\")\n    img = pyvips.Image.thumbnail(image_pth, 2048).numpy()\n    img = img.transpose([2, 0, 1])\n    print(img.shape)\n    np.save(os.path.join(resize_dir, f\"{image_id}.npy\"), img)\n    del img\n    gc.collect()\n</code></pre>\n<p>To be updated:</p>\n<p>an open sourced model which can at least have a reasonable performance (not sure if I'm about to do it, felt very hard so far).</p>",
  "messages": [
    {
      "id": 1883900,
      "postDate": "2022-08-04T07:10:54.770Z",
      "content": "<p>I'm working on building a MONAI based pipeline for this challenge, and any progresses will be updated in this topic.<br>\nI don't have any experience in pathology image, thus I expect to learn a lot during the following serval weeks : )</p>\n<p>update 1:<br>\nI created a notebook that shows how to use MONAI transforms to load whole slide images:<br>\n<a href=\"https://www.kaggle.com/code/yiheng/load-wsi-and-extract-patches-with-monai/notebook\" target=\"_blank\">https://www.kaggle.com/code/yiheng/load-wsi-and-extract-patches-with-monai/notebook</a><br>\na whole slide image will be divided into multiple patches.<br>\nIt can be used as the input of a Multiple Instance Learning (MIL) model, which expected the shape of:<br>\n<code>[Batch, Number of patches, Channel, H, W]</code></p>\n<p>Update 2:<br>\n<br>\nI met several issues, such as data loading and poor model performance.<br>\nNow, I resized the training set into <code>n * 4096</code> (the longest dimension is 4096), and the dataset has been uploaded:<br>\n<a href=\"https://www.kaggle.com/datasets/yiheng/mayo-longest-4096\" target=\"_blank\">https://www.kaggle.com/datasets/yiheng/mayo-longest-4096</a></p>\n<p>For reference, I refered to: <a href=\"https://www.kaggle.com/code/analokamus/how-to-use-pyvips-offline\" target=\"_blank\">https://www.kaggle.com/code/analokamus/how-to-use-pyvips-offline</a> and use <code>pyvips</code> to load the image first, and then resized into <code>n * 4096</code>. The function I used is:</p>\n<pre><code>def save_resized_img(image_id):\n    image_pth = os.path.join(test_img_dir, image_id + \".tif\")\n    img = pyvips.Image.thumbnail(image_pth, 2048).numpy()\n    img = img.transpose([2, 0, 1])\n    print(img.shape)\n    np.save(os.path.join(resize_dir, f\"{image_id}.npy\"), img)\n    del img\n    gc.collect()\n</code></pre>\n<p>To be updated:</p>\n<p>an open sourced model which can at least have a reasonable performance (not sure if I'm about to do it, felt very hard so far).</p>",
      "rawMarkdown": "I'm working on building a MONAI based pipeline for this challenge, and any progresses will be updated in this topic.\nI don't have any experience in pathology image, thus I expect to learn a lot during the following serval weeks : )\n\nupdate 1:\nI created a notebook that shows how to use MONAI transforms to load whole slide images:\nhttps://www.kaggle.com/code/yiheng/load-wsi-and-extract-patches-with-monai/notebook\na whole slide image will be divided into multiple patches.\nIt can be used as the input of a Multiple Instance Learning (MIL) model, which expected the shape of:\n`[Batch, Number of patches, Channel, H, W]`\n\nUpdate 2:\n~~I will build a MILModel based training pipeline.~~\nI met several issues, such as data loading and poor model performance.\nNow, I resized the training set into `n * 4096` (the longest dimension is 4096), and the dataset has been uploaded:\nhttps://www.kaggle.com/datasets/yiheng/mayo-longest-4096\n\nFor reference, I refered to: https://www.kaggle.com/code/analokamus/how-to-use-pyvips-offline and use `pyvips` to load the image first, and then resized into `n * 4096`. The function I used is:\n```\ndef save_resized_img(image_id):\n    image_pth = os.path.join(test_img_dir, image_id + \".tif\")\n    img = pyvips.Image.thumbnail(image_pth, 2048).numpy()\n    img = img.transpose([2, 0, 1])\n    print(img.shape)\n    np.save(os.path.join(resize_dir, f\"{image_id}.npy\"), img)\n    del img\n    gc.collect()\n```\n\nTo be updated:\n\nan open sourced model which can at least have a reasonable performance (not sure if I'm about to do it, felt very hard so far).",
      "votes": 11
    },
    {
      "id": 1884140,
      "postDate": "2022-08-04T10:16:44.977Z",
      "content": "<p>Kudos to your willingness to learn and develop your skills for the assignment! Good luck!</p>",
      "rawMarkdown": "Kudos to your willingness to learn and develop your skills for the assignment! Good luck!"
    }
  ],
  "comments": [
    {
      "id": 1884140,
      "author_name": "Ravi Ramakrishnan",
      "author_url": "",
      "post_date": "2022-08-04T10:16:44.977000",
      "content": "<p>Kudos to your willingness to learn and develop your skills for the assignment! Good luck!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1883900": "I'm working on building a MONAI based pipeline for this challenge, and any progresses will be updated in this topic.\nI don't have any experience in pathology image, thus I expect to learn a lot during the following serval weeks : )\n\nupdate 1:\nI created a notebook that shows how to use MONAI transforms to load whole slide images:\nhttps://www.kaggle.com/code/yiheng/load-wsi-and-extract-patches-with-monai/notebook\na whole slide image will be divided into multiple patches.\nIt can be used as the input of a Multiple Instance Learning (MIL) model, which expected the shape of:\n`[Batch, Number of patches, Channel, H, W]`\n\nUpdate 2:\n~~I will build a MILModel based training pipeline.~~\nI met several issues, such as data loading and poor model performance.\nNow, I resized the training set into `n * 4096` (the longest dimension is 4096), and the dataset has been uploaded:\nhttps://www.kaggle.com/datasets/yiheng/mayo-longest-4096\n\nFor reference, I refered to: https://www.kaggle.com/code/analokamus/how-to-use-pyvips-offline and use `pyvips` to load the image first, and then resized into `n * 4096`. The function I used is:\n```\ndef save_resized_img(image_id):\n    image_pth = os.path.join(test_img_dir, image_id + \".tif\")\n    img = pyvips.Image.thumbnail(image_pth, 2048).numpy()\n    img = img.transpose([2, 0, 1])\n    print(img.shape)\n    np.save(os.path.join(resize_dir, f\"{image_id}.npy\"), img)\n    del img\n    gc.collect()\n```\n\nTo be updated:\n\nan open sourced model which can at least have a reasonable performance (not sure if I'm about to do it, felt very hard so far).",
    "1884140": "Kudos to your willingness to learn and develop your skills for the assignment! Good luck!"
  }
}