{
  "id": 335976,
  "title": "How I Deal With WSI (Whole Slide Images aka. RAM Crashingly Large Images)",
  "url": "/competitions/mayo-clinic-strip-ai/discussion/335976",
  "author_name": "Darien Schettler",
  "post_date": "2022-07-08T18:48:36.076000",
  "votes": 58,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Hi there, I've seen a lot of different posts already about how to deal with the data in this competition as the images are quite large and can go OOM in RAM quite easily. Even if you had an incredibly powerful system, it would still be to your benefit to learn how to interact with these objects/images in a smart way.</p>\n<p>Let's discuss the two tools I'll be using (for now) in this competition to handle/manipulate these images (WSI).</p>\n<h2>1. <a href=\"https://openslide.org/api/python/\" target=\"_blank\"><strong>OpenSlide</strong></a></h2>\n<p>This tool is insanely useful when dealing with WSI data. I won't get into the details of WHY that do not pertain to this competition, but trust me when I say it is very useful for many pipelines related to WSI. So what makes it useful? Why use it?</p>\n<ul>\n<li>This is the main one for this comp: <strong>You can grab any patch from the image without loading the entire image into memory.</strong><ul>\n<li>This means I could tile the image without ever using more than the tile-size image's worth of RAM.</li></ul></li>\n<li>You can \"open\" an image and inspect pieces of metadata (ie. just load the header) without loading the actual image file into memory. This will let you access any metadata included in the image and at the least will give you access to the image size.</li>\n</ul>\n<h2>2. <a href=\"https://github.com/libvips/pyvips\" target=\"_blank\"><strong>pyvips</strong></a> and <a href=\"https://github.com/libvips\" target=\"_blank\"><strong>libvips</strong></a></h2>\n<p>These are new to me, but they are what allowed me to be able to quickly (relatively speaking) resize the full-size images down to a more manageable size. (3-90 seconds depending on original image resolution and desired downscale). They use virtual memory and some other stuff I don't understand at this point in time. I encourage you to read up about them.</p>\n<p>Currently, I'm installing them over the internet, but I will most likely create a 'local' Kaggle dataset so that I can use these tools without the internet.</p>\n<h2>3. ADDENDUM - <a href=\"https://www.kaggle.com/jirkaborovec\" target=\"_blank\">@jirkaborovec</a>'s <a href=\"https://www.kaggle.com/code/jirkaborovec/bloodclots-classif-eda-load-crop-images\" target=\"_blank\">Notebook Showing How to Use PIL! and more!</a></h2>\n<p>I didn't think PIL would be fast enough or memory savvy enough to handle this, but it appears as though I was wrong!<br>\nPlease check out the above notebook for details regarding using PIL to resize the images (and normalize).</p>\n<hr>\n<p>To see these tools in action check out the simple notebook I made:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/dschettler8845/mcsai-how-to-interact-with-large-tif-files\" target=\"_blank\"><strong>MCSAI - How To Interact With Large .tif Files</strong></a></li>\n</ul>\n<hr>\n<p><br></p>\n<p>A side note: I plan to use these tools to create a low-res PNG dataset that can be used alongside the OpenSlide patching tool to first explore (EDA) and later model the data. I'm still not sure yet how to approach this task, but I feel much more confident now that I have these tools. I hope this helps you too!</p>",
  "messages": [
    {
      "id": 1848611,
      "postDate": "2022-07-08T18:48:36.077Z",
      "content": "<p>Hi there, I've seen a lot of different posts already about how to deal with the data in this competition as the images are quite large and can go OOM in RAM quite easily. Even if you had an incredibly powerful system, it would still be to your benefit to learn how to interact with these objects/images in a smart way.</p>\n<p>Let's discuss the two tools I'll be using (for now) in this competition to handle/manipulate these images (WSI).</p>\n<h2>1. <a href=\"https://openslide.org/api/python/\" target=\"_blank\"><strong>OpenSlide</strong></a></h2>\n<p>This tool is insanely useful when dealing with WSI data. I won't get into the details of WHY that do not pertain to this competition, but trust me when I say it is very useful for many pipelines related to WSI. So what makes it useful? Why use it?</p>\n<ul>\n<li>This is the main one for this comp: <strong>You can grab any patch from the image without loading the entire image into memory.</strong><ul>\n<li>This means I could tile the image without ever using more than the tile-size image's worth of RAM.</li></ul></li>\n<li>You can \"open\" an image and inspect pieces of metadata (ie. just load the header) without loading the actual image file into memory. This will let you access any metadata included in the image and at the least will give you access to the image size.</li>\n</ul>\n<h2>2. <a href=\"https://github.com/libvips/pyvips\" target=\"_blank\"><strong>pyvips</strong></a> and <a href=\"https://github.com/libvips\" target=\"_blank\"><strong>libvips</strong></a></h2>\n<p>These are new to me, but they are what allowed me to be able to quickly (relatively speaking) resize the full-size images down to a more manageable size. (3-90 seconds depending on original image resolution and desired downscale). They use virtual memory and some other stuff I don't understand at this point in time. I encourage you to read up about them.</p>\n<p>Currently, I'm installing them over the internet, but I will most likely create a 'local' Kaggle dataset so that I can use these tools without the internet.</p>\n<h2>3. ADDENDUM - <a href=\"https://www.kaggle.com/jirkaborovec\" target=\"_blank\">@jirkaborovec</a>'s <a href=\"https://www.kaggle.com/code/jirkaborovec/bloodclots-classif-eda-load-crop-images\" target=\"_blank\">Notebook Showing How to Use PIL! and more!</a></h2>\n<p>I didn't think PIL would be fast enough or memory savvy enough to handle this, but it appears as though I was wrong!<br>\nPlease check out the above notebook for details regarding using PIL to resize the images (and normalize).</p>\n<hr>\n<p>To see these tools in action check out the simple notebook I made:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/dschettler8845/mcsai-how-to-interact-with-large-tif-files\" target=\"_blank\"><strong>MCSAI - How To Interact With Large .tif Files</strong></a></li>\n</ul>\n<hr>\n<p><br></p>\n<p>A side note: I plan to use these tools to create a low-res PNG dataset that can be used alongside the OpenSlide patching tool to first explore (EDA) and later model the data. I'm still not sure yet how to approach this task, but I feel much more confident now that I have these tools. I hope this helps you too!</p>",
      "rawMarkdown": "Hi there, I've seen a lot of different posts already about how to deal with the data in this competition as the images are quite large and can go OOM in RAM quite easily. Even if you had an incredibly powerful system, it would still be to your benefit to learn how to interact with these objects/images in a smart way.\n\nLet's discuss the two tools I'll be using (for now) in this competition to handle/manipulate these images (WSI).\n\n## 1. [**OpenSlide**](https://openslide.org/api/python/)\n\nThis tool is insanely useful when dealing with WSI data. I won't get into the details of WHY that do not pertain to this competition, but trust me when I say it is very useful for many pipelines related to WSI. So what makes it useful? Why use it?\n* This is the main one for this comp: **You can grab any patch from the image without loading the entire image into memory.**\n  * This means I could tile the image without ever using more than the tile-size image's worth of RAM.\n* You can \"open\" an image and inspect pieces of metadata (ie. just load the header) without loading the actual image file into memory. This will let you access any metadata included in the image and at the least will give you access to the image size.\n\n## 2. [**pyvips**](https://github.com/libvips/pyvips) and [**libvips**](https://github.com/libvips)\n\nThese are new to me, but they are what allowed me to be able to quickly (relatively speaking) resize the full-size images down to a more manageable size. (3-90 seconds depending on original image resolution and desired downscale). They use virtual memory and some other stuff I don't understand at this point in time. I encourage you to read up about them.\n\nCurrently, I'm installing them over the internet, but I will most likely create a 'local' Kaggle dataset so that I can use these tools without the internet.\n\n## 3. ADDENDUM - @jirkaborovec's [Notebook Showing How to Use PIL! and more!](https://www.kaggle.com/code/jirkaborovec/bloodclots-classif-eda-load-crop-images)\n\nI didn't think PIL would be fast enough or memory savvy enough to handle this, but it appears as though I was wrong!\nPlease check out the above notebook for details regarding using PIL to resize the images (and normalize).\n\n---\n\nTo see these tools in action check out the simple notebook I made:\n* [**MCSAI - How To Interact With Large .tif Files**](https://www.kaggle.com/code/dschettler8845/mcsai-how-to-interact-with-large-tif-files)\n\n---\n\n<br>\n\nA side note: I plan to use these tools to create a low-res PNG dataset that can be used alongside the OpenSlide patching tool to first explore (EDA) and later model the data. I'm still not sure yet how to approach this task, but I feel much more confident now that I have these tools. I hope this helps you too!",
      "votes": 58
    },
    {
      "id": 1848781,
      "postDate": "2022-07-08T22:21:21.213Z",
      "content": "<p>The newer notebooks does not work well with older pyvips installation which is publically available for installation on kaggle, I did using conda -download-only to download them into my working directory, then downloading all .bz2 files, and then making a dataset out of them to install it without internet. (I got the latest version to work!) (My first sub is a trained model b7 on full images, score is 0.7 because it is very easy to overfit)</p>\n<p>To my understanding, pyvips works with multiprocessing, and loading 1 part at a time, compressing it and keeping it into memory (saves massive memory, and because of multiprocessing, still finishes faster than other libraries), the newer version lets use load the tiff file into a PIL Image file on a dimension we want [using pyvips.Image.thumbnail(path) ], it takes less than 3 minutes to load test set (4 images given) and save them into 1024x1024x3 npy, my guess is it would take more or less about 2 hours to load the ~200 hidden test set images. </p>\n<p>Even after manually removing the PIL's file processing limit (using Image.MAX_IMAGE_PIXELS = None), the file can not be loaded into kaggle's env because of low ram, it just crashes the kernel at some images.</p>",
      "rawMarkdown": "The newer notebooks does not work well with older pyvips installation which is publically available for installation on kaggle, I did using conda -download-only to download them into my working directory, then downloading all .bz2 files, and then making a dataset out of them to install it without internet. (I got the latest version to work!) (My first sub is a trained model b7 on full images, score is 0.7 because it is very easy to overfit)\n\nTo my understanding, pyvips works with multiprocessing, and loading 1 part at a time, compressing it and keeping it into memory (saves massive memory, and because of multiprocessing, still finishes faster than other libraries), the newer version lets use load the tiff file into a PIL Image file on a dimension we want [using pyvips.Image.thumbnail(path) ], it takes less than 3 minutes to load test set (4 images given) and save them into 1024x1024x3 npy, my guess is it would take more or less about 2 hours to load the ~200 hidden test set images. \n\nEven after manually removing the PIL's file processing limit (using Image.MAX_IMAGE_PIXELS = None), the file can not be loaded into kaggle's env because of low ram, it just crashes the kernel at some images.",
      "votes": 3,
      "replies": [
        {
          "id": 1860451,
          "postDate": "2022-07-18T10:23:39.797Z",
          "content": "<p>Can you make pyvips installation dataset public ?</p>",
          "rawMarkdown": "Can you make pyvips installation dataset public ?",
          "votes": 2
        },
        {
          "id": 1860972,
          "postDate": "2022-07-18T17:30:12.407Z",
          "content": "<p>It is already public in this competition by another good fellow, please check 'code' tab</p>",
          "rawMarkdown": "It is already public in this competition by another good fellow, please check 'code' tab"
        },
        {
          "id": 1860980,
          "postDate": "2022-07-18T17:35:28.193Z",
          "content": "<p><a href=\"https://www.kaggle.com/code/dschettler8845/use-pyvips-offline-v-2-1-13\" target=\"_blank\">My version</a> is strictly for the old version. Is there an example showing how to with the new version ? If so can you link it <a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> </p>",
          "rawMarkdown": "[My version](https://www.kaggle.com/code/dschettler8845/use-pyvips-offline-v-2-1-13) is strictly for the old version. Is there an example showing how to with the new version ? If so can you link it @harshitsheoran ",
          "votes": 2
        },
        {
          "id": 1861015,
          "postDate": "2022-07-18T18:03:48.653Z",
          "content": "<p><a href=\"https://www.kaggle.com/code/analokamus/how-to-use-pyvips-offline\" target=\"_blank\">https://www.kaggle.com/code/analokamus/how-to-use-pyvips-offline</a><br>\nposted yesterday so you might not have found it…</p>",
          "rawMarkdown": "https://www.kaggle.com/code/analokamus/how-to-use-pyvips-offline\nposted yesterday so you might not have found it...",
          "votes": 2
        },
        {
          "id": 1861824,
          "postDate": "2022-07-19T09:12:38.040Z",
          "content": "<p>If you still have problems with the installation, you can add the output of my notebook to your dataset : <a href=\"https://www.kaggle.com/code/ahmedelfazouan/conda-pyvips/notebook\" target=\"_blank\">https://www.kaggle.com/code/ahmedelfazouan/conda-pyvips/notebook</a><br>\nand then execute this command : <strong>!conda install ../input/conda-pyvips/*.tar.bz2</strong></p>",
          "rawMarkdown": "If you still have problems with the installation, you can add the output of my notebook to your dataset : https://www.kaggle.com/code/ahmedelfazouan/conda-pyvips/notebook\nand then execute this command : **!conda install ../input/conda-pyvips/*.tar.bz2**",
          "votes": 2
        },
        {
          "id": 1868744,
          "postDate": "2022-07-24T07:30:41.407Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1869893,
      "postDate": "2022-07-25T06:02:04.950Z",
      "content": "<p>UPDATE: just run and convert all test images in offline mode 🎉</p>\n<blockquote>\n  <p><a href=\"https://www.kaggle.com/code/jirkaborovec/bloodclots-classif-fake-prediction?scriptVersionId=101664294\" target=\"_blank\">https://www.kaggle.com/code/jirkaborovec/bloodclots-classif-fake-prediction?scriptVersionId=101664294</a></p>\n</blockquote>",
      "rawMarkdown": "UPDATE: just run and convert all test images in offline mode 🎉\n> https://www.kaggle.com/code/jirkaborovec/bloodclots-classif-fake-prediction?scriptVersionId=101664294",
      "votes": 1
    },
    {
      "id": 1860844,
      "postDate": "2022-07-18T16:00:50.393Z",
      "content": "<p>Thank's for the information.  I will surely use this concept in my future work.</p>",
      "rawMarkdown": "Thank's for the information.  I will surely use this concept in my future work.",
      "votes": 1
    },
    {
      "id": 1849826,
      "postDate": "2022-07-09T20:37:14.197Z",
      "content": "<p>I was able to load and convert almost all images except 3, the threshold id 4.100.000.000 pixels</p>",
      "rawMarkdown": "I was able to load and convert almost all images except 3, the threshold id 4.100.000.000 pixels",
      "votes": 1
    },
    {
      "id": 1903903,
      "postDate": "2022-08-17T18:32:22.560Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a>, thanks, I tried both OpenSlide and pyvips as well. That's why I have a question, have you experienced any problem with growing RAM while using <code>pyvips</code>? </p>\n<p>I delete useless images and use <code>gc</code> at every iteration, even set <code>pyvips.cache_set_max(0)</code> to not to store any cache but it still seems to grow significantly. What is more strange is that <code>psutil.Process(os.getpid()).memory_info().rss</code> seems not to match with the displayed usage of the RAM on the picture… While it is displayed that I used almost 7 GB, <code>psutil.Process(os.getpid()).memory_info().rss</code> shows that there were no changes and has 400MB all the time…</p>\n<p>Maybe somebody else have similar Problem?</p>",
      "rawMarkdown": "Hi @dschettler8845, thanks, I tried both OpenSlide and pyvips as well. That's why I have a question, have you experienced any problem with growing RAM while using `pyvips`? \n\nI delete useless images and use `gc` at every iteration, even set `pyvips.cache_set_max(0)` to not to store any cache but it still seems to grow significantly. What is more strange is that `psutil.Process(os.getpid()).memory_info().rss` seems not to match with the displayed usage of the RAM on the picture... While it is displayed that I used almost 7 GB, `psutil.Process(os.getpid()).memory_info().rss` shows that there were no changes and has 400MB all the time...\n\nMaybe somebody else have similar Problem?"
    },
    {
      "id": 1888105,
      "postDate": "2022-08-07T11:15:43.733Z",
      "content": "<p>Hi Darien,<br>\nGreat work by the way, Upvoted.<br>\nI have one question though. Does loading and resizing the image through this command pyvips.Image.new_from_file(img_path).resize(1/resize_factor).numpy() actually solve the memory problem.</p>\n<p>Earlier I used skimage.io.imread to read the image and cv2.resize to resize, but I got \"Notebook Exceed Maximum Allowed Compute\" after submission.</p>\n<p>After using pyvips I am seeing continious RAM usage rise in my notebook similar as before. I am afraid to submit because of the 1 submission limit per day until I am sure the tiff image loading problem is solved.</p>\n<p>I ran my notebook on the training set using skimage.io.imread and cv2.resize and there is no problem. But something changes during submission, maybe a huge image with size not seen before</p>\n<p>Any comments ?</p>",
      "rawMarkdown": "Hi Darien,\nGreat work by the way, Upvoted.\nI have one question though. Does loading and resizing the image through this command pyvips.Image.new_from_file(img_path).resize(1/resize_factor).numpy() actually solve the memory problem.\n\nEarlier I used skimage.io.imread to read the image and cv2.resize to resize, but I got \"Notebook Exceed Maximum Allowed Compute\" after submission.\n\nAfter using pyvips I am seeing continious RAM usage rise in my notebook similar as before. I am afraid to submit because of the 1 submission limit per day until I am sure the tiff image loading problem is solved.\n\nI ran my notebook on the training set using skimage.io.imread and cv2.resize and there is no problem. But something changes during submission, maybe a huge image with size not seen before\n\nAny comments ?"
    },
    {
      "id": 1865142,
      "postDate": "2022-07-21T14:56:12.707Z",
      "content": "<p>Have you experimented with <a href=\"https://docs.monai.io/en/stable/data.html#openslidewsireader\" target=\"_blank\">MONAI:openslidewsireader</a> and reading images per patch and composing back together already called patches?</p>",
      "rawMarkdown": "Have you experimented with [MONAI:openslidewsireader](https://docs.monai.io/en/stable/data.html#openslidewsireader) and reading images per patch and composing back together already called patches?"
    }
  ],
  "comments": [
    {
      "id": 1848781,
      "author_name": "Harshit Sheoran",
      "author_url": "",
      "post_date": "2022-07-08T22:21:21.213000",
      "content": "<p>The newer notebooks does not work well with older pyvips installation which is publically available for installation on kaggle, I did using conda -download-only to download them into my working directory, then downloading all .bz2 files, and then making a dataset out of them to install it without internet. (I got the latest version to work!) (My first sub is a trained model b7 on full images, score is 0.7 because it is very easy to overfit)</p>\n<p>To my understanding, pyvips works with multiprocessing, and loading 1 part at a time, compressing it and keeping it into memory (saves massive memory, and because of multiprocessing, still finishes faster than other libraries), the newer version lets use load the tiff file into a PIL Image file on a dimension we want [using pyvips.Image.thumbnail(path) ], it takes less than 3 minutes to load test set (4 images given) and save them into 1024x1024x3 npy, my guess is it would take more or less about 2 hours to load the ~200 hidden test set images. </p>\n<p>Even after manually removing the PIL's file processing limit (using Image.MAX_IMAGE_PIXELS = None), the file can not be loaded into kaggle's env because of low ram, it just crashes the kernel at some images.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1860451,
          "author_name": "Ahmed El Fazouani",
          "author_url": "",
          "post_date": "2022-07-18T10:23:39.797000",
          "content": "<p>Can you make pyvips installation dataset public ?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1860972,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2022-07-18T17:30:12.407000",
          "content": "<p>It is already public in this competition by another good fellow, please check 'code' tab</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1860980,
          "author_name": "Darien Schettler",
          "author_url": "",
          "post_date": "2022-07-18T17:35:28.193000",
          "content": "<p><a href=\"https://www.kaggle.com/code/dschettler8845/use-pyvips-offline-v-2-1-13\" target=\"_blank\">My version</a> is strictly for the old version. Is there an example showing how to with the new version ? If so can you link it <a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1861015,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2022-07-18T18:03:48.653000",
          "content": "<p><a href=\"https://www.kaggle.com/code/analokamus/how-to-use-pyvips-offline\" target=\"_blank\">https://www.kaggle.com/code/analokamus/how-to-use-pyvips-offline</a><br>\nposted yesterday so you might not have found it…</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1861824,
          "author_name": "Ahmed El Fazouani",
          "author_url": "",
          "post_date": "2022-07-19T09:12:38.040000",
          "content": "<p>If you still have problems with the installation, you can add the output of my notebook to your dataset : <a href=\"https://www.kaggle.com/code/ahmedelfazouan/conda-pyvips/notebook\" target=\"_blank\">https://www.kaggle.com/code/ahmedelfazouan/conda-pyvips/notebook</a><br>\nand then execute this command : <strong>!conda install ../input/conda-pyvips/*.tar.bz2</strong></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1868744,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-07-24T07:30:41.407000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1869893,
      "author_name": "Jirka",
      "author_url": "",
      "post_date": "2022-07-25T06:02:04.950000",
      "content": "<p>UPDATE: just run and convert all test images in offline mode 🎉</p>\n<blockquote>\n  <p><a href=\"https://www.kaggle.com/code/jirkaborovec/bloodclots-classif-fake-prediction?scriptVersionId=101664294\" target=\"_blank\">https://www.kaggle.com/code/jirkaborovec/bloodclots-classif-fake-prediction?scriptVersionId=101664294</a></p>\n</blockquote>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1860844,
      "author_name": "Abdul Basit",
      "author_url": "",
      "post_date": "2022-07-18T16:00:50.393000",
      "content": "<p>Thank's for the information.  I will surely use this concept in my future work.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1849826,
      "author_name": "Jirka",
      "author_url": "",
      "post_date": "2022-07-09T20:37:14.197000",
      "content": "<p>I was able to load and convert almost all images except 3, the threshold id 4.100.000.000 pixels</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1903903,
      "author_name": "Kostiantyn Lavronenko",
      "author_url": "",
      "post_date": "2022-08-17T18:32:22.560000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a>, thanks, I tried both OpenSlide and pyvips as well. That's why I have a question, have you experienced any problem with growing RAM while using <code>pyvips</code>? </p>\n<p>I delete useless images and use <code>gc</code> at every iteration, even set <code>pyvips.cache_set_max(0)</code> to not to store any cache but it still seems to grow significantly. What is more strange is that <code>psutil.Process(os.getpid()).memory_info().rss</code> seems not to match with the displayed usage of the RAM on the picture… While it is displayed that I used almost 7 GB, <code>psutil.Process(os.getpid()).memory_info().rss</code> shows that there were no changes and has 400MB all the time…</p>\n<p>Maybe somebody else have similar Problem?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1888105,
      "author_name": "Abhisek Dash",
      "author_url": "",
      "post_date": "2022-08-07T11:15:43.733000",
      "content": "<p>Hi Darien,<br>\nGreat work by the way, Upvoted.<br>\nI have one question though. Does loading and resizing the image through this command pyvips.Image.new_from_file(img_path).resize(1/resize_factor).numpy() actually solve the memory problem.</p>\n<p>Earlier I used skimage.io.imread to read the image and cv2.resize to resize, but I got \"Notebook Exceed Maximum Allowed Compute\" after submission.</p>\n<p>After using pyvips I am seeing continious RAM usage rise in my notebook similar as before. I am afraid to submit because of the 1 submission limit per day until I am sure the tiff image loading problem is solved.</p>\n<p>I ran my notebook on the training set using skimage.io.imread and cv2.resize and there is no problem. But something changes during submission, maybe a huge image with size not seen before</p>\n<p>Any comments ?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1865142,
      "author_name": "Jirka",
      "author_url": "",
      "post_date": "2022-07-21T14:56:12.707000",
      "content": "<p>Have you experimented with <a href=\"https://docs.monai.io/en/stable/data.html#openslidewsireader\" target=\"_blank\">MONAI:openslidewsireader</a> and reading images per patch and composing back together already called patches?</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1848611": "Hi there, I've seen a lot of different posts already about how to deal with the data in this competition as the images are quite large and can go OOM in RAM quite easily. Even if you had an incredibly powerful system, it would still be to your benefit to learn how to interact with these objects/images in a smart way.\n\nLet's discuss the two tools I'll be using (for now) in this competition to handle/manipulate these images (WSI).\n\n## 1. [**OpenSlide**](https://openslide.org/api/python/)\n\nThis tool is insanely useful when dealing with WSI data. I won't get into the details of WHY that do not pertain to this competition, but trust me when I say it is very useful for many pipelines related to WSI. So what makes it useful? Why use it?\n* This is the main one for this comp: **You can grab any patch from the image without loading the entire image into memory.**\n  * This means I could tile the image without ever using more than the tile-size image's worth of RAM.\n* You can \"open\" an image and inspect pieces of metadata (ie. just load the header) without loading the actual image file into memory. This will let you access any metadata included in the image and at the least will give you access to the image size.\n\n## 2. [**pyvips**](https://github.com/libvips/pyvips) and [**libvips**](https://github.com/libvips)\n\nThese are new to me, but they are what allowed me to be able to quickly (relatively speaking) resize the full-size images down to a more manageable size. (3-90 seconds depending on original image resolution and desired downscale). They use virtual memory and some other stuff I don't understand at this point in time. I encourage you to read up about them.\n\nCurrently, I'm installing them over the internet, but I will most likely create a 'local' Kaggle dataset so that I can use these tools without the internet.\n\n## 3. ADDENDUM - @jirkaborovec's [Notebook Showing How to Use PIL! and more!](https://www.kaggle.com/code/jirkaborovec/bloodclots-classif-eda-load-crop-images)\n\nI didn't think PIL would be fast enough or memory savvy enough to handle this, but it appears as though I was wrong!\nPlease check out the above notebook for details regarding using PIL to resize the images (and normalize).\n\n---\n\nTo see these tools in action check out the simple notebook I made:\n* [**MCSAI - How To Interact With Large .tif Files**](https://www.kaggle.com/code/dschettler8845/mcsai-how-to-interact-with-large-tif-files)\n\n---\n\n<br>\n\nA side note: I plan to use these tools to create a low-res PNG dataset that can be used alongside the OpenSlide patching tool to first explore (EDA) and later model the data. I'm still not sure yet how to approach this task, but I feel much more confident now that I have these tools. I hope this helps you too!",
    "1848781": "The newer notebooks does not work well with older pyvips installation which is publically available for installation on kaggle, I did using conda -download-only to download them into my working directory, then downloading all .bz2 files, and then making a dataset out of them to install it without internet. (I got the latest version to work!) (My first sub is a trained model b7 on full images, score is 0.7 because it is very easy to overfit)\n\nTo my understanding, pyvips works with multiprocessing, and loading 1 part at a time, compressing it and keeping it into memory (saves massive memory, and because of multiprocessing, still finishes faster than other libraries), the newer version lets use load the tiff file into a PIL Image file on a dimension we want [using pyvips.Image.thumbnail(path) ], it takes less than 3 minutes to load test set (4 images given) and save them into 1024x1024x3 npy, my guess is it would take more or less about 2 hours to load the ~200 hidden test set images. \n\nEven after manually removing the PIL's file processing limit (using Image.MAX_IMAGE_PIXELS = None), the file can not be loaded into kaggle's env because of low ram, it just crashes the kernel at some images.",
    "1869893": "UPDATE: just run and convert all test images in offline mode 🎉\n> https://www.kaggle.com/code/jirkaborovec/bloodclots-classif-fake-prediction?scriptVersionId=101664294",
    "1860844": "Thank's for the information.  I will surely use this concept in my future work.",
    "1849826": "I was able to load and convert almost all images except 3, the threshold id 4.100.000.000 pixels",
    "1903903": "Hi @dschettler8845, thanks, I tried both OpenSlide and pyvips as well. That's why I have a question, have you experienced any problem with growing RAM while using `pyvips`? \n\nI delete useless images and use `gc` at every iteration, even set `pyvips.cache_set_max(0)` to not to store any cache but it still seems to grow significantly. What is more strange is that `psutil.Process(os.getpid()).memory_info().rss` seems not to match with the displayed usage of the RAM on the picture... While it is displayed that I used almost 7 GB, `psutil.Process(os.getpid()).memory_info().rss` shows that there were no changes and has 400MB all the time...\n\nMaybe somebody else have similar Problem?",
    "1888105": "Hi Darien,\nGreat work by the way, Upvoted.\nI have one question though. Does loading and resizing the image through this command pyvips.Image.new_from_file(img_path).resize(1/resize_factor).numpy() actually solve the memory problem.\n\nEarlier I used skimage.io.imread to read the image and cv2.resize to resize, but I got \"Notebook Exceed Maximum Allowed Compute\" after submission.\n\nAfter using pyvips I am seeing continious RAM usage rise in my notebook similar as before. I am afraid to submit because of the 1 submission limit per day until I am sure the tiff image loading problem is solved.\n\nI ran my notebook on the training set using skimage.io.imread and cv2.resize and there is no problem. But something changes during submission, maybe a huge image with size not seen before\n\nAny comments ?",
    "1865142": "Have you experimented with [MONAI:openslidewsireader](https://docs.monai.io/en/stable/data.html#openslidewsireader) and reading images per patch and composing back together already called patches?"
  }
}