{
  "id": 451907,
  "title": "submission `Notebook Out of Disk` with pyVips",
  "url": "/competitions/UBC-OCEAN/discussion/451907",
  "author_name": "Jirka",
  "post_date": "2023-10-31T02:23:28.796000",
  "votes": 7,
  "comment_count": 18,
  "views": 0,
  "content": "<p>I was just facing this strange error that I had not seen while training. After some debugging, it seems that you really need to set <code>os.environ['VIPS_DISC_THRESHOLD'] = '15gb'</code> or any other high number because the default value makes it swap early nd the file is not space efficient…</p>\n<p>Ref: <a href=\"https://www.kaggle.com/code/jirkaborovec/cancer-subtype-lightning-torch-inference-tiles\" target=\"_blank\">https://www.kaggle.com/code/jirkaborovec/cancer-subtype-lightning-torch-inference-tiles</a></p>",
  "messages": [
    {
      "id": 2505959,
      "postDate": "2023-10-31T02:23:28.797Z",
      "content": "<p>I was just facing this strange error that I had not seen while training. After some debugging, it seems that you really need to set <code>os.environ['VIPS_DISC_THRESHOLD'] = '15gb'</code> or any other high number because the default value makes it swap early nd the file is not space efficient…</p>\n<p>Ref: <a href=\"https://www.kaggle.com/code/jirkaborovec/cancer-subtype-lightning-torch-inference-tiles\" target=\"_blank\">https://www.kaggle.com/code/jirkaborovec/cancer-subtype-lightning-torch-inference-tiles</a></p>",
      "rawMarkdown": "I was just facing this strange error that I had not seen while training. After some debugging, it seems that you really need to set `os.environ['VIPS_DISC_THRESHOLD'] = '15gb'` or any other high number because the default value makes it swap early nd the file is not space efficient...\n\nRef: https://www.kaggle.com/code/jirkaborovec/cancer-subtype-lightning-torch-inference-tiles",
      "votes": 7
    },
    {
      "id": 2506825,
      "postDate": "2023-10-31T15:03:24.537Z",
      "content": "<p>In testing dataset contains 2000  large images. we have to take tile of each image. How many tiles of single image you have taken. Please comment it? </p>",
      "rawMarkdown": "In testing dataset contains 2000  large images. we have to take tile of each image. How many tiles of single image you have taken. Please comment it? ",
      "votes": 1,
      "replies": [
        {
          "id": 2506829,
          "postDate": "2023-10-31T15:07:49.843Z",
          "content": "<p>running tiles 2048px scaled down to 512px and limiting max 50 tiles per image and all passe within the 12h limit…<br>\n<a href=\"https://www.kaggle.com/code/jirkaborovec/cancer-subtype-lightning-torch-inference-tiles\" target=\"_blank\">https://www.kaggle.com/code/jirkaborovec/cancer-subtype-lightning-torch-inference-tiles</a></p>",
          "rawMarkdown": "running tiles 2048px scaled down to 512px and limiting max 50 tiles per image and all passe within the 12h limit...\nhttps://www.kaggle.com/code/jirkaborovec/cancer-subtype-lightning-torch-inference-tiles",
          "votes": 2,
          "replies": [
            {
              "id": 2522096,
              "postDate": "2023-11-12T11:40:10.597Z",
              "content": "<p>Have you tried not limiting the maximum number of patches? I'm using a weakly supervised multi-instance method, which requires converting all patches into feature vectors and saving them as .pt files. If I only use 50 patches, the results would be quite poor</p>",
              "rawMarkdown": "Have you tried not limiting the maximum number of patches? I'm using a weakly supervised multi-instance method, which requires converting all patches into feature vectors and saving them as .pt files. If I only use 50 patches, the results would be quite poor",
              "votes": 2
            },
            {
              "id": 2522149,
              "postDate": "2023-11-12T12:24:05.003Z",
              "content": "<p>Yes I tried but ran out of time, and it is impossible to say bow much as there is no progress status from the time of termination… MIL is something I wanted to try too, do you have a reference?</p>",
              "rawMarkdown": "Yes I tried but ran out of time, and it is impossible to say bow much as there is no progress status from the time of termination... MIL is something I wanted to try too, do you have a reference?"
            },
            {
              "id": 2522161,
              "postDate": "2023-11-12T12:46:00.933Z",
              "content": "<p>I've reviewed your code, and for the MIL approach, if the patches are saved directly in memory and then embedded into .pt files for storage, it seems that subsampling does not actually speed up the process. I believe that in your method, the main acceleration from subsampling comes from reducing the I/O operations involved in saving to the temp directory.</p>",
              "rawMarkdown": "I've reviewed your code, and for the MIL approach, if the patches are saved directly in memory and then embedded into .pt files for storage, it seems that subsampling does not actually speed up the process. I believe that in your method, the main acceleration from subsampling comes from reducing the I/O operations involved in saving to the temp directory.",
              "votes": 1
            },
            {
              "id": 2522192,
              "postDate": "2023-11-12T13:22:29.890Z",
              "content": "<p>yes the speedup is only because of I/O reduction, and also this way it is easier to run two in parallel</p>",
              "rawMarkdown": "yes the speedup is only because of I/O reduction, and also this way it is easier to run two in parallel"
            }
          ]
        },
        {
          "id": 2509636,
          "postDate": "2023-11-02T13:36:26.767Z",
          "content": "<p>Why do you know that the test set contains 2000 large images？</p>",
          "rawMarkdown": "Why do you know that the test set contains 2000 large images？",
          "replies": [
            {
              "id": 2509726,
              "postDate": "2023-11-02T14:28:49.247Z",
              "content": "<p>I think it is written in the competition details.</p>\n<blockquote>\n  <p>The test set contains images from different source hospitals than the train set, with the largest area images almost 100,000 x 50,000 pixels. We strongly recommend taking an expansive approach to thinking about the scenarios your error handling should manage, including differences in image dimensions, quality, slide staining techniques, and more. Expect <strong>roughly 2,000 images in the test set</strong>, the majority of which are TMAs. The total size is 550 GB so simply loading the data will be time consuming. Be warned that the test set was specifically constructed to assess how well models generalize.</p>\n</blockquote>",
              "rawMarkdown": "I think it is written in the competition details.\n\n> The test set contains images from different source hospitals than the train set, with the largest area images almost 100,000 x 50,000 pixels. We strongly recommend taking an expansive approach to thinking about the scenarios your error handling should manage, including differences in image dimensions, quality, slide staining techniques, and more. Expect **roughly 2,000 images in the test set**, the majority of which are TMAs. The total size is 550 GB so simply loading the data will be time consuming. Be warned that the test set was specifically constructed to assess how well models generalize."
            },
            {
              "id": 2509779,
              "postDate": "2023-11-02T14:49:50.727Z",
              "content": "<p>It's my fault, thank you!</p>",
              "rawMarkdown": "It's my fault, thank you!"
            }
          ]
        }
      ]
    },
    {
      "id": 2595584,
      "postDate": "2024-01-10T14:10:35.647Z",
      "content": "<p>Sorry, как вы решили эту проблему ? When I write <br>\n\"os.environ['VIPS_CONCURRENCY'] = '4'<br>\nos.environ['VIPS_DISC_THRESHOLD'] = '15gb' \", I get Out of memory.</p>\n<p>Can you help me please ?</p>",
      "rawMarkdown": "Sorry, как вы решили эту проблему ? When I write \n\"os.environ['VIPS_CONCURRENCY'] = '4'\nos.environ['VIPS_DISC_THRESHOLD'] = '15gb' \", I get Out of memory.\n\nCan you help me please ?"
    },
    {
      "id": 2506444,
      "postDate": "2023-10-31T10:12:21.237Z",
      "content": "<p>Where do you create files? If in \"Output\" (I assume that answer is yes) you should change writing directory to …. eg. /tmp  <br>\nSimply set the output directory to /tmp/images and save the tiles in it.</p>",
      "rawMarkdown": "Where do you create files? If in \"Output\" (I assume that answer is yes) you should change writing directory to .... eg. /tmp  \nSimply set the output directory to /tmp/images and save the tiles in it.",
      "replies": [
        {
          "id": 2506512,
          "postDate": "2023-10-31T10:53:31.063Z",
          "content": "<p>I think it is by default in the system temp, so as you say…</p>",
          "rawMarkdown": "I think it is by default in the system temp, so as you say...",
          "replies": [
            {
              "id": 2507094,
              "postDate": "2023-10-31T19:12:35.057Z",
              "content": "<p>Do you assume or do you know?</p>\n<p><img src=\"https://i.ibb.co/34SKCzg/001.jpg\" alt=\"\"></p>",
              "rawMarkdown": "Do you assume or do you know?\n\n![](https://i.ibb.co/34SKCzg/001.jpg)"
            },
            {
              "id": 2507117,
              "postDate": "2023-10-31T19:23:14.687Z",
              "content": "<p><code>/tmp</code> is limited to 40GB also :) but yes likely it is the user's home</p>",
              "rawMarkdown": "`/tmp` is limited to 40GB also :) but yes likely it is the user's home"
            }
          ]
        }
      ]
    },
    {
      "id": 2506267,
      "postDate": "2023-10-31T07:33:18.367Z",
      "content": "<p>I also meet this problem.When I cropped the train_images on my server,I met the similar problem.I see the livvips used almost all of my  disk.So I had to use your dataset for training.Also my submission took a lot of time,I also what the solution as well.<br>\nI tried to use pyvips.cache_set_max(),but it didn't work well.But I think the way is right.</p>",
      "rawMarkdown": "I also meet this problem.When I cropped the train_images on my server,I met the similar problem.I see the livvips used almost all of my  disk.So I had to use your dataset for training.Also my submission took a lot of time,I also what the solution as well.\nI tried to use pyvips.cache_set_max(),but it didn't work well.But I think the way is right."
    },
    {
      "id": 2506179,
      "postDate": "2023-10-31T06:33:03.377Z",
      "content": "<p>I guess you have do inference in memory. Why do you need to save images to disk?</p>",
      "rawMarkdown": "I guess you have do inference in memory. Why do you need to save images to disk?",
      "replies": [
        {
          "id": 2506266,
          "postDate": "2023-10-31T07:32:23.360Z",
          "content": "<p>In general, when splitting WSI to tiles, you may have more tiles than the memory, especially as the in-memory image could be in a data type that is hungry, so you may want to decompose the image to tiles and then load it from folders…<br>\nAlso true that this additional IO traffic may slow down the whole inference so when we limit nb of tiles all can be in-memory…</p>\n<p>In addition this file am talking about is build-in by Vips when the loading image exceed the RAM</p>",
          "rawMarkdown": "In general, when splitting WSI to tiles, you may have more tiles than the memory, especially as the in-memory image could be in a data type that is hungry, so you may want to decompose the image to tiles and then load it from folders...\nAlso true that this additional IO traffic may slow down the whole inference so when we limit nb of tiles all can be in-memory...\n\nIn addition this file am talking about is build-in by Vips when the loading image exceed the RAM"
        }
      ]
    },
    {
      "id": 2506249,
      "postDate": "2023-10-31T07:21:30.110Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2506825,
      "author_name": "sunil thite",
      "author_url": "",
      "post_date": "2023-10-31T15:03:24.537000",
      "content": "<p>In testing dataset contains 2000  large images. we have to take tile of each image. How many tiles of single image you have taken. Please comment it? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2506829,
          "author_name": "Jirka",
          "author_url": "",
          "post_date": "2023-10-31T15:07:49.843000",
          "content": "<p>running tiles 2048px scaled down to 512px and limiting max 50 tiles per image and all passe within the 12h limit…<br>\n<a href=\"https://www.kaggle.com/code/jirkaborovec/cancer-subtype-lightning-torch-inference-tiles\" target=\"_blank\">https://www.kaggle.com/code/jirkaborovec/cancer-subtype-lightning-torch-inference-tiles</a></p>",
          "votes": 2,
          "replies": [
            {
              "id": 2522096,
              "author_name": "Huang Jin Feng",
              "author_url": "",
              "post_date": "2023-11-12T11:40:10.597000",
              "content": "<p>Have you tried not limiting the maximum number of patches? I'm using a weakly supervised multi-instance method, which requires converting all patches into feature vectors and saving them as .pt files. If I only use 50 patches, the results would be quite poor</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2522149,
              "author_name": "Jirka",
              "author_url": "",
              "post_date": "2023-11-12T12:24:05.003000",
              "content": "<p>Yes I tried but ran out of time, and it is impossible to say bow much as there is no progress status from the time of termination… MIL is something I wanted to try too, do you have a reference?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2522161,
              "author_name": "Huang Jin Feng",
              "author_url": "",
              "post_date": "2023-11-12T12:46:00.933000",
              "content": "<p>I've reviewed your code, and for the MIL approach, if the patches are saved directly in memory and then embedded into .pt files for storage, it seems that subsampling does not actually speed up the process. I believe that in your method, the main acceleration from subsampling comes from reducing the I/O operations involved in saving to the temp directory.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2522192,
              "author_name": "Jirka",
              "author_url": "",
              "post_date": "2023-11-12T13:22:29.890000",
              "content": "<p>yes the speedup is only because of I/O reduction, and also this way it is easier to run two in parallel</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2509636,
          "author_name": "WangXuC",
          "author_url": "",
          "post_date": "2023-11-02T13:36:26.767000",
          "content": "<p>Why do you know that the test set contains 2000 large images？</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2509726,
              "author_name": "Jirka",
              "author_url": "",
              "post_date": "2023-11-02T14:28:49.247000",
              "content": "<p>I think it is written in the competition details.</p>\n<blockquote>\n  <p>The test set contains images from different source hospitals than the train set, with the largest area images almost 100,000 x 50,000 pixels. We strongly recommend taking an expansive approach to thinking about the scenarios your error handling should manage, including differences in image dimensions, quality, slide staining techniques, and more. Expect <strong>roughly 2,000 images in the test set</strong>, the majority of which are TMAs. The total size is 550 GB so simply loading the data will be time consuming. Be warned that the test set was specifically constructed to assess how well models generalize.</p>\n</blockquote>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2509779,
              "author_name": "WangXuC",
              "author_url": "",
              "post_date": "2023-11-02T14:49:50.727000",
              "content": "<p>It's my fault, thank you!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2595584,
      "author_name": "Rinat",
      "author_url": "",
      "post_date": "2024-01-10T14:10:35.647000",
      "content": "<p>Sorry, как вы решили эту проблему ? When I write <br>\n\"os.environ['VIPS_CONCURRENCY'] = '4'<br>\nos.environ['VIPS_DISC_THRESHOLD'] = '15gb' \", I get Out of memory.</p>\n<p>Can you help me please ?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2506444,
      "author_name": "Remek Kinas",
      "author_url": "",
      "post_date": "2023-10-31T10:12:21.237000",
      "content": "<p>Where do you create files? If in \"Output\" (I assume that answer is yes) you should change writing directory to …. eg. /tmp  <br>\nSimply set the output directory to /tmp/images and save the tiles in it.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2506512,
          "author_name": "Jirka",
          "author_url": "",
          "post_date": "2023-10-31T10:53:31.063000",
          "content": "<p>I think it is by default in the system temp, so as you say…</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2507094,
              "author_name": "Remek Kinas",
              "author_url": "",
              "post_date": "2023-10-31T19:12:35.057000",
              "content": "<p>Do you assume or do you know?</p>\n<p><img src=\"https://i.ibb.co/34SKCzg/001.jpg\" alt=\"\"></p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2507117,
              "author_name": "Jirka",
              "author_url": "",
              "post_date": "2023-10-31T19:23:14.687000",
              "content": "<p><code>/tmp</code> is limited to 40GB also :) but yes likely it is the user's home</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2506267,
      "author_name": "LuoZiqian",
      "author_url": "",
      "post_date": "2023-10-31T07:33:18.367000",
      "content": "<p>I also meet this problem.When I cropped the train_images on my server,I met the similar problem.I see the livvips used almost all of my  disk.So I had to use your dataset for training.Also my submission took a lot of time,I also what the solution as well.<br>\nI tried to use pyvips.cache_set_max(),but it didn't work well.But I think the way is right.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2506179,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2023-10-31T06:33:03.377000",
      "content": "<p>I guess you have do inference in memory. Why do you need to save images to disk?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2506266,
          "author_name": "Jirka",
          "author_url": "",
          "post_date": "2023-10-31T07:32:23.360000",
          "content": "<p>In general, when splitting WSI to tiles, you may have more tiles than the memory, especially as the in-memory image could be in a data type that is hungry, so you may want to decompose the image to tiles and then load it from folders…<br>\nAlso true that this additional IO traffic may slow down the whole inference so when we limit nb of tiles all can be in-memory…</p>\n<p>In addition this file am talking about is build-in by Vips when the loading image exceed the RAM</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2506249,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-10-31T07:21:30.110000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2505959": "I was just facing this strange error that I had not seen while training. After some debugging, it seems that you really need to set `os.environ['VIPS_DISC_THRESHOLD'] = '15gb'` or any other high number because the default value makes it swap early nd the file is not space efficient...\n\nRef: https://www.kaggle.com/code/jirkaborovec/cancer-subtype-lightning-torch-inference-tiles",
    "2506825": "In testing dataset contains 2000  large images. we have to take tile of each image. How many tiles of single image you have taken. Please comment it? ",
    "2595584": "Sorry, как вы решили эту проблему ? When I write \n\"os.environ['VIPS_CONCURRENCY'] = '4'\nos.environ['VIPS_DISC_THRESHOLD'] = '15gb' \", I get Out of memory.\n\nCan you help me please ?",
    "2506444": "Where do you create files? If in \"Output\" (I assume that answer is yes) you should change writing directory to .... eg. /tmp  \nSimply set the output directory to /tmp/images and save the tiles in it.",
    "2506267": "I also meet this problem.When I cropped the train_images on my server,I met the similar problem.I see the livvips used almost all of my  disk.So I had to use your dataset for training.Also my submission took a lot of time,I also what the solution as well.\nI tried to use pyvips.cache_set_max(),but it didn't work well.But I think the way is right.",
    "2506179": "I guess you have do inference in memory. Why do you need to save images to disk?",
    "2506249": ""
  }
}