{
  "id": 336484,
  "title": "\"Notebook Exceeded Allowed Compute\" while submission",
  "url": "/competitions/mayo-clinic-strip-ai/discussion/336484",
  "author_name": "Jirka",
  "post_date": "2022-07-11T11:51:32.902000",
  "votes": 9,
  "comment_count": 26,
  "views": 0,
  "content": "<p>I am getting this strange error even though my run with provided test images works fine and I have set a limit of image size so it won't kill the small RAM… is there something dramatically different with the used machine for submission? Also, the submission details state we have 9 hours runtime, and this kernel dies after 2 hours or so…</p>\n<blockquote>\n  <p>Submission kernel: <a href=\"https://www.kaggle.com/code/jirkaborovec/bloodclots-classif-baseline-flash-effnet\" target=\"_blank\">https://www.kaggle.com/code/jirkaborovec/bloodclots-classif-baseline-flash-effnet</a><br>\n  Conversion kernel: <a href=\"https://www.kaggle.com/code/jirkaborovec/bloodclots-classif-eda-load-crop-images\" target=\"_blank\">https://www.kaggle.com/code/jirkaborovec/bloodclots-classif-eda-load-crop-images</a></p>\n</blockquote>\n<p>to show that I can convert all images up to 4e9 pixels without problem…</p>\n<p><strong>UPDATE:</strong></p>\n<p>a good finding from: <a href=\"https://www.kaggle.com/competitions/mayo-clinic-strip-ai/discussion/336819\" target=\"_blank\">https://www.kaggle.com/competitions/mayo-clinic-strip-ai/discussion/336819</a></p>\n<blockquote>\n  <p>If you don't use gpu, you have 16 Gb of RAM, but if you use gpu, it's reduced to 13 Gb on the kaggle notebook.</p>\n</blockquote>\n<p>Also, I was able to load and covert images up to 1.5e9 pixels with GPU</p>",
  "messages": [
    {
      "id": 1851612,
      "postDate": "2022-07-11T11:51:32.903Z",
      "content": "<p>I am getting this strange error even though my run with provided test images works fine and I have set a limit of image size so it won't kill the small RAM… is there something dramatically different with the used machine for submission? Also, the submission details state we have 9 hours runtime, and this kernel dies after 2 hours or so…</p>\n<blockquote>\n  <p>Submission kernel: <a href=\"https://www.kaggle.com/code/jirkaborovec/bloodclots-classif-baseline-flash-effnet\" target=\"_blank\">https://www.kaggle.com/code/jirkaborovec/bloodclots-classif-baseline-flash-effnet</a><br>\n  Conversion kernel: <a href=\"https://www.kaggle.com/code/jirkaborovec/bloodclots-classif-eda-load-crop-images\" target=\"_blank\">https://www.kaggle.com/code/jirkaborovec/bloodclots-classif-eda-load-crop-images</a></p>\n</blockquote>\n<p>to show that I can convert all images up to 4e9 pixels without problem…</p>\n<p><strong>UPDATE:</strong></p>\n<p>a good finding from: <a href=\"https://www.kaggle.com/competitions/mayo-clinic-strip-ai/discussion/336819\" target=\"_blank\">https://www.kaggle.com/competitions/mayo-clinic-strip-ai/discussion/336819</a></p>\n<blockquote>\n  <p>If you don't use gpu, you have 16 Gb of RAM, but if you use gpu, it's reduced to 13 Gb on the kaggle notebook.</p>\n</blockquote>\n<p>Also, I was able to load and covert images up to 1.5e9 pixels with GPU</p>",
      "rawMarkdown": "I am getting this strange error even though my run with provided test images works fine and I have set a limit of image size so it won't kill the small RAM... is there something dramatically different with the used machine for submission? Also, the submission details state we have 9 hours runtime, and this kernel dies after 2 hours or so...\n\n> Submission kernel: https://www.kaggle.com/code/jirkaborovec/bloodclots-classif-baseline-flash-effnet\n> Conversion kernel: https://www.kaggle.com/code/jirkaborovec/bloodclots-classif-eda-load-crop-images\n\nto show that I can convert all images up to 4e9 pixels without problem...\n\n**UPDATE:**\n\na good finding from: https://www.kaggle.com/competitions/mayo-clinic-strip-ai/discussion/336819\n\n> If you don't use gpu, you have 16 Gb of RAM, but if you use gpu, it's reduced to 13 Gb on the kaggle notebook.\n\nAlso, I was able to load and covert images up to 1.5e9 pixels with GPU",
      "votes": 9
    },
    {
      "id": 1907793,
      "postDate": "2022-08-21T04:52:48.877Z",
      "content": "<p>Hello! This is probably a re-working of Topic Author's solution, but I would like to make a few comments.</p>\n<p>I too encountered the \"Notebook Exceeded Allowed Compute\" in my submission pipeline.<br>\nIt seems that the OOM was occurring when loading images using the <code>tifffile</code>.</p>\n<p>Since the RAM limit for a GPU instance is 13 GB, I used the train data to investigate the maximum file size (width x height) that could not be loaded due to this limitation.<br>\n<a href=\"https://www.kaggle.com/code/bobfromjapan/mayo-filesize-inspection\" target=\"_blank\">https://www.kaggle.com/code/bobfromjapan/mayo-filesize-inspection</a></p>\n<p>The results showed that the largest image in train data, <code>6baf51_0.tif</code>, has a size of <code>4896084492</code>, which would exceed the 13GB System RAM of the GPU instance.<br>\nThe second largest image, <code>b894f4_0.tif</code>, is <code>4131662535</code> and can be loaded.</p>\n<p>I tried to skip processing when the size exceeds <code>4131662535</code> as a workaround to avoid the \"Notebook Exceeded Allowed Compute\". Below is an example.</p>\n<pre><code>too_big_for_process = []\nfor path in test_image_paths:\n\n    print(path)\n    slide = OpenSlide(path)\n\n    if slide.dimensions[0]*slide.dimensions[1] &lt; 4131662535:\n        image_id = os.path.splitext(os.path.basename(path))[0]\n        image = tiff.imread(path)\n        print(f\"{image.shape}\")\n        cv2.imwrite(os.path.join(output_dir, f\"{image_id}.jpg\"), image[::scale,::scale,::-1])\n        del image\n        gc.collect()\n    else:\n        print(\"Skip process for avoiding OOM possibility.\")\n        too_big_for_process.append(path)\n</code></pre>\n<p>After the inference, if there were images that could not be processed, the value of prediction was given as a constant. This avoids the \"Notebook Exceeded Allowed Compute\" and the submission succeeds!</p>\n<pre><code>if len(too_big_for_process)&gt;0:\n    for i in too_big_for_process:\n        preds.append((i, 0.82, 0.28))\n</code></pre>",
      "rawMarkdown": "Hello! This is probably a re-working of Topic Author's solution, but I would like to make a few comments.\n\nI too encountered the \"Notebook Exceeded Allowed Compute\" in my submission pipeline.\nIt seems that the OOM was occurring when loading images using the `tifffile`.\n\nSince the RAM limit for a GPU instance is 13 GB, I used the train data to investigate the maximum file size (width x height) that could not be loaded due to this limitation.\nhttps://www.kaggle.com/code/bobfromjapan/mayo-filesize-inspection\n\nThe results showed that the largest image in train data, `6baf51_0.tif`, has a size of `4896084492`, which would exceed the 13GB System RAM of the GPU instance.\nThe second largest image, `b894f4_0.tif`, is `4131662535` and can be loaded.\n\nI tried to skip processing when the size exceeds `4131662535` as a workaround to avoid the \"Notebook Exceeded Allowed Compute\". Below is an example.\n\n```\ntoo_big_for_process = []\nfor path in test_image_paths:\n\n    print(path)\n    slide = OpenSlide(path)\n\n    if slide.dimensions[0]*slide.dimensions[1] < 4131662535:\n        image_id = os.path.splitext(os.path.basename(path))[0]\n        image = tiff.imread(path)\n        print(f\"{image.shape}\")\n        cv2.imwrite(os.path.join(output_dir, f\"{image_id}.jpg\"), image[::scale,::scale,::-1])\n        del image\n        gc.collect()\n    else:\n        print(\"Skip process for avoiding OOM possibility.\")\n        too_big_for_process.append(path)\n\n```\nAfter the inference, if there were images that could not be processed, the value of prediction was given as a constant. This avoids the \"Notebook Exceeded Allowed Compute\" and the submission succeeds!\n\n```\nif len(too_big_for_process)>0:\n    for i in too_big_for_process:\n        preds.append((i, 0.82, 0.28))\n```"
    },
    {
      "id": 1853611,
      "postDate": "2022-07-13T02:12:34.897Z",
      "content": "<p>I ran into a time error, made it 20% faster to barely finish it in time with 1 fold (HUGE OVERFIT BTW, leaderboard is nothing as training on my set), although I am unsure, what will happen if my run 5 folds instead </p>",
      "rawMarkdown": "I ran into a time error, made it 20% faster to barely finish it in time with 1 fold (HUGE OVERFIT BTW, leaderboard is nothing as training on my set), although I am unsure, what will happen if my run 5 folds instead ",
      "replies": [
        {
          "id": 1853807,
          "postDate": "2022-07-13T06:51:09.570Z",
          "content": "<p>So you say it is timeout issue? </p>",
          "rawMarkdown": "So you say it is timeout issue? "
        },
        {
          "id": 1853996,
          "postDate": "2022-07-13T11:20:06.410Z",
          "content": "<p>I had to make tiles, which if I load images carefully avoiding memory errors, takes too much time…</p>",
          "rawMarkdown": "I had to make tiles, which if I load images carefully avoiding memory errors, takes too much time..."
        },
        {
          "id": 1854026,
          "postDate": "2022-07-13T11:45:35.370Z",
          "content": "<p>what does too much time mean? mine is killed after bout 2h of runtime…<br>\neven they state:</p>\n<blockquote>\n  <p>CPU Notebook &lt;= 9 hours run-time<br>\n  GPU Notebook &lt;= 9 hours run-time<br>\n  Internet access disabled</p>\n</blockquote>",
          "rawMarkdown": "what does too much time mean? mine is killed after bout 2h of runtime...\neven they state:\n\n> CPU Notebook <= 9 hours run-time\nGPU Notebook <= 9 hours run-time\nInternet access disabled"
        },
        {
          "id": 1854098,
          "postDate": "2022-07-13T12:35:34.157Z",
          "content": "<p>It means, mine took more than 9 hours to executue, so if there occurs no error, instead of showing me the score it will show 'Notebook Timeout Error' </p>",
          "rawMarkdown": "It means, mine took more than 9 hours to executue, so if there occurs no error, instead of showing me the score it will show 'Notebook Timeout Error' "
        },
        {
          "id": 1854295,
          "postDate": "2022-07-13T15:28:50.637Z",
          "content": "<p>I see, the \"Notebook Timeout Error\" is quite clear… :)</p>",
          "rawMarkdown": "I see, the \"Notebook Timeout Error\" is quite clear... :)"
        }
      ]
    },
    {
      "id": 1853553,
      "postDate": "2022-07-13T00:25:39.647Z",
      "content": "<p>I think that the test is huge.. I keep getting errors every day when I try to run a submission..<br>\nMaybe try converting them in batches and you should see an improvement.</p>",
      "rawMarkdown": "I think that the test is huge.. I keep getting errors every day when I try to run a submission..\nMaybe try converting them in batches and you should see an improvement.\n\n",
      "replies": [
        {
          "id": 1853805,
          "postDate": "2022-07-13T06:49:30.587Z",
          "content": "<p>Hi, may you elaborate more on what you mean converting in batches? </p>",
          "rawMarkdown": "Hi, may you elaborate more on what you mean converting in batches? "
        }
      ]
    },
    {
      "id": 1852084,
      "postDate": "2022-07-11T18:41:59.003Z",
      "content": "<p>The real test set is MUCH larger than the provided test set. You need to debug with a representative dataset.</p>\n<p>I recommend you do the following to debug:</p>\n<ul>\n<li>instead of using the provided test data use a portion of the training data equivalent to the size  of the hidden test set</li>\n<li>Use something like debug==len(ss_df)==4 and then use that to initialize a more representative dataset.</li>\n<li>if I had to guess I’d say PIL is holding some RAM and causing an eventual OOM. </li>\n</ul>",
      "rawMarkdown": "The real test set is MUCH larger than the provided test set. You need to debug with a representative dataset.\n\nI recommend you do the following to debug:\n\n- instead of using the provided test data use a portion of the training data equivalent to the size  of the hidden test set\n- Use something like debug==len(ss_df)==4 and then use that to initialize a more representative dataset.\n- if I had to guess I’d say PIL is holding some RAM and causing an eventual OOM. \n",
      "replies": [
        {
          "id": 1852138,
          "postDate": "2022-07-11T20:00:16.430Z",
          "content": "<p>Yes, that is what I tried and I do not have any problem with the train dataset as I skipped all too large image even before loading… </p>",
          "rawMarkdown": "Yes, that is what I tried and I do not have any problem with the train dataset as I skipped all too large image even before loading... "
        },
        {
          "id": 1852157,
          "postDate": "2022-07-11T20:46:57.130Z",
          "content": "<p>So if you run your pipeline on a random sample of 280 training images (Save and Run All), you don't encounter any errors?</p>",
          "rawMarkdown": "So if you run your pipeline on a random sample of 280 training images (Save and Run All), you don't encounter any errors?"
        },
        {
          "id": 1852175,
          "postDate": "2022-07-11T21:08:24.773Z",
          "content": "<p>correct, I can convert all 751 out of 754 images from the training dataset on the limited instance and the three remaining are safely skipped… 👀</p>",
          "rawMarkdown": "correct, I can convert all 751 out of 754 images from the training dataset on the limited instance and the three remaining are safely skipped... 👀"
        },
        {
          "id": 1852372,
          "postDate": "2022-07-12T03:08:03.150Z",
          "content": "<p>Hmm, I'm not sure. Is this the notebook that crashes?</p>\n<p><a href=\"https://www.kaggle.com/code/jirkaborovec/bloodclots-classif-baseline-flash-effnet\" target=\"_blank\">https://www.kaggle.com/code/jirkaborovec/bloodclots-classif-baseline-flash-effnet</a></p>",
          "rawMarkdown": "Hmm, I'm not sure. Is this the notebook that crashes?\n\nhttps://www.kaggle.com/code/jirkaborovec/bloodclots-classif-baseline-flash-effnet"
        },
        {
          "id": 1852570,
          "postDate": "2022-07-12T06:56:09.523Z",
          "content": "<p>Yes, this one </p>",
          "rawMarkdown": "Yes, this one "
        }
      ]
    },
    {
      "id": 1851861,
      "postDate": "2022-07-11T15:30:42.160Z",
      "content": "<p>Image.load itself will fail if the tiff is too big… I had my successful submission by using pyvips</p>",
      "rawMarkdown": "Image.load itself will fail if the tiff is too big... I had my successful submission by using pyvips",
      "replies": [
        {
          "id": 1851879,
          "postDate": "2022-07-11T15:45:06.597Z",
          "content": "<p>I see but that is not part of the environment so you have to download the OS package and then save it as a dataset…<br>\nbut I do not use <code>Image.load</code>, just <code>Image.open</code>, so do you think that these images can't be even opened? :(</p>",
          "rawMarkdown": "I see but that is not part of the environment so you have to download the OS package and then save it as a dataset...\nbut I do not use `Image.load`, just `Image.open`, so do you think that these images can't be even opened? :("
        },
        {
          "id": 1851966,
          "postDate": "2022-07-11T17:15:54.103Z",
          "content": "<p>Sry, I meant, Image.open opens the whole image, Image.load does not exist, typo, yes I think these images can not even be opened</p>",
          "rawMarkdown": "Sry, I meant, Image.open opens the whole image, Image.load does not exist, typo, yes I think these images can not even be opened"
        },
        {
          "id": 1851979,
          "postDate": "2022-07-11T17:23:34.740Z",
          "content": "<p>what I see, <code>Image.open</code> doe snot take almost any memory, it does when you call <code>Image.read</code></p>",
          "rawMarkdown": "what I see, `Image.open` doe snot take almost any memory, it does when you call `Image.read`",
          "votes": 1
        },
        {
          "id": 1852002,
          "postDate": "2022-07-11T17:34:34.903Z",
          "content": "<p>I saw your submission notebook, I suggest you use pylibs anyways, it works for me…, although I have currently submitted my submission notebook, it will verify if pylibs works with very big sizes if this submission does not give an error (I first used it, imediately converting images to HigherThan1024x1024, it did not give me a submission error)</p>",
          "rawMarkdown": "I saw your submission notebook, I suggest you use pylibs anyways, it works for me..., although I have currently submitted my submission notebook, it will verify if pylibs works with very big sizes if this submission does not give an error (I first used it, imediately converting images to HigherThan1024x1024, it did not give me a submission error)"
        },
        {
          "id": 1852140,
          "postDate": "2022-07-11T20:01:13.247Z",
          "content": "<p>Do you have/ mind sharing your notebook with this loading? =) </p>",
          "rawMarkdown": "Do you have/ mind sharing your notebook with this loading? =) "
        },
        {
          "id": 1856272,
          "postDate": "2022-07-15T08:26:50.497Z",
          "content": "<p>just verified that <code>Image.open</code> does not load any data and I can apply it to all images in the test set :)</p>\n<blockquote>\n  <p><a href=\"https://www.kaggle.com/code/jirkaborovec/bloodclots-classif-fake-stats-prediction\" target=\"_blank\">https://www.kaggle.com/code/jirkaborovec/bloodclots-classif-fake-stats-prediction</a></p>\n</blockquote>",
          "rawMarkdown": "just verified that `Image.open` does not load any data and I can apply it to all images in the test set :)\n> https://www.kaggle.com/code/jirkaborovec/bloodclots-classif-fake-stats-prediction"
        },
        {
          "id": 1894131,
          "postDate": "2022-08-11T09:37:49.917Z",
          "content": "<p>when using pyvips, have you found memory leak? <br>\nonly run this line for every WSI, the memory will increase to 13g quikly<br>\n<code>img = pyvips.Image.thumbnail(path, width).numpy()</code></p>",
          "rawMarkdown": "when using pyvips, have you found memory leak? \nonly run this line for every WSI, the memory will increase to 13g quikly\n`img = pyvips.Image.thumbnail(path, width).numpy()`"
        },
        {
          "id": 1894137,
          "postDate": "2022-08-11T09:44:45.810Z",
          "content": "<p>Have you tried to set this env. variable? <code>os.environ['VIPS_DISC_THRESHOLD'] = '7gb'</code></p>",
          "rawMarkdown": "Have you tried to set this env. variable? `os.environ['VIPS_DISC_THRESHOLD'] = '7gb'`"
        },
        {
          "id": 1895340,
          "postDate": "2022-08-12T04:48:03.343Z",
          "content": "<p>no…<br>\ni think memory will also increase to 7gb, even with <code>os.environ['VIPS_DISC_THRESHOLD'] = '7gb'</code> , it's a memory leak bug of pyvips libary.<br>\neven though memory increases to 13g, it will not throw OOM error, the code still can run to the final submission.</p>",
          "rawMarkdown": "no...\ni think memory will also increase to 7gb, even with ` os.environ['VIPS_DISC_THRESHOLD'] = '7gb' ` , it's a memory leak bug of pyvips libary.\neven though memory increases to 13g, it will not throw OOM error, the code still can run to the final submission."
        },
        {
          "id": 1903142,
          "postDate": "2022-08-17T06:23:12.663Z",
          "content": "<p>How many threads are you running? Because eventually the 7gb would be used baby each not together…</p>",
          "rawMarkdown": "How many threads are you running? Because eventually the 7gb would be used baby each not together..."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1907793,
      "author_name": "bobfromjapan",
      "author_url": "",
      "post_date": "2022-08-21T04:52:48.877000",
      "content": "<p>Hello! This is probably a re-working of Topic Author's solution, but I would like to make a few comments.</p>\n<p>I too encountered the \"Notebook Exceeded Allowed Compute\" in my submission pipeline.<br>\nIt seems that the OOM was occurring when loading images using the <code>tifffile</code>.</p>\n<p>Since the RAM limit for a GPU instance is 13 GB, I used the train data to investigate the maximum file size (width x height) that could not be loaded due to this limitation.<br>\n<a href=\"https://www.kaggle.com/code/bobfromjapan/mayo-filesize-inspection\" target=\"_blank\">https://www.kaggle.com/code/bobfromjapan/mayo-filesize-inspection</a></p>\n<p>The results showed that the largest image in train data, <code>6baf51_0.tif</code>, has a size of <code>4896084492</code>, which would exceed the 13GB System RAM of the GPU instance.<br>\nThe second largest image, <code>b894f4_0.tif</code>, is <code>4131662535</code> and can be loaded.</p>\n<p>I tried to skip processing when the size exceeds <code>4131662535</code> as a workaround to avoid the \"Notebook Exceeded Allowed Compute\". Below is an example.</p>\n<pre><code>too_big_for_process = []\nfor path in test_image_paths:\n\n    print(path)\n    slide = OpenSlide(path)\n\n    if slide.dimensions[0]*slide.dimensions[1] &lt; 4131662535:\n        image_id = os.path.splitext(os.path.basename(path))[0]\n        image = tiff.imread(path)\n        print(f\"{image.shape}\")\n        cv2.imwrite(os.path.join(output_dir, f\"{image_id}.jpg\"), image[::scale,::scale,::-1])\n        del image\n        gc.collect()\n    else:\n        print(\"Skip process for avoiding OOM possibility.\")\n        too_big_for_process.append(path)\n</code></pre>\n<p>After the inference, if there were images that could not be processed, the value of prediction was given as a constant. This avoids the \"Notebook Exceeded Allowed Compute\" and the submission succeeds!</p>\n<pre><code>if len(too_big_for_process)&gt;0:\n    for i in too_big_for_process:\n        preds.append((i, 0.82, 0.28))\n</code></pre>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1853611,
      "author_name": "Harshit Sheoran",
      "author_url": "",
      "post_date": "2022-07-13T02:12:34.897000",
      "content": "<p>I ran into a time error, made it 20% faster to barely finish it in time with 1 fold (HUGE OVERFIT BTW, leaderboard is nothing as training on my set), although I am unsure, what will happen if my run 5 folds instead </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1853807,
          "author_name": "Jirka",
          "author_url": "",
          "post_date": "2022-07-13T06:51:09.570000",
          "content": "<p>So you say it is timeout issue? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1853996,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2022-07-13T11:20:06.410000",
          "content": "<p>I had to make tiles, which if I load images carefully avoiding memory errors, takes too much time…</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1854026,
          "author_name": "Jirka",
          "author_url": "",
          "post_date": "2022-07-13T11:45:35.370000",
          "content": "<p>what does too much time mean? mine is killed after bout 2h of runtime…<br>\neven they state:</p>\n<blockquote>\n  <p>CPU Notebook &lt;= 9 hours run-time<br>\n  GPU Notebook &lt;= 9 hours run-time<br>\n  Internet access disabled</p>\n</blockquote>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1854098,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2022-07-13T12:35:34.157000",
          "content": "<p>It means, mine took more than 9 hours to executue, so if there occurs no error, instead of showing me the score it will show 'Notebook Timeout Error' </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1854295,
          "author_name": "Jirka",
          "author_url": "",
          "post_date": "2022-07-13T15:28:50.637000",
          "content": "<p>I see, the \"Notebook Timeout Error\" is quite clear… :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1853553,
      "author_name": "The Devastator",
      "author_url": "",
      "post_date": "2022-07-13T00:25:39.647000",
      "content": "<p>I think that the test is huge.. I keep getting errors every day when I try to run a submission..<br>\nMaybe try converting them in batches and you should see an improvement.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1853805,
          "author_name": "Jirka",
          "author_url": "",
          "post_date": "2022-07-13T06:49:30.587000",
          "content": "<p>Hi, may you elaborate more on what you mean converting in batches? </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1852084,
      "author_name": "Darien Schettler",
      "author_url": "",
      "post_date": "2022-07-11T18:41:59.003000",
      "content": "<p>The real test set is MUCH larger than the provided test set. You need to debug with a representative dataset.</p>\n<p>I recommend you do the following to debug:</p>\n<ul>\n<li>instead of using the provided test data use a portion of the training data equivalent to the size  of the hidden test set</li>\n<li>Use something like debug==len(ss_df)==4 and then use that to initialize a more representative dataset.</li>\n<li>if I had to guess I’d say PIL is holding some RAM and causing an eventual OOM. </li>\n</ul>",
      "votes": 0,
      "replies": [
        {
          "id": 1852138,
          "author_name": "Jirka",
          "author_url": "",
          "post_date": "2022-07-11T20:00:16.430000",
          "content": "<p>Yes, that is what I tried and I do not have any problem with the train dataset as I skipped all too large image even before loading… </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1852157,
          "author_name": "Darien Schettler",
          "author_url": "",
          "post_date": "2022-07-11T20:46:57.130000",
          "content": "<p>So if you run your pipeline on a random sample of 280 training images (Save and Run All), you don't encounter any errors?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1852175,
          "author_name": "Jirka",
          "author_url": "",
          "post_date": "2022-07-11T21:08:24.773000",
          "content": "<p>correct, I can convert all 751 out of 754 images from the training dataset on the limited instance and the three remaining are safely skipped… 👀</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1852372,
          "author_name": "Darien Schettler",
          "author_url": "",
          "post_date": "2022-07-12T03:08:03.150000",
          "content": "<p>Hmm, I'm not sure. Is this the notebook that crashes?</p>\n<p><a href=\"https://www.kaggle.com/code/jirkaborovec/bloodclots-classif-baseline-flash-effnet\" target=\"_blank\">https://www.kaggle.com/code/jirkaborovec/bloodclots-classif-baseline-flash-effnet</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1852570,
          "author_name": "Jirka",
          "author_url": "",
          "post_date": "2022-07-12T06:56:09.523000",
          "content": "<p>Yes, this one </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1851861,
      "author_name": "Harshit Sheoran",
      "author_url": "",
      "post_date": "2022-07-11T15:30:42.160000",
      "content": "<p>Image.load itself will fail if the tiff is too big… I had my successful submission by using pyvips</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1851879,
          "author_name": "Jirka",
          "author_url": "",
          "post_date": "2022-07-11T15:45:06.597000",
          "content": "<p>I see but that is not part of the environment so you have to download the OS package and then save it as a dataset…<br>\nbut I do not use <code>Image.load</code>, just <code>Image.open</code>, so do you think that these images can't be even opened? :(</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1851966,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2022-07-11T17:15:54.103000",
          "content": "<p>Sry, I meant, Image.open opens the whole image, Image.load does not exist, typo, yes I think these images can not even be opened</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1851979,
          "author_name": "Jirka",
          "author_url": "",
          "post_date": "2022-07-11T17:23:34.740000",
          "content": "<p>what I see, <code>Image.open</code> doe snot take almost any memory, it does when you call <code>Image.read</code></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1852002,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2022-07-11T17:34:34.903000",
          "content": "<p>I saw your submission notebook, I suggest you use pylibs anyways, it works for me…, although I have currently submitted my submission notebook, it will verify if pylibs works with very big sizes if this submission does not give an error (I first used it, imediately converting images to HigherThan1024x1024, it did not give me a submission error)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1852140,
          "author_name": "Jirka",
          "author_url": "",
          "post_date": "2022-07-11T20:01:13.247000",
          "content": "<p>Do you have/ mind sharing your notebook with this loading? =) </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1856272,
          "author_name": "Jirka",
          "author_url": "",
          "post_date": "2022-07-15T08:26:50.497000",
          "content": "<p>just verified that <code>Image.open</code> does not load any data and I can apply it to all images in the test set :)</p>\n<blockquote>\n  <p><a href=\"https://www.kaggle.com/code/jirkaborovec/bloodclots-classif-fake-stats-prediction\" target=\"_blank\">https://www.kaggle.com/code/jirkaborovec/bloodclots-classif-fake-stats-prediction</a></p>\n</blockquote>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1894131,
          "author_name": "ddt0106",
          "author_url": "",
          "post_date": "2022-08-11T09:37:49.917000",
          "content": "<p>when using pyvips, have you found memory leak? <br>\nonly run this line for every WSI, the memory will increase to 13g quikly<br>\n<code>img = pyvips.Image.thumbnail(path, width).numpy()</code></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1894137,
          "author_name": "Jirka",
          "author_url": "",
          "post_date": "2022-08-11T09:44:45.810000",
          "content": "<p>Have you tried to set this env. variable? <code>os.environ['VIPS_DISC_THRESHOLD'] = '7gb'</code></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1895340,
          "author_name": "ddt0106",
          "author_url": "",
          "post_date": "2022-08-12T04:48:03.343000",
          "content": "<p>no…<br>\ni think memory will also increase to 7gb, even with <code>os.environ['VIPS_DISC_THRESHOLD'] = '7gb'</code> , it's a memory leak bug of pyvips libary.<br>\neven though memory increases to 13g, it will not throw OOM error, the code still can run to the final submission.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1903142,
          "author_name": "Jirka",
          "author_url": "",
          "post_date": "2022-08-17T06:23:12.663000",
          "content": "<p>How many threads are you running? Because eventually the 7gb would be used baby each not together…</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1851612": "I am getting this strange error even though my run with provided test images works fine and I have set a limit of image size so it won't kill the small RAM... is there something dramatically different with the used machine for submission? Also, the submission details state we have 9 hours runtime, and this kernel dies after 2 hours or so...\n\n> Submission kernel: https://www.kaggle.com/code/jirkaborovec/bloodclots-classif-baseline-flash-effnet\n> Conversion kernel: https://www.kaggle.com/code/jirkaborovec/bloodclots-classif-eda-load-crop-images\n\nto show that I can convert all images up to 4e9 pixels without problem...\n\n**UPDATE:**\n\na good finding from: https://www.kaggle.com/competitions/mayo-clinic-strip-ai/discussion/336819\n\n> If you don't use gpu, you have 16 Gb of RAM, but if you use gpu, it's reduced to 13 Gb on the kaggle notebook.\n\nAlso, I was able to load and covert images up to 1.5e9 pixels with GPU",
    "1907793": "Hello! This is probably a re-working of Topic Author's solution, but I would like to make a few comments.\n\nI too encountered the \"Notebook Exceeded Allowed Compute\" in my submission pipeline.\nIt seems that the OOM was occurring when loading images using the `tifffile`.\n\nSince the RAM limit for a GPU instance is 13 GB, I used the train data to investigate the maximum file size (width x height) that could not be loaded due to this limitation.\nhttps://www.kaggle.com/code/bobfromjapan/mayo-filesize-inspection\n\nThe results showed that the largest image in train data, `6baf51_0.tif`, has a size of `4896084492`, which would exceed the 13GB System RAM of the GPU instance.\nThe second largest image, `b894f4_0.tif`, is `4131662535` and can be loaded.\n\nI tried to skip processing when the size exceeds `4131662535` as a workaround to avoid the \"Notebook Exceeded Allowed Compute\". Below is an example.\n\n```\ntoo_big_for_process = []\nfor path in test_image_paths:\n\n    print(path)\n    slide = OpenSlide(path)\n\n    if slide.dimensions[0]*slide.dimensions[1] < 4131662535:\n        image_id = os.path.splitext(os.path.basename(path))[0]\n        image = tiff.imread(path)\n        print(f\"{image.shape}\")\n        cv2.imwrite(os.path.join(output_dir, f\"{image_id}.jpg\"), image[::scale,::scale,::-1])\n        del image\n        gc.collect()\n    else:\n        print(\"Skip process for avoiding OOM possibility.\")\n        too_big_for_process.append(path)\n\n```\nAfter the inference, if there were images that could not be processed, the value of prediction was given as a constant. This avoids the \"Notebook Exceeded Allowed Compute\" and the submission succeeds!\n\n```\nif len(too_big_for_process)>0:\n    for i in too_big_for_process:\n        preds.append((i, 0.82, 0.28))\n```",
    "1853611": "I ran into a time error, made it 20% faster to barely finish it in time with 1 fold (HUGE OVERFIT BTW, leaderboard is nothing as training on my set), although I am unsure, what will happen if my run 5 folds instead ",
    "1853553": "I think that the test is huge.. I keep getting errors every day when I try to run a submission..\nMaybe try converting them in batches and you should see an improvement.\n\n",
    "1852084": "The real test set is MUCH larger than the provided test set. You need to debug with a representative dataset.\n\nI recommend you do the following to debug:\n\n- instead of using the provided test data use a portion of the training data equivalent to the size  of the hidden test set\n- Use something like debug==len(ss_df)==4 and then use that to initialize a more representative dataset.\n- if I had to guess I’d say PIL is holding some RAM and causing an eventual OOM. \n",
    "1851861": "Image.load itself will fail if the tiff is too big... I had my successful submission by using pyvips"
  }
}