{
  "id": 461775,
  "title": "Only one WSI prevents me from completing a full submission...",
  "url": "/competitions/UBC-OCEAN/discussion/461775",
  "author_name": "Patrick Robitaille",
  "post_date": "2023-12-16T10:43:32.124000",
  "votes": 5,
  "comment_count": 13,
  "views": 0,
  "content": "<p>I have spent a lot of time in the last week to troubleshoot \"in the dark\" my inference code in order to complete successfully a full submission. Unfortunately, I keep getting the frustrating and non-informative \"Notebook Threw Exception\" result.</p>\n<p>I have tested my inference code by running it using the full training set, and it completes without any exceptions raised. When faced with these repeated 'Notebook Threw Exception\" results, I bubble-wrapped my code with try-except statements at various locations of possible failure and also ensured the following:</p>\n<ul>\n<li>That padding is applied when cropping patches that might expand outside the boundaries of the TMA/WSI (probably more important for TMAs than WSIs)</li>\n<li>That the number of patches from a WSI is limited to a certain maximum to avoid out-of-memory problems with the GPU. While I managed to have this maximum at about 700 on my local PC (RTX 4080), I found that it needed to be below 500 on the Kaggle GPU. I have tried to set this at 450, then 400, but I still get the same result.</li>\n<li>That I have mechanisms in place at various locations in my code to ensure the rejection of patches that are too uniform in color. Only one of these among selected patches from a WSI can cause a model to protest fairly quickly.</li>\n</ul>\n<p>After much probing, I have been able to conclude that the sole culprit for the \"Notebook Threw Exception\" results is one single WSI with the following characteristics (after separating the test images between WSIs and TMAs, processing the WSIs first):</p>\n<ul>\n<li>If I process the WSIs in ascending order of size (width x height), the notebook returns the exception at about 3.5 hours of execution; in descending order of size, it crashes at about 4.5 hours. So, the WSI is of median-to-average size, which goes against the out-of-memory possibility.</li>\n<li>If I process the WSIs in ascending image_id, the notebook crashes after 82 minutes (the code is still running for descending order, but I would think that it would be around 6.5 hours). So, the image_id is probably 1XXXX, with the first X possibly in the low digits.</li>\n</ul>\n<p>I'm really at a loss and short of ideas to fix this problem, especially when I have to do this \"in the dark\" with no other information than a stupid \"Notebook Threw Exception\" message. Does anybody have any suggestions on what I could try next? Has anybody come across this type of situation?</p>\n<p>(Sidenote: when s**t happens, it happens a lot; my PC is on life-support since Monday, most likely due to a faulty motherboard; obviously, all my files for this competition were on the hard drive, yet to be backed up, facing the likelihood of irremediable oblivion; I'm anxiously waiting for a call from the tech heads looking into it. So, any help from you guys would be REALLY appreciated)</p>",
  "messages": [
    {
      "id": 2563500,
      "postDate": "2023-12-16T10:43:32.123Z",
      "content": "<p>I have spent a lot of time in the last week to troubleshoot \"in the dark\" my inference code in order to complete successfully a full submission. Unfortunately, I keep getting the frustrating and non-informative \"Notebook Threw Exception\" result.</p>\n<p>I have tested my inference code by running it using the full training set, and it completes without any exceptions raised. When faced with these repeated 'Notebook Threw Exception\" results, I bubble-wrapped my code with try-except statements at various locations of possible failure and also ensured the following:</p>\n<ul>\n<li>That padding is applied when cropping patches that might expand outside the boundaries of the TMA/WSI (probably more important for TMAs than WSIs)</li>\n<li>That the number of patches from a WSI is limited to a certain maximum to avoid out-of-memory problems with the GPU. While I managed to have this maximum at about 700 on my local PC (RTX 4080), I found that it needed to be below 500 on the Kaggle GPU. I have tried to set this at 450, then 400, but I still get the same result.</li>\n<li>That I have mechanisms in place at various locations in my code to ensure the rejection of patches that are too uniform in color. Only one of these among selected patches from a WSI can cause a model to protest fairly quickly.</li>\n</ul>\n<p>After much probing, I have been able to conclude that the sole culprit for the \"Notebook Threw Exception\" results is one single WSI with the following characteristics (after separating the test images between WSIs and TMAs, processing the WSIs first):</p>\n<ul>\n<li>If I process the WSIs in ascending order of size (width x height), the notebook returns the exception at about 3.5 hours of execution; in descending order of size, it crashes at about 4.5 hours. So, the WSI is of median-to-average size, which goes against the out-of-memory possibility.</li>\n<li>If I process the WSIs in ascending image_id, the notebook crashes after 82 minutes (the code is still running for descending order, but I would think that it would be around 6.5 hours). So, the image_id is probably 1XXXX, with the first X possibly in the low digits.</li>\n</ul>\n<p>I'm really at a loss and short of ideas to fix this problem, especially when I have to do this \"in the dark\" with no other information than a stupid \"Notebook Threw Exception\" message. Does anybody have any suggestions on what I could try next? Has anybody come across this type of situation?</p>\n<p>(Sidenote: when s**t happens, it happens a lot; my PC is on life-support since Monday, most likely due to a faulty motherboard; obviously, all my files for this competition were on the hard drive, yet to be backed up, facing the likelihood of irremediable oblivion; I'm anxiously waiting for a call from the tech heads looking into it. So, any help from you guys would be REALLY appreciated)</p>",
      "rawMarkdown": "I have spent a lot of time in the last week to troubleshoot \"in the dark\" my inference code in order to complete successfully a full submission. Unfortunately, I keep getting the frustrating and non-informative \"Notebook Threw Exception\" result.\n\nI have tested my inference code by running it using the full training set, and it completes without any exceptions raised. When faced with these repeated 'Notebook Threw Exception\" results, I bubble-wrapped my code with try-except statements at various locations of possible failure and also ensured the following:\n\n- That padding is applied when cropping patches that might expand outside the boundaries of the TMA/WSI (probably more important for TMAs than WSIs)\n- That the number of patches from a WSI is limited to a certain maximum to avoid out-of-memory problems with the GPU. While I managed to have this maximum at about 700 on my local PC (RTX 4080), I found that it needed to be below 500 on the Kaggle GPU. I have tried to set this at 450, then 400, but I still get the same result.\n- That I have mechanisms in place at various locations in my code to ensure the rejection of patches that are too uniform in color. Only one of these among selected patches from a WSI can cause a model to protest fairly quickly.\n\nAfter much probing, I have been able to conclude that the sole culprit for the \"Notebook Threw Exception\" results is one single WSI with the following characteristics (after separating the test images between WSIs and TMAs, processing the WSIs first):\n\n- If I process the WSIs in ascending order of size (width x height), the notebook returns the exception at about 3.5 hours of execution; in descending order of size, it crashes at about 4.5 hours. So, the WSI is of median-to-average size, which goes against the out-of-memory possibility.\n- If I process the WSIs in ascending image_id, the notebook crashes after 82 minutes (the code is still running for descending order, but I would think that it would be around 6.5 hours). So, the image_id is probably 1XXXX, with the first X possibly in the low digits.\n\nI'm really at a loss and short of ideas to fix this problem, especially when I have to do this \"in the dark\" with no other information than a stupid \"Notebook Threw Exception\" message. Does anybody have any suggestions on what I could try next? Has anybody come across this type of situation?\n\n(Sidenote: when s**t happens, it happens a lot; my PC is on life-support since Monday, most likely due to a faulty motherboard; obviously, all my files for this competition were on the hard drive, yet to be backed up, facing the likelihood of irremediable oblivion; I'm anxiously waiting for a call from the tech heads looking into it. So, any help from you guys would be REALLY appreciated)",
      "votes": 4
    },
    {
      "id": 2565756,
      "postDate": "2023-12-18T09:27:06.540Z",
      "content": "<p>A direct approach maybe: You have successful submissions, try to replicate the changes you have made step by step (from working submission to the failing).</p>\n<p>Let me also share an indirect approach I had to take for a similar experience:<br>\nIn my case I wasn't even getting any exception, where the problem was somewhere within parallel processing script so the exception was stopping that specific process but the notebook was continuing with the subsequent cells/code, which for me meant a zero score (because rest of the code was actually good enough to open and save sample_submission without the labels as 'submission.csv'.)</p>\n<p>All experiments with train set (full trial, wsi only, tma only, sorted, reversesorted) were running OK.</p>\n<p>The problem appeared to be the first stage model that I was using to create a segmentation mask for the second stage classifier to predict. With train set it was OK, because this baseline segmentator was doing fine to segment variants of the WSI's from the same data distribution as it was trained. </p>\n<p>But obviously it was having problems with the test set, so I made sure there is a fallback option if the output of the segmentation processing did not result in required input for the next stage.</p>\n<p>Be careful I am not saying the segmentation model was not doing its thing but maybe the cropareas calculation afterwards was the problem such as having 0 width/height rectangles. </p>\n<p>I hope you can find it sooner than later, good luck!</p>",
      "rawMarkdown": "A direct approach maybe: You have successful submissions, try to replicate the changes you have made step by step (from working submission to the failing).\n\nLet me also share an indirect approach I had to take for a similar experience:\nIn my case I wasn't even getting any exception, where the problem was somewhere within parallel processing script so the exception was stopping that specific process but the notebook was continuing with the subsequent cells/code, which for me meant a zero score (because rest of the code was actually good enough to open and save sample_submission without the labels as 'submission.csv'.)\n\nAll experiments with train set (full trial, wsi only, tma only, sorted, reversesorted) were running OK.\n\nThe problem appeared to be the first stage model that I was using to create a segmentation mask for the second stage classifier to predict. With train set it was OK, because this baseline segmentator was doing fine to segment variants of the WSI's from the same data distribution as it was trained. \n\nBut obviously it was having problems with the test set, so I made sure there is a fallback option if the output of the segmentation processing did not result in required input for the next stage.\n\nBe careful I am not saying the segmentation model was not doing its thing but maybe the cropareas calculation afterwards was the problem such as having 0 width/height rectangles. \n\nI hope you can find it sooner than later, good luck!",
      "votes": 1,
      "replies": [
        {
          "id": 2565992,
          "postDate": "2023-12-18T11:54:21.717Z",
          "content": "<p>Thanks Guner. My successful submissions so far were only using the TMAs in the test set, that's how I know that it is a WSI that is the problem. I can't really comment on your specific experience, as I am not using the segmentation masks. However, I know that the issue might lie in my patch selection pipeline (cropping is definitely no longer an issue for me). I have tried about a dozen possibilities, yet I have not been able to put the finger on it. It is a very marginal edge case, since I can process ~800-1000 different WSIs without a problem… just not this one very specific WSI.</p>",
          "rawMarkdown": "Thanks Guner. My successful submissions so far were only using the TMAs in the test set, that's how I know that it is a WSI that is the problem. I can't really comment on your specific experience, as I am not using the segmentation masks. However, I know that the issue might lie in my patch selection pipeline (cropping is definitely no longer an issue for me). I have tried about a dozen possibilities, yet I have not been able to put the finger on it. It is a very marginal edge case, since I can process ~800-1000 different WSIs without a problem... just not this one very specific WSI."
        }
      ]
    },
    {
      "id": 2563663,
      "postDate": "2023-12-16T13:52:24.057Z",
      "content": "<p>Is it possible that some WSI has 4 channels, since it's saved with <code>.png</code> format? I have never tested, just a guess. </p>",
      "rawMarkdown": "Is it possible that some WSI has 4 channels, since it's saved with `.png` format? I have never tested, just a guess. ",
      "votes": 1,
      "replies": [
        {
          "id": 2564165,
          "postDate": "2023-12-16T21:34:12.183Z",
          "content": "<p>My code already takes care of that possibility. And such an error would probably be caught easily with a try-except statement. My problem appears to be trickier than that.</p>",
          "rawMarkdown": "My code already takes care of that possibility. And such an error would probably be caught easily with a try-except statement. My problem appears to be trickier than that."
        }
      ]
    },
    {
      "id": 2568309,
      "postDate": "2023-12-20T12:08:43.513Z",
      "content": "<p>I got some problems on very few WSI also… what I did is just randomly assign a label (with the distribution of the training set) to problematic images. At least the notebooks run, and the impact on the metric is minimal </p>",
      "rawMarkdown": "I got some problems on very few WSI also… what I did is just randomly assign a label (with the distribution of the training set) to problematic images. At least the notebooks run, and the impact on the metric is minimal ",
      "votes": 2,
      "replies": [
        {
          "id": 2568656,
          "postDate": "2023-12-20T18:21:17.650Z",
          "content": "<p>it seems like you are good to go :-)</p>",
          "rawMarkdown": "it seems like you are good to go :-)"
        }
      ]
    },
    {
      "id": 2563595,
      "postDate": "2023-12-16T12:56:32.467Z",
      "content": "<p>I understand your frustration, a few other also had similar problem, <br>\nwhat I suggested is that, take your model, and run it on whole training set, exactly like how you run your submission notebook, but do it on train.csv and train images, don't submit it, but just run it to monitor any error</p>\n<p>this could potentially help</p>",
      "rawMarkdown": "I understand your frustration, a few other also had similar problem, \nwhat I suggested is that, take your model, and run it on whole training set, exactly like how you run your submission notebook, but do it on train.csv and train images, don't submit it, but just run it to monitor any error\n\nthis could potentially help",
      "replies": [
        {
          "id": 2564163,
          "postDate": "2023-12-16T21:31:05.053Z",
          "content": "<p>Thanks, but I have already done that, as mentioned in my initial comment.</p>",
          "rawMarkdown": "Thanks, but I have already done that, as mentioned in my initial comment."
        }
      ]
    },
    {
      "id": 2563511,
      "postDate": "2023-12-16T11:00:27.113Z",
      "content": "<p>I should also mention that I wrapped my model call into a try-except statement:</p>\n<p><code>try:\n            preds = model(X)\n        except:\n(do something else)\n</code></p>\n<p>but it doesn't seem to catch any exception. So, it seems that it could still be a GPU memory-related problem (that would be a RuntimeError) and/or an error that would cause a system exit.</p>",
      "rawMarkdown": "I should also mention that I wrapped my model call into a try-except statement:\n\n`try:\n            preds = model(X)\n        except:\n(do something else)\n`\n\nbut it doesn't seem to catch any exception. So, it seems that it could still be a GPU memory-related problem (that would be a RuntimeError) and/or an error that would cause a system exit."
    },
    {
      "id": 2564344,
      "postDate": "2023-12-17T03:38:56.980Z",
      "content": "<p>This solved post may be helpful.</p>\n<p><a href=\"https://www.kaggle.com/competitions/UBC-OCEAN/discussion/461341\" target=\"_blank\">https://www.kaggle.com/competitions/UBC-OCEAN/discussion/461341</a></p>\n<p>Try setting the submission notebook batch size to 1.<br>\nIn my case, \"Notebook Throw Exception\" occurred when the batch size was 2 or more.</p>\n<p>Just in case, please try setting num_workers of DataLoader to 1 as well.</p>\n<p>If you are using PIL like me, it worked if you set Image.MAX_IMAGE_PIXELS as below and then changed the image loading code as below.<br>\n(This part needs to be modified to suit your code)</p>\n<pre><code> PIL  Image\nImage.MAX_IMAGE_PIXELS =  * \n</code></pre>\n<pre><code> ():\n      () -&gt; :\n         img_path = self.df.iloc[idx][]\n\n         :\n             tile = Image.(img_path)\n         :\n             \n             \n             tile = Image.fromarray(np.zeros((, , )).astype(np.uint8))\n</code></pre>\n<p>Please refer to the URL below for the entire code.</p>\n<p><a href=\"https://www.kaggle.com/code/yamitomo/j010-minimum-submit-ok\" target=\"_blank\">https://www.kaggle.com/code/yamitomo/j010-minimum-submit-ok</a></p>",
      "rawMarkdown": "This solved post may be helpful.\n\nhttps://www.kaggle.com/competitions/UBC-OCEAN/discussion/461341\n\nTry setting the submission notebook batch size to 1.\nIn my case, \"Notebook Throw Exception\" occurred when the batch size was 2 or more.\n\nJust in case, please try setting num_workers of DataLoader to 1 as well.\n\nIf you are using PIL like me, it worked if you set Image.MAX_IMAGE_PIXELS as below and then changed the image loading code as below.\n(This part needs to be modified to suit your code)\n\n```python\nfrom PIL import Image\nImage.MAX_IMAGE_PIXELS = 7000 * 7000\n```\n\n```python\nclass UBCDatasetInfer(Dataset):\n     def __getitem__(self, idx: int) -> tuple:\n         img_path = self.df.iloc[idx][\"path\"]\n        \n         try:\n             tile = Image.open(img_path)\n         except:\n             # thumb_path = self.df.iloc[idx][\"thumb_path\"]\n             # tile = Image.open(thumb_path)\n             tile = Image.fromarray(np.zeros((1000, 1000, 3)).astype(np.uint8))\n```\n\nPlease refer to the URL below for the entire code.\n\nhttps://www.kaggle.com/code/yamitomo/j010-minimum-submit-ok",
      "isDeleted": true,
      "replies": [
        {
          "id": 2564476,
          "postDate": "2023-12-17T06:25:30.040Z",
          "content": "<p>Thanks for the suggestions. I was already using <code>batch_size = 1</code> and <code>num_workers = 1</code>. I am also setting <code>Image.MAX_IMAGE_PIXELS = None</code>, this gets rid of any warning/error messages with respect to image size.</p>\n<p>I've done a bit more probing and have identified my model call (<code>preds = model(X)</code>) as the source of the problem. Now I have to figure out why this call works for about 2500 cases and fails for only 1…</p>",
          "rawMarkdown": "Thanks for the suggestions. I was already using `batch_size = 1` and `num_workers = 1`. I am also setting `Image.MAX_IMAGE_PIXELS = None`, this gets rid of any warning/error messages with respect to image size.\n\nI've done a bit more probing and have identified my model call (`preds = model(X)`) as the source of the problem. Now I have to figure out why this call works for about 2500 cases and fails for only 1...",
          "votes": 1,
          "replies": [
            {
              "id": 2564579,
              "postDate": "2023-12-17T08:10:03.337Z",
              "content": "<p>If you set Image.MAX_IMAGE_PIXELS = None, when loading images, there is no upper limit on the number of pixels.</p>\n<p>Therefore, when loading a large image, it will crash due to excessive memory usage.</p>\n<p>If you set Image.MAX_IMAGE_PIXELS = 7000 * 7000 like in my code, when a large image comes in the image loading process of the second code, it will not be loaded and an exception will occur.</p>\n<p>All that's left to do is handle the exception.</p>\n<p>I hope you find this information useful.</p>",
              "rawMarkdown": "If you set Image.MAX_IMAGE_PIXELS = None, when loading images, there is no upper limit on the number of pixels.\n\nTherefore, when loading a large image, it will crash due to excessive memory usage.\n\nIf you set Image.MAX_IMAGE_PIXELS = 7000 * 7000 like in my code, when a large image comes in the image loading process of the second code, it will not be loaded and an exception will occur.\n\nAll that's left to do is handle the exception.\n\nI hope you find this information useful.",
              "votes": 1,
              "isDeleted": true
            },
            {
              "id": 2565390,
              "postDate": "2023-12-18T02:24:42.020Z",
              "content": "<p>As mentioned above, the WSI that causes the problem is of median-to-average size. Setting <code>Image.MAX_IMAGE_PIXELS = None</code> in the context of this competition is perfectly safe. I also mentioned previously that I ran the inference code on the full training set and on the WSIs in the test set in descending order of size; for the latter case, the notebook threw the exception after around 4.5 hours of execution. If excessive RAM usage was the issue, the code would have crashed right at the start. </p>",
              "rawMarkdown": "As mentioned above, the WSI that causes the problem is of median-to-average size. Setting `Image.MAX_IMAGE_PIXELS = None` in the context of this competition is perfectly safe. I also mentioned previously that I ran the inference code on the full training set and on the WSIs in the test set in descending order of size; for the latter case, the notebook threw the exception after around 4.5 hours of execution. If excessive RAM usage was the issue, the code would have crashed right at the start. ",
              "votes": 1
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2565756,
      "author_name": "GUNER",
      "author_url": "",
      "post_date": "2023-12-18T09:27:06.540000",
      "content": "<p>A direct approach maybe: You have successful submissions, try to replicate the changes you have made step by step (from working submission to the failing).</p>\n<p>Let me also share an indirect approach I had to take for a similar experience:<br>\nIn my case I wasn't even getting any exception, where the problem was somewhere within parallel processing script so the exception was stopping that specific process but the notebook was continuing with the subsequent cells/code, which for me meant a zero score (because rest of the code was actually good enough to open and save sample_submission without the labels as 'submission.csv'.)</p>\n<p>All experiments with train set (full trial, wsi only, tma only, sorted, reversesorted) were running OK.</p>\n<p>The problem appeared to be the first stage model that I was using to create a segmentation mask for the second stage classifier to predict. With train set it was OK, because this baseline segmentator was doing fine to segment variants of the WSI's from the same data distribution as it was trained. </p>\n<p>But obviously it was having problems with the test set, so I made sure there is a fallback option if the output of the segmentation processing did not result in required input for the next stage.</p>\n<p>Be careful I am not saying the segmentation model was not doing its thing but maybe the cropareas calculation afterwards was the problem such as having 0 width/height rectangles. </p>\n<p>I hope you can find it sooner than later, good luck!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2565992,
          "author_name": "Patrick Robitaille",
          "author_url": "",
          "post_date": "2023-12-18T11:54:21.717000",
          "content": "<p>Thanks Guner. My successful submissions so far were only using the TMAs in the test set, that's how I know that it is a WSI that is the problem. I can't really comment on your specific experience, as I am not using the segmentation masks. However, I know that the issue might lie in my patch selection pipeline (cropping is definitely no longer an issue for me). I have tried about a dozen possibilities, yet I have not been able to put the finger on it. It is a very marginal edge case, since I can process ~800-1000 different WSIs without a problem… just not this one very specific WSI.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2563663,
      "author_name": "ForcewithMe",
      "author_url": "",
      "post_date": "2023-12-16T13:52:24.057000",
      "content": "<p>Is it possible that some WSI has 4 channels, since it's saved with <code>.png</code> format? I have never tested, just a guess. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2564165,
          "author_name": "Patrick Robitaille",
          "author_url": "",
          "post_date": "2023-12-16T21:34:12.183000",
          "content": "<p>My code already takes care of that possibility. And such an error would probably be caught easily with a try-except statement. My problem appears to be trickier than that.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2568309,
      "author_name": "prototype",
      "author_url": "",
      "post_date": "2023-12-20T12:08:43.513000",
      "content": "<p>I got some problems on very few WSI also… what I did is just randomly assign a label (with the distribution of the training set) to problematic images. At least the notebooks run, and the impact on the metric is minimal </p>",
      "votes": 2,
      "replies": [
        {
          "id": 2568656,
          "author_name": "GUNER",
          "author_url": "",
          "post_date": "2023-12-20T18:21:17.650000",
          "content": "<p>it seems like you are good to go :-)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2563595,
      "author_name": "Ali",
      "author_url": "",
      "post_date": "2023-12-16T12:56:32.467000",
      "content": "<p>I understand your frustration, a few other also had similar problem, <br>\nwhat I suggested is that, take your model, and run it on whole training set, exactly like how you run your submission notebook, but do it on train.csv and train images, don't submit it, but just run it to monitor any error</p>\n<p>this could potentially help</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2564163,
          "author_name": "Patrick Robitaille",
          "author_url": "",
          "post_date": "2023-12-16T21:31:05.053000",
          "content": "<p>Thanks, but I have already done that, as mentioned in my initial comment.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2563511,
      "author_name": "Patrick Robitaille",
      "author_url": "",
      "post_date": "2023-12-16T11:00:27.113000",
      "content": "<p>I should also mention that I wrapped my model call into a try-except statement:</p>\n<p><code>try:\n            preds = model(X)\n        except:\n(do something else)\n</code></p>\n<p>but it doesn't seem to catch any exception. So, it seems that it could still be a GPU memory-related problem (that would be a RuntimeError) and/or an error that would cause a system exit.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2564344,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-12-17T03:38:56.980000",
      "content": "<p>This solved post may be helpful.</p>\n<p><a href=\"https://www.kaggle.com/competitions/UBC-OCEAN/discussion/461341\" target=\"_blank\">https://www.kaggle.com/competitions/UBC-OCEAN/discussion/461341</a></p>\n<p>Try setting the submission notebook batch size to 1.<br>\nIn my case, \"Notebook Throw Exception\" occurred when the batch size was 2 or more.</p>\n<p>Just in case, please try setting num_workers of DataLoader to 1 as well.</p>\n<p>If you are using PIL like me, it worked if you set Image.MAX_IMAGE_PIXELS as below and then changed the image loading code as below.<br>\n(This part needs to be modified to suit your code)</p>\n<pre><code> PIL  Image\nImage.MAX_IMAGE_PIXELS =  * \n</code></pre>\n<pre><code> ():\n      () -&gt; :\n         img_path = self.df.iloc[idx][]\n\n         :\n             tile = Image.(img_path)\n         :\n             \n             \n             tile = Image.fromarray(np.zeros((, , )).astype(np.uint8))\n</code></pre>\n<p>Please refer to the URL below for the entire code.</p>\n<p><a href=\"https://www.kaggle.com/code/yamitomo/j010-minimum-submit-ok\" target=\"_blank\">https://www.kaggle.com/code/yamitomo/j010-minimum-submit-ok</a></p>",
      "votes": 0,
      "replies": [
        {
          "id": 2564476,
          "author_name": "Patrick Robitaille",
          "author_url": "",
          "post_date": "2023-12-17T06:25:30.040000",
          "content": "<p>Thanks for the suggestions. I was already using <code>batch_size = 1</code> and <code>num_workers = 1</code>. I am also setting <code>Image.MAX_IMAGE_PIXELS = None</code>, this gets rid of any warning/error messages with respect to image size.</p>\n<p>I've done a bit more probing and have identified my model call (<code>preds = model(X)</code>) as the source of the problem. Now I have to figure out why this call works for about 2500 cases and fails for only 1…</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2564579,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-12-17T08:10:03.337000",
              "content": "<p>If you set Image.MAX_IMAGE_PIXELS = None, when loading images, there is no upper limit on the number of pixels.</p>\n<p>Therefore, when loading a large image, it will crash due to excessive memory usage.</p>\n<p>If you set Image.MAX_IMAGE_PIXELS = 7000 * 7000 like in my code, when a large image comes in the image loading process of the second code, it will not be loaded and an exception will occur.</p>\n<p>All that's left to do is handle the exception.</p>\n<p>I hope you find this information useful.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2565390,
              "author_name": "Patrick Robitaille",
              "author_url": "",
              "post_date": "2023-12-18T02:24:42.020000",
              "content": "<p>As mentioned above, the WSI that causes the problem is of median-to-average size. Setting <code>Image.MAX_IMAGE_PIXELS = None</code> in the context of this competition is perfectly safe. I also mentioned previously that I ran the inference code on the full training set and on the WSIs in the test set in descending order of size; for the latter case, the notebook threw the exception after around 4.5 hours of execution. If excessive RAM usage was the issue, the code would have crashed right at the start. </p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2563500": "I have spent a lot of time in the last week to troubleshoot \"in the dark\" my inference code in order to complete successfully a full submission. Unfortunately, I keep getting the frustrating and non-informative \"Notebook Threw Exception\" result.\n\nI have tested my inference code by running it using the full training set, and it completes without any exceptions raised. When faced with these repeated 'Notebook Threw Exception\" results, I bubble-wrapped my code with try-except statements at various locations of possible failure and also ensured the following:\n\n- That padding is applied when cropping patches that might expand outside the boundaries of the TMA/WSI (probably more important for TMAs than WSIs)\n- That the number of patches from a WSI is limited to a certain maximum to avoid out-of-memory problems with the GPU. While I managed to have this maximum at about 700 on my local PC (RTX 4080), I found that it needed to be below 500 on the Kaggle GPU. I have tried to set this at 450, then 400, but I still get the same result.\n- That I have mechanisms in place at various locations in my code to ensure the rejection of patches that are too uniform in color. Only one of these among selected patches from a WSI can cause a model to protest fairly quickly.\n\nAfter much probing, I have been able to conclude that the sole culprit for the \"Notebook Threw Exception\" results is one single WSI with the following characteristics (after separating the test images between WSIs and TMAs, processing the WSIs first):\n\n- If I process the WSIs in ascending order of size (width x height), the notebook returns the exception at about 3.5 hours of execution; in descending order of size, it crashes at about 4.5 hours. So, the WSI is of median-to-average size, which goes against the out-of-memory possibility.\n- If I process the WSIs in ascending image_id, the notebook crashes after 82 minutes (the code is still running for descending order, but I would think that it would be around 6.5 hours). So, the image_id is probably 1XXXX, with the first X possibly in the low digits.\n\nI'm really at a loss and short of ideas to fix this problem, especially when I have to do this \"in the dark\" with no other information than a stupid \"Notebook Threw Exception\" message. Does anybody have any suggestions on what I could try next? Has anybody come across this type of situation?\n\n(Sidenote: when s**t happens, it happens a lot; my PC is on life-support since Monday, most likely due to a faulty motherboard; obviously, all my files for this competition were on the hard drive, yet to be backed up, facing the likelihood of irremediable oblivion; I'm anxiously waiting for a call from the tech heads looking into it. So, any help from you guys would be REALLY appreciated)",
    "2565756": "A direct approach maybe: You have successful submissions, try to replicate the changes you have made step by step (from working submission to the failing).\n\nLet me also share an indirect approach I had to take for a similar experience:\nIn my case I wasn't even getting any exception, where the problem was somewhere within parallel processing script so the exception was stopping that specific process but the notebook was continuing with the subsequent cells/code, which for me meant a zero score (because rest of the code was actually good enough to open and save sample_submission without the labels as 'submission.csv'.)\n\nAll experiments with train set (full trial, wsi only, tma only, sorted, reversesorted) were running OK.\n\nThe problem appeared to be the first stage model that I was using to create a segmentation mask for the second stage classifier to predict. With train set it was OK, because this baseline segmentator was doing fine to segment variants of the WSI's from the same data distribution as it was trained. \n\nBut obviously it was having problems with the test set, so I made sure there is a fallback option if the output of the segmentation processing did not result in required input for the next stage.\n\nBe careful I am not saying the segmentation model was not doing its thing but maybe the cropareas calculation afterwards was the problem such as having 0 width/height rectangles. \n\nI hope you can find it sooner than later, good luck!",
    "2563663": "Is it possible that some WSI has 4 channels, since it's saved with `.png` format? I have never tested, just a guess. ",
    "2568309": "I got some problems on very few WSI also… what I did is just randomly assign a label (with the distribution of the training set) to problematic images. At least the notebooks run, and the impact on the metric is minimal ",
    "2563595": "I understand your frustration, a few other also had similar problem, \nwhat I suggested is that, take your model, and run it on whole training set, exactly like how you run your submission notebook, but do it on train.csv and train images, don't submit it, but just run it to monitor any error\n\nthis could potentially help",
    "2563511": "I should also mention that I wrapped my model call into a try-except statement:\n\n`try:\n            preds = model(X)\n        except:\n(do something else)\n`\n\nbut it doesn't seem to catch any exception. So, it seems that it could still be a GPU memory-related problem (that would be a RuntimeError) and/or an error that would cause a system exit.",
    "2564344": "This solved post may be helpful.\n\nhttps://www.kaggle.com/competitions/UBC-OCEAN/discussion/461341\n\nTry setting the submission notebook batch size to 1.\nIn my case, \"Notebook Throw Exception\" occurred when the batch size was 2 or more.\n\nJust in case, please try setting num_workers of DataLoader to 1 as well.\n\nIf you are using PIL like me, it worked if you set Image.MAX_IMAGE_PIXELS as below and then changed the image loading code as below.\n(This part needs to be modified to suit your code)\n\n```python\nfrom PIL import Image\nImage.MAX_IMAGE_PIXELS = 7000 * 7000\n```\n\n```python\nclass UBCDatasetInfer(Dataset):\n     def __getitem__(self, idx: int) -> tuple:\n         img_path = self.df.iloc[idx][\"path\"]\n        \n         try:\n             tile = Image.open(img_path)\n         except:\n             # thumb_path = self.df.iloc[idx][\"thumb_path\"]\n             # tile = Image.open(thumb_path)\n             tile = Image.fromarray(np.zeros((1000, 1000, 3)).astype(np.uint8))\n```\n\nPlease refer to the URL below for the entire code.\n\nhttps://www.kaggle.com/code/yamitomo/j010-minimum-submit-ok"
  }
}