{
  "id": 464324,
  "title": "What are the most common reasons for submission fail after successful notebook run and how to inspect them?",
  "url": "/competitions/UBC-OCEAN/discussion/464324",
  "author_name": "Khasia",
  "post_date": "2023-12-29T20:56:22.688000",
  "votes": 5,
  "comment_count": 5,
  "views": 0,
  "content": "<p>It seams that logs we see on kaggle notebook belong to the first run of the notebook and it looks like the scoring process starts only after this run. </p>\n<p>This is why I am trying to understand what are the most common reasons for submission fail after successful notebook run and how to inspect them. </p>\n<p>My test submission was successful but when I try to submit my finalized model it fails: <br>\n      - My finalized model uses less time for inference and it seems the timeout should not be a problem.<br>\n      - I changed batch size from 8 to 30</p>\n<p>Please help ✊ It's my first competition, worked a lot on this and hope to get to the point when I can submit it! 🤗</p>\n<p>** Update:<br>\nUsed train dataset instead of test dataset for identifying a problem. One problem was varying batch sizes when last batch was loaded. A single test image based data did not give me possibility to see that problem. So, one good idea is to use train dataset for troubleshooting.</p>\n<p>What else can I try?</p>",
  "messages": [
    {
      "id": 2579178,
      "postDate": "2023-12-29T20:56:22.690Z",
      "content": "<p>It seams that logs we see on kaggle notebook belong to the first run of the notebook and it looks like the scoring process starts only after this run. </p>\n<p>This is why I am trying to understand what are the most common reasons for submission fail after successful notebook run and how to inspect them. </p>\n<p>My test submission was successful but when I try to submit my finalized model it fails: <br>\n      - My finalized model uses less time for inference and it seems the timeout should not be a problem.<br>\n      - I changed batch size from 8 to 30</p>\n<p>Please help ✊ It's my first competition, worked a lot on this and hope to get to the point when I can submit it! 🤗</p>\n<p>** Update:<br>\nUsed train dataset instead of test dataset for identifying a problem. One problem was varying batch sizes when last batch was loaded. A single test image based data did not give me possibility to see that problem. So, one good idea is to use train dataset for troubleshooting.</p>\n<p>What else can I try?</p>",
      "rawMarkdown": "It seams that logs we see on kaggle notebook belong to the first run of the notebook and it looks like the scoring process starts only after this run. \n\nThis is why I am trying to understand what are the most common reasons for submission fail after successful notebook run and how to inspect them. \n\nMy test submission was successful but when I try to submit my finalized model it fails: \n      - My finalized model uses less time for inference and it seems the timeout should not be a problem.\n      - I changed batch size from 8 to 30\n\n Please help ✊ It's my first competition, worked a lot on this and hope to get to the point when I can submit it! 🤗\n\n** Update:\nUsed train dataset instead of test dataset for identifying a problem. One problem was varying batch sizes when last batch was loaded. A single test image based data did not give me possibility to see that problem. So, one good idea is to use train dataset for troubleshooting.\n\nWhat else can I try?",
      "votes": 5
    },
    {
      "id": 2586061,
      "postDate": "2024-01-04T00:18:15.757Z",
      "content": "<p>I tried your tip, pretending the training data is test data by deleting the label, and executed my notebook as if it was in test mode. The run was successful, but still it threw submission error \"Notebook Threw Exception\".</p>\n<p>As there are many different types of error possible, I surrounded my code with try-catch blocks, where in the catch block I <em>intentionally</em> trigger another known error. I managed to narrow down the cause to a few lines of numpy statements (vstack the prediction batches, argmax the prediction, label_encoder.inverse_transform it back to text), that were identical in all my submissions (and the others were successful). The output is always the same shape regardless of my models, and all inputs are \"protected\" by nan_to_num wrappers. Meaning I still have no idea what went wrong and I'm stuck at 7XX-th place.</p>",
      "rawMarkdown": "I tried your tip, pretending the training data is test data by deleting the label, and executed my notebook as if it was in test mode. The run was successful, but still it threw submission error \"Notebook Threw Exception\".\n\nAs there are many different types of error possible, I surrounded my code with try-catch blocks, where in the catch block I *intentionally* trigger another known error. I managed to narrow down the cause to a few lines of numpy statements (vstack the prediction batches, argmax the prediction, label_encoder.inverse_transform it back to text), that were identical in all my submissions (and the others were successful). The output is always the same shape regardless of my models, and all inputs are \"protected\" by nan_to_num wrappers. Meaning I still have no idea what went wrong and I'm stuck at 7XX-th place.",
      "votes": 1,
      "replies": [
        {
          "id": 2587226,
          "postDate": "2024-01-04T16:15:36.197Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 2587267,
          "postDate": "2024-01-04T16:40:04.657Z",
          "content": "<p>I got also bad ranking :))) let's team up next time! </p>",
          "rawMarkdown": "I got also bad ranking :))) let's team up next time! "
        }
      ]
    },
    {
      "id": 2584058,
      "postDate": "2024-01-02T16:44:37.440Z",
      "content": "<p>When you say batch size do you mean the number of WSIs you are processing at once?  If that is the case I recommend using a batch size of 1 because the WSIs can be very large.  Other than that I recommend running the training set as if it were the test set. The test set here is about 200GB smaller than the training set so if you can run the training set in time then the test set will complete.  </p>\n<p>Next, what is the message it is giving you when it fails to submit?  Out of memory, timeout, kaggle error, etc.?</p>",
      "rawMarkdown": "When you say batch size do you mean the number of WSIs you are processing at once?  If that is the case I recommend using a batch size of 1 because the WSIs can be very large.  Other than that I recommend running the training set as if it were the test set. The test set here is about 200GB smaller than the training set so if you can run the training set in time then the test set will complete.  \n\nNext, what is the message it is giving you when it fails to submit?  Out of memory, timeout, kaggle error, etc.?",
      "replies": [
        {
          "id": 2584825,
          "postDate": "2024-01-03T07:51:32.250Z",
          "content": "<p>Thanks Connor!<br>\nI already have run train set like test set and exactly that helped me in finding batch related problem mentioned above.<br>\nI am accessing patches of one WSI one at a time (so that patches of same one image make a batch).<br>\nI did not notice kaggle related errors you mentioned, but may be that is exactly what I have to find and look at.</p>",
          "rawMarkdown": "Thanks Connor!\nI already have run train set like test set and exactly that helped me in finding batch related problem mentioned above.\nI am accessing patches of one WSI one at a time (so that patches of same one image make a batch).\nI did not notice kaggle related errors you mentioned, but may be that is exactly what I have to find and look at."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2586061,
      "author_name": "tinkei",
      "author_url": "",
      "post_date": "2024-01-04T00:18:15.757000",
      "content": "<p>I tried your tip, pretending the training data is test data by deleting the label, and executed my notebook as if it was in test mode. The run was successful, but still it threw submission error \"Notebook Threw Exception\".</p>\n<p>As there are many different types of error possible, I surrounded my code with try-catch blocks, where in the catch block I <em>intentionally</em> trigger another known error. I managed to narrow down the cause to a few lines of numpy statements (vstack the prediction batches, argmax the prediction, label_encoder.inverse_transform it back to text), that were identical in all my submissions (and the others were successful). The output is always the same shape regardless of my models, and all inputs are \"protected\" by nan_to_num wrappers. Meaning I still have no idea what went wrong and I'm stuck at 7XX-th place.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2587226,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-01-04T16:15:36.197000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2587267,
          "author_name": "Khasia",
          "author_url": "",
          "post_date": "2024-01-04T16:40:04.657000",
          "content": "<p>I got also bad ranking :))) let's team up next time! </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2584058,
      "author_name": "Connor",
      "author_url": "",
      "post_date": "2024-01-02T16:44:37.440000",
      "content": "<p>When you say batch size do you mean the number of WSIs you are processing at once?  If that is the case I recommend using a batch size of 1 because the WSIs can be very large.  Other than that I recommend running the training set as if it were the test set. The test set here is about 200GB smaller than the training set so if you can run the training set in time then the test set will complete.  </p>\n<p>Next, what is the message it is giving you when it fails to submit?  Out of memory, timeout, kaggle error, etc.?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2584825,
          "author_name": "Khasia",
          "author_url": "",
          "post_date": "2024-01-03T07:51:32.250000",
          "content": "<p>Thanks Connor!<br>\nI already have run train set like test set and exactly that helped me in finding batch related problem mentioned above.<br>\nI am accessing patches of one WSI one at a time (so that patches of same one image make a batch).<br>\nI did not notice kaggle related errors you mentioned, but may be that is exactly what I have to find and look at.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2579178": "It seams that logs we see on kaggle notebook belong to the first run of the notebook and it looks like the scoring process starts only after this run. \n\nThis is why I am trying to understand what are the most common reasons for submission fail after successful notebook run and how to inspect them. \n\nMy test submission was successful but when I try to submit my finalized model it fails: \n      - My finalized model uses less time for inference and it seems the timeout should not be a problem.\n      - I changed batch size from 8 to 30\n\n Please help ✊ It's my first competition, worked a lot on this and hope to get to the point when I can submit it! 🤗\n\n** Update:\nUsed train dataset instead of test dataset for identifying a problem. One problem was varying batch sizes when last batch was loaded. A single test image based data did not give me possibility to see that problem. So, one good idea is to use train dataset for troubleshooting.\n\nWhat else can I try?",
    "2586061": "I tried your tip, pretending the training data is test data by deleting the label, and executed my notebook as if it was in test mode. The run was successful, but still it threw submission error \"Notebook Threw Exception\".\n\nAs there are many different types of error possible, I surrounded my code with try-catch blocks, where in the catch block I *intentionally* trigger another known error. I managed to narrow down the cause to a few lines of numpy statements (vstack the prediction batches, argmax the prediction, label_encoder.inverse_transform it back to text), that were identical in all my submissions (and the others were successful). The output is always the same shape regardless of my models, and all inputs are \"protected\" by nan_to_num wrappers. Meaning I still have no idea what went wrong and I'm stuck at 7XX-th place.",
    "2584058": "When you say batch size do you mean the number of WSIs you are processing at once?  If that is the case I recommend using a batch size of 1 because the WSIs can be very large.  Other than that I recommend running the training set as if it were the test set. The test set here is about 200GB smaller than the training set so if you can run the training set in time then the test set will complete.  \n\nNext, what is the message it is giving you when it fails to submit?  Out of memory, timeout, kaggle error, etc.?"
  }
}