{
  "id": 465598,
  "title": "\"Notebook Out of Disk\" submission error",
  "url": "/competitions/UBC-OCEAN/discussion/465598",
  "author_name": "tinkei",
  "post_date": "2024-01-04T20:17:01.374000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Now that the competition is over, would the organizers kindly share the cause of my notebook submission errors?</p>\n<p>My final submission has a \"Notebook Out of Disk\" error, which might sound obvious to everyone, except that there is a grand total of zero result when you search this exact error message on Google. Previously it was getting a generic \"Notebook Threw Exception\".</p>\n<p>My model trains and predicts only using thumbnails, rather unlikely that it would use up disk storage. The model is an FCN with a MobileNetV3 backbone which has a tiny 11M parameter count. Model weights and pretrained weights are hosted in a Dataset. Shared codes common to all other (successful) submissions are hosted in another Dataset. With a 3:7 train-validation split (not a typo, because it trains purely on segmentation masks that everybody claims to be useless), my model achieves a validation F1 macro of 56% (simple average over predicted class per pixel).</p>\n<p>By intentionally writing empty strings in submission.csv whenever I catch an exception, which would trigger a \"Submission Scoring Error\", I narrowed down the cause to a bunch of stack, argmax, inverse_transform code in the prediction step that is identical across all my (successful) submissions. I varied the batch size from 1 sample (to catch BatchNorm / drop_last errors) to 4x the final submission batch size (so it's not OoM) by predicting the training set and everything worked on Kaggle Notebook, right up till submission, and after it had been running for almost 2 hours. My while-loops where I sample for quality tiles from the thumbnail (i.e. no black background) are protected by a max-iter of 10 retries or 4x batch size. I am frankly tired and frustrated from all these hidden errors.</p>",
  "messages": [
    {
      "id": 2587541,
      "postDate": "2024-01-04T20:17:01.373Z",
      "content": "<p>Now that the competition is over, would the organizers kindly share the cause of my notebook submission errors?</p>\n<p>My final submission has a \"Notebook Out of Disk\" error, which might sound obvious to everyone, except that there is a grand total of zero result when you search this exact error message on Google. Previously it was getting a generic \"Notebook Threw Exception\".</p>\n<p>My model trains and predicts only using thumbnails, rather unlikely that it would use up disk storage. The model is an FCN with a MobileNetV3 backbone which has a tiny 11M parameter count. Model weights and pretrained weights are hosted in a Dataset. Shared codes common to all other (successful) submissions are hosted in another Dataset. With a 3:7 train-validation split (not a typo, because it trains purely on segmentation masks that everybody claims to be useless), my model achieves a validation F1 macro of 56% (simple average over predicted class per pixel).</p>\n<p>By intentionally writing empty strings in submission.csv whenever I catch an exception, which would trigger a \"Submission Scoring Error\", I narrowed down the cause to a bunch of stack, argmax, inverse_transform code in the prediction step that is identical across all my (successful) submissions. I varied the batch size from 1 sample (to catch BatchNorm / drop_last errors) to 4x the final submission batch size (so it's not OoM) by predicting the training set and everything worked on Kaggle Notebook, right up till submission, and after it had been running for almost 2 hours. My while-loops where I sample for quality tiles from the thumbnail (i.e. no black background) are protected by a max-iter of 10 retries or 4x batch size. I am frankly tired and frustrated from all these hidden errors.</p>",
      "rawMarkdown": "Now that the competition is over, would the organizers kindly share the cause of my notebook submission errors?\n\nMy final submission has a \"Notebook Out of Disk\" error, which might sound obvious to everyone, except that there is a grand total of zero result when you search this exact error message on Google. Previously it was getting a generic \"Notebook Threw Exception\".\n\nMy model trains and predicts only using thumbnails, rather unlikely that it would use up disk storage. The model is an FCN with a MobileNetV3 backbone which has a tiny 11M parameter count. Model weights and pretrained weights are hosted in a Dataset. Shared codes common to all other (successful) submissions are hosted in another Dataset. With a 3:7 train-validation split (not a typo, because it trains purely on segmentation masks that everybody claims to be useless), my model achieves a validation F1 macro of 56% (simple average over predicted class per pixel).\n\nBy intentionally writing empty strings in submission.csv whenever I catch an exception, which would trigger a \"Submission Scoring Error\", I narrowed down the cause to a bunch of stack, argmax, inverse_transform code in the prediction step that is identical across all my (successful) submissions. I varied the batch size from 1 sample (to catch BatchNorm / drop_last errors) to 4x the final submission batch size (so it's not OoM) by predicting the training set and everything worked on Kaggle Notebook, right up till submission, and after it had been running for almost 2 hours. My while-loops where I sample for quality tiles from the thumbnail (i.e. no black background) are protected by a max-iter of 10 retries or 4x batch size. I am frankly tired and frustrated from all these hidden errors.",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2587541": "Now that the competition is over, would the organizers kindly share the cause of my notebook submission errors?\n\nMy final submission has a \"Notebook Out of Disk\" error, which might sound obvious to everyone, except that there is a grand total of zero result when you search this exact error message on Google. Previously it was getting a generic \"Notebook Threw Exception\".\n\nMy model trains and predicts only using thumbnails, rather unlikely that it would use up disk storage. The model is an FCN with a MobileNetV3 backbone which has a tiny 11M parameter count. Model weights and pretrained weights are hosted in a Dataset. Shared codes common to all other (successful) submissions are hosted in another Dataset. With a 3:7 train-validation split (not a typo, because it trains purely on segmentation masks that everybody claims to be useless), my model achieves a validation F1 macro of 56% (simple average over predicted class per pixel).\n\nBy intentionally writing empty strings in submission.csv whenever I catch an exception, which would trigger a \"Submission Scoring Error\", I narrowed down the cause to a bunch of stack, argmax, inverse_transform code in the prediction step that is identical across all my (successful) submissions. I varied the batch size from 1 sample (to catch BatchNorm / drop_last errors) to 4x the final submission batch size (so it's not OoM) by predicting the training set and everything worked on Kaggle Notebook, right up till submission, and after it had been running for almost 2 hours. My while-loops where I sample for quality tiles from the thumbnail (i.e. no black background) are protected by a max-iter of 10 retries or 4x batch size. I am frankly tired and frustrated from all these hidden errors."
  }
}