{
  "id": 463261,
  "title": "How to deal with \"Notebook Threw Exception\"?",
  "url": "/competitions/UBC-OCEAN/discussion/463261",
  "author_name": "Orzlala",
  "post_date": "2023-12-24T09:35:45.512000",
  "votes": 3,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I've submitted three different pieces of inference code using GPU100 today, all of which successfully ran offline, but all of which Threw the same error \"Notebook Threw Exception\" after being submitted to the background.<br>\nAfter the first error I found that the image_id in the submission.csv submission was of type object, so I converted it to int64.<br>\nI removed joblib's parallel computation after the second error, but also failed to commit.<br>\nDuring this period, I found that each time an error occurred in the background, it took the same amount of time as it took to run all the code locally. Does this mean that the error did not occur during the prediction process of the model, but in the calculation of the score? (I'm not sure if the background uses it directly after getting the submission.csv to calculate the score).<br>\nTherefore, I use the official evaluation metrics in the inference code to evaluate the output of the model, but the inference code can run successfully on the notebook. This makes me not know how to deal with this mistake.😭</p>\n<hr>\n<p>I have successfully solved this problem by progressively splitting the redundant inference code to find the root cause. An implicit error about the index of the Dataset causes the data to fail to load into the DataFrame with multiple test data.</p>",
  "messages": [
    {
      "id": 2572481,
      "postDate": "2023-12-24T09:35:45.513Z",
      "content": "<p>I've submitted three different pieces of inference code using GPU100 today, all of which successfully ran offline, but all of which Threw the same error \"Notebook Threw Exception\" after being submitted to the background.<br>\nAfter the first error I found that the image_id in the submission.csv submission was of type object, so I converted it to int64.<br>\nI removed joblib's parallel computation after the second error, but also failed to commit.<br>\nDuring this period, I found that each time an error occurred in the background, it took the same amount of time as it took to run all the code locally. Does this mean that the error did not occur during the prediction process of the model, but in the calculation of the score? (I'm not sure if the background uses it directly after getting the submission.csv to calculate the score).<br>\nTherefore, I use the official evaluation metrics in the inference code to evaluate the output of the model, but the inference code can run successfully on the notebook. This makes me not know how to deal with this mistake.😭</p>\n<hr>\n<p>I have successfully solved this problem by progressively splitting the redundant inference code to find the root cause. An implicit error about the index of the Dataset causes the data to fail to load into the DataFrame with multiple test data.</p>",
      "rawMarkdown": "I've submitted three different pieces of inference code using GPU100 today, all of which successfully ran offline, but all of which Threw the same error \"Notebook Threw Exception\" after being submitted to the background.\nAfter the first error I found that the image_id in the submission.csv submission was of type object, so I converted it to int64.\nI removed joblib's parallel computation after the second error, but also failed to commit.\nDuring this period, I found that each time an error occurred in the background, it took the same amount of time as it took to run all the code locally. Does this mean that the error did not occur during the prediction process of the model, but in the calculation of the score? (I'm not sure if the background uses it directly after getting the submission.csv to calculate the score).\nTherefore, I use the official evaluation metrics in the inference code to evaluate the output of the model, but the inference code can run successfully on the notebook. This makes me not know how to deal with this mistake.😭\n\n---\nI have successfully solved this problem by progressively splitting the redundant inference code to find the root cause. An implicit error about the index of the Dataset causes the data to fail to load into the DataFrame with multiple test data.",
      "votes": 3
    },
    {
      "id": 2578086,
      "postDate": "2023-12-29T03:30:09.520Z",
      "content": "<p>Hi,<br>\nI have a problem that sounds similar. <br>\nHow did you technically solve it?<br>\nThank</p>",
      "rawMarkdown": "Hi,\nI have a problem that sounds similar. \nHow did you technically solve it?\nThank",
      "replies": [
        {
          "id": 2578410,
          "postDate": "2023-12-29T08:44:06.123Z",
          "content": "<p>I think the reason for such problems in most cases is often some kind of problem in the code written by myself. The problem I had at that time was an error in reading data about the Dataset. My solution was to construct different forms of data as the input of the Dataset for testing. Then I managed to figure out that there was an error. So I think you can take a similar approach to testing your code to find problems and make your code more robust. Hope you can resolve this issue soon!🥰</p>",
          "rawMarkdown": "I think the reason for such problems in most cases is often some kind of problem in the code written by myself. The problem I had at that time was an error in reading data about the Dataset. My solution was to construct different forms of data as the input of the Dataset for testing. Then I managed to figure out that there was an error. So I think you can take a similar approach to testing your code to find problems and make your code more robust. Hope you can resolve this issue soon!🥰",
          "replies": [
            {
              "id": 2581529,
              "postDate": "2023-12-31T19:47:27.500Z",
              "content": "<p>Thank you very much for your answer <br>\nUnfortunately, I didn't find the problem at the moment. In competition environment is very difficult understand where is the problem and the origin.</p>\n<p>Happy new year ;)</p>",
              "rawMarkdown": "Thank you very much for your answer \nUnfortunately, I didn't find the problem at the moment. In competition environment is very difficult understand where is the problem and the origin.\n\n\nHappy new year ;)\n"
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2578086,
      "author_name": "GuidoRo",
      "author_url": "",
      "post_date": "2023-12-29T03:30:09.520000",
      "content": "<p>Hi,<br>\nI have a problem that sounds similar. <br>\nHow did you technically solve it?<br>\nThank</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2578410,
          "author_name": "Orzlala",
          "author_url": "",
          "post_date": "2023-12-29T08:44:06.123000",
          "content": "<p>I think the reason for such problems in most cases is often some kind of problem in the code written by myself. The problem I had at that time was an error in reading data about the Dataset. My solution was to construct different forms of data as the input of the Dataset for testing. Then I managed to figure out that there was an error. So I think you can take a similar approach to testing your code to find problems and make your code more robust. Hope you can resolve this issue soon!🥰</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2581529,
              "author_name": "GuidoRo",
              "author_url": "",
              "post_date": "2023-12-31T19:47:27.500000",
              "content": "<p>Thank you very much for your answer <br>\nUnfortunately, I didn't find the problem at the moment. In competition environment is very difficult understand where is the problem and the origin.</p>\n<p>Happy new year ;)</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2572481": "I've submitted three different pieces of inference code using GPU100 today, all of which successfully ran offline, but all of which Threw the same error \"Notebook Threw Exception\" after being submitted to the background.\nAfter the first error I found that the image_id in the submission.csv submission was of type object, so I converted it to int64.\nI removed joblib's parallel computation after the second error, but also failed to commit.\nDuring this period, I found that each time an error occurred in the background, it took the same amount of time as it took to run all the code locally. Does this mean that the error did not occur during the prediction process of the model, but in the calculation of the score? (I'm not sure if the background uses it directly after getting the submission.csv to calculate the score).\nTherefore, I use the official evaluation metrics in the inference code to evaluate the output of the model, but the inference code can run successfully on the notebook. This makes me not know how to deal with this mistake.😭\n\n---\nI have successfully solved this problem by progressively splitting the redundant inference code to find the root cause. An implicit error about the index of the Dataset causes the data to fail to load into the DataFrame with multiple test data.",
    "2578086": "Hi,\nI have a problem that sounds similar. \nHow did you technically solve it?\nThank"
  }
}