{
  "id": 186355,
  "title": "Submission timeout",
  "url": "/competitions/rsna-str-pulmonary-embolism-detection/discussion/186355",
  "author_name": "Marco Stefani",
  "post_date": "2020-09-24T07:51:01.856000",
  "votes": 2,
  "comment_count": 5,
  "views": 0,
  "content": "<p>My notebook took 1 hour to commit. The full test set is about 4 times the public test set. It should take about 4 hours to submit but it keeps going in timeout, meaning it takes more than 9 hours. Why? Am I missing something?</p>\n<p>Update: I was using multithreading, the workers silently failed because of a pydicom error while reading .dcm files, and the main thread was stuck waiting the end of the workers. I don't know exactly how that was possible but doing all the computations on the main thread made the notebook finish and threw an exception. Then, installing gdcm resolved the pydicom reading exception. Only the private test set contains images not readable without gdcm</p>",
  "messages": [
    {
      "id": 1025324,
      "postDate": "2020-09-24T13:32:24.113Z",
      "content": "<p><a href=\"https://www.kaggle.com/neuronffh\" target=\"_blank\">@neuronffh</a>  I head the same issue. I found out the problem is the memory limitation. In my pipeline I had a loop which loaded all the dicom metadata into a Pandas DataFrame. This loop consumed a lot of memory which caused the notebook to use the disk swap space which really slowed the script and it timed - out. After improving the loop and saving only the necessary data, the notebook finished after the expected time. </p>",
      "rawMarkdown": "@neuronffh  I head the same issue. I found out the problem is the memory limitation. In my pipeline I had a loop which loaded all the dicom metadata into a Pandas DataFrame. This loop consumed a lot of memory which caused the notebook to use the disk swap space which really slowed the script and it timed - out. After improving the loop and saving only the necessary data, the notebook finished after the expected time. ",
      "votes": 1,
      "replies": [
        {
          "id": 1025692,
          "postDate": "2020-09-24T17:56:22.650Z",
          "content": "<p>Thank you a lot for your suggestion!<br>\nI'm now looking at my code to see how it manages memory and whether ram is a problem or not</p>",
          "rawMarkdown": "Thank you a lot for your suggestion!\nI'm now looking at my code to see how it manages memory and whether ram is a problem or not"
        }
      ]
    },
    {
      "id": 1024902,
      "postDate": "2020-09-24T07:51:01.857Z",
      "content": "<p>My notebook took 1 hour to commit. The full test set is about 4 times the public test set. It should take about 4 hours to submit but it keeps going in timeout, meaning it takes more than 9 hours. Why? Am I missing something?</p>\n<p>Update: I was using multithreading, the workers silently failed because of a pydicom error while reading .dcm files, and the main thread was stuck waiting the end of the workers. I don't know exactly how that was possible but doing all the computations on the main thread made the notebook finish and threw an exception. Then, installing gdcm resolved the pydicom reading exception. Only the private test set contains images not readable without gdcm</p>",
      "rawMarkdown": "My notebook took 1 hour to commit. The full test set is about 4 times the public test set. It should take about 4 hours to submit but it keeps going in timeout, meaning it takes more than 9 hours. Why? Am I missing something?\n\nUpdate: I was using multithreading, the workers silently failed because of a pydicom error while reading .dcm files, and the main thread was stuck waiting the end of the workers. I don't know exactly how that was possible but doing all the computations on the main thread made the notebook finish and threw an exception. Then, installing gdcm resolved the pydicom reading exception. Only the private test set contains images not readable without gdcm\n",
      "votes": 2
    },
    {
      "id": 1025250,
      "postDate": "2020-09-24T12:52:07.997Z",
      "content": "<p>I have exactly the same problem and I genuinely don't know how to solve it…<br>\nMade some changes this morning to streamline it as much as possible, submitted again, this time (after a still ridiculous amount of time spent processing it) I get a \"submission scoring error\" - \"no additional details provided for this error\".<br>\nAnybody know something about that?</p>",
      "rawMarkdown": "I have exactly the same problem and I genuinely don't know how to solve it...\nMade some changes this morning to streamline it as much as possible, submitted again, this time (after a still ridiculous amount of time spent processing it) I get a \"submission scoring error\" - \"no additional details provided for this error\".\nAnybody know something about that?",
      "replies": [
        {
          "id": 1025332,
          "postDate": "2020-09-24T13:36:44.030Z",
          "content": "<p>Look <a href=\"https://www.kaggle.com/code-competition-debugging\" target=\"_blank\">here</a> \"submission scoring error\" means your submission file is not in the correct format:</p>\n<ul>\n<li>You may be missing the header</li>\n<li>Maybe you don't have all the required id's</li>\n</ul>\n<p>You might want to compare your submission file to sample_submission.csv</p>",
          "rawMarkdown": "Look [here](https://www.kaggle.com/code-competition-debugging) \"submission scoring error\" means your submission file is not in the correct format:\n* You may be missing the header\n* Maybe you don't have all the required id's\n\nYou might want to compare your submission file to sample_submission.csv"
        },
        {
          "id": 1025424,
          "postDate": "2020-09-24T14:41:07.110Z",
          "content": "<p>Yeah I think I'll just resort to loading the sample submission file and replacing the values, instead of writing my own submission file… thanks!</p>",
          "rawMarkdown": "Yeah I think I'll just resort to loading the sample submission file and replacing the values, instead of writing my own submission file... thanks!"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1025324,
      "author_name": "yuval reina",
      "author_url": "",
      "post_date": "2020-09-24T13:32:24.113000",
      "content": "<p><a href=\"https://www.kaggle.com/neuronffh\" target=\"_blank\">@neuronffh</a>  I head the same issue. I found out the problem is the memory limitation. In my pipeline I had a loop which loaded all the dicom metadata into a Pandas DataFrame. This loop consumed a lot of memory which caused the notebook to use the disk swap space which really slowed the script and it timed - out. After improving the loop and saving only the necessary data, the notebook finished after the expected time. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1025692,
          "author_name": "Marco Stefani",
          "author_url": "",
          "post_date": "2020-09-24T17:56:22.650000",
          "content": "<p>Thank you a lot for your suggestion!<br>\nI'm now looking at my code to see how it manages memory and whether ram is a problem or not</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1025250,
      "author_name": "Alex Bader",
      "author_url": "",
      "post_date": "2020-09-24T12:52:07.997000",
      "content": "<p>I have exactly the same problem and I genuinely don't know how to solve it…<br>\nMade some changes this morning to streamline it as much as possible, submitted again, this time (after a still ridiculous amount of time spent processing it) I get a \"submission scoring error\" - \"no additional details provided for this error\".<br>\nAnybody know something about that?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1025332,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2020-09-24T13:36:44.030000",
          "content": "<p>Look <a href=\"https://www.kaggle.com/code-competition-debugging\" target=\"_blank\">here</a> \"submission scoring error\" means your submission file is not in the correct format:</p>\n<ul>\n<li>You may be missing the header</li>\n<li>Maybe you don't have all the required id's</li>\n</ul>\n<p>You might want to compare your submission file to sample_submission.csv</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1025424,
          "author_name": "Alex Bader",
          "author_url": "",
          "post_date": "2020-09-24T14:41:07.110000",
          "content": "<p>Yeah I think I'll just resort to loading the sample submission file and replacing the values, instead of writing my own submission file… thanks!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1025324": "@neuronffh  I head the same issue. I found out the problem is the memory limitation. In my pipeline I had a loop which loaded all the dicom metadata into a Pandas DataFrame. This loop consumed a lot of memory which caused the notebook to use the disk swap space which really slowed the script and it timed - out. After improving the loop and saving only the necessary data, the notebook finished after the expected time. ",
    "1024902": "My notebook took 1 hour to commit. The full test set is about 4 times the public test set. It should take about 4 hours to submit but it keeps going in timeout, meaning it takes more than 9 hours. Why? Am I missing something?\n\nUpdate: I was using multithreading, the workers silently failed because of a pydicom error while reading .dcm files, and the main thread was stuck waiting the end of the workers. I don't know exactly how that was possible but doing all the computations on the main thread made the notebook finish and threw an exception. Then, installing gdcm resolved the pydicom reading exception. Only the private test set contains images not readable without gdcm\n",
    "1025250": "I have exactly the same problem and I genuinely don't know how to solve it...\nMade some changes this morning to streamline it as much as possible, submitted again, this time (after a still ridiculous amount of time spent processing it) I get a \"submission scoring error\" - \"no additional details provided for this error\".\nAnybody know something about that?"
  }
}