{
  "id": 185461,
  "title": "submission errors",
  "url": "/competitions/rsna-str-pulmonary-embolism-detection/discussion/185461",
  "author_name": "qzw",
  "post_date": "2020-09-21T01:14:01.465000",
  "votes": 3,
  "comment_count": 11,
  "views": 0,
  "content": "<p>hi all,</p>\n<p>I create a notebook, which can run successfully in the public data set and save submission.csv in the working dir.</p>\n<p>But after submitting to the complete, It aways can't return the right score. The status shows <strong>Submission CSV Not Found</strong>.</p>\n<p>I don't know which step is wrong. And I can't see the debug information. </p>\n<p>Is there someone who can help me to solve this problem?</p>\n<p>Thanks</p>",
  "messages": [
    {
      "id": 1020139,
      "postDate": "2020-09-21T01:14:01.467Z",
      "content": "<p>hi all,</p>\n<p>I create a notebook, which can run successfully in the public data set and save submission.csv in the working dir.</p>\n<p>But after submitting to the complete, It aways can't return the right score. The status shows <strong>Submission CSV Not Found</strong>.</p>\n<p>I don't know which step is wrong. And I can't see the debug information. </p>\n<p>Is there someone who can help me to solve this problem?</p>\n<p>Thanks</p>",
      "rawMarkdown": "hi all,\n\nI create a notebook, which can run successfully in the public data set and save submission.csv in the working dir.\n\nBut after submitting to the complete, It aways can't return the right score. The status shows **Submission CSV Not Found**.\n\nI don't know which step is wrong. And I can't see the debug information. \n\nIs there someone who can help me to solve this problem?\n\nThanks",
      "votes": 3
    },
    {
      "id": 1046182,
      "postDate": "2020-10-11T12:21:08.880Z",
      "content": "<p>As a basic sanity test, I have created a notebook to check if submission.csv can be properly generated in the Kaggle runtime.</p>\n<p>Please take a look here and copy as needed - <a href=\"https://www.kaggle.com/abhimahule/practice-submission\" target=\"_blank\">https://www.kaggle.com/abhimahule/practice-submission</a></p>",
      "rawMarkdown": "As a basic sanity test, I have created a notebook to check if submission.csv can be properly generated in the Kaggle runtime.\n\nPlease take a look here and copy as needed - https://www.kaggle.com/abhimahule/practice-submission"
    },
    {
      "id": 1045202,
      "postDate": "2020-10-10T12:35:13.427Z",
      "content": "<p>This is because fake code is used. You use fake code to generate Submission CSV files, but your fake code only runs in the fake environment, and it does not run in the real submission environment. You should make this code run in both environments.</p>",
      "rawMarkdown": "This is because fake code is used. You use fake code to generate Submission CSV files, but your fake code only runs in the fake environment, and it does not run in the real submission environment. You should make this code run in both environments.",
      "replies": [
        {
          "id": 1045252,
          "postDate": "2020-10-10T13:06:00.450Z",
          "content": "<p>can you explain this little bit more ,what do you mean by fake code ??</p>",
          "rawMarkdown": "can you explain this little bit more ,what do you mean by fake code ??"
        }
      ]
    },
    {
      "id": 1037946,
      "postDate": "2020-10-05T12:45:28.977Z",
      "content": "<p>hii <a href=\"https://www.kaggle.com/qzw\" target=\"_blank\">@qzw</a> I am facing same problem, please give any suggestion!</p>",
      "rawMarkdown": "hii @qzw I am facing same problem, please give any suggestion!"
    },
    {
      "id": 1020580,
      "postDate": "2020-09-21T09:23:08.390Z",
      "content": "<p><a href=\"https://www.kaggle.com/teeyee314\" target=\"_blank\">@teeyee314</a> </p>\n<p>Thanks,<br>\nI think so.</p>\n<p>But I am still confused about my submission.<br>\nJust as we thought, we didn't need to change anything if our notebook can run correctly.<br>\nHowever, my notebook can run correctly before submitting it to complete. <br>\nWhy they can't generate submission.csv in the private dataset?</p>\n<p>I just use **../input/rsna-str-pulmonary-embolism-detection/test.csv **and my model weights.</p>",
      "rawMarkdown": "@teeyee314 \n\nThanks,\nI think so.\n\nBut I am still confused about my submission.\nJust as we thought, we didn't need to change anything if our notebook can run correctly.\nHowever, my notebook can run correctly before submitting it to complete. \nWhy they can't generate submission.csv in the private dataset?\n\nI just use **../input/rsna-str-pulmonary-embolism-detection/test.csv **and my model weights.\n",
      "replies": [
        {
          "id": 1020634,
          "postDate": "2020-09-21T10:11:45.427Z",
          "content": "<p>create a notebook for inferencing the model, built a pipeline so store result of images present in the test dataset in the sample submission format which is given. Keep in mind to create a file named <code>submission.csv</code> in same format as the sample submission file. when you commit the notebook it will create a file named `submission.csv' in output and the submit option will be available there beside it. Just submit that file it will re-run your notebook on a complete test dataset consisting of private and public test set and will give you your score.</p>",
          "rawMarkdown": "create a notebook for inferencing the model, built a pipeline so store result of images present in the test dataset in the sample submission format which is given. Keep in mind to create a file named `submission.csv` in same format as the sample submission file. when you commit the notebook it will create a file named `submission.csv' in output and the submit option will be available there beside it. Just submit that file it will re-run your notebook on a complete test dataset consisting of private and public test set and will give you your score."
        },
        {
          "id": 1020823,
          "postDate": "2020-09-21T13:03:14.907Z",
          "content": "<p>Thoughts:</p>\n<p>Are you timing out? Does the notebook fail relatively quickly or after 9 hours?<br>\nMake sure you can run your notebook with Internet turned off.<br>\nThe Private Test data is larger than the Public Test data. Are you reading it all into memory at the same time? Any chance you are running out of memory?<br>\nMake sure your submitted file has GPU turned on if you use it for inference.<br>\nDon't assume anything about the private test data<br>\n--I think there is only one SeriesInstanceUID per StudyInstanceUID, but I don't know if we know for sure.<br>\n--DICOM data has studies that don't start at Image number 1. There are gaps in the image numbers (internal Dicom metadata).</p>\n<p>Build your submission file by reading in sample_submission.csv and replacing labels with your calculated values. That way you are guaranteed that all rows are present, even if for some reason your code doesn't calculate a value for a given Study or SOP. Using this method, you can make a submission that only reads 100 lines of the test.csv file. Your score will be poor, but if it succeeds you know your code is structured correctly. Then you are dealing with Memory or Timeout issues. You can also do your testing much faster.</p>\n<p>Other thoughts - sometimes a Batch Size works most of the time, but fails after a while. Lower your Batch Size if that could be the problem.</p>\n<p>How long does Inference on the entire dataset take you? In truth, I don't think I've even tried to run a full inference on the dataset. With pre-processing (which I might not be doing in the most efficient way), I can only inference 80,000 images in 9 hours.</p>\n<p>A final option to at least submit a valid file and get an idea of how you are doing is to run the model without Commit against the public test.csv file (as a bonus, you can use the TPU). Then put that submission in a dataset and write a notebook to read that submission, update the 40% of the private sample_submission.csv that it covers and submit that. Leave default values for the rest of the sample_submission.csv file. That way, at least you can see which direction your score is going and confirm that your submission format is correct.</p>\n<p>Also, look closely at your public submission file. Verify it has all required rows and that there are no NAN anywhere.</p>\n<p>Best of luck,</p>\n<p>-Rich</p>",
          "rawMarkdown": "Thoughts:\n\nAre you timing out? Does the notebook fail relatively quickly or after 9 hours?\nMake sure you can run your notebook with Internet turned off.\nThe Private Test data is larger than the Public Test data. Are you reading it all into memory at the same time? Any chance you are running out of memory?\nMake sure your submitted file has GPU turned on if you use it for inference.\nDon't assume anything about the private test data\n--I think there is only one SeriesInstanceUID per StudyInstanceUID, but I don't know if we know for sure.\n--DICOM data has studies that don't start at Image number 1. There are gaps in the image numbers (internal Dicom metadata).\n\nBuild your submission file by reading in sample_submission.csv and replacing labels with your calculated values. That way you are guaranteed that all rows are present, even if for some reason your code doesn't calculate a value for a given Study or SOP. Using this method, you can make a submission that only reads 100 lines of the test.csv file. Your score will be poor, but if it succeeds you know your code is structured correctly. Then you are dealing with Memory or Timeout issues. You can also do your testing much faster.\n\nOther thoughts - sometimes a Batch Size works most of the time, but fails after a while. Lower your Batch Size if that could be the problem.\n\nHow long does Inference on the entire dataset take you? In truth, I don't think I've even tried to run a full inference on the dataset. With pre-processing (which I might not be doing in the most efficient way), I can only inference 80,000 images in 9 hours.\n\nA final option to at least submit a valid file and get an idea of how you are doing is to run the model without Commit against the public test.csv file (as a bonus, you can use the TPU). Then put that submission in a dataset and write a notebook to read that submission, update the 40% of the private sample_submission.csv that it covers and submit that. Leave default values for the rest of the sample_submission.csv file. That way, at least you can see which direction your score is going and confirm that your submission format is correct.\n\nAlso, look closely at your public submission file. Verify it has all required rows and that there are no NAN anywhere.\n\nBest of luck,\n\n-Rich",
          "votes": 2
        }
      ]
    },
    {
      "id": 1020440,
      "postDate": "2020-09-21T06:47:52.290Z",
      "content": "<p><a href=\"https://www.kaggle.com/teeyee314\" target=\"_blank\">@teeyee314</a> </p>\n<p>hi, thanks for your advice.</p>\n<p>I have a question. <br>\nThe test dataset csv file is ** ../input/rsna-str-pulmonary-embolism-detection/test.csv**.</p>\n<p>Is that mean we just write this path and read the <strong>test.csv</strong> when we submitted the notebook to the complete?<br>\nAnd the test.csv path will be replaced by <strong>test.csv in the private dataset</strong>.</p>",
      "rawMarkdown": "@teeyee314 \n\nhi, thanks for your advice.\n\nI have a question. \nThe test dataset csv file is ** ../input/rsna-str-pulmonary-embolism-detection/test.csv**.\n\nIs that mean we just write this path and read the **test.csv** when we submitted the notebook to the complete?\nAnd the test.csv path will be replaced by **test.csv in the private dataset**.",
      "replies": [
        {
          "id": 1020508,
          "postDate": "2020-09-21T07:59:10.753Z",
          "content": "<p>if i'm not mistaken, all files in the directory is replaced from public directory to private directory. the file names would be the same - stuff like sample_submission.csv, test.csv, etc. that would mean you shouldn't need to change anything if your submission pipeline is written correctly.</p>",
          "rawMarkdown": "if i'm not mistaken, all files in the directory is replaced from public directory to private directory. the file names would be the same - stuff like sample_submission.csv, test.csv, etc. that would mean you shouldn't need to change anything if your submission pipeline is written correctly."
        }
      ]
    },
    {
      "id": 1020366,
      "postDate": "2020-09-21T05:57:52.897Z",
      "content": "<p>My guess is you didn't successfully create a <code>submission.csv</code> file. Where in the execution of the code that threw an error that caused the code block to run the submission.csv file is not clear. So for example I had a network timeout error and I thought it might have been caused by gdcm error being thrown, after adding gdcm, I re-ran and still got the same error. I then re-wrote the datagenerator and prediction code using an alternate method from what I had previously and I was able to get a successful submission. </p>\n<p>My advice. Work backwards from the code that executes the submission.csv. <em>What information does that submission.csv line depend on?</em></p>",
      "rawMarkdown": "My guess is you didn't successfully create a `submission.csv` file. Where in the execution of the code that threw an error that caused the code block to run the submission.csv file is not clear. So for example I had a network timeout error and I thought it might have been caused by gdcm error being thrown, after adding gdcm, I re-ran and still got the same error. I then re-wrote the datagenerator and prediction code using an alternate method from what I had previously and I was able to get a successful submission. \n\nMy advice. Work backwards from the code that executes the submission.csv. *What information does that submission.csv line depend on?*",
      "replies": [
        {
          "id": 1020537,
          "postDate": "2020-09-21T08:30:26.813Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1046182,
      "author_name": "Abhi Mahule",
      "author_url": "",
      "post_date": "2020-10-11T12:21:08.880000",
      "content": "<p>As a basic sanity test, I have created a notebook to check if submission.csv can be properly generated in the Kaggle runtime.</p>\n<p>Please take a look here and copy as needed - <a href=\"https://www.kaggle.com/abhimahule/practice-submission\" target=\"_blank\">https://www.kaggle.com/abhimahule/practice-submission</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1045202,
      "author_name": "zhugekongan",
      "author_url": "",
      "post_date": "2020-10-10T12:35:13.427000",
      "content": "<p>This is because fake code is used. You use fake code to generate Submission CSV files, but your fake code only runs in the fake environment, and it does not run in the real submission environment. You should make this code run in both environments.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1045252,
          "author_name": "Shivam",
          "author_url": "",
          "post_date": "2020-10-10T13:06:00.450000",
          "content": "<p>can you explain this little bit more ,what do you mean by fake code ??</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1037946,
      "author_name": "Shivam",
      "author_url": "",
      "post_date": "2020-10-05T12:45:28.977000",
      "content": "<p>hii <a href=\"https://www.kaggle.com/qzw\" target=\"_blank\">@qzw</a> I am facing same problem, please give any suggestion!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1020580,
      "author_name": "qzw",
      "author_url": "",
      "post_date": "2020-09-21T09:23:08.390000",
      "content": "<p><a href=\"https://www.kaggle.com/teeyee314\" target=\"_blank\">@teeyee314</a> </p>\n<p>Thanks,<br>\nI think so.</p>\n<p>But I am still confused about my submission.<br>\nJust as we thought, we didn't need to change anything if our notebook can run correctly.<br>\nHowever, my notebook can run correctly before submitting it to complete. <br>\nWhy they can't generate submission.csv in the private dataset?</p>\n<p>I just use **../input/rsna-str-pulmonary-embolism-detection/test.csv **and my model weights.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1020634,
          "author_name": "SumanSudhir",
          "author_url": "",
          "post_date": "2020-09-21T10:11:45.427000",
          "content": "<p>create a notebook for inferencing the model, built a pipeline so store result of images present in the test dataset in the sample submission format which is given. Keep in mind to create a file named <code>submission.csv</code> in same format as the sample submission file. when you commit the notebook it will create a file named `submission.csv' in output and the submit option will be available there beside it. Just submit that file it will re-run your notebook on a complete test dataset consisting of private and public test set and will give you your score.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1020823,
          "author_name": "quadcore/Richard Epstein",
          "author_url": "",
          "post_date": "2020-09-21T13:03:14.907000",
          "content": "<p>Thoughts:</p>\n<p>Are you timing out? Does the notebook fail relatively quickly or after 9 hours?<br>\nMake sure you can run your notebook with Internet turned off.<br>\nThe Private Test data is larger than the Public Test data. Are you reading it all into memory at the same time? Any chance you are running out of memory?<br>\nMake sure your submitted file has GPU turned on if you use it for inference.<br>\nDon't assume anything about the private test data<br>\n--I think there is only one SeriesInstanceUID per StudyInstanceUID, but I don't know if we know for sure.<br>\n--DICOM data has studies that don't start at Image number 1. There are gaps in the image numbers (internal Dicom metadata).</p>\n<p>Build your submission file by reading in sample_submission.csv and replacing labels with your calculated values. That way you are guaranteed that all rows are present, even if for some reason your code doesn't calculate a value for a given Study or SOP. Using this method, you can make a submission that only reads 100 lines of the test.csv file. Your score will be poor, but if it succeeds you know your code is structured correctly. Then you are dealing with Memory or Timeout issues. You can also do your testing much faster.</p>\n<p>Other thoughts - sometimes a Batch Size works most of the time, but fails after a while. Lower your Batch Size if that could be the problem.</p>\n<p>How long does Inference on the entire dataset take you? In truth, I don't think I've even tried to run a full inference on the dataset. With pre-processing (which I might not be doing in the most efficient way), I can only inference 80,000 images in 9 hours.</p>\n<p>A final option to at least submit a valid file and get an idea of how you are doing is to run the model without Commit against the public test.csv file (as a bonus, you can use the TPU). Then put that submission in a dataset and write a notebook to read that submission, update the 40% of the private sample_submission.csv that it covers and submit that. Leave default values for the rest of the sample_submission.csv file. That way, at least you can see which direction your score is going and confirm that your submission format is correct.</p>\n<p>Also, look closely at your public submission file. Verify it has all required rows and that there are no NAN anywhere.</p>\n<p>Best of luck,</p>\n<p>-Rich</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1020440,
      "author_name": "qzw",
      "author_url": "",
      "post_date": "2020-09-21T06:47:52.290000",
      "content": "<p><a href=\"https://www.kaggle.com/teeyee314\" target=\"_blank\">@teeyee314</a> </p>\n<p>hi, thanks for your advice.</p>\n<p>I have a question. <br>\nThe test dataset csv file is ** ../input/rsna-str-pulmonary-embolism-detection/test.csv**.</p>\n<p>Is that mean we just write this path and read the <strong>test.csv</strong> when we submitted the notebook to the complete?<br>\nAnd the test.csv path will be replaced by <strong>test.csv in the private dataset</strong>.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1020508,
          "author_name": "Tim Yee",
          "author_url": "",
          "post_date": "2020-09-21T07:59:10.753000",
          "content": "<p>if i'm not mistaken, all files in the directory is replaced from public directory to private directory. the file names would be the same - stuff like sample_submission.csv, test.csv, etc. that would mean you shouldn't need to change anything if your submission pipeline is written correctly.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1020366,
      "author_name": "Tim Yee",
      "author_url": "",
      "post_date": "2020-09-21T05:57:52.897000",
      "content": "<p>My guess is you didn't successfully create a <code>submission.csv</code> file. Where in the execution of the code that threw an error that caused the code block to run the submission.csv file is not clear. So for example I had a network timeout error and I thought it might have been caused by gdcm error being thrown, after adding gdcm, I re-ran and still got the same error. I then re-wrote the datagenerator and prediction code using an alternate method from what I had previously and I was able to get a successful submission. </p>\n<p>My advice. Work backwards from the code that executes the submission.csv. <em>What information does that submission.csv line depend on?</em></p>",
      "votes": 0,
      "replies": [
        {
          "id": 1020537,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-09-21T08:30:26.813000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1020139": "hi all,\n\nI create a notebook, which can run successfully in the public data set and save submission.csv in the working dir.\n\nBut after submitting to the complete, It aways can't return the right score. The status shows **Submission CSV Not Found**.\n\nI don't know which step is wrong. And I can't see the debug information. \n\nIs there someone who can help me to solve this problem?\n\nThanks",
    "1046182": "As a basic sanity test, I have created a notebook to check if submission.csv can be properly generated in the Kaggle runtime.\n\nPlease take a look here and copy as needed - https://www.kaggle.com/abhimahule/practice-submission",
    "1045202": "This is because fake code is used. You use fake code to generate Submission CSV files, but your fake code only runs in the fake environment, and it does not run in the real submission environment. You should make this code run in both environments.",
    "1037946": "hii @qzw I am facing same problem, please give any suggestion!",
    "1020580": "@teeyee314 \n\nThanks,\nI think so.\n\nBut I am still confused about my submission.\nJust as we thought, we didn't need to change anything if our notebook can run correctly.\nHowever, my notebook can run correctly before submitting it to complete. \nWhy they can't generate submission.csv in the private dataset?\n\nI just use **../input/rsna-str-pulmonary-embolism-detection/test.csv **and my model weights.\n",
    "1020440": "@teeyee314 \n\nhi, thanks for your advice.\n\nI have a question. \nThe test dataset csv file is ** ../input/rsna-str-pulmonary-embolism-detection/test.csv**.\n\nIs that mean we just write this path and read the **test.csv** when we submitted the notebook to the complete?\nAnd the test.csv path will be replaced by **test.csv in the private dataset**.",
    "1020366": "My guess is you didn't successfully create a `submission.csv` file. Where in the execution of the code that threw an error that caused the code block to run the submission.csv file is not clear. So for example I had a network timeout error and I thought it might have been caused by gdcm error being thrown, after adding gdcm, I re-ran and still got the same error. I then re-wrote the datagenerator and prediction code using an alternate method from what I had previously and I was able to get a successful submission. \n\nMy advice. Work backwards from the code that executes the submission.csv. *What information does that submission.csv line depend on?*"
  }
}