{
  "id": 187060,
  "title": "anyone has a working submission notebook",
  "url": "/competitions/rsna-str-pulmonary-embolism-detection/discussion/187060",
  "author_name": "Yee Ng",
  "post_date": "2020-09-27T10:14:32.734000",
  "votes": 3,
  "comment_count": 16,
  "views": 0,
  "content": "<p>Hi all,</p>\n<p>My team is new to a code competition, so we will be grateful for any assistance. Can anyone share a working submission/inference notebook I can learn from?</p>\n<p>I created one <a href=\"url\" target=\"_blank\">https://www.kaggle.com/yeeseng/rsna-pe-submission-notebook</a> but for some reason, it keeps giving me \"score submission error\". The bizarre thing is that I didn't even use the predictions of my model inference although there is a block of codes inferring a pre-trained model in the loop. Instead I tried submitting 0.5 for all labels and the notebook still erred out. When I commented the code block doing the model inference (not touching the ids and the labels which are left at 0.5), the submission strangely becomes successful. </p>\n<p><a href=\"url\" target=\"_blank\">https://www.kaggle.com/yeeseng/rsna-pe-submission-notebook</a></p>",
  "messages": [
    {
      "id": 1028917,
      "postDate": "2020-09-27T10:14:32.733Z",
      "content": "<p>Hi all,</p>\n<p>My team is new to a code competition, so we will be grateful for any assistance. Can anyone share a working submission/inference notebook I can learn from?</p>\n<p>I created one <a href=\"url\" target=\"_blank\">https://www.kaggle.com/yeeseng/rsna-pe-submission-notebook</a> but for some reason, it keeps giving me \"score submission error\". The bizarre thing is that I didn't even use the predictions of my model inference although there is a block of codes inferring a pre-trained model in the loop. Instead I tried submitting 0.5 for all labels and the notebook still erred out. When I commented the code block doing the model inference (not touching the ids and the labels which are left at 0.5), the submission strangely becomes successful. </p>\n<p><a href=\"url\" target=\"_blank\">https://www.kaggle.com/yeeseng/rsna-pe-submission-notebook</a></p>",
      "rawMarkdown": "Hi all,\n\nMy team is new to a code competition, so we will be grateful for any assistance. Can anyone share a working submission/inference notebook I can learn from?\n\nI created one [https://www.kaggle.com/yeeseng/rsna-pe-submission-notebook](url) but for some reason, it keeps giving me \"score submission error\". The bizarre thing is that I didn't even use the predictions of my model inference although there is a block of codes inferring a pre-trained model in the loop. Instead I tried submitting 0.5 for all labels and the notebook still erred out. When I commented the code block doing the model inference (not touching the ids and the labels which are left at 0.5), the submission strangely becomes successful. \n\n[https://www.kaggle.com/yeeseng/rsna-pe-submission-notebook](url)",
      "votes": 3
    },
    {
      "id": 1030195,
      "postDate": "2020-09-28T13:19:59.153Z",
      "content": "<p>Maybe you can try to write submission file with pandas?<br>\nI tried your submission file, after I read it by pandas, the label column can't be read to local memory.<br>\nSo I guess there are something wrong with how you write the label column.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3617078%2F8a4ae4b56f8711bf6f76fb2777dc8725%2F2020-09-28%209.17.52.png?generation=1601299144809410&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Maybe you can try to write submission file with pandas?\nI tried your submission file, after I read it by pandas, the label column can't be read to local memory.\nSo I guess there are something wrong with how you write the label column.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3617078%2F8a4ae4b56f8711bf6f76fb2777dc8725%2F2020-09-28%209.17.52.png?generation=1601299144809410&alt=media)",
      "votes": 1,
      "replies": [
        {
          "id": 1030262,
          "postDate": "2020-09-28T14:08:00.173Z",
          "content": "<p>I used pandas to write to file in earlier versions. I am trying to write to file directly on more recent versions just to see if it changes anything. </p>\n<p>Fixed it - there shouldn't be spaces in a comma separated values.</p>",
          "rawMarkdown": "I used pandas to write to file in earlier versions. I am trying to write to file directly on more recent versions just to see if it changes anything. \n\nFixed it - there shouldn't be spaces in a comma separated values.",
          "votes": 1
        },
        {
          "id": 1030734,
          "postDate": "2020-09-28T22:51:11.910Z",
          "content": "<p>Nice! good to hear that.</p>",
          "rawMarkdown": "Nice! good to hear that.\n"
        }
      ]
    },
    {
      "id": 1028967,
      "postDate": "2020-09-27T11:08:43.333Z",
      "content": "<p>Try removing the directory when creating your submission file. </p>\n<p>Just submission.csv</p>\n<p>No directory path</p>\n<p>Rich</p>",
      "rawMarkdown": "Try removing the directory when creating your submission file. \n\nJust submission.csv\n\nNo directory path\n\nRich",
      "votes": 1,
      "replies": [
        {
          "id": 1028979,
          "postDate": "2020-09-27T11:23:04.587Z",
          "content": "<p>Thanks. I will try that.</p>",
          "rawMarkdown": "Thanks. I will try that."
        },
        {
          "id": 1029602,
          "postDate": "2020-09-27T23:18:57.693Z",
          "content": "<p>Nope… not working :(</p>",
          "rawMarkdown": "Nope... not working :(",
          "replies": [
            {
              "id": 1029611,
              "postDate": "2020-09-28T00:06:38.077Z",
              "content": "<p>A few thoughts:</p>\n<p>Always possible that dcmread or calls to RescaleSlope or RescaleIntercept are failing. Consider putting try/except blocks around them to catch any errors. No guarantee that RescaleSlope and RescaleIntercept are present in every file. That being said, I don't think that is your problem.</p>\n<p>Maybe you are running out of memory with your scoreList.append command. Append (to the extent I understand Python), creates a new copy each time. Since the total test data is about 3-4 times the public test data, maybe you are running out of memory. Consider simply writing the submission file to disk as you go. No memory issues. Or some other method of appending.</p>\n<p>To test, run your outer FOR loop 4 times. For speed, maybe skip your baseModel predict call.</p>",
              "rawMarkdown": "A few thoughts:\n\nAlways possible that dcmread or calls to RescaleSlope or RescaleIntercept are failing. Consider putting try/except blocks around them to catch any errors. No guarantee that RescaleSlope and RescaleIntercept are present in every file. That being said, I don't think that is your problem.\n\nMaybe you are running out of memory with your scoreList.append command. Append (to the extent I understand Python), creates a new copy each time. Since the total test data is about 3-4 times the public test data, maybe you are running out of memory. Consider simply writing the submission file to disk as you go. No memory issues. Or some other method of appending.\n\nTo test, run your outer FOR loop 4 times. For speed, maybe skip your baseModel predict call.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 1029756,
      "postDate": "2020-09-28T05:35:38.250Z",
      "content": "<p><a href=\"https://www.kaggle.com/yeeseng\" target=\"_blank\">@yeeseng</a> Another thing, as you code is written, you can't know if the submission file have wrong values, or if the prediction cell fails.<br>\nput the lines:</p>\n<pre><code>print(\"totalEntries\",len(scoreDF))\nscoreDF.to_csv('submission.csv', index=True)\n</code></pre>\n<p>in the prediction cell, so if the loop fails, your notebook will save no submission file and the error you'll get will be <code>Submission CSV Not Found</code><br>\n(When a cell throws an error, the notebook continue and run other cells) </p>",
      "rawMarkdown": "@yeeseng Another thing, as you code is written, you can't know if the submission file have wrong values, or if the prediction cell fails.\nput the lines:\n\n```\nprint(\"totalEntries\",len(scoreDF))\nscoreDF.to_csv('submission.csv', index=True)\n\n```\nin the prediction cell, so if the loop fails, your notebook will save no submission file and the error you'll get will be `Submission CSV Not Found`\n(When a cell throws an error, the notebook continue and run other cells) ",
      "votes": 2,
      "replies": [
        {
          "id": 1030265,
          "postDate": "2020-09-28T14:08:30.077Z",
          "content": "<p>Good point. I will do that.</p>",
          "rawMarkdown": "Good point. I will do that."
        }
      ]
    },
    {
      "id": 1029745,
      "postDate": "2020-09-28T05:25:45.707Z",
      "content": "<p><a href=\"https://www.kaggle.com/yeeseng\" target=\"_blank\">@yeeseng</a> looking at your code it seems you'll get scoring error if the output of you base model will be nan for one of the inputs. try checking for nan and replace with 0.5 if it happens.</p>",
      "rawMarkdown": "@yeeseng looking at your code it seems you'll get scoring error if the output of you base model will be nan for one of the inputs. try checking for nan and replace with 0.5 if it happens.",
      "votes": 2
    },
    {
      "id": 1038003,
      "postDate": "2020-10-05T13:23:53.707Z",
      "content": "<p>How is everyone loading the test data for analysis?</p>\n<p>I'm currently loading the 'test.csv' to get the list of UIDs, construct the filepaths to load them. For me, this generates a working submission csv file. Although it currently only contains UIDs of the public test set, but Julia Elliot said the contents will be swapped out with UIDs of the private test set. <a href=\"url\" target=\"_blank\">https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/182251#1031679</a></p>\n<p>Did anyone try to do an os walk of the test directory? For some reason, when I do that, I get a submission score error. Not that it matters since there's a workaround, but curious why that is.</p>",
      "rawMarkdown": "How is everyone loading the test data for analysis?\n\nI'm currently loading the 'test.csv' to get the list of UIDs, construct the filepaths to load them. For me, this generates a working submission csv file. Although it currently only contains UIDs of the public test set, but Julia Elliot said the contents will be swapped out with UIDs of the private test set. [https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/182251#1031679](url)\n\nDid anyone try to do an os walk of the test directory? For some reason, when I do that, I get a submission score error. Not that it matters since there's a workaround, but curious why that is.",
      "replies": [
        {
          "id": 1038370,
          "postDate": "2020-10-05T18:52:27.330Z",
          "content": "<p>I am using test.csv to get the UIDs. Got \"Submission Scoring Error\" when I increased the batch size. So you can check how many images you are appending. P.S. haven't tried os.walk</p>",
          "rawMarkdown": "I am using test.csv to get the UIDs. Got \"Submission Scoring Error\" when I increased the batch size. So you can check how many images you are appending. P.S. haven't tried os.walk"
        }
      ]
    },
    {
      "id": 1031090,
      "postDate": "2020-09-29T08:34:21.900Z",
      "content": "<p>So the error went away if instead of walking the files in the directory, I load the 'test.csv' to get the list of the ids, and then load the files that way.</p>\n<p>If the 'test.csv' would be updated with the private test set details when running the notebook against the private test set, just like the public test set is replaced by the private test set, I think my notebook will work.</p>\n<p>Thanks everyone for your help!</p>",
      "rawMarkdown": "So the error went away if instead of walking the files in the directory, I load the 'test.csv' to get the list of the ids, and then load the files that way.\n\nIf the 'test.csv' would be updated with the private test set details when running the notebook against the private test set, just like the public test set is replaced by the private test set, I think my notebook will work.\n\nThanks everyone for your help!"
    },
    {
      "id": 1030675,
      "postDate": "2020-09-28T20:34:14.580Z",
      "content": "<p>My notebook is working well on the public test dataset. I have also run it on 420000 images from the training dataset but when I am making a submission the notebook after running for approximately 2 hrs is throwing error <code>Submission CSV Not Found</code>.  I don't know what exactly is going wrong as I am not getting any error on the public test dataset and even on 420K(approx whole test dataset) train images.</p>",
      "rawMarkdown": "My notebook is working well on the public test dataset. I have also run it on 420000 images from the training dataset but when I am making a submission the notebook after running for approximately 2 hrs is throwing error `Submission CSV Not Found`.  I don't know what exactly is going wrong as I am not getting any error on the public test dataset and even on 420K(approx whole test dataset) train images.",
      "replies": [
        {
          "id": 1030682,
          "postDate": "2020-09-28T20:39:49.063Z",
          "content": "<p>Maybe you have an error while reading or processing one of the test images. <br>\nYou might want to add <code>try: except:</code> around some of your loading or processing.</p>",
          "rawMarkdown": "Maybe you have an error while reading or processing one of the test images. \nYou might want to add `try: except:` around some of your loading or processing.",
          "votes": 1
        },
        {
          "id": 1030755,
          "postDate": "2020-09-28T23:35:44.210Z",
          "content": "<p>Hey, thanks it worked but I don't know that if it is DCM file why is it showing an error with pydicom, and for how many images from the private test set exactly I am doing wrong prediction due to this error. </p>",
          "rawMarkdown": "Hey, thanks it worked but I don't know that if it is DCM file why is it showing an error with pydicom, and for how many images from the private test set exactly I am doing wrong prediction due to this error. "
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1030195,
      "author_name": "Tsai29",
      "author_url": "",
      "post_date": "2020-09-28T13:19:59.153000",
      "content": "<p>Maybe you can try to write submission file with pandas?<br>\nI tried your submission file, after I read it by pandas, the label column can't be read to local memory.<br>\nSo I guess there are something wrong with how you write the label column.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3617078%2F8a4ae4b56f8711bf6f76fb2777dc8725%2F2020-09-28%209.17.52.png?generation=1601299144809410&amp;alt=media\" alt=\"\"></p>",
      "votes": 1,
      "replies": [
        {
          "id": 1030262,
          "author_name": "Yee Ng",
          "author_url": "",
          "post_date": "2020-09-28T14:08:00.173000",
          "content": "<p>I used pandas to write to file in earlier versions. I am trying to write to file directly on more recent versions just to see if it changes anything. </p>\n<p>Fixed it - there shouldn't be spaces in a comma separated values.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1030734,
          "author_name": "Tsai29",
          "author_url": "",
          "post_date": "2020-09-28T22:51:11.910000",
          "content": "<p>Nice! good to hear that.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1028967,
      "author_name": "quadcore/Richard Epstein",
      "author_url": "",
      "post_date": "2020-09-27T11:08:43.333000",
      "content": "<p>Try removing the directory when creating your submission file. </p>\n<p>Just submission.csv</p>\n<p>No directory path</p>\n<p>Rich</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1028979,
          "author_name": "Yee Ng",
          "author_url": "",
          "post_date": "2020-09-27T11:23:04.587000",
          "content": "<p>Thanks. I will try that.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1029602,
          "author_name": "Yee Ng",
          "author_url": "",
          "post_date": "2020-09-27T23:18:57.693000",
          "content": "<p>Nope… not working :(</p>",
          "votes": 0,
          "replies": [
            {
              "id": 1029611,
              "author_name": "quadcore/Richard Epstein",
              "author_url": "",
              "post_date": "2020-09-28T00:06:38.077000",
              "content": "<p>A few thoughts:</p>\n<p>Always possible that dcmread or calls to RescaleSlope or RescaleIntercept are failing. Consider putting try/except blocks around them to catch any errors. No guarantee that RescaleSlope and RescaleIntercept are present in every file. That being said, I don't think that is your problem.</p>\n<p>Maybe you are running out of memory with your scoreList.append command. Append (to the extent I understand Python), creates a new copy each time. Since the total test data is about 3-4 times the public test data, maybe you are running out of memory. Consider simply writing the submission file to disk as you go. No memory issues. Or some other method of appending.</p>\n<p>To test, run your outer FOR loop 4 times. For speed, maybe skip your baseModel predict call.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 1029756,
      "author_name": "yuval reina",
      "author_url": "",
      "post_date": "2020-09-28T05:35:38.250000",
      "content": "<p><a href=\"https://www.kaggle.com/yeeseng\" target=\"_blank\">@yeeseng</a> Another thing, as you code is written, you can't know if the submission file have wrong values, or if the prediction cell fails.<br>\nput the lines:</p>\n<pre><code>print(\"totalEntries\",len(scoreDF))\nscoreDF.to_csv('submission.csv', index=True)\n</code></pre>\n<p>in the prediction cell, so if the loop fails, your notebook will save no submission file and the error you'll get will be <code>Submission CSV Not Found</code><br>\n(When a cell throws an error, the notebook continue and run other cells) </p>",
      "votes": 2,
      "replies": [
        {
          "id": 1030265,
          "author_name": "Yee Ng",
          "author_url": "",
          "post_date": "2020-09-28T14:08:30.077000",
          "content": "<p>Good point. I will do that.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1029745,
      "author_name": "yuval reina",
      "author_url": "",
      "post_date": "2020-09-28T05:25:45.707000",
      "content": "<p><a href=\"https://www.kaggle.com/yeeseng\" target=\"_blank\">@yeeseng</a> looking at your code it seems you'll get scoring error if the output of you base model will be nan for one of the inputs. try checking for nan and replace with 0.5 if it happens.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1038003,
      "author_name": "Yee Ng",
      "author_url": "",
      "post_date": "2020-10-05T13:23:53.707000",
      "content": "<p>How is everyone loading the test data for analysis?</p>\n<p>I'm currently loading the 'test.csv' to get the list of UIDs, construct the filepaths to load them. For me, this generates a working submission csv file. Although it currently only contains UIDs of the public test set, but Julia Elliot said the contents will be swapped out with UIDs of the private test set. <a href=\"url\" target=\"_blank\">https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/182251#1031679</a></p>\n<p>Did anyone try to do an os walk of the test directory? For some reason, when I do that, I get a submission score error. Not that it matters since there's a workaround, but curious why that is.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1038370,
          "author_name": "Zhanseri Ikram",
          "author_url": "",
          "post_date": "2020-10-05T18:52:27.330000",
          "content": "<p>I am using test.csv to get the UIDs. Got \"Submission Scoring Error\" when I increased the batch size. So you can check how many images you are appending. P.S. haven't tried os.walk</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1031090,
      "author_name": "Yee Ng",
      "author_url": "",
      "post_date": "2020-09-29T08:34:21.900000",
      "content": "<p>So the error went away if instead of walking the files in the directory, I load the 'test.csv' to get the list of the ids, and then load the files that way.</p>\n<p>If the 'test.csv' would be updated with the private test set details when running the notebook against the private test set, just like the public test set is replaced by the private test set, I think my notebook will work.</p>\n<p>Thanks everyone for your help!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1030675,
      "author_name": "SumanSudhir",
      "author_url": "",
      "post_date": "2020-09-28T20:34:14.580000",
      "content": "<p>My notebook is working well on the public test dataset. I have also run it on 420000 images from the training dataset but when I am making a submission the notebook after running for approximately 2 hrs is throwing error <code>Submission CSV Not Found</code>.  I don't know what exactly is going wrong as I am not getting any error on the public test dataset and even on 420K(approx whole test dataset) train images.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1030682,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2020-09-28T20:39:49.063000",
          "content": "<p>Maybe you have an error while reading or processing one of the test images. <br>\nYou might want to add <code>try: except:</code> around some of your loading or processing.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1030755,
          "author_name": "SumanSudhir",
          "author_url": "",
          "post_date": "2020-09-28T23:35:44.210000",
          "content": "<p>Hey, thanks it worked but I don't know that if it is DCM file why is it showing an error with pydicom, and for how many images from the private test set exactly I am doing wrong prediction due to this error. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1028917": "Hi all,\n\nMy team is new to a code competition, so we will be grateful for any assistance. Can anyone share a working submission/inference notebook I can learn from?\n\nI created one [https://www.kaggle.com/yeeseng/rsna-pe-submission-notebook](url) but for some reason, it keeps giving me \"score submission error\". The bizarre thing is that I didn't even use the predictions of my model inference although there is a block of codes inferring a pre-trained model in the loop. Instead I tried submitting 0.5 for all labels and the notebook still erred out. When I commented the code block doing the model inference (not touching the ids and the labels which are left at 0.5), the submission strangely becomes successful. \n\n[https://www.kaggle.com/yeeseng/rsna-pe-submission-notebook](url)",
    "1030195": "Maybe you can try to write submission file with pandas?\nI tried your submission file, after I read it by pandas, the label column can't be read to local memory.\nSo I guess there are something wrong with how you write the label column.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3617078%2F8a4ae4b56f8711bf6f76fb2777dc8725%2F2020-09-28%209.17.52.png?generation=1601299144809410&alt=media)",
    "1028967": "Try removing the directory when creating your submission file. \n\nJust submission.csv\n\nNo directory path\n\nRich",
    "1029756": "@yeeseng Another thing, as you code is written, you can't know if the submission file have wrong values, or if the prediction cell fails.\nput the lines:\n\n```\nprint(\"totalEntries\",len(scoreDF))\nscoreDF.to_csv('submission.csv', index=True)\n\n```\nin the prediction cell, so if the loop fails, your notebook will save no submission file and the error you'll get will be `Submission CSV Not Found`\n(When a cell throws an error, the notebook continue and run other cells) ",
    "1029745": "@yeeseng looking at your code it seems you'll get scoring error if the output of you base model will be nan for one of the inputs. try checking for nan and replace with 0.5 if it happens.",
    "1038003": "How is everyone loading the test data for analysis?\n\nI'm currently loading the 'test.csv' to get the list of UIDs, construct the filepaths to load them. For me, this generates a working submission csv file. Although it currently only contains UIDs of the public test set, but Julia Elliot said the contents will be swapped out with UIDs of the private test set. [https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/182251#1031679](url)\n\nDid anyone try to do an os walk of the test directory? For some reason, when I do that, I get a submission score error. Not that it matters since there's a workaround, but curious why that is.",
    "1031090": "So the error went away if instead of walking the files in the directory, I load the 'test.csv' to get the list of the ids, and then load the files that way.\n\nIf the 'test.csv' would be updated with the private test set details when running the notebook against the private test set, just like the public test set is replaced by the private test set, I think my notebook will work.\n\nThanks everyone for your help!",
    "1030675": "My notebook is working well on the public test dataset. I have also run it on 420000 images from the training dataset but when I am making a submission the notebook after running for approximately 2 hrs is throwing error `Submission CSV Not Found`.  I don't know what exactly is going wrong as I am not getting any error on the public test dataset and even on 420K(approx whole test dataset) train images."
  }
}