{
  "id": 432745,
  "title": "[Solved] Submission scoring error",
  "url": "/competitions/rsna-2023-abdominal-trauma-detection/discussion/432745",
  "author_name": "Priya Nagda",
  "post_date": "2023-08-18T18:07:37.373000",
  "votes": 1,
  "comment_count": 19,
  "views": 0,
  "content": "<p>After an hour of running :/<br>\nAny pointers on what could be the issue?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7037441%2F44071aea474c2c3ad0207a3f1f228447%2FScreenshot%202023-08-18%20at%2011.34.55%20PM.png?generation=1692382025922257&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7037441%2Fe46f25696c6ef1f16775e496937bdf1e%2FScreenshot%202023-08-18%20at%2011.34.24%20PM.png?generation=1692382039233228&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 2397255,
      "postDate": "2023-08-18T21:09:14.213Z",
      "content": "<p>Hi, probably is due to the exponent numbers inside the pd dataframe, you can try with this code:</p>\n<p><code>sub_df.to_csv(\"submission.csv\", index=False, float_format='%.20f')</code></p>",
      "rawMarkdown": "Hi, probably is due to the exponent numbers inside the pd dataframe, you can try with this code:\n\n``sub_df.to_csv(\"submission.csv\", index=False, float_format='%.20f')``",
      "votes": 1,
      "replies": [
        {
          "id": 2397443,
          "postDate": "2023-08-19T03:25:46.763Z",
          "content": "<p>tried this, still gives the same error <br>\nalso for anyone else trying out the same thing, <code>sub_df.to_csv(\"submission.csv\", index=False, float_format='%.20f')</code> doesn't work for me. I had to use <code>sub_df.round(3).to_csv(\"submission.csv\", index=False)</code> just in case it doesn't work for you too,</p>",
          "rawMarkdown": "tried this, still gives the same error \nalso for anyone else trying out the same thing, `sub_df.to_csv(\"submission.csv\", index=False, float_format='%.20f')` doesn't work for me. I had to use `sub_df.round(3).to_csv(\"submission.csv\", index=False)` just in case it doesn't work for you too,",
          "votes": 1
        }
      ]
    },
    {
      "id": 2397106,
      "postDate": "2023-08-18T18:07:37.373Z",
      "content": "<p>After an hour of running :/<br>\nAny pointers on what could be the issue?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7037441%2F44071aea474c2c3ad0207a3f1f228447%2FScreenshot%202023-08-18%20at%2011.34.55%20PM.png?generation=1692382025922257&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7037441%2Fe46f25696c6ef1f16775e496937bdf1e%2FScreenshot%202023-08-18%20at%2011.34.24%20PM.png?generation=1692382039233228&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "After an hour of running :/\nAny pointers on what could be the issue?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7037441%2F44071aea474c2c3ad0207a3f1f228447%2FScreenshot%202023-08-18%20at%2011.34.55%20PM.png?generation=1692382025922257&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7037441%2Fe46f25696c6ef1f16775e496937bdf1e%2FScreenshot%202023-08-18%20at%2011.34.24%20PM.png?generation=1692382039233228&alt=media)",
      "votes": 1
    },
    {
      "id": 2475786,
      "postDate": "2023-10-10T06:45:49.373Z",
      "content": "<p>hi, can you show us your result, I just want to see the format of the submission.csv, you could fill whatever values you what</p>",
      "rawMarkdown": "hi, can you show us your result, I just want to see the format of the submission.csv, you could fill whatever values you what",
      "replies": [
        {
          "id": 2476048,
          "postDate": "2023-10-10T09:54:57.987Z",
          "content": "<p>As I figured it out…the format is same as sample_submission…no matter if you  round values or not….</p>\n<p>MAIN ISSUE seems to be that EVERY SINGLE test patient should be present in the submission file!!!! Order also doesnt matter as the scoring is done on group weighted loss average(please have a look at the scoring script!)..<br>\nThe most likely reason (that's what was in my case)..if there is an exception in reading DICOM files , one might be skipping it…that means that patient-id is skipped altogether in submission file which LEADS to 'Submission Scoring Error'…because \"solution\" is coming in with all the test patients and submission file will not have the skipped patients… at least this was my case…</p>",
          "rawMarkdown": "As I figured it out...the format is same as sample_submission...no matter if you  round values or not....\n\nMAIN ISSUE seems to be that EVERY SINGLE test patient should be present in the submission file!!!! Order also doesnt matter as the scoring is done on group weighted loss average(please have a look at the scoring script!)..\nThe most likely reason (that's what was in my case)..if there is an exception in reading DICOM files , one might be skipping it...that means that patient-id is skipped altogether in submission file which LEADS to 'Submission Scoring Error'...because \"solution\" is coming in with all the test patients and submission file will not have the skipped patients... at least this was my case...\n\n",
          "replies": [
            {
              "id": 2476070,
              "postDate": "2023-10-10T10:13:37.113Z",
              "content": "<p>I use the file name as my patient_id, and check the number of series_id,I am not sure is it work, but I still have the  Submission scoring error problems</p>",
              "rawMarkdown": "I use the file name as my patient_id, and check the number of series_id,I am not sure is it work, but I still have the  Submission scoring error problems"
            }
          ]
        }
      ]
    },
    {
      "id": 2404412,
      "postDate": "2023-08-23T08:37:06.327Z",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/priyanagda\" target=\"_blank\">@priyanagda</a> <br>\nHow did you solve it<br>\nI am stuck for no reason</p>",
      "rawMarkdown": "Hey @priyanagda \nHow did you solve it\nI am stuck for no reason",
      "replies": [
        {
          "id": 2404538,
          "postDate": "2023-08-23T09:56:10.793Z",
          "content": "<p>could you share your notebook?</p>",
          "rawMarkdown": "could you share your notebook?",
          "replies": [
            {
              "id": 2405598,
              "postDate": "2023-08-24T02:26:55.130Z",
              "content": "<pre><code>pd.set_option(, .)\ndf_sub=pd.read_csv()\ndf_sub[[col  col  df_sub.columns  col !=]]= df_sub[[col  col  df_sub.columns  col !=]].astype()\nsub_df = pd.read_csv()\nsub_df = sub_df[[]]\nsub_df = sub_df.merge(df_sub, on=, how=)\n\n\nsub_df.to_csv(,index=)\nsub_df.head()\n</code></pre>\n<p>This is all what I did <a href=\"https://www.kaggle.com/priyanagda\" target=\"_blank\">@priyanagda</a> </p>",
              "rawMarkdown": "```python\npd.set_option('display.float_format', '{:.16f}'.format)\ndf_sub=pd.read_csv('/kaggle/input/rsna-submissions-hemanth/RSNA-ATD(1).csv')\ndf_sub[[col for col in df_sub.columns if col !=\"patient_id\"]]= df_sub[[col for col in df_sub.columns if col !=\"patient_id\"]].astype('float64')\nsub_df = pd.read_csv('/kaggle/input/rsna-2023-abdominal-trauma-detection/sample_submission.csv')\nsub_df = sub_df[['patient_id']]\nsub_df = sub_df.merge(df_sub, on='patient_id', how='left')\n\n# Store submission\nsub_df.to_csv('submission.csv',index=False)\nsub_df.head()\n```\n\nThis is all what I did @priyanagda "
            },
            {
              "id": 2406010,
              "postDate": "2023-08-24T07:29:47.797Z",
              "content": "<p>Hard to say anything since I don't know what your submission file looks like. My assumptions:</p>\n<ul>\n<li>You probably have more than one entry/row for a single patient.</li>\n<li>NaN values in predictions (try something like fillna(0) and submitting, if this solves it you know the error)</li>\n<li>Keep only the patient_id column and the target cols to be on the safe side</li>\n</ul>",
              "rawMarkdown": "Hard to say anything since I don't know what your submission file looks like. My assumptions:\n- You probably have more than one entry/row for a single patient.\n- NaN values in predictions (try something like fillna(0) and submitting, if this solves it you know the error)\n- Keep only the patient_id column and the target cols to be on the safe side"
            }
          ]
        }
      ]
    },
    {
      "id": 2397536,
      "postDate": "2023-08-19T05:33:18.163Z",
      "content": "<p>Better help possible if you make your notebook public and put a link here!</p>",
      "rawMarkdown": "Better help possible if you make your notebook public and put a link here!",
      "replies": [
        {
          "id": 2398656,
          "postDate": "2023-08-19T19:48:46.533Z",
          "content": "<p>Hi, <a href=\"https://www.kaggle.com/code/priyanagda/atd-inference\" target=\"_blank\">this</a> is the notebook I am trying to submit.</p>",
          "rawMarkdown": "Hi, [this](https://www.kaggle.com/code/priyanagda/atd-inference) is the notebook I am trying to submit.",
          "replies": [
            {
              "id": 2399062,
              "postDate": "2023-08-20T06:01:39.543Z",
              "content": "<p>Did not see anything obvious - was going to fork it and run a larger <a href=\"https://www.kaggle.com/datasets/pcjimmmy/rsna-7-patients-test\" target=\"_blank\">test set</a> that I have created, but your model pth file is private - make it public also.</p>\n<p>As a first guess I would change the batch size.  With the test set we have you can load all the images for a patient (of course there is only 1) and not happen upon a memory error that results in incomplete csv file being created.  But a patient with 700 slices in a DICOM would likely cause memory error (if you stride =1).</p>",
              "rawMarkdown": "Did not see anything obvious - was going to fork it and run a larger [test set](https://www.kaggle.com/datasets/pcjimmmy/rsna-7-patients-test) that I have created, but your model pth file is private - make it public also.\n\nAs a first guess I would change the batch size.  With the test set we have you can load all the images for a patient (of course there is only 1) and not happen upon a memory error that results in incomplete csv file being created.  But a patient with 700 slices in a DICOM would likely cause memory error (if you stride =1)."
            },
            {
              "id": 2399206,
              "postDate": "2023-08-20T07:52:46.877Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2399210,
              "postDate": "2023-08-20T07:53:49.793Z",
              "content": "<p>I believe the notebook is running fine, because if it would have been a memory error, it would have given Notebook threw an exception, isn't it?</p>",
              "rawMarkdown": "I believe the notebook is running fine, because if it would have been a memory error, it would have given Notebook threw an exception, isn't it?"
            },
            {
              "id": 2400266,
              "postDate": "2023-08-21T00:01:40.183Z",
              "content": "<p>Hmm your notebook link is 404 error.  Title now says Solved - assume that means you figured it out.</p>",
              "rawMarkdown": "Hmm your notebook link is 404 error.  Title now says Solved - assume that means you figured it out."
            },
            {
              "id": 2467929,
              "postDate": "2023-10-05T03:08:46.103Z",
              "content": "<p>And how did all solve the \"Submission Scoring Problem\"  finally…I seem to have no 'luck' with that…deliberately using the term 'luck' as there is no specificity about the error…I seem to have more 'try…catch\" than the actual inference code…</p>",
              "rawMarkdown": "And how did all solve the \"Submission Scoring Problem\"  finally...I seem to have no 'luck' with that...deliberately using the term 'luck' as there is no specificity about the error...I seem to have more 'try...catch\" than the actual inference code..."
            }
          ]
        }
      ]
    },
    {
      "id": 2397475,
      "postDate": "2023-08-19T04:16:43.990Z",
      "content": "<p>Given its a submission scoring error, you can generate predictions for the train set, and then use this notebook (<a href=\"https://www.kaggle.com/code/metric/rsna-trauma-metric/notebook\" target=\"_blank\">https://www.kaggle.com/code/metric/rsna-trauma-metric/notebook</a>) to compute the score on the train set. Check if there are any errors.</p>",
      "rawMarkdown": "Given its a submission scoring error, you can generate predictions for the train set, and then use this notebook (https://www.kaggle.com/code/metric/rsna-trauma-metric/notebook) to compute the score on the train set. Check if there are any errors.",
      "replies": [
        {
          "id": 2413667,
          "postDate": "2023-08-29T05:00:50.523Z",
          "content": "<p>I had this error and like you suggested. I was able to successfully make submission. Thank you for this comment!</p>",
          "rawMarkdown": "I had this error and like you suggested. I was able to successfully make submission. Thank you for this comment!",
          "votes": 1
        }
      ]
    },
    {
      "id": 2416753,
      "postDate": "2023-08-31T07:38:42.180Z",
      "content": "<p>del solution[row_id_column_name]<br>\ndel submission[row_id_column_name]</p>\n<p>what is row_id_column_name ??? </p>",
      "rawMarkdown": "del solution[row_id_column_name]\ndel submission[row_id_column_name]\n\nwhat is row_id_column_name ??? ",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2397255,
      "author_name": "Antonio Félix",
      "author_url": "",
      "post_date": "2023-08-18T21:09:14.213000",
      "content": "<p>Hi, probably is due to the exponent numbers inside the pd dataframe, you can try with this code:</p>\n<p><code>sub_df.to_csv(\"submission.csv\", index=False, float_format='%.20f')</code></p>",
      "votes": 1,
      "replies": [
        {
          "id": 2397443,
          "author_name": "Priya Nagda",
          "author_url": "",
          "post_date": "2023-08-19T03:25:46.763000",
          "content": "<p>tried this, still gives the same error <br>\nalso for anyone else trying out the same thing, <code>sub_df.to_csv(\"submission.csv\", index=False, float_format='%.20f')</code> doesn't work for me. I had to use <code>sub_df.round(3).to_csv(\"submission.csv\", index=False)</code> just in case it doesn't work for you too,</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2475786,
      "author_name": "john",
      "author_url": "",
      "post_date": "2023-10-10T06:45:49.373000",
      "content": "<p>hi, can you show us your result, I just want to see the format of the submission.csv, you could fill whatever values you what</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2476048,
          "author_name": "Sunil Krishnan",
          "author_url": "",
          "post_date": "2023-10-10T09:54:57.987000",
          "content": "<p>As I figured it out…the format is same as sample_submission…no matter if you  round values or not….</p>\n<p>MAIN ISSUE seems to be that EVERY SINGLE test patient should be present in the submission file!!!! Order also doesnt matter as the scoring is done on group weighted loss average(please have a look at the scoring script!)..<br>\nThe most likely reason (that's what was in my case)..if there is an exception in reading DICOM files , one might be skipping it…that means that patient-id is skipped altogether in submission file which LEADS to 'Submission Scoring Error'…because \"solution\" is coming in with all the test patients and submission file will not have the skipped patients… at least this was my case…</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2476070,
              "author_name": "john",
              "author_url": "",
              "post_date": "2023-10-10T10:13:37.113000",
              "content": "<p>I use the file name as my patient_id, and check the number of series_id,I am not sure is it work, but I still have the  Submission scoring error problems</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2404412,
      "author_name": "Hemanth Harikrishnan",
      "author_url": "",
      "post_date": "2023-08-23T08:37:06.327000",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/priyanagda\" target=\"_blank\">@priyanagda</a> <br>\nHow did you solve it<br>\nI am stuck for no reason</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2404538,
          "author_name": "Priya Nagda",
          "author_url": "",
          "post_date": "2023-08-23T09:56:10.793000",
          "content": "<p>could you share your notebook?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2405598,
              "author_name": "Hemanth Harikrishnan",
              "author_url": "",
              "post_date": "2023-08-24T02:26:55.130000",
              "content": "<pre><code>pd.set_option(, .)\ndf_sub=pd.read_csv()\ndf_sub[[col  col  df_sub.columns  col !=]]= df_sub[[col  col  df_sub.columns  col !=]].astype()\nsub_df = pd.read_csv()\nsub_df = sub_df[[]]\nsub_df = sub_df.merge(df_sub, on=, how=)\n\n\nsub_df.to_csv(,index=)\nsub_df.head()\n</code></pre>\n<p>This is all what I did <a href=\"https://www.kaggle.com/priyanagda\" target=\"_blank\">@priyanagda</a> </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2406010,
              "author_name": "Priya Nagda",
              "author_url": "",
              "post_date": "2023-08-24T07:29:47.797000",
              "content": "<p>Hard to say anything since I don't know what your submission file looks like. My assumptions:</p>\n<ul>\n<li>You probably have more than one entry/row for a single patient.</li>\n<li>NaN values in predictions (try something like fillna(0) and submitting, if this solves it you know the error)</li>\n<li>Keep only the patient_id column and the target cols to be on the safe side</li>\n</ul>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2397536,
      "author_name": "PC Jimmmy",
      "author_url": "",
      "post_date": "2023-08-19T05:33:18.163000",
      "content": "<p>Better help possible if you make your notebook public and put a link here!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2398656,
          "author_name": "Priya Nagda",
          "author_url": "",
          "post_date": "2023-08-19T19:48:46.533000",
          "content": "<p>Hi, <a href=\"https://www.kaggle.com/code/priyanagda/atd-inference\" target=\"_blank\">this</a> is the notebook I am trying to submit.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2399062,
              "author_name": "PC Jimmmy",
              "author_url": "",
              "post_date": "2023-08-20T06:01:39.543000",
              "content": "<p>Did not see anything obvious - was going to fork it and run a larger <a href=\"https://www.kaggle.com/datasets/pcjimmmy/rsna-7-patients-test\" target=\"_blank\">test set</a> that I have created, but your model pth file is private - make it public also.</p>\n<p>As a first guess I would change the batch size.  With the test set we have you can load all the images for a patient (of course there is only 1) and not happen upon a memory error that results in incomplete csv file being created.  But a patient with 700 slices in a DICOM would likely cause memory error (if you stride =1).</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2399206,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-08-20T07:52:46.877000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2399210,
              "author_name": "Priya Nagda",
              "author_url": "",
              "post_date": "2023-08-20T07:53:49.793000",
              "content": "<p>I believe the notebook is running fine, because if it would have been a memory error, it would have given Notebook threw an exception, isn't it?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2400266,
              "author_name": "PC Jimmmy",
              "author_url": "",
              "post_date": "2023-08-21T00:01:40.183000",
              "content": "<p>Hmm your notebook link is 404 error.  Title now says Solved - assume that means you figured it out.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2467929,
              "author_name": "Sunil Krishnan",
              "author_url": "",
              "post_date": "2023-10-05T03:08:46.103000",
              "content": "<p>And how did all solve the \"Submission Scoring Problem\"  finally…I seem to have no 'luck' with that…deliberately using the term 'luck' as there is no specificity about the error…I seem to have more 'try…catch\" than the actual inference code…</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2397475,
      "author_name": "Jebastin Nadar",
      "author_url": "",
      "post_date": "2023-08-19T04:16:43.990000",
      "content": "<p>Given its a submission scoring error, you can generate predictions for the train set, and then use this notebook (<a href=\"https://www.kaggle.com/code/metric/rsna-trauma-metric/notebook\" target=\"_blank\">https://www.kaggle.com/code/metric/rsna-trauma-metric/notebook</a>) to compute the score on the train set. Check if there are any errors.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2413667,
          "author_name": "Franklin Shih0617",
          "author_url": "",
          "post_date": "2023-08-29T05:00:50.523000",
          "content": "<p>I had this error and like you suggested. I was able to successfully make submission. Thank you for this comment!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2416753,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-08-31T07:38:42.180000",
      "content": "<p>del solution[row_id_column_name]<br>\ndel submission[row_id_column_name]</p>\n<p>what is row_id_column_name ??? </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2397255": "Hi, probably is due to the exponent numbers inside the pd dataframe, you can try with this code:\n\n``sub_df.to_csv(\"submission.csv\", index=False, float_format='%.20f')``",
    "2397106": "After an hour of running :/\nAny pointers on what could be the issue?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7037441%2F44071aea474c2c3ad0207a3f1f228447%2FScreenshot%202023-08-18%20at%2011.34.55%20PM.png?generation=1692382025922257&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7037441%2Fe46f25696c6ef1f16775e496937bdf1e%2FScreenshot%202023-08-18%20at%2011.34.24%20PM.png?generation=1692382039233228&alt=media)",
    "2475786": "hi, can you show us your result, I just want to see the format of the submission.csv, you could fill whatever values you what",
    "2404412": "Hey @priyanagda \nHow did you solve it\nI am stuck for no reason",
    "2397536": "Better help possible if you make your notebook public and put a link here!",
    "2397475": "Given its a submission scoring error, you can generate predictions for the train set, and then use this notebook (https://www.kaggle.com/code/metric/rsna-trauma-metric/notebook) to compute the score on the train set. Check if there are any errors.",
    "2416753": "del solution[row_id_column_name]\ndel submission[row_id_column_name]\n\nwhat is row_id_column_name ??? "
  }
}