{
  "id": 369113,
  "title": "Why is my submission failing?",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/369113",
  "author_name": "Radek Osmulski",
  "post_date": "2022-11-29T01:59:41.801000",
  "votes": 8,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Hey,</p>\n<p>Towards the end of my EDA : <a href=\"https://www.kaggle.com/code/radek1/initial-eda-a-first-look-at-the-data/notebook?scriptVersionId=112383297\" target=\"_blank\">📊 Initial EDA -- a first look at the data 🚀</a> I am outputting the predictions as follows:</p>\n<pre><code>submission = pd.DataFrame(data={'prediction_id': test_csv['prediction_id'], 'cancer': np.random.rand(test_csv.shape[0])})\nsubmission.to_csv('submission.csv', index=False)\n</code></pre>\n<p>However, regardless of whether I output floats in the range of <code>[0,1)</code> as predictions, or all zeros, I am met with <code>submission scoring error</code>.</p>\n<p>Could I please ask for help in figuring out what might be going wrong there? </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F4d0dc870b8879102e50463007ea302ab%2Fsubmission_scoring_error.png?generation=1669687168872049&amp;alt=media\" alt=\"\"></p>\n<p>Am I doing something wrong or is there some issue with the machinery for scoring?</p>\n<p>Thank you for your help!</p>",
  "messages": [
    {
      "id": 2047781,
      "postDate": "2022-11-29T03:30:11.060Z",
      "content": "<p>Maybe, but I think it's because your submission.csv has duplicates for prediction_id.</p>",
      "rawMarkdown": "Maybe, but I think it's because your submission.csv has duplicates for prediction_id.",
      "votes": 7,
      "replies": [
        {
          "id": 2047784,
          "postDate": "2022-11-29T03:32:40.687Z",
          "content": "<p>Great suggestion, thank you!!!! will give this a go, that makes sense!</p>",
          "rawMarkdown": "Great suggestion, thank you!!!! will give this a go, that makes sense!",
          "votes": 1
        },
        {
          "id": 2047792,
          "postDate": "2022-11-29T03:42:23.403Z",
          "content": "<p>That was it, thank you very much!!!!!!!!! 🙏</p>",
          "rawMarkdown": "That was it, thank you very much!!!!!!!!! 🙏",
          "votes": 2
        }
      ]
    },
    {
      "id": 2047727,
      "postDate": "2022-11-29T01:59:41.800Z",
      "content": "<p>Hey,</p>\n<p>Towards the end of my EDA : <a href=\"https://www.kaggle.com/code/radek1/initial-eda-a-first-look-at-the-data/notebook?scriptVersionId=112383297\" target=\"_blank\">📊 Initial EDA -- a first look at the data 🚀</a> I am outputting the predictions as follows:</p>\n<pre><code>submission = pd.DataFrame(data={'prediction_id': test_csv['prediction_id'], 'cancer': np.random.rand(test_csv.shape[0])})\nsubmission.to_csv('submission.csv', index=False)\n</code></pre>\n<p>However, regardless of whether I output floats in the range of <code>[0,1)</code> as predictions, or all zeros, I am met with <code>submission scoring error</code>.</p>\n<p>Could I please ask for help in figuring out what might be going wrong there? </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F4d0dc870b8879102e50463007ea302ab%2Fsubmission_scoring_error.png?generation=1669687168872049&amp;alt=media\" alt=\"\"></p>\n<p>Am I doing something wrong or is there some issue with the machinery for scoring?</p>\n<p>Thank you for your help!</p>",
      "rawMarkdown": "Hey,\n\nTowards the end of my EDA : [📊 Initial EDA -- a first look at the data 🚀](https://www.kaggle.com/code/radek1/initial-eda-a-first-look-at-the-data/notebook?scriptVersionId=112383297) I am outputting the predictions as follows:\n\n```\nsubmission = pd.DataFrame(data={'prediction_id': test_csv['prediction_id'], 'cancer': np.random.rand(test_csv.shape[0])})\nsubmission.to_csv('submission.csv', index=False)\n```\n\nHowever, regardless of whether I output floats in the range of `[0,1)` as predictions, or all zeros, I am met with `submission scoring error`.\n\nCould I please ask for help in figuring out what might be going wrong there? \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F4d0dc870b8879102e50463007ea302ab%2Fsubmission_scoring_error.png?generation=1669687168872049&alt=media)\n\nAm I doing something wrong or is there some issue with the machinery for scoring?\n\nThank you for your help!",
      "votes": 8
    },
    {
      "id": 2054782,
      "postDate": "2022-12-04T13:04:40.187Z",
      "content": "<p>you can simply prune duplicates by using max cancer predictions:<br>\n<code>df_preds = df_preds.groupby('prediction_id').max()</code><br>\nsee: <a href=\"https://www.kaggle.com/code/jirkaborovec/mammography-eda-loading-dicom\" target=\"_blank\">https://www.kaggle.com/code/jirkaborovec/mammography-eda-loading-dicom</a></p>",
      "rawMarkdown": "you can simply prune duplicates by using max cancer predictions:\n`df_preds = df_preds.groupby('prediction_id').max()`\nsee: https://www.kaggle.com/code/jirkaborovec/mammography-eda-loading-dicom",
      "votes": 1,
      "replies": [
        {
          "id": 2156447,
          "postDate": "2023-02-23T10:45:55.360Z",
          "content": "<p>Hi, I have the same issue and I used the code above. I have unique prediction ID values and the csv is comma separated. However I am still getting the error message. However, when I submit the sample submission file it works. Would anyone be able to help me in this matter.<br>\nThanks &amp; Best Regards<br>\nAMJS</p>",
          "rawMarkdown": "Hi, I have the same issue and I used the code above. I have unique prediction ID values and the csv is comma separated. However I am still getting the error message. However, when I submit the sample submission file it works. Would anyone be able to help me in this matter.\nThanks & Best Regards\nAMJS"
        }
      ]
    },
    {
      "id": 2050198,
      "postDate": "2022-11-30T15:27:21.760Z",
      "content": "<p>Did you find any solution? My submission.csv looks like this, why is it still failing?</p>\n<p>prediction_id cancer<br>\n10008_R 0.123<br>\n10008_L 0.123</p>",
      "rawMarkdown": "Did you find any solution? My submission.csv looks like this, why is it still failing?\n\nprediction_id cancer\n10008_R 0.123\n10008_L 0.123",
      "votes": 1,
      "replies": [
        {
          "id": 2050536,
          "postDate": "2022-11-30T20:20:15.950Z",
          "content": "<p>these need to be comma seperated, they need to be a proper csv file</p>\n<p>the solution for me was removing duplicates, duplicated predictions for the same <code>prediction_id</code></p>",
          "rawMarkdown": "these need to be comma seperated, they need to be a proper csv file\n\nthe solution for me was removing duplicates, duplicated predictions for the same `prediction_id`"
        }
      ]
    },
    {
      "id": 2047780,
      "postDate": "2022-11-29T03:29:40.123Z",
      "content": "<p>Do you drop duplicated prediction_id?</p>",
      "rawMarkdown": "Do you drop duplicated prediction_id?",
      "votes": 1,
      "replies": [
        {
          "id": 2047785,
          "postDate": "2022-11-29T03:32:59.820Z",
          "content": "<p>thank you <a href=\"https://www.kaggle.com/tomooinubushi\" target=\"_blank\">@tomooinubushi</a>, that makes sense 🙂 Trying it now!</p>",
          "rawMarkdown": "thank you @tomooinubushi, that makes sense 🙂 Trying it now!",
          "votes": 2
        }
      ]
    },
    {
      "id": 2047771,
      "postDate": "2022-11-29T03:13:10.363Z",
      "content": "<p>Maybe you should name your submission file sample_submission.csv</p>",
      "rawMarkdown": "Maybe you should name your submission file sample_submission.csv",
      "replies": [
        {
          "id": 2047782,
          "postDate": "2022-11-29T03:30:35.547Z",
          "content": "<p>Thank you for the suggestion, <a href=\"https://www.kaggle.com/xiuqi0\" target=\"_blank\">@xiuqi0</a>! That is an interesting thought, I think you might be right! Will give it a shot.</p>\n<p>I looked at solutions from other competitions and think everything works there with the submission file being named <code>submission.csv</code> there. But might be not everything is set up properly here just yet 🙂</p>\n<p>Trying now! </p>",
          "rawMarkdown": "Thank you for the suggestion, @xiuqi0! That is an interesting thought, I think you might be right! Will give it a shot.\n\nI looked at solutions from other competitions and think everything works there with the submission file being named `submission.csv` there. But might be not everything is set up properly here just yet 🙂\n\nTrying now! "
        }
      ]
    },
    {
      "id": 2053295,
      "postDate": "2022-12-03T05:16:05.650Z",
      "content": "<p>Wait, but combination prediction_id = \"patient_id + laterality\" is not unique, it corresponds to different image_id's.<br>\n How did you remove duplicated prediction_id's? </p>",
      "rawMarkdown": "Wait, but combination prediction_id = \"patient_id + laterality\" is not unique, it corresponds to different image_id's.\n How did you remove duplicated prediction_id's? ",
      "replies": [
        {
          "id": 2053348,
          "postDate": "2022-12-03T06:25:28.003Z",
          "content": "<p>there is <code>prediction_id</code> that you can make sure is unique and all is well 🙂</p>",
          "rawMarkdown": "there is `prediction_id` that you can make sure is unique and all is well 🙂",
          "votes": 1
        },
        {
          "id": 2053858,
          "postDate": "2022-12-03T16:56:41.163Z",
          "content": "<p>Great, it works! thanks!!</p>",
          "rawMarkdown": "Great, it works! thanks!!"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2047781,
      "author_name": "YYama",
      "author_url": "",
      "post_date": "2022-11-29T03:30:11.060000",
      "content": "<p>Maybe, but I think it's because your submission.csv has duplicates for prediction_id.</p>",
      "votes": 7,
      "replies": [
        {
          "id": 2047784,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-11-29T03:32:40.687000",
          "content": "<p>Great suggestion, thank you!!!! will give this a go, that makes sense!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2047792,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-11-29T03:42:23.403000",
          "content": "<p>That was it, thank you very much!!!!!!!!! 🙏</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2054782,
      "author_name": "Jirka",
      "author_url": "",
      "post_date": "2022-12-04T13:04:40.187000",
      "content": "<p>you can simply prune duplicates by using max cancer predictions:<br>\n<code>df_preds = df_preds.groupby('prediction_id').max()</code><br>\nsee: <a href=\"https://www.kaggle.com/code/jirkaborovec/mammography-eda-loading-dicom\" target=\"_blank\">https://www.kaggle.com/code/jirkaborovec/mammography-eda-loading-dicom</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 2156447,
          "author_name": "Schroter",
          "author_url": "",
          "post_date": "2023-02-23T10:45:55.360000",
          "content": "<p>Hi, I have the same issue and I used the code above. I have unique prediction ID values and the csv is comma separated. However I am still getting the error message. However, when I submit the sample submission file it works. Would anyone be able to help me in this matter.<br>\nThanks &amp; Best Regards<br>\nAMJS</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2050198,
      "author_name": "Hey24sheep",
      "author_url": "",
      "post_date": "2022-11-30T15:27:21.760000",
      "content": "<p>Did you find any solution? My submission.csv looks like this, why is it still failing?</p>\n<p>prediction_id cancer<br>\n10008_R 0.123<br>\n10008_L 0.123</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2050536,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-11-30T20:20:15.950000",
          "content": "<p>these need to be comma seperated, they need to be a proper csv file</p>\n<p>the solution for me was removing duplicates, duplicated predictions for the same <code>prediction_id</code></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2047780,
      "author_name": "tomoo inubushi",
      "author_url": "",
      "post_date": "2022-11-29T03:29:40.123000",
      "content": "<p>Do you drop duplicated prediction_id?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2047785,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-11-29T03:32:59.820000",
          "content": "<p>thank you <a href=\"https://www.kaggle.com/tomooinubushi\" target=\"_blank\">@tomooinubushi</a>, that makes sense 🙂 Trying it now!</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2047771,
      "author_name": "XIUQI",
      "author_url": "",
      "post_date": "2022-11-29T03:13:10.363000",
      "content": "<p>Maybe you should name your submission file sample_submission.csv</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2047782,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-11-29T03:30:35.547000",
          "content": "<p>Thank you for the suggestion, <a href=\"https://www.kaggle.com/xiuqi0\" target=\"_blank\">@xiuqi0</a>! That is an interesting thought, I think you might be right! Will give it a shot.</p>\n<p>I looked at solutions from other competitions and think everything works there with the submission file being named <code>submission.csv</code> there. But might be not everything is set up properly here just yet 🙂</p>\n<p>Trying now! </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2053295,
      "author_name": "Boris Polishchuk",
      "author_url": "",
      "post_date": "2022-12-03T05:16:05.650000",
      "content": "<p>Wait, but combination prediction_id = \"patient_id + laterality\" is not unique, it corresponds to different image_id's.<br>\n How did you remove duplicated prediction_id's? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2053348,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-12-03T06:25:28.003000",
          "content": "<p>there is <code>prediction_id</code> that you can make sure is unique and all is well 🙂</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2053858,
          "author_name": "Boris Polishchuk",
          "author_url": "",
          "post_date": "2022-12-03T16:56:41.163000",
          "content": "<p>Great, it works! thanks!!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2047781": "Maybe, but I think it's because your submission.csv has duplicates for prediction_id.",
    "2047727": "Hey,\n\nTowards the end of my EDA : [📊 Initial EDA -- a first look at the data 🚀](https://www.kaggle.com/code/radek1/initial-eda-a-first-look-at-the-data/notebook?scriptVersionId=112383297) I am outputting the predictions as follows:\n\n```\nsubmission = pd.DataFrame(data={'prediction_id': test_csv['prediction_id'], 'cancer': np.random.rand(test_csv.shape[0])})\nsubmission.to_csv('submission.csv', index=False)\n```\n\nHowever, regardless of whether I output floats in the range of `[0,1)` as predictions, or all zeros, I am met with `submission scoring error`.\n\nCould I please ask for help in figuring out what might be going wrong there? \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F4d0dc870b8879102e50463007ea302ab%2Fsubmission_scoring_error.png?generation=1669687168872049&alt=media)\n\nAm I doing something wrong or is there some issue with the machinery for scoring?\n\nThank you for your help!",
    "2054782": "you can simply prune duplicates by using max cancer predictions:\n`df_preds = df_preds.groupby('prediction_id').max()`\nsee: https://www.kaggle.com/code/jirkaborovec/mammography-eda-loading-dicom",
    "2050198": "Did you find any solution? My submission.csv looks like this, why is it still failing?\n\nprediction_id cancer\n10008_R 0.123\n10008_L 0.123",
    "2047780": "Do you drop duplicated prediction_id?",
    "2047771": "Maybe you should name your submission file sample_submission.csv",
    "2053295": "Wait, but combination prediction_id = \"patient_id + laterality\" is not unique, it corresponds to different image_id's.\n How did you remove duplicated prediction_id's? "
  }
}