{
  "id": 369603,
  "title": "Why Submission Scoring Error?",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/369603",
  "author_name": "Hey24sheep",
  "post_date": "2022-11-30T16:38:13.127000",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Why am I getting \"Submission Scoring Error\". What is wrong with my submission file? Can anyone please help me here? </p>\n<p>Note : I am removing duplicates</p>\n<p>My submission file looks like this. <br>\n <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9062237%2F90bcc8167c3420d25588cdadf525d2d7%2FCapture4.PNG?generation=1669825857366616&amp;alt=media\" alt=\"\"></p>\n<p>Original \"SampleSubmission\" has this info<br>\n<code>&lt;class 'pandas.core.frame.DataFrame'&gt;\nRangeIndex: 2 entries, 0 to 1\nData columns (total 2 columns):\n #   Column         Non-Null Count  Dtype  \n---  ------         --------------  -----  \n 0   prediction_id  2 non-null      object \n 1   cancer         2 non-null      float64\ndtypes: float64(1), object(1)\nmemory usage: 160.0+ bytes</code></p>\n<p>My Submission file has this info<br>\n<code>&lt;class 'pandas.core.frame.DataFrame'&gt;\nRangeIndex: 2 entries, 0 to 1\nData columns (total 2 columns):\n #   Column         Non-Null Count  Dtype  \n---  ------         --------------  -----  \n 0   prediction_id  2 non-null      object \n 1   cancer         2 non-null      float64\ndtypes: float64(1), object(1)\nmemory usage: 160.0+ bytes</code></p>",
  "messages": [
    {
      "id": 2050742,
      "postDate": "2022-12-01T00:28:48.030Z",
      "content": "<p>This should work</p>\n<pre><code>csv_file = '/kaggle/input/rsna-breast-cancer-detection/test.csv'\ntest_df = pd.read_csv(csv_file)\ntest_df.loc[:, 'cancer'] = np.random.uniform(0,1,len(test_df)) #  dummy probability values\nsubmit_df = test_df[['prediction_id', 'cancer']]\nsubmit_df = submit_df.groupby('prediction_id').mean()  #dummy aggregation method\nsubmit_df = submit_df.sort_index()\nsubmit_df.to_csv('submission.csv',index=True)\n\n\nnote that\ntest_df.loc[:, 'prediction_id'] = test_df.patient_id.astype(str) + '_' + test_df.laterality\n</code></pre>\n<p><br>\nNote: this is for illustration only. your model should take in multiple images (e.g. same/cross view,  same/cross-laterality  attention) and produce a single prediction per  (patient, laterality)</p>",
      "rawMarkdown": "\nThis should work\n```\n \ncsv_file = '/kaggle/input/rsna-breast-cancer-detection/test.csv'\ntest_df = pd.read_csv(csv_file)\ntest_df.loc[:, 'cancer'] = np.random.uniform(0,1,len(test_df)) #  dummy probability values\n\nsubmit_df = test_df[['prediction_id', 'cancer']]\nsubmit_df = submit_df.groupby('prediction_id').mean()  #dummy aggregation method\nsubmit_df = submit_df.sort_index()\nsubmit_df.to_csv('submission.csv',index=True)\n   \n   \nnote that\n test_df.loc[:, 'prediction_id'] = test_df.patient_id.astype(str) + '_' + test_df.laterality\n\n```  \nNote: this is for illustration only. your model should take in multiple images (e.g. same/cross view,  same/cross-laterality  attention) and produce a single prediction per  (patient, laterality)",
      "votes": 1,
      "replies": [
        {
          "id": 2051523,
          "postDate": "2022-12-01T13:35:35.007Z",
          "content": "<p>Thanks, this helped me find the issue. I was hardcoding public test patient in my submission df as the prediction ID. I fixed that thanks to your code.</p>",
          "rawMarkdown": "Thanks, this helped me find the issue. I was hardcoding public test patient in my submission df as the prediction ID. I fixed that thanks to your code."
        },
        {
          "id": 2078573,
          "postDate": "2022-12-28T12:39:54.427Z",
          "content": "<p><a href=\"https://www.kaggle.com/code/alaamadi/first-submission-with-smaller-dataset\" target=\"_blank\">https://www.kaggle.com/code/alaamadi/first-submission-with-smaller-dataset</a><br>\nmay you please check mine?</p>",
          "rawMarkdown": "https://www.kaggle.com/code/alaamadi/first-submission-with-smaller-dataset\nmay you please check mine?\n"
        },
        {
          "id": 2121074,
          "postDate": "2023-01-30T03:40:59.950Z",
          "content": "<p>Why have you done \"<em>index=True</em>\" when sample submission does not have index column in the file?</p>",
          "rawMarkdown": "Why have you done \"*index=True*\" when sample submission does not have index column in the file?\n"
        }
      ]
    },
    {
      "id": 2050307,
      "postDate": "2022-11-30T16:38:13.127Z",
      "content": "<p>Why am I getting \"Submission Scoring Error\". What is wrong with my submission file? Can anyone please help me here? </p>\n<p>Note : I am removing duplicates</p>\n<p>My submission file looks like this. <br>\n <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9062237%2F90bcc8167c3420d25588cdadf525d2d7%2FCapture4.PNG?generation=1669825857366616&amp;alt=media\" alt=\"\"></p>\n<p>Original \"SampleSubmission\" has this info<br>\n<code>&lt;class 'pandas.core.frame.DataFrame'&gt;\nRangeIndex: 2 entries, 0 to 1\nData columns (total 2 columns):\n #   Column         Non-Null Count  Dtype  \n---  ------         --------------  -----  \n 0   prediction_id  2 non-null      object \n 1   cancer         2 non-null      float64\ndtypes: float64(1), object(1)\nmemory usage: 160.0+ bytes</code></p>\n<p>My Submission file has this info<br>\n<code>&lt;class 'pandas.core.frame.DataFrame'&gt;\nRangeIndex: 2 entries, 0 to 1\nData columns (total 2 columns):\n #   Column         Non-Null Count  Dtype  \n---  ------         --------------  -----  \n 0   prediction_id  2 non-null      object \n 1   cancer         2 non-null      float64\ndtypes: float64(1), object(1)\nmemory usage: 160.0+ bytes</code></p>",
      "rawMarkdown": "Why am I getting \"Submission Scoring Error\". What is wrong with my submission file? Can anyone please help me here? \n\nNote : I am removing duplicates\n\nMy submission file looks like this. \n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9062237%2F90bcc8167c3420d25588cdadf525d2d7%2FCapture4.PNG?generation=1669825857366616&alt=media)\n\nOriginal \"SampleSubmission\" has this info\n`<class 'pandas.core.frame.DataFrame'>\nRangeIndex: 2 entries, 0 to 1\nData columns (total 2 columns):\n #   Column         Non-Null Count  Dtype  \n---  ------         --------------  -----  \n 0   prediction_id  2 non-null      object \n 1   cancer         2 non-null      float64\ndtypes: float64(1), object(1)\nmemory usage: 160.0+ bytes`\n\nMy Submission file has this info\n`<class 'pandas.core.frame.DataFrame'>\nRangeIndex: 2 entries, 0 to 1\nData columns (total 2 columns):\n #   Column         Non-Null Count  Dtype  \n---  ------         --------------  -----  \n 0   prediction_id  2 non-null      object \n 1   cancer         2 non-null      float64\ndtypes: float64(1), object(1)\nmemory usage: 160.0+ bytes`",
      "votes": 1
    },
    {
      "id": 2050576,
      "postDate": "2022-11-30T21:06:32.633Z",
      "content": "<p>It will accelerate the help if you share your notebook or a simple version that fails in the same way.I</p>",
      "rawMarkdown": "It will accelerate the help if you share your notebook or a simple version that fails in the same way.I"
    }
  ],
  "comments": [
    {
      "id": 2050742,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-12-01T00:28:48.030000",
      "content": "<p>This should work</p>\n<pre><code>csv_file = '/kaggle/input/rsna-breast-cancer-detection/test.csv'\ntest_df = pd.read_csv(csv_file)\ntest_df.loc[:, 'cancer'] = np.random.uniform(0,1,len(test_df)) #  dummy probability values\nsubmit_df = test_df[['prediction_id', 'cancer']]\nsubmit_df = submit_df.groupby('prediction_id').mean()  #dummy aggregation method\nsubmit_df = submit_df.sort_index()\nsubmit_df.to_csv('submission.csv',index=True)\n\n\nnote that\ntest_df.loc[:, 'prediction_id'] = test_df.patient_id.astype(str) + '_' + test_df.laterality\n</code></pre>\n<p><br>\nNote: this is for illustration only. your model should take in multiple images (e.g. same/cross view,  same/cross-laterality  attention) and produce a single prediction per  (patient, laterality)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2051523,
          "author_name": "Hey24sheep",
          "author_url": "",
          "post_date": "2022-12-01T13:35:35.007000",
          "content": "<p>Thanks, this helped me find the issue. I was hardcoding public test patient in my submission df as the prediction ID. I fixed that thanks to your code.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2078573,
          "author_name": "Alaa Madi",
          "author_url": "",
          "post_date": "2022-12-28T12:39:54.427000",
          "content": "<p><a href=\"https://www.kaggle.com/code/alaamadi/first-submission-with-smaller-dataset\" target=\"_blank\">https://www.kaggle.com/code/alaamadi/first-submission-with-smaller-dataset</a><br>\nmay you please check mine?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2121074,
          "author_name": "Rajat Dhawan",
          "author_url": "",
          "post_date": "2023-01-30T03:40:59.950000",
          "content": "<p>Why have you done \"<em>index=True</em>\" when sample submission does not have index column in the file?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2050576,
      "author_name": "Darien Schettler",
      "author_url": "",
      "post_date": "2022-11-30T21:06:32.633000",
      "content": "<p>It will accelerate the help if you share your notebook or a simple version that fails in the same way.I</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2050742": "\nThis should work\n```\n \ncsv_file = '/kaggle/input/rsna-breast-cancer-detection/test.csv'\ntest_df = pd.read_csv(csv_file)\ntest_df.loc[:, 'cancer'] = np.random.uniform(0,1,len(test_df)) #  dummy probability values\n\nsubmit_df = test_df[['prediction_id', 'cancer']]\nsubmit_df = submit_df.groupby('prediction_id').mean()  #dummy aggregation method\nsubmit_df = submit_df.sort_index()\nsubmit_df.to_csv('submission.csv',index=True)\n   \n   \nnote that\n test_df.loc[:, 'prediction_id'] = test_df.patient_id.astype(str) + '_' + test_df.laterality\n\n```  \nNote: this is for illustration only. your model should take in multiple images (e.g. same/cross view,  same/cross-laterality  attention) and produce a single prediction per  (patient, laterality)",
    "2050307": "Why am I getting \"Submission Scoring Error\". What is wrong with my submission file? Can anyone please help me here? \n\nNote : I am removing duplicates\n\nMy submission file looks like this. \n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9062237%2F90bcc8167c3420d25588cdadf525d2d7%2FCapture4.PNG?generation=1669825857366616&alt=media)\n\nOriginal \"SampleSubmission\" has this info\n`<class 'pandas.core.frame.DataFrame'>\nRangeIndex: 2 entries, 0 to 1\nData columns (total 2 columns):\n #   Column         Non-Null Count  Dtype  \n---  ------         --------------  -----  \n 0   prediction_id  2 non-null      object \n 1   cancer         2 non-null      float64\ndtypes: float64(1), object(1)\nmemory usage: 160.0+ bytes`\n\nMy Submission file has this info\n`<class 'pandas.core.frame.DataFrame'>\nRangeIndex: 2 entries, 0 to 1\nData columns (total 2 columns):\n #   Column         Non-Null Count  Dtype  \n---  ------         --------------  -----  \n 0   prediction_id  2 non-null      object \n 1   cancer         2 non-null      float64\ndtypes: float64(1), object(1)\nmemory usage: 160.0+ bytes`",
    "2050576": "It will accelerate the help if you share your notebook or a simple version that fails in the same way.I"
  }
}