{
  "id": 381427,
  "title": "Competition submission failed",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/381427",
  "author_name": "Julian Macnamara",
  "post_date": "2023-01-26T16:27:44.215000",
  "votes": 1,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Can anybody shed any light on why my submission is failing when the notebook I'm trying to submit is executing correctly?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5387387%2F8dcffba636054032fb16e91eaf0401af%2FRSNA%202023-01-26.png?generation=1674749766949063&amp;alt=media\" alt=\"\"></p>\n<p>Contents of kaggle/working are:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5387387%2F5e3578f4f52ee37f9761e833bb05a944%2Fkaggle%202023-01-26.png?generation=1674749822021649&amp;alt=media\" alt=\"\"></p>\n<p>My submission.csv file looks like this</p>\n<p>prediction_id    cancer<br>\n10008_L    0.030661125<br>\n10008_R    0.011916713</p>\n<p>The notebook URL is <a href=\"https://www.kaggle.com/code/julianmacnamara/rsna-screening-mammography-breast-cancer-detection\" target=\"_blank\">https://www.kaggle.com/code/julianmacnamara/rsna-screening-mammography-breast-cancer-detection</a></p>\n<p>Many thanks in advance for any help anybody can offer</p>",
  "messages": [
    {
      "id": 2118749,
      "postDate": "2023-01-28T08:40:03.213Z",
      "content": "<p>I revisited the code and realised I'd inadvertently made 'prediction_id' an index and I should have included \"as_index=False\" in my .groupby statement as shown below</p>\n<p>`# Select only the 'prediction_id' and 'cancer' columns<br>\nresulting_df = merged_df[['prediction_id', 'cancer']]</p>\n<p>resulting_df = resulting_df.groupby('prediction_id', as_index=False).mean()</p>\n<p>resulting_df = resulting_df.sort_index()`</p>\n<p>However, when I ran this I still got a \"Submission score error\"</p>\n<p>I was able to resolve this by following Abdoue's suggestion below</p>",
      "rawMarkdown": "I revisited the code and realised I'd inadvertently made 'prediction_id' an index and I should have included \"as_index=False\" in my .groupby statement as shown below\n\n`# Select only the 'prediction_id' and 'cancer' columns\nresulting_df = merged_df[['prediction_id', 'cancer']]\n\nresulting_df = resulting_df.groupby('prediction_id', as_index=False).mean()\n\nresulting_df = resulting_df.sort_index()`\n\nHowever, when I ran this I still got a \"Submission score error\"\n\nI was able to resolve this by following Abdoue's suggestion below",
      "votes": 1
    },
    {
      "id": 2116664,
      "postDate": "2023-01-26T16:27:44.217Z",
      "content": "<p>Can anybody shed any light on why my submission is failing when the notebook I'm trying to submit is executing correctly?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5387387%2F8dcffba636054032fb16e91eaf0401af%2FRSNA%202023-01-26.png?generation=1674749766949063&amp;alt=media\" alt=\"\"></p>\n<p>Contents of kaggle/working are:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5387387%2F5e3578f4f52ee37f9761e833bb05a944%2Fkaggle%202023-01-26.png?generation=1674749822021649&amp;alt=media\" alt=\"\"></p>\n<p>My submission.csv file looks like this</p>\n<p>prediction_id    cancer<br>\n10008_L    0.030661125<br>\n10008_R    0.011916713</p>\n<p>The notebook URL is <a href=\"https://www.kaggle.com/code/julianmacnamara/rsna-screening-mammography-breast-cancer-detection\" target=\"_blank\">https://www.kaggle.com/code/julianmacnamara/rsna-screening-mammography-breast-cancer-detection</a></p>\n<p>Many thanks in advance for any help anybody can offer</p>",
      "rawMarkdown": "Can anybody shed any light on why my submission is failing when the notebook I'm trying to submit is executing correctly?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5387387%2F8dcffba636054032fb16e91eaf0401af%2FRSNA%202023-01-26.png?generation=1674749766949063&alt=media)\n\nContents of kaggle/working are:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5387387%2F5e3578f4f52ee37f9761e833bb05a944%2Fkaggle%202023-01-26.png?generation=1674749822021649&alt=media)\n\nMy submission.csv file looks like this\n\nprediction_id\tcancer\n10008_L\t0.030661125\n10008_R\t0.011916713\n\nThe notebook URL is https://www.kaggle.com/code/julianmacnamara/rsna-screening-mammography-breast-cancer-detection\n\nMany thanks in advance for any help anybody can offer\n\n",
      "votes": 1
    },
    {
      "id": 2117480,
      "postDate": "2023-01-27T10:01:45.227Z",
      "content": "<p>Many thanks to everbody for their kind words and suggestions. I've tried them all as well as checking <a href=\"https://www.kaggle.com/code-competition-debugging\" target=\"_blank\">this</a> but all to no avail</p>\n<p>\"Insanity is doing the same thing over and over again and expecting different results.\" may be mis-attributed to Einstein but I still find it true in this case. However I've run out of things to try so am going to retire gracefully which is a shame. I was never going to win but, as a novice, I'd like to have seen how my ideas stacked up against other competitors</p>",
      "rawMarkdown": "Many thanks to everbody for their kind words and suggestions. I've tried them all as well as checking [this](https://www.kaggle.com/code-competition-debugging) but all to no avail\n\n \"Insanity is doing the same thing over and over again and expecting different results.\" may be mis-attributed to Einstein but I still find it true in this case. However I've run out of things to try so am going to retire gracefully which is a shame. I was never going to win but, as a novice, I'd like to have seen how my ideas stacked up against other competitors",
      "votes": 2
    },
    {
      "id": 2117960,
      "postDate": "2023-01-27T17:02:33Z",
      "content": "<p>What was the error message? </p>\n<p>If it was \"Submission score error\" I had the same problem. Solved it by mapping my predictions onto a dataframe created directly from the provided test_csv. This ensures that all values of 'prediction_id' are there, if for whatever reason a prediction doesn't make it through the pipeline. </p>\n<pre><code>sub_df = pd.concat([prediction_df, full_df], join=) \n\nsub_df = sub_df.sort_values(by=[]).reset_index(drop=) \n\nsub_df = sub_df.fillna() \n</code></pre>",
      "rawMarkdown": "What was the error message? \n\nIf it was \"Submission score error\" I had the same problem. Solved it by mapping my predictions onto a dataframe created directly from the provided test_csv. This ensures that all values of 'prediction_id' are there, if for whatever reason a prediction doesn't make it through the pipeline. \n\n```python\nsub_df = pd.concat([prediction_df, full_df], join='outer') # This eliminates any missing prediction ids\n\nsub_df = sub_df.sort_values(by=['prediction_id']).reset_index(drop=True) # Reordering by prediction id as required\n\nsub_df = sub_df.fillna(0) # Fill missing prediction values\n```",
      "replies": [
        {
          "id": 2118739,
          "postDate": "2023-01-28T08:33:36.853Z",
          "content": "<p>Many thanks Abdou<br>\nI don't understand why this should work but it does </p>",
          "rawMarkdown": "Many thanks Abdou\nI don't understand why this should work but it does ",
          "votes": 1
        }
      ]
    },
    {
      "id": 2117104,
      "postDate": "2023-01-27T02:22:03.723Z",
      "content": "<p>Do you have other files except submission.csv in the output folder? If so, try to make sure only submission.csv in that folder.</p>",
      "rawMarkdown": "Do you have other files except submission.csv in the output folder? If so, try to make sure only submission.csv in that folder."
    },
    {
      "id": 2117096,
      "postDate": "2023-01-27T01:51:37.130Z",
      "content": "<p>Check on submissions (at the right of \"Leaderboard\"/Rules/Team), there is a little more information about why it has failed.</p>\n<p>Usually messages:<br>\n-Submission score error<br>\n-Notebook timeout<br>\n-Notebook run out of memory<br>\n-Notebook Threw exception</p>\n<ul>\n<li>Not sure, but I wouldn't use \"patient_id\" as an index, considered resetting the index before generating the CSV file.</li>\n</ul>",
      "rawMarkdown": "Check on submissions (at the right of \"Leaderboard\"/Rules/Team), there is a little more information about why it has failed.\n\nUsually messages:\n-Submission score error\n-Notebook timeout\n-Notebook run out of memory\n-Notebook Threw exception\n\n* Not sure, but I wouldn't use \"patient_id\" as an index, considered resetting the index before generating the CSV file.",
      "replies": [
        {
          "id": 2118953,
          "postDate": "2023-01-28T12:06:33.503Z",
          "content": "<p>Hi Alfredo</p>\n<p>Thanks for making me improve my understanding on indexing</p>\n<p>All the best</p>\n<p>Julian</p>",
          "rawMarkdown": "Hi Alfredo\n\nThanks for making me improve my understanding on indexing\n\nAll the best\n\nJulian"
        }
      ]
    },
    {
      "id": 2116943,
      "postDate": "2023-01-26T20:37:21.427Z",
      "content": "<p>I hope anyone could answer that and help you with the submission.</p>",
      "rawMarkdown": "I hope anyone could answer that and help you with the submission.",
      "replies": [
        {
          "id": 2116952,
          "postDate": "2023-01-26T20:46:34.087Z",
          "content": "<p>Thank you. Much appreciated </p>\n<p>All the best </p>\n<p>Julian </p>",
          "rawMarkdown": "Thank you. Much appreciated \n\nAll the best \n\nJulian ",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2118749,
      "author_name": "Julian Macnamara",
      "author_url": "",
      "post_date": "2023-01-28T08:40:03.213000",
      "content": "<p>I revisited the code and realised I'd inadvertently made 'prediction_id' an index and I should have included \"as_index=False\" in my .groupby statement as shown below</p>\n<p>`# Select only the 'prediction_id' and 'cancer' columns<br>\nresulting_df = merged_df[['prediction_id', 'cancer']]</p>\n<p>resulting_df = resulting_df.groupby('prediction_id', as_index=False).mean()</p>\n<p>resulting_df = resulting_df.sort_index()`</p>\n<p>However, when I ran this I still got a \"Submission score error\"</p>\n<p>I was able to resolve this by following Abdoue's suggestion below</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2117480,
      "author_name": "Julian Macnamara",
      "author_url": "",
      "post_date": "2023-01-27T10:01:45.227000",
      "content": "<p>Many thanks to everbody for their kind words and suggestions. I've tried them all as well as checking <a href=\"https://www.kaggle.com/code-competition-debugging\" target=\"_blank\">this</a> but all to no avail</p>\n<p>\"Insanity is doing the same thing over and over again and expecting different results.\" may be mis-attributed to Einstein but I still find it true in this case. However I've run out of things to try so am going to retire gracefully which is a shame. I was never going to win but, as a novice, I'd like to have seen how my ideas stacked up against other competitors</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2117960,
      "author_name": "AbdouE",
      "author_url": "",
      "post_date": "2023-01-27T17:02:33",
      "content": "<p>What was the error message? </p>\n<p>If it was \"Submission score error\" I had the same problem. Solved it by mapping my predictions onto a dataframe created directly from the provided test_csv. This ensures that all values of 'prediction_id' are there, if for whatever reason a prediction doesn't make it through the pipeline. </p>\n<pre><code>sub_df = pd.concat([prediction_df, full_df], join=) \n\nsub_df = sub_df.sort_values(by=[]).reset_index(drop=) \n\nsub_df = sub_df.fillna() \n</code></pre>",
      "votes": 0,
      "replies": [
        {
          "id": 2118739,
          "author_name": "Julian Macnamara",
          "author_url": "",
          "post_date": "2023-01-28T08:33:36.853000",
          "content": "<p>Many thanks Abdou<br>\nI don't understand why this should work but it does </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2117104,
      "author_name": "Jijie Li",
      "author_url": "",
      "post_date": "2023-01-27T02:22:03.723000",
      "content": "<p>Do you have other files except submission.csv in the output folder? If so, try to make sure only submission.csv in that folder.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2117096,
      "author_name": "Alfredo Maussa",
      "author_url": "",
      "post_date": "2023-01-27T01:51:37.130000",
      "content": "<p>Check on submissions (at the right of \"Leaderboard\"/Rules/Team), there is a little more information about why it has failed.</p>\n<p>Usually messages:<br>\n-Submission score error<br>\n-Notebook timeout<br>\n-Notebook run out of memory<br>\n-Notebook Threw exception</p>\n<ul>\n<li>Not sure, but I wouldn't use \"patient_id\" as an index, considered resetting the index before generating the CSV file.</li>\n</ul>",
      "votes": 0,
      "replies": [
        {
          "id": 2118953,
          "author_name": "Julian Macnamara",
          "author_url": "",
          "post_date": "2023-01-28T12:06:33.503000",
          "content": "<p>Hi Alfredo</p>\n<p>Thanks for making me improve my understanding on indexing</p>\n<p>All the best</p>\n<p>Julian</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2116943,
      "author_name": "Marília Prata",
      "author_url": "",
      "post_date": "2023-01-26T20:37:21.427000",
      "content": "<p>I hope anyone could answer that and help you with the submission.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2116952,
          "author_name": "Julian Macnamara",
          "author_url": "",
          "post_date": "2023-01-26T20:46:34.087000",
          "content": "<p>Thank you. Much appreciated </p>\n<p>All the best </p>\n<p>Julian </p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2118749": "I revisited the code and realised I'd inadvertently made 'prediction_id' an index and I should have included \"as_index=False\" in my .groupby statement as shown below\n\n`# Select only the 'prediction_id' and 'cancer' columns\nresulting_df = merged_df[['prediction_id', 'cancer']]\n\nresulting_df = resulting_df.groupby('prediction_id', as_index=False).mean()\n\nresulting_df = resulting_df.sort_index()`\n\nHowever, when I ran this I still got a \"Submission score error\"\n\nI was able to resolve this by following Abdoue's suggestion below",
    "2116664": "Can anybody shed any light on why my submission is failing when the notebook I'm trying to submit is executing correctly?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5387387%2F8dcffba636054032fb16e91eaf0401af%2FRSNA%202023-01-26.png?generation=1674749766949063&alt=media)\n\nContents of kaggle/working are:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5387387%2F5e3578f4f52ee37f9761e833bb05a944%2Fkaggle%202023-01-26.png?generation=1674749822021649&alt=media)\n\nMy submission.csv file looks like this\n\nprediction_id\tcancer\n10008_L\t0.030661125\n10008_R\t0.011916713\n\nThe notebook URL is https://www.kaggle.com/code/julianmacnamara/rsna-screening-mammography-breast-cancer-detection\n\nMany thanks in advance for any help anybody can offer\n\n",
    "2117480": "Many thanks to everbody for their kind words and suggestions. I've tried them all as well as checking [this](https://www.kaggle.com/code-competition-debugging) but all to no avail\n\n \"Insanity is doing the same thing over and over again and expecting different results.\" may be mis-attributed to Einstein but I still find it true in this case. However I've run out of things to try so am going to retire gracefully which is a shame. I was never going to win but, as a novice, I'd like to have seen how my ideas stacked up against other competitors",
    "2117960": "What was the error message? \n\nIf it was \"Submission score error\" I had the same problem. Solved it by mapping my predictions onto a dataframe created directly from the provided test_csv. This ensures that all values of 'prediction_id' are there, if for whatever reason a prediction doesn't make it through the pipeline. \n\n```python\nsub_df = pd.concat([prediction_df, full_df], join='outer') # This eliminates any missing prediction ids\n\nsub_df = sub_df.sort_values(by=['prediction_id']).reset_index(drop=True) # Reordering by prediction id as required\n\nsub_df = sub_df.fillna(0) # Fill missing prediction values\n```",
    "2117104": "Do you have other files except submission.csv in the output folder? If so, try to make sure only submission.csv in that folder.",
    "2117096": "Check on submissions (at the right of \"Leaderboard\"/Rules/Team), there is a little more information about why it has failed.\n\nUsually messages:\n-Submission score error\n-Notebook timeout\n-Notebook run out of memory\n-Notebook Threw exception\n\n* Not sure, but I wouldn't use \"patient_id\" as an index, considered resetting the index before generating the CSV file.",
    "2116943": "I hope anyone could answer that and help you with the submission."
  }
}