{
  "id": 441077,
  "title": "R submission error: “Submission scoring error\"",
  "url": "/competitions/rsna-2023-abdominal-trauma-detection/discussion/441077",
  "author_name": "RickPack",
  "post_date": "2023-09-17T13:53:56.455000",
  "votes": 1,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Has anyone been able to submit successfully using an R notebook? I keep seeing \"Submission scoring error\".</p>\n<p>I used str() to examine the structure of the submission vs. the sample submissions, and others I have submitted, and all looks fine (e.g., 3 rows, 14 numeric columns).</p>\n<p>Notebook visible at <a href=\"https://www.kaggle.com/rickpack/r-kerascv-wtmean-submissions-averaged\" target=\"_blank\">https://www.kaggle.com/rickpack/r-kerascv-wtmean-submissions-averaged</a></p>",
  "messages": [
    {
      "id": 2446803,
      "postDate": "2023-09-19T16:13:51.510Z",
      "content": "<p>I have almost the same issue. If I hardcode each submission with the set in the <code>sample_submission.cvs</code> file</p>\n<p>[, 0.5,0.5,0.5,0.5,0.3333333333333333,0.3333333333333333,0.3333333333333333,0.3333333333333333,0.3333333333333333,0.3333333333333333,0.3333333333333333,0.3333333333333333,0.3333333333333333]</p>\n<p>I get green. But if I hardcode to the following, I got \"Submission scoring error\":</p>\n<p>[, 1.0,1.0,1.0,1.0,1.0,1.0,1.0,1.0,1.0,1.0,1.0,1.0,1.0]</p>\n<p><a href=\"https://www.kaggle.com/artberz\" target=\"_blank\">@artberz</a> can you please share how you set your arrays and then you convert to csv? I am really stuck!</p>",
      "rawMarkdown": "I have almost the same issue. If I hardcode each submission with the set in the `sample_submission.cvs` file\n\n[<patient_id>, 0.5,0.5,0.5,0.5,0.3333333333333333,0.3333333333333333,0.3333333333333333,0.3333333333333333,0.3333333333333333,0.3333333333333333,0.3333333333333333,0.3333333333333333,0.3333333333333333]\n\nI get green. But if I hardcode to the following, I got \"Submission scoring error\":\n\n[<patient_id>, 1.0,1.0,1.0,1.0,1.0,1.0,1.0,1.0,1.0,1.0,1.0,1.0,1.0]\n\n@artberz can you please share how you set your arrays and then you convert to csv? I am really stuck!",
      "votes": 1,
      "replies": [
        {
          "id": 2446822,
          "postDate": "2023-09-19T16:25:05.987Z",
          "content": "<p>I think that's likely a separate issue. If you push 1.0 as the prediction for every single column, you're saying that there's a 100% likelyhood that there's a healthy (kidney_healthy = 1.0 = 100%) person with a high_grade kidney injury (kidney_high = 1.0 = 100%) and a low_grade kidney injury (kidney_low = 1.0 = 100%). That's impossible, since a healthy kidney cannot have an injury, and you cannot have both a low and high grade kidney injury, as far as I have seen in the data. In any case, it follows the same pattern with all the columns, you cannot have bowel_healthy and bowel_injury simultaneously. </p>",
          "rawMarkdown": "I think that's likely a separate issue. If you push 1.0 as the prediction for every single column, you're saying that there's a 100% likelyhood that there's a healthy (kidney_healthy = 1.0 = 100%) person with a high_grade kidney injury (kidney_high = 1.0 = 100%) and a low_grade kidney injury (kidney_low = 1.0 = 100%). That's impossible, since a healthy kidney cannot have an injury, and you cannot have both a low and high grade kidney injury, as far as I have seen in the data. In any case, it follows the same pattern with all the columns, you cannot have bowel_healthy and bowel_injury simultaneously. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 2443050,
      "postDate": "2023-09-17T13:53:56.457Z",
      "content": "<p>Has anyone been able to submit successfully using an R notebook? I keep seeing \"Submission scoring error\".</p>\n<p>I used str() to examine the structure of the submission vs. the sample submissions, and others I have submitted, and all looks fine (e.g., 3 rows, 14 numeric columns).</p>\n<p>Notebook visible at <a href=\"https://www.kaggle.com/rickpack/r-kerascv-wtmean-submissions-averaged\" target=\"_blank\">https://www.kaggle.com/rickpack/r-kerascv-wtmean-submissions-averaged</a></p>",
      "rawMarkdown": "Has anyone been able to submit successfully using an R notebook? I keep seeing \"Submission scoring error\".\n\nI used str() to examine the structure of the submission vs. the sample submissions, and others I have submitted, and all looks fine (e.g., 3 rows, 14 numeric columns).\n\nNotebook visible at https://www.kaggle.com/rickpack/r-kerascv-wtmean-submissions-averaged",
      "votes": 1
    },
    {
      "id": 2445532,
      "postDate": "2023-09-18T22:59:09.370Z",
      "content": "<p>From a quick glance over, it isn't quite clear to me where your test patient_ids are getting injected into the submission.csv file. If you aren't getting them from the test_images part of the data, that might be an issue since they won't get pulled properly. I made the mistake of submitting just the baseline sample submission (just reading the file and submitting it) instead of pulling in the list of test_ids from test images, which for the sample test set is 3 patient_ids long, but for the hidden test set is over 1000 patient_ids long. If that isn't the issue, I've found it helpful to break down each part of the process until I understand it completely, and that usually gets me the answer of why things aren't working. Good luck debugging!</p>",
      "rawMarkdown": "From a quick glance over, it isn't quite clear to me where your test patient_ids are getting injected into the submission.csv file. If you aren't getting them from the test_images part of the data, that might be an issue since they won't get pulled properly. I made the mistake of submitting just the baseline sample submission (just reading the file and submitting it) instead of pulling in the list of test_ids from test images, which for the sample test set is 3 patient_ids long, but for the hidden test set is over 1000 patient_ids long. If that isn't the issue, I've found it helpful to break down each part of the process until I understand it completely, and that usually gets me the answer of why things aren't working. Good luck debugging!",
      "votes": 2,
      "replies": [
        {
          "id": 2445623,
          "postDate": "2023-09-19T01:54:37.940Z",
          "content": "<p>Thank you, are you saying that the submission should have more rows than 3? I have the patient_id column, copied from the sample submission CSV.</p>",
          "rawMarkdown": "Thank you, are you saying that the submission should have more rows than 3? I have the patient_id column, copied from the sample submission CSV.",
          "votes": 1,
          "replies": [
            {
              "id": 2445739,
              "postDate": "2023-09-19T03:58:50.403Z",
              "content": "<p>Nope, the sample submission will always be 3 patient ids, but it matters how your code obtains them. If you're just taking the sample csv file that they give you and plugging numbers in, every time you run your code, it will just go to the sample submission and put numbers in there. Whereas, if you code in a way for the code to:</p>\n<ul>\n<li>Go into the test_images folder under competition data</li>\n<li>Grab patient id</li>\n<li>do other stuff, like plug in those numbers, format, and save as submission.csv<br>\nIt will always work, because it goes the same path when you submit. When you submit, there's the main block of code that runs which is your main notebook, the stuff you can see, and then if it finishes running, test_images gets switched out, where there's roughly 1000 patient_ids. If you have the pathway built, it works just as well, but if you don't, you're just getting scored on those 3, because that's all that your notebook is coded to be able to see. I'm not skilled enough in R programming language to really fully understand your notebook or run it, but I hope that makes sense. Let me know if you have any other questions. </li>\n</ul>",
              "rawMarkdown": "Nope, the sample submission will always be 3 patient ids, but it matters how your code obtains them. If you're just taking the sample csv file that they give you and plugging numbers in, every time you run your code, it will just go to the sample submission and put numbers in there. Whereas, if you code in a way for the code to:\n- Go into the test_images folder under competition data\n- Grab patient id\n- do other stuff, like plug in those numbers, format, and save as submission.csv\nIt will always work, because it goes the same path when you submit. When you submit, there's the main block of code that runs which is your main notebook, the stuff you can see, and then if it finishes running, test_images gets switched out, where there's roughly 1000 patient_ids. If you have the pathway built, it works just as well, but if you don't, you're just getting scored on those 3, because that's all that your notebook is coded to be able to see. I'm not skilled enough in R programming language to really fully understand your notebook or run it, but I hope that makes sense. Let me know if you have any other questions. ",
              "votes": 1
            },
            {
              "id": 2456580,
              "postDate": "2023-09-26T09:47:39.610Z",
              "content": "<p>Art, this pointed me to the solution and helped me understand the competition better. Thank you for your investment of time!</p>",
              "rawMarkdown": "Art, this pointed me to the solution and helped me understand the competition better. Thank you for your investment of time!",
              "votes": 1
            },
            {
              "id": 2457000,
              "postDate": "2023-09-26T14:43:58.647Z",
              "content": "<p>Glad I could help :) .</p>",
              "rawMarkdown": "Glad I could help :) .",
              "votes": 1
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2446803,
      "author_name": "Byzantine Monk",
      "author_url": "",
      "post_date": "2023-09-19T16:13:51.510000",
      "content": "<p>I have almost the same issue. If I hardcode each submission with the set in the <code>sample_submission.cvs</code> file</p>\n<p>[, 0.5,0.5,0.5,0.5,0.3333333333333333,0.3333333333333333,0.3333333333333333,0.3333333333333333,0.3333333333333333,0.3333333333333333,0.3333333333333333,0.3333333333333333,0.3333333333333333]</p>\n<p>I get green. But if I hardcode to the following, I got \"Submission scoring error\":</p>\n<p>[, 1.0,1.0,1.0,1.0,1.0,1.0,1.0,1.0,1.0,1.0,1.0,1.0,1.0]</p>\n<p><a href=\"https://www.kaggle.com/artberz\" target=\"_blank\">@artberz</a> can you please share how you set your arrays and then you convert to csv? I am really stuck!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2446822,
          "author_name": "Art. Berz.",
          "author_url": "",
          "post_date": "2023-09-19T16:25:05.987000",
          "content": "<p>I think that's likely a separate issue. If you push 1.0 as the prediction for every single column, you're saying that there's a 100% likelyhood that there's a healthy (kidney_healthy = 1.0 = 100%) person with a high_grade kidney injury (kidney_high = 1.0 = 100%) and a low_grade kidney injury (kidney_low = 1.0 = 100%). That's impossible, since a healthy kidney cannot have an injury, and you cannot have both a low and high grade kidney injury, as far as I have seen in the data. In any case, it follows the same pattern with all the columns, you cannot have bowel_healthy and bowel_injury simultaneously. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2445532,
      "author_name": "Art. Berz.",
      "author_url": "",
      "post_date": "2023-09-18T22:59:09.370000",
      "content": "<p>From a quick glance over, it isn't quite clear to me where your test patient_ids are getting injected into the submission.csv file. If you aren't getting them from the test_images part of the data, that might be an issue since they won't get pulled properly. I made the mistake of submitting just the baseline sample submission (just reading the file and submitting it) instead of pulling in the list of test_ids from test images, which for the sample test set is 3 patient_ids long, but for the hidden test set is over 1000 patient_ids long. If that isn't the issue, I've found it helpful to break down each part of the process until I understand it completely, and that usually gets me the answer of why things aren't working. Good luck debugging!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2445623,
          "author_name": "RickPack",
          "author_url": "",
          "post_date": "2023-09-19T01:54:37.940000",
          "content": "<p>Thank you, are you saying that the submission should have more rows than 3? I have the patient_id column, copied from the sample submission CSV.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2445739,
              "author_name": "Art. Berz.",
              "author_url": "",
              "post_date": "2023-09-19T03:58:50.403000",
              "content": "<p>Nope, the sample submission will always be 3 patient ids, but it matters how your code obtains them. If you're just taking the sample csv file that they give you and plugging numbers in, every time you run your code, it will just go to the sample submission and put numbers in there. Whereas, if you code in a way for the code to:</p>\n<ul>\n<li>Go into the test_images folder under competition data</li>\n<li>Grab patient id</li>\n<li>do other stuff, like plug in those numbers, format, and save as submission.csv<br>\nIt will always work, because it goes the same path when you submit. When you submit, there's the main block of code that runs which is your main notebook, the stuff you can see, and then if it finishes running, test_images gets switched out, where there's roughly 1000 patient_ids. If you have the pathway built, it works just as well, but if you don't, you're just getting scored on those 3, because that's all that your notebook is coded to be able to see. I'm not skilled enough in R programming language to really fully understand your notebook or run it, but I hope that makes sense. Let me know if you have any other questions. </li>\n</ul>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2456580,
              "author_name": "RickPack",
              "author_url": "",
              "post_date": "2023-09-26T09:47:39.610000",
              "content": "<p>Art, this pointed me to the solution and helped me understand the competition better. Thank you for your investment of time!</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2457000,
              "author_name": "Art. Berz.",
              "author_url": "",
              "post_date": "2023-09-26T14:43:58.647000",
              "content": "<p>Glad I could help :) .</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2446803": "I have almost the same issue. If I hardcode each submission with the set in the `sample_submission.cvs` file\n\n[<patient_id>, 0.5,0.5,0.5,0.5,0.3333333333333333,0.3333333333333333,0.3333333333333333,0.3333333333333333,0.3333333333333333,0.3333333333333333,0.3333333333333333,0.3333333333333333,0.3333333333333333]\n\nI get green. But if I hardcode to the following, I got \"Submission scoring error\":\n\n[<patient_id>, 1.0,1.0,1.0,1.0,1.0,1.0,1.0,1.0,1.0,1.0,1.0,1.0,1.0]\n\n@artberz can you please share how you set your arrays and then you convert to csv? I am really stuck!",
    "2443050": "Has anyone been able to submit successfully using an R notebook? I keep seeing \"Submission scoring error\".\n\nI used str() to examine the structure of the submission vs. the sample submissions, and others I have submitted, and all looks fine (e.g., 3 rows, 14 numeric columns).\n\nNotebook visible at https://www.kaggle.com/rickpack/r-kerascv-wtmean-submissions-averaged",
    "2445532": "From a quick glance over, it isn't quite clear to me where your test patient_ids are getting injected into the submission.csv file. If you aren't getting them from the test_images part of the data, that might be an issue since they won't get pulled properly. I made the mistake of submitting just the baseline sample submission (just reading the file and submitting it) instead of pulling in the list of test_ids from test images, which for the sample test set is 3 patient_ids long, but for the hidden test set is over 1000 patient_ids long. If that isn't the issue, I've found it helpful to break down each part of the process until I understand it completely, and that usually gets me the answer of why things aren't working. Good luck debugging!"
  }
}