{
  "id": 342189,
  "title": "🤔Discrepancy between IDs in test_images folder and test.csv🤔",
  "url": "/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/342189",
  "author_name": "Wonjun Kim",
  "post_date": "2022-08-06T00:41:19.178000",
  "votes": 4,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi! I don't understand the discrepancy between the contents of the test_images folder and the entries in the test.csv file</p>\n<p>The test_images folder contains CT scans for the following IDs</p>\n<ul>\n<li>1.2.826.0.1.3680043.22327</li>\n<li>1.2.826.0.1.3680043.25399</li>\n<li>1.2.826.0.1.3680043.5876</li>\n</ul>\n<p>But the test.csv file (and the sample_submission.csv file as well) has rows for the following IDs</p>\n<ul>\n<li>1.2.826.0.1.3680043.10197</li>\n<li>1.2.826.0.1.3680043.10454</li>\n<li>1.2.826.0.1.3680043.10690</li>\n</ul>\n<p>My understanding is that we build models to make predictions for images in the test_images folder and submit those predictions for a spot on the public leaderboard. But this difference between test_images folder and test.csv/sample_submission.csv file has me confused.</p>\n<p>How should I go about putting my submission together? Do I just make predictions for C1 fractures for IDs in the test_images folder and take the results and paste it to the IDs in the test.csv file? I'm new to kaggle and maybe I'm missing something? Any help would be fantabulous :D Thanks in advance 😃</p>",
  "messages": [
    {
      "id": 1888514,
      "postDate": "2022-08-07T16:29:49.703Z",
      "content": "<p>You don't actually submit predictions from the test.csv or test_image directory.  When you submit, you actually submit a notebook that has to read a test.csv and test_image directory provided at run time.  It will have way more than 3 image folders in it.  Somewhere else they say how many images (maybe 2000?).   Your notebook will have to create a submission file from that.  Thus there is no way for you to directly look at the test images.  The test_image directory and the test.csv file are just examples so you can check your formatting.  Why the ID's don't match between the two samples, I don't know.  If you look at the TF-RSNA-Efficient-Baseline notebook, you will see that the author substitutes the image directories provided for the ones in test.csv just so he can create a (dummy) submission file locally.</p>",
      "rawMarkdown": "You don't actually submit predictions from the test.csv or test_image directory.  When you submit, you actually submit a notebook that has to read a test.csv and test_image directory provided at run time.  It will have way more than 3 image folders in it.  Somewhere else they say how many images (maybe 2000?).   Your notebook will have to create a submission file from that.  Thus there is no way for you to directly look at the test images.  The test_image directory and the test.csv file are just examples so you can check your formatting.  Why the ID's don't match between the two samples, I don't know.  If you look at the TF-RSNA-Efficient-Baseline notebook, you will see that the author substitutes the image directories provided for the ones in test.csv just so he can create a (dummy) submission file locally.",
      "votes": 6,
      "replies": [
        {
          "id": 1890609,
          "postDate": "2022-08-08T22:32:56.347Z",
          "content": "<p>Oh I see! Thank you for clearing up my misunderstanding :D</p>",
          "rawMarkdown": "Oh I see! Thank you for clearing up my misunderstanding :D"
        }
      ]
    },
    {
      "id": 1886550,
      "postDate": "2022-08-06T00:41:19.180Z",
      "content": "<p>Hi! I don't understand the discrepancy between the contents of the test_images folder and the entries in the test.csv file</p>\n<p>The test_images folder contains CT scans for the following IDs</p>\n<ul>\n<li>1.2.826.0.1.3680043.22327</li>\n<li>1.2.826.0.1.3680043.25399</li>\n<li>1.2.826.0.1.3680043.5876</li>\n</ul>\n<p>But the test.csv file (and the sample_submission.csv file as well) has rows for the following IDs</p>\n<ul>\n<li>1.2.826.0.1.3680043.10197</li>\n<li>1.2.826.0.1.3680043.10454</li>\n<li>1.2.826.0.1.3680043.10690</li>\n</ul>\n<p>My understanding is that we build models to make predictions for images in the test_images folder and submit those predictions for a spot on the public leaderboard. But this difference between test_images folder and test.csv/sample_submission.csv file has me confused.</p>\n<p>How should I go about putting my submission together? Do I just make predictions for C1 fractures for IDs in the test_images folder and take the results and paste it to the IDs in the test.csv file? I'm new to kaggle and maybe I'm missing something? Any help would be fantabulous :D Thanks in advance 😃</p>",
      "rawMarkdown": "Hi! I don't understand the discrepancy between the contents of the test_images folder and the entries in the test.csv file\n\nThe test_images folder contains CT scans for the following IDs\n- 1.2.826.0.1.3680043.22327\n- 1.2.826.0.1.3680043.25399\n- 1.2.826.0.1.3680043.5876\n\nBut the test.csv file (and the sample_submission.csv file as well) has rows for the following IDs\n- 1.2.826.0.1.3680043.10197\n- 1.2.826.0.1.3680043.10454\n- 1.2.826.0.1.3680043.10690\n\nMy understanding is that we build models to make predictions for images in the test_images folder and submit those predictions for a spot on the public leaderboard. But this difference between test_images folder and test.csv/sample_submission.csv file has me confused.\n\nHow should I go about putting my submission together? Do I just make predictions for C1 fractures for IDs in the test_images folder and take the results and paste it to the IDs in the test.csv file? I'm new to kaggle and maybe I'm missing something? Any help would be fantabulous :D Thanks in advance 😃",
      "votes": 4
    }
  ],
  "comments": [
    {
      "id": 1888514,
      "author_name": "SolverWorld",
      "author_url": "",
      "post_date": "2022-08-07T16:29:49.703000",
      "content": "<p>You don't actually submit predictions from the test.csv or test_image directory.  When you submit, you actually submit a notebook that has to read a test.csv and test_image directory provided at run time.  It will have way more than 3 image folders in it.  Somewhere else they say how many images (maybe 2000?).   Your notebook will have to create a submission file from that.  Thus there is no way for you to directly look at the test images.  The test_image directory and the test.csv file are just examples so you can check your formatting.  Why the ID's don't match between the two samples, I don't know.  If you look at the TF-RSNA-Efficient-Baseline notebook, you will see that the author substitutes the image directories provided for the ones in test.csv just so he can create a (dummy) submission file locally.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 1890609,
          "author_name": "Wonjun Kim",
          "author_url": "",
          "post_date": "2022-08-08T22:32:56.347000",
          "content": "<p>Oh I see! Thank you for clearing up my misunderstanding :D</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1888514": "You don't actually submit predictions from the test.csv or test_image directory.  When you submit, you actually submit a notebook that has to read a test.csv and test_image directory provided at run time.  It will have way more than 3 image folders in it.  Somewhere else they say how many images (maybe 2000?).   Your notebook will have to create a submission file from that.  Thus there is no way for you to directly look at the test images.  The test_image directory and the test.csv file are just examples so you can check your formatting.  Why the ID's don't match between the two samples, I don't know.  If you look at the TF-RSNA-Efficient-Baseline notebook, you will see that the author substitutes the image directories provided for the ones in test.csv just so he can create a (dummy) submission file locally.",
    "1886550": "Hi! I don't understand the discrepancy between the contents of the test_images folder and the entries in the test.csv file\n\nThe test_images folder contains CT scans for the following IDs\n- 1.2.826.0.1.3680043.22327\n- 1.2.826.0.1.3680043.25399\n- 1.2.826.0.1.3680043.5876\n\nBut the test.csv file (and the sample_submission.csv file as well) has rows for the following IDs\n- 1.2.826.0.1.3680043.10197\n- 1.2.826.0.1.3680043.10454\n- 1.2.826.0.1.3680043.10690\n\nMy understanding is that we build models to make predictions for images in the test_images folder and submit those predictions for a spot on the public leaderboard. But this difference between test_images folder and test.csv/sample_submission.csv file has me confused.\n\nHow should I go about putting my submission together? Do I just make predictions for C1 fractures for IDs in the test_images folder and take the results and paste it to the IDs in the test.csv file? I'm new to kaggle and maybe I'm missing something? Any help would be fantabulous :D Thanks in advance 😃"
  }
}