{
  "id": 340729,
  "title": "Paths don't exist on test.csv..",
  "url": "/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/340729",
  "author_name": "The Devastator",
  "post_date": "2022-07-30T17:17:30.754000",
  "votes": 10,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hello,</p>\n<p>I was trying to build a baseline for the competition and I found out that the small <code>test.csv</code> we have for testing is not small at all. <br>\nAlso, it doesn't contain many of the paths that exist in the test folder, and the other way around: The CSV file contains many paths that simply are not there. </p>\n<p>I assume that the csv file should have been the \"small test\" we get to test ourselves prior to submission but it doesn't look like it is the case.</p>\n<p>This situation makes it harder to submit since we need to automatically recognize within our code if we are running on the real test or simply in a \"commit session\".</p>\n<p>Is this a mistake? </p>\n<p>Thank you in advance. </p>",
  "messages": [
    {
      "id": 1877481,
      "postDate": "2022-07-30T17:17:30.753Z",
      "content": "<p>Hello,</p>\n<p>I was trying to build a baseline for the competition and I found out that the small <code>test.csv</code> we have for testing is not small at all. <br>\nAlso, it doesn't contain many of the paths that exist in the test folder, and the other way around: The CSV file contains many paths that simply are not there. </p>\n<p>I assume that the csv file should have been the \"small test\" we get to test ourselves prior to submission but it doesn't look like it is the case.</p>\n<p>This situation makes it harder to submit since we need to automatically recognize within our code if we are running on the real test or simply in a \"commit session\".</p>\n<p>Is this a mistake? </p>\n<p>Thank you in advance. </p>",
      "rawMarkdown": "Hello,\n\nI was trying to build a baseline for the competition and I found out that the small `test.csv` we have for testing is not small at all. \nAlso, it doesn't contain many of the paths that exist in the test folder, and the other way around: The CSV file contains many paths that simply are not there. \n\nI assume that the csv file should have been the \"small test\" we get to test ourselves prior to submission but it doesn't look like it is the case.\n\nThis situation makes it harder to submit since we need to automatically recognize within our code if we are running on the real test or simply in a \"commit session\".\n\nIs this a mistake? \n\nThank you in advance. \n\n\n",
      "votes": 9
    },
    {
      "id": 1877718,
      "postDate": "2022-07-31T00:30:35.667Z",
      "content": "<p>The test.csv we currently have 3 rows, it will be about 1500 rows while we submit, even if you think it is too big it is no problem because the only thing we need from test.csv is the column StudyInstanceUID, if we do <code>pd.unique(test.StudyInstanceUID)</code> we will get about 1500 values while submitting, and 3 values while running in notebook.. then you can use these 3 values to glob the folders in test_images, even if there are more rows in test, unique StudyInstanceUID is only about 1500 values, and even if there are more than 1500 folders in test_images, we only have about 1500 to predict on.. (each of those 1500 folder contains like 200-500 images, so yes, the test set is not small at all)</p>\n<p>If you feel the need to recognize if we are running on real test or commit session, if you can just do <code>IS_REAL = 1 if len(test)==3 else 0</code></p>",
      "rawMarkdown": "The test.csv we currently have 3 rows, it will be about 1500 rows while we submit, even if you think it is too big it is no problem because the only thing we need from test.csv is the column StudyInstanceUID, if we do `pd.unique(test.StudyInstanceUID)` we will get about 1500 values while submitting, and 3 values while running in notebook.. then you can use these 3 values to glob the folders in test_images, even if there are more rows in test, unique StudyInstanceUID is only about 1500 values, and even if there are more than 1500 folders in test_images, we only have about 1500 to predict on.. (each of those 1500 folder contains like 200-500 images, so yes, the test set is not small at all)\n\nIf you feel the need to recognize if we are running on real test or commit session, if you can just do `IS_REAL = 1 if len(test)==3 else 0`",
      "votes": 8,
      "replies": [
        {
          "id": 1877831,
          "postDate": "2022-07-31T03:16:41.243Z",
          "content": "<p>I ended up doing pretty much what you said but it would have been much easier on everyone if the test_df would only contain the 3 values that are present when we commit. <br>\nAs it is done in every other competition.. </p>",
          "rawMarkdown": "I ended up doing pretty much what you said but it would have been much easier on everyone if the test_df would only contain the 3 values that are present when we commit. \nAs it is done in every other competition.. ",
          "votes": 6
        },
        {
          "id": 1878105,
          "postDate": "2022-07-31T07:58:57.890Z",
          "content": "<p>I agree with this, the current structure introduces a higher level of complexity and could create some misgiving too.</p>",
          "rawMarkdown": "I agree with this, the current structure introduces a higher level of complexity and could create some misgiving too.",
          "votes": 1
        },
        {
          "id": 1985236,
          "postDate": "2022-10-13T07:35:03.043Z",
          "content": "<p>For some reason when I submit, the notebook still runs with the original test.csv with 3 rows, and the same with test_images. Am I doing something wrong? How can I access the hidden test set?</p>",
          "rawMarkdown": "For some reason when I submit, the notebook still runs with the original test.csv with 3 rows, and the same with test_images. Am I doing something wrong? How can I access the hidden test set?"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1877718,
      "author_name": "Harshit Sheoran",
      "author_url": "",
      "post_date": "2022-07-31T00:30:35.667000",
      "content": "<p>The test.csv we currently have 3 rows, it will be about 1500 rows while we submit, even if you think it is too big it is no problem because the only thing we need from test.csv is the column StudyInstanceUID, if we do <code>pd.unique(test.StudyInstanceUID)</code> we will get about 1500 values while submitting, and 3 values while running in notebook.. then you can use these 3 values to glob the folders in test_images, even if there are more rows in test, unique StudyInstanceUID is only about 1500 values, and even if there are more than 1500 folders in test_images, we only have about 1500 to predict on.. (each of those 1500 folder contains like 200-500 images, so yes, the test set is not small at all)</p>\n<p>If you feel the need to recognize if we are running on real test or commit session, if you can just do <code>IS_REAL = 1 if len(test)==3 else 0</code></p>",
      "votes": 8,
      "replies": [
        {
          "id": 1877831,
          "author_name": "The Devastator",
          "author_url": "",
          "post_date": "2022-07-31T03:16:41.243000",
          "content": "<p>I ended up doing pretty much what you said but it would have been much easier on everyone if the test_df would only contain the 3 values that are present when we commit. <br>\nAs it is done in every other competition.. </p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 1878105,
          "author_name": "Ravi Ramakrishnan",
          "author_url": "",
          "post_date": "2022-07-31T07:58:57.890000",
          "content": "<p>I agree with this, the current structure introduces a higher level of complexity and could create some misgiving too.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1985236,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-10-13T07:35:03.043000",
          "content": "<p>For some reason when I submit, the notebook still runs with the original test.csv with 3 rows, and the same with test_images. Am I doing something wrong? How can I access the hidden test set?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1877481": "Hello,\n\nI was trying to build a baseline for the competition and I found out that the small `test.csv` we have for testing is not small at all. \nAlso, it doesn't contain many of the paths that exist in the test folder, and the other way around: The CSV file contains many paths that simply are not there. \n\nI assume that the csv file should have been the \"small test\" we get to test ourselves prior to submission but it doesn't look like it is the case.\n\nThis situation makes it harder to submit since we need to automatically recognize within our code if we are running on the real test or simply in a \"commit session\".\n\nIs this a mistake? \n\nThank you in advance. \n\n\n",
    "1877718": "The test.csv we currently have 3 rows, it will be about 1500 rows while we submit, even if you think it is too big it is no problem because the only thing we need from test.csv is the column StudyInstanceUID, if we do `pd.unique(test.StudyInstanceUID)` we will get about 1500 values while submitting, and 3 values while running in notebook.. then you can use these 3 values to glob the folders in test_images, even if there are more rows in test, unique StudyInstanceUID is only about 1500 values, and even if there are more than 1500 folders in test_images, we only have about 1500 to predict on.. (each of those 1500 folder contains like 200-500 images, so yes, the test set is not small at all)\n\nIf you feel the need to recognize if we are running on real test or commit session, if you can just do `IS_REAL = 1 if len(test)==3 else 0`"
  }
}