{
  "id": 432113,
  "title": "Why does the test set contain only single-slice (2D) DICOM images for each patient?",
  "url": "/competitions/rsna-2023-abdominal-trauma-detection/discussion/432113",
  "author_name": "wuxiaolongg",
  "post_date": "2023-08-16T07:17:10.217000",
  "votes": 3,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Hello, competition host. We have noticed that the current test set consists of single DICOM (dcm) files, while the training set contains multiple DICOM files for each patient. This has led us to design our models to accommodate the 2D CT data in the test set, making it challenging to leverage the 3D CT data. Could you please clarify if the test set will be changed to 3D CT data in the later stages of the competition?</p>",
  "messages": [
    {
      "id": 2394354,
      "postDate": "2023-08-16T21:13:47.247Z",
      "content": "<p>I see a lot of confusion around this, so I will explain it completely</p>\n<p>When we make a notebook on kaggle for submission, we will only get 1 dicom file in the test set, but when we submit it, the test_images directory will be replaced with the actual test set which will have around 1100 test patients with every patient having one or more studies and each study with 100s of dicoms just like the format we have in train_images directory, our inference code / submission notebook will be ran on all of the test patients<br>\nTwo scores will be calculated, one will be using 36% of our predictions which will be shown to us after submission is finished and we will be able to see this on the public leaderboard and the other 64% will be hidden from us and will be revealed when the competition ends, only this hidden score which we call private leaderboard is our final standing, based on which we will receive medals</p>",
      "rawMarkdown": "I see a lot of confusion around this, so I will explain it completely\n\nWhen we make a notebook on kaggle for submission, we will only get 1 dicom file in the test set, but when we submit it, the test_images directory will be replaced with the actual test set which will have around 1100 test patients with every patient having one or more studies and each study with 100s of dicoms just like the format we have in train_images directory, our inference code / submission notebook will be ran on all of the test patients\nTwo scores will be calculated, one will be using 36% of our predictions which will be shown to us after submission is finished and we will be able to see this on the public leaderboard and the other 64% will be hidden from us and will be revealed when the competition ends, only this hidden score which we call private leaderboard is our final standing, based on which we will receive medals",
      "votes": 8,
      "replies": [
        {
          "id": 2394370,
          "postDate": "2023-08-16T21:25:30.997Z",
          "content": "<p><a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> is correct.</p>",
          "rawMarkdown": "@harshitsheoran is correct."
        },
        {
          "id": 2394922,
          "postDate": "2023-08-17T07:38:30.007Z",
          "content": "<p>We greatly appreciate your response. This is extremely helpful to us.</p>",
          "rawMarkdown": "We greatly appreciate your response. This is extremely helpful to us."
        }
      ]
    },
    {
      "id": 2393195,
      "postDate": "2023-08-16T07:17:10.217Z",
      "content": "<p>Hello, competition host. We have noticed that the current test set consists of single DICOM (dcm) files, while the training set contains multiple DICOM files for each patient. This has led us to design our models to accommodate the 2D CT data in the test set, making it challenging to leverage the 3D CT data. Could you please clarify if the test set will be changed to 3D CT data in the later stages of the competition?</p>",
      "rawMarkdown": "Hello, competition host. We have noticed that the current test set consists of single DICOM (dcm) files, while the training set contains multiple DICOM files for each patient. This has led us to design our models to accommodate the 2D CT data in the test set, making it challenging to leverage the 3D CT data. Could you please clarify if the test set will be changed to 3D CT data in the later stages of the competition?",
      "votes": 3
    },
    {
      "id": 2393947,
      "postDate": "2023-08-16T16:06:56.407Z",
      "content": "<p>The public test set data is a place holder only, and as such apparently kaggle decided 1 slice will be enough.</p>\n<p>The real and private test data set appears to be pretty similar to the training set.   There are a number of posts on this topic including a rant by me.</p>\n<p>If you need to confirm that your infer notebook works than you will need to create your own testing data.  I have created such a <a href=\"https://www.kaggle.com/datasets/pcjimmmy/rsna-7-patients-test\" target=\"_blank\">set</a>, but have not validated that it works on kaggle (using it on my local machine only so far).</p>",
      "rawMarkdown": "The public test set data is a place holder only, and as such apparently kaggle decided 1 slice will be enough.\n\nThe real and private test data set appears to be pretty similar to the training set.   There are a number of posts on this topic including a rant by me.\n\n If you need to confirm that your infer notebook works than you will need to create your own testing data.  I have created such a [set](https://www.kaggle.com/datasets/pcjimmmy/rsna-7-patients-test), but have not validated that it works on kaggle (using it on my local machine only so far).",
      "votes": 4
    },
    {
      "id": 2393200,
      "postDate": "2023-08-16T07:19:37.367Z",
      "content": "<p>I have the same question.</p>",
      "rawMarkdown": "I have the same question.",
      "votes": 2
    },
    {
      "id": 2393795,
      "postDate": "2023-08-16T14:28:06.093Z",
      "content": "<p>This reason might be due to some dicom files containing 3D data instead of 2D, as stated by <a href=\"https://www.kaggle.com/YYama\" target=\"_blank\">@YYama</a>. According to the <a href=\"https://dicom.nema.org/dicom/2013/output/chtml/part04/sect_B.5.html\" target=\"_blank\">DICOM SOP class uid specifications</a>, and GPT. However, I cannot verify it by probing the test data, for more details you can check:<br>\n<a href=\"https://www.kaggle.com/code/aritrag/kerascv-starter-notebook-infer\" target=\"_blank\">official starter notebook</a><br>\n<a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/431836\" target=\"_blank\">my thread</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6066624%2F5f32b1ef3cd060966f08c1ac78a73a7b%2Fdicom.png?generation=1692195992035735&amp;alt=media\" alt=\"GPT\"></p>",
      "rawMarkdown": "This reason might be due to some dicom files containing 3D data instead of 2D, as stated by @YYama. According to the [DICOM SOP class uid specifications](https://dicom.nema.org/dicom/2013/output/chtml/part04/sect_B.5.html), and GPT. However, I cannot verify it by probing the test data, for more details you can check:\n[official starter notebook](https://www.kaggle.com/code/aritrag/kerascv-starter-notebook-infer)\n[my thread](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/431836)\n\n![GPT](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6066624%2F5f32b1ef3cd060966f08c1ac78a73a7b%2Fdicom.png?generation=1692195992035735&alt=media)\n",
      "replies": [
        {
          "id": 2393849,
          "postDate": "2023-08-16T15:00:53.430Z",
          "content": "<p>Hello, and thank you for your response.<br>\nI visualized the DICOM images from the test dataset and found that they only contain 2D information.</p>",
          "rawMarkdown": "Hello, and thank you for your response.\nI visualized the DICOM images from the test dataset and found that they only contain 2D information.",
          "votes": 1,
          "replies": [
            {
              "id": 2393959,
              "postDate": "2023-08-16T16:11:44.137Z",
              "content": "<p>Thanks for the information. Have you tried probing the test dataset to try out what the actual test images look like, such as adding an assert statement to check listdir has at least say 5 .dcm files? I tried probing the test set to check whether the actual test set contains 3D images in a single .dcm but somehow my code throws an error.</p>",
              "rawMarkdown": "Thanks for the information. Have you tried probing the test dataset to try out what the actual test images look like, such as adding an assert statement to check listdir has at least say 5 .dcm files? I tried probing the test set to check whether the actual test set contains 3D images in a single .dcm but somehow my code throws an error."
            }
          ]
        }
      ]
    },
    {
      "id": 2393605,
      "postDate": "2023-08-16T12:02:18.700Z",
      "content": "<p>The 2D DICOM data also stores the z-coordinate, so by arranging according to this z-coordinate, you can obtain 3D data.</p>",
      "rawMarkdown": "The 2D DICOM data also stores the z-coordinate, so by arranging according to this z-coordinate, you can obtain 3D data.",
      "replies": [
        {
          "id": 2393848,
          "postDate": "2023-08-16T14:59:46.773Z",
          "content": "<p>Hello, and thank you for your response.<br>\nI understand that DICOM images contain z-axis information. However, even with the z-axis information, having just one DICOM image clearly won't allow me to obtain three-dimensional abdominal data.</p>",
          "rawMarkdown": "Hello, and thank you for your response.\nI understand that DICOM images contain z-axis information. However, even with the z-axis information, having just one DICOM image clearly won't allow me to obtain three-dimensional abdominal data."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2394354,
      "author_name": "Harshit Sheoran",
      "author_url": "",
      "post_date": "2023-08-16T21:13:47.247000",
      "content": "<p>I see a lot of confusion around this, so I will explain it completely</p>\n<p>When we make a notebook on kaggle for submission, we will only get 1 dicom file in the test set, but when we submit it, the test_images directory will be replaced with the actual test set which will have around 1100 test patients with every patient having one or more studies and each study with 100s of dicoms just like the format we have in train_images directory, our inference code / submission notebook will be ran on all of the test patients<br>\nTwo scores will be calculated, one will be using 36% of our predictions which will be shown to us after submission is finished and we will be able to see this on the public leaderboard and the other 64% will be hidden from us and will be revealed when the competition ends, only this hidden score which we call private leaderboard is our final standing, based on which we will receive medals</p>",
      "votes": 8,
      "replies": [
        {
          "id": 2394370,
          "author_name": "Sohier Dane",
          "author_url": "",
          "post_date": "2023-08-16T21:25:30.997000",
          "content": "<p><a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> is correct.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2394922,
          "author_name": "wuxiaolongg",
          "author_url": "",
          "post_date": "2023-08-17T07:38:30.007000",
          "content": "<p>We greatly appreciate your response. This is extremely helpful to us.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2393947,
      "author_name": "PC Jimmmy",
      "author_url": "",
      "post_date": "2023-08-16T16:06:56.407000",
      "content": "<p>The public test set data is a place holder only, and as such apparently kaggle decided 1 slice will be enough.</p>\n<p>The real and private test data set appears to be pretty similar to the training set.   There are a number of posts on this topic including a rant by me.</p>\n<p>If you need to confirm that your infer notebook works than you will need to create your own testing data.  I have created such a <a href=\"https://www.kaggle.com/datasets/pcjimmmy/rsna-7-patients-test\" target=\"_blank\">set</a>, but have not validated that it works on kaggle (using it on my local machine only so far).</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 2393200,
      "author_name": "Wang Yu-shen",
      "author_url": "",
      "post_date": "2023-08-16T07:19:37.367000",
      "content": "<p>I have the same question.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2393795,
      "author_name": "Chau YH",
      "author_url": "",
      "post_date": "2023-08-16T14:28:06.093000",
      "content": "<p>This reason might be due to some dicom files containing 3D data instead of 2D, as stated by <a href=\"https://www.kaggle.com/YYama\" target=\"_blank\">@YYama</a>. According to the <a href=\"https://dicom.nema.org/dicom/2013/output/chtml/part04/sect_B.5.html\" target=\"_blank\">DICOM SOP class uid specifications</a>, and GPT. However, I cannot verify it by probing the test data, for more details you can check:<br>\n<a href=\"https://www.kaggle.com/code/aritrag/kerascv-starter-notebook-infer\" target=\"_blank\">official starter notebook</a><br>\n<a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/431836\" target=\"_blank\">my thread</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6066624%2F5f32b1ef3cd060966f08c1ac78a73a7b%2Fdicom.png?generation=1692195992035735&amp;alt=media\" alt=\"GPT\"></p>",
      "votes": 0,
      "replies": [
        {
          "id": 2393849,
          "author_name": "wuxiaolongg",
          "author_url": "",
          "post_date": "2023-08-16T15:00:53.430000",
          "content": "<p>Hello, and thank you for your response.<br>\nI visualized the DICOM images from the test dataset and found that they only contain 2D information.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2393959,
              "author_name": "Chau YH",
              "author_url": "",
              "post_date": "2023-08-16T16:11:44.137000",
              "content": "<p>Thanks for the information. Have you tried probing the test dataset to try out what the actual test images look like, such as adding an assert statement to check listdir has at least say 5 .dcm files? I tried probing the test set to check whether the actual test set contains 3D images in a single .dcm but somehow my code throws an error.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2393605,
      "author_name": "YYama",
      "author_url": "",
      "post_date": "2023-08-16T12:02:18.700000",
      "content": "<p>The 2D DICOM data also stores the z-coordinate, so by arranging according to this z-coordinate, you can obtain 3D data.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2393848,
          "author_name": "wuxiaolongg",
          "author_url": "",
          "post_date": "2023-08-16T14:59:46.773000",
          "content": "<p>Hello, and thank you for your response.<br>\nI understand that DICOM images contain z-axis information. However, even with the z-axis information, having just one DICOM image clearly won't allow me to obtain three-dimensional abdominal data.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2394354": "I see a lot of confusion around this, so I will explain it completely\n\nWhen we make a notebook on kaggle for submission, we will only get 1 dicom file in the test set, but when we submit it, the test_images directory will be replaced with the actual test set which will have around 1100 test patients with every patient having one or more studies and each study with 100s of dicoms just like the format we have in train_images directory, our inference code / submission notebook will be ran on all of the test patients\nTwo scores will be calculated, one will be using 36% of our predictions which will be shown to us after submission is finished and we will be able to see this on the public leaderboard and the other 64% will be hidden from us and will be revealed when the competition ends, only this hidden score which we call private leaderboard is our final standing, based on which we will receive medals",
    "2393195": "Hello, competition host. We have noticed that the current test set consists of single DICOM (dcm) files, while the training set contains multiple DICOM files for each patient. This has led us to design our models to accommodate the 2D CT data in the test set, making it challenging to leverage the 3D CT data. Could you please clarify if the test set will be changed to 3D CT data in the later stages of the competition?",
    "2393947": "The public test set data is a place holder only, and as such apparently kaggle decided 1 slice will be enough.\n\nThe real and private test data set appears to be pretty similar to the training set.   There are a number of posts on this topic including a rant by me.\n\n If you need to confirm that your infer notebook works than you will need to create your own testing data.  I have created such a [set](https://www.kaggle.com/datasets/pcjimmmy/rsna-7-patients-test), but have not validated that it works on kaggle (using it on my local machine only so far).",
    "2393200": "I have the same question.",
    "2393795": "This reason might be due to some dicom files containing 3D data instead of 2D, as stated by @YYama. According to the [DICOM SOP class uid specifications](https://dicom.nema.org/dicom/2013/output/chtml/part04/sect_B.5.html), and GPT. However, I cannot verify it by probing the test data, for more details you can check:\n[official starter notebook](https://www.kaggle.com/code/aritrag/kerascv-starter-notebook-infer)\n[my thread](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/431836)\n\n![GPT](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6066624%2F5f32b1ef3cd060966f08c1ac78a73a7b%2Fdicom.png?generation=1692195992035735&alt=media)\n",
    "2393605": "The 2D DICOM data also stores the z-coordinate, so by arranging according to this z-coordinate, you can obtain 3D data."
  }
}