{
  "id": 435305,
  "title": "Test set incomplete ? ",
  "url": "/competitions/rsna-2023-abdominal-trauma-detection/discussion/435305",
  "author_name": "Julien Genzling",
  "post_date": "2023-08-28T20:02:13.090000",
  "votes": 2,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hello. <br>\nI don't understand why we only have 3 dicom files (=3 images) in our test dataset. Shouldn't we dispose of more images to make accurate predictions ? More precisely, it seems difficult to predict extravasation or bowel health issues, since these two can occur in any place and will probably not show on one image. Also, I don't understand why it says in the presentation of the competition (Data) that we should \"expect to see roughly 1,100 patients in the test set\". </p>\n<p>My best guess is that the real test set is replaced when we submit the notebook. However, it seems to me that this causes issues since our model, at inference time, should be adapted to have just one image as an input. This becomes a problem when we train on 2.5D or 3D data. </p>\n<p>If anyone could take some time to clarify these points for me, I would be very grateful. </p>",
  "messages": [
    {
      "id": 2413361,
      "postDate": "2023-08-28T20:17:01.193Z",
      "content": "<blockquote>\n  <p>My best guess is that the real test set is replaced when we submit the notebook. </p>\n</blockquote>\n<p>Yes, given folder is simply a placeholder for you to test your model on. Real test case is hidden and your notebook will be run on entire test set when submitted. As mentioned in the Data tab of the competition - <code>Expect to see roughly 1,100 patients in the test set.</code></p>\n<blockquote>\n  <p>should be adapted to have just one image as an input. This becomes a problem when we train on 2.5D or 3D data.</p>\n</blockquote>\n<p>I use either one of the below code and it works fine. In case of only one image available it, will return copies of the same image.<br>\n<code>indices = np.quantile(list(range(shape[2])), np.linspace(0., 1., Z)).round().astype(int)</code><br>\nor </p>\n<pre><code> ():\n    series = glob()\n    paths = []\n     series_path  series:\n        paths.extend(glob())  \n    N = (paths) - \n    res_paths = []\n     i  p:\n        idx = (N * i)\n        res_paths.append(paths[idx])\n     res_paths\n</code></pre>",
      "rawMarkdown": "> My best guess is that the real test set is replaced when we submit the notebook. \n\nYes, given folder is simply a placeholder for you to test your model on. Real test case is hidden and your notebook will be run on entire test set when submitted. As mentioned in the Data tab of the competition - ` Expect to see roughly 1,100 patients in the test set.`\n\n> should be adapted to have just one image as an input. This becomes a problem when we train on 2.5D or 3D data.\n\nI use either one of the below code and it works fine. In case of only one image available it, will return copies of the same image.\n`indices = np.quantile(list(range(shape[2])), np.linspace(0., 1., Z)).round().astype(int)`\nor \n```\ndef getNthPercentileFile(patient_id, p=[0.25, 0.50, 0.75]):\n    series = glob(f\"{BASE_PATH}/test_images/{patient_id}/*\")\n    paths = []\n    for series_path in series:\n        paths.extend(glob(f\"{series_path}/*\"))  # Paths need to be sorted by scan/instance id here though.\n    N = len(paths) - 1\n    res_paths = []\n    for i in p:\n        idx = int(N * i)\n        res_paths.append(paths[idx])\n    return res_paths\n```",
      "votes": 2,
      "replies": [
        {
          "id": 2413377,
          "postDate": "2023-08-28T20:35:27.177Z",
          "content": "<p>Thank you :)</p>",
          "rawMarkdown": "Thank you :)"
        },
        {
          "id": 2442071,
          "postDate": "2023-09-16T17:20:25.600Z",
          "rawMarkdown": "",
          "isDeleted": true,
          "replies": [
            {
              "id": 2442824,
              "postDate": "2023-09-17T10:31:33.320Z",
              "content": "<p>yes true, this was just an example. <br>\nI later changed it to be unique to series_id and patient_id and then club the results over patient by mean or something similar but forgot to edit it here!<br>\nGood shout for people who come to this post now :)</p>",
              "rawMarkdown": "yes true, this was just an example. \nI later changed it to be unique to series_id and patient_id and then club the results over patient by mean or something similar but forgot to edit it here!\nGood shout for people who come to this post now :)"
            }
          ]
        }
      ]
    },
    {
      "id": 2413348,
      "postDate": "2023-08-28T20:02:13.090Z",
      "content": "<p>Hello. <br>\nI don't understand why we only have 3 dicom files (=3 images) in our test dataset. Shouldn't we dispose of more images to make accurate predictions ? More precisely, it seems difficult to predict extravasation or bowel health issues, since these two can occur in any place and will probably not show on one image. Also, I don't understand why it says in the presentation of the competition (Data) that we should \"expect to see roughly 1,100 patients in the test set\". </p>\n<p>My best guess is that the real test set is replaced when we submit the notebook. However, it seems to me that this causes issues since our model, at inference time, should be adapted to have just one image as an input. This becomes a problem when we train on 2.5D or 3D data. </p>\n<p>If anyone could take some time to clarify these points for me, I would be very grateful. </p>",
      "rawMarkdown": "Hello. \nI don't understand why we only have 3 dicom files (=3 images) in our test dataset. Shouldn't we dispose of more images to make accurate predictions ? More precisely, it seems difficult to predict extravasation or bowel health issues, since these two can occur in any place and will probably not show on one image. Also, I don't understand why it says in the presentation of the competition (Data) that we should \"expect to see roughly 1,100 patients in the test set\". \n\nMy best guess is that the real test set is replaced when we submit the notebook. However, it seems to me that this causes issues since our model, at inference time, should be adapted to have just one image as an input. This becomes a problem when we train on 2.5D or 3D data. \n\nIf anyone could take some time to clarify these points for me, I would be very grateful. \n",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 2413361,
      "author_name": "Priya Nagda",
      "author_url": "",
      "post_date": "2023-08-28T20:17:01.193000",
      "content": "<blockquote>\n  <p>My best guess is that the real test set is replaced when we submit the notebook. </p>\n</blockquote>\n<p>Yes, given folder is simply a placeholder for you to test your model on. Real test case is hidden and your notebook will be run on entire test set when submitted. As mentioned in the Data tab of the competition - <code>Expect to see roughly 1,100 patients in the test set.</code></p>\n<blockquote>\n  <p>should be adapted to have just one image as an input. This becomes a problem when we train on 2.5D or 3D data.</p>\n</blockquote>\n<p>I use either one of the below code and it works fine. In case of only one image available it, will return copies of the same image.<br>\n<code>indices = np.quantile(list(range(shape[2])), np.linspace(0., 1., Z)).round().astype(int)</code><br>\nor </p>\n<pre><code> ():\n    series = glob()\n    paths = []\n     series_path  series:\n        paths.extend(glob())  \n    N = (paths) - \n    res_paths = []\n     i  p:\n        idx = (N * i)\n        res_paths.append(paths[idx])\n     res_paths\n</code></pre>",
      "votes": 2,
      "replies": [
        {
          "id": 2413377,
          "author_name": "Julien Genzling",
          "author_url": "",
          "post_date": "2023-08-28T20:35:27.177000",
          "content": "<p>Thank you :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2442071,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-09-16T17:20:25.600000",
          "content": "",
          "votes": 0,
          "replies": [
            {
              "id": 2442824,
              "author_name": "Priya Nagda",
              "author_url": "",
              "post_date": "2023-09-17T10:31:33.320000",
              "content": "<p>yes true, this was just an example. <br>\nI later changed it to be unique to series_id and patient_id and then club the results over patient by mean or something similar but forgot to edit it here!<br>\nGood shout for people who come to this post now :)</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2413361": "> My best guess is that the real test set is replaced when we submit the notebook. \n\nYes, given folder is simply a placeholder for you to test your model on. Real test case is hidden and your notebook will be run on entire test set when submitted. As mentioned in the Data tab of the competition - ` Expect to see roughly 1,100 patients in the test set.`\n\n> should be adapted to have just one image as an input. This becomes a problem when we train on 2.5D or 3D data.\n\nI use either one of the below code and it works fine. In case of only one image available it, will return copies of the same image.\n`indices = np.quantile(list(range(shape[2])), np.linspace(0., 1., Z)).round().astype(int)`\nor \n```\ndef getNthPercentileFile(patient_id, p=[0.25, 0.50, 0.75]):\n    series = glob(f\"{BASE_PATH}/test_images/{patient_id}/*\")\n    paths = []\n    for series_path in series:\n        paths.extend(glob(f\"{series_path}/*\"))  # Paths need to be sorted by scan/instance id here though.\n    N = len(paths) - 1\n    res_paths = []\n    for i in p:\n        idx = int(N * i)\n        res_paths.append(paths[idx])\n    return res_paths\n```",
    "2413348": "Hello. \nI don't understand why we only have 3 dicom files (=3 images) in our test dataset. Shouldn't we dispose of more images to make accurate predictions ? More precisely, it seems difficult to predict extravasation or bowel health issues, since these two can occur in any place and will probably not show on one image. Also, I don't understand why it says in the presentation of the competition (Data) that we should \"expect to see roughly 1,100 patients in the test set\". \n\nMy best guess is that the real test set is replaced when we submit the notebook. However, it seems to me that this causes issues since our model, at inference time, should be adapted to have just one image as an input. This becomes a problem when we train on 2.5D or 3D data. \n\nIf anyone could take some time to clarify these points for me, I would be very grateful. \n"
  }
}