{
  "id": 432253,
  "title": "Concerns about the time consumption of data preprocessing",
  "url": "/competitions/rsna-2023-abdominal-trauma-detection/discussion/432253",
  "author_name": "NorthM344",
  "post_date": "2023-08-16T16:02:02.063000",
  "votes": 4,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I'm trying to build a preprocssing pipeline for the images, but I find it takes a large amount of time. </p>\n<p>Considering that we must perform temporary preprocessing of the test images in the submission notebook, we must control the time required for the preprocessing pipeline to prevent runtime exceeding (Notice that the private test set is approximately twice the size of the public test set, so we must reserve enough time on public LB). I am worried that this will prevent us from using some useful preprocessing methods.</p>\n<p>I'm currently trying to apply resampling with <code>SimpleITK</code>, and it takes about 30s - 1min to process one 3D image.  It seems that just resampling will take too much time. Do you have any suggestions on how to perform resampling faster?</p>",
  "messages": [
    {
      "id": 2393940,
      "postDate": "2023-08-16T16:02:02.063Z",
      "content": "<p>I'm trying to build a preprocssing pipeline for the images, but I find it takes a large amount of time. </p>\n<p>Considering that we must perform temporary preprocessing of the test images in the submission notebook, we must control the time required for the preprocessing pipeline to prevent runtime exceeding (Notice that the private test set is approximately twice the size of the public test set, so we must reserve enough time on public LB). I am worried that this will prevent us from using some useful preprocessing methods.</p>\n<p>I'm currently trying to apply resampling with <code>SimpleITK</code>, and it takes about 30s - 1min to process one 3D image.  It seems that just resampling will take too much time. Do you have any suggestions on how to perform resampling faster?</p>",
      "rawMarkdown": "I'm trying to build a preprocssing pipeline for the images, but I find it takes a large amount of time. \n\nConsidering that we must perform temporary preprocessing of the test images in the submission notebook, we must control the time required for the preprocessing pipeline to prevent runtime exceeding (Notice that the private test set is approximately twice the size of the public test set, so we must reserve enough time on public LB). I am worried that this will prevent us from using some useful preprocessing methods.\n\nI'm currently trying to apply resampling with `SimpleITK`, and it takes about 30s - 1min to process one 3D image.  It seems that just resampling will take too much time. Do you have any suggestions on how to perform resampling faster?",
      "votes": 4
    },
    {
      "id": 2393986,
      "postDate": "2023-08-16T16:34:19.880Z",
      "content": "<p>You'll need to use GPU's and multiprocessing for faster data preprocessing. My current pipeline takes ~5s for one 3D volume. I'm pretty sure it can be optimised further to 3-4s.</p>",
      "rawMarkdown": "You'll need to use GPU's and multiprocessing for faster data preprocessing. My current pipeline takes ~5s for one 3D volume. I'm pretty sure it can be optimised further to 3-4s.",
      "votes": 1,
      "replies": [
        {
          "id": 2393997,
          "postDate": "2023-08-16T16:40:03.220Z",
          "content": "<p><code>SimpleITK</code> doesn't support GPU acceleration. I need to find other methods to do that.</p>",
          "rawMarkdown": "`SimpleITK` doesn't support GPU acceleration. I need to find other methods to do that."
        }
      ]
    },
    {
      "id": 2393961,
      "postDate": "2023-08-16T16:13:30.817Z",
      "content": "<blockquote>\n  <p>[train/test]_images/[patient_id]/[series_id]/[image_instance_number].dcm The CT scan data, in DICOM format. Scans from dozens of different CT machines have been reprocessed to use the run length encoded lossless compression format but retain other differences such as the number of bits per pixel, pixel range, and pixel representation. Expect to see roughly 1,100 patients in the test set.&gt;</p>\n</blockquote>\n<p>You misstated the size of the private test - the above is quoted from the Data tab.  But your very correct that 1 minute per image is pretty slow (some of the 1,100 patients will almost certainly have two or more image sets).</p>",
      "rawMarkdown": ">[train/test]_images/[patient_id]/[series_id]/[image_instance_number].dcm The CT scan data, in DICOM format. Scans from dozens of different CT machines have been reprocessed to use the run length encoded lossless compression format but retain other differences such as the number of bits per pixel, pixel range, and pixel representation. Expect to see roughly 1,100 patients in the test set.>\n\nYou misstated the size of the private test - the above is quoted from the Data tab.  But your very correct that 1 minute per image is pretty slow (some of the 1,100 patients will almost certainly have two or more image sets).\n\n",
      "votes": 1,
      "replies": [
        {
          "id": 2393966,
          "postDate": "2023-08-16T16:21:57.417Z",
          "content": "<blockquote>\n  <p>This leaderboard is calculated with approximately 36% of the test data. The final results will be based on the other 64%</p>\n</blockquote>\n<p>I mean that the 64% test set of private LB is about twice that of public LB's 36% dataset </p>",
          "rawMarkdown": ">This leaderboard is calculated with approximately 36% of the test data. The final results will be based on the other 64%\n\nI mean that the 64% test set of private LB is about twice that of public LB's 36% dataset "
        },
        {
          "id": 2393975,
          "postDate": "2023-08-16T16:28:17.117Z",
          "content": "<p>When you submit to get a leader board score kaggle runs the entire private test data (both 36% and 64% portions)  So if your notebook completes without running out of time your OK.   </p>\n<p>They score both but only show you the 36% value.  </p>",
          "rawMarkdown": "When you submit to get a leader board score kaggle runs the entire private test data (both 36% and 64% portions)  So if your notebook completes without running out of time your OK.   \n\nThey score both but only show you the 36% value.  ",
          "votes": 1,
          "replies": [
            {
              "id": 2393992,
              "postDate": "2023-08-16T16:37:44.227Z",
              "content": "<p>I see. Thank you for telling me. This means that preprocessing requires more time. It seems that I need to find a faster resampling method. </p>",
              "rawMarkdown": "I see. Thank you for telling me. This means that preprocessing requires more time. It seems that I need to find a faster resampling method. "
            }
          ]
        }
      ]
    },
    {
      "id": 2399175,
      "postDate": "2023-08-20T07:19:45.763Z",
      "content": "<p>It's not resampling, I'd say - SimpleITK takes 25-30 seconds per series to load the data. If you use Pydicom standard reader and stack images into 3D \"manually\", you should reduce the processing time per series to 2-2.5 seconds plus another 1-1.5 seconds for resampling, padding, and other transformations.</p>\n<p>You may also want to use CuPy library for resampling - with GPU support it should save some processing time as well, but the most time-consuming part for now is, likely, the reading of DICOM files per se.</p>",
      "rawMarkdown": "It's not resampling, I'd say - SimpleITK takes 25-30 seconds per series to load the data. If you use Pydicom standard reader and stack images into 3D \"manually\", you should reduce the processing time per series to 2-2.5 seconds plus another 1-1.5 seconds for resampling, padding, and other transformations.\n\nYou may also want to use CuPy library for resampling - with GPU support it should save some processing time as well, but the most time-consuming part for now is, likely, the reading of DICOM files per se.",
      "replies": [
        {
          "id": 2399478,
          "postDate": "2023-08-20T11:25:18.113Z",
          "content": "<p>Yes you are right. I found this problem few days ago and I have changed to use pydicom. It works well now</p>",
          "rawMarkdown": "Yes you are right. I found this problem few days ago and I have changed to use pydicom. It works well now",
          "votes": 3
        }
      ]
    },
    {
      "id": 2393964,
      "postDate": "2023-08-16T16:19:23.720Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2393986,
      "author_name": "Jebastin Nadar",
      "author_url": "",
      "post_date": "2023-08-16T16:34:19.880000",
      "content": "<p>You'll need to use GPU's and multiprocessing for faster data preprocessing. My current pipeline takes ~5s for one 3D volume. I'm pretty sure it can be optimised further to 3-4s.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2393997,
          "author_name": "NorthM344",
          "author_url": "",
          "post_date": "2023-08-16T16:40:03.220000",
          "content": "<p><code>SimpleITK</code> doesn't support GPU acceleration. I need to find other methods to do that.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2393961,
      "author_name": "PC Jimmmy",
      "author_url": "",
      "post_date": "2023-08-16T16:13:30.817000",
      "content": "<blockquote>\n  <p>[train/test]_images/[patient_id]/[series_id]/[image_instance_number].dcm The CT scan data, in DICOM format. Scans from dozens of different CT machines have been reprocessed to use the run length encoded lossless compression format but retain other differences such as the number of bits per pixel, pixel range, and pixel representation. Expect to see roughly 1,100 patients in the test set.&gt;</p>\n</blockquote>\n<p>You misstated the size of the private test - the above is quoted from the Data tab.  But your very correct that 1 minute per image is pretty slow (some of the 1,100 patients will almost certainly have two or more image sets).</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2393966,
          "author_name": "NorthM344",
          "author_url": "",
          "post_date": "2023-08-16T16:21:57.417000",
          "content": "<blockquote>\n  <p>This leaderboard is calculated with approximately 36% of the test data. The final results will be based on the other 64%</p>\n</blockquote>\n<p>I mean that the 64% test set of private LB is about twice that of public LB's 36% dataset </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2393975,
          "author_name": "PC Jimmmy",
          "author_url": "",
          "post_date": "2023-08-16T16:28:17.117000",
          "content": "<p>When you submit to get a leader board score kaggle runs the entire private test data (both 36% and 64% portions)  So if your notebook completes without running out of time your OK.   </p>\n<p>They score both but only show you the 36% value.  </p>",
          "votes": 1,
          "replies": [
            {
              "id": 2393992,
              "author_name": "NorthM344",
              "author_url": "",
              "post_date": "2023-08-16T16:37:44.227000",
              "content": "<p>I see. Thank you for telling me. This means that preprocessing requires more time. It seems that I need to find a faster resampling method. </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2399175,
      "author_name": "Victor Shlepov",
      "author_url": "",
      "post_date": "2023-08-20T07:19:45.763000",
      "content": "<p>It's not resampling, I'd say - SimpleITK takes 25-30 seconds per series to load the data. If you use Pydicom standard reader and stack images into 3D \"manually\", you should reduce the processing time per series to 2-2.5 seconds plus another 1-1.5 seconds for resampling, padding, and other transformations.</p>\n<p>You may also want to use CuPy library for resampling - with GPU support it should save some processing time as well, but the most time-consuming part for now is, likely, the reading of DICOM files per se.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2399478,
          "author_name": "NorthM344",
          "author_url": "",
          "post_date": "2023-08-20T11:25:18.113000",
          "content": "<p>Yes you are right. I found this problem few days ago and I have changed to use pydicom. It works well now</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2393964,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-08-16T16:19:23.720000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2393940": "I'm trying to build a preprocssing pipeline for the images, but I find it takes a large amount of time. \n\nConsidering that we must perform temporary preprocessing of the test images in the submission notebook, we must control the time required for the preprocessing pipeline to prevent runtime exceeding (Notice that the private test set is approximately twice the size of the public test set, so we must reserve enough time on public LB). I am worried that this will prevent us from using some useful preprocessing methods.\n\nI'm currently trying to apply resampling with `SimpleITK`, and it takes about 30s - 1min to process one 3D image.  It seems that just resampling will take too much time. Do you have any suggestions on how to perform resampling faster?",
    "2393986": "You'll need to use GPU's and multiprocessing for faster data preprocessing. My current pipeline takes ~5s for one 3D volume. I'm pretty sure it can be optimised further to 3-4s.",
    "2393961": ">[train/test]_images/[patient_id]/[series_id]/[image_instance_number].dcm The CT scan data, in DICOM format. Scans from dozens of different CT machines have been reprocessed to use the run length encoded lossless compression format but retain other differences such as the number of bits per pixel, pixel range, and pixel representation. Expect to see roughly 1,100 patients in the test set.>\n\nYou misstated the size of the private test - the above is quoted from the Data tab.  But your very correct that 1 minute per image is pretty slow (some of the 1,100 patients will almost certainly have two or more image sets).\n\n",
    "2399175": "It's not resampling, I'd say - SimpleITK takes 25-30 seconds per series to load the data. If you use Pydicom standard reader and stack images into 3D \"manually\", you should reduce the processing time per series to 2-2.5 seconds plus another 1-1.5 seconds for resampling, padding, and other transformations.\n\nYou may also want to use CuPy library for resampling - with GPU support it should save some processing time as well, but the most time-consuming part for now is, likely, the reading of DICOM files per se.",
    "2393964": ""
  }
}