{
  "id": 369684,
  "title": "how long does it take to convert dicom to png? ",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/369684",
  "author_name": "hengck23",
  "post_date": "2022-12-01T02:02:45.675000",
  "votes": 25,
  "comment_count": 37,
  "views": 0,
  "content": "<p>I wonder if i have done anything wrong (e.g. wrong lib version, …).<br>\nIt takes about 6 hours just to convert the hidden test dicom files with the code:</p>\n<pre><code>#https://www.kaggle.com/code/theoviel/dicom-resized-png-jpg\ndef dcm_to_png(dcm_file, image_size=image_size, image_dir=image_dir):\n    patient_id = dcm_file.split('/')[-2]\n    image_id = dcm_file.split('/')[-1][:-4]\n\n    dicom = pydicom.dcmread(dcm_file)\n    img = dicom.pixel_array \n    img = (img - img.min()) / (img.max() - img.min()) \n    if dicom.PhotometricInterpretation == \"MONOCHROME1\":\n        img = 1 - img\n\n    img = cv2.resize(img, (image_size, image_size),cv2.INTER_LINEAR) \n    img = (img * 255).astype(np.uint8)\n    cv2.imwrite(image_dir + '/' + f'{patient_id}_{image_id}.png', img)\n\nprint('convert dcm_file ...', len(dcm_file))           \nParallel(n_jobs=4)(\n    delayed(dcm_to_png)(f, image_size=image_size, image_dir=image_dir)\n    for f in tqdm(dcm_file)\n)\n</code></pre>",
  "messages": [
    {
      "id": 2050784,
      "postDate": "2022-12-01T02:02:45.677Z",
      "content": "<p>I wonder if i have done anything wrong (e.g. wrong lib version, …).<br>\nIt takes about 6 hours just to convert the hidden test dicom files with the code:</p>\n<pre><code>#https://www.kaggle.com/code/theoviel/dicom-resized-png-jpg\ndef dcm_to_png(dcm_file, image_size=image_size, image_dir=image_dir):\n    patient_id = dcm_file.split('/')[-2]\n    image_id = dcm_file.split('/')[-1][:-4]\n\n    dicom = pydicom.dcmread(dcm_file)\n    img = dicom.pixel_array \n    img = (img - img.min()) / (img.max() - img.min()) \n    if dicom.PhotometricInterpretation == \"MONOCHROME1\":\n        img = 1 - img\n\n    img = cv2.resize(img, (image_size, image_size),cv2.INTER_LINEAR) \n    img = (img * 255).astype(np.uint8)\n    cv2.imwrite(image_dir + '/' + f'{patient_id}_{image_id}.png', img)\n\nprint('convert dcm_file ...', len(dcm_file))           \nParallel(n_jobs=4)(\n    delayed(dcm_to_png)(f, image_size=image_size, image_dir=image_dir)\n    for f in tqdm(dcm_file)\n)\n</code></pre>",
      "rawMarkdown": "I wonder if i have done anything wrong (e.g. wrong lib version, ...).\nIt takes about 6 hours just to convert the hidden test dicom files with the code:\n\n```\n#https://www.kaggle.com/code/theoviel/dicom-resized-png-jpg\ndef dcm_to_png(dcm_file, image_size=image_size, image_dir=image_dir):\n    patient_id = dcm_file.split('/')[-2]\n    image_id = dcm_file.split('/')[-1][:-4]\n\n    dicom = pydicom.dcmread(dcm_file)\n    img = dicom.pixel_array \n    img = (img - img.min()) / (img.max() - img.min()) \n    if dicom.PhotometricInterpretation == \"MONOCHROME1\":\n        img = 1 - img\n\n    img = cv2.resize(img, (image_size, image_size),cv2.INTER_LINEAR) \n    img = (img * 255).astype(np.uint8)\n    cv2.imwrite(image_dir + '/' + f'{patient_id}_{image_id}.png', img)\n\nprint('convert dcm_file ...', len(dcm_file))           \nParallel(n_jobs=4)(\n    delayed(dcm_to_png)(f, image_size=image_size, image_dir=image_dir)\n    for f in tqdm(dcm_file)\n)\n```",
      "votes": 25
    },
    {
      "id": 2051694,
      "postDate": "2022-12-01T15:18:15.113Z",
      "content": "<p>Maybe remove the step of writing the JPG to disk and just pass the pixels directly to your data loader. There's no real point in storing the JPGs.</p>\n<p>I don't think we'll have time to export ALL the images in the test set and infer on them .. especially not with an ensemble/multiple models.</p>\n<p>It seems like we'll only be able to have a single model inferring on a one (maybe two) views per patient.</p>",
      "rawMarkdown": "Maybe remove the step of writing the JPG to disk and just pass the pixels directly to your data loader. There's no real point in storing the JPGs.\n\nI don't think we'll have time to export ALL the images in the test set and infer on them .. especially not with an ensemble/multiple models.\n\nIt seems like we'll only be able to have a single model inferring on a one (maybe two) views per patient.",
      "votes": 3
    },
    {
      "id": 2058037,
      "postDate": "2022-12-07T15:07:53.627Z",
      "content": "<p>too long 😭😭 for me at least 9 hours on the train set</p>",
      "rawMarkdown": "too long 😭😭 for me at least 9 hours on the train set",
      "votes": 1
    },
    {
      "id": 2057267,
      "postDate": "2022-12-06T23:09:16.780Z",
      "content": "<p>What about this lib <br>\n<a href=\"https://github.com/tsangel/dicomsdl\" target=\"_blank\">https://github.com/tsangel/dicomsdl</a><br>\n🤔</p>",
      "rawMarkdown": "What about this lib \nhttps://github.com/tsangel/dicomsdl\n🤔",
      "votes": 1,
      "replies": [
        {
          "id": 2057282,
          "postDate": "2022-12-06T23:50:49.367Z",
          "content": "<p>Seems faster with a quick test.  Assuming it works with all the images and opencv, probably a better option to try</p>\n<pre><code> dicomsdl  dicom, time\n pydicom  dcmread\nf = \ndset = dicom.(f)\nst = time.time()\ndset.pixelData()\n(time.time() - st)\nst = time.time()\nimage_dataset = dcmread(f).pixel_array\n(time.time() - st)\n</code></pre>\n<p>1.2729527950286865<br>\n1.6648890972137451</p>",
          "rawMarkdown": "Seems faster with a quick test.  Assuming it works with all the images and opencv, probably a better option to try\n\n```python\nimport dicomsdl as dicom, time\nfrom pydicom import dcmread\nf = \"/kaggle/input/rsna-breast-cancer-detection/train_images/10006/1459541791.dcm\"\ndset = dicom.open(f)\nst = time.time()\ndset.pixelData()\nprint(time.time() - st)\nst = time.time()\nimage_dataset = dcmread(f).pixel_array\nprint(time.time() - st)\n```\n\n1.2729527950286865\n1.6648890972137451\n",
          "votes": 6
        },
        {
          "id": 2057290,
          "postDate": "2022-12-07T00:01:59.043Z",
          "content": "<p>Curiously the pixelData from dicomsdl is 2x what you get from dicom.</p>\n<pre><code>st = time.time()\npa = dset.pixelData()\n(time.time() - st, pa.nbytes)\nst = time.time()\nimage_dataset = dcmread(f).pixel_array\n(time.time() - st, image_dataset.nbytes*)\n</code></pre>\n<p>1.2884056568145752 105279300<br>\n1.5564007759094238 105279300</p>",
          "rawMarkdown": "Curiously the pixelData from dicomsdl is 2x what you get from dicom.\n\n```python\nst = time.time()\npa = dset.pixelData()\nprint(time.time() - st, pa.nbytes)\nst = time.time()\nimage_dataset = dcmread(f).pixel_array\nprint(time.time() - st, image_dataset.nbytes*2)\n```\n1.2884056568145752 105279300\n1.5564007759094238 105279300"
        },
        {
          "id": 2057478,
          "postDate": "2022-12-07T06:28:54.880Z",
          "content": "<p>I haven't checked it, but I think one of them is a <code>float32</code> the other one is <code>uint16</code>.</p>",
          "rawMarkdown": "I haven't checked it, but I think one of them is a `float32` the other one is `uint16`."
        },
        {
          "id": 2058064,
          "postDate": "2022-12-07T15:41:11.593Z",
          "content": "<p>I checked this library …. improvement:</p>\n<p>Kaggle GPU env / 100 dicom images to process (images processed sequentialy - no Parallel):</p>\n<ul>\n<li>pydicom -&gt; 106,23 sec</li>\n<li>dicomsdl -&gt; 69 sec</li>\n</ul>\n<p>Kaggle GPU env / 500 images (images processed sequentialy - no Parallel):</p>\n<ul>\n<li>pydicom -&gt;</li>\n<li>dicomsdl -&gt; 381.98 sec</li>\n</ul>\n<p>GPU / 500 images (Parallel - 2 jobs):</p>\n<ul>\n<li>pydicom -&gt; 396.74 sec</li>\n<li>dicomsdl -&gt; 243.39 sec </li>\n</ul>\n<p>But I have to check result after exporting (images). </p>\n<p>I published quick implementation to experimtnt this: <a href=\"https://www.kaggle.com/remekkinas/fast-dicom-processing/\" target=\"_blank\">https://www.kaggle.com/remekkinas/fast-dicom-processing/</a></p>",
          "rawMarkdown": "I checked this library .... improvement:\n\nKaggle GPU env / 100 dicom images to process (images processed sequentialy - no Parallel):\n\n- pydicom -> 106,23 sec\n- dicomsdl -> 69 sec\n\nKaggle GPU env / 500 images (images processed sequentialy - no Parallel):\n- pydicom ->\n- dicomsdl -> 381.98 sec\n\nGPU / 500 images (Parallel - 2 jobs):\n- pydicom -> 396.74 sec\n- dicomsdl -> 243.39 sec \n\n\nBut I have to check result after exporting (images). \n\nI published quick implementation to experimtnt this: https://www.kaggle.com/remekkinas/fast-dicom-processing/",
          "votes": 6
        }
      ]
    },
    {
      "id": 2050834,
      "postDate": "2022-12-01T02:57:10.927Z",
      "content": "<p>yes,the training dataset contains 54000+images, it takes 9hours,the test set contains 8000 patients,per patient contains 4 images.,so the hidden testset contains 32000 images,it takes 6~7 hours</p>",
      "rawMarkdown": "yes,the training dataset contains 54000+images, it takes 9hours,the test set contains 8000 patients,per patient contains 4 images.,so the hidden testset contains 32000 images,it takes 6~7 hours",
      "votes": 1,
      "replies": [
        {
          "id": 2050837,
          "postDate": "2022-12-01T02:59:05.737Z",
          "content": "<p>Maybe we need to find a faster way to convert dicom images to png, jpg etc….</p>",
          "rawMarkdown": "Maybe we need to find a faster way to convert dicom images to png, jpg etc...."
        },
        {
          "id": 2050840,
          "postDate": "2022-12-01T03:09:43.060Z",
          "content": "<p>Yes，this important，</p>",
          "rawMarkdown": "Yes，this important，"
        },
        {
          "id": 2050861,
          "postDate": "2022-12-01T03:41:17.407Z",
          "content": "<p>maybe kaggle need to change how code submission works.<br>\nnow we are wasting 6-7 hours doing reperepetitive task (same dcm and png conversion code) for each submission, hogging and wasting gpu resource.</p>\n<p>i wonder if we can create and cache some pre-processed hidden dataset for future use. maybe this can be fufuture kaggle product upgrade</p>",
          "rawMarkdown": "maybe kaggle need to change how code submission works.\nnow we are wasting 6-7 hours doing reperepetitive task (same dcm and png conversion code) for each submission, hogging and wasting gpu resource.\n\ni wonder if we can create and cache some pre-processed hidden dataset for future use. maybe this can be fufuture kaggle product upgrade",
          "votes": 24
        },
        {
          "id": 2051213,
          "postDate": "2022-12-01T09:31:48.087Z",
          "content": "<p>Mammograms are usually of very high-resolution so the conversion is slow …<br>\nI'm wondering if Kaggle can provide us with extra test folders containing raw pixel array/ converted PNGs.</p>\n<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> Is this possible to happen in this competition ? </p>",
          "rawMarkdown": "Mammograms are usually of very high-resolution so the conversion is slow ...\nI'm wondering if Kaggle can provide us with extra test folders containing raw pixel array/ converted PNGs.\n\n@sohier Is this possible to happen in this competition ? ",
          "votes": 8
        },
        {
          "id": 2056179,
          "postDate": "2022-12-05T20:24:22.877Z",
          "content": "<p>what if kaggle offered a converted test set with a high resolution PNG format? e.g 1024x1024 pngs</p>",
          "rawMarkdown": "what if kaggle offered a converted test set with a high resolution PNG format? e.g 1024x1024 pngs",
          "votes": 1
        },
        {
          "id": 2056399,
          "postDate": "2022-12-06T04:56:15.460Z",
          "content": "<p>I agree with you. This is not ECO competition if 66% of submission time is file processing. We burn a lot of resources each time we want to submit and check models we trained. I know that this is part of solution but even cut hidden dataset to 50% will help our planet 😊</p>",
          "rawMarkdown": "I agree with you. This is not ECO competition if 66% of submission time is file processing. We burn a lot of resources each time we want to submit and check models we trained. I know that this is part of solution but even cut hidden dataset to 50% will help our planet 😊",
          "votes": 2
        },
        {
          "id": 2056947,
          "postDate": "2022-12-06T15:48:58.973Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 2056971,
          "postDate": "2022-12-06T16:10:44.803Z",
          "content": "<p><a href=\"https://www.kaggle.com/andy2709\" target=\"_blank\">@andy2709</a> Unfortunately, it looks like raw pixel arrays are 10x+ in size, which might just push the problem over into IO.  Does that look right to you?</p>\n<p>edit to add:  see my comments above.  If we can get kaggle to convert the images to a more friendly compression format for speed, we might strike the right balance between IO and compute.  lz4 is 2x in size, but .12s load times versus the 1.6s with dicom, a 10x speedup.   And that was just first blush optimization.</p>",
          "rawMarkdown": "@andy2709 Unfortunately, it looks like raw pixel arrays are 10x+ in size, which might just push the problem over into IO.  Does that look right to you?\n\nedit to add:  see my comments above.  If we can get kaggle to convert the images to a more friendly compression format for speed, we might strike the right balance between IO and compute.  lz4 is 2x in size, but .12s load times versus the 1.6s with dicom, a 10x speedup.   And that was just first blush optimization."
        },
        {
          "id": 2060801,
          "postDate": "2022-12-10T12:54:36.417Z",
          "content": "<p>Maybe create an ETL staging area</p>",
          "rawMarkdown": "Maybe create an ETL staging area"
        }
      ]
    },
    {
      "id": 2050860,
      "postDate": "2022-12-01T03:40:56.663Z",
      "content": "<p>1 hour to convert data 4x downscale .npy on my pc, could have been faster, I used only 8 threads… If it really takes 6 hours just to convert test, its gonna be a massive pain</p>",
      "rawMarkdown": "1 hour to convert data 4x downscale .npy on my pc, could have been faster, I used only 8 threads... If it really takes 6 hours just to convert test, its gonna be a massive pain",
      "replies": [
        {
          "id": 2051273,
          "postDate": "2022-12-01T10:07:04.403Z",
          "content": "<p>Maybe using cv2.imwrite is slow because it compress the array, using np.save could be faster?</p>\n<p>I will check and come back.</p>\n<p>Update: It is almost the same, bottleneck comes from loading the images.</p>",
          "rawMarkdown": "Maybe using cv2.imwrite is slow because it compress the array, using np.save could be faster?\n\nI will check and come back.\n\nUpdate: It is almost the same, bottleneck comes from loading the images.",
          "votes": 3
        }
      ]
    },
    {
      "id": 2052026,
      "postDate": "2022-12-01T20:22:20.137Z",
      "content": "<p>I think the bottleneck is the IO for reading the DICOM files on Kaggle notebook. A few hours in my experience are usually what it takes for other similar competitions just to read all the DICOMS. Try not doing anything else and rerun the codes just to read DICOM files. How much time did you save by not converting into png and saving to file? If it is not a lot of time, then it doesn't really penalize you that much by including this step in you preprocessing pipeline. Everyone has to read each DICOM file one way or the other..</p>",
      "rawMarkdown": "I think the bottleneck is the IO for reading the DICOM files on Kaggle notebook. A few hours in my experience are usually what it takes for other similar competitions just to read all the DICOMS. Try not doing anything else and rerun the codes just to read DICOM files. How much time did you save by not converting into png and saving to file? If it is not a lot of time, then it doesn't really penalize you that much by including this step in you preprocessing pipeline. Everyone has to read each DICOM file one way or the other..",
      "votes": 1,
      "replies": [
        {
          "id": 2056248,
          "postDate": "2022-12-05T22:46:56.883Z",
          "content": "<p>Parallel jobs should deal with any disk IO bottleneck.</p>",
          "rawMarkdown": "Parallel jobs should deal with any disk IO bottleneck."
        },
        {
          "id": 2056345,
          "postDate": "2022-12-06T03:18:06.493Z",
          "content": "<p>Nope, if your IO is limited at 100MBps (just for example), multiple threads/cores/processes won't let you read any faster than 100MBps</p>",
          "rawMarkdown": "Nope, if your IO is limited at 100MBps (just for example), multiple threads/cores/processes won't let you read any faster than 100MBps",
          "votes": 1
        },
        {
          "id": 2056682,
          "postDate": "2022-12-06T11:10:39.717Z",
          "content": "<p>the slowest part in Heng's and our scripts is: img = dicom.pixel_array, in which JPEG decoding happens. In other medical comps, this wasn't a major bottleneck since the pixel array are of manageable sizes (300x300, 512x512 etc) </p>",
          "rawMarkdown": "the slowest part in Heng's and our scripts is: img = dicom.pixel_array, in which JPEG decoding happens. In other medical comps, this wasn't a major bottleneck since the pixel array are of manageable sizes (300x300, 512x512 etc) ",
          "votes": 1
        },
        {
          "id": 2056753,
          "postDate": "2022-12-06T12:11:55.003Z",
          "content": "<p>Once you mention it I remember NVIDIA DALI framework supports GPU JPEG decoding, but it doesn't support dicom format directly, I wonder if we can find a way to extract non-decoded JPEG data from dicom file and then decode it on GPU</p>",
          "rawMarkdown": "Once you mention it I remember NVIDIA DALI framework supports GPU JPEG decoding, but it doesn't support dicom format directly, I wonder if we can find a way to extract non-decoded JPEG data from dicom file and then decode it on GPU",
          "votes": 1
        },
        {
          "id": 2056754,
          "postDate": "2022-12-06T12:12:56.713Z",
          "content": "<p>check also <a href=\"https://github.com/UsingNet/nvjpeg-python\" target=\"_blank\">https://github.com/UsingNet/nvjpeg-python</a></p>",
          "rawMarkdown": "check also https://github.com/UsingNet/nvjpeg-python"
        },
        {
          "id": 2056780,
          "postDate": "2022-12-06T12:26:13.600Z",
          "content": "<p>read raw uncompresse data …. this one?<br>\ndimg.PixelData <br>\n(PixelData - The raw byte string that is stored in the DICOM file)</p>\n<p><a href=\"https://github.com/pydicom/pydicom/issues/1274\" target=\"_blank\">https://github.com/pydicom/pydicom/issues/1274</a><br>\nsee also:<br>\n<a href=\"https://towardsdatascience.com/understanding-dicoms-835cd2e57d0b\" target=\"_blank\">https://towardsdatascience.com/understanding-dicoms-835cd2e57d0b</a></p>\n<p>This is not possible. The closest what you can do is dcmread(filename, specific_tags=[\"PixelData\"]). This would only read the contents of the PixelData tag (and SpecificCharacterSet, which is always read) and skip over all other tags. It will still skip over the tags one by one to read the tag length to be able to seek to the next tag.</p>",
          "rawMarkdown": "read raw uncompresse data .... this one?\ndimg.PixelData \n(PixelData - The raw byte string that is stored in the DICOM file)\n\nhttps://github.com/pydicom/pydicom/issues/1274\nsee also:\nhttps://towardsdatascience.com/understanding-dicoms-835cd2e57d0b\n\nThis is not possible. The closest what you can do is dcmread(filename, specific_tags=[\"PixelData\"]). This would only read the contents of the PixelData tag (and SpecificCharacterSet, which is always read) and skip over all other tags. It will still skip over the tags one by one to read the tag length to be able to seek to the next tag.\n\n",
          "votes": 2
        },
        {
          "id": 2056913,
          "postDate": "2022-12-06T15:09:26.733Z",
          "content": "<p><a href=\"https://www.kaggle.com/martynoveduard\" target=\"_blank\">@martynoveduard</a> </p>\n<p>See my comments below with regards to GPU decoding.  There are caveats</p>\n<p><a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> </p>\n<p>\"Nope\"..?  Glad to see some confidence.  ;)  However, you can profile where the hotspot is yourself, see here - </p>\n<p><a href=\"https://www.kaggle.com/code/kaggleqrdl/profile-dicom-resized-png-jpg?scriptVersionId=113109718\" target=\"_blank\">https://www.kaggle.com/code/kaggleqrdl/profile-dicom-resized-png-jpg?scriptVersionId=113109718</a></p>\n<p>read 0.1030178 , <br>\npixel_array 1.49143059,<br>\nresize 0.00801535, <br>\nwrite 0.00662226])</p>\n<p>And as <a href=\"https://www.kaggle.com/andy2709\" target=\"_blank\">@andy2709</a> mentions it appears to be in pixel array, possibly during decompress.   </p>\n<p>If so, maybe we could get an uncompressed / decoded dataset to work with (ie, saved pixel_array files).  Then indeed the read will be a bottleneck!  Though that can probably be addressed by putting augmentation in joblib.</p>",
          "rawMarkdown": "@martynoveduard \n\nSee my comments below with regards to GPU decoding.  There are caveats\n\n@harshitsheoran \n\n\"Nope\"..?  Glad to see some confidence.  ;)  However, you can profile where the hotspot is yourself, see here - \n\nhttps://www.kaggle.com/code/kaggleqrdl/profile-dicom-resized-png-jpg?scriptVersionId=113109718\n\nread 0.1030178 , \npixel_array 1.49143059,\nresize 0.00801535, \nwrite 0.00662226])\n\nAnd as @andy2709 mentions it appears to be in pixel array, possibly during decompress.   \n\nIf so, maybe we could get an uncompressed / decoded dataset to work with (ie, saved pixel_array files).  Then indeed the read will be a bottleneck!  Though that can probably be addressed by putting augmentation in joblib."
        },
        {
          "id": 2056953,
          "postDate": "2022-12-06T15:53:51.017Z",
          "content": "<p><a href=\"https://www.kaggle.com/kaggleqrdl\" target=\"_blank\">@kaggleqrdl</a>  I see, Thanks for pointing out my mistake, .pixel_array does not actually load anything from IO, it recreates the image from the base64 encoding. Good to know, because the dicom files were always so large I assumed that they hold metadata and the array too, but they only hold the encoding</p>",
          "rawMarkdown": "@kaggleqrdl  I see, Thanks for pointing out my mistake, .pixel_array does not actually load anything from IO, it recreates the image from the base64 encoding. Good to know, because the dicom files were always so large I assumed that they hold metadata and the array too, but they only hold the encoding"
        },
        {
          "id": 2056964,
          "postDate": "2022-12-06T16:00:41.537Z",
          "content": "<p><a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> </p>\n<p>Your point was an interesting one though, do you know the IO bandwidth constraints for Kaggle?</p>\n<p>Saved raw pixel arrays might be a no go for that reason .. I'm seeing a 10x+ size increase for the raw data being reported by np.nbytes</p>\n<p>I'm guessing the problem here is all in decompress.</p>",
          "rawMarkdown": "@harshitsheoran \n\nYour point was an interesting one though, do you know the IO bandwidth constraints for Kaggle?\n\nSaved raw pixel arrays might be a no go for that reason .. I'm seeing a 10x+ size increase for the raw data being reported by np.nbytes\n\nI'm guessing the problem here is all in decompress.\n\n\n"
        },
        {
          "id": 2056985,
          "postDate": "2022-12-06T16:29:45.763Z",
          "content": "<p>Maybe a faster compression format?  Get a dataset with compressed format that decompresses faster.</p>\n<p>For example, just using numpy savez I'm seeing .19s decompress load times.  There might be even faster formats out there.  It's about 1.5x the size of dicom, but that should be doable.</p>\n<p>lz4 is 2x in size, but a .12s decompress of file versus the 1.4s dicom pixel_array.  I'm just dumping/loading the compressed/decompressed pickle of the pixel_array of a single image there.   Some pretty basic stuff, I'm sure we can get more clever.</p>",
          "rawMarkdown": "Maybe a faster compression format?  Get a dataset with compressed format that decompresses faster.\n\nFor example, just using numpy savez I'm seeing .19s decompress load times.  There might be even faster formats out there.  It's about 1.5x the size of dicom, but that should be doable.\n\nlz4 is 2x in size, but a .12s decompress of file versus the 1.4s dicom pixel_array.  I'm just dumping/loading the compressed/decompressed pickle of the pixel_array of a single image there.   Some pretty basic stuff, I'm sure we can get more clever."
        },
        {
          "id": 2057008,
          "postDate": "2022-12-06T16:56:04.567Z",
          "content": "<p>Saving multiple imgs creates a 2x boost in time to load, but the 2x size still and issue with compressed pickle files.</p>\n<pre><code>imgs = []\ntsize = \nst = time.time()\n f  train_images[:]:\n    tsize = tsize + os.path.getsize(f)\n    imgs.append(pydicom.dcmread(f).pixel_array)\n(time.time() - st, tsize/**)\n</code></pre>\n<p>86.36614727973938 668.18131668.18131  (62s without the img append)</p>\n<pre><code> lz4.frame.(, )  f:\n    pickle.dump(imgs, f)\nt = time.time()\n lz4.frame.(, )  f:\n    arr = pickle.load(f)\n(time.time() - t)\n</code></pre>\n<p>6.029420375823975</p>\n<p>62 seconds to load via dicom.pixel_array, so a 10x speedup on a single process.  </p>\n<p>But..</p>\n<p><code>os.path.getsize(\"imgs.lz4\")/10**6</code><br>\n1124.482061  </p>\n<p>So, just under a 2x size increase with pickle.  Maybe there are more efficient serialization formats for np arrays.</p>\n<p>If we could just get all the imgs in a few blobs (not sure how many cores on submission machines), that <em>might</em> be interesting.   Big files can be troublesome sometimes.  Also, IO constraints as mentioned start being a factor probably, so need to think about how best to manage that.</p>",
          "rawMarkdown": "Saving multiple imgs creates a 2x boost in time to load, but the 2x size still and issue with compressed pickle files.\n\n```python\nimgs = []\ntsize = 0\nst = time.time()\nfor f in train_images[:100]:\n    tsize = tsize + os.path.getsize(f)\n    imgs.append(pydicom.dcmread(f).pixel_array)\nprint(time.time() - st, tsize/10**6)\n```\n86.36614727973938 668.18131668.18131  (62s without the img append)\n```python\nwith lz4.frame.open('imgs.lz4', 'wb') as f:\n    pickle.dump(imgs, f)\nt = time.time()\nwith lz4.frame.open('imgs.lz4', 'rb') as f:\n    arr = pickle.load(f)\nprint(time.time() - t)\n```\n6.029420375823975\n\n62 seconds to load via dicom.pixel_array, so a 10x speedup on a single process.  \n\nBut..\n\n`os.path.getsize(\"imgs.lz4\")/10**6`\n1124.482061  \n\nSo, just under a 2x size increase with pickle.  Maybe there are more efficient serialization formats for np arrays.\n\nIf we could just get all the imgs in a few blobs (not sure how many cores on submission machines), that *might* be interesting.   Big files can be troublesome sometimes.  Also, IO constraints as mentioned start being a factor probably, so need to think about how best to manage that."
        },
        {
          "id": 2057046,
          "postDate": "2022-12-06T17:45:45.887Z",
          "content": "<blockquote>\n  <p>not sure how many cores on submission machines</p>\n</blockquote>\n<p>2 cores if you are using a GPU accelerator, 4 if you use CPU only</p>",
          "rawMarkdown": "> not sure how many cores on submission machines\n\n2 cores if you are using a GPU accelerator, 4 if you use CPU only",
          "votes": 1
        },
        {
          "id": 2057054,
          "postDate": "2022-12-06T17:53:18.710Z",
          "content": "<p>It's actually 2 cores with hyperthreading, so 4 siblings on the CPU machines that I regularly get.    But submission machines might be different.</p>\n<p>You can see this in procinfo</p>\n<p>processor    : 0<br>\nvendor_id    : GenuineIntel<br>\ncpu family    : 6<br>\nmodel        : 79<br>\nmodel name    : Intel(R) Xeon(R) CPU @ 2.20GHz<br>\nstepping    : 0<br>\nmicrocode    : 0x1<br>\ncpu MHz        : 2199.998<br>\ncache size    : 56320 KB<br>\nphysical id    : 0<br>\nsiblings    : 4<br>\ncore id        : 0<br>\ncpu cores    : 2</p>\n<p>Hyperthreading comes with caveats in terms of performance when doing multiprocessing.</p>\n<p>edit to add:<br>\nActually, more accurately, we get 4 processors each with 2 cores.  I was just looking at the first processor.  Does that mean we have 8 cores?  On a p100 instance, it's two processors each with 1 core.</p>",
          "rawMarkdown": "It's actually 2 cores with hyperthreading, so 4 siblings on the CPU machines that I regularly get.    But submission machines might be different.\n\nYou can see this in procinfo\n\nprocessor\t: 0\nvendor_id\t: GenuineIntel\ncpu family\t: 6\nmodel\t\t: 79\nmodel name\t: Intel(R) Xeon(R) CPU @ 2.20GHz\nstepping\t: 0\nmicrocode\t: 0x1\ncpu MHz\t\t: 2199.998\ncache size\t: 56320 KB\nphysical id\t: 0\nsiblings\t: 4\ncore id\t\t: 0\ncpu cores\t: 2\n\nHyperthreading comes with caveats in terms of performance when doing multiprocessing.\n\nedit to add:\nActually, more accurately, we get 4 processors each with 2 cores.  I was just looking at the first processor.  Does that mean we have 8 cores?  On a p100 instance, it's two processors each with 1 core."
        }
      ]
    },
    {
      "id": 2056244,
      "postDate": "2022-12-05T22:31:00.073Z",
      "content": "<p>It might be faster to convert using gpu and then infer from there to skip memory/disk/gpu IO.  We have the 2 GPU modes now, though not sure if they get allocated for submission runs.  At the very least I don't know if we'll get the multiple cores on submission.  Have to ask.</p>\n<p><a href=\"https://stackoverflow.com/questions/58779746/image-resize-on-gpu-is-slower-than-cvresize\" target=\"_blank\">https://stackoverflow.com/questions/58779746/image-resize-on-gpu-is-slower-than-cvresize</a></p>\n<p>This may only be beneficial with larger resolutions.  512x512 is a lot less data to load into the gpu.  Loading in large batches rather than image by image might help  This may require augmentation on GPU as well, which might also provide some speedup.</p>",
      "rawMarkdown": "It might be faster to convert using gpu and then infer from there to skip memory/disk/gpu IO.  We have the 2 GPU modes now, though not sure if they get allocated for submission runs.  At the very least I don't know if we'll get the multiple cores on submission.  Have to ask.\n\nhttps://stackoverflow.com/questions/58779746/image-resize-on-gpu-is-slower-than-cvresize\n\nThis may only be beneficial with larger resolutions.  512x512 is a lot less data to load into the gpu.  Loading in large batches rather than image by image might help  This may require augmentation on GPU as well, which might also provide some speedup.\n\n"
    },
    {
      "id": 2051812,
      "postDate": "2022-12-01T16:23:53.967Z",
      "content": "<p>Yeah, it took 6 hours for converting all files.</p>",
      "rawMarkdown": "Yeah, it took 6 hours for converting all files."
    },
    {
      "id": 2050820,
      "postDate": "2022-12-01T02:45:18.230Z",
      "content": "<p>Test dataset is not public, Is it possible to infer with converting images?</p>",
      "rawMarkdown": "Test dataset is not public, Is it possible to infer with converting images?",
      "replies": [
        {
          "id": 2050835,
          "postDate": "2022-12-01T02:57:16.777Z",
          "content": "<p>Well, my best guess is that using one submission and do the same process for the hidden test set will give you how much it takes. </p>",
          "rawMarkdown": "Well, my best guess is that using one submission and do the same process for the hidden test set will give you how much it takes. ",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2051694,
      "author_name": "David Roberts",
      "author_url": "",
      "post_date": "2022-12-01T15:18:15.113000",
      "content": "<p>Maybe remove the step of writing the JPG to disk and just pass the pixels directly to your data loader. There's no real point in storing the JPGs.</p>\n<p>I don't think we'll have time to export ALL the images in the test set and infer on them .. especially not with an ensemble/multiple models.</p>\n<p>It seems like we'll only be able to have a single model inferring on a one (maybe two) views per patient.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2058037,
      "author_name": "pineapple",
      "author_url": "",
      "post_date": "2022-12-07T15:07:53.627000",
      "content": "<p>too long 😭😭 for me at least 9 hours on the train set</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2057267,
      "author_name": "AleNic",
      "author_url": "",
      "post_date": "2022-12-06T23:09:16.780000",
      "content": "<p>What about this lib <br>\n<a href=\"https://github.com/tsangel/dicomsdl\" target=\"_blank\">https://github.com/tsangel/dicomsdl</a><br>\n🤔</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2057282,
          "author_name": "@kaggleqrdl",
          "author_url": "",
          "post_date": "2022-12-06T23:50:49.367000",
          "content": "<p>Seems faster with a quick test.  Assuming it works with all the images and opencv, probably a better option to try</p>\n<pre><code> dicomsdl  dicom, time\n pydicom  dcmread\nf = \ndset = dicom.(f)\nst = time.time()\ndset.pixelData()\n(time.time() - st)\nst = time.time()\nimage_dataset = dcmread(f).pixel_array\n(time.time() - st)\n</code></pre>\n<p>1.2729527950286865<br>\n1.6648890972137451</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 2057290,
          "author_name": "@kaggleqrdl",
          "author_url": "",
          "post_date": "2022-12-07T00:01:59.043000",
          "content": "<p>Curiously the pixelData from dicomsdl is 2x what you get from dicom.</p>\n<pre><code>st = time.time()\npa = dset.pixelData()\n(time.time() - st, pa.nbytes)\nst = time.time()\nimage_dataset = dcmread(f).pixel_array\n(time.time() - st, image_dataset.nbytes*)\n</code></pre>\n<p>1.2884056568145752 105279300<br>\n1.5564007759094238 105279300</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2057478,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2022-12-07T06:28:54.880000",
          "content": "<p>I haven't checked it, but I think one of them is a <code>float32</code> the other one is <code>uint16</code>.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2058064,
          "author_name": "Remek Kinas",
          "author_url": "",
          "post_date": "2022-12-07T15:41:11.593000",
          "content": "<p>I checked this library …. improvement:</p>\n<p>Kaggle GPU env / 100 dicom images to process (images processed sequentialy - no Parallel):</p>\n<ul>\n<li>pydicom -&gt; 106,23 sec</li>\n<li>dicomsdl -&gt; 69 sec</li>\n</ul>\n<p>Kaggle GPU env / 500 images (images processed sequentialy - no Parallel):</p>\n<ul>\n<li>pydicom -&gt;</li>\n<li>dicomsdl -&gt; 381.98 sec</li>\n</ul>\n<p>GPU / 500 images (Parallel - 2 jobs):</p>\n<ul>\n<li>pydicom -&gt; 396.74 sec</li>\n<li>dicomsdl -&gt; 243.39 sec </li>\n</ul>\n<p>But I have to check result after exporting (images). </p>\n<p>I published quick implementation to experimtnt this: <a href=\"https://www.kaggle.com/remekkinas/fast-dicom-processing/\" target=\"_blank\">https://www.kaggle.com/remekkinas/fast-dicom-processing/</a></p>",
          "votes": 6,
          "replies": []
        }
      ]
    },
    {
      "id": 2050834,
      "author_name": "yanqiangmiffy",
      "author_url": "",
      "post_date": "2022-12-01T02:57:10.927000",
      "content": "<p>yes,the training dataset contains 54000+images, it takes 9hours,the test set contains 8000 patients,per patient contains 4 images.,so the hidden testset contains 32000 images,it takes 6~7 hours</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2050837,
          "author_name": "Martin Kovacevic Buvinic",
          "author_url": "",
          "post_date": "2022-12-01T02:59:05.737000",
          "content": "<p>Maybe we need to find a faster way to convert dicom images to png, jpg etc….</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2050840,
          "author_name": "yanqiangmiffy",
          "author_url": "",
          "post_date": "2022-12-01T03:09:43.060000",
          "content": "<p>Yes，this important，</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2050861,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2022-12-01T03:41:17.407000",
          "content": "<p>maybe kaggle need to change how code submission works.<br>\nnow we are wasting 6-7 hours doing reperepetitive task (same dcm and png conversion code) for each submission, hogging and wasting gpu resource.</p>\n<p>i wonder if we can create and cache some pre-processed hidden dataset for future use. maybe this can be fufuture kaggle product upgrade</p>",
          "votes": 24,
          "replies": []
        },
        {
          "id": 2051213,
          "author_name": "NguyenThanhNhan",
          "author_url": "",
          "post_date": "2022-12-01T09:31:48.087000",
          "content": "<p>Mammograms are usually of very high-resolution so the conversion is slow …<br>\nI'm wondering if Kaggle can provide us with extra test folders containing raw pixel array/ converted PNGs.</p>\n<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> Is this possible to happen in this competition ? </p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 2056179,
          "author_name": "moth",
          "author_url": "",
          "post_date": "2022-12-05T20:24:22.877000",
          "content": "<p>what if kaggle offered a converted test set with a high resolution PNG format? e.g 1024x1024 pngs</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2056399,
          "author_name": "Remek Kinas",
          "author_url": "",
          "post_date": "2022-12-06T04:56:15.460000",
          "content": "<p>I agree with you. This is not ECO competition if 66% of submission time is file processing. We burn a lot of resources each time we want to submit and check models we trained. I know that this is part of solution but even cut hidden dataset to 50% will help our planet 😊</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2056947,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-12-06T15:48:58.973000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2056971,
          "author_name": "@kaggleqrdl",
          "author_url": "",
          "post_date": "2022-12-06T16:10:44.803000",
          "content": "<p><a href=\"https://www.kaggle.com/andy2709\" target=\"_blank\">@andy2709</a> Unfortunately, it looks like raw pixel arrays are 10x+ in size, which might just push the problem over into IO.  Does that look right to you?</p>\n<p>edit to add:  see my comments above.  If we can get kaggle to convert the images to a more friendly compression format for speed, we might strike the right balance between IO and compute.  lz4 is 2x in size, but .12s load times versus the 1.6s with dicom, a 10x speedup.   And that was just first blush optimization.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2060801,
          "author_name": "Robson",
          "author_url": "",
          "post_date": "2022-12-10T12:54:36.417000",
          "content": "<p>Maybe create an ETL staging area</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2050860,
      "author_name": "Harshit Sheoran",
      "author_url": "",
      "post_date": "2022-12-01T03:40:56.663000",
      "content": "<p>1 hour to convert data 4x downscale .npy on my pc, could have been faster, I used only 8 threads… If it really takes 6 hours just to convert test, its gonna be a massive pain</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2051273,
          "author_name": "Martin Kovacevic Buvinic",
          "author_url": "",
          "post_date": "2022-12-01T10:07:04.403000",
          "content": "<p>Maybe using cv2.imwrite is slow because it compress the array, using np.save could be faster?</p>\n<p>I will check and come back.</p>\n<p>Update: It is almost the same, bottleneck comes from loading the images.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2052026,
      "author_name": "Yee Ng",
      "author_url": "",
      "post_date": "2022-12-01T20:22:20.137000",
      "content": "<p>I think the bottleneck is the IO for reading the DICOM files on Kaggle notebook. A few hours in my experience are usually what it takes for other similar competitions just to read all the DICOMS. Try not doing anything else and rerun the codes just to read DICOM files. How much time did you save by not converting into png and saving to file? If it is not a lot of time, then it doesn't really penalize you that much by including this step in you preprocessing pipeline. Everyone has to read each DICOM file one way or the other..</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2056248,
          "author_name": "@kaggleqrdl",
          "author_url": "",
          "post_date": "2022-12-05T22:46:56.883000",
          "content": "<p>Parallel jobs should deal with any disk IO bottleneck.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2056345,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2022-12-06T03:18:06.493000",
          "content": "<p>Nope, if your IO is limited at 100MBps (just for example), multiple threads/cores/processes won't let you read any faster than 100MBps</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2056682,
          "author_name": "NguyenThanhNhan",
          "author_url": "",
          "post_date": "2022-12-06T11:10:39.717000",
          "content": "<p>the slowest part in Heng's and our scripts is: img = dicom.pixel_array, in which JPEG decoding happens. In other medical comps, this wasn't a major bottleneck since the pixel array are of manageable sizes (300x300, 512x512 etc) </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2056753,
          "author_name": "slime",
          "author_url": "",
          "post_date": "2022-12-06T12:11:55.003000",
          "content": "<p>Once you mention it I remember NVIDIA DALI framework supports GPU JPEG decoding, but it doesn't support dicom format directly, I wonder if we can find a way to extract non-decoded JPEG data from dicom file and then decode it on GPU</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2056754,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2022-12-06T12:12:56.713000",
          "content": "<p>check also <a href=\"https://github.com/UsingNet/nvjpeg-python\" target=\"_blank\">https://github.com/UsingNet/nvjpeg-python</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2056780,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2022-12-06T12:26:13.600000",
          "content": "<p>read raw uncompresse data …. this one?<br>\ndimg.PixelData <br>\n(PixelData - The raw byte string that is stored in the DICOM file)</p>\n<p><a href=\"https://github.com/pydicom/pydicom/issues/1274\" target=\"_blank\">https://github.com/pydicom/pydicom/issues/1274</a><br>\nsee also:<br>\n<a href=\"https://towardsdatascience.com/understanding-dicoms-835cd2e57d0b\" target=\"_blank\">https://towardsdatascience.com/understanding-dicoms-835cd2e57d0b</a></p>\n<p>This is not possible. The closest what you can do is dcmread(filename, specific_tags=[\"PixelData\"]). This would only read the contents of the PixelData tag (and SpecificCharacterSet, which is always read) and skip over all other tags. It will still skip over the tags one by one to read the tag length to be able to seek to the next tag.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2056913,
          "author_name": "@kaggleqrdl",
          "author_url": "",
          "post_date": "2022-12-06T15:09:26.733000",
          "content": "<p><a href=\"https://www.kaggle.com/martynoveduard\" target=\"_blank\">@martynoveduard</a> </p>\n<p>See my comments below with regards to GPU decoding.  There are caveats</p>\n<p><a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> </p>\n<p>\"Nope\"..?  Glad to see some confidence.  ;)  However, you can profile where the hotspot is yourself, see here - </p>\n<p><a href=\"https://www.kaggle.com/code/kaggleqrdl/profile-dicom-resized-png-jpg?scriptVersionId=113109718\" target=\"_blank\">https://www.kaggle.com/code/kaggleqrdl/profile-dicom-resized-png-jpg?scriptVersionId=113109718</a></p>\n<p>read 0.1030178 , <br>\npixel_array 1.49143059,<br>\nresize 0.00801535, <br>\nwrite 0.00662226])</p>\n<p>And as <a href=\"https://www.kaggle.com/andy2709\" target=\"_blank\">@andy2709</a> mentions it appears to be in pixel array, possibly during decompress.   </p>\n<p>If so, maybe we could get an uncompressed / decoded dataset to work with (ie, saved pixel_array files).  Then indeed the read will be a bottleneck!  Though that can probably be addressed by putting augmentation in joblib.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2056953,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2022-12-06T15:53:51.017000",
          "content": "<p><a href=\"https://www.kaggle.com/kaggleqrdl\" target=\"_blank\">@kaggleqrdl</a>  I see, Thanks for pointing out my mistake, .pixel_array does not actually load anything from IO, it recreates the image from the base64 encoding. Good to know, because the dicom files were always so large I assumed that they hold metadata and the array too, but they only hold the encoding</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2056964,
          "author_name": "@kaggleqrdl",
          "author_url": "",
          "post_date": "2022-12-06T16:00:41.537000",
          "content": "<p><a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> </p>\n<p>Your point was an interesting one though, do you know the IO bandwidth constraints for Kaggle?</p>\n<p>Saved raw pixel arrays might be a no go for that reason .. I'm seeing a 10x+ size increase for the raw data being reported by np.nbytes</p>\n<p>I'm guessing the problem here is all in decompress.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2056985,
          "author_name": "@kaggleqrdl",
          "author_url": "",
          "post_date": "2022-12-06T16:29:45.763000",
          "content": "<p>Maybe a faster compression format?  Get a dataset with compressed format that decompresses faster.</p>\n<p>For example, just using numpy savez I'm seeing .19s decompress load times.  There might be even faster formats out there.  It's about 1.5x the size of dicom, but that should be doable.</p>\n<p>lz4 is 2x in size, but a .12s decompress of file versus the 1.4s dicom pixel_array.  I'm just dumping/loading the compressed/decompressed pickle of the pixel_array of a single image there.   Some pretty basic stuff, I'm sure we can get more clever.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2057008,
          "author_name": "@kaggleqrdl",
          "author_url": "",
          "post_date": "2022-12-06T16:56:04.567000",
          "content": "<p>Saving multiple imgs creates a 2x boost in time to load, but the 2x size still and issue with compressed pickle files.</p>\n<pre><code>imgs = []\ntsize = \nst = time.time()\n f  train_images[:]:\n    tsize = tsize + os.path.getsize(f)\n    imgs.append(pydicom.dcmread(f).pixel_array)\n(time.time() - st, tsize/**)\n</code></pre>\n<p>86.36614727973938 668.18131668.18131  (62s without the img append)</p>\n<pre><code> lz4.frame.(, )  f:\n    pickle.dump(imgs, f)\nt = time.time()\n lz4.frame.(, )  f:\n    arr = pickle.load(f)\n(time.time() - t)\n</code></pre>\n<p>6.029420375823975</p>\n<p>62 seconds to load via dicom.pixel_array, so a 10x speedup on a single process.  </p>\n<p>But..</p>\n<p><code>os.path.getsize(\"imgs.lz4\")/10**6</code><br>\n1124.482061  </p>\n<p>So, just under a 2x size increase with pickle.  Maybe there are more efficient serialization formats for np arrays.</p>\n<p>If we could just get all the imgs in a few blobs (not sure how many cores on submission machines), that <em>might</em> be interesting.   Big files can be troublesome sometimes.  Also, IO constraints as mentioned start being a factor probably, so need to think about how best to manage that.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2057046,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2022-12-06T17:45:45.887000",
          "content": "<blockquote>\n  <p>not sure how many cores on submission machines</p>\n</blockquote>\n<p>2 cores if you are using a GPU accelerator, 4 if you use CPU only</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2057054,
          "author_name": "@kaggleqrdl",
          "author_url": "",
          "post_date": "2022-12-06T17:53:18.710000",
          "content": "<p>It's actually 2 cores with hyperthreading, so 4 siblings on the CPU machines that I regularly get.    But submission machines might be different.</p>\n<p>You can see this in procinfo</p>\n<p>processor    : 0<br>\nvendor_id    : GenuineIntel<br>\ncpu family    : 6<br>\nmodel        : 79<br>\nmodel name    : Intel(R) Xeon(R) CPU @ 2.20GHz<br>\nstepping    : 0<br>\nmicrocode    : 0x1<br>\ncpu MHz        : 2199.998<br>\ncache size    : 56320 KB<br>\nphysical id    : 0<br>\nsiblings    : 4<br>\ncore id        : 0<br>\ncpu cores    : 2</p>\n<p>Hyperthreading comes with caveats in terms of performance when doing multiprocessing.</p>\n<p>edit to add:<br>\nActually, more accurately, we get 4 processors each with 2 cores.  I was just looking at the first processor.  Does that mean we have 8 cores?  On a p100 instance, it's two processors each with 1 core.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2056244,
      "author_name": "@kaggleqrdl",
      "author_url": "",
      "post_date": "2022-12-05T22:31:00.073000",
      "content": "<p>It might be faster to convert using gpu and then infer from there to skip memory/disk/gpu IO.  We have the 2 GPU modes now, though not sure if they get allocated for submission runs.  At the very least I don't know if we'll get the multiple cores on submission.  Have to ask.</p>\n<p><a href=\"https://stackoverflow.com/questions/58779746/image-resize-on-gpu-is-slower-than-cvresize\" target=\"_blank\">https://stackoverflow.com/questions/58779746/image-resize-on-gpu-is-slower-than-cvresize</a></p>\n<p>This may only be beneficial with larger resolutions.  512x512 is a lot less data to load into the gpu.  Loading in large batches rather than image by image might help  This may require augmentation on GPU as well, which might also provide some speedup.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2051812,
      "author_name": "Khang Duong",
      "author_url": "",
      "post_date": "2022-12-01T16:23:53.967000",
      "content": "<p>Yeah, it took 6 hours for converting all files.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2050820,
      "author_name": "dragon zhang",
      "author_url": "",
      "post_date": "2022-12-01T02:45:18.230000",
      "content": "<p>Test dataset is not public, Is it possible to infer with converting images?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2050835,
          "author_name": "Martin Kovacevic Buvinic",
          "author_url": "",
          "post_date": "2022-12-01T02:57:16.777000",
          "content": "<p>Well, my best guess is that using one submission and do the same process for the hidden test set will give you how much it takes. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2050784": "I wonder if i have done anything wrong (e.g. wrong lib version, ...).\nIt takes about 6 hours just to convert the hidden test dicom files with the code:\n\n```\n#https://www.kaggle.com/code/theoviel/dicom-resized-png-jpg\ndef dcm_to_png(dcm_file, image_size=image_size, image_dir=image_dir):\n    patient_id = dcm_file.split('/')[-2]\n    image_id = dcm_file.split('/')[-1][:-4]\n\n    dicom = pydicom.dcmread(dcm_file)\n    img = dicom.pixel_array \n    img = (img - img.min()) / (img.max() - img.min()) \n    if dicom.PhotometricInterpretation == \"MONOCHROME1\":\n        img = 1 - img\n\n    img = cv2.resize(img, (image_size, image_size),cv2.INTER_LINEAR) \n    img = (img * 255).astype(np.uint8)\n    cv2.imwrite(image_dir + '/' + f'{patient_id}_{image_id}.png', img)\n\nprint('convert dcm_file ...', len(dcm_file))           \nParallel(n_jobs=4)(\n    delayed(dcm_to_png)(f, image_size=image_size, image_dir=image_dir)\n    for f in tqdm(dcm_file)\n)\n```",
    "2051694": "Maybe remove the step of writing the JPG to disk and just pass the pixels directly to your data loader. There's no real point in storing the JPGs.\n\nI don't think we'll have time to export ALL the images in the test set and infer on them .. especially not with an ensemble/multiple models.\n\nIt seems like we'll only be able to have a single model inferring on a one (maybe two) views per patient.",
    "2058037": "too long 😭😭 for me at least 9 hours on the train set",
    "2057267": "What about this lib \nhttps://github.com/tsangel/dicomsdl\n🤔",
    "2050834": "yes,the training dataset contains 54000+images, it takes 9hours,the test set contains 8000 patients,per patient contains 4 images.,so the hidden testset contains 32000 images,it takes 6~7 hours",
    "2050860": "1 hour to convert data 4x downscale .npy on my pc, could have been faster, I used only 8 threads... If it really takes 6 hours just to convert test, its gonna be a massive pain",
    "2052026": "I think the bottleneck is the IO for reading the DICOM files on Kaggle notebook. A few hours in my experience are usually what it takes for other similar competitions just to read all the DICOMS. Try not doing anything else and rerun the codes just to read DICOM files. How much time did you save by not converting into png and saving to file? If it is not a lot of time, then it doesn't really penalize you that much by including this step in you preprocessing pipeline. Everyone has to read each DICOM file one way or the other..",
    "2056244": "It might be faster to convert using gpu and then infer from there to skip memory/disk/gpu IO.  We have the 2 GPU modes now, though not sure if they get allocated for submission runs.  At the very least I don't know if we'll get the multiple cores on submission.  Have to ask.\n\nhttps://stackoverflow.com/questions/58779746/image-resize-on-gpu-is-slower-than-cvresize\n\nThis may only be beneficial with larger resolutions.  512x512 is a lot less data to load into the gpu.  Loading in large batches rather than image by image might help  This may require augmentation on GPU as well, which might also provide some speedup.\n\n",
    "2051812": "Yeah, it took 6 hours for converting all files.",
    "2050820": "Test dataset is not public, Is it possible to infer with converting images?"
  }
}