{
  "id": 371033,
  "title": "Fast dicom export and processing (1.6-2x faster) 💪💪 ",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/371033",
  "author_name": "Remek Kinas",
  "post_date": "2022-12-07T16:32:10.115000",
  "votes": 51,
  "comment_count": 29,
  "views": 0,
  "content": "<p><strong>This is not my dicovery but decided to create separate topic because this is probably important discovery</strong>. <a href=\"https://www.kaggle.com/kaggleqrdl\" target=\"_blank\">@kaggleqrdl</a> described ( <a href=\"https://www.kaggle.com/alenic\" target=\"_blank\">@alenic</a> idea ) it  here: <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369684#2057282\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369684#2057282</a> I checked only speed of libraries and implemented notebook to show result and provide way to experiment.<br>\nIt seems that we got ~1,6x-2x improvemend in dicom processing speed.</p>\n<p>I created quick implementation how to speed up dicom processing significantly.</p>\n<p>GPU / 500 images (Parallel - 2 jobs):</p>\n<ul>\n<li>pydicom -&gt; 396.74 sec</li>\n<li>dicomsdl -&gt; 243.39 sec </li>\n</ul>\n<p>My estimation is that now processing 32.000 photos is about: 4h30min (but could be wrong)</p>\n<p><a href=\"https://www.kaggle.com/code/remekkinas/fast-dicom-processing\" target=\"_blank\">https://www.kaggle.com/code/remekkinas/fast-dicom-processing</a></p>",
  "messages": [
    {
      "id": 2058150,
      "postDate": "2022-12-07T16:32:10.117Z",
      "content": "<p><strong>This is not my dicovery but decided to create separate topic because this is probably important discovery</strong>. <a href=\"https://www.kaggle.com/kaggleqrdl\" target=\"_blank\">@kaggleqrdl</a> described ( <a href=\"https://www.kaggle.com/alenic\" target=\"_blank\">@alenic</a> idea ) it  here: <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369684#2057282\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369684#2057282</a> I checked only speed of libraries and implemented notebook to show result and provide way to experiment.<br>\nIt seems that we got ~1,6x-2x improvemend in dicom processing speed.</p>\n<p>I created quick implementation how to speed up dicom processing significantly.</p>\n<p>GPU / 500 images (Parallel - 2 jobs):</p>\n<ul>\n<li>pydicom -&gt; 396.74 sec</li>\n<li>dicomsdl -&gt; 243.39 sec </li>\n</ul>\n<p>My estimation is that now processing 32.000 photos is about: 4h30min (but could be wrong)</p>\n<p><a href=\"https://www.kaggle.com/code/remekkinas/fast-dicom-processing\" target=\"_blank\">https://www.kaggle.com/code/remekkinas/fast-dicom-processing</a></p>",
      "rawMarkdown": "**This is not my dicovery but decided to create separate topic because this is probably important discovery**. @kaggleqrdl described ( @alenic idea ) it  here: https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369684#2057282 I checked only speed of libraries and implemented notebook to show result and provide way to experiment.\nIt seems that we got ~1,6x-2x improvemend in dicom processing speed.\n\nI created quick implementation how to speed up dicom processing significantly.\n\nGPU / 500 images (Parallel - 2 jobs):\n- pydicom -> 396.74 sec\n- dicomsdl -> 243.39 sec \n\nMy estimation is that now processing 32.000 photos is about: 4h30min (but could be wrong)\n\nhttps://www.kaggle.com/code/remekkinas/fast-dicom-processing\n\n\n",
      "votes": 49
    },
    {
      "id": 2058426,
      "postDate": "2022-12-08T00:04:56.313Z",
      "content": "<p>if you want to keep the previous code the same:</p>\n<pre><code>    # https://github.com/tsangel/dicomsdl/blob/master/tutorials/timeit_test.ipynb\n    dicom = pydicom.dcmread(f)\n    image0 = dicom.pixel_array\n\n    dicom1 = dicomsdl.open(f)\n    image1 = dicom1.pixelData(storedvalue=True)\n\n    #---- check\n    diff = image0.astype(np.float32)-image1.astype(np.float32)\n    print(diff.max(), diff.mean())\n\n    plt.scatter(image1.reshape(-1), image0.reshape(-1))\n    plt.show()\n</code></pre>\n<p>to improve cpu to gpu transfer (e.g. for large 1024 or 2048) at inference, use 8bit byte and single channel. <br>\nconvert to float() and expand to 3 channel in your model forward function (which is in cuda)</p>",
      "rawMarkdown": "if you want to keep the previous code the same:\n\n```\n    # https://github.com/tsangel/dicomsdl/blob/master/tutorials/timeit_test.ipynb\n    dicom = pydicom.dcmread(f)\n    image0 = dicom.pixel_array\n\n    dicom1 = dicomsdl.open(f)\n    image1 = dicom1.pixelData(storedvalue=True)\n \n    #---- check\n    diff = image0.astype(np.float32)-image1.astype(np.float32)\n    print(diff.max(), diff.mean())\n    \n    plt.scatter(image1.reshape(-1), image0.reshape(-1))\n    plt.show()\n\n```\n\nto improve cpu to gpu transfer (e.g. for large 1024 or 2048) at inference, use 8bit byte and single channel. \nconvert to float() and expand to 3 channel in your model forward function (which is in cuda)\n",
      "votes": 6,
      "replies": [
        {
          "id": 2060484,
          "postDate": "2022-12-10T02:02:59.553Z",
          "content": "<p>refer to python/dicomsdl/__init__.py,  you can modify </p>\n<pre><code>/python/dicomsdl/__init__.py\ndef __dataset__to_pil_image(self, index=0):\n</code></pre>\n<ol>\n<li>change min, max for normalistion  so that winding is not used. This ensures that the converted image is the same as previous pydicom code.</li>\n<li>output as np array instead of PIL image</li>\n</ol>\n<p>or better still, use</p>\n<pre><code>__dataset__pixelData__\n\n storedvalue (bool): True for get stored values; pixel values before LUT\n                        transformation using RescaleSlope and RescaleIntercept.\n  Returns:\n    Numpy array containing pixel values of `index`'th image if dataset holds\n    multiframe data. If `storedvalue` is False, RescaleSlope and\n    RescaleIntercept are applied to pixel values.\n    ( pixel values = stored values * RescsaleSlope + RescaleIntercept )\n</code></pre>\n<p>further, the function below is c++ code and may be faster than numpy?</p>\n<pre><code>util.convert_to_uint8(outarr, data8, xmin, xmax)\n</code></pre>",
          "rawMarkdown": "refer to python/dicomsdl/\\__init__.py,  you can modify \n```\n/python/dicomsdl/__init__.py\ndef __dataset__to_pil_image(self, index=0):\n\n```\n1. change min, max for normalistion  so that winding is not used. This ensures that the converted image is the same as previous pydicom code.\n2. output as np array instead of PIL image\n\n\nor better still, use\n```\n__dataset__pixelData__\n\n storedvalue (bool): True for get stored values; pixel values before LUT\n                        transformation using RescaleSlope and RescaleIntercept.\n  Returns:\n    Numpy array containing pixel values of `index`'th image if dataset holds\n    multiframe data. If `storedvalue` is False, RescaleSlope and\n    RescaleIntercept are applied to pixel values.\n    ( pixel values = stored values * RescsaleSlope + RescaleIntercept )\n```\n\n\nfurther, the function below is c++ code and may be faster than numpy?\n```\nutil.convert_to_uint8(outarr, data8, xmin, xmax)\n```",
          "votes": 1
        },
        {
          "id": 2060676,
          "postDate": "2022-12-10T09:35:14.153Z",
          "content": "<p>Thank you for your feedback - I will look into code and improve it. <br>\nNow (new version of notebook) it seems for me that images generated by dicomsdl are close to the pydicom - but this is observation based on model scoring only (test validation for both - pydicom and discomsdl). It increased speed significantly so now even 6 models (eg. resnet50d) on blend finished in 6h.</p>",
          "rawMarkdown": "Thank you for your feedback - I will look into code and improve it. \nNow (new version of notebook) it seems for me that images generated by dicomsdl are close to the pydicom - but this is observation based on model scoring only (test validation for both - pydicom and discomsdl). It increased speed significantly so now even 6 models (eg. resnet50d) on blend finished in 6h."
        },
        {
          "id": 2060688,
          "postDate": "2022-12-10T09:56:38.163Z",
          "content": "<p>i verify that the difference is 1, -1 for uint8 (range = 0 to 255):</p>\n<p><a href=\"https://www.kaggle.com/code/hengck23/notebook8875e48da3\" target=\"_blank\">https://www.kaggle.com/code/hengck23/notebook8875e48da3</a></p>",
          "rawMarkdown": "i verify that the difference is 1, -1 for uint8 (range = 0 to 255):\n\nhttps://www.kaggle.com/code/hengck23/notebook8875e48da3",
          "votes": 1
        },
        {
          "id": 2060704,
          "postDate": "2022-12-10T10:28:43.143Z",
          "content": "<p>I see (good comparision) - difference is tiny (if we can say such way in ML 😄). It should not influence on our data.</p>",
          "rawMarkdown": "I see (good comparision) - difference is tiny (if we can say such way in ML 😄). It should not influence on our data."
        }
      ]
    },
    {
      "id": 2093811,
      "postDate": "2023-01-10T10:38:24.693Z",
      "content": "<p>Hi there!<br>\nThank you <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> for finding this awesome DicomSDL!</p>\n<p>Now, I might be wrong, but I think that the only problem with DicomSDL is that it misses the apply_voi_lut (or in this competition's case), apply_windowing functions that Pydicom has.</p>\n<p>I've dug into the source code of Pydicom to see what happens exactly, and I've posted some <a href=\"https://www.kaggle.com/code/bobdegraaf/dicomsdl-voi-lut\" target=\"_blank\">results here</a>.</p>\n<p>I think that if you use DicomSDL, you should check it out! :)</p>",
      "rawMarkdown": "Hi there!\nThank you @remekkinas for finding this awesome DicomSDL!\n\nNow, I might be wrong, but I think that the only problem with DicomSDL is that it misses the apply_voi_lut (or in this competition's case), apply_windowing functions that Pydicom has.\n\nI've dug into the source code of Pydicom to see what happens exactly, and I've posted some [results here](https://www.kaggle.com/code/bobdegraaf/dicomsdl-voi-lut).\n\nI think that if you use DicomSDL, you should check it out! :)",
      "votes": 3,
      "replies": [
        {
          "id": 2093813,
          "postDate": "2023-01-10T10:42:02.333Z",
          "content": "<p>Thank you for your contribution. Yes, you are right - I was trying to apply lut but with no effect. Thank you for providing solution. 👍</p>",
          "rawMarkdown": "Thank you for your contribution. Yes, you are right - I was trying to apply lut but with no effect. Thank you for providing solution. 👍"
        }
      ]
    },
    {
      "id": 2058156,
      "postDate": "2022-12-07T16:35:24.787Z",
      "content": "<p>All credit goes to <a href=\"https://www.kaggle.com/alenic\" target=\"_blank\">@alenic</a> .. nice call out!  This is what makes Kaggle great.  </p>\n<p>It's worth noting that dicomsdl decompresses into 2x the amount of bytes that pidicom does.  Does this mean that dicomsdl is more accurate or more noisy?  Or is there some redundancy in there?</p>\n<pre><code>st = time.time()\npa = dset.pixelData()\n(time.time() - st, pa.nbytes)\nst = time.time()\nimage_dataset = dcmread(f).pixel_array\n(time.time() - st, image_dataset.nbytes*)\n</code></pre>\n<p>1.2884056568145752 105279300<br>\n1.5564007759094238 105279300</p>",
      "rawMarkdown": "All credit goes to @alenic .. nice call out!  This is what makes Kaggle great.  \n\nIt's worth noting that dicomsdl decompresses into 2x the amount of bytes that pidicom does.  Does this mean that dicomsdl is more accurate or more noisy?  Or is there some redundancy in there?\n\n```python\nst = time.time()\npa = dset.pixelData()\nprint(time.time() - st, pa.nbytes)\nst = time.time()\nimage_dataset = dcmread(f).pixel_array\nprint(time.time() - st, image_dataset.nbytes*2)\n```\n1.2884056568145752 105279300\n1.5564007759094238 105279300",
      "votes": 3,
      "replies": [
        {
          "id": 2058160,
          "postDate": "2022-12-07T16:38:07.057Z",
          "content": "<p>I understad. I changed in description.<br>\nBut … your work is really awesome. This was nightmare to process files more then 60% of submission time.<br>\nGreat work guys!</p>",
          "rawMarkdown": "I understad. I changed in description.\nBut ... your work is really awesome. This was nightmare to process files more then 60% of submission time.\nGreat work guys!",
          "votes": 1
        }
      ]
    },
    {
      "id": 2060420,
      "postDate": "2022-12-09T22:25:19.490Z",
      "content": "<p>Hi, </p>\n<p>I just tried it for my submission (that failed because I saved the submission.csv in the wrong file 😑) and it took 4 hours with your script. With the previous script I used with pydicom it took 6-7 hours. So there indeed is a very nice improvement !</p>\n<p>Thanks !</p>",
      "rawMarkdown": "Hi, \n\nI just tried it for my submission (that failed because I saved the submission.csv in the wrong file 😑) and it took 4 hours with your script. With the previous script I used with pydicom it took 6-7 hours. So there indeed is a very nice improvement !\n\nThanks !",
      "votes": 4,
      "replies": [
        {
          "id": 2060679,
          "postDate": "2022-12-10T09:40:51.373Z",
          "content": "<p>Good to know! I managed to save a lot of time in my inference script as well. Look it works and we have a lot of time for next stages of pipeline.</p>",
          "rawMarkdown": "Good to know! I managed to save a lot of time in my inference script as well. Look it works and we have a lot of time for next stages of pipeline.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2091919,
      "postDate": "2023-01-08T21:09:11.683Z",
      "content": "<p>I tried my preprocessing with dicomsdl and the loading is really 1.6 times faster! Thank you <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a><br>\nUnfortunately I also wanted to use the functions from pydicom apply_windowing/apply_voi_lut but they do not work with the files loaded with dicomsdl. The dicomsdl library doesn't contain the code for the similar functions. 😟<br>\nDid you manage to use the windowing function or did you write a custom one?</p>",
      "rawMarkdown": "I tried my preprocessing with dicomsdl and the loading is really 1.6 times faster! Thank you @remekkinas\nUnfortunately I also wanted to use the functions from pydicom apply_windowing/apply_voi_lut but they do not work with the files loaded with dicomsdl. The dicomsdl library doesn't contain the code for the similar functions. 😟\nDid you manage to use the windowing function or did you write a custom one?",
      "votes": 1,
      "replies": [
        {
          "id": 2091943,
          "postDate": "2023-01-08T21:46:47.820Z",
          "content": "<p>It does not support. I was experimenting with this but unfortunately have not managed to apply windowing. 😬</p>",
          "rawMarkdown": "It does not support. I was experimenting with this but unfortunately have not managed to apply windowing. 😬",
          "replies": [
            {
              "id": 2091957,
              "postDate": "2023-01-08T22:10:18.107Z",
              "content": "<p><a href=\"https://www.kaggle.com/paulbacher\" target=\"_blank\">@paulbacher</a> <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> </p>\n<p>Looking at the <a href=\"https://github.com/pydicom/pydicom/blob/54c5e493a299e4853bfb46de2e0ddfdd0d285e92/pydicom/pixel_data_handlers/util.py#L468\" target=\"_blank\">source code here</a> it looks pretty easy to re-implement. You can use dicomsdl to grab the same attributes and do the calculation. I was able to do it within the <a href=\"https://www.kaggle.com/code/outwrest/yolov5-roi-batch-dali-preprocessing-pipeline\" target=\"_blank\">same kernel</a> I released before and build on top of it. I fed attributes through and did the calculation on the GPU. The great think about DALI is that I am able to write python functions using CuPy which is basically NumPy but GPU-accelerated so all the data remains on the GPU. I saw only a little minor decrease in throughput.</p>\n<p>Best of luck!</p>",
              "rawMarkdown": "@paulbacher @remekkinas \n\nLooking at the [source code here](https://github.com/pydicom/pydicom/blob/54c5e493a299e4853bfb46de2e0ddfdd0d285e92/pydicom/pixel_data_handlers/util.py#L468) it looks pretty easy to re-implement. You can use dicomsdl to grab the same attributes and do the calculation. I was able to do it within the [same kernel](https://www.kaggle.com/code/outwrest/yolov5-roi-batch-dali-preprocessing-pipeline) I released before and build on top of it. I fed attributes through and did the calculation on the GPU. The great think about DALI is that I am able to write python functions using CuPy which is basically NumPy but GPU-accelerated so all the data remains on the GPU. I saw only a little minor decrease in throughput.\n\nBest of luck!",
              "votes": 3
            },
            {
              "id": 2091965,
              "postDate": "2023-01-08T22:25:56.473Z",
              "content": "<pre><code># dicomsdl reader\ndef normalised_to_8bit(image, photometric_interpretation):\n    xmin = image.min()\n    xmax = image.max() \n    norm = np.empty_like(image, dtype=np.uint8)\n    dicomsdl.util.convert_to_uint8(image, norm, xmin, xmax)\n    if photometric_interpretation == 'MONOCHROME1':\n        norm = 255 - norm\n    return norm\n\ndef dicomsdl_to_numpy_image(ds, index=0): \n    info = ds.getPixelDataInfo()\n    if info['SamplesPerPixel'] != 1:\n        raise RuntimeError('SamplesPerPixel != 1')\n    shape = [info['Rows'], info['Cols']]\n    dtype = info['dtype']\n    outarr = np.empty(shape, dtype=dtype)\n    ds.copyFrameData(index, outarr)\n    return outarr\n\ndef dicomsdl_parallel_process(d, dcm_dir, image_dir, image_height, is_voi_lut):\n    dcm_file = f'{dcm_dir}/{d.patient_id}/{d.image_id}.dcm'\n    ds = dicomsdl.open(dcm_file)\n    image = dicomsdl_to_numpy_image(ds)\n    if is_voi_lut:\n        dc = pydicom.dcmread(dcm_file)\n        image = apply_voi_lut(image, dc)\n        image = image.astype(np.float32)\n    image = normalised_to_8bit(image, ds.PhotometricInterpretation)  # +1\n\n    # resize and save as png\n    ...\n\nParallel(n_jobs=n_jobs)(\n        delayed(dicomsdl_parallel_process)(d, dcm_dir, image_dir, image_height, is_voi_lut)\n        for t,d in tqdm(df.iterrows())\n )\n</code></pre>",
              "rawMarkdown": "```\n# dicomsdl reader\ndef normalised_to_8bit(image, photometric_interpretation):\n    xmin = image.min()\n    xmax = image.max() \n    norm = np.empty_like(image, dtype=np.uint8)\n    dicomsdl.util.convert_to_uint8(image, norm, xmin, xmax)\n    if photometric_interpretation == 'MONOCHROME1':\n        norm = 255 - norm\n    return norm\n\ndef dicomsdl_to_numpy_image(ds, index=0): \n    info = ds.getPixelDataInfo()\n    if info['SamplesPerPixel'] != 1:\n        raise RuntimeError('SamplesPerPixel != 1')\n    shape = [info['Rows'], info['Cols']]\n    dtype = info['dtype']\n    outarr = np.empty(shape, dtype=dtype)\n    ds.copyFrameData(index, outarr)\n    return outarr\n\ndef dicomsdl_parallel_process(d, dcm_dir, image_dir, image_height, is_voi_lut):\n    dcm_file = f'{dcm_dir}/{d.patient_id}/{d.image_id}.dcm'\n    ds = dicomsdl.open(dcm_file)\n    image = dicomsdl_to_numpy_image(ds)\n    if is_voi_lut:\n        dc = pydicom.dcmread(dcm_file)\n        image = apply_voi_lut(image, dc)\n        image = image.astype(np.float32)\n    image = normalised_to_8bit(image, ds.PhotometricInterpretation)  # +1\n\n    # resize and save as png\n    ...\n\nParallel(n_jobs=n_jobs)(\n        delayed(dicomsdl_parallel_process)(d, dcm_dir, image_dir, image_height, is_voi_lut)\n        for t,d in tqdm(df.iterrows())\n )\n\n```\n",
              "votes": 1
            },
            {
              "id": 2092188,
              "postDate": "2023-01-09T07:38:17.457Z",
              "content": "<p>Great! Thank you! 👍</p>",
              "rawMarkdown": "Great! Thank you! 👍"
            },
            {
              "id": 2092475,
              "postDate": "2023-01-09T11:07:23.117Z",
              "content": "<p><a href=\"https://www.kaggle.com/outwrest\" target=\"_blank\">@outwrest</a> Thank you, this is what I did and it's solved 👍<br>\n<a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>  Thank you for this sharing 👍</p>",
              "rawMarkdown": "@outwrest Thank you, this is what I did and it's solved 👍\n@hengck23  Thank you for this sharing 👍"
            }
          ]
        }
      ]
    },
    {
      "id": 2058675,
      "postDate": "2022-12-08T05:54:43.030Z",
      "content": "<p>You can look at my notebook having multiple packages (latest) for offline use. I have added \"dicomsdl\" to it as well.</p>\n<p>Here is my <a href=\"https://www.kaggle.com/code/hey24sheep/frozen-packages-for-offline-use\" target=\"_blank\">notebook</a></p>",
      "rawMarkdown": "You can look at my notebook having multiple packages (latest) for offline use. I have added \"dicomsdl\" to it as well.\n\nHere is my [notebook](https://www.kaggle.com/code/hey24sheep/frozen-packages-for-offline-use)",
      "votes": 1
    },
    {
      "id": 2058198,
      "postDate": "2022-12-07T17:14:44.993Z",
      "content": "<p>I think you also should include normalization in the dicomsdl function, so the images match per-pixel, but maybe it's unnecesary since the processing after JPEG decoding takes little time, nice!</p>",
      "rawMarkdown": "I think you also should include normalization in the dicomsdl function, so the images match per-pixel, but maybe it's unnecesary since the processing after JPEG decoding takes little time, nice!",
      "votes": 1,
      "replies": [
        {
          "id": 2060705,
          "postDate": "2022-12-10T10:30:18.007Z",
          "content": "<p><a href=\"https://www.kaggle.com/martynoveduard\" target=\"_blank\">@martynoveduard</a> than you for your commend. Exactly! I changed it in new notebook. <br>\nGood submission! Great score!</p>",
          "rawMarkdown": "@martynoveduard than you for your commend. Exactly! I changed it in new notebook. \nGood submission! Great score!"
        }
      ]
    },
    {
      "id": 2058788,
      "postDate": "2022-12-08T08:01:24Z",
      "content": "<p>What do you think about further speeding it up by using an CPU notebook and therefore doubling the number of workers? The inference should be reasonably fast on a CPU too….</p>",
      "rawMarkdown": "What do you think about further speeding it up by using an CPU notebook and therefore doubling the number of workers? The inference should be reasonably fast on a CPU too....",
      "votes": 1,
      "replies": [
        {
          "id": 2060680,
          "postDate": "2022-12-10T09:43:18.430Z",
          "content": "<p>I checked it - you are right CPU allow for 4 concurent processes. It speed up file processing twice but … model blending and TTA will require a lot of GPU resources to process 32.000 images. I will stay with GPU.</p>",
          "rawMarkdown": "I checked it - you are right CPU allow for 4 concurent processes. It speed up file processing twice but ... model blending and TTA will require a lot of GPU resources to process 32.000 images. I will stay with GPU."
        }
      ]
    },
    {
      "id": 2160767,
      "postDate": "2023-02-27T00:56:26.923Z",
      "content": "<p>Hi there!<br>\nThank you <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> for finding this awesome DicomSDL!<br>\nI have just one question and any help well be highly appreciated! </p>\n<p>When I run '!pip list' in the cell, it shows that 'dicomsdl' was successfully downloaded and was one of my python site-packages; however, when I run the whole notebook by clicking either 'save version' or 'submit', it always raiseed: ModuleNotFoundError: No module named 'dicomsdl'.</p>\n<p>I wonder how you solve this problem? Thank you so much!</p>",
      "rawMarkdown": "Hi there!\nThank you @remekkinas for finding this awesome DicomSDL!\nI have just one question and any help well be highly appreciated! \n\nWhen I run '!pip list' in the cell, it shows that 'dicomsdl' was successfully downloaded and was one of my python site-packages; however, when I run the whole notebook by clicking either 'save version' or 'submit', it always raiseed: ModuleNotFoundError: No module named 'dicomsdl'.\n\nI wonder how you solve this problem? Thank you so much!"
    },
    {
      "id": 2058251,
      "postDate": "2022-12-07T18:27:39.133Z",
      "content": "<p>Hey Remek, what do you think about a pipeline which utilizes both GPU and processors at the same time?</p>\n<p>My thinking is that with the current pipelines being used, we are idle for several hours on the GPU so we aren't optimizing our compute usage.  </p>\n<p>eg:</p>\n<p><a href=\"https://www.kaggle.com/code/vslaykovsky/infer-effnetv2-aux-targets-weighted-loss-thres\" target=\"_blank\">https://www.kaggle.com/code/vslaykovsky/infer-effnetv2-aux-targets-weighted-loss-thres</a></p>\n<p>This hits the job lib and doesn't do anything else until all the images have been resized.</p>\n<p>I think the architecture would be something like a job pool filled with conversion / augmentation / feeding images to the GPU.  </p>\n<p>I'm not a big GPU expert so it'd be great if folks can chime in here.  I know there is a IO bottleneck with the pipeline between cpu and GPU, and there are also bottlenecks in CPU memory, but given the disk IO happening I don't think the latter is too much of an issue.  Not sure what kind of bottleneck can occur between sending data to GPU / decoding at the same time.  CPU caches can be problematic when it comes to performance.</p>\n<p>At the very least, if done right, this submission time should be &lt;= CPU usage time, with all GPU time being free as its happening during CPU utilization.</p>\n<p>How much of a win that would be, I'm not sure.</p>",
      "rawMarkdown": "Hey Remek, what do you think about a pipeline which utilizes both GPU and processors at the same time?\n\nMy thinking is that with the current pipelines being used, we are idle for several hours on the GPU so we aren't optimizing our compute usage.  \n\neg:\n\nhttps://www.kaggle.com/code/vslaykovsky/infer-effnetv2-aux-targets-weighted-loss-thres\n\nThis hits the job lib and doesn't do anything else until all the images have been resized.\n\n\nI think the architecture would be something like a job pool filled with conversion / augmentation / feeding images to the GPU.  \n\nI'm not a big GPU expert so it'd be great if folks can chime in here.  I know there is a IO bottleneck with the pipeline between cpu and GPU, and there are also bottlenecks in CPU memory, but given the disk IO happening I don't think the latter is too much of an issue.  Not sure what kind of bottleneck can occur between sending data to GPU / decoding at the same time.  CPU caches can be problematic when it comes to performance.\n\nAt the very least, if done right, this submission time should be <= CPU usage time, with all GPU time being free as its happening during CPU utilization.\n\nHow much of a win that would be, I'm not sure.\n\n",
      "replies": [
        {
          "id": 2058269,
          "postDate": "2022-12-07T18:46:55.680Z",
          "content": "<p>nvidia has dali, which seems to be along these lines.</p>\n<p><a href=\"https://developer.nvidia.com/blog/rapid-data-pre-processing-with-nvidia-dali/\" target=\"_blank\">https://developer.nvidia.com/blog/rapid-data-pre-processing-with-nvidia-dali/</a></p>\n<p>more about dali from our friends at tds:<br>\n<a href=\"https://towardsdatascience.com/overcoming-data-preprocessing-bottlenecks-with-tensorflow-data-service-nvidia-dali-and-other-d6321917f851\" target=\"_blank\">https://towardsdatascience.com/overcoming-data-preprocessing-bottlenecks-with-tensorflow-data-service-nvidia-dali-and-other-d6321917f851</a></p>\n<blockquote>\n  <p>A CPU bottleneck occurs when the GPU resource is under utilized as a result of one, or more of the CPUs, having reached maximum utilization. In this situation, the GPU will be partially idle while it waits for the CPU to pass in training data. This is an undesired state. Being that the GPU is, typically, the most expensive resource in the system, your goal should always be to maximize its utilization.</p>\n</blockquote>\n<p>notebook here - <br>\n<a href=\"https://www.kaggle.com/code/hirune924/nvidia-dali-the-fastest-data-loading\" target=\"_blank\">https://www.kaggle.com/code/hirune924/nvidia-dali-the-fastest-data-loading</a></p>",
          "rawMarkdown": "nvidia has dali, which seems to be along these lines.\n\nhttps://developer.nvidia.com/blog/rapid-data-pre-processing-with-nvidia-dali/\n\nmore about dali from our friends at tds:\nhttps://towardsdatascience.com/overcoming-data-preprocessing-bottlenecks-with-tensorflow-data-service-nvidia-dali-and-other-d6321917f851\n\n> A CPU bottleneck occurs when the GPU resource is under utilized as a result of one, or more of the CPUs, having reached maximum utilization. In this situation, the GPU will be partially idle while it waits for the CPU to pass in training data. This is an undesired state. Being that the GPU is, typically, the most expensive resource in the system, your goal should always be to maximize its utilization.\n\n\nnotebook here - \nhttps://www.kaggle.com/code/hirune924/nvidia-dali-the-fastest-data-loading"
        },
        {
          "id": 2058327,
          "postDate": "2022-12-07T20:11:08.297Z",
          "content": "<p>I am working on it. First test is promising 😄</p>",
          "rawMarkdown": "I am working on it. First test is promising 😄",
          "votes": 1
        },
        {
          "id": 2058346,
          "postDate": "2022-12-07T20:34:24.263Z",
          "content": "<p>That's really awesome.  Should be cool.  I know psi was concerned about this.  Not sure if he's using dali.</p>",
          "rawMarkdown": "That's really awesome.  Should be cool.  I know psi was concerned about this.  Not sure if he's using dali."
        },
        {
          "id": 2058386,
          "postDate": "2022-12-07T21:51:23.860Z",
          "content": "<p>Now I am prototyping using Ray.<br>\nI will look into Dali for sure.</p>",
          "rawMarkdown": "Now I am prototyping using Ray.\nI will look into Dali for sure.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2059770,
      "postDate": "2022-12-09T07:51:38.443Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2058426,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-12-08T00:04:56.313000",
      "content": "<p>if you want to keep the previous code the same:</p>\n<pre><code>    # https://github.com/tsangel/dicomsdl/blob/master/tutorials/timeit_test.ipynb\n    dicom = pydicom.dcmread(f)\n    image0 = dicom.pixel_array\n\n    dicom1 = dicomsdl.open(f)\n    image1 = dicom1.pixelData(storedvalue=True)\n\n    #---- check\n    diff = image0.astype(np.float32)-image1.astype(np.float32)\n    print(diff.max(), diff.mean())\n\n    plt.scatter(image1.reshape(-1), image0.reshape(-1))\n    plt.show()\n</code></pre>\n<p>to improve cpu to gpu transfer (e.g. for large 1024 or 2048) at inference, use 8bit byte and single channel. <br>\nconvert to float() and expand to 3 channel in your model forward function (which is in cuda)</p>",
      "votes": 6,
      "replies": [
        {
          "id": 2060484,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2022-12-10T02:02:59.553000",
          "content": "<p>refer to python/dicomsdl/__init__.py,  you can modify </p>\n<pre><code>/python/dicomsdl/__init__.py\ndef __dataset__to_pil_image(self, index=0):\n</code></pre>\n<ol>\n<li>change min, max for normalistion  so that winding is not used. This ensures that the converted image is the same as previous pydicom code.</li>\n<li>output as np array instead of PIL image</li>\n</ol>\n<p>or better still, use</p>\n<pre><code>__dataset__pixelData__\n\n storedvalue (bool): True for get stored values; pixel values before LUT\n                        transformation using RescaleSlope and RescaleIntercept.\n  Returns:\n    Numpy array containing pixel values of `index`'th image if dataset holds\n    multiframe data. If `storedvalue` is False, RescaleSlope and\n    RescaleIntercept are applied to pixel values.\n    ( pixel values = stored values * RescsaleSlope + RescaleIntercept )\n</code></pre>\n<p>further, the function below is c++ code and may be faster than numpy?</p>\n<pre><code>util.convert_to_uint8(outarr, data8, xmin, xmax)\n</code></pre>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2060676,
          "author_name": "Remek Kinas",
          "author_url": "",
          "post_date": "2022-12-10T09:35:14.153000",
          "content": "<p>Thank you for your feedback - I will look into code and improve it. <br>\nNow (new version of notebook) it seems for me that images generated by dicomsdl are close to the pydicom - but this is observation based on model scoring only (test validation for both - pydicom and discomsdl). It increased speed significantly so now even 6 models (eg. resnet50d) on blend finished in 6h.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2060688,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2022-12-10T09:56:38.163000",
          "content": "<p>i verify that the difference is 1, -1 for uint8 (range = 0 to 255):</p>\n<p><a href=\"https://www.kaggle.com/code/hengck23/notebook8875e48da3\" target=\"_blank\">https://www.kaggle.com/code/hengck23/notebook8875e48da3</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2060704,
          "author_name": "Remek Kinas",
          "author_url": "",
          "post_date": "2022-12-10T10:28:43.143000",
          "content": "<p>I see (good comparision) - difference is tiny (if we can say such way in ML 😄). It should not influence on our data.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2093811,
      "author_name": "Bob de Graaf",
      "author_url": "",
      "post_date": "2023-01-10T10:38:24.693000",
      "content": "<p>Hi there!<br>\nThank you <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> for finding this awesome DicomSDL!</p>\n<p>Now, I might be wrong, but I think that the only problem with DicomSDL is that it misses the apply_voi_lut (or in this competition's case), apply_windowing functions that Pydicom has.</p>\n<p>I've dug into the source code of Pydicom to see what happens exactly, and I've posted some <a href=\"https://www.kaggle.com/code/bobdegraaf/dicomsdl-voi-lut\" target=\"_blank\">results here</a>.</p>\n<p>I think that if you use DicomSDL, you should check it out! :)</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2093813,
          "author_name": "Remek Kinas",
          "author_url": "",
          "post_date": "2023-01-10T10:42:02.333000",
          "content": "<p>Thank you for your contribution. Yes, you are right - I was trying to apply lut but with no effect. Thank you for providing solution. 👍</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2058156,
      "author_name": "@kaggleqrdl",
      "author_url": "",
      "post_date": "2022-12-07T16:35:24.787000",
      "content": "<p>All credit goes to <a href=\"https://www.kaggle.com/alenic\" target=\"_blank\">@alenic</a> .. nice call out!  This is what makes Kaggle great.  </p>\n<p>It's worth noting that dicomsdl decompresses into 2x the amount of bytes that pidicom does.  Does this mean that dicomsdl is more accurate or more noisy?  Or is there some redundancy in there?</p>\n<pre><code>st = time.time()\npa = dset.pixelData()\n(time.time() - st, pa.nbytes)\nst = time.time()\nimage_dataset = dcmread(f).pixel_array\n(time.time() - st, image_dataset.nbytes*)\n</code></pre>\n<p>1.2884056568145752 105279300<br>\n1.5564007759094238 105279300</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2058160,
          "author_name": "Remek Kinas",
          "author_url": "",
          "post_date": "2022-12-07T16:38:07.057000",
          "content": "<p>I understad. I changed in description.<br>\nBut … your work is really awesome. This was nightmare to process files more then 60% of submission time.<br>\nGreat work guys!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2060420,
      "author_name": "Natyu",
      "author_url": "",
      "post_date": "2022-12-09T22:25:19.490000",
      "content": "<p>Hi, </p>\n<p>I just tried it for my submission (that failed because I saved the submission.csv in the wrong file 😑) and it took 4 hours with your script. With the previous script I used with pydicom it took 6-7 hours. So there indeed is a very nice improvement !</p>\n<p>Thanks !</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2060679,
          "author_name": "Remek Kinas",
          "author_url": "",
          "post_date": "2022-12-10T09:40:51.373000",
          "content": "<p>Good to know! I managed to save a lot of time in my inference script as well. Look it works and we have a lot of time for next stages of pipeline.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2091919,
      "author_name": "Paul Bacher",
      "author_url": "",
      "post_date": "2023-01-08T21:09:11.683000",
      "content": "<p>I tried my preprocessing with dicomsdl and the loading is really 1.6 times faster! Thank you <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a><br>\nUnfortunately I also wanted to use the functions from pydicom apply_windowing/apply_voi_lut but they do not work with the files loaded with dicomsdl. The dicomsdl library doesn't contain the code for the similar functions. 😟<br>\nDid you manage to use the windowing function or did you write a custom one?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2091943,
          "author_name": "Remek Kinas",
          "author_url": "",
          "post_date": "2023-01-08T21:46:47.820000",
          "content": "<p>It does not support. I was experimenting with this but unfortunately have not managed to apply windowing. 😬</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2091957,
              "author_name": "outwrest",
              "author_url": "",
              "post_date": "2023-01-08T22:10:18.107000",
              "content": "<p><a href=\"https://www.kaggle.com/paulbacher\" target=\"_blank\">@paulbacher</a> <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> </p>\n<p>Looking at the <a href=\"https://github.com/pydicom/pydicom/blob/54c5e493a299e4853bfb46de2e0ddfdd0d285e92/pydicom/pixel_data_handlers/util.py#L468\" target=\"_blank\">source code here</a> it looks pretty easy to re-implement. You can use dicomsdl to grab the same attributes and do the calculation. I was able to do it within the <a href=\"https://www.kaggle.com/code/outwrest/yolov5-roi-batch-dali-preprocessing-pipeline\" target=\"_blank\">same kernel</a> I released before and build on top of it. I fed attributes through and did the calculation on the GPU. The great think about DALI is that I am able to write python functions using CuPy which is basically NumPy but GPU-accelerated so all the data remains on the GPU. I saw only a little minor decrease in throughput.</p>\n<p>Best of luck!</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2091965,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-01-08T22:25:56.473000",
              "content": "<pre><code># dicomsdl reader\ndef normalised_to_8bit(image, photometric_interpretation):\n    xmin = image.min()\n    xmax = image.max() \n    norm = np.empty_like(image, dtype=np.uint8)\n    dicomsdl.util.convert_to_uint8(image, norm, xmin, xmax)\n    if photometric_interpretation == 'MONOCHROME1':\n        norm = 255 - norm\n    return norm\n\ndef dicomsdl_to_numpy_image(ds, index=0): \n    info = ds.getPixelDataInfo()\n    if info['SamplesPerPixel'] != 1:\n        raise RuntimeError('SamplesPerPixel != 1')\n    shape = [info['Rows'], info['Cols']]\n    dtype = info['dtype']\n    outarr = np.empty(shape, dtype=dtype)\n    ds.copyFrameData(index, outarr)\n    return outarr\n\ndef dicomsdl_parallel_process(d, dcm_dir, image_dir, image_height, is_voi_lut):\n    dcm_file = f'{dcm_dir}/{d.patient_id}/{d.image_id}.dcm'\n    ds = dicomsdl.open(dcm_file)\n    image = dicomsdl_to_numpy_image(ds)\n    if is_voi_lut:\n        dc = pydicom.dcmread(dcm_file)\n        image = apply_voi_lut(image, dc)\n        image = image.astype(np.float32)\n    image = normalised_to_8bit(image, ds.PhotometricInterpretation)  # +1\n\n    # resize and save as png\n    ...\n\nParallel(n_jobs=n_jobs)(\n        delayed(dicomsdl_parallel_process)(d, dcm_dir, image_dir, image_height, is_voi_lut)\n        for t,d in tqdm(df.iterrows())\n )\n</code></pre>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2092188,
              "author_name": "Remek Kinas",
              "author_url": "",
              "post_date": "2023-01-09T07:38:17.457000",
              "content": "<p>Great! Thank you! 👍</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2092475,
              "author_name": "Paul Bacher",
              "author_url": "",
              "post_date": "2023-01-09T11:07:23.117000",
              "content": "<p><a href=\"https://www.kaggle.com/outwrest\" target=\"_blank\">@outwrest</a> Thank you, this is what I did and it's solved 👍<br>\n<a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>  Thank you for this sharing 👍</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2058675,
      "author_name": "Hey24sheep",
      "author_url": "",
      "post_date": "2022-12-08T05:54:43.030000",
      "content": "<p>You can look at my notebook having multiple packages (latest) for offline use. I have added \"dicomsdl\" to it as well.</p>\n<p>Here is my <a href=\"https://www.kaggle.com/code/hey24sheep/frozen-packages-for-offline-use\" target=\"_blank\">notebook</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2058198,
      "author_name": "slime",
      "author_url": "",
      "post_date": "2022-12-07T17:14:44.993000",
      "content": "<p>I think you also should include normalization in the dicomsdl function, so the images match per-pixel, but maybe it's unnecesary since the processing after JPEG decoding takes little time, nice!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2060705,
          "author_name": "Remek Kinas",
          "author_url": "",
          "post_date": "2022-12-10T10:30:18.007000",
          "content": "<p><a href=\"https://www.kaggle.com/martynoveduard\" target=\"_blank\">@martynoveduard</a> than you for your commend. Exactly! I changed it in new notebook. <br>\nGood submission! Great score!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2058788,
      "author_name": "Marius ",
      "author_url": "",
      "post_date": "2022-12-08T08:01:24",
      "content": "<p>What do you think about further speeding it up by using an CPU notebook and therefore doubling the number of workers? The inference should be reasonably fast on a CPU too….</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2060680,
          "author_name": "Remek Kinas",
          "author_url": "",
          "post_date": "2022-12-10T09:43:18.430000",
          "content": "<p>I checked it - you are right CPU allow for 4 concurent processes. It speed up file processing twice but … model blending and TTA will require a lot of GPU resources to process 32.000 images. I will stay with GPU.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2160767,
      "author_name": "Eric006",
      "author_url": "",
      "post_date": "2023-02-27T00:56:26.923000",
      "content": "<p>Hi there!<br>\nThank you <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> for finding this awesome DicomSDL!<br>\nI have just one question and any help well be highly appreciated! </p>\n<p>When I run '!pip list' in the cell, it shows that 'dicomsdl' was successfully downloaded and was one of my python site-packages; however, when I run the whole notebook by clicking either 'save version' or 'submit', it always raiseed: ModuleNotFoundError: No module named 'dicomsdl'.</p>\n<p>I wonder how you solve this problem? Thank you so much!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2058251,
      "author_name": "@kaggleqrdl",
      "author_url": "",
      "post_date": "2022-12-07T18:27:39.133000",
      "content": "<p>Hey Remek, what do you think about a pipeline which utilizes both GPU and processors at the same time?</p>\n<p>My thinking is that with the current pipelines being used, we are idle for several hours on the GPU so we aren't optimizing our compute usage.  </p>\n<p>eg:</p>\n<p><a href=\"https://www.kaggle.com/code/vslaykovsky/infer-effnetv2-aux-targets-weighted-loss-thres\" target=\"_blank\">https://www.kaggle.com/code/vslaykovsky/infer-effnetv2-aux-targets-weighted-loss-thres</a></p>\n<p>This hits the job lib and doesn't do anything else until all the images have been resized.</p>\n<p>I think the architecture would be something like a job pool filled with conversion / augmentation / feeding images to the GPU.  </p>\n<p>I'm not a big GPU expert so it'd be great if folks can chime in here.  I know there is a IO bottleneck with the pipeline between cpu and GPU, and there are also bottlenecks in CPU memory, but given the disk IO happening I don't think the latter is too much of an issue.  Not sure what kind of bottleneck can occur between sending data to GPU / decoding at the same time.  CPU caches can be problematic when it comes to performance.</p>\n<p>At the very least, if done right, this submission time should be &lt;= CPU usage time, with all GPU time being free as its happening during CPU utilization.</p>\n<p>How much of a win that would be, I'm not sure.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2058269,
          "author_name": "@kaggleqrdl",
          "author_url": "",
          "post_date": "2022-12-07T18:46:55.680000",
          "content": "<p>nvidia has dali, which seems to be along these lines.</p>\n<p><a href=\"https://developer.nvidia.com/blog/rapid-data-pre-processing-with-nvidia-dali/\" target=\"_blank\">https://developer.nvidia.com/blog/rapid-data-pre-processing-with-nvidia-dali/</a></p>\n<p>more about dali from our friends at tds:<br>\n<a href=\"https://towardsdatascience.com/overcoming-data-preprocessing-bottlenecks-with-tensorflow-data-service-nvidia-dali-and-other-d6321917f851\" target=\"_blank\">https://towardsdatascience.com/overcoming-data-preprocessing-bottlenecks-with-tensorflow-data-service-nvidia-dali-and-other-d6321917f851</a></p>\n<blockquote>\n  <p>A CPU bottleneck occurs when the GPU resource is under utilized as a result of one, or more of the CPUs, having reached maximum utilization. In this situation, the GPU will be partially idle while it waits for the CPU to pass in training data. This is an undesired state. Being that the GPU is, typically, the most expensive resource in the system, your goal should always be to maximize its utilization.</p>\n</blockquote>\n<p>notebook here - <br>\n<a href=\"https://www.kaggle.com/code/hirune924/nvidia-dali-the-fastest-data-loading\" target=\"_blank\">https://www.kaggle.com/code/hirune924/nvidia-dali-the-fastest-data-loading</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2058327,
          "author_name": "Remek Kinas",
          "author_url": "",
          "post_date": "2022-12-07T20:11:08.297000",
          "content": "<p>I am working on it. First test is promising 😄</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2058346,
          "author_name": "@kaggleqrdl",
          "author_url": "",
          "post_date": "2022-12-07T20:34:24.263000",
          "content": "<p>That's really awesome.  Should be cool.  I know psi was concerned about this.  Not sure if he's using dali.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2058386,
          "author_name": "Remek Kinas",
          "author_url": "",
          "post_date": "2022-12-07T21:51:23.860000",
          "content": "<p>Now I am prototyping using Ray.<br>\nI will look into Dali for sure.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2059770,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-12-09T07:51:38.443000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2058150": "**This is not my dicovery but decided to create separate topic because this is probably important discovery**. @kaggleqrdl described ( @alenic idea ) it  here: https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369684#2057282 I checked only speed of libraries and implemented notebook to show result and provide way to experiment.\nIt seems that we got ~1,6x-2x improvemend in dicom processing speed.\n\nI created quick implementation how to speed up dicom processing significantly.\n\nGPU / 500 images (Parallel - 2 jobs):\n- pydicom -> 396.74 sec\n- dicomsdl -> 243.39 sec \n\nMy estimation is that now processing 32.000 photos is about: 4h30min (but could be wrong)\n\nhttps://www.kaggle.com/code/remekkinas/fast-dicom-processing\n\n\n",
    "2058426": "if you want to keep the previous code the same:\n\n```\n    # https://github.com/tsangel/dicomsdl/blob/master/tutorials/timeit_test.ipynb\n    dicom = pydicom.dcmread(f)\n    image0 = dicom.pixel_array\n\n    dicom1 = dicomsdl.open(f)\n    image1 = dicom1.pixelData(storedvalue=True)\n \n    #---- check\n    diff = image0.astype(np.float32)-image1.astype(np.float32)\n    print(diff.max(), diff.mean())\n    \n    plt.scatter(image1.reshape(-1), image0.reshape(-1))\n    plt.show()\n\n```\n\nto improve cpu to gpu transfer (e.g. for large 1024 or 2048) at inference, use 8bit byte and single channel. \nconvert to float() and expand to 3 channel in your model forward function (which is in cuda)\n",
    "2093811": "Hi there!\nThank you @remekkinas for finding this awesome DicomSDL!\n\nNow, I might be wrong, but I think that the only problem with DicomSDL is that it misses the apply_voi_lut (or in this competition's case), apply_windowing functions that Pydicom has.\n\nI've dug into the source code of Pydicom to see what happens exactly, and I've posted some [results here](https://www.kaggle.com/code/bobdegraaf/dicomsdl-voi-lut).\n\nI think that if you use DicomSDL, you should check it out! :)",
    "2058156": "All credit goes to @alenic .. nice call out!  This is what makes Kaggle great.  \n\nIt's worth noting that dicomsdl decompresses into 2x the amount of bytes that pidicom does.  Does this mean that dicomsdl is more accurate or more noisy?  Or is there some redundancy in there?\n\n```python\nst = time.time()\npa = dset.pixelData()\nprint(time.time() - st, pa.nbytes)\nst = time.time()\nimage_dataset = dcmread(f).pixel_array\nprint(time.time() - st, image_dataset.nbytes*2)\n```\n1.2884056568145752 105279300\n1.5564007759094238 105279300",
    "2060420": "Hi, \n\nI just tried it for my submission (that failed because I saved the submission.csv in the wrong file 😑) and it took 4 hours with your script. With the previous script I used with pydicom it took 6-7 hours. So there indeed is a very nice improvement !\n\nThanks !",
    "2091919": "I tried my preprocessing with dicomsdl and the loading is really 1.6 times faster! Thank you @remekkinas\nUnfortunately I also wanted to use the functions from pydicom apply_windowing/apply_voi_lut but they do not work with the files loaded with dicomsdl. The dicomsdl library doesn't contain the code for the similar functions. 😟\nDid you manage to use the windowing function or did you write a custom one?",
    "2058675": "You can look at my notebook having multiple packages (latest) for offline use. I have added \"dicomsdl\" to it as well.\n\nHere is my [notebook](https://www.kaggle.com/code/hey24sheep/frozen-packages-for-offline-use)",
    "2058198": "I think you also should include normalization in the dicomsdl function, so the images match per-pixel, but maybe it's unnecesary since the processing after JPEG decoding takes little time, nice!",
    "2058788": "What do you think about further speeding it up by using an CPU notebook and therefore doubling the number of workers? The inference should be reasonably fast on a CPU too....",
    "2160767": "Hi there!\nThank you @remekkinas for finding this awesome DicomSDL!\nI have just one question and any help well be highly appreciated! \n\nWhen I run '!pip list' in the cell, it shows that 'dicomsdl' was successfully downloaded and was one of my python site-packages; however, when I run the whole notebook by clicking either 'save version' or 'submit', it always raiseed: ModuleNotFoundError: No module named 'dicomsdl'.\n\nI wonder how you solve this problem? Thank you so much!",
    "2058251": "Hey Remek, what do you think about a pipeline which utilizes both GPU and processors at the same time?\n\nMy thinking is that with the current pipelines being used, we are idle for several hours on the GPU so we aren't optimizing our compute usage.  \n\neg:\n\nhttps://www.kaggle.com/code/vslaykovsky/infer-effnetv2-aux-targets-weighted-loss-thres\n\nThis hits the job lib and doesn't do anything else until all the images have been resized.\n\n\nI think the architecture would be something like a job pool filled with conversion / augmentation / feeding images to the GPU.  \n\nI'm not a big GPU expert so it'd be great if folks can chime in here.  I know there is a IO bottleneck with the pipeline between cpu and GPU, and there are also bottlenecks in CPU memory, but given the disk IO happening I don't think the latter is too much of an issue.  Not sure what kind of bottleneck can occur between sending data to GPU / decoding at the same time.  CPU caches can be problematic when it comes to performance.\n\nAt the very least, if done right, this submission time should be <= CPU usage time, with all GPU time being free as its happening during CPU utilization.\n\nHow much of a win that would be, I'm not sure.\n\n",
    "2059770": ""
  }
}