{
  "id": 372275,
  "title": "Easy load the image with nvJPEG2000(5x faster)",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/372275",
  "author_name": "Chenglu",
  "post_date": "2022-12-15T07:22:19.351000",
  "votes": 31,
  "comment_count": 17,
  "views": 0,
  "content": "<p>I write a <a href=\"https://www.kaggle.com/code/snaker/easy-load-the-image-with-nvjpeg2000\" target=\"_blank\">notebook</a> to demonstrate how to load the dicom file with nvJPEG2000.</p>\n<p>The trick is implemented by Python extension, the source code of that extension is at: <a href=\"https://github.com/louis-she/nvjpeg2k-python\" target=\"_blank\">https://github.com/louis-she/nvjpeg2k-python</a> . I built it with Python3.7 and CUDA11 so that can be used in Kaggle's default environment. For how to compile the extension in local machine, see the readme from the repo.</p>\n<p>Thanks to the post <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/371534\" target=\"_blank\">17x dicom decode speedup on GPU for jpeg2000 encodings</a>.</p>",
  "messages": [
    {
      "id": 2065903,
      "postDate": "2022-12-15T07:22:19.350Z",
      "content": "<p>I write a <a href=\"https://www.kaggle.com/code/snaker/easy-load-the-image-with-nvjpeg2000\" target=\"_blank\">notebook</a> to demonstrate how to load the dicom file with nvJPEG2000.</p>\n<p>The trick is implemented by Python extension, the source code of that extension is at: <a href=\"https://github.com/louis-she/nvjpeg2k-python\" target=\"_blank\">https://github.com/louis-she/nvjpeg2k-python</a> . I built it with Python3.7 and CUDA11 so that can be used in Kaggle's default environment. For how to compile the extension in local machine, see the readme from the repo.</p>\n<p>Thanks to the post <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/371534\" target=\"_blank\">17x dicom decode speedup on GPU for jpeg2000 encodings</a>.</p>",
      "rawMarkdown": "I write a [notebook](https://www.kaggle.com/code/snaker/easy-load-the-image-with-nvjpeg2000) to demonstrate how to load the dicom file with nvJPEG2000.\n\nThe trick is implemented by Python extension, the source code of that extension is at: https://github.com/louis-she/nvjpeg2k-python . I built it with Python3.7 and CUDA11 so that can be used in Kaggle's default environment. For how to compile the extension in local machine, see the readme from the repo.\n\nThanks to the post [17x dicom decode speedup on GPU for jpeg2000 encodings](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/371534).",
      "votes": 30
    },
    {
      "id": 2065982,
      "postDate": "2022-12-15T08:59:31.603Z",
      "content": "<p>** for local compilation only **</p>\n<p>before compile the python extension, you have to verify that nvJPEG2000  is installed correctly.</p>\n<ol>\n<li>download and install nvJPEG2000 from <a href=\"https://developer.nvidia.com/nvjpeg\" target=\"_blank\">https://developer.nvidia.com/nvjpeg</a></li>\n<li>verify correct installation by compile the example at <a href=\"https://github.com/NVIDIA/CUDALibrarySamples/tree/master/nvJPEG2000/nvJPEG2000-Decoder\" target=\"_blank\">https://github.com/NVIDIA/CUDALibrarySamples/tree/master/nvJPEG2000/nvJPEG2000-Decoder</a></li>\n</ol>\n<p>to locate the where you have installed nvJPEG2000, i use:</p>\n<pre><code>locate nvjpeg2k\n/usr/include/nvjpeg2k.h\n/usr/include/nvjpeg2k_version.h\n</code></pre>\n<p>hence my DNVJPEG2K_PATH=/usr</p>\n<p>then,</p>\n<pre><code>mkdir build\ncd build \nexport CUDACXX=nvcc\ncmake ..  -DNVJPEG2K_PATH=/usr\nmake\n</code></pre>\n<p>Finally run the example</p>\n<pre><code>./nvjpeg2000_decode_sample -i ../images/2k_image_lossless/2k_lossless.jp2 -o .\n</code></pre>\n<p>expected output</p>\n<pre><code>Decoding images in directory: ../images/2k_image_lossless/2k_lossless.jp2, total 1, batchsize 1\nTotal decoding time: 0.0242717\nAvg decoding time per image: 0.0242717\nAvg images per sec: 41.2002\nAvg decoding time per batch: 0.0242717\n</code></pre>\n<hr>\n<p>I use pycharm + andaconda. <br>\nTo link to the correct pybind11, </p>\n<pre><code>0. I first check the pybind11 version that is already installed in my system:\n/home/titanx/hengck/opt/anaconda3.9/lib/python3.9/site-packages/pybind11\n1.  git clone the same version from https://github.com/pybind/pybind11 to the direct extern/pybind11\n2 . use the terminal from pycharm IDE interface\n3. run the commands:\n\nexport NVJPEG2K_PATH=/usr\nexport CUDACXX=nvcc\ncmake .. \\\n  -DCMAKE_BUILD_TYPE=Debug \\\n  -DNVJPEG2K_PATH=/usr \\\n  -DNVJPEG2K_LIB=/usr/lib/x86_64-linux-gnu/libnvjpeg2k_static.a\n</code></pre>\n<p>you should see</p>\n<pre><code>-- pybind11 v2.11.0 dev1\n-- Found PythonInterp: /home/titanx/hengck/opt/anaconda3.9/bin/python (found suitable version \"3.9.12\", minimum required is \"3.6\") \n-- Found PythonLibs: /home/titanx/hengck/opt/anaconda3.9/lib/libpython3.9.so\n-- Performing Test HAS_FLTO\n-- Performing Test HAS_FLTO - Success\n\n-- Configuring done\n-- Generating done\n-- Build files have been written to: /home/titanx/hengck/share1/kaggle/2022/rsna-breast-mammography/tool/nvjpeg2k/nvjpeg2k-python-main/build\n</code></pre>\n<p>you can download and decompress:<br>\n<a href=\"https://ubuntu.pkgs.org/22.04/cuda-amd64/libnvjpeg2k0_0.6.0.28-1_amd64.deb.html\" target=\"_blank\">https://ubuntu.pkgs.org/22.04/cuda-amd64/libnvjpeg2k0_0.6.0.28-1_amd64.deb.html</a><br>\n<a href=\"https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2004/x86_64/libnvjpeg2k0_0.6.0.28-1_amd64.deb\" target=\"_blank\">https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2004/x86_64/libnvjpeg2k0_0.6.0.28-1_amd64.deb</a></p>\n<p>or i use this:</p>\n<pre><code>  sudo apt-key adv --fetch-keys http://developer.download.nvidia.com/compute/cuda/repos/ubuntu1804/x86_64/3bf863cc.pub\n  sudo add-apt-repository \"deb http://developer.download.nvidia.com/compute/cuda/repos/ubuntu1804/x86_64/ /\"\n  sudo apt update\n\n  sudo apt-get install libnvjpeg2k0 libnvjpeg2k-dev\n</code></pre>",
      "rawMarkdown": "** for local compilation only **\n\nbefore compile the python extension, you have to verify that nvJPEG2000  is installed correctly.\n\n1. download and install nvJPEG2000 from https://developer.nvidia.com/nvjpeg\n2. verify correct installation by compile the example at https://github.com/NVIDIA/CUDALibrarySamples/tree/master/nvJPEG2000/nvJPEG2000-Decoder\n\n\nto locate the where you have installed nvJPEG2000, i use:\n```\nlocate nvjpeg2k\n/usr/include/nvjpeg2k.h\n/usr/include/nvjpeg2k_version.h\n```\n\nhence my DNVJPEG2K_PATH=/usr\n\nthen,\n\n```\nmkdir build\ncd build \nexport CUDACXX=nvcc\ncmake ..  -DNVJPEG2K_PATH=/usr\nmake\n```\nFinally run the example\n```\n./nvjpeg2000_decode_sample -i ../images/2k_image_lossless/2k_lossless.jp2 -o .\n```\n\nexpected output\n\n```\nDecoding images in directory: ../images/2k_image_lossless/2k_lossless.jp2, total 1, batchsize 1\nTotal decoding time: 0.0242717\nAvg decoding time per image: 0.0242717\nAvg images per sec: 41.2002\nAvg decoding time per batch: 0.0242717\n\n\n```\n\n----\n\nI use pycharm + andaconda. \nTo link to the correct pybind11, \n\n```\n\n0. I first check the pybind11 version that is already installed in my system:\n/home/titanx/hengck/opt/anaconda3.9/lib/python3.9/site-packages/pybind11\n1.  git clone the same version from https://github.com/pybind/pybind11 to the direct extern/pybind11\n2 . use the terminal from pycharm IDE interface\n3. run the commands:\n\nexport NVJPEG2K_PATH=/usr\nexport CUDACXX=nvcc\ncmake .. \\\n  -DCMAKE_BUILD_TYPE=Debug \\\n  -DNVJPEG2K_PATH=/usr \\\n  -DNVJPEG2K_LIB=/usr/lib/x86_64-linux-gnu/libnvjpeg2k_static.a\n\n```\nyou should see\n\n```\n-- pybind11 v2.11.0 dev1\n-- Found PythonInterp: /home/titanx/hengck/opt/anaconda3.9/bin/python (found suitable version \"3.9.12\", minimum required is \"3.6\") \n-- Found PythonLibs: /home/titanx/hengck/opt/anaconda3.9/lib/libpython3.9.so\n-- Performing Test HAS_FLTO\n-- Performing Test HAS_FLTO - Success\n\n-- Configuring done\n-- Generating done\n-- Build files have been written to: /home/titanx/hengck/share1/kaggle/2022/rsna-breast-mammography/tool/nvjpeg2k/nvjpeg2k-python-main/build\n\n\n```\n\nyou can download and decompress:\nhttps://ubuntu.pkgs.org/22.04/cuda-amd64/libnvjpeg2k0_0.6.0.28-1_amd64.deb.html\nhttps://developer.download.nvidia.com/compute/cuda/repos/ubuntu2004/x86_64/libnvjpeg2k0_0.6.0.28-1_amd64.deb\n\nor i use this:\n\n```\n  sudo apt-key adv --fetch-keys http://developer.download.nvidia.com/compute/cuda/repos/ubuntu1804/x86_64/3bf863cc.pub\n  sudo add-apt-repository \"deb http://developer.download.nvidia.com/compute/cuda/repos/ubuntu1804/x86_64/ /\"\n  sudo apt update\n\n  sudo apt-get install libnvjpeg2k0 libnvjpeg2k-dev\n\n```\n",
      "votes": 3
    },
    {
      "id": 2079508,
      "postDate": "2022-12-29T11:25:00.750Z",
      "content": "<p>If one wants to use the extension in the multi-process. Here's the tip:</p>\n<p><strong>Do not</strong> create the <code>Decoder</code> instance in main process, instead, create the instance in sub-processes. There are unexpected behaviors when the <code>Decoder</code> is initialized in a process that will fork sub-processes. </p>",
      "rawMarkdown": "If one wants to use the extension in the multi-process. Here's the tip:\n\n**Do not** create the `Decoder` instance in main process, instead, create the instance in sub-processes. There are unexpected behaviors when the `Decoder` is initialized in a process that will fork sub-processes. ",
      "votes": 2,
      "replies": [
        {
          "id": 2079516,
          "postDate": "2022-12-29T11:36:14.623Z",
          "content": "<p>Creating the decoder in sub-processes could slow down the process of decoding dcm?</p>",
          "rawMarkdown": "Creating the decoder in sub-processes could slow down the process of decoding dcm?",
          "replies": [
            {
              "id": 2079522,
              "postDate": "2022-12-29T11:43:24.637Z",
              "content": "<p>There is no performance difference between creating <code>Decoder</code> in main process or sub process. The reason why we should create the <code>Decoder</code> in subprocesses is to prevent the \"unexpected behavior\".</p>\n<p>The \"unexpected behavior\" is that if create a <code>Decoder</code> in main process and fork sub processes, the decoder will not produce right results, you will see many 0s or 65536s in the decoded image. I don't known the root cause of this cause I'm not a CUDA expert.</p>\n<p>The whole thing only applies to those who wants to use the extension concurrently, cause Kaggle's kernel has 2 CPU cores. If there are no plans to use multi processing, then there are no such issue of the extension.</p>",
              "rawMarkdown": "There is no performance difference between creating `Decoder` in main process or sub process. The reason why we should create the `Decoder` in subprocesses is to prevent the \"unexpected behavior\".\n\nThe \"unexpected behavior\" is that if create a `Decoder` in main process and fork sub processes, the decoder will not produce right results, you will see many 0s or 65536s in the decoded image. I don't known the root cause of this cause I'm not a CUDA expert.\n\nThe whole thing only applies to those who wants to use the extension concurrently, cause Kaggle's kernel has 2 CPU cores. If there are no plans to use multi processing, then there are no such issue of the extension.",
              "votes": 1
            },
            {
              "id": 2079534,
              "postDate": "2022-12-29T11:56:31.840Z",
              "content": "<p>Yes I'd noticed this when using the Decoder in a pytorch dataloader.</p>\n<p>If I run with <code>n_workers=0</code> then it works perfectly. However, when &gt;0, ie using multiprocessing, the results are wrong.</p>\n<p>I then tried to instantiate the decoder in <code>__getitem__()</code>, to ensure it is instantiated in the subprocess, rather than the main process which is then forked. However, this doesn't solve the issue for me - I still get the wrong image.</p>\n<p>My suspicion is that this is an issue to do with using cuda on the main process before the fork (see <a href=\"https://github.com/NVIDIA/DALI/issues/3826\" target=\"_blank\">here</a>).</p>\n<p>If that's the case, then an 'inline' approach, of processing DICOMs on the fly in a dataloader during inference with any background processes, rather than all in advance, looks difficult.</p>\n<p>A way around this, if so, it to use the \"spawn\" rather than \"fork\" approach to multiprocessing, but this is painful, in my experience, in notebooks.</p>",
              "rawMarkdown": "Yes I'd noticed this when using the Decoder in a pytorch dataloader.\n\nIf I run with `n_workers=0` then it works perfectly. However, when >0, ie using multiprocessing, the results are wrong.\n\nI then tried to instantiate the decoder in `__getitem__()`, to ensure it is instantiated in the subprocess, rather than the main process which is then forked. However, this doesn't solve the issue for me - I still get the wrong image.\n\nMy suspicion is that this is an issue to do with using cuda on the main process before the fork (see [here](https://github.com/NVIDIA/DALI/issues/3826)).\n\nIf that's the case, then an 'inline' approach, of processing DICOMs on the fly in a dataloader during inference with any background processes, rather than all in advance, looks difficult.\n\nA way around this, if so, it to use the \"spawn\" rather than \"fork\" approach to multiprocessing, but this is painful, in my experience, in notebooks."
            }
          ]
        },
        {
          "id": 2080226,
          "postDate": "2022-12-30T01:18:12.317Z",
          "content": "<p>There's also an easy way to do this with DALI which doesn't create the need for a DALI pipeline or a separate Decoder instance.  DALI has an eager mode (works in multi-process mode) which can decode and return an image in just one line (just use a nightly build of DALI). </p>\n<pre><code> nvidia.dali.experimental  eager\n nvidia.dali.types  types\n nvidia.dali.types  DALIDataType\n\nimage = eager.experimental.decoders.image([bitstream], device=, output_type=types.ANY_DATA, dtype=DALIDataType.UINT16)\n\nimage = image.as_cpu().at().squeeze(-)  \n</code></pre>",
          "rawMarkdown": "There's also an easy way to do this with DALI which doesn't create the need for a DALI pipeline or a separate Decoder instance.  DALI has an eager mode (works in multi-process mode) which can decode and return an image in just one line (just use a nightly build of DALI). \n\n\n```python\nfrom nvidia.dali.experimental import eager\nimport nvidia.dali.types as types\nfrom nvidia.dali.types import DALIDataType\n\nimage = eager.experimental.decoders.image([bitstream], device='gpu', output_type=types.ANY_DATA, dtype=DALIDataType.UINT16)\n\nimage = image.as_cpu().at(0).squeeze(-1)  #move back to cpu\n```\n  ",
          "votes": 3,
          "replies": [
            {
              "id": 2080525,
              "postDate": "2022-12-30T08:48:36.090Z",
              "content": "<p>Got error with the code <a href=\"https://www.kaggle.com/tivfrvqhs5\" target=\"_blank\">@tivfrvqhs5</a> provided</p>\n<pre><code>        dcmfile = pydicom.dcmread(f)\n         dcmfile.file_meta.TransferSyntaxUID == :\n             (f, )  fp:\n                raw = DicomBytesIO(fp.read())\n                ds = pydicom.dcmread(raw)\n            offset = ds.PixelData.find()  \n            hackedbitstream = ()\n            hackedbitstream.extend(ds.PixelData[offset:])\n            image = eager.experimental.decoders.image([hackedbitstream], device=, output_type=types.ANY_DATA, dtype=DALIDataType.UINT16)\n            image = image.as_cpu().at().squeeze(-)  \n</code></pre>\n<p>Using !pip install --extra-index-url <a href=\"https://developer.download.nvidia.com/compute/redist/nightly\" target=\"_blank\">https://developer.download.nvidia.com/compute/redist/nightly</a> --upgrade nvidia-dali-nightly-cuda110</p>\n<pre><code>Traceback (most recent call last):\n  File , line ,  _process_worker\n    r = call_item()\n  File , line ,  __call__\n     self.fn(*self.args, **self.kwargs)\n  File , line ,  __call__\n     self.func(*args, **kwargs)\n  File , line ,  __call__\n     func, args, kwargs  self.items]\n  File , line ,  &lt;listcomp&gt;\n     func, args, kwargs  self.items]\n  File , line ,  process_nv2000\n  File , line ,  wrapper\n    inputs, kwargs, op_name, wrapper_name, _callable_op_factory.disqualified_arguments)\n  File , line ,  _prep_args\n    inputs = _prep_inputs(inputs, batch_size)\n  File , line ,  _prep_inputs\n    inputs[i] = _transform_data_to_tensorlist(, batch_size)\n  File , line ,  _transform_data_to_tensorlist\n    data = _prep_data_for_feed_input(data, batch_size, layout, device_id)\n  File , line ,  _prep_data_for_feed_input\n    _check_data_batch(data, batch_size, layout)\n  File , line ,  _check_data_batch\n    shape, uniform = _get_batch_shape(data)\n  File , line ,  _get_batch_shape\n     (data[].shape):\nAttributeError:   has no attribute \n</code></pre>\n<p>If this works with joblib and 2 jobs, that'd be cleaner but  it won't speed things up afaict, best I can get is about 4400s processing 16K jpeg2000 + 16K lossless. </p>\n<p>I think you might be able to get around the lack of multiprocessor support in <a href=\"https://www.kaggle.com/snaker\" target=\"_blank\">@snaker</a> solution by using the approach I implement in the notebook - <a href=\"https://www.kaggle.com/code/kaggleqrdl/baseline-32000-nvjpeg2000-dicomsdl\" target=\"_blank\">https://www.kaggle.com/code/kaggleqrdl/baseline-32000-nvjpeg2000-dicomsdl</a> .. basically lock one joblib process to GPU and the other to dicomsdl, and once GPU is finished flip over to both processes using dicomsdl.  The trick is utilizing both processors all the time.  snaker solution is faster than dicomsdl, but only 2x.  It's also cleaner than dali .. which is rather messy and using GPU for everything isn't as flexible.</p>\n<p>I haven't tested it to see if the GPU output makes sense, though, but will try to get around to that given that folks are seeing issues.</p>\n<p>If this works, than the right approach is probably decode about K (where K is large but not enough to take up all your disk, eg &gt; batching in 100), and then infer on that - assuming your augmentation is on CPU and can be split across joblib.</p>\n<p>The open question imho - is the best approach do as much as possible on GPU with dali, including augmentation?  GPU agumentation doesn't have as much support as CPU augmentation and is harder to do, but this comp may come down to the fact that we're having to infer on 32K images and the team with the fastest processing (plus best models) may be the one that wins.</p>\n<p>It's notable that the results on the leaderboard are far off what folks are getting in the literature.  This is not promising for eventual outcome, but we're still early in the comp and one is hopeful.  My guess is the reason is partly because we're trying to infer on so many images and we're constrained by the 9 hour submit window.</p>",
              "rawMarkdown": "Got error with the code @tivfrvqhs5 provided\n```python\n        dcmfile = pydicom.dcmread(f)\n        if dcmfile.file_meta.TransferSyntaxUID == '1.2.840.10008.1.2.4.90':\n            with open(f, 'rb') as fp:\n                raw = DicomBytesIO(fp.read())\n                ds = pydicom.dcmread(raw)\n            offset = ds.PixelData.find(b\"\\x00\\x00\\x00\\x0C\")  #<---- the jpeg2000 header info we're looking for\n            hackedbitstream = bytearray()\n            hackedbitstream.extend(ds.PixelData[offset:])\n            image = eager.experimental.decoders.image([hackedbitstream], device='gpu', output_type=types.ANY_DATA, dtype=DALIDataType.UINT16)\n            image = image.as_cpu().at(0).squeeze(-1)  #move back to cpu    \n```\nUsing !pip install --extra-index-url https://developer.download.nvidia.com/compute/redist/nightly --upgrade nvidia-dali-nightly-cuda110\n\n```python\nTraceback (most recent call last):\n  File \"/opt/conda/lib/python3.7/site-packages/joblib/externals/loky/process_executor.py\", line 431, in _process_worker\n    r = call_item()\n  File \"/opt/conda/lib/python3.7/site-packages/joblib/externals/loky/process_executor.py\", line 285, in __call__\n    return self.fn(*self.args, **self.kwargs)\n  File \"/opt/conda/lib/python3.7/site-packages/joblib/_parallel_backends.py\", line 595, in __call__\n    return self.func(*args, **kwargs)\n  File \"/opt/conda/lib/python3.7/site-packages/joblib/parallel.py\", line 263, in __call__\n    for func, args, kwargs in self.items]\n  File \"/opt/conda/lib/python3.7/site-packages/joblib/parallel.py\", line 263, in <listcomp>\n    for func, args, kwargs in self.items]\n  File \"/tmp/ipykernel_24/3740975955.py\", line 44, in process_nv2000\n  File \"/opt/conda/lib/python3.7/site-packages/nvidia/dali/_utils/eager_utils.py\", line 684, in wrapper\n    inputs, kwargs, op_name, wrapper_name, _callable_op_factory.disqualified_arguments)\n  File \"/opt/conda/lib/python3.7/site-packages/nvidia/dali/_utils/eager_utils.py\", line 655, in _prep_args\n    inputs = _prep_inputs(inputs, batch_size)\n  File \"/opt/conda/lib/python3.7/site-packages/nvidia/dali/_utils/eager_utils.py\", line 636, in _prep_inputs\n    inputs[i] = _transform_data_to_tensorlist(input, batch_size)\n  File \"/opt/conda/lib/python3.7/site-packages/nvidia/dali/_utils/eager_utils.py\", line 80, in _transform_data_to_tensorlist\n    data = _prep_data_for_feed_input(data, batch_size, layout, device_id)\n  File \"/opt/conda/lib/python3.7/site-packages/nvidia/dali/external_source.py\", line 83, in _prep_data_for_feed_input\n    _check_data_batch(data, batch_size, layout)\n  File \"/opt/conda/lib/python3.7/site-packages/nvidia/dali/external_source.py\", line 43, in _check_data_batch\n    shape, uniform = _get_batch_shape(data)\n  File \"/opt/conda/lib/python3.7/site-packages/nvidia/dali/external_source.py\", line 31, in _get_batch_shape\n    if callable(data[0].shape):\nAttributeError: 'bytearray' object has no attribute 'shape'\n```\n\nIf this works with joblib and 2 jobs, that'd be cleaner but  it won't speed things up afaict, best I can get is about 4400s processing 16K jpeg2000 + 16K lossless. \n\n I think you might be able to get around the lack of multiprocessor support in @snaker solution by using the approach I implement in the notebook - https://www.kaggle.com/code/kaggleqrdl/baseline-32000-nvjpeg2000-dicomsdl .. basically lock one joblib process to GPU and the other to dicomsdl, and once GPU is finished flip over to both processes using dicomsdl.  The trick is utilizing both processors all the time.  snaker solution is faster than dicomsdl, but only 2x.  It's also cleaner than dali .. which is rather messy and using GPU for everything isn't as flexible.\n\nI haven't tested it to see if the GPU output makes sense, though, but will try to get around to that given that folks are seeing issues.\n\nIf this works, than the right approach is probably decode about K (where K is large but not enough to take up all your disk, eg > batching in 100), and then infer on that - assuming your augmentation is on CPU and can be split across joblib.\n\nThe open question imho - is the best approach do as much as possible on GPU with dali, including augmentation?  GPU agumentation doesn't have as much support as CPU augmentation and is harder to do, but this comp may come down to the fact that we're having to infer on 32K images and the team with the fastest processing (plus best models) may be the one that wins.\n\nIt's notable that the results on the leaderboard are far off what folks are getting in the literature.  This is not promising for eventual outcome, but we're still early in the comp and one is hopeful.  My guess is the reason is partly because we're trying to infer on so many images and we're constrained by the 9 hour submit window.",
              "votes": 1
            },
            {
              "id": 2080752,
              "postDate": "2022-12-30T12:59:21.677Z",
              "content": "<p>With DALI eager mode you need to wrap the byte string in a np array.</p>\n<p><code>bitstream = np.array(bytearray(ds.PixelData[offset:]), dtype=np.uint8)</code></p>",
              "rawMarkdown": "With DALI eager mode you need to wrap the byte string in a np array.\n\n`bitstream = np.array(bytearray(ds.PixelData[offset:]), dtype=np.uint8)`",
              "votes": 2
            },
            {
              "id": 2080769,
              "postDate": "2022-12-30T13:22:15.043Z",
              "content": "<p>Seems to work, but multiprocessing doesn't help much.  dali might be locking on gpu.  Still, it's a bit cleaner in that you don't need the custom built decoder library.</p>\n<p>edit, hmm, let me try the t4x2</p>\n<p>seems faster with t4x2 and multiprocessing, faster than single cpu + nvjpeg2000 lib.   </p>\n<p>Hard to tell exactly though, as I think caching is getting in the way of results and repeat testing.  Clearing it doesn't seem to help, might have to do something more aggressive.  Right now I'm just restarting the VMs.  </p>\n<p>Interestingly t4x2 doesn't show GPU being used, might be a glitch with those vms.   Also, not seeing 2x speedups, just marginal.</p>\n<p>Need to test the images to see if they're actually being produced though.   If it works, this could shave off around 10-20 minutes or so off decoding time.</p>\n<p>edit - I looked at a bunch of images, they all seem to process fine.  Not sure how this if failing for folks though, might be a non deterministic / race condition.  Would have to test every image.</p>",
              "rawMarkdown": "Seems to work, but multiprocessing doesn't help much.  dali might be locking on gpu.  Still, it's a bit cleaner in that you don't need the custom built decoder library.\n\nedit, hmm, let me try the t4x2\n\nseems faster with t4x2 and multiprocessing, faster than single cpu + nvjpeg2000 lib.   \n\nHard to tell exactly though, as I think caching is getting in the way of results and repeat testing.  Clearing it doesn't seem to help, might have to do something more aggressive.  Right now I'm just restarting the VMs.  \n\nInterestingly t4x2 doesn't show GPU being used, might be a glitch with those vms.   Also, not seeing 2x speedups, just marginal.\n\nNeed to test the images to see if they're actually being produced though.   If it works, this could shave off around 10-20 minutes or so off decoding time.\n\nedit - I looked at a bunch of images, they all seem to process fine.  Not sure how this if failing for folks though, might be a non deterministic / race condition.  Would have to test every image.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2066687,
      "postDate": "2022-12-16T00:25:25.487Z",
      "content": "<p>I've made the dataset used in the Notebook public which I forgot in the first place. The code can be run directely without internet.</p>",
      "rawMarkdown": "I've made the dataset used in the Notebook public which I forgot in the first place. The code can be run directely without internet."
    },
    {
      "id": 2066323,
      "postDate": "2022-12-15T15:25:09.913Z",
      "content": "<p>Brilliant work, very impressive.</p>",
      "rawMarkdown": "Brilliant work, very impressive."
    },
    {
      "id": 2066108,
      "postDate": "2022-12-15T12:12:28.587Z",
      "content": "<p>not sure if this is an error or my environment is different.<br>\ni have to hack the code a bit to make it work at my local system.</p>\n<p>nvjpeg2k.cpp</p>\n<pre><code>line 27\n/*\n  pybind11::array_t&lt;uint16_t&gt; decode(std::string path) { \n    std::ostringstream sstream;\n    std::ifstream fs(path.c_str());\n    sstream &lt;&lt; fs.rdbuf();\n    const std::string str(sstream.str());\n    auto length = str.length();\n    auto bitstream_buffer = (const unsigned char *)str.c_str();\n*/\n\n\n  pybind11::array_t&lt;uint16_t&gt; decode(pybind11::array_t&lt;uint8_t&gt; ss) { \n    pybind11::buffer_info info = ss.request();\n    auto length = info.size;\n    auto bitstream_buffer = (const unsigned char *)(info.ptr);\n\n....\n\n\nline 91\n    //result.reshape({image_info.image_height, image_info.image_width});\n    result.resize({image_info.image_height, image_info.image_width});\n    return result;\n</code></pre>\n<p>in the python code, i use:</p>\n<pre><code>    dcm_file =  ' xxx.dcm'\n    ds = pydicom.dcmread(dcm_file)\n\n    offset = ds.PixelData.find(b'\\x00\\x00\\x00\\x0C')  \n    jpeg_stream =  np.array(bytearray(ds.PixelData[offset:]), np.uint8)\n    image = decoder.decode(jpeg_stream)\n\n    print(image.shape)\n    plt.imshow(image)\n    plt.show()\n</code></pre>",
      "rawMarkdown": "not sure if this is an error or my environment is different.\ni have to hack the code a bit to make it work at my local system.\n\n\nnvjpeg2k.cpp\n```\n\nline 27\n/*\n  pybind11::array_t<uint16_t> decode(std::string path) { \n    std::ostringstream sstream;\n    std::ifstream fs(path.c_str());\n    sstream << fs.rdbuf();\n    const std::string str(sstream.str());\n    auto length = str.length();\n    auto bitstream_buffer = (const unsigned char *)str.c_str();\n*/\n\n\n  pybind11::array_t<uint16_t> decode(pybind11::array_t<uint8_t> ss) { \n    pybind11::buffer_info info = ss.request();\n    auto length = info.size;\n    auto bitstream_buffer = (const unsigned char *)(info.ptr);\n\n....\n\n\nline 91\n    //result.reshape({image_info.image_height, image_info.image_width});\n    result.resize({image_info.image_height, image_info.image_width});\n    return result;\n\n\n```\n\nin the python code, i use:\n\n```\n    dcm_file =  ' xxx.dcm'\n    ds = pydicom.dcmread(dcm_file)\n\n    offset = ds.PixelData.find(b'\\x00\\x00\\x00\\x0C')  \n    jpeg_stream =  np.array(bytearray(ds.PixelData[offset:]), np.uint8)\n    image = decoder.decode(jpeg_stream)\n\n    print(image.shape)\n    plt.imshow(image)\n    plt.show()\n\n```\n"
    },
    {
      "id": 2065986,
      "postDate": "2022-12-15T09:03:45.210Z",
      "content": "<p>I have update the README in the repo to show how to compile the extension in the local machine.</p>",
      "rawMarkdown": "I have update the README in the repo to show how to compile the extension in the local machine."
    },
    {
      "id": 2065934,
      "postDate": "2022-12-15T08:12:30.903Z",
      "content": "<p>Thanks!</p>\n<p>any instructions to make for loca machine for the GitHub repo?</p>",
      "rawMarkdown": "Thanks!\n\nany instructions to make for loca machine for the GitHub repo?",
      "replies": [
        {
          "id": 2065952,
          "postDate": "2022-12-15T08:29:34.350Z",
          "content": "<p>I'm working on the <code>readme.md</code>. Will let you know when it is finished.</p>",
          "rawMarkdown": "I'm working on the `readme.md`. Will let you know when it is finished.",
          "replies": [
            {
              "id": 2066689,
              "postDate": "2022-12-16T00:26:43.030Z",
              "content": "<p><a href=\"https://www.kaggle.com/vipscu\" target=\"_blank\">@vipscu</a> Thanks again.</p>\n<p>you can avoid reading the dicom twice( pydicom.dcmread(path) and pydicom.dcmread(raw)), using my fix above</p>\n<pre><code>    ds = pydicom.dcmread(dcm_file)\n    offset = ds.PixelData.find(b'\\x00\\x00\\x00\\x0C')  \n    #jpeg_stream =  np.array(bytearray(ds.PixelData[offset:]), np.uint8)\n    jpeg_stream =  bytearray(ds.PixelData[offset:])\n    image = decoder.decode(jpeg_stream)\n</code></pre>",
              "rawMarkdown": "@vipscu Thanks again.\n\nyou can avoid reading the dicom twice( pydicom.dcmread(path) and pydicom.dcmread(raw)), using my fix above\n\n```\n\n    ds = pydicom.dcmread(dcm_file)\n    offset = ds.PixelData.find(b'\\x00\\x00\\x00\\x0C')  \n    #jpeg_stream =  np.array(bytearray(ds.PixelData[offset:]), np.uint8)\n    jpeg_stream =  bytearray(ds.PixelData[offset:])\n    image = decoder.decode(jpeg_stream)\n\n```"
            }
          ]
        }
      ]
    },
    {
      "id": 2074133,
      "postDate": "2022-12-23T18:33:50.197Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2065982,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-12-15T08:59:31.603000",
      "content": "<p>** for local compilation only **</p>\n<p>before compile the python extension, you have to verify that nvJPEG2000  is installed correctly.</p>\n<ol>\n<li>download and install nvJPEG2000 from <a href=\"https://developer.nvidia.com/nvjpeg\" target=\"_blank\">https://developer.nvidia.com/nvjpeg</a></li>\n<li>verify correct installation by compile the example at <a href=\"https://github.com/NVIDIA/CUDALibrarySamples/tree/master/nvJPEG2000/nvJPEG2000-Decoder\" target=\"_blank\">https://github.com/NVIDIA/CUDALibrarySamples/tree/master/nvJPEG2000/nvJPEG2000-Decoder</a></li>\n</ol>\n<p>to locate the where you have installed nvJPEG2000, i use:</p>\n<pre><code>locate nvjpeg2k\n/usr/include/nvjpeg2k.h\n/usr/include/nvjpeg2k_version.h\n</code></pre>\n<p>hence my DNVJPEG2K_PATH=/usr</p>\n<p>then,</p>\n<pre><code>mkdir build\ncd build \nexport CUDACXX=nvcc\ncmake ..  -DNVJPEG2K_PATH=/usr\nmake\n</code></pre>\n<p>Finally run the example</p>\n<pre><code>./nvjpeg2000_decode_sample -i ../images/2k_image_lossless/2k_lossless.jp2 -o .\n</code></pre>\n<p>expected output</p>\n<pre><code>Decoding images in directory: ../images/2k_image_lossless/2k_lossless.jp2, total 1, batchsize 1\nTotal decoding time: 0.0242717\nAvg decoding time per image: 0.0242717\nAvg images per sec: 41.2002\nAvg decoding time per batch: 0.0242717\n</code></pre>\n<hr>\n<p>I use pycharm + andaconda. <br>\nTo link to the correct pybind11, </p>\n<pre><code>0. I first check the pybind11 version that is already installed in my system:\n/home/titanx/hengck/opt/anaconda3.9/lib/python3.9/site-packages/pybind11\n1.  git clone the same version from https://github.com/pybind/pybind11 to the direct extern/pybind11\n2 . use the terminal from pycharm IDE interface\n3. run the commands:\n\nexport NVJPEG2K_PATH=/usr\nexport CUDACXX=nvcc\ncmake .. \\\n  -DCMAKE_BUILD_TYPE=Debug \\\n  -DNVJPEG2K_PATH=/usr \\\n  -DNVJPEG2K_LIB=/usr/lib/x86_64-linux-gnu/libnvjpeg2k_static.a\n</code></pre>\n<p>you should see</p>\n<pre><code>-- pybind11 v2.11.0 dev1\n-- Found PythonInterp: /home/titanx/hengck/opt/anaconda3.9/bin/python (found suitable version \"3.9.12\", minimum required is \"3.6\") \n-- Found PythonLibs: /home/titanx/hengck/opt/anaconda3.9/lib/libpython3.9.so\n-- Performing Test HAS_FLTO\n-- Performing Test HAS_FLTO - Success\n\n-- Configuring done\n-- Generating done\n-- Build files have been written to: /home/titanx/hengck/share1/kaggle/2022/rsna-breast-mammography/tool/nvjpeg2k/nvjpeg2k-python-main/build\n</code></pre>\n<p>you can download and decompress:<br>\n<a href=\"https://ubuntu.pkgs.org/22.04/cuda-amd64/libnvjpeg2k0_0.6.0.28-1_amd64.deb.html\" target=\"_blank\">https://ubuntu.pkgs.org/22.04/cuda-amd64/libnvjpeg2k0_0.6.0.28-1_amd64.deb.html</a><br>\n<a href=\"https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2004/x86_64/libnvjpeg2k0_0.6.0.28-1_amd64.deb\" target=\"_blank\">https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2004/x86_64/libnvjpeg2k0_0.6.0.28-1_amd64.deb</a></p>\n<p>or i use this:</p>\n<pre><code>  sudo apt-key adv --fetch-keys http://developer.download.nvidia.com/compute/cuda/repos/ubuntu1804/x86_64/3bf863cc.pub\n  sudo add-apt-repository \"deb http://developer.download.nvidia.com/compute/cuda/repos/ubuntu1804/x86_64/ /\"\n  sudo apt update\n\n  sudo apt-get install libnvjpeg2k0 libnvjpeg2k-dev\n</code></pre>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2079508,
      "author_name": "Chenglu",
      "author_url": "",
      "post_date": "2022-12-29T11:25:00.750000",
      "content": "<p>If one wants to use the extension in the multi-process. Here's the tip:</p>\n<p><strong>Do not</strong> create the <code>Decoder</code> instance in main process, instead, create the instance in sub-processes. There are unexpected behaviors when the <code>Decoder</code> is initialized in a process that will fork sub-processes. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 2079516,
          "author_name": "Rasoul Mojtahedzadeh",
          "author_url": "",
          "post_date": "2022-12-29T11:36:14.623000",
          "content": "<p>Creating the decoder in sub-processes could slow down the process of decoding dcm?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2079522,
              "author_name": "Chenglu",
              "author_url": "",
              "post_date": "2022-12-29T11:43:24.637000",
              "content": "<p>There is no performance difference between creating <code>Decoder</code> in main process or sub process. The reason why we should create the <code>Decoder</code> in subprocesses is to prevent the \"unexpected behavior\".</p>\n<p>The \"unexpected behavior\" is that if create a <code>Decoder</code> in main process and fork sub processes, the decoder will not produce right results, you will see many 0s or 65536s in the decoded image. I don't known the root cause of this cause I'm not a CUDA expert.</p>\n<p>The whole thing only applies to those who wants to use the extension concurrently, cause Kaggle's kernel has 2 CPU cores. If there are no plans to use multi processing, then there are no such issue of the extension.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2079534,
              "author_name": "James Howard",
              "author_url": "",
              "post_date": "2022-12-29T11:56:31.840000",
              "content": "<p>Yes I'd noticed this when using the Decoder in a pytorch dataloader.</p>\n<p>If I run with <code>n_workers=0</code> then it works perfectly. However, when &gt;0, ie using multiprocessing, the results are wrong.</p>\n<p>I then tried to instantiate the decoder in <code>__getitem__()</code>, to ensure it is instantiated in the subprocess, rather than the main process which is then forked. However, this doesn't solve the issue for me - I still get the wrong image.</p>\n<p>My suspicion is that this is an issue to do with using cuda on the main process before the fork (see <a href=\"https://github.com/NVIDIA/DALI/issues/3826\" target=\"_blank\">here</a>).</p>\n<p>If that's the case, then an 'inline' approach, of processing DICOMs on the fly in a dataloader during inference with any background processes, rather than all in advance, looks difficult.</p>\n<p>A way around this, if so, it to use the \"spawn\" rather than \"fork\" approach to multiprocessing, but this is painful, in my experience, in notebooks.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2080226,
          "author_name": "David Austin",
          "author_url": "",
          "post_date": "2022-12-30T01:18:12.317000",
          "content": "<p>There's also an easy way to do this with DALI which doesn't create the need for a DALI pipeline or a separate Decoder instance.  DALI has an eager mode (works in multi-process mode) which can decode and return an image in just one line (just use a nightly build of DALI). </p>\n<pre><code> nvidia.dali.experimental  eager\n nvidia.dali.types  types\n nvidia.dali.types  DALIDataType\n\nimage = eager.experimental.decoders.image([bitstream], device=, output_type=types.ANY_DATA, dtype=DALIDataType.UINT16)\n\nimage = image.as_cpu().at().squeeze(-)  \n</code></pre>",
          "votes": 3,
          "replies": [
            {
              "id": 2080525,
              "author_name": "@kaggleqrdl",
              "author_url": "",
              "post_date": "2022-12-30T08:48:36.090000",
              "content": "<p>Got error with the code <a href=\"https://www.kaggle.com/tivfrvqhs5\" target=\"_blank\">@tivfrvqhs5</a> provided</p>\n<pre><code>        dcmfile = pydicom.dcmread(f)\n         dcmfile.file_meta.TransferSyntaxUID == :\n             (f, )  fp:\n                raw = DicomBytesIO(fp.read())\n                ds = pydicom.dcmread(raw)\n            offset = ds.PixelData.find()  \n            hackedbitstream = ()\n            hackedbitstream.extend(ds.PixelData[offset:])\n            image = eager.experimental.decoders.image([hackedbitstream], device=, output_type=types.ANY_DATA, dtype=DALIDataType.UINT16)\n            image = image.as_cpu().at().squeeze(-)  \n</code></pre>\n<p>Using !pip install --extra-index-url <a href=\"https://developer.download.nvidia.com/compute/redist/nightly\" target=\"_blank\">https://developer.download.nvidia.com/compute/redist/nightly</a> --upgrade nvidia-dali-nightly-cuda110</p>\n<pre><code>Traceback (most recent call last):\n  File , line ,  _process_worker\n    r = call_item()\n  File , line ,  __call__\n     self.fn(*self.args, **self.kwargs)\n  File , line ,  __call__\n     self.func(*args, **kwargs)\n  File , line ,  __call__\n     func, args, kwargs  self.items]\n  File , line ,  &lt;listcomp&gt;\n     func, args, kwargs  self.items]\n  File , line ,  process_nv2000\n  File , line ,  wrapper\n    inputs, kwargs, op_name, wrapper_name, _callable_op_factory.disqualified_arguments)\n  File , line ,  _prep_args\n    inputs = _prep_inputs(inputs, batch_size)\n  File , line ,  _prep_inputs\n    inputs[i] = _transform_data_to_tensorlist(, batch_size)\n  File , line ,  _transform_data_to_tensorlist\n    data = _prep_data_for_feed_input(data, batch_size, layout, device_id)\n  File , line ,  _prep_data_for_feed_input\n    _check_data_batch(data, batch_size, layout)\n  File , line ,  _check_data_batch\n    shape, uniform = _get_batch_shape(data)\n  File , line ,  _get_batch_shape\n     (data[].shape):\nAttributeError:   has no attribute \n</code></pre>\n<p>If this works with joblib and 2 jobs, that'd be cleaner but  it won't speed things up afaict, best I can get is about 4400s processing 16K jpeg2000 + 16K lossless. </p>\n<p>I think you might be able to get around the lack of multiprocessor support in <a href=\"https://www.kaggle.com/snaker\" target=\"_blank\">@snaker</a> solution by using the approach I implement in the notebook - <a href=\"https://www.kaggle.com/code/kaggleqrdl/baseline-32000-nvjpeg2000-dicomsdl\" target=\"_blank\">https://www.kaggle.com/code/kaggleqrdl/baseline-32000-nvjpeg2000-dicomsdl</a> .. basically lock one joblib process to GPU and the other to dicomsdl, and once GPU is finished flip over to both processes using dicomsdl.  The trick is utilizing both processors all the time.  snaker solution is faster than dicomsdl, but only 2x.  It's also cleaner than dali .. which is rather messy and using GPU for everything isn't as flexible.</p>\n<p>I haven't tested it to see if the GPU output makes sense, though, but will try to get around to that given that folks are seeing issues.</p>\n<p>If this works, than the right approach is probably decode about K (where K is large but not enough to take up all your disk, eg &gt; batching in 100), and then infer on that - assuming your augmentation is on CPU and can be split across joblib.</p>\n<p>The open question imho - is the best approach do as much as possible on GPU with dali, including augmentation?  GPU agumentation doesn't have as much support as CPU augmentation and is harder to do, but this comp may come down to the fact that we're having to infer on 32K images and the team with the fastest processing (plus best models) may be the one that wins.</p>\n<p>It's notable that the results on the leaderboard are far off what folks are getting in the literature.  This is not promising for eventual outcome, but we're still early in the comp and one is hopeful.  My guess is the reason is partly because we're trying to infer on so many images and we're constrained by the 9 hour submit window.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2080752,
              "author_name": "David Austin",
              "author_url": "",
              "post_date": "2022-12-30T12:59:21.677000",
              "content": "<p>With DALI eager mode you need to wrap the byte string in a np array.</p>\n<p><code>bitstream = np.array(bytearray(ds.PixelData[offset:]), dtype=np.uint8)</code></p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2080769,
              "author_name": "@kaggleqrdl",
              "author_url": "",
              "post_date": "2022-12-30T13:22:15.043000",
              "content": "<p>Seems to work, but multiprocessing doesn't help much.  dali might be locking on gpu.  Still, it's a bit cleaner in that you don't need the custom built decoder library.</p>\n<p>edit, hmm, let me try the t4x2</p>\n<p>seems faster with t4x2 and multiprocessing, faster than single cpu + nvjpeg2000 lib.   </p>\n<p>Hard to tell exactly though, as I think caching is getting in the way of results and repeat testing.  Clearing it doesn't seem to help, might have to do something more aggressive.  Right now I'm just restarting the VMs.  </p>\n<p>Interestingly t4x2 doesn't show GPU being used, might be a glitch with those vms.   Also, not seeing 2x speedups, just marginal.</p>\n<p>Need to test the images to see if they're actually being produced though.   If it works, this could shave off around 10-20 minutes or so off decoding time.</p>\n<p>edit - I looked at a bunch of images, they all seem to process fine.  Not sure how this if failing for folks though, might be a non deterministic / race condition.  Would have to test every image.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2066687,
      "author_name": "Chenglu",
      "author_url": "",
      "post_date": "2022-12-16T00:25:25.487000",
      "content": "<p>I've made the dataset used in the Notebook public which I forgot in the first place. The code can be run directely without internet.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2066323,
      "author_name": "James Howard",
      "author_url": "",
      "post_date": "2022-12-15T15:25:09.913000",
      "content": "<p>Brilliant work, very impressive.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2066108,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-12-15T12:12:28.587000",
      "content": "<p>not sure if this is an error or my environment is different.<br>\ni have to hack the code a bit to make it work at my local system.</p>\n<p>nvjpeg2k.cpp</p>\n<pre><code>line 27\n/*\n  pybind11::array_t&lt;uint16_t&gt; decode(std::string path) { \n    std::ostringstream sstream;\n    std::ifstream fs(path.c_str());\n    sstream &lt;&lt; fs.rdbuf();\n    const std::string str(sstream.str());\n    auto length = str.length();\n    auto bitstream_buffer = (const unsigned char *)str.c_str();\n*/\n\n\n  pybind11::array_t&lt;uint16_t&gt; decode(pybind11::array_t&lt;uint8_t&gt; ss) { \n    pybind11::buffer_info info = ss.request();\n    auto length = info.size;\n    auto bitstream_buffer = (const unsigned char *)(info.ptr);\n\n....\n\n\nline 91\n    //result.reshape({image_info.image_height, image_info.image_width});\n    result.resize({image_info.image_height, image_info.image_width});\n    return result;\n</code></pre>\n<p>in the python code, i use:</p>\n<pre><code>    dcm_file =  ' xxx.dcm'\n    ds = pydicom.dcmread(dcm_file)\n\n    offset = ds.PixelData.find(b'\\x00\\x00\\x00\\x0C')  \n    jpeg_stream =  np.array(bytearray(ds.PixelData[offset:]), np.uint8)\n    image = decoder.decode(jpeg_stream)\n\n    print(image.shape)\n    plt.imshow(image)\n    plt.show()\n</code></pre>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2065986,
      "author_name": "Chenglu",
      "author_url": "",
      "post_date": "2022-12-15T09:03:45.210000",
      "content": "<p>I have update the README in the repo to show how to compile the extension in the local machine.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2065934,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-12-15T08:12:30.903000",
      "content": "<p>Thanks!</p>\n<p>any instructions to make for loca machine for the GitHub repo?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2065952,
          "author_name": "Chenglu",
          "author_url": "",
          "post_date": "2022-12-15T08:29:34.350000",
          "content": "<p>I'm working on the <code>readme.md</code>. Will let you know when it is finished.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2066689,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2022-12-16T00:26:43.030000",
              "content": "<p><a href=\"https://www.kaggle.com/vipscu\" target=\"_blank\">@vipscu</a> Thanks again.</p>\n<p>you can avoid reading the dicom twice( pydicom.dcmread(path) and pydicom.dcmread(raw)), using my fix above</p>\n<pre><code>    ds = pydicom.dcmread(dcm_file)\n    offset = ds.PixelData.find(b'\\x00\\x00\\x00\\x0C')  \n    #jpeg_stream =  np.array(bytearray(ds.PixelData[offset:]), np.uint8)\n    jpeg_stream =  bytearray(ds.PixelData[offset:])\n    image = decoder.decode(jpeg_stream)\n</code></pre>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2074133,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-12-23T18:33:50.197000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2065903": "I write a [notebook](https://www.kaggle.com/code/snaker/easy-load-the-image-with-nvjpeg2000) to demonstrate how to load the dicom file with nvJPEG2000.\n\nThe trick is implemented by Python extension, the source code of that extension is at: https://github.com/louis-she/nvjpeg2k-python . I built it with Python3.7 and CUDA11 so that can be used in Kaggle's default environment. For how to compile the extension in local machine, see the readme from the repo.\n\nThanks to the post [17x dicom decode speedup on GPU for jpeg2000 encodings](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/371534).",
    "2065982": "** for local compilation only **\n\nbefore compile the python extension, you have to verify that nvJPEG2000  is installed correctly.\n\n1. download and install nvJPEG2000 from https://developer.nvidia.com/nvjpeg\n2. verify correct installation by compile the example at https://github.com/NVIDIA/CUDALibrarySamples/tree/master/nvJPEG2000/nvJPEG2000-Decoder\n\n\nto locate the where you have installed nvJPEG2000, i use:\n```\nlocate nvjpeg2k\n/usr/include/nvjpeg2k.h\n/usr/include/nvjpeg2k_version.h\n```\n\nhence my DNVJPEG2K_PATH=/usr\n\nthen,\n\n```\nmkdir build\ncd build \nexport CUDACXX=nvcc\ncmake ..  -DNVJPEG2K_PATH=/usr\nmake\n```\nFinally run the example\n```\n./nvjpeg2000_decode_sample -i ../images/2k_image_lossless/2k_lossless.jp2 -o .\n```\n\nexpected output\n\n```\nDecoding images in directory: ../images/2k_image_lossless/2k_lossless.jp2, total 1, batchsize 1\nTotal decoding time: 0.0242717\nAvg decoding time per image: 0.0242717\nAvg images per sec: 41.2002\nAvg decoding time per batch: 0.0242717\n\n\n```\n\n----\n\nI use pycharm + andaconda. \nTo link to the correct pybind11, \n\n```\n\n0. I first check the pybind11 version that is already installed in my system:\n/home/titanx/hengck/opt/anaconda3.9/lib/python3.9/site-packages/pybind11\n1.  git clone the same version from https://github.com/pybind/pybind11 to the direct extern/pybind11\n2 . use the terminal from pycharm IDE interface\n3. run the commands:\n\nexport NVJPEG2K_PATH=/usr\nexport CUDACXX=nvcc\ncmake .. \\\n  -DCMAKE_BUILD_TYPE=Debug \\\n  -DNVJPEG2K_PATH=/usr \\\n  -DNVJPEG2K_LIB=/usr/lib/x86_64-linux-gnu/libnvjpeg2k_static.a\n\n```\nyou should see\n\n```\n-- pybind11 v2.11.0 dev1\n-- Found PythonInterp: /home/titanx/hengck/opt/anaconda3.9/bin/python (found suitable version \"3.9.12\", minimum required is \"3.6\") \n-- Found PythonLibs: /home/titanx/hengck/opt/anaconda3.9/lib/libpython3.9.so\n-- Performing Test HAS_FLTO\n-- Performing Test HAS_FLTO - Success\n\n-- Configuring done\n-- Generating done\n-- Build files have been written to: /home/titanx/hengck/share1/kaggle/2022/rsna-breast-mammography/tool/nvjpeg2k/nvjpeg2k-python-main/build\n\n\n```\n\nyou can download and decompress:\nhttps://ubuntu.pkgs.org/22.04/cuda-amd64/libnvjpeg2k0_0.6.0.28-1_amd64.deb.html\nhttps://developer.download.nvidia.com/compute/cuda/repos/ubuntu2004/x86_64/libnvjpeg2k0_0.6.0.28-1_amd64.deb\n\nor i use this:\n\n```\n  sudo apt-key adv --fetch-keys http://developer.download.nvidia.com/compute/cuda/repos/ubuntu1804/x86_64/3bf863cc.pub\n  sudo add-apt-repository \"deb http://developer.download.nvidia.com/compute/cuda/repos/ubuntu1804/x86_64/ /\"\n  sudo apt update\n\n  sudo apt-get install libnvjpeg2k0 libnvjpeg2k-dev\n\n```\n",
    "2079508": "If one wants to use the extension in the multi-process. Here's the tip:\n\n**Do not** create the `Decoder` instance in main process, instead, create the instance in sub-processes. There are unexpected behaviors when the `Decoder` is initialized in a process that will fork sub-processes. ",
    "2066687": "I've made the dataset used in the Notebook public which I forgot in the first place. The code can be run directely without internet.",
    "2066323": "Brilliant work, very impressive.",
    "2066108": "not sure if this is an error or my environment is different.\ni have to hack the code a bit to make it work at my local system.\n\n\nnvjpeg2k.cpp\n```\n\nline 27\n/*\n  pybind11::array_t<uint16_t> decode(std::string path) { \n    std::ostringstream sstream;\n    std::ifstream fs(path.c_str());\n    sstream << fs.rdbuf();\n    const std::string str(sstream.str());\n    auto length = str.length();\n    auto bitstream_buffer = (const unsigned char *)str.c_str();\n*/\n\n\n  pybind11::array_t<uint16_t> decode(pybind11::array_t<uint8_t> ss) { \n    pybind11::buffer_info info = ss.request();\n    auto length = info.size;\n    auto bitstream_buffer = (const unsigned char *)(info.ptr);\n\n....\n\n\nline 91\n    //result.reshape({image_info.image_height, image_info.image_width});\n    result.resize({image_info.image_height, image_info.image_width});\n    return result;\n\n\n```\n\nin the python code, i use:\n\n```\n    dcm_file =  ' xxx.dcm'\n    ds = pydicom.dcmread(dcm_file)\n\n    offset = ds.PixelData.find(b'\\x00\\x00\\x00\\x0C')  \n    jpeg_stream =  np.array(bytearray(ds.PixelData[offset:]), np.uint8)\n    image = decoder.decode(jpeg_stream)\n\n    print(image.shape)\n    plt.imshow(image)\n    plt.show()\n\n```\n",
    "2065986": "I have update the README in the repo to show how to compile the extension in the local machine.",
    "2065934": "Thanks!\n\nany instructions to make for loca machine for the GitHub repo?",
    "2074133": ""
  }
}