{
  "id": 371981,
  "title": "Faster Dicom Processing on GPU",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/371981",
  "author_name": "Theo Viel",
  "post_date": "2022-12-13T14:23:01.067000",
  "votes": 37,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I've made an inference notebook public, that runs in about 4h !</p>\n<p>Trick used is shared by <a href=\"https://www.kaggle.com/tivfrvqhs5\" target=\"_blank\">@tivfrvqhs5</a> <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/371534\" target=\"_blank\">here</a>, basically you can process jpeg-compressed dicoms on GPU with NVIDIA Dali. </p>\n<p>The cool thing is, the output is already on GPU, so you can save even more time by processing the images on GPU as well : resizing and rescaling is super fast then.</p>\n<p>Implementation is a bit tricky, I had to modify a file in the Dali package for it to work with uint16, and process everything per chunk otherwise your run out of disk space.</p>\n<p><strong>Code is here :</strong></p>\n<blockquote>\n  <p><a href=\"https://www.kaggle.com/code/theoviel/rsna-breast-baseline-faster-inference-with-dali\" target=\"_blank\">https://www.kaggle.com/code/theoviel/rsna-breast-baseline-faster-inference-with-dali</a></p>\n</blockquote>\n<p>The runtime can be further improved by using batching, but I'm already happy with a 2x speed up :)</p>\n<p>EDIT: Added <a href=\"https://www.kaggle.com/code/hengck23/combine-dali-and-dicomsdl-for-reading-dicom-files\" target=\"_blank\">dicomsdl</a>, I think I gain an extra hour with it. Thanks <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> !</p>\n<p>(The goal of the kernel is not to share a 10th place model, which is why I deleted the dataset, but rather to give everyone a working pipeline to use GPUs efficiently &amp; save submission time. I did not expect the submission to score 0.46 to be honest. I may re-upload the weights later once people have reached higher LB scores.)</p>",
  "messages": [
    {
      "id": 2064105,
      "postDate": "2022-12-13T14:23:01.067Z",
      "content": "<p>I've made an inference notebook public, that runs in about 4h !</p>\n<p>Trick used is shared by <a href=\"https://www.kaggle.com/tivfrvqhs5\" target=\"_blank\">@tivfrvqhs5</a> <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/371534\" target=\"_blank\">here</a>, basically you can process jpeg-compressed dicoms on GPU with NVIDIA Dali. </p>\n<p>The cool thing is, the output is already on GPU, so you can save even more time by processing the images on GPU as well : resizing and rescaling is super fast then.</p>\n<p>Implementation is a bit tricky, I had to modify a file in the Dali package for it to work with uint16, and process everything per chunk otherwise your run out of disk space.</p>\n<p><strong>Code is here :</strong></p>\n<blockquote>\n  <p><a href=\"https://www.kaggle.com/code/theoviel/rsna-breast-baseline-faster-inference-with-dali\" target=\"_blank\">https://www.kaggle.com/code/theoviel/rsna-breast-baseline-faster-inference-with-dali</a></p>\n</blockquote>\n<p>The runtime can be further improved by using batching, but I'm already happy with a 2x speed up :)</p>\n<p>EDIT: Added <a href=\"https://www.kaggle.com/code/hengck23/combine-dali-and-dicomsdl-for-reading-dicom-files\" target=\"_blank\">dicomsdl</a>, I think I gain an extra hour with it. Thanks <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> !</p>\n<p>(The goal of the kernel is not to share a 10th place model, which is why I deleted the dataset, but rather to give everyone a working pipeline to use GPUs efficiently &amp; save submission time. I did not expect the submission to score 0.46 to be honest. I may re-upload the weights later once people have reached higher LB scores.)</p>",
      "rawMarkdown": "I've made an inference notebook public, that runs in about 4h !\n\nTrick used is shared by @tivfrvqhs5 [here](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/371534), basically you can process jpeg-compressed dicoms on GPU with NVIDIA Dali. \n\nThe cool thing is, the output is already on GPU, so you can save even more time by processing the images on GPU as well : resizing and rescaling is super fast then.\n\nImplementation is a bit tricky, I had to modify a file in the Dali package for it to work with uint16, and process everything per chunk otherwise your run out of disk space.\n\n**Code is here :**\n> https://www.kaggle.com/code/theoviel/rsna-breast-baseline-faster-inference-with-dali\n\nThe runtime can be further improved by using batching, but I'm already happy with a 2x speed up :)\n\nEDIT: Added [dicomsdl](https://www.kaggle.com/code/hengck23/combine-dali-and-dicomsdl-for-reading-dicom-files), I think I gain an extra hour with it. Thanks @hengck23 !\n\n\n(The goal of the kernel is not to share a 10th place model, which is why I deleted the dataset, but rather to give everyone a working pipeline to use GPUs efficiently & save submission time. I did not expect the submission to score 0.46 to be honest. I may re-upload the weights later once people have reached higher LB scores.)",
      "votes": 36
    },
    {
      "id": 2065132,
      "postDate": "2022-12-14T10:29:47.280Z",
      "content": "<p>if you are using 8bit png, dicomsdl has a faster function  dicomsdl.util.convert_to_uint8</p>\n<pre><code>outarr = np.empty(shape, dtype=dtype)\nself.copyFrameData(index, outarr)\n\nxmin = outarr.min()\nxmax = outarr.max()\ndata8 = np.empty_like(outarr, dtype=np.uint8)\nutil.convert_to_uint8(outarr, data8, xmin, xmax)  #equaivalent to       outarr =( (outarr - min) / (max - min)  * 255).astype(np.uint8)\n</code></pre>\n<p>.</p>",
      "rawMarkdown": "if you are using 8bit png, dicomsdl has a faster function  dicomsdl.util.convert_to_uint8\n\n```\noutarr = np.empty(shape, dtype=dtype)\nself.copyFrameData(index, outarr)\n\nxmin = outarr.min()\nxmax = outarr.max()\ndata8 = np.empty_like(outarr, dtype=np.uint8)\nutil.convert_to_uint8(outarr, data8, xmin, xmax)  #equaivalent to       outarr =( (outarr - min) / (max - min)  * 255).astype(np.uint8)\n\n```\n.\n ",
      "votes": 3
    },
    {
      "id": 2074261,
      "postDate": "2022-12-23T23:13:41.743Z",
      "content": "<p>I did some debugging and I found out that the tensors before normalisation to 0,1 are in the range from 0 to 4095. It indicates that in fact 12 bits are used instead of 16. Does anyone know if this is a typical behaviour, or is it specific to this solution based on Dali?</p>",
      "rawMarkdown": "I did some debugging and I found out that the tensors before normalisation to 0,1 are in the range from 0 to 4095. It indicates that in fact 12 bits are used instead of 16. Does anyone know if this is a typical behaviour, or is it specific to this solution based on Dali?",
      "votes": 1
    },
    {
      "id": 2074135,
      "postDate": "2022-12-23T18:35:33.773Z",
      "content": "<p>I am getting different results than with the standard method. Has anyone experienced the same issue?<br>\nI am using interpolation cv.inter_nearest in my preprocessing, which I set respectively (as well as torch interpolation to \"nearest\" which according to the official torch documentation is the same as the method implemented in CV). The results are different and I don't know if the only way is to re-create the dataset using this method + re-train my models, or are you aware of any possible solution? I thought about a potential loss of precision in some conversions throughout the process, but I couldn't solve it.</p>",
      "rawMarkdown": "I am getting different results than with the standard method. Has anyone experienced the same issue?\nI am using interpolation cv.inter_nearest in my preprocessing, which I set respectively (as well as torch interpolation to \"nearest\" which according to the official torch documentation is the same as the method implemented in CV). The results are different and I don't know if the only way is to re-create the dataset using this method + re-train my models, or are you aware of any possible solution? I thought about a potential loss of precision in some conversions throughout the process, but I couldn't solve it.",
      "votes": 1,
      "replies": [
        {
          "id": 2158675,
          "postDate": "2023-02-25T04:13:04.450Z",
          "content": "<p>Hi did you manage to resolve this issue? I used a preprocessing pipeline that does aspect ratio resize to 1024x512 as pngs. Does that mean I have to retrain my models to be in sync with DALI's preprocessing?</p>",
          "rawMarkdown": "Hi did you manage to resolve this issue? I used a preprocessing pipeline that does aspect ratio resize to 1024x512 as pngs. Does that mean I have to retrain my models to be in sync with DALI's preprocessing?"
        }
      ]
    },
    {
      "id": 2065119,
      "postDate": "2022-12-14T10:18:23.440Z",
      "content": "<p>Did you train the model with the prev. dicom jpgs ds. and did inference on the new framework/config.  ?</p>",
      "rawMarkdown": "Did you train the model with the prev. dicom jpgs ds. and did inference on the new framework/config.  ?",
      "votes": 1,
      "replies": [
        {
          "id": 2065123,
          "postDate": "2022-12-14T10:21:08.733Z",
          "content": "<p>Yep, the new preprocessing creates the same images as the previous ones, so you can do that freely !</p>",
          "rawMarkdown": "Yep, the new preprocessing creates the same images as the previous ones, so you can do that freely !",
          "votes": 1
        }
      ]
    },
    {
      "id": 2064531,
      "postDate": "2022-12-13T21:04:47.820Z",
      "content": "<p>Does it make sense for incomplete notebooks to rank high on sort by score?</p>\n<p><a href=\"https://www.kaggle.com/discussions/product-feedback/371249\" target=\"_blank\">https://www.kaggle.com/discussions/product-feedback/371249</a></p>",
      "rawMarkdown": "Does it make sense for incomplete notebooks to rank high on sort by score?\n\nhttps://www.kaggle.com/discussions/product-feedback/371249"
    },
    {
      "id": 2064130,
      "postDate": "2022-12-13T14:38:31.597Z",
      "content": "<p>\"Also, dicomsdl does not bring any speed up\"</p>\n<p>there is no improvement if you use CPU notebook.<br>\nthere is speedup of about 1.5 if you choose P100 GPU notebook.</p>\n<p>for dual T4 GPU, i have not tried</p>",
      "rawMarkdown": "\"Also, dicomsdl does not bring any speed up\"\n\nthere is no improvement if you use CPU notebook.\nthere is speedup of about 1.5 if you choose P100 GPU notebook.\n\nfor dual T4 GPU, i have not tried",
      "replies": [
        {
          "id": 2064157,
          "postDate": "2022-12-13T14:56:26.867Z",
          "content": "<blockquote>\n  <p>there is no improvement if you use CPU notebook.<br>\n  there is speedup of about 1.5 if you choose P100 GPU notebook.</p>\n</blockquote>\n<p>that does not really make much sense, either it gives a speedup or not</p>\n<p>the reason could rather be that you are getting different CPU processors, which can be randomly assigned to your kernel</p>",
          "rawMarkdown": "> there is no improvement if you use CPU notebook.\nthere is speedup of about 1.5 if you choose P100 GPU notebook.\n\nthat does not really make much sense, either it gives a speedup or not\n\nthe reason could rather be that you are getting different CPU processors, which can be randomly assigned to your kernel",
          "votes": 1
        },
        {
          "id": 2064227,
          "postDate": "2022-12-13T15:48:39.193Z",
          "content": "<p>Ran some more experiments.</p>\n<p>dicomsdl is faster inside loops, despite being slower outside.<br>\nThis seems true both on cpu and gpu.</p>\n<p>So we should use dicomsdl =)</p>",
          "rawMarkdown": "Ran some more experiments.\n\ndicomsdl is faster inside loops, despite being slower outside.\nThis seems true both on cpu and gpu.\n\nSo we should use dicomsdl =)",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2065132,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-12-14T10:29:47.280000",
      "content": "<p>if you are using 8bit png, dicomsdl has a faster function  dicomsdl.util.convert_to_uint8</p>\n<pre><code>outarr = np.empty(shape, dtype=dtype)\nself.copyFrameData(index, outarr)\n\nxmin = outarr.min()\nxmax = outarr.max()\ndata8 = np.empty_like(outarr, dtype=np.uint8)\nutil.convert_to_uint8(outarr, data8, xmin, xmax)  #equaivalent to       outarr =( (outarr - min) / (max - min)  * 255).astype(np.uint8)\n</code></pre>\n<p>.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2074261,
      "author_name": "Michał Choiński",
      "author_url": "",
      "post_date": "2022-12-23T23:13:41.743000",
      "content": "<p>I did some debugging and I found out that the tensors before normalisation to 0,1 are in the range from 0 to 4095. It indicates that in fact 12 bits are used instead of 16. Does anyone know if this is a typical behaviour, or is it specific to this solution based on Dali?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2074135,
      "author_name": "Michał Choiński",
      "author_url": "",
      "post_date": "2022-12-23T18:35:33.773000",
      "content": "<p>I am getting different results than with the standard method. Has anyone experienced the same issue?<br>\nI am using interpolation cv.inter_nearest in my preprocessing, which I set respectively (as well as torch interpolation to \"nearest\" which according to the official torch documentation is the same as the method implemented in CV). The results are different and I don't know if the only way is to re-create the dataset using this method + re-train my models, or are you aware of any possible solution? I thought about a potential loss of precision in some conversions throughout the process, but I couldn't solve it.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2158675,
          "author_name": "gao-hongnan",
          "author_url": "",
          "post_date": "2023-02-25T04:13:04.450000",
          "content": "<p>Hi did you manage to resolve this issue? I used a preprocessing pipeline that does aspect ratio resize to 1024x512 as pngs. Does that mean I have to retrain my models to be in sync with DALI's preprocessing?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2065119,
      "author_name": "Kirderf",
      "author_url": "",
      "post_date": "2022-12-14T10:18:23.440000",
      "content": "<p>Did you train the model with the prev. dicom jpgs ds. and did inference on the new framework/config.  ?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2065123,
          "author_name": "Theo Viel",
          "author_url": "",
          "post_date": "2022-12-14T10:21:08.733000",
          "content": "<p>Yep, the new preprocessing creates the same images as the previous ones, so you can do that freely !</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2064531,
      "author_name": "@kaggleqrdl",
      "author_url": "",
      "post_date": "2022-12-13T21:04:47.820000",
      "content": "<p>Does it make sense for incomplete notebooks to rank high on sort by score?</p>\n<p><a href=\"https://www.kaggle.com/discussions/product-feedback/371249\" target=\"_blank\">https://www.kaggle.com/discussions/product-feedback/371249</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2064130,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-12-13T14:38:31.597000",
      "content": "<p>\"Also, dicomsdl does not bring any speed up\"</p>\n<p>there is no improvement if you use CPU notebook.<br>\nthere is speedup of about 1.5 if you choose P100 GPU notebook.</p>\n<p>for dual T4 GPU, i have not tried</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2064157,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2022-12-13T14:56:26.867000",
          "content": "<blockquote>\n  <p>there is no improvement if you use CPU notebook.<br>\n  there is speedup of about 1.5 if you choose P100 GPU notebook.</p>\n</blockquote>\n<p>that does not really make much sense, either it gives a speedup or not</p>\n<p>the reason could rather be that you are getting different CPU processors, which can be randomly assigned to your kernel</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2064227,
          "author_name": "Theo Viel",
          "author_url": "",
          "post_date": "2022-12-13T15:48:39.193000",
          "content": "<p>Ran some more experiments.</p>\n<p>dicomsdl is faster inside loops, despite being slower outside.<br>\nThis seems true both on cpu and gpu.</p>\n<p>So we should use dicomsdl =)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2064105": "I've made an inference notebook public, that runs in about 4h !\n\nTrick used is shared by @tivfrvqhs5 [here](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/371534), basically you can process jpeg-compressed dicoms on GPU with NVIDIA Dali. \n\nThe cool thing is, the output is already on GPU, so you can save even more time by processing the images on GPU as well : resizing and rescaling is super fast then.\n\nImplementation is a bit tricky, I had to modify a file in the Dali package for it to work with uint16, and process everything per chunk otherwise your run out of disk space.\n\n**Code is here :**\n> https://www.kaggle.com/code/theoviel/rsna-breast-baseline-faster-inference-with-dali\n\nThe runtime can be further improved by using batching, but I'm already happy with a 2x speed up :)\n\nEDIT: Added [dicomsdl](https://www.kaggle.com/code/hengck23/combine-dali-and-dicomsdl-for-reading-dicom-files), I think I gain an extra hour with it. Thanks @hengck23 !\n\n\n(The goal of the kernel is not to share a 10th place model, which is why I deleted the dataset, but rather to give everyone a working pipeline to use GPUs efficiently & save submission time. I did not expect the submission to score 0.46 to be honest. I may re-upload the weights later once people have reached higher LB scores.)",
    "2065132": "if you are using 8bit png, dicomsdl has a faster function  dicomsdl.util.convert_to_uint8\n\n```\noutarr = np.empty(shape, dtype=dtype)\nself.copyFrameData(index, outarr)\n\nxmin = outarr.min()\nxmax = outarr.max()\ndata8 = np.empty_like(outarr, dtype=np.uint8)\nutil.convert_to_uint8(outarr, data8, xmin, xmax)  #equaivalent to       outarr =( (outarr - min) / (max - min)  * 255).astype(np.uint8)\n\n```\n.\n ",
    "2074261": "I did some debugging and I found out that the tensors before normalisation to 0,1 are in the range from 0 to 4095. It indicates that in fact 12 bits are used instead of 16. Does anyone know if this is a typical behaviour, or is it specific to this solution based on Dali?",
    "2074135": "I am getting different results than with the standard method. Has anyone experienced the same issue?\nI am using interpolation cv.inter_nearest in my preprocessing, which I set respectively (as well as torch interpolation to \"nearest\" which according to the official torch documentation is the same as the method implemented in CV). The results are different and I don't know if the only way is to re-create the dataset using this method + re-train my models, or are you aware of any possible solution? I thought about a potential loss of precision in some conversions throughout the process, but I couldn't solve it.",
    "2065119": "Did you train the model with the prev. dicom jpgs ds. and did inference on the new framework/config.  ?",
    "2064531": "Does it make sense for incomplete notebooks to rank high on sort by score?\n\nhttps://www.kaggle.com/discussions/product-feedback/371249",
    "2064130": "\"Also, dicomsdl does not bring any speed up\"\n\nthere is no improvement if you use CPU notebook.\nthere is speedup of about 1.5 if you choose P100 GPU notebook.\n\nfor dual T4 GPU, i have not tried"
  }
}