{
  "id": 373896,
  "title": "Performance optimization for decoding libraries, seems like an unecessary digression",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/373896",
  "author_name": "@kaggleqrdl",
  "post_date": "2022-12-24T01:21:25.374000",
  "votes": 7,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I've been spending a lot of time looking into performance optimizations.  The dali work already done was a huge benefit, but I believe there is more to be done as we have ~32000 images we're decoding on submission and each hour we save on decoding means another hour we can spend on inferrence.  </p>\n<p>A lot of the ideas I've been looking into could really benefit from that extra hour.</p>\n<p>Was the comp designed for this reason?  Perhaps for scanning large libraries of images? If so, then I suppose it makes sense and no need to read any further.</p>\n<p>If not, then I think we're doing the hosts a fairly big disservice and it would be great to get the images transformed at least into jpeg2000 or something faster.</p>",
  "messages": [
    {
      "id": 2074310,
      "postDate": "2022-12-24T01:21:25.373Z",
      "content": "<p>I've been spending a lot of time looking into performance optimizations.  The dali work already done was a huge benefit, but I believe there is more to be done as we have ~32000 images we're decoding on submission and each hour we save on decoding means another hour we can spend on inferrence.  </p>\n<p>A lot of the ideas I've been looking into could really benefit from that extra hour.</p>\n<p>Was the comp designed for this reason?  Perhaps for scanning large libraries of images? If so, then I suppose it makes sense and no need to read any further.</p>\n<p>If not, then I think we're doing the hosts a fairly big disservice and it would be great to get the images transformed at least into jpeg2000 or something faster.</p>",
      "rawMarkdown": "I've been spending a lot of time looking into performance optimizations.  The dali work already done was a huge benefit, but I believe there is more to be done as we have ~32000 images we're decoding on submission and each hour we save on decoding means another hour we can spend on inferrence.  \n\nA lot of the ideas I've been looking into could really benefit from that extra hour.\n\nWas the comp designed for this reason?  Perhaps for scanning large libraries of images? If so, then I suppose it makes sense and no need to read any further.\n\nIf not, then I think we're doing the hosts a fairly big disservice and it would be great to get the images transformed at least into jpeg2000 or something faster.\n\n\n",
      "votes": 7
    },
    {
      "id": 2074316,
      "postDate": "2022-12-24T02:03:51.577Z",
      "content": "<p>I think it would be reasonable to run a one-time conversion to jpeg2000 on the dcms (or at least provide it in an alternative directory). I mean every competitor has to spend time doing <strong>the exact same</strong> JPEG decoding on <strong>every</strong> dcm file. We only have two formats in the competition dataset, if wider support is needed, it would be pretty easy to extend to other formats in the open-source version for an E2E solution.</p>",
      "rawMarkdown": "I think it would be reasonable to run a one-time conversion to jpeg2000 on the dcms (or at least provide it in an alternative directory). I mean every competitor has to spend time doing **the exact same** JPEG decoding on **every** dcm file. We only have two formats in the competition dataset, if wider support is needed, it would be pretty easy to extend to other formats in the open-source version for an E2E solution.",
      "votes": 4,
      "replies": [
        {
          "id": 2076240,
          "postDate": "2022-12-26T09:18:03.870Z",
          "content": "<p>Yeah, with the 2x GPU arch, if everything was jpeg2000 this actually would be very fast.  nvjpeg is 2x as fast as dicomsdl.</p>",
          "rawMarkdown": "Yeah, with the 2x GPU arch, if everything was jpeg2000 this actually would be very fast.  nvjpeg is 2x as fast as dicomsdl."
        }
      ]
    },
    {
      "id": 2074464,
      "postDate": "2022-12-24T07:43:36.490Z",
      "content": "<p>Maybe it (file processing) was one of the goal of this competition. Now it looks like problem is under control and we can focus on cancer classification problem.</p>",
      "rawMarkdown": "Maybe it (file processing) was one of the goal of this competition. Now it looks like problem is under control and we can focus on cancer classification problem.",
      "votes": 1,
      "replies": [
        {
          "id": 2074471,
          "postDate": "2022-12-24T07:54:23.730Z",
          "content": "<p>Remek, just as a datapoint, how long does it take you to decode 32K files using 2 cpus?  Ideally not in GPU space (or at least copy back to CPU) as there's some augmentation stuff that needs to be done off gpu.</p>\n<p>I think if we can get decoding down to around 30 to 45 minutes, that should be enough. The problem I keep running into is that anything I do, I have to keep making 'dumber' because it's going to go over time - so anything I can eek out in decoding helps.</p>",
          "rawMarkdown": "Remek, just as a datapoint, how long does it take you to decode 32K files using 2 cpus?  Ideally not in GPU space (or at least copy back to CPU) as there's some augmentation stuff that needs to be done off gpu.\n\nI think if we can get decoding down to around 30 to 45 minutes, that should be enough. The problem I keep running into is that anything I do, I have to keep making 'dumber' because it's going to go over time - so anything I can eek out in decoding helps.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2076185,
      "postDate": "2022-12-26T07:52:44.227Z",
      "content": "<p>I've created a baseline version here, it takes 8253.6s or 2.3 hours.    <a href=\"https://www.kaggle.com/code/kaggleqrdl/baseline-read-32000-images\" target=\"_blank\">https://www.kaggle.com/code/kaggleqrdl/baseline-read-32000-images</a></p>\n<p>It uses only dicomsdl.  The nvjpeg2000 notebook is fairly fast and that's the next thing I'm going to try.</p>\n<p>Dali is cool, but my concern is that it limits you to what you can do in terms of augmentation if you have to do everything on the GPU.  That said, it's possible this might be the only way to seriously cut down time.   </p>\n<p>Basically, I am going to try to get another hour cut off this time.  It's my belief that an extra hour could be the difference between a winning notebook and one that doesn't perform as well.</p>\n<p>For those who've downvoted this thread, it'd be great if you could explain why you've downvoted as I am a bit puzzled.  Sharing a notebook which out performs the above would be great as well.</p>\n<p><a href=\"https://www.kaggle.com/code/kaggleqrdl/baseline-32000-nvjpeg2000-dicomsdl\" target=\"_blank\">https://www.kaggle.com/code/kaggleqrdl/baseline-32000-nvjpeg2000-dicomsdl</a>   Using nvjpeg, can reduce this to about 1/2 the time, or 1.15 hours.</p>",
      "rawMarkdown": "I've created a baseline version here, it takes 8253.6s or 2.3 hours.    https://www.kaggle.com/code/kaggleqrdl/baseline-read-32000-images\n\nIt uses only dicomsdl.  The nvjpeg2000 notebook is fairly fast and that's the next thing I'm going to try.\n\nDali is cool, but my concern is that it limits you to what you can do in terms of augmentation if you have to do everything on the GPU.  That said, it's possible this might be the only way to seriously cut down time.   \n\nBasically, I am going to try to get another hour cut off this time.  It's my belief that an extra hour could be the difference between a winning notebook and one that doesn't perform as well.\n\nFor those who've downvoted this thread, it'd be great if you could explain why you've downvoted as I am a bit puzzled.  Sharing a notebook which out performs the above would be great as well.\n\nhttps://www.kaggle.com/code/kaggleqrdl/baseline-32000-nvjpeg2000-dicomsdl   Using nvjpeg, can reduce this to about 1/2 the time, or 1.15 hours.",
      "votes": 2
    },
    {
      "id": 2125370,
      "postDate": "2023-02-01T16:06:13.613Z",
      "content": "<p>thanks for sharing </p>",
      "rawMarkdown": "thanks for sharing \n"
    }
  ],
  "comments": [
    {
      "id": 2074316,
      "author_name": "outwrest",
      "author_url": "",
      "post_date": "2022-12-24T02:03:51.577000",
      "content": "<p>I think it would be reasonable to run a one-time conversion to jpeg2000 on the dcms (or at least provide it in an alternative directory). I mean every competitor has to spend time doing <strong>the exact same</strong> JPEG decoding on <strong>every</strong> dcm file. We only have two formats in the competition dataset, if wider support is needed, it would be pretty easy to extend to other formats in the open-source version for an E2E solution.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2076240,
          "author_name": "@kaggleqrdl",
          "author_url": "",
          "post_date": "2022-12-26T09:18:03.870000",
          "content": "<p>Yeah, with the 2x GPU arch, if everything was jpeg2000 this actually would be very fast.  nvjpeg is 2x as fast as dicomsdl.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2074464,
      "author_name": "Remek Kinas",
      "author_url": "",
      "post_date": "2022-12-24T07:43:36.490000",
      "content": "<p>Maybe it (file processing) was one of the goal of this competition. Now it looks like problem is under control and we can focus on cancer classification problem.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2074471,
          "author_name": "@kaggleqrdl",
          "author_url": "",
          "post_date": "2022-12-24T07:54:23.730000",
          "content": "<p>Remek, just as a datapoint, how long does it take you to decode 32K files using 2 cpus?  Ideally not in GPU space (or at least copy back to CPU) as there's some augmentation stuff that needs to be done off gpu.</p>\n<p>I think if we can get decoding down to around 30 to 45 minutes, that should be enough. The problem I keep running into is that anything I do, I have to keep making 'dumber' because it's going to go over time - so anything I can eek out in decoding helps.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2076185,
      "author_name": "@kaggleqrdl",
      "author_url": "",
      "post_date": "2022-12-26T07:52:44.227000",
      "content": "<p>I've created a baseline version here, it takes 8253.6s or 2.3 hours.    <a href=\"https://www.kaggle.com/code/kaggleqrdl/baseline-read-32000-images\" target=\"_blank\">https://www.kaggle.com/code/kaggleqrdl/baseline-read-32000-images</a></p>\n<p>It uses only dicomsdl.  The nvjpeg2000 notebook is fairly fast and that's the next thing I'm going to try.</p>\n<p>Dali is cool, but my concern is that it limits you to what you can do in terms of augmentation if you have to do everything on the GPU.  That said, it's possible this might be the only way to seriously cut down time.   </p>\n<p>Basically, I am going to try to get another hour cut off this time.  It's my belief that an extra hour could be the difference between a winning notebook and one that doesn't perform as well.</p>\n<p>For those who've downvoted this thread, it'd be great if you could explain why you've downvoted as I am a bit puzzled.  Sharing a notebook which out performs the above would be great as well.</p>\n<p><a href=\"https://www.kaggle.com/code/kaggleqrdl/baseline-32000-nvjpeg2000-dicomsdl\" target=\"_blank\">https://www.kaggle.com/code/kaggleqrdl/baseline-32000-nvjpeg2000-dicomsdl</a>   Using nvjpeg, can reduce this to about 1/2 the time, or 1.15 hours.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2125370,
      "author_name": "Nguyen Thi Thanh Hoa",
      "author_url": "",
      "post_date": "2023-02-01T16:06:13.613000",
      "content": "<p>thanks for sharing </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2074310": "I've been spending a lot of time looking into performance optimizations.  The dali work already done was a huge benefit, but I believe there is more to be done as we have ~32000 images we're decoding on submission and each hour we save on decoding means another hour we can spend on inferrence.  \n\nA lot of the ideas I've been looking into could really benefit from that extra hour.\n\nWas the comp designed for this reason?  Perhaps for scanning large libraries of images? If so, then I suppose it makes sense and no need to read any further.\n\nIf not, then I think we're doing the hosts a fairly big disservice and it would be great to get the images transformed at least into jpeg2000 or something faster.\n\n\n",
    "2074316": "I think it would be reasonable to run a one-time conversion to jpeg2000 on the dcms (or at least provide it in an alternative directory). I mean every competitor has to spend time doing **the exact same** JPEG decoding on **every** dcm file. We only have two formats in the competition dataset, if wider support is needed, it would be pretty easy to extend to other formats in the open-source version for an E2E solution.",
    "2074464": "Maybe it (file processing) was one of the goal of this competition. Now it looks like problem is under control and we can focus on cancer classification problem.",
    "2076185": "I've created a baseline version here, it takes 8253.6s or 2.3 hours.    https://www.kaggle.com/code/kaggleqrdl/baseline-read-32000-images\n\nIt uses only dicomsdl.  The nvjpeg2000 notebook is fairly fast and that's the next thing I'm going to try.\n\nDali is cool, but my concern is that it limits you to what you can do in terms of augmentation if you have to do everything on the GPU.  That said, it's possible this might be the only way to seriously cut down time.   \n\nBasically, I am going to try to get another hour cut off this time.  It's my belief that an extra hour could be the difference between a winning notebook and one that doesn't perform as well.\n\nFor those who've downvoted this thread, it'd be great if you could explain why you've downvoted as I am a bit puzzled.  Sharing a notebook which out performs the above would be great as well.\n\nhttps://www.kaggle.com/code/kaggleqrdl/baseline-32000-nvjpeg2000-dicomsdl   Using nvjpeg, can reduce this to about 1/2 the time, or 1.15 hours.",
    "2125370": "thanks for sharing \n"
  }
}