{
  "id": 371407,
  "title": "Scoring Time Trouble",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/371407",
  "author_name": "happymentee",
  "post_date": "2022-12-09T19:03:42.580000",
  "votes": 0,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hi, I'm new to Kaggle, so I am slightly confused about scoring time. <br>\nI just used the following:</p>\n<ul>\n<li>A single model (28MB). </li>\n<li>Data processing: Dicom -&gt; 256x256 array -&gt; 256x256x3 array -&gt; 3x256x256 tensor + normalize.</li>\n<li>batch_size: 32<br>\nBut my scoring time is about &gt;7h. Is this sound right? <br>\nIf not, what may be the reason for this problem</li>\n</ul>",
  "messages": [
    {
      "id": 2060367,
      "postDate": "2022-12-09T20:14:27.513Z",
      "content": "<p>Yeah, it's an issue.    As an example,  Vlad's notebook takes about 10 minutes to train but &gt;7h to infer.  :p   </p>\n<p>You can speed up decompressing/decoding by using the dicomsdl library, which helps somewhat.  See here for more info - <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/371033\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/371033</a></p>\n<p>The problem is that all the images have been compressed with dicom standard which is very slow to decompress.  Dicom is better in size than some other compression standards which are much much faster, but takes up to 2x the amount of space.   This may be an issue for Kaggle networks / storage infra, or they don't want to disrupt an ongoing comp to fix it.  Not sure!</p>\n<p>My guess is that dicom was designed with the idea in mind that folks would just decompress once, and iterate on that.  However, in this code comp, we're iterating many times on the compressed data.  It's notable that even the dali folks <a href=\"https://github.com/NVIDIA/DALI/issues/3774#issuecomment-1281350830\" target=\"_blank\">won't support dicom</a> because they don't see this as a valid use case.</p>\n<p>What's curious though is that we have 5 submissions per day.  Given the compute cost, you'd think that would be pushed down to 1 or 2 and folks would just rely on their cross validation score (which folks seem to be doing).  What do they want us to do with all those extra subs?  Probe the leaderboard?</p>",
      "rawMarkdown": "Yeah, it's an issue.    As an example,  Vlad's notebook takes about 10 minutes to train but >7h to infer.  :p   \n\nYou can speed up decompressing/decoding by using the dicomsdl library, which helps somewhat.  See here for more info - https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/371033\n\n\nThe problem is that all the images have been compressed with dicom standard which is very slow to decompress.  Dicom is better in size than some other compression standards which are much much faster, but takes up to 2x the amount of space.   This may be an issue for Kaggle networks / storage infra, or they don't want to disrupt an ongoing comp to fix it.  Not sure!\n\nMy guess is that dicom was designed with the idea in mind that folks would just decompress once, and iterate on that.  However, in this code comp, we're iterating many times on the compressed data.  It's notable that even the dali folks [won't support dicom](https://github.com/NVIDIA/DALI/issues/3774#issuecomment-1281350830) because they don't see this as a valid use case.\n\nWhat's curious though is that we have 5 submissions per day.  Given the compute cost, you'd think that would be pushed down to 1 or 2 and folks would just rely on their cross validation score (which folks seem to be doing).  What do they want us to do with all those extra subs?  Probe the leaderboard?\n\n\n",
      "votes": 1
    },
    {
      "id": 2060337,
      "postDate": "2022-12-09T19:03:42.580Z",
      "content": "<p>Hi, I'm new to Kaggle, so I am slightly confused about scoring time. <br>\nI just used the following:</p>\n<ul>\n<li>A single model (28MB). </li>\n<li>Data processing: Dicom -&gt; 256x256 array -&gt; 256x256x3 array -&gt; 3x256x256 tensor + normalize.</li>\n<li>batch_size: 32<br>\nBut my scoring time is about &gt;7h. Is this sound right? <br>\nIf not, what may be the reason for this problem</li>\n</ul>",
      "rawMarkdown": "Hi, I'm new to Kaggle, so I am slightly confused about scoring time. \nI just used the following:\n- A single model (28MB). \n- Data processing: Dicom -> 256x256 array -> 256x256x3 array -> 3x256x256 tensor + normalize.\n- batch_size: 32\nBut my scoring time is about >7h. Is this sound right? \nIf not, what may be the reason for this problem"
    }
  ],
  "comments": [
    {
      "id": 2060367,
      "author_name": "@kaggleqrdl",
      "author_url": "",
      "post_date": "2022-12-09T20:14:27.513000",
      "content": "<p>Yeah, it's an issue.    As an example,  Vlad's notebook takes about 10 minutes to train but &gt;7h to infer.  :p   </p>\n<p>You can speed up decompressing/decoding by using the dicomsdl library, which helps somewhat.  See here for more info - <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/371033\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/371033</a></p>\n<p>The problem is that all the images have been compressed with dicom standard which is very slow to decompress.  Dicom is better in size than some other compression standards which are much much faster, but takes up to 2x the amount of space.   This may be an issue for Kaggle networks / storage infra, or they don't want to disrupt an ongoing comp to fix it.  Not sure!</p>\n<p>My guess is that dicom was designed with the idea in mind that folks would just decompress once, and iterate on that.  However, in this code comp, we're iterating many times on the compressed data.  It's notable that even the dali folks <a href=\"https://github.com/NVIDIA/DALI/issues/3774#issuecomment-1281350830\" target=\"_blank\">won't support dicom</a> because they don't see this as a valid use case.</p>\n<p>What's curious though is that we have 5 submissions per day.  Given the compute cost, you'd think that would be pushed down to 1 or 2 and folks would just rely on their cross validation score (which folks seem to be doing).  What do they want us to do with all those extra subs?  Probe the leaderboard?</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2060367": "Yeah, it's an issue.    As an example,  Vlad's notebook takes about 10 minutes to train but >7h to infer.  :p   \n\nYou can speed up decompressing/decoding by using the dicomsdl library, which helps somewhat.  See here for more info - https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/371033\n\n\nThe problem is that all the images have been compressed with dicom standard which is very slow to decompress.  Dicom is better in size than some other compression standards which are much much faster, but takes up to 2x the amount of space.   This may be an issue for Kaggle networks / storage infra, or they don't want to disrupt an ongoing comp to fix it.  Not sure!\n\nMy guess is that dicom was designed with the idea in mind that folks would just decompress once, and iterate on that.  However, in this code comp, we're iterating many times on the compressed data.  It's notable that even the dali folks [won't support dicom](https://github.com/NVIDIA/DALI/issues/3774#issuecomment-1281350830) because they don't see this as a valid use case.\n\nWhat's curious though is that we have 5 submissions per day.  Given the compute cost, you'd think that would be pushed down to 1 or 2 and folks would just rely on their cross validation score (which folks seem to be doing).  What do they want us to do with all those extra subs?  Probe the leaderboard?\n\n\n",
    "2060337": "Hi, I'm new to Kaggle, so I am slightly confused about scoring time. \nI just used the following:\n- A single model (28MB). \n- Data processing: Dicom -> 256x256 array -> 256x256x3 array -> 3x256x256 tensor + normalize.\n- batch_size: 32\nBut my scoring time is about >7h. Is this sound right? \nIf not, what may be the reason for this problem"
  }
}