{
  "id": 427217,
  "title": "Standardizing Unusual Dicoms",
  "url": "/competitions/rsna-2023-abdominal-trauma-detection/discussion/427217",
  "author_name": "Hui Ming Lin",
  "post_date": "2023-07-26T23:25:03.827000",
  "votes": 45,
  "comment_count": 12,
  "views": 0,
  "content": "<h1>Unusual DICOM Files</h1>\n<p>The RSNA 2023 Abdominal Trauma Detection Challenge showcases an exceptionally heterogeneous DICOM data structure due to the large number of institutions that donated data.</p>\n<p>Certain DICOM pixel arrays will appear unusual when opened using the pydicom library.</p>\n<p>Every DICOM file contains a field called <a href=\"https://dicom.innolitics.com/ciods/rt-dose/image-pixel/00280103\" target=\"_blank\">Pixel Representation</a> indicating if the pixel datatype is unsigned(0) or signed (1).</p>\n<p>Some DICOM files with PixelRepresentation == 1 exhibit visually unusual characteristics. You can use the [train/test]_dicom_tags.parquet files to identify images likely to be affected.</p>\n<h1>Potential Solution</h1>\n<p>The following function is a potential solution to the described problem.</p>\n<pre><code> standardize_pixel_array(dcm: pydicom.dataset.) -&gt; np.ndarray:\n    #   pixel_array   == \n    pixel_array = dcm.pixel_array\n     dcm. == :\n        bit_shift = dcm. - dcm.\n        d = pixel_array.dtype \n        pixel_array = (pixel_array &lt;&lt; bit_shift).(dtype) &gt;&gt;  bit_shift\n    return pixel_array\n</code></pre>\n<p>Please see this <a href=\"https://www.kaggle.com/code/huiminglin/rsna-2023-abdomen-dicom-fix\" target=\"_blank\">notebook</a> for example implementation. We did not apply this processing step to the competition dataset as we did not have time to confirm that it properly resolves the issue in all cases and does not have any unintended side effects.</p>",
  "messages": [
    {
      "id": 2360696,
      "postDate": "2023-07-26T23:25:03.827Z",
      "content": "<h1>Unusual DICOM Files</h1>\n<p>The RSNA 2023 Abdominal Trauma Detection Challenge showcases an exceptionally heterogeneous DICOM data structure due to the large number of institutions that donated data.</p>\n<p>Certain DICOM pixel arrays will appear unusual when opened using the pydicom library.</p>\n<p>Every DICOM file contains a field called <a href=\"https://dicom.innolitics.com/ciods/rt-dose/image-pixel/00280103\" target=\"_blank\">Pixel Representation</a> indicating if the pixel datatype is unsigned(0) or signed (1).</p>\n<p>Some DICOM files with PixelRepresentation == 1 exhibit visually unusual characteristics. You can use the [train/test]_dicom_tags.parquet files to identify images likely to be affected.</p>\n<h1>Potential Solution</h1>\n<p>The following function is a potential solution to the described problem.</p>\n<pre><code> standardize_pixel_array(dcm: pydicom.dataset.) -&gt; np.ndarray:\n    #   pixel_array   == \n    pixel_array = dcm.pixel_array\n     dcm. == :\n        bit_shift = dcm. - dcm.\n        d = pixel_array.dtype \n        pixel_array = (pixel_array &lt;&lt; bit_shift).(dtype) &gt;&gt;  bit_shift\n    return pixel_array\n</code></pre>\n<p>Please see this <a href=\"https://www.kaggle.com/code/huiminglin/rsna-2023-abdomen-dicom-fix\" target=\"_blank\">notebook</a> for example implementation. We did not apply this processing step to the competition dataset as we did not have time to confirm that it properly resolves the issue in all cases and does not have any unintended side effects.</p>",
      "rawMarkdown": "# Unusual DICOM Files\n\nThe RSNA 2023 Abdominal Trauma Detection Challenge showcases an exceptionally heterogeneous DICOM data structure due to the large number of institutions that donated data.\n\nCertain DICOM pixel arrays will appear unusual when opened using the pydicom library.\n\nEvery DICOM file contains a field called [Pixel Representation](https://dicom.innolitics.com/ciods/rt-dose/image-pixel/00280103) indicating if the pixel datatype is unsigned(0) or signed (1).\n\nSome DICOM files with PixelRepresentation == 1 exhibit visually unusual characteristics. You can use the [train/test]_dicom_tags.parquet files to identify images likely to be affected.\n\n\n# Potential Solution\n\nThe following function is a potential solution to the described problem.\n\n```\ndef standardize_pixel_array(dcm: pydicom.dataset.FileDataset) -> np.ndarray:\n    # Correct DICOM pixel_array if PixelRepresentation == 1.\n    pixel_array = dcm.pixel_array\n    if dcm.PixelRepresentation == 1:\n        bit_shift = dcm.BitsAllocated - dcm.BitsStored\n        dtype = pixel_array.dtype \n        pixel_array = (pixel_array << bit_shift).astype(dtype) >>  bit_shift\n    return pixel_array\n```\n\n\nPlease see this [notebook](https://www.kaggle.com/code/huiminglin/rsna-2023-abdomen-dicom-fix) for example implementation. We did not apply this processing step to the competition dataset as we did not have time to confirm that it properly resolves the issue in all cases and does not have any unintended side effects.",
      "votes": 45
    },
    {
      "id": 2370495,
      "postDate": "2023-08-02T13:08:33.580Z",
      "content": "<p>Why the <code>apply_modality_lut</code> is applied only when <code>dcm.PixelRepresentation == 1</code>?<br>\nI think this should be applied to both, otherwise, neither.</p>\n<pre><code>    if dcm.PixelRepresentation == 1:\n            bit_shift = dcm.BitsAllocated - dcm.BitsStored\n            dtype = pixel_array.dtype \n\n\n\n    return pixel_array\n</code></pre>\n<p>pixel_array</p>",
      "rawMarkdown": "Why the `apply_modality_lut` is applied only when `dcm.PixelRepresentation == 1`?\nI think this should be applied to both, otherwise, neither.\n\n```diff\n    if dcm.PixelRepresentation == 1:\n            bit_shift = dcm.BitsAllocated - dcm.BitsStored\n            dtype = pixel_array.dtype \n!            pixel_array = (pixel_array << bit_shift).astype(dtype) >>  bit_shift\n-            pixel_array = pydicom.pixel_data_handlers.util.apply_modality_lut(new_array, dcm)\n+    pixel_array = pydicom.pixel_data_handlers.util.apply_modality_lut(pixel_array, dcm)\n    return pixel_array\n```\npixel_array",
      "votes": 6,
      "replies": [
        {
          "id": 2370520,
          "postDate": "2023-08-02T13:23:55.530Z",
          "content": "<p>i think so, too. apply_modality_lut function is used to convert pixel array to Hounsfield Unit (HU). it has nothing to do with PixelRepresentation (signed or unsigned data)</p>",
          "rawMarkdown": "i think so, too. apply_modality_lut function is used to convert pixel array to Hounsfield Unit (HU). it has nothing to do with PixelRepresentation (signed or unsigned data)",
          "votes": 1
        },
        {
          "id": 2370672,
          "postDate": "2023-08-02T15:06:50.460Z",
          "content": "<p>You are absolutely correct, thank you for pointing this out. I have removed apply_modality_lut from the function. The notebook has also been updated.</p>",
          "rawMarkdown": "You are absolutely correct, thank you for pointing this out. I have removed apply_modality_lut from the function. The notebook has also been updated."
        }
      ]
    },
    {
      "id": 2373646,
      "postDate": "2023-08-04T11:09:42.310Z",
      "content": "<p>I noticed a particular quirk of the dataset. DICOMs where <code>PixelRepresentation == 1</code> always returns <code>pixel_array</code> as a <code>np.int16</code> and other DICOMs always returns <code>pixel_array</code> as a <code>np.uint16</code>. Does this affect normalization in any way?</p>\n<p>Edit: By normalization, I meant the final output of the <code>standardize_pixel_array</code> function. Because the function will still return <code>np.int16</code> for <code>PixelRepresentation == 1</code> and <code>np.uint16</code> otherwise.</p>",
      "rawMarkdown": "I noticed a particular quirk of the dataset. DICOMs where `PixelRepresentation == 1` always returns `pixel_array` as a `np.int16` and other DICOMs always returns `pixel_array` as a `np.uint16`. Does this affect normalization in any way?\n\nEdit: By normalization, I meant the final output of the `standardize_pixel_array` function. Because the function will still return `np.int16` for `PixelRepresentation == 1` and `np.uint16` otherwise.",
      "votes": 4
    },
    {
      "id": 2426792,
      "postDate": "2023-09-06T20:30:07.313Z",
      "content": "<p>What do you mean by the visually unusual characteristics?</p>",
      "rawMarkdown": "What do you mean by the visually unusual characteristics?",
      "votes": 1
    },
    {
      "id": 2383574,
      "postDate": "2023-08-10T13:00:09.300Z",
      "content": "<p>please how to view dataset images</p>",
      "rawMarkdown": " please how to view dataset images\n"
    },
    {
      "id": 2465637,
      "postDate": "2023-10-03T06:50:21.100Z",
      "content": "<p>As a beginner in such competitions, I've taken a look at multiple notebooks. I'm questioning whether this will directly influence my results or if it's more about comprehending pixel representations.</p>",
      "rawMarkdown": "As a beginner in such competitions, I've taken a look at multiple notebooks. I'm questioning whether this will directly influence my results or if it's more about comprehending pixel representations."
    },
    {
      "id": 2383052,
      "postDate": "2023-08-10T06:47:55.457Z",
      "content": "<p>I don't understand how shifting bits would change the output. You left shift bits by some amount and then you right shift with the same amount which returns the same exact array. What am I missing?</p>",
      "rawMarkdown": "I don't understand how shifting bits would change the output. You left shift bits by some amount and then you right shift with the same amount which returns the same exact array. What am I missing?",
      "replies": [
        {
          "id": 2384908,
          "postDate": "2023-08-11T05:51:26.940Z",
          "content": "<p>Does this example help you?</p>\n<pre><code> numpy  np\n\nBitsAllocated=\nBitsStored=\nbit_shift = BitsAllocated - BitsStored\n\nx = np.array().astype(np.int8)  \nx_new = (x &lt;&lt; bit_shift).astype(np.int8) &gt;&gt;  bit_shift\n\n(np.binary_repr(x, BitsAllocated), )\n(np.binary_repr(x_new, BitsAllocated), )\n\n\n\n</code></pre>",
          "rawMarkdown": "Does this example help you?\n\n```py\nimport numpy as np\n\nBitsAllocated=8\nBitsStored=6\nbit_shift = BitsAllocated - BitsStored\n\nx = np.array(103).astype(np.int8)  # int8: -128:127\nx_new = (x << bit_shift).astype(np.int8) >>  bit_shift\n\nprint(np.binary_repr(x, BitsAllocated), f\"({x})\")\nprint(np.binary_repr(x_new, BitsAllocated), f\"({x_new})\")\n\n#> 01100111 (103)\n#> 11100111 (-25)\n```",
          "votes": 4,
          "replies": [
            {
              "id": 2385390,
              "postDate": "2023-08-11T09:56:19.730Z",
              "content": "<p>Thank you. I think I understood the problem. Even though those images are labeled as signed 16-bit integers (PixelRepresentation == 1), their pixel values are starting from 0. If the values are starting from 0, then we are not using some of the allocated bits.</p>\n<p><img src=\"https://i.ibb.co/9h6nLf6/Screenshot-from-2023-08-11-10-52-15.png\" alt=\"1\"></p>\n<p>When we left shift the bits, we make more bits used for a signed data type.</p>\n<p><img src=\"https://i.ibb.co/r6HXxxm/Screenshot-from-2023-08-11-10-52-35.png\" alt=\"2\"></p>\n<p>When we right shift with the same amount again, we truncate the last unused bits on the right.</p>\n<p><img src=\"https://i.ibb.co/qmGBHyj/Screenshot-from-2023-08-11-10-52-44.png\" alt=\"3\"> </p>",
              "rawMarkdown": "Thank you. I think I understood the problem. Even though those images are labeled as signed 16-bit integers (PixelRepresentation == 1), their pixel values are starting from 0. If the values are starting from 0, then we are not using some of the allocated bits.\n\n![1](https://i.ibb.co/9h6nLf6/Screenshot-from-2023-08-11-10-52-15.png)\n\nWhen we left shift the bits, we make more bits used for a signed data type.\n\n![2](https://i.ibb.co/r6HXxxm/Screenshot-from-2023-08-11-10-52-35.png)\n\nWhen we right shift with the same amount again, we truncate the last unused bits on the right.\n\n![3](https://i.ibb.co/qmGBHyj/Screenshot-from-2023-08-11-10-52-44.png) ",
              "votes": 6
            },
            {
              "id": 2388847,
              "postDate": "2023-08-13T16:46:55.617Z",
              "content": "<p>Need to check the BitsAllocated and BitsStored DICOM tags. Most CTs are actually 10-12 bits deep. Since there isn't a 10bit datatype, we must use 16 bit dtypes, then calculate the actual pixel width of the image.</p>",
              "rawMarkdown": "Need to check the BitsAllocated and BitsStored DICOM tags. Most CTs are actually 10-12 bits deep. Since there isn't a 10bit datatype, we must use 16 bit dtypes, then calculate the actual pixel width of the image.",
              "votes": 4
            }
          ]
        }
      ]
    },
    {
      "id": 2465634,
      "postDate": "2023-10-03T06:47:41.610Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2370495,
      "author_name": "i-aki.y",
      "author_url": "",
      "post_date": "2023-08-02T13:08:33.580000",
      "content": "<p>Why the <code>apply_modality_lut</code> is applied only when <code>dcm.PixelRepresentation == 1</code>?<br>\nI think this should be applied to both, otherwise, neither.</p>\n<pre><code>    if dcm.PixelRepresentation == 1:\n            bit_shift = dcm.BitsAllocated - dcm.BitsStored\n            dtype = pixel_array.dtype \n\n\n\n    return pixel_array\n</code></pre>\n<p>pixel_array</p>",
      "votes": 6,
      "replies": [
        {
          "id": 2370520,
          "author_name": "Vietnamese-NHNAM",
          "author_url": "",
          "post_date": "2023-08-02T13:23:55.530000",
          "content": "<p>i think so, too. apply_modality_lut function is used to convert pixel array to Hounsfield Unit (HU). it has nothing to do with PixelRepresentation (signed or unsigned data)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2370672,
          "author_name": "Hui Ming Lin",
          "author_url": "",
          "post_date": "2023-08-02T15:06:50.460000",
          "content": "<p>You are absolutely correct, thank you for pointing this out. I have removed apply_modality_lut from the function. The notebook has also been updated.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2373646,
      "author_name": "coderRKJ",
      "author_url": "",
      "post_date": "2023-08-04T11:09:42.310000",
      "content": "<p>I noticed a particular quirk of the dataset. DICOMs where <code>PixelRepresentation == 1</code> always returns <code>pixel_array</code> as a <code>np.int16</code> and other DICOMs always returns <code>pixel_array</code> as a <code>np.uint16</code>. Does this affect normalization in any way?</p>\n<p>Edit: By normalization, I meant the final output of the <code>standardize_pixel_array</code> function. Because the function will still return <code>np.int16</code> for <code>PixelRepresentation == 1</code> and <code>np.uint16</code> otherwise.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 2426792,
      "author_name": "Gyula Maloveczky4",
      "author_url": "",
      "post_date": "2023-09-06T20:30:07.313000",
      "content": "<p>What do you mean by the visually unusual characteristics?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2383574,
      "author_name": "martial taga",
      "author_url": "",
      "post_date": "2023-08-10T13:00:09.300000",
      "content": "<p>please how to view dataset images</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2465637,
      "author_name": "USMAN SAFDAR",
      "author_url": "",
      "post_date": "2023-10-03T06:50:21.100000",
      "content": "<p>As a beginner in such competitions, I've taken a look at multiple notebooks. I'm questioning whether this will directly influence my results or if it's more about comprehending pixel representations.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2383052,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2023-08-10T06:47:55.457000",
      "content": "<p>I don't understand how shifting bits would change the output. You left shift bits by some amount and then you right shift with the same amount which returns the same exact array. What am I missing?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2384908,
          "author_name": "i-aki.y",
          "author_url": "",
          "post_date": "2023-08-11T05:51:26.940000",
          "content": "<p>Does this example help you?</p>\n<pre><code> numpy  np\n\nBitsAllocated=\nBitsStored=\nbit_shift = BitsAllocated - BitsStored\n\nx = np.array().astype(np.int8)  \nx_new = (x &lt;&lt; bit_shift).astype(np.int8) &gt;&gt;  bit_shift\n\n(np.binary_repr(x, BitsAllocated), )\n(np.binary_repr(x_new, BitsAllocated), )\n\n\n\n</code></pre>",
          "votes": 4,
          "replies": [
            {
              "id": 2385390,
              "author_name": "Gunes Evitan",
              "author_url": "",
              "post_date": "2023-08-11T09:56:19.730000",
              "content": "<p>Thank you. I think I understood the problem. Even though those images are labeled as signed 16-bit integers (PixelRepresentation == 1), their pixel values are starting from 0. If the values are starting from 0, then we are not using some of the allocated bits.</p>\n<p><img src=\"https://i.ibb.co/9h6nLf6/Screenshot-from-2023-08-11-10-52-15.png\" alt=\"1\"></p>\n<p>When we left shift the bits, we make more bits used for a signed data type.</p>\n<p><img src=\"https://i.ibb.co/r6HXxxm/Screenshot-from-2023-08-11-10-52-35.png\" alt=\"2\"></p>\n<p>When we right shift with the same amount again, we truncate the last unused bits on the right.</p>\n<p><img src=\"https://i.ibb.co/qmGBHyj/Screenshot-from-2023-08-11-10-52-44.png\" alt=\"3\"> </p>",
              "votes": 6,
              "replies": []
            },
            {
              "id": 2388847,
              "author_name": "David Roberts",
              "author_url": "",
              "post_date": "2023-08-13T16:46:55.617000",
              "content": "<p>Need to check the BitsAllocated and BitsStored DICOM tags. Most CTs are actually 10-12 bits deep. Since there isn't a 10bit datatype, we must use 16 bit dtypes, then calculate the actual pixel width of the image.</p>",
              "votes": 4,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2465634,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-10-03T06:47:41.610000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2360696": "# Unusual DICOM Files\n\nThe RSNA 2023 Abdominal Trauma Detection Challenge showcases an exceptionally heterogeneous DICOM data structure due to the large number of institutions that donated data.\n\nCertain DICOM pixel arrays will appear unusual when opened using the pydicom library.\n\nEvery DICOM file contains a field called [Pixel Representation](https://dicom.innolitics.com/ciods/rt-dose/image-pixel/00280103) indicating if the pixel datatype is unsigned(0) or signed (1).\n\nSome DICOM files with PixelRepresentation == 1 exhibit visually unusual characteristics. You can use the [train/test]_dicom_tags.parquet files to identify images likely to be affected.\n\n\n# Potential Solution\n\nThe following function is a potential solution to the described problem.\n\n```\ndef standardize_pixel_array(dcm: pydicom.dataset.FileDataset) -> np.ndarray:\n    # Correct DICOM pixel_array if PixelRepresentation == 1.\n    pixel_array = dcm.pixel_array\n    if dcm.PixelRepresentation == 1:\n        bit_shift = dcm.BitsAllocated - dcm.BitsStored\n        dtype = pixel_array.dtype \n        pixel_array = (pixel_array << bit_shift).astype(dtype) >>  bit_shift\n    return pixel_array\n```\n\n\nPlease see this [notebook](https://www.kaggle.com/code/huiminglin/rsna-2023-abdomen-dicom-fix) for example implementation. We did not apply this processing step to the competition dataset as we did not have time to confirm that it properly resolves the issue in all cases and does not have any unintended side effects.",
    "2370495": "Why the `apply_modality_lut` is applied only when `dcm.PixelRepresentation == 1`?\nI think this should be applied to both, otherwise, neither.\n\n```diff\n    if dcm.PixelRepresentation == 1:\n            bit_shift = dcm.BitsAllocated - dcm.BitsStored\n            dtype = pixel_array.dtype \n!            pixel_array = (pixel_array << bit_shift).astype(dtype) >>  bit_shift\n-            pixel_array = pydicom.pixel_data_handlers.util.apply_modality_lut(new_array, dcm)\n+    pixel_array = pydicom.pixel_data_handlers.util.apply_modality_lut(pixel_array, dcm)\n    return pixel_array\n```\npixel_array",
    "2373646": "I noticed a particular quirk of the dataset. DICOMs where `PixelRepresentation == 1` always returns `pixel_array` as a `np.int16` and other DICOMs always returns `pixel_array` as a `np.uint16`. Does this affect normalization in any way?\n\nEdit: By normalization, I meant the final output of the `standardize_pixel_array` function. Because the function will still return `np.int16` for `PixelRepresentation == 1` and `np.uint16` otherwise.",
    "2426792": "What do you mean by the visually unusual characteristics?",
    "2383574": " please how to view dataset images\n",
    "2465637": "As a beginner in such competitions, I've taken a look at multiple notebooks. I'm questioning whether this will directly influence my results or if it's more about comprehending pixel representations.",
    "2383052": "I don't understand how shifting bits would change the output. You left shift bits by some amount and then you right shift with the same amount which returns the same exact array. What am I missing?",
    "2465634": ""
  }
}