{
  "id": 358659,
  "title": "DICOM in TensorFlow",
  "url": "/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/358659",
  "author_name": "",
  "post_date": "2022-10-08T19:17:26",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I'm trying to create a Tensorflow Dataset that loads the DICOM images one batch at a time, but I'm running into issues. I started by loading (in the dataset pipeline) the images using tfio.image.decode_dicom_image(path), but every time an image is loaded it prints \"W: DcmMetaInfo: No Group Length available in Meta Information Header\", which makes training difficult. </p>\n<p>I then tried pydicom.dcmread(path).pixel_array in the pipeline, but that runs into an error, as the dataset pipeline's graph turns the path into a tensor of dtype string (as opposed to just a string) that the dcmread function cannot accept.</p>\n<p>Help with either of these two issues would be appreciated.</p>",
  "messages": [
    {
      "id": 1980332,
      "postDate": "2022-10-10T04:59:49.147Z",
      "content": "<p>I also get the same warning message when loading these DCM files using tensorflow io. I couldn't figure a way to silence those messages. Here is a code snippet that I use to load DCM files</p>\n<pre><code>def load_dicom(path):\n    img=dicom.dcmread(path)\n    img.PhotometricInterpretation = 'YBR_FULL'\n    data = img.pixel_array    \n    data = data - np.min(data)\n    if np.max(data) != 0:\n        data = data / np.max(data)\n    data = data.astype(np.float32)\n    return data\n\ndef px_array(filename):\n    return tf.convert_to_tensor(load_dicom(filename.numpy().decode(\"utf-8\") ))\n\ndef read_dcm_tf(filename):\n    img = tf.py_function(px_array, inp=[filename], Tout=tf.float32)   \n    img = tf.stack([img,img,img], axis = 2)\n    img = tf.image.resize(img, [IMAGE_SIZE, IMAGE_SIZE])\n    return img\n\ndef process_image(filename, target=None):\n    image = read_dcm_tf(filename)\n    if (target is not None):\n        return image, target\n    else:\n        return image\n\nds1 = (tf.data.Dataset.from_tensor_slices((your_list_of_dcm_files))\n       .map(process_image, num_parallel_calls=tf.data.experimental.AUTOTUNE)\n       .cache()\n       .batch(16)\n       .prefetch(tf.data.experimental.AUTOTUNE))\n\ndef display_batch(batch, size=2):\n    imgs = batch\n    for img_idx in range(size):\n        plt.imshow(imgs[img_idx,...].numpy())\n        plt.show()\n\n\nds = ds1.unbatch().batch(4)\nbatch = next(iter(ds1))\ndisplay_batch(batch, 2)\n</code></pre>",
      "rawMarkdown": "I also get the same warning message when loading these DCM files using tensorflow io. I couldn't figure a way to silence those messages. Here is a code snippet that I use to load DCM files\n\n```\ndef load_dicom(path):\n    img=dicom.dcmread(path)\n    img.PhotometricInterpretation = 'YBR_FULL'\n    data = img.pixel_array    \n    data = data - np.min(data)\n    if np.max(data) != 0:\n        data = data / np.max(data)\n    data = data.astype(np.float32)\n    return data\n\ndef px_array(filename):\n    return tf.convert_to_tensor(load_dicom(filename.numpy().decode(\"utf-8\") ))\n\ndef read_dcm_tf(filename):\n    img = tf.py_function(px_array, inp=[filename], Tout=tf.float32)   \n    img = tf.stack([img,img,img], axis = 2)\n    img = tf.image.resize(img, [IMAGE_SIZE, IMAGE_SIZE])\n    return img\n\ndef process_image(filename, target=None):\n    image = read_dcm_tf(filename)\n    if (target is not None):\n        return image, target\n    else:\n        return image\n\t\t\nds1 = (tf.data.Dataset.from_tensor_slices((your_list_of_dcm_files))\n       .map(process_image, num_parallel_calls=tf.data.experimental.AUTOTUNE)\n       .cache()\n       .batch(16)\n       .prefetch(tf.data.experimental.AUTOTUNE))\n\ndef display_batch(batch, size=2):\n    imgs = batch\n    for img_idx in range(size):\n        plt.imshow(imgs[img_idx,...].numpy())\n        plt.show()\n        \n\nds = ds1.unbatch().batch(4)\nbatch = next(iter(ds1))\ndisplay_batch(batch, 2)\n```",
      "votes": 3
    },
    {
      "id": 1978575,
      "postDate": "2022-10-08T19:17:26Z",
      "content": "<p>I'm trying to create a Tensorflow Dataset that loads the DICOM images one batch at a time, but I'm running into issues. I started by loading (in the dataset pipeline) the images using tfio.image.decode_dicom_image(path), but every time an image is loaded it prints \"W: DcmMetaInfo: No Group Length available in Meta Information Header\", which makes training difficult. </p>\n<p>I then tried pydicom.dcmread(path).pixel_array in the pipeline, but that runs into an error, as the dataset pipeline's graph turns the path into a tensor of dtype string (as opposed to just a string) that the dcmread function cannot accept.</p>\n<p>Help with either of these two issues would be appreciated.</p>",
      "rawMarkdown": "I'm trying to create a Tensorflow Dataset that loads the DICOM images one batch at a time, but I'm running into issues. I started by loading (in the dataset pipeline) the images using tfio.image.decode_dicom_image(path), but every time an image is loaded it prints \"W: DcmMetaInfo: No Group Length available in Meta Information Header\", which makes training difficult. \n\nI then tried pydicom.dcmread(path).pixel_array in the pipeline, but that runs into an error, as the dataset pipeline's graph turns the path into a tensor of dtype string (as opposed to just a string) that the dcmread function cannot accept.\n\nHelp with either of these two issues would be appreciated.",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 1980332,
      "author_name": "Parag",
      "author_url": "",
      "post_date": "2022-10-10T04:59:49.147000",
      "content": "<p>I also get the same warning message when loading these DCM files using tensorflow io. I couldn't figure a way to silence those messages. Here is a code snippet that I use to load DCM files</p>\n<pre><code>def load_dicom(path):\n    img=dicom.dcmread(path)\n    img.PhotometricInterpretation = 'YBR_FULL'\n    data = img.pixel_array    \n    data = data - np.min(data)\n    if np.max(data) != 0:\n        data = data / np.max(data)\n    data = data.astype(np.float32)\n    return data\n\ndef px_array(filename):\n    return tf.convert_to_tensor(load_dicom(filename.numpy().decode(\"utf-8\") ))\n\ndef read_dcm_tf(filename):\n    img = tf.py_function(px_array, inp=[filename], Tout=tf.float32)   \n    img = tf.stack([img,img,img], axis = 2)\n    img = tf.image.resize(img, [IMAGE_SIZE, IMAGE_SIZE])\n    return img\n\ndef process_image(filename, target=None):\n    image = read_dcm_tf(filename)\n    if (target is not None):\n        return image, target\n    else:\n        return image\n\nds1 = (tf.data.Dataset.from_tensor_slices((your_list_of_dcm_files))\n       .map(process_image, num_parallel_calls=tf.data.experimental.AUTOTUNE)\n       .cache()\n       .batch(16)\n       .prefetch(tf.data.experimental.AUTOTUNE))\n\ndef display_batch(batch, size=2):\n    imgs = batch\n    for img_idx in range(size):\n        plt.imshow(imgs[img_idx,...].numpy())\n        plt.show()\n\n\nds = ds1.unbatch().batch(4)\nbatch = next(iter(ds1))\ndisplay_batch(batch, 2)\n</code></pre>",
      "votes": 3,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1980332": "I also get the same warning message when loading these DCM files using tensorflow io. I couldn't figure a way to silence those messages. Here is a code snippet that I use to load DCM files\n\n```\ndef load_dicom(path):\n    img=dicom.dcmread(path)\n    img.PhotometricInterpretation = 'YBR_FULL'\n    data = img.pixel_array    \n    data = data - np.min(data)\n    if np.max(data) != 0:\n        data = data / np.max(data)\n    data = data.astype(np.float32)\n    return data\n\ndef px_array(filename):\n    return tf.convert_to_tensor(load_dicom(filename.numpy().decode(\"utf-8\") ))\n\ndef read_dcm_tf(filename):\n    img = tf.py_function(px_array, inp=[filename], Tout=tf.float32)   \n    img = tf.stack([img,img,img], axis = 2)\n    img = tf.image.resize(img, [IMAGE_SIZE, IMAGE_SIZE])\n    return img\n\ndef process_image(filename, target=None):\n    image = read_dcm_tf(filename)\n    if (target is not None):\n        return image, target\n    else:\n        return image\n\t\t\nds1 = (tf.data.Dataset.from_tensor_slices((your_list_of_dcm_files))\n       .map(process_image, num_parallel_calls=tf.data.experimental.AUTOTUNE)\n       .cache()\n       .batch(16)\n       .prefetch(tf.data.experimental.AUTOTUNE))\n\ndef display_batch(batch, size=2):\n    imgs = batch\n    for img_idx in range(size):\n        plt.imshow(imgs[img_idx,...].numpy())\n        plt.show()\n        \n\nds = ds1.unbatch().batch(4)\nbatch = next(iter(ds1))\ndisplay_batch(batch, 2)\n```",
    "1978575": "I'm trying to create a Tensorflow Dataset that loads the DICOM images one batch at a time, but I'm running into issues. I started by loading (in the dataset pipeline) the images using tfio.image.decode_dicom_image(path), but every time an image is loaded it prints \"W: DcmMetaInfo: No Group Length available in Meta Information Header\", which makes training difficult. \n\nI then tried pydicom.dcmread(path).pixel_array in the pipeline, but that runs into an error, as the dataset pipeline's graph turns the path into a tensor of dtype string (as opposed to just a string) that the dcmread function cannot accept.\n\nHelp with either of these two issues would be appreciated."
  }
}