{
  "id": 113692,
  "title": "Creating a dataset made of pixels from dicom images",
  "url": "/competitions/rsna-intracranial-hemorrhage-detection/discussion/113692",
  "author_name": "Patrick",
  "post_date": "2019-10-21T14:13:55.169000",
  "votes": 1,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hello to everyone. Any hint on how to quickly create a dataset made of pixels from dicom images? For me it is taking many days(72 days) to create the dataset when my RAM is set to 64GB and processor is intel core i7:. Did you compress images? If so to how many pixels are you compressing images? </p>",
  "messages": [
    {
      "id": 654148,
      "postDate": "2019-10-21T14:13:55.170Z",
      "content": "<p>Hello to everyone. Any hint on how to quickly create a dataset made of pixels from dicom images? For me it is taking many days(72 days) to create the dataset when my RAM is set to 64GB and processor is intel core i7:. Did you compress images? If so to how many pixels are you compressing images? </p>",
      "rawMarkdown": "Hello to everyone. Any hint on how to quickly create a dataset made of pixels from dicom images? For me it is taking many days(72 days) to create the dataset when my RAM is set to 64GB and processor is intel core i7:. Did you compress images? If so to how many pixels are you compressing images? ",
      "votes": 1
    },
    {
      "id": 655736,
      "postDate": "2019-10-23T12:52:45.410Z",
      "content": "<p>Another question: I am new to dicom image processing. When I extract image pixels using image_array = ds.pixel_array, \nI get some negative values and even positive values that are not in the range of [0,255]. How do you deal with these pixels values? I f any good complete material I would read, please I welcome it.</p>",
      "rawMarkdown": "Another question: I am new to dicom image processing. When I extract image pixels using image_array = ds.pixel_array, \nI get some negative values and even positive values that are not in the range of [0,255]. How do you deal with these pixels values? I f any good complete material I would read, please I welcome it."
    },
    {
      "id": 655723,
      "postDate": "2019-10-23T12:40:28.600Z",
      "content": "<p>The following is my simple logic: I am extracting image pixels as a 2d dimensional matrix (image_array = ds.pixel_array), convert this matrix into a line(row) matrix, then pack or append these row matrices   to finally build X as a train set. Hence X will be  a matrix of shape(n,m) where n is the number of images in training set and m=image_width*image_height. </p>",
      "rawMarkdown": "The following is my simple logic: I am extracting image pixels as a 2d dimensional matrix (image_array = ds.pixel_array), convert this matrix into a line(row) matrix, then pack or append these row matrices   to finally build X as a train set. Hence X will be  a matrix of shape(n,m) where n is the number of images in training set and m=image_width*image_height. "
    },
    {
      "id": 654458,
      "postDate": "2019-10-21T22:11:57.077Z",
      "content": "<p>What do you mean by dataset? Do you want extract images from dicom files? You can read images directly in your data loader \n<code>\ndef load_dicom(path):\n        img = pydicom.read_file(path).pixel_array.astype('float32')\n        ...\n        return img\n</code>\nor preprocessusing multiprocessing as here :  <a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/110223#latest-651299\">https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/110223#latest-651299</a>\nAlso it will take much more time if your drives are HDD</p>",
      "rawMarkdown": "What do you mean by dataset? Do you want extract images from dicom files? You can read images directly in your data loader \n```\ndef load_dicom(path):\n        img = pydicom.read_file(path).pixel_array.astype('float32')\n        ...\n        return img\n```\nor preprocessusing multiprocessing as here :  https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/110223#latest-651299\nAlso it will take much more time if your drives are HDD\n",
      "replies": [
        {
          "id": 655737,
          "postDate": "2019-10-23T12:53:12.663Z",
          "content": "<p>The following is my simple logic: I am extracting image pixels as a 2d dimensional matrix (imagearray = ds.pixelarray), convert this matrix into a line(row) matrix, then pack or append these row matrices to finally build X as a train set. Hence X will be a matrix of shape(n,m) where n is the number of images in training set and m=imagewidth*imageheight.</p>",
          "rawMarkdown": "The following is my simple logic: I am extracting image pixels as a 2d dimensional matrix (imagearray = ds.pixelarray), convert this matrix into a line(row) matrix, then pack or append these row matrices to finally build X as a train set. Hence X will be a matrix of shape(n,m) where n is the number of images in training set and m=imagewidth*imageheight."
        }
      ]
    },
    {
      "id": 654427,
      "postDate": "2019-10-21T21:13:58.033Z",
      "content": "<p>Are you storing your pixel arrays in memory all at the same time? Because that would seem to be the only reason you're running out of memory with 64GB RAM. When you extract the <code>img = dcm.pixel_array</code>, you follow it up by writing to a file immediately <code>cv2.imwrite(filename, img)</code>. Every image you reuse the img variable and never append the img to a running list of pixels. I am able to extract the full train/test in around 3hours without multiprocessing reading and writing to the same ssd.</p>",
      "rawMarkdown": "Are you storing your pixel arrays in memory all at the same time? Because that would seem to be the only reason you're running out of memory with 64GB RAM. When you extract the `img = dcm.pixel_array`, you follow it up by writing to a file immediately `cv2.imwrite(filename, img)`. Every image you reuse the img variable and never append the img to a running list of pixels. I am able to extract the full train/test in around 3hours without multiprocessing reading and writing to the same ssd.",
      "replies": [
        {
          "id": 655739,
          "postDate": "2019-10-23T12:53:26.613Z",
          "content": "<p>The following is my simple logic: I am extracting image pixels as a 2d dimensional matrix (imagearray = ds.pixelarray), convert this matrix into a line(row) matrix, then pack or append these row matrices to finally build X as a train set. Hence X will be a matrix of shape(n,m) where n is the number of images in training set and m=imagewidth*imageheight.</p>",
          "rawMarkdown": "The following is my simple logic: I am extracting image pixels as a 2d dimensional matrix (imagearray = ds.pixelarray), convert this matrix into a line(row) matrix, then pack or append these row matrices to finally build X as a train set. Hence X will be a matrix of shape(n,m) where n is the number of images in training set and m=imagewidth*imageheight."
        },
        {
          "id": 655870,
          "postDate": "2019-10-23T15:44:13.917Z",
          "content": "<p>Yes, I figured you were appending pixel_array data and building X as training set IN MEMORY. That is why you ran out of RAM. You cannot store training X in memory. Here's the simple math - 512 x 512  * 2bits * ~500,000 images = 262,144,000,000 bits. Think about how you can store that in 64GB memory. The work around to that is loading image files from persistent disk using a dataloader for pytorch or imagegenerator with keras.</p>",
          "rawMarkdown": "Yes, I figured you were appending pixel_array data and building X as training set IN MEMORY. That is why you ran out of RAM. You cannot store training X in memory. Here's the simple math - 512 x 512  * 2bits * ~500,000 images = 262,144,000,000 bits. Think about how you can store that in 64GB memory. The work around to that is loading image files from persistent disk using a dataloader for pytorch or imagegenerator with keras.",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 655736,
      "author_name": "Patrick",
      "author_url": "",
      "post_date": "2019-10-23T12:52:45.410000",
      "content": "<p>Another question: I am new to dicom image processing. When I extract image pixels using image_array = ds.pixel_array, \nI get some negative values and even positive values that are not in the range of [0,255]. How do you deal with these pixels values? I f any good complete material I would read, please I welcome it.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 655723,
      "author_name": "Patrick",
      "author_url": "",
      "post_date": "2019-10-23T12:40:28.600000",
      "content": "<p>The following is my simple logic: I am extracting image pixels as a 2d dimensional matrix (image_array = ds.pixel_array), convert this matrix into a line(row) matrix, then pack or append these row matrices   to finally build X as a train set. Hence X will be  a matrix of shape(n,m) where n is the number of images in training set and m=image_width*image_height. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 654458,
      "author_name": "Oleg Yaroshevskiy",
      "author_url": "",
      "post_date": "2019-10-21T22:11:57.077000",
      "content": "<p>What do you mean by dataset? Do you want extract images from dicom files? You can read images directly in your data loader \n<code>\ndef load_dicom(path):\n        img = pydicom.read_file(path).pixel_array.astype('float32')\n        ...\n        return img\n</code>\nor preprocessusing multiprocessing as here :  <a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/110223#latest-651299\">https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/110223#latest-651299</a>\nAlso it will take much more time if your drives are HDD</p>",
      "votes": 0,
      "replies": [
        {
          "id": 655737,
          "author_name": "Patrick",
          "author_url": "",
          "post_date": "2019-10-23T12:53:12.663000",
          "content": "<p>The following is my simple logic: I am extracting image pixels as a 2d dimensional matrix (imagearray = ds.pixelarray), convert this matrix into a line(row) matrix, then pack or append these row matrices to finally build X as a train set. Hence X will be a matrix of shape(n,m) where n is the number of images in training set and m=imagewidth*imageheight.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 654427,
      "author_name": "Tim Yee",
      "author_url": "",
      "post_date": "2019-10-21T21:13:58.033000",
      "content": "<p>Are you storing your pixel arrays in memory all at the same time? Because that would seem to be the only reason you're running out of memory with 64GB RAM. When you extract the <code>img = dcm.pixel_array</code>, you follow it up by writing to a file immediately <code>cv2.imwrite(filename, img)</code>. Every image you reuse the img variable and never append the img to a running list of pixels. I am able to extract the full train/test in around 3hours without multiprocessing reading and writing to the same ssd.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 655739,
          "author_name": "Patrick",
          "author_url": "",
          "post_date": "2019-10-23T12:53:26.613000",
          "content": "<p>The following is my simple logic: I am extracting image pixels as a 2d dimensional matrix (imagearray = ds.pixelarray), convert this matrix into a line(row) matrix, then pack or append these row matrices to finally build X as a train set. Hence X will be a matrix of shape(n,m) where n is the number of images in training set and m=imagewidth*imageheight.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 655870,
          "author_name": "Tim Yee",
          "author_url": "",
          "post_date": "2019-10-23T15:44:13.917000",
          "content": "<p>Yes, I figured you were appending pixel_array data and building X as training set IN MEMORY. That is why you ran out of RAM. You cannot store training X in memory. Here's the simple math - 512 x 512  * 2bits * ~500,000 images = 262,144,000,000 bits. Think about how you can store that in 64GB memory. The work around to that is loading image files from persistent disk using a dataloader for pytorch or imagegenerator with keras.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "654148": "Hello to everyone. Any hint on how to quickly create a dataset made of pixels from dicom images? For me it is taking many days(72 days) to create the dataset when my RAM is set to 64GB and processor is intel core i7:. Did you compress images? If so to how many pixels are you compressing images? ",
    "655736": "Another question: I am new to dicom image processing. When I extract image pixels using image_array = ds.pixel_array, \nI get some negative values and even positive values that are not in the range of [0,255]. How do you deal with these pixels values? I f any good complete material I would read, please I welcome it.",
    "655723": "The following is my simple logic: I am extracting image pixels as a 2d dimensional matrix (image_array = ds.pixel_array), convert this matrix into a line(row) matrix, then pack or append these row matrices   to finally build X as a train set. Hence X will be  a matrix of shape(n,m) where n is the number of images in training set and m=image_width*image_height. ",
    "654458": "What do you mean by dataset? Do you want extract images from dicom files? You can read images directly in your data loader \n```\ndef load_dicom(path):\n        img = pydicom.read_file(path).pixel_array.astype('float32')\n        ...\n        return img\n```\nor preprocessusing multiprocessing as here :  https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/110223#latest-651299\nAlso it will take much more time if your drives are HDD\n",
    "654427": "Are you storing your pixel arrays in memory all at the same time? Because that would seem to be the only reason you're running out of memory with 64GB RAM. When you extract the `img = dcm.pixel_array`, you follow it up by writing to a file immediately `cv2.imwrite(filename, img)`. Every image you reuse the img variable and never append the img to a running list of pixels. I am able to extract the full train/test in around 3hours without multiprocessing reading and writing to the same ssd."
  }
}