{
  "id": 110491,
  "title": "Is it ok to use PNG images converted from DICOM files for training models?",
  "url": "/competitions/rsna-intracranial-hemorrhage-detection/discussion/110491",
  "author_name": "Manideep",
  "post_date": "2019-09-28T10:48:53.140000",
  "votes": 3,
  "comment_count": 4,
  "views": 0,
  "content": "<p>My doubt is whether it's ok to use the PNG/JPEG images converted from DICOM files (like <a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/110223#latest-635747\">this data</a>) for training the Deep Learning models. Will there be any information loss during the conversion which would probably affect the training?​​​</p>",
  "messages": [
    {
      "id": 635848,
      "postDate": "2019-09-28T10:48:53.140Z",
      "content": "<p>My doubt is whether it's ok to use the PNG/JPEG images converted from DICOM files (like <a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/110223#latest-635747\">this data</a>) for training the Deep Learning models. Will there be any information loss during the conversion which would probably affect the training?​​​</p>",
      "rawMarkdown": "My doubt is whether it's ok to use the PNG/JPEG images converted from DICOM files (like [this data](https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/110223#latest-635747)) for training the Deep Learning models. Will there be any information loss during the conversion which would probably affect the training?​​​",
      "votes": 2
    },
    {
      "id": 639805,
      "postDate": "2019-10-03T15:23:10.213Z",
      "content": "<p>I was wondering the same and I'm still not sure what to use, here are some pro/cons:\n1. I For me PNG is a lot faster than DICOM. But I also used already preprocessed pngs, so that might not be valid actually. Well and PNG uses much less space. \n2. I am not quite sure on this, but PNG's use 256 colors (ints), which already entails possible compression of the data range. Heavily depends on the windowing settings, though, with a level of 40 and length 80 you should be fine.  I'm also not sure if the min-max normalization using the image data, which I saw in a few notebooks is really helpful. It shouldn't be too much of an issue, but it is (theoretically) possible to have values in a slice that are not inside the windowing area. <br>\n3. Because of 2. you possibly have to set or use a predefined window, so you won't be able to fiddle around with it anymore. </p>\n\n<p>A different question, does anybody know whether saving preprocessed data as numpy arrays has any drawbacks? Regarding loading speed etc.? I mean, when using any kind of image loading function, the data is usually returned as np.array anyways. </p>",
      "rawMarkdown": "I was wondering the same and I'm still not sure what to use, here are some pro/cons:\n1. I For me PNG is a lot faster than DICOM. But I also used already preprocessed pngs, so that might not be valid actually. Well and PNG uses much less space. \n2. I am not quite sure on this, but PNG's use 256 colors (ints), which already entails possible compression of the data range. Heavily depends on the windowing settings, though, with a level of 40 and length 80 you should be fine.  I'm also not sure if the min-max normalization using the image data, which I saw in a few notebooks is really helpful. It shouldn't be too much of an issue, but it is (theoretically) possible to have values in a slice that are not inside the windowing area.  \n3. Because of 2. you possibly have to set or use a predefined window, so you won't be able to fiddle around with it anymore. \n\nA different question, does anybody know whether saving preprocessed data as numpy arrays has any drawbacks? Regarding loading speed etc.? I mean, when using any kind of image loading function, the data is usually returned as np.array anyways. "
    },
    {
      "id": 636185,
      "postDate": "2019-09-29T00:52:32.007Z",
      "content": "<p>png is lossless compression while jpg is lossy compression. so png should retain integrity. The author of the dataset you linked specifically used jpg for speed as he mentioned.</p>",
      "rawMarkdown": "png is lossless compression while jpg is lossy compression. so png should retain integrity. The author of the dataset you linked specifically used jpg for speed as he mentioned."
    },
    {
      "id": 636183,
      "postDate": "2019-09-29T00:31:46.367Z",
      "content": "<p>I think someone should evaluate both bit depths and share the results! As for me, I'm going with the full depth. It is unclear how PNG would be more useful than DCM, given the wealth of libraries that handle it.</p>",
      "rawMarkdown": "I think someone should evaluate both bit depths and share the results! As for me, I'm going with the full depth. It is unclear how PNG would be more useful than DCM, given the wealth of libraries that handle it."
    },
    {
      "id": 635871,
      "postDate": "2019-09-28T11:40:14.453Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 639805,
      "author_name": "srs",
      "author_url": "",
      "post_date": "2019-10-03T15:23:10.213000",
      "content": "<p>I was wondering the same and I'm still not sure what to use, here are some pro/cons:\n1. I For me PNG is a lot faster than DICOM. But I also used already preprocessed pngs, so that might not be valid actually. Well and PNG uses much less space. \n2. I am not quite sure on this, but PNG's use 256 colors (ints), which already entails possible compression of the data range. Heavily depends on the windowing settings, though, with a level of 40 and length 80 you should be fine.  I'm also not sure if the min-max normalization using the image data, which I saw in a few notebooks is really helpful. It shouldn't be too much of an issue, but it is (theoretically) possible to have values in a slice that are not inside the windowing area. <br>\n3. Because of 2. you possibly have to set or use a predefined window, so you won't be able to fiddle around with it anymore. </p>\n\n<p>A different question, does anybody know whether saving preprocessed data as numpy arrays has any drawbacks? Regarding loading speed etc.? I mean, when using any kind of image loading function, the data is usually returned as np.array anyways. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 636185,
      "author_name": "Tim Yee",
      "author_url": "",
      "post_date": "2019-09-29T00:52:32.007000",
      "content": "<p>png is lossless compression while jpg is lossy compression. so png should retain integrity. The author of the dataset you linked specifically used jpg for speed as he mentioned.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 636183,
      "author_name": "Tom H.",
      "author_url": "",
      "post_date": "2019-09-29T00:31:46.367000",
      "content": "<p>I think someone should evaluate both bit depths and share the results! As for me, I'm going with the full depth. It is unclear how PNG would be more useful than DCM, given the wealth of libraries that handle it.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 635871,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-28T11:40:14.453000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "635848": "My doubt is whether it's ok to use the PNG/JPEG images converted from DICOM files (like [this data](https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/110223#latest-635747)) for training the Deep Learning models. Will there be any information loss during the conversion which would probably affect the training?​​​",
    "639805": "I was wondering the same and I'm still not sure what to use, here are some pro/cons:\n1. I For me PNG is a lot faster than DICOM. But I also used already preprocessed pngs, so that might not be valid actually. Well and PNG uses much less space. \n2. I am not quite sure on this, but PNG's use 256 colors (ints), which already entails possible compression of the data range. Heavily depends on the windowing settings, though, with a level of 40 and length 80 you should be fine.  I'm also not sure if the min-max normalization using the image data, which I saw in a few notebooks is really helpful. It shouldn't be too much of an issue, but it is (theoretically) possible to have values in a slice that are not inside the windowing area.  \n3. Because of 2. you possibly have to set or use a predefined window, so you won't be able to fiddle around with it anymore. \n\nA different question, does anybody know whether saving preprocessed data as numpy arrays has any drawbacks? Regarding loading speed etc.? I mean, when using any kind of image loading function, the data is usually returned as np.array anyways. ",
    "636185": "png is lossless compression while jpg is lossy compression. so png should retain integrity. The author of the dataset you linked specifically used jpg for speed as he mentioned.",
    "636183": "I think someone should evaluate both bit depths and share the results! As for me, I'm going with the full depth. It is unclear how PNG would be more useful than DCM, given the wealth of libraries that handle it.",
    "635871": ""
  }
}