{
  "id": 487228,
  "title": "Data modification description",
  "url": "/competitions/spr-head-ct-age-prediction-challenge/discussion/487228",
  "author_name": "neuroguide",
  "post_date": "2024-03-28T07:55:30.159000",
  "votes": 2,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi there, <br>\nI'm looking at the training data and have a few questions:</p>\n<ol>\n<li>Can you list the modifications of the raw CT data in the dataset?</li>\n<li>Why are the dicom tags removed so that the most common dicom to nifti converters (itk and dcm2niix) do not recognize data as dicom? (I understand anonymization, but have a problem with spatial reconstruction)</li>\n<li>What are the parameters of the noise in the images?</li>\n<li>I guess that you have computed a mask for the head region of interest to modify background noise. Are you going to share it?</li>\n<li>Did the mask affect the skull bones in any way? (Bone changes in the frontal bones and temporal bones are associated with age)</li>\n<li>Can you provide the data in nifti format? if not can you provide the range of resolutions and acquisition protocols for the datasets</li>\n<li>Do you think that the results will be generelizable without knowing the above?</li>\n</ol>",
  "messages": [
    {
      "id": 2720742,
      "postDate": "2024-03-28T14:39:38.300Z",
      "content": "<p>Hi, thank you for your great questions!</p>\n<p>Regarding the DICOM tags, you are correct. </p>\n<p>We are missing some DICOM tags to convert to nifti. I suspect the missing tag is Spacing Between Slices, because the other ones related to spacial resolution are present. </p>\n<p>I suspect the missing tag will make training 3D models more challenging, but not as much for 2D models. The LB results are good so far, but I suspect they could be even better if we had the missing tag. </p>\n<p>I will try to get access to the original files to read that tag. Let me know if you miss any other tag.</p>\n<p>Regarding the transformation in the pixel data, the complete code for facial deidentification is here: <a href=\"https://github.com/kitamura-felipe/face_deid_ct\" target=\"_blank\">https://github.com/kitamura-felipe/face_deid_ct</a></p>\n<p>In a nutshell, the code fills the air with pixels that have values in the range of those of the patient's skin. By doing that, we reduce the chance that the process can be perfectly reverted. In some cases, it removed part of the skull as well. That is a limitation of this method.</p>\n<p>We believe that models trained with images using that defacing technique will generalize if we apply the same function to newer studies on which we want to do inference. The results of the LB show that is true so far.</p>\n<p>Thank you for bringing that up!</p>",
      "rawMarkdown": "Hi, thank you for your great questions!\n\nRegarding the DICOM tags, you are correct. \n\nWe are missing some DICOM tags to convert to nifti. I suspect the missing tag is Spacing Between Slices, because the other ones related to spacial resolution are present. \n\nI suspect the missing tag will make training 3D models more challenging, but not as much for 2D models. The LB results are good so far, but I suspect they could be even better if we had the missing tag. \n\nI will try to get access to the original files to read that tag. Let me know if you miss any other tag.\n\nRegarding the transformation in the pixel data, the complete code for facial deidentification is here: https://github.com/kitamura-felipe/face_deid_ct\n\nIn a nutshell, the code fills the air with pixels that have values in the range of those of the patient's skin. By doing that, we reduce the chance that the process can be perfectly reverted. In some cases, it removed part of the skull as well. That is a limitation of this method.\n\nWe believe that models trained with images using that defacing technique will generalize if we apply the same function to newer studies on which we want to do inference. The results of the LB show that is true so far.\n\nThank you for bringing that up!",
      "votes": 1,
      "replies": [
        {
          "id": 2720785,
          "postDate": "2024-03-28T15:16:16.743Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 2722030,
          "postDate": "2024-03-29T11:17:19.897Z",
          "content": "<p>Only some of the studies had the tag SpaceBetweenSlices. I am uploading two CSV files with that information for the training and test sets. It should be available in 4 hours.</p>",
          "rawMarkdown": "Only some of the studies had the tag SpaceBetweenSlices. I am uploading two CSV files with that information for the training and test sets. It should be available in 4 hours.",
          "votes": 2
        },
        {
          "id": 2722045,
          "postDate": "2024-03-29T11:25:15.430Z",
          "content": "<p>we can create 3d array from dicoms .</p>\n<pre><code> ():\n    \n    img = (img*slope +intercept) \n    img_min = window_center - window_width// \n    img_max = window_center + window_width// \n    img[img&lt;img_min] = img_min \n    img[img&gt;img_max] = img_max \n     rescale: \n        img = (img - img_min) / (img_max - img_min)* \n     img\n\n ():\n    \n    image = image - np.(image)\n     image / np.(image)\n ():\n        \n    :\n        image = dcm.pixel_array\n    :\n        \n     window_center     window_width   :\n\n        image = window_image(image, window_center, window_width, dcm.RescaleIntercept, dcm.RescaleSlope)\n\n     dcm.PhotometricInterpretation == :\n        image = np.invert(image)\n\n    normalized_image = normalize_image(image)\n     normalized_image\n ():\n\n    t_paths = (glob.glob(os.path.join(path, ))[:], key= x: (x.split()[-].split()[]))\n\n\n    dicoms = [pydicom.dcmread(d, force= )  d  t_paths]\n    sereis_instance_uid = [d.SeriesInstanceUID  d  dicoms]\n    ((sereis_instance_uid))\n\n    images = {i:[]  i  (sereis_instance_uid)}\n    z_pos = {i:[]  i  (sereis_instance_uid)}\n     dcm,s_id  (dicoms,sereis_instance_uid):\n        d = load_dicom(dcm,,)\n         d  :\n            \n        images[s_id].append(d)\n        z_pos[s_id].append((dcm.ImagePositionPatient[-]))\n\n    \n     s_id  (sereis_instance_uid):\n        images[s_id] = np.stack(images[s_id], -)[:,:,np.argsort(z_pos[s_id])]\n        z_pos[s_id] = np.sort(z_pos[s_id])\n     images,z_pos\n</code></pre>",
          "rawMarkdown": "we can create 3d array from dicoms .\n\n```\ndef window_image(img, window_center,window_width, intercept, slope, rescale=True):\n    '''\n    This fucntion came from this notebook https://www.kaggle.com/code/redwankarimsony/ct-scans-dicom-files-windowing-explained\n    If you want to understand more about windowing the referenced notebook is a good read.\n    '''\n    img = (img*slope +intercept) #for translation adjustments given in the dicom file. \n    img_min = window_center - window_width//2 #minimum HU level\n    img_max = window_center + window_width//2 #maximum HU level\n    img[img<img_min] = img_min #set img_min for all HU levels less than minimum HU level\n    img[img>img_max] = img_max #set img_max for all HU levels higher than maximum HU level\n    if rescale: \n        img = (img - img_min) / (img_max - img_min)*255.0 \n    return img\n\ndef normalize_image(image):\n    \"\"\"\n    Normalize image to the range [0, 1].\n    \"\"\"\n    image = image - np.min(image)\n    return image / np.max(image)\ndef load_dicom(dcm, window_center=None, window_width=None):\n    \"\"\"\n    Process a DICOM file and save it as a PNG file.\n    \"\"\"    \n    try:\n        image = dcm.pixel_array\n    except:\n        return\n    if window_center is not None and window_width is not None:\n        \n        image = window_image(image, window_center, window_width, dcm.RescaleIntercept, dcm.RescaleSlope)\n\n    if dcm.PhotometricInterpretation == \"MONOCHROME1\":\n        image = np.invert(image)\n    \n    normalized_image = normalize_image(image)\n    return normalized_image\ndef load_dicom_line_par(path):\n\n    t_paths = sorted(glob.glob(os.path.join(path, \"*\"))[1:], key=lambda x: int(x.split('/')[-1].split(\".\")[0]))\n\n\n    dicoms = [pydicom.dcmread(d, force= True) for d in t_paths]\n    sereis_instance_uid = [d.SeriesInstanceUID for d in dicoms]\n    print(set(sereis_instance_uid))\n        \n    images = {i:[] for i in set(sereis_instance_uid)}\n    z_pos = {i:[] for i in set(sereis_instance_uid)}\n    for dcm,s_id in zip(dicoms,sereis_instance_uid):\n        d = load_dicom(dcm,40,80)\n        if d is None:\n            continue\n        images[s_id].append(d)\n        z_pos[s_id].append(float(dcm.ImagePositionPatient[-1]))\n        \n    #z_pos = [float(d.ImagePositionPatient[-1]) for d in dicoms]\n    for s_id in set(sereis_instance_uid):\n        images[s_id] = np.stack(images[s_id], -1)[:,:,np.argsort(z_pos[s_id])]\n        z_pos[s_id] = np.sort(z_pos[s_id])\n    return images,z_pos\n\n```\n",
          "votes": 3,
          "replies": [
            {
              "id": 2723431,
              "postDate": "2024-03-30T07:57:22.397Z",
              "content": "<p>Hi Felipe, <br>\nThank you for help and explanation. I have explored less common converters and the Fiji is recognizing a lot of the sets. here is an example:<br>\n(Fiji Is Just) ImageJ 2.14.0/1.54f; Java 1.8.0_202 [64-bit]; </p>\n<p>Title: 002360<br>\nWidth:  249.9999 mm (512)<br>\nHeight:  249.9999 mm (512)<br>\nDepth:  173.75 mm (139)<br>\nSize:  70MB<br>\nResolution:  2.0480 pixels per mm<br>\nVoxel size: 0.4883x0.4883x1.25 mm^3<br>\nID: -212<br>\nBits per pixel: 16 (unsigned, grayscale LUT)<br>\nDisplay range: -15 - 95<br>\nRaw pixel value range: 33069 - 35553<br>\nImage: 82/139 (1.2.840.12345.4549484626668932613792536314089807458768514401)<br>\nNo threshold<br>\nScaleToFit: false</p>\n<p>So the voxel size can be read, but the typical location for shape definition has been erased like you mentioned.<br>\nIs the gender encoded in the dicom real or false?</p>",
              "rawMarkdown": "Hi Felipe, \nThank you for help and explanation. I have explored less common converters and the Fiji is recognizing a lot of the sets. here is an example:\n(Fiji Is Just) ImageJ 2.14.0/1.54f; Java 1.8.0_202 [64-bit]; \n \nTitle: 002360\nWidth:  249.9999 mm (512)\nHeight:  249.9999 mm (512)\nDepth:  173.75 mm (139)\nSize:  70MB\nResolution:  2.0480 pixels per mm\nVoxel size: 0.4883x0.4883x1.25 mm^3\nID: -212\nBits per pixel: 16 (unsigned, grayscale LUT)\nDisplay range: -15 - 95\nRaw pixel value range: 33069 - 35553\nImage: 82/139 (1.2.840.12345.4549484626668932613792536314089807458768514401)\nNo threshold\nScaleToFit: false\n\nSo the voxel size can be read, but the typical location for shape definition has been erased like you mentioned.\nIs the gender encoded in the dicom real or false?\n\n\n"
            },
            {
              "id": 2726054,
              "postDate": "2024-03-31T23:57:56.103Z",
              "content": "<p>Great to hear you are making progress.</p>\n<p>The gender is real and can be used as independent variable.</p>",
              "rawMarkdown": "Great to hear you are making progress.\n\nThe gender is real and can be used as independent variable."
            }
          ]
        }
      ]
    },
    {
      "id": 2720280,
      "postDate": "2024-03-28T07:55:30.160Z",
      "content": "<p>Hi there, <br>\nI'm looking at the training data and have a few questions:</p>\n<ol>\n<li>Can you list the modifications of the raw CT data in the dataset?</li>\n<li>Why are the dicom tags removed so that the most common dicom to nifti converters (itk and dcm2niix) do not recognize data as dicom? (I understand anonymization, but have a problem with spatial reconstruction)</li>\n<li>What are the parameters of the noise in the images?</li>\n<li>I guess that you have computed a mask for the head region of interest to modify background noise. Are you going to share it?</li>\n<li>Did the mask affect the skull bones in any way? (Bone changes in the frontal bones and temporal bones are associated with age)</li>\n<li>Can you provide the data in nifti format? if not can you provide the range of resolutions and acquisition protocols for the datasets</li>\n<li>Do you think that the results will be generelizable without knowing the above?</li>\n</ol>",
      "rawMarkdown": "Hi there, \nI'm looking at the training data and have a few questions:\n1. Can you list the modifications of the raw CT data in the dataset?\n2. Why are the dicom tags removed so that the most common dicom to nifti converters (itk and dcm2niix) do not recognize data as dicom? (I understand anonymization, but have a problem with spatial reconstruction)\n3. What are the parameters of the noise in the images?\n4. I guess that you have computed a mask for the head region of interest to modify background noise. Are you going to share it?\n5. Did the mask affect the skull bones in any way? (Bone changes in the frontal bones and temporal bones are associated with age)\n6. Can you provide the data in nifti format? if not can you provide the range of resolutions and acquisition protocols for the datasets\n7. Do you think that the results will be generelizable without knowing the above?\n",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 2720742,
      "author_name": "FelipeKitamura, MD, PhD",
      "author_url": "",
      "post_date": "2024-03-28T14:39:38.300000",
      "content": "<p>Hi, thank you for your great questions!</p>\n<p>Regarding the DICOM tags, you are correct. </p>\n<p>We are missing some DICOM tags to convert to nifti. I suspect the missing tag is Spacing Between Slices, because the other ones related to spacial resolution are present. </p>\n<p>I suspect the missing tag will make training 3D models more challenging, but not as much for 2D models. The LB results are good so far, but I suspect they could be even better if we had the missing tag. </p>\n<p>I will try to get access to the original files to read that tag. Let me know if you miss any other tag.</p>\n<p>Regarding the transformation in the pixel data, the complete code for facial deidentification is here: <a href=\"https://github.com/kitamura-felipe/face_deid_ct\" target=\"_blank\">https://github.com/kitamura-felipe/face_deid_ct</a></p>\n<p>In a nutshell, the code fills the air with pixels that have values in the range of those of the patient's skin. By doing that, we reduce the chance that the process can be perfectly reverted. In some cases, it removed part of the skull as well. That is a limitation of this method.</p>\n<p>We believe that models trained with images using that defacing technique will generalize if we apply the same function to newer studies on which we want to do inference. The results of the LB show that is true so far.</p>\n<p>Thank you for bringing that up!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2720785,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-03-28T15:16:16.743000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2722030,
          "author_name": "FelipeKitamura, MD, PhD",
          "author_url": "",
          "post_date": "2024-03-29T11:17:19.897000",
          "content": "<p>Only some of the studies had the tag SpaceBetweenSlices. I am uploading two CSV files with that information for the training and test sets. It should be available in 4 hours.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2722045,
          "author_name": "patriot",
          "author_url": "",
          "post_date": "2024-03-29T11:25:15.430000",
          "content": "<p>we can create 3d array from dicoms .</p>\n<pre><code> ():\n    \n    img = (img*slope +intercept) \n    img_min = window_center - window_width// \n    img_max = window_center + window_width// \n    img[img&lt;img_min] = img_min \n    img[img&gt;img_max] = img_max \n     rescale: \n        img = (img - img_min) / (img_max - img_min)* \n     img\n\n ():\n    \n    image = image - np.(image)\n     image / np.(image)\n ():\n        \n    :\n        image = dcm.pixel_array\n    :\n        \n     window_center     window_width   :\n\n        image = window_image(image, window_center, window_width, dcm.RescaleIntercept, dcm.RescaleSlope)\n\n     dcm.PhotometricInterpretation == :\n        image = np.invert(image)\n\n    normalized_image = normalize_image(image)\n     normalized_image\n ():\n\n    t_paths = (glob.glob(os.path.join(path, ))[:], key= x: (x.split()[-].split()[]))\n\n\n    dicoms = [pydicom.dcmread(d, force= )  d  t_paths]\n    sereis_instance_uid = [d.SeriesInstanceUID  d  dicoms]\n    ((sereis_instance_uid))\n\n    images = {i:[]  i  (sereis_instance_uid)}\n    z_pos = {i:[]  i  (sereis_instance_uid)}\n     dcm,s_id  (dicoms,sereis_instance_uid):\n        d = load_dicom(dcm,,)\n         d  :\n            \n        images[s_id].append(d)\n        z_pos[s_id].append((dcm.ImagePositionPatient[-]))\n\n    \n     s_id  (sereis_instance_uid):\n        images[s_id] = np.stack(images[s_id], -)[:,:,np.argsort(z_pos[s_id])]\n        z_pos[s_id] = np.sort(z_pos[s_id])\n     images,z_pos\n</code></pre>",
          "votes": 3,
          "replies": [
            {
              "id": 2723431,
              "author_name": "neuroguide",
              "author_url": "",
              "post_date": "2024-03-30T07:57:22.397000",
              "content": "<p>Hi Felipe, <br>\nThank you for help and explanation. I have explored less common converters and the Fiji is recognizing a lot of the sets. here is an example:<br>\n(Fiji Is Just) ImageJ 2.14.0/1.54f; Java 1.8.0_202 [64-bit]; </p>\n<p>Title: 002360<br>\nWidth:  249.9999 mm (512)<br>\nHeight:  249.9999 mm (512)<br>\nDepth:  173.75 mm (139)<br>\nSize:  70MB<br>\nResolution:  2.0480 pixels per mm<br>\nVoxel size: 0.4883x0.4883x1.25 mm^3<br>\nID: -212<br>\nBits per pixel: 16 (unsigned, grayscale LUT)<br>\nDisplay range: -15 - 95<br>\nRaw pixel value range: 33069 - 35553<br>\nImage: 82/139 (1.2.840.12345.4549484626668932613792536314089807458768514401)<br>\nNo threshold<br>\nScaleToFit: false</p>\n<p>So the voxel size can be read, but the typical location for shape definition has been erased like you mentioned.<br>\nIs the gender encoded in the dicom real or false?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2726054,
              "author_name": "FelipeKitamura, MD, PhD",
              "author_url": "",
              "post_date": "2024-03-31T23:57:56.103000",
              "content": "<p>Great to hear you are making progress.</p>\n<p>The gender is real and can be used as independent variable.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2720742": "Hi, thank you for your great questions!\n\nRegarding the DICOM tags, you are correct. \n\nWe are missing some DICOM tags to convert to nifti. I suspect the missing tag is Spacing Between Slices, because the other ones related to spacial resolution are present. \n\nI suspect the missing tag will make training 3D models more challenging, but not as much for 2D models. The LB results are good so far, but I suspect they could be even better if we had the missing tag. \n\nI will try to get access to the original files to read that tag. Let me know if you miss any other tag.\n\nRegarding the transformation in the pixel data, the complete code for facial deidentification is here: https://github.com/kitamura-felipe/face_deid_ct\n\nIn a nutshell, the code fills the air with pixels that have values in the range of those of the patient's skin. By doing that, we reduce the chance that the process can be perfectly reverted. In some cases, it removed part of the skull as well. That is a limitation of this method.\n\nWe believe that models trained with images using that defacing technique will generalize if we apply the same function to newer studies on which we want to do inference. The results of the LB show that is true so far.\n\nThank you for bringing that up!",
    "2720280": "Hi there, \nI'm looking at the training data and have a few questions:\n1. Can you list the modifications of the raw CT data in the dataset?\n2. Why are the dicom tags removed so that the most common dicom to nifti converters (itk and dcm2niix) do not recognize data as dicom? (I understand anonymization, but have a problem with spatial reconstruction)\n3. What are the parameters of the noise in the images?\n4. I guess that you have computed a mask for the head region of interest to modify background noise. Are you going to share it?\n5. Did the mask affect the skull bones in any way? (Bone changes in the frontal bones and temporal bones are associated with age)\n6. Can you provide the data in nifti format? if not can you provide the range of resolutions and acquisition protocols for the datasets\n7. Do you think that the results will be generelizable without knowing the above?\n"
  }
}