{
  "id": 109953,
  "title": "Rule clarification - \"pixel data only\" and what we can do with Dicom metadata",
  "url": "/competitions/rsna-intracranial-hemorrhage-detection/discussion/109953",
  "author_name": "Al-Khwârizmî",
  "post_date": "2019-09-23T17:58:55.163000",
  "votes": 38,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Hi <a href=\"/philculliton\">@philculliton</a> \nThis question was already posted in another thread, but I want to go further and elaborate on what we can do with dicom data and asking if you can clarify if it's permitted or not. Thanks !</p>\n\n<p>In this competition only few dicom metadata were given ! however, if used properly, it can become a very useful tool</p>\n\n<p><strong>1) (0010, 0020) Patient ID</strong>  : using patient ID you can gather together all the images for the same patient. Doing this helps us to understand the nature of data and the problem:\nTotal patient in train: 17079\nTotal patient in test: 2144\nThe number of patients found in train and test (duplicated):  285</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F609324%2Fe09921c46d09b23340c3076226ac46ce%2F1.PNG?generation=1569258502815669&amp;alt=media\" alt=\"\"></p>\n\n<p><strong>2)(0020, 0032)  Image Position (Patient)</strong>  : Using Image Position Patient you can sort the images and then perform a 3D CT reconstruction : \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F609324%2Fbe4ece92f98e4c65ee159c081b64d828%2F2.png?generation=1569258815738964&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F609324%2Fa8a838ee0bc256c75995268b270e56d9%2F3.png?generation=1569258829361398&amp;alt=media\" alt=\"\"></p>\n\n<p>As you can see now the images contain more than just the brain ! which explains why a lot of images are labeled 0 (all slices outside of the brain by default are labeled 0)</p>\n\n<p>Also, you can see that maybe some image transformation were applied to some images (maybe for anonymisation) like in the following example:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F609324%2F8d2867677ab79fe22c6c3bbfa3c9fe09%2F4.png?generation=1569258973163026&amp;alt=media\" alt=\"\"></p>\n\n<p>ps: you can apply some transformation here to remove the table from the images.</p>\n\n<p>*<em>3) (0020, 000d) Study Instance UID *</em> : In this competition the same patient may have different studies ! I spot some patients labeled as healthy with 4 studies (4scans) which increase the data imbalance in the dataset. Some patients have more than 100 images and other more than 200 (ID0369d8b8 or ID03962d75). I've found 365 in total, and I think there are more.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F609324%2F18a0543cc132c3781053f6ab06f1a978%2F5.PNG?generation=1569261458673592&amp;alt=media\" alt=\"\"></p>\n\n<p><strong>4) Slice thickness (the missing information)</strong>: Once you sorted your images and seperate them using Study Instance UID, you can find the slice thickness : \n slice_thickness = np.abs(slices[i].ImagePositionPatient[2] - slices[i+1].ImagePositionPatient[2])</p>\n\n<p><strong>5) (0028, 0030) Pixel Spacing</strong>  : Once you have the slice thickness and using the Pixel Spacing you can resample the images to obtain an isotropic resolution</p>\n\n<p><strong>6)(0020, 0037) Image Orientation (Patient)</strong>  : with this you can find the orientation of the image in the plane ! you can use it to make all the patients in the same orientation</p>\n\n<p>*<em>7)(0028, 0103) Pixel Representation *</em>: to tell you if it's unsigned (0) or signed integer (1) ! this is why sometimes you get some weird values  when you transform the image in HU ! </p>\n\n<p><strong>8) (0028, 1052) Rescale Intercept   Intercept:</strong> used with slope to transform the image to HU - there are 4 different intercepts in this database</p>\n\n<p>This is what come to my mind, but there are other things you can do using dicom data ! So I think we need more clarification about this rule \"Submission predictions must be based entirely on the pixel data in the provided datasets.\"</p>",
  "messages": [
    {
      "id": 632536,
      "postDate": "2019-09-23T17:58:55.163Z",
      "content": "<p>Hi <a href=\"/philculliton\">@philculliton</a> \nThis question was already posted in another thread, but I want to go further and elaborate on what we can do with dicom data and asking if you can clarify if it's permitted or not. Thanks !</p>\n\n<p>In this competition only few dicom metadata were given ! however, if used properly, it can become a very useful tool</p>\n\n<p><strong>1) (0010, 0020) Patient ID</strong>  : using patient ID you can gather together all the images for the same patient. Doing this helps us to understand the nature of data and the problem:\nTotal patient in train: 17079\nTotal patient in test: 2144\nThe number of patients found in train and test (duplicated):  285</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F609324%2Fe09921c46d09b23340c3076226ac46ce%2F1.PNG?generation=1569258502815669&amp;alt=media\" alt=\"\"></p>\n\n<p><strong>2)(0020, 0032)  Image Position (Patient)</strong>  : Using Image Position Patient you can sort the images and then perform a 3D CT reconstruction : \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F609324%2Fbe4ece92f98e4c65ee159c081b64d828%2F2.png?generation=1569258815738964&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F609324%2Fa8a838ee0bc256c75995268b270e56d9%2F3.png?generation=1569258829361398&amp;alt=media\" alt=\"\"></p>\n\n<p>As you can see now the images contain more than just the brain ! which explains why a lot of images are labeled 0 (all slices outside of the brain by default are labeled 0)</p>\n\n<p>Also, you can see that maybe some image transformation were applied to some images (maybe for anonymisation) like in the following example:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F609324%2F8d2867677ab79fe22c6c3bbfa3c9fe09%2F4.png?generation=1569258973163026&amp;alt=media\" alt=\"\"></p>\n\n<p>ps: you can apply some transformation here to remove the table from the images.</p>\n\n<p>*<em>3) (0020, 000d) Study Instance UID *</em> : In this competition the same patient may have different studies ! I spot some patients labeled as healthy with 4 studies (4scans) which increase the data imbalance in the dataset. Some patients have more than 100 images and other more than 200 (ID0369d8b8 or ID03962d75). I've found 365 in total, and I think there are more.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F609324%2F18a0543cc132c3781053f6ab06f1a978%2F5.PNG?generation=1569261458673592&amp;alt=media\" alt=\"\"></p>\n\n<p><strong>4) Slice thickness (the missing information)</strong>: Once you sorted your images and seperate them using Study Instance UID, you can find the slice thickness : \n slice_thickness = np.abs(slices[i].ImagePositionPatient[2] - slices[i+1].ImagePositionPatient[2])</p>\n\n<p><strong>5) (0028, 0030) Pixel Spacing</strong>  : Once you have the slice thickness and using the Pixel Spacing you can resample the images to obtain an isotropic resolution</p>\n\n<p><strong>6)(0020, 0037) Image Orientation (Patient)</strong>  : with this you can find the orientation of the image in the plane ! you can use it to make all the patients in the same orientation</p>\n\n<p>*<em>7)(0028, 0103) Pixel Representation *</em>: to tell you if it's unsigned (0) or signed integer (1) ! this is why sometimes you get some weird values  when you transform the image in HU ! </p>\n\n<p><strong>8) (0028, 1052) Rescale Intercept   Intercept:</strong> used with slope to transform the image to HU - there are 4 different intercepts in this database</p>\n\n<p>This is what come to my mind, but there are other things you can do using dicom data ! So I think we need more clarification about this rule \"Submission predictions must be based entirely on the pixel data in the provided datasets.\"</p>",
      "rawMarkdown": "Hi @philculliton \nThis question was already posted in another thread, but I want to go further and elaborate on what we can do with dicom data and asking if you can clarify if it's permitted or not. Thanks !\n\nIn this competition only few dicom metadata were given ! however, if used properly, it can become a very useful tool\n\n**1) (0010, 0020) Patient ID**  : using patient ID you can gather together all the images for the same patient. Doing this helps us to understand the nature of data and the problem:\nTotal patient in train: 17079\nTotal patient in test: 2144\nThe number of patients found in train and test (duplicated):  285\n\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F609324%2Fe09921c46d09b23340c3076226ac46ce%2F1.PNG?generation=1569258502815669&amp;alt=media)\n\n\n**2)(0020, 0032)  Image Position (Patient)**  : Using Image Position Patient you can sort the images and then perform a 3D CT reconstruction : \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F609324%2Fbe4ece92f98e4c65ee159c081b64d828%2F2.png?generation=1569258815738964&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F609324%2Fa8a838ee0bc256c75995268b270e56d9%2F3.png?generation=1569258829361398&amp;alt=media)\n\n\nAs you can see now the images contain more than just the brain ! which explains why a lot of images are labeled 0 (all slices outside of the brain by default are labeled 0)\n\nAlso, you can see that maybe some image transformation were applied to some images (maybe for anonymisation) like in the following example:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F609324%2F8d2867677ab79fe22c6c3bbfa3c9fe09%2F4.png?generation=1569258973163026&amp;alt=media)\n\n\nps: you can apply some transformation here to remove the table from the images.\n\n**3) (0020, 000d) Study Instance UID ** : In this competition the same patient may have different studies ! I spot some patients labeled as healthy with 4 studies (4scans) which increase the data imbalance in the dataset. Some patients have more than 100 images and other more than 200 (ID0369d8b8 or ID03962d75). I've found 365 in total, and I think there are more.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F609324%2F18a0543cc132c3781053f6ab06f1a978%2F5.PNG?generation=1569261458673592&amp;alt=media)\n\n\n**4) Slice thickness (the missing information)**: Once you sorted your images and seperate them using Study Instance UID, you can find the slice thickness : \n slice_thickness = np.abs(slices[i].ImagePositionPatient[2] - slices[i+1].ImagePositionPatient[2])\n\n**5) (0028, 0030) Pixel Spacing**  : Once you have the slice thickness and using the Pixel Spacing you can resample the images to obtain an isotropic resolution\n\n**6)(0020, 0037) Image Orientation (Patient)**  : with this you can find the orientation of the image in the plane ! you can use it to make all the patients in the same orientation\n\n**7)(0028, 0103) Pixel Representation **: to tell you if it's unsigned (0) or signed integer (1) ! this is why sometimes you get some weird values  when you transform the image in HU ! \n\n**8) (0028, 1052) Rescale Intercept   Intercept:** used with slope to transform the image to HU - there are 4 different intercepts in this database\n\nThis is what come to my mind, but there are other things you can do using dicom data ! So I think we need more clarification about this rule \"Submission predictions must be based entirely on the pixel data in the provided datasets.\"",
      "votes": 38
    },
    {
      "id": 632564,
      "postDate": "2019-09-23T18:51:23.527Z",
      "content": "<p>Some patients have multiple Study Instance UID  and Series Instance UID, most likely these are patients who had multiple studies at different times. \nI agree, we need clarification. I understand that we should not use metadata to try to find data leaks, but using them to analyze images from the same series together should be allowed. As a radiologist, I do not review each image by itself, but in the context of the study. For instance, analyzing adjacent images together help to distinguish partial volume artefacts from real findings.  </p>",
      "rawMarkdown": "Some patients have multiple Study Instance UID  and Series Instance UID, most likely these are patients who had multiple studies at different times. \nI agree, we need clarification. I understand that we should not use metadata to try to find data leaks, but using them to analyze images from the same series together should be allowed. As a radiologist, I do not review each image by itself, but in the context of the study. For instance, analyzing adjacent images together help to distinguish partial volume artefacts from real findings.  \n",
      "votes": 6
    },
    {
      "id": 633722,
      "postDate": "2019-09-25T10:26:07.620Z",
      "content": "<p>Pixel data implies using 2D images only. If we are allowed to use 3D data, perhaps the rules need to be expanded to voxel data too?</p>",
      "rawMarkdown": "Pixel data implies using 2D images only. If we are allowed to use 3D data, perhaps the rules need to be expanded to voxel data too?",
      "votes": 1
    },
    {
      "id": 633440,
      "postDate": "2019-09-24T23:57:57.743Z",
      "content": "<p>I agree, clarification is key.   I'm unsure if we can assume that the validation data has something of a whole scan in there versus just some slices in there.  While I understand a slice-based approach, I find it would find it strange as the CT is taken of the head (or some part of it) and not one slice at a time.  From the whole brain you can still get a slice-specific bleed indicator, but using the whole head seems to make more sense to me.</p>",
      "rawMarkdown": "I agree, clarification is key.   I'm unsure if we can assume that the validation data has something of a whole scan in there versus just some slices in there.  While I understand a slice-based approach, I find it would find it strange as the CT is taken of the head (or some part of it) and not one slice at a time.  From the whole brain you can still get a slice-specific bleed indicator, but using the whole head seems to make more sense to me.",
      "votes": 1
    },
    {
      "id": 637528,
      "postDate": "2019-10-01T05:13:10.957Z",
      "content": "<p>We’ve confirmed with the host that <strong>metadata can be used for pre-processing of images, but they cannot be features in your model or used to change or label predictions post-modeling</strong>. This should help resolve many of the more specific questions that have been raised recently on what types of things can be done with the metadata.</p>",
      "rawMarkdown": "We’ve confirmed with the host that **metadata can be used for pre-processing of images, but they cannot be features in your model or used to change or label predictions post-modeling**. This should help resolve many of the more specific questions that have been raised recently on what types of things can be done with the metadata.",
      "votes": -2,
      "replies": [
        {
          "id": 637768,
          "postDate": "2019-10-01T08:42:44.640Z",
          "content": "<p>Maybe it is just me, but I still cannot figure out if it is allowed to group slices from the same study images together, or not, i.e. to use 3d-convolutions.</p>",
          "rawMarkdown": "Maybe it is just me, but I still cannot figure out if it is allowed to group slices from the same study images together, or not, i.e. to use 3d-convolutions.",
          "votes": 4
        }
      ]
    },
    {
      "id": 643473,
      "postDate": "2019-10-07T14:34:53.587Z",
      "content": "<p>We can use 3D convolution as per the new policy, right?</p>",
      "rawMarkdown": "We can use 3D convolution as per the new policy, right?"
    },
    {
      "id": 639253,
      "postDate": "2019-10-03T01:38:52.007Z",
      "content": "<p>Nice job.</p>",
      "rawMarkdown": "Nice job."
    },
    {
      "id": 637791,
      "postDate": "2019-10-01T09:03:07.967Z",
      "content": "<p>Thank you for the great summary of the metadata!</p>\n\n<blockquote>\n  <p>5) (0028, 0030) Pixel Spacing : Once you have the slice thickness and using the Pixel Spacing you can resample the images to obtain an isotropic resolution</p>\n</blockquote>\n\n<p>Could you please elaborate on this one?</p>",
      "rawMarkdown": "Thank you for the great summary of the metadata!\n\n&gt; 5) (0028, 0030) Pixel Spacing : Once you have the slice thickness and using the Pixel Spacing you can resample the images to obtain an isotropic resolution\n\nCould you please elaborate on this one?"
    },
    {
      "id": 636088,
      "postDate": "2019-09-28T19:50:20.290Z",
      "content": "<p>About slice thickness, it is correct that there should be 3 components in the vector, for width, height, and depth? And there are only 2 elements available? So it is not possible to reconstruct pixel density along the depth dimension. Or should it be considered as <code>1</code>?</p>",
      "rawMarkdown": "About slice thickness, it is correct that there should be 3 components in the vector, for width, height, and depth? And there are only 2 elements available? So it is not possible to reconstruct pixel density along the depth dimension. Or should it be considered as `1`?"
    },
    {
      "id": 635611,
      "postDate": "2019-09-27T23:24:36.407Z",
      "content": "<p>Thanks for this, Al-Khwârizmî! I think everyone should be exploring using the 3D reconstruction!</p>",
      "rawMarkdown": "Thanks for this, Al-Khwârizmî! I think everyone should be exploring using the 3D reconstruction!"
    },
    {
      "id": 633959,
      "postDate": "2019-09-25T16:14:07.310Z",
      "content": "<p>So I am thinking how the resampling of images will affect accuracy. Will it be better or different pixel spacing makes model more robust?</p>",
      "rawMarkdown": "So I am thinking how the resampling of images will affect accuracy. Will it be better or different pixel spacing makes model more robust?\n"
    }
  ],
  "comments": [
    {
      "id": 632564,
      "author_name": "Amil Gentili",
      "author_url": "",
      "post_date": "2019-09-23T18:51:23.527000",
      "content": "<p>Some patients have multiple Study Instance UID  and Series Instance UID, most likely these are patients who had multiple studies at different times. \nI agree, we need clarification. I understand that we should not use metadata to try to find data leaks, but using them to analyze images from the same series together should be allowed. As a radiologist, I do not review each image by itself, but in the context of the study. For instance, analyzing adjacent images together help to distinguish partial volume artefacts from real findings.  </p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 633722,
      "author_name": "datasaurus",
      "author_url": "",
      "post_date": "2019-09-25T10:26:07.620000",
      "content": "<p>Pixel data implies using 2D images only. If we are allowed to use 3D data, perhaps the rules need to be expanded to voxel data too?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 633440,
      "author_name": "John M",
      "author_url": "",
      "post_date": "2019-09-24T23:57:57.743000",
      "content": "<p>I agree, clarification is key.   I'm unsure if we can assume that the validation data has something of a whole scan in there versus just some slices in there.  While I understand a slice-based approach, I find it would find it strange as the CT is taken of the head (or some part of it) and not one slice at a time.  From the whole brain you can still get a slice-specific bleed indicator, but using the whole head seems to make more sense to me.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 637528,
      "author_name": "Julia Elliott",
      "author_url": "",
      "post_date": "2019-10-01T05:13:10.957000",
      "content": "<p>We’ve confirmed with the host that <strong>metadata can be used for pre-processing of images, but they cannot be features in your model or used to change or label predictions post-modeling</strong>. This should help resolve many of the more specific questions that have been raised recently on what types of things can be done with the metadata.</p>",
      "votes": -2,
      "replies": [
        {
          "id": 637768,
          "author_name": "Ilia Zaitsev",
          "author_url": "",
          "post_date": "2019-10-01T08:42:44.640000",
          "content": "<p>Maybe it is just me, but I still cannot figure out if it is allowed to group slices from the same study images together, or not, i.e. to use 3d-convolutions.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 643473,
      "author_name": "tandelDipak",
      "author_url": "",
      "post_date": "2019-10-07T14:34:53.587000",
      "content": "<p>We can use 3D convolution as per the new policy, right?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 639253,
      "author_name": "Bo Peng",
      "author_url": "",
      "post_date": "2019-10-03T01:38:52.007000",
      "content": "<p>Nice job.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 637791,
      "author_name": "Dmytro Panchenko",
      "author_url": "",
      "post_date": "2019-10-01T09:03:07.967000",
      "content": "<p>Thank you for the great summary of the metadata!</p>\n\n<blockquote>\n  <p>5) (0028, 0030) Pixel Spacing : Once you have the slice thickness and using the Pixel Spacing you can resample the images to obtain an isotropic resolution</p>\n</blockquote>\n\n<p>Could you please elaborate on this one?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 636088,
      "author_name": "Ilia Zaitsev",
      "author_url": "",
      "post_date": "2019-09-28T19:50:20.290000",
      "content": "<p>About slice thickness, it is correct that there should be 3 components in the vector, for width, height, and depth? And there are only 2 elements available? So it is not possible to reconstruct pixel density along the depth dimension. Or should it be considered as <code>1</code>?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 635611,
      "author_name": "Tom H.",
      "author_url": "",
      "post_date": "2019-09-27T23:24:36.407000",
      "content": "<p>Thanks for this, Al-Khwârizmî! I think everyone should be exploring using the 3D reconstruction!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 633959,
      "author_name": "Kuda",
      "author_url": "",
      "post_date": "2019-09-25T16:14:07.310000",
      "content": "<p>So I am thinking how the resampling of images will affect accuracy. Will it be better or different pixel spacing makes model more robust?</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "632536": "Hi @philculliton \nThis question was already posted in another thread, but I want to go further and elaborate on what we can do with dicom data and asking if you can clarify if it's permitted or not. Thanks !\n\nIn this competition only few dicom metadata were given ! however, if used properly, it can become a very useful tool\n\n**1) (0010, 0020) Patient ID**  : using patient ID you can gather together all the images for the same patient. Doing this helps us to understand the nature of data and the problem:\nTotal patient in train: 17079\nTotal patient in test: 2144\nThe number of patients found in train and test (duplicated):  285\n\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F609324%2Fe09921c46d09b23340c3076226ac46ce%2F1.PNG?generation=1569258502815669&amp;alt=media)\n\n\n**2)(0020, 0032)  Image Position (Patient)**  : Using Image Position Patient you can sort the images and then perform a 3D CT reconstruction : \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F609324%2Fbe4ece92f98e4c65ee159c081b64d828%2F2.png?generation=1569258815738964&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F609324%2Fa8a838ee0bc256c75995268b270e56d9%2F3.png?generation=1569258829361398&amp;alt=media)\n\n\nAs you can see now the images contain more than just the brain ! which explains why a lot of images are labeled 0 (all slices outside of the brain by default are labeled 0)\n\nAlso, you can see that maybe some image transformation were applied to some images (maybe for anonymisation) like in the following example:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F609324%2F8d2867677ab79fe22c6c3bbfa3c9fe09%2F4.png?generation=1569258973163026&amp;alt=media)\n\n\nps: you can apply some transformation here to remove the table from the images.\n\n**3) (0020, 000d) Study Instance UID ** : In this competition the same patient may have different studies ! I spot some patients labeled as healthy with 4 studies (4scans) which increase the data imbalance in the dataset. Some patients have more than 100 images and other more than 200 (ID0369d8b8 or ID03962d75). I've found 365 in total, and I think there are more.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F609324%2F18a0543cc132c3781053f6ab06f1a978%2F5.PNG?generation=1569261458673592&amp;alt=media)\n\n\n**4) Slice thickness (the missing information)**: Once you sorted your images and seperate them using Study Instance UID, you can find the slice thickness : \n slice_thickness = np.abs(slices[i].ImagePositionPatient[2] - slices[i+1].ImagePositionPatient[2])\n\n**5) (0028, 0030) Pixel Spacing**  : Once you have the slice thickness and using the Pixel Spacing you can resample the images to obtain an isotropic resolution\n\n**6)(0020, 0037) Image Orientation (Patient)**  : with this you can find the orientation of the image in the plane ! you can use it to make all the patients in the same orientation\n\n**7)(0028, 0103) Pixel Representation **: to tell you if it's unsigned (0) or signed integer (1) ! this is why sometimes you get some weird values  when you transform the image in HU ! \n\n**8) (0028, 1052) Rescale Intercept   Intercept:** used with slope to transform the image to HU - there are 4 different intercepts in this database\n\nThis is what come to my mind, but there are other things you can do using dicom data ! So I think we need more clarification about this rule \"Submission predictions must be based entirely on the pixel data in the provided datasets.\"",
    "632564": "Some patients have multiple Study Instance UID  and Series Instance UID, most likely these are patients who had multiple studies at different times. \nI agree, we need clarification. I understand that we should not use metadata to try to find data leaks, but using them to analyze images from the same series together should be allowed. As a radiologist, I do not review each image by itself, but in the context of the study. For instance, analyzing adjacent images together help to distinguish partial volume artefacts from real findings.  \n",
    "633722": "Pixel data implies using 2D images only. If we are allowed to use 3D data, perhaps the rules need to be expanded to voxel data too?",
    "633440": "I agree, clarification is key.   I'm unsure if we can assume that the validation data has something of a whole scan in there versus just some slices in there.  While I understand a slice-based approach, I find it would find it strange as the CT is taken of the head (or some part of it) and not one slice at a time.  From the whole brain you can still get a slice-specific bleed indicator, but using the whole head seems to make more sense to me.",
    "637528": "We’ve confirmed with the host that **metadata can be used for pre-processing of images, but they cannot be features in your model or used to change or label predictions post-modeling**. This should help resolve many of the more specific questions that have been raised recently on what types of things can be done with the metadata.",
    "643473": "We can use 3D convolution as per the new policy, right?",
    "639253": "Nice job.",
    "637791": "Thank you for the great summary of the metadata!\n\n&gt; 5) (0028, 0030) Pixel Spacing : Once you have the slice thickness and using the Pixel Spacing you can resample the images to obtain an isotropic resolution\n\nCould you please elaborate on this one?",
    "636088": "About slice thickness, it is correct that there should be 3 components in the vector, for width, height, and depth? And there are only 2 elements available? So it is not possible to reconstruct pixel density along the depth dimension. Or should it be considered as `1`?",
    "635611": "Thanks for this, Al-Khwârizmî! I think everyone should be exploring using the 3D reconstruction!",
    "633959": "So I am thinking how the resampling of images will affect accuracy. Will it be better or different pixel spacing makes model more robust?\n"
  }
}