{
  "id": 182736,
  "title": "Relationship between StudyInstanceUID,SeriesInstanceUID and SOPInstanceUID ",
  "url": "/competitions/rsna-str-pulmonary-embolism-detection/discussion/182736",
  "author_name": "Gelesh G Omathil",
  "post_date": "2020-09-14T06:07:13.534000",
  "votes": 3,
  "comment_count": 3,
  "views": 0,
  "content": "<p>May I know the relationship between<br>\nRelationship between <strong><em><em>StudyInstanceUID, SeriesInstanceUID and SOPInstanceUID</em></em></strong> .</p>\n<p>I am looking for some explanation like :<br>\n<em>One patient scan would be a StudyInstanceUID, and also a SeriesInstanceUID .\nMultiple Scan of Same Patient images collected in a sequence uniquely identified by SOPInstanceUID</em><br>\nOr could some one explain </p>",
  "messages": [
    {
      "id": 1009881,
      "postDate": "2020-09-14T10:42:00.330Z",
      "content": "<p>UIDs are \"unique identifiers\" and are basically ID numbers for things in the DICOM world.</p>\n<p>StudyInstanceUID - a unique identifier for a single imaging study. In this case a CT scan.</p>\n<p>SeriesInstanceUID - during a CT scan, you can take one or more sets of images. Each set of images is called a \"Series\". In this case, only one Series for each Study (so this directory level isn't meaningful for this contest, but is maintained because it is a common data structure to organized images).</p>\n<p>SOPInstanceUID - SOP means \"Service-Object Pair\" and is another basic DICOM idea. In this case, it refers to each image. </p>\n<p>In the real world, these look something like this:</p>\n<p>1.1.1.1.2.2264008.11756025286.32766</p>\n<p>and there is meaning encoded in the numbers. For this contest, they are replaced with random hexadecimal numbers. The SOPInstanceUID names are not in any specific order, and you need to look at the DICOM metadata to put the images in order.</p>\n<p>In the real world, a Patient could have multiple Studies (say a yearly CT scan). In this contest, we are only given one study per patient.</p>\n<p>In the real world, a CT scan would include at least two series. First, a \"Scout\" which is like an X-ray picture and let's the technologist see where the top and bottom of the lungs are so they can plan the study. Then images of the chest going slice by slice from top to bottom (or bottom to top).</p>\n<p>A CTA of the chest for Pulmonary Embolism could include a Series that is used for timing how long intravenous contrast takes to get to the pulmonary arteries. In this contest we are given only one series per patient, so we don't have to figure out which series to use for a Study.</p>\n<p>For a CT Scan, the images (SOPInstanceUID) can be stacked together to make up a 3D volume.</p>\n<p>For DICOM, there is no requirement that images are organized in nested directories. Inside each DICOM file (in addition to image data) is MetaData that identifies the Patient, date of study, scanner used, position in a three-dimensional space that relates to other images, etc. There are typically dozens or hundreds of data items. Most have been anonymized for this study. Even if all the images were in one giant directory, you could use the metadata to reassemble the Studies, Series and Images.</p>\n<p>Pydicom is a library that you can use in Python to inspect the metadata.</p>\n<p>A couple of Metadata items of interest:</p>\n<p>Pixel size - unlike a photograph, CT images have the real size of the image recorded. So you can actually calculate true measurements of structures in each image.</p>\n<p>Slice Thickness - since the image is of a physical object, it has a thickness of how much of the chest was imaged in that slice. Comparable to a loaf of bread being sliced. Slices could be thick or thin. Typical is 1.25 mm. Some cases are up to 5 mm. Varies by machine and local imaging style.</p>\n<p>Rows/Cols - all the Train images are 512 x 512. A typical value for a modern CT scanner. I wouldn't presume that the Test data is also 512 x 512, but if I had to guess, it would be. Best to write your code to handle other sizes.</p>\n<p>Patient position - FFP (Feet first prone), FFS (feet first supine), HFS (head first supine). Refers to the position of the patient in the machine. In the train dataset:</p>\n<p>FFP    242 <br>\nFFS    1732839<br>\nHFS    57513   </p>\n<p>In theory, this could affect the patient position. The FFP (prone) images suggest the patient is lying on their belly. Supposedly you would want to flip these images (but I haven't looked at the DICOM to see if this is true).</p>\n<p>InstanceNumber - a unique number for each image within a Series. In these series, usually starts at 1 and counts up in the physical order of the images. There are Studies that don't start at zero. I don't know if there are cases where numbers are skipped. Could count from top to bottom or bottom to top. I think the examples I checked are all top to bottom, but that might not be 100%.</p>\n<p>Image Patient Position - three floating point numbers that define a relative location in the real world. Image Patient Position [2] is the \"Z-axis\" and can be used to put the images in order.</p>\n<p>Image Orientation - another parameter that defines the orientation of the patient/images in space. Might be relevant for building three dimensional models and defining top/bottom of patient. I haven't checked.</p>\n<p>You can use these parameters to reassemble the images in order and build a 3D model. Note that you are not guaranteed that the images give you a physically complete image. Nothing stops you from having gaps in the images. This would not be normal processing, but since we cannot see the real test data, we cannot check this.</p>\n<p>In theory, if you put them in order by InstanceNumber or Image Patient Position [2], the change in position would equal the Slice Thickness. But in practice, you could have gaps or even overlaps (5 mm slices taken every 2.5 mm). How far you go to cover the possibilities depends on your model needs. If you look at the data in the Pulmonary Fibrosis competition, the CT images have missing images, gaps in images, repeated imaging through the patient in one Study, etc.</p>\n<p>Long answer to a short question.</p>\n<p>-Rich</p>",
      "rawMarkdown": "UIDs are \"unique identifiers\" and are basically ID numbers for things in the DICOM world.\n\nStudyInstanceUID - a unique identifier for a single imaging study. In this case a CT scan.\n\nSeriesInstanceUID - during a CT scan, you can take one or more sets of images. Each set of images is called a \"Series\". In this case, only one Series for each Study (so this directory level isn't meaningful for this contest, but is maintained because it is a common data structure to organized images).\n\nSOPInstanceUID - SOP means \"Service-Object Pair\" and is another basic DICOM idea. In this case, it refers to each image. \n\nIn the real world, these look something like this:\n\n1.1.1.1.2.2264008.11756025286.32766\n\nand there is meaning encoded in the numbers. For this contest, they are replaced with random hexadecimal numbers. The SOPInstanceUID names are not in any specific order, and you need to look at the DICOM metadata to put the images in order.\n\nIn the real world, a Patient could have multiple Studies (say a yearly CT scan). In this contest, we are only given one study per patient.\n\nIn the real world, a CT scan would include at least two series. First, a \"Scout\" which is like an X-ray picture and let's the technologist see where the top and bottom of the lungs are so they can plan the study. Then images of the chest going slice by slice from top to bottom (or bottom to top).\n\nA CTA of the chest for Pulmonary Embolism could include a Series that is used for timing how long intravenous contrast takes to get to the pulmonary arteries. In this contest we are given only one series per patient, so we don't have to figure out which series to use for a Study.\n\nFor a CT Scan, the images (SOPInstanceUID) can be stacked together to make up a 3D volume.\n\nFor DICOM, there is no requirement that images are organized in nested directories. Inside each DICOM file (in addition to image data) is MetaData that identifies the Patient, date of study, scanner used, position in a three-dimensional space that relates to other images, etc. There are typically dozens or hundreds of data items. Most have been anonymized for this study. Even if all the images were in one giant directory, you could use the metadata to reassemble the Studies, Series and Images.\n\nPydicom is a library that you can use in Python to inspect the metadata.\n\nA couple of Metadata items of interest:\n\nPixel size - unlike a photograph, CT images have the real size of the image recorded. So you can actually calculate true measurements of structures in each image.\n\nSlice Thickness - since the image is of a physical object, it has a thickness of how much of the chest was imaged in that slice. Comparable to a loaf of bread being sliced. Slices could be thick or thin. Typical is 1.25 mm. Some cases are up to 5 mm. Varies by machine and local imaging style.\n\nRows/Cols - all the Train images are 512 x 512. A typical value for a modern CT scanner. I wouldn't presume that the Test data is also 512 x 512, but if I had to guess, it would be. Best to write your code to handle other sizes.\n\nPatient position - FFP (Feet first prone), FFS (feet first supine), HFS (head first supine). Refers to the position of the patient in the machine. In the train dataset:\n\nFFP\t242\t\nFFS\t1732839\nHFS\t57513\t\n\nIn theory, this could affect the patient position. The FFP (prone) images suggest the patient is lying on their belly. Supposedly you would want to flip these images (but I haven't looked at the DICOM to see if this is true).\n\nInstanceNumber - a unique number for each image within a Series. In these series, usually starts at 1 and counts up in the physical order of the images. There are Studies that don't start at zero. I don't know if there are cases where numbers are skipped. Could count from top to bottom or bottom to top. I think the examples I checked are all top to bottom, but that might not be 100%.\n\nImage Patient Position - three floating point numbers that define a relative location in the real world. Image Patient Position [2] is the \"Z-axis\" and can be used to put the images in order.\n\nImage Orientation - another parameter that defines the orientation of the patient/images in space. Might be relevant for building three dimensional models and defining top/bottom of patient. I haven't checked.\n\nYou can use these parameters to reassemble the images in order and build a 3D model. Note that you are not guaranteed that the images give you a physically complete image. Nothing stops you from having gaps in the images. This would not be normal processing, but since we cannot see the real test data, we cannot check this.\n\nIn theory, if you put them in order by InstanceNumber or Image Patient Position [2], the change in position would equal the Slice Thickness. But in practice, you could have gaps or even overlaps (5 mm slices taken every 2.5 mm). How far you go to cover the possibilities depends on your model needs. If you look at the data in the Pulmonary Fibrosis competition, the CT images have missing images, gaps in images, repeated imaging through the patient in one Study, etc.\n\nLong answer to a short question.\n\n-Rich\n",
      "votes": 9,
      "replies": [
        {
          "id": 1015238,
          "postDate": "2020-09-18T04:13:49.250Z",
          "content": "<p>Thank you for this wonderful and detailed explanation. This is really helpful.:)</p>",
          "rawMarkdown": "Thank you for this wonderful and detailed explanation. This is really helpful.:)"
        }
      ]
    },
    {
      "id": 1009606,
      "postDate": "2020-09-14T06:07:13.533Z",
      "content": "<p>May I know the relationship between<br>\nRelationship between <strong><em><em>StudyInstanceUID, SeriesInstanceUID and SOPInstanceUID</em></em></strong> .</p>\n<p>I am looking for some explanation like :<br>\n<em>One patient scan would be a StudyInstanceUID, and also a SeriesInstanceUID .\nMultiple Scan of Same Patient images collected in a sequence uniquely identified by SOPInstanceUID</em><br>\nOr could some one explain </p>",
      "rawMarkdown": "May I know the relationship between\nRelationship between ****StudyInstanceUID, SeriesInstanceUID and SOPInstanceUID**** .\n\nI am looking for some explanation like :\n*One patient scan would be a StudyInstanceUID, and also a SeriesInstanceUID .\nMultiple Scan of Same Patient images collected in a sequence uniquely identified by SOPInstanceUID*\nOr could some one explain ",
      "votes": 3
    },
    {
      "id": 1009827,
      "postDate": "2020-09-14T09:17:36.733Z",
      "content": "<p>Well, in short terms:</p>\n<p>Each <strong>StudyInstanceUID</strong> contains one or more <strong>SeriesInstanceUID</strong><br>\nEach <strong>SeriesInstanceUID</strong> contrains one or more <strong>SOPInstanceUID</strong><br>\nEach <strong>SOPInstanceUID</strong> appended with <code>.dcm</code> is the filename for each DICOM file.</p>\n<p>One Instance can be many series of scans. One scan will have many '.dcm' files.</p>",
      "rawMarkdown": "Well, in short terms:\n\nEach **StudyInstanceUID** contains one or more **SeriesInstanceUID**\nEach **SeriesInstanceUID** contrains one or more **SOPInstanceUID**\nEach **SOPInstanceUID** appended with `.dcm` is the filename for each DICOM file.\n\nOne Instance can be many series of scans. One scan will have many '.dcm' files.",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 1009881,
      "author_name": "quadcore/Richard Epstein",
      "author_url": "",
      "post_date": "2020-09-14T10:42:00.330000",
      "content": "<p>UIDs are \"unique identifiers\" and are basically ID numbers for things in the DICOM world.</p>\n<p>StudyInstanceUID - a unique identifier for a single imaging study. In this case a CT scan.</p>\n<p>SeriesInstanceUID - during a CT scan, you can take one or more sets of images. Each set of images is called a \"Series\". In this case, only one Series for each Study (so this directory level isn't meaningful for this contest, but is maintained because it is a common data structure to organized images).</p>\n<p>SOPInstanceUID - SOP means \"Service-Object Pair\" and is another basic DICOM idea. In this case, it refers to each image. </p>\n<p>In the real world, these look something like this:</p>\n<p>1.1.1.1.2.2264008.11756025286.32766</p>\n<p>and there is meaning encoded in the numbers. For this contest, they are replaced with random hexadecimal numbers. The SOPInstanceUID names are not in any specific order, and you need to look at the DICOM metadata to put the images in order.</p>\n<p>In the real world, a Patient could have multiple Studies (say a yearly CT scan). In this contest, we are only given one study per patient.</p>\n<p>In the real world, a CT scan would include at least two series. First, a \"Scout\" which is like an X-ray picture and let's the technologist see where the top and bottom of the lungs are so they can plan the study. Then images of the chest going slice by slice from top to bottom (or bottom to top).</p>\n<p>A CTA of the chest for Pulmonary Embolism could include a Series that is used for timing how long intravenous contrast takes to get to the pulmonary arteries. In this contest we are given only one series per patient, so we don't have to figure out which series to use for a Study.</p>\n<p>For a CT Scan, the images (SOPInstanceUID) can be stacked together to make up a 3D volume.</p>\n<p>For DICOM, there is no requirement that images are organized in nested directories. Inside each DICOM file (in addition to image data) is MetaData that identifies the Patient, date of study, scanner used, position in a three-dimensional space that relates to other images, etc. There are typically dozens or hundreds of data items. Most have been anonymized for this study. Even if all the images were in one giant directory, you could use the metadata to reassemble the Studies, Series and Images.</p>\n<p>Pydicom is a library that you can use in Python to inspect the metadata.</p>\n<p>A couple of Metadata items of interest:</p>\n<p>Pixel size - unlike a photograph, CT images have the real size of the image recorded. So you can actually calculate true measurements of structures in each image.</p>\n<p>Slice Thickness - since the image is of a physical object, it has a thickness of how much of the chest was imaged in that slice. Comparable to a loaf of bread being sliced. Slices could be thick or thin. Typical is 1.25 mm. Some cases are up to 5 mm. Varies by machine and local imaging style.</p>\n<p>Rows/Cols - all the Train images are 512 x 512. A typical value for a modern CT scanner. I wouldn't presume that the Test data is also 512 x 512, but if I had to guess, it would be. Best to write your code to handle other sizes.</p>\n<p>Patient position - FFP (Feet first prone), FFS (feet first supine), HFS (head first supine). Refers to the position of the patient in the machine. In the train dataset:</p>\n<p>FFP    242 <br>\nFFS    1732839<br>\nHFS    57513   </p>\n<p>In theory, this could affect the patient position. The FFP (prone) images suggest the patient is lying on their belly. Supposedly you would want to flip these images (but I haven't looked at the DICOM to see if this is true).</p>\n<p>InstanceNumber - a unique number for each image within a Series. In these series, usually starts at 1 and counts up in the physical order of the images. There are Studies that don't start at zero. I don't know if there are cases where numbers are skipped. Could count from top to bottom or bottom to top. I think the examples I checked are all top to bottom, but that might not be 100%.</p>\n<p>Image Patient Position - three floating point numbers that define a relative location in the real world. Image Patient Position [2] is the \"Z-axis\" and can be used to put the images in order.</p>\n<p>Image Orientation - another parameter that defines the orientation of the patient/images in space. Might be relevant for building three dimensional models and defining top/bottom of patient. I haven't checked.</p>\n<p>You can use these parameters to reassemble the images in order and build a 3D model. Note that you are not guaranteed that the images give you a physically complete image. Nothing stops you from having gaps in the images. This would not be normal processing, but since we cannot see the real test data, we cannot check this.</p>\n<p>In theory, if you put them in order by InstanceNumber or Image Patient Position [2], the change in position would equal the Slice Thickness. But in practice, you could have gaps or even overlaps (5 mm slices taken every 2.5 mm). How far you go to cover the possibilities depends on your model needs. If you look at the data in the Pulmonary Fibrosis competition, the CT images have missing images, gaps in images, repeated imaging through the patient in one Study, etc.</p>\n<p>Long answer to a short question.</p>\n<p>-Rich</p>",
      "votes": 9,
      "replies": [
        {
          "id": 1015238,
          "author_name": "Aman Arora",
          "author_url": "",
          "post_date": "2020-09-18T04:13:49.250000",
          "content": "<p>Thank you for this wonderful and detailed explanation. This is really helpful.:)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1009827,
      "author_name": "Venturillo JE",
      "author_url": "",
      "post_date": "2020-09-14T09:17:36.733000",
      "content": "<p>Well, in short terms:</p>\n<p>Each <strong>StudyInstanceUID</strong> contains one or more <strong>SeriesInstanceUID</strong><br>\nEach <strong>SeriesInstanceUID</strong> contrains one or more <strong>SOPInstanceUID</strong><br>\nEach <strong>SOPInstanceUID</strong> appended with <code>.dcm</code> is the filename for each DICOM file.</p>\n<p>One Instance can be many series of scans. One scan will have many '.dcm' files.</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1009881": "UIDs are \"unique identifiers\" and are basically ID numbers for things in the DICOM world.\n\nStudyInstanceUID - a unique identifier for a single imaging study. In this case a CT scan.\n\nSeriesInstanceUID - during a CT scan, you can take one or more sets of images. Each set of images is called a \"Series\". In this case, only one Series for each Study (so this directory level isn't meaningful for this contest, but is maintained because it is a common data structure to organized images).\n\nSOPInstanceUID - SOP means \"Service-Object Pair\" and is another basic DICOM idea. In this case, it refers to each image. \n\nIn the real world, these look something like this:\n\n1.1.1.1.2.2264008.11756025286.32766\n\nand there is meaning encoded in the numbers. For this contest, they are replaced with random hexadecimal numbers. The SOPInstanceUID names are not in any specific order, and you need to look at the DICOM metadata to put the images in order.\n\nIn the real world, a Patient could have multiple Studies (say a yearly CT scan). In this contest, we are only given one study per patient.\n\nIn the real world, a CT scan would include at least two series. First, a \"Scout\" which is like an X-ray picture and let's the technologist see where the top and bottom of the lungs are so they can plan the study. Then images of the chest going slice by slice from top to bottom (or bottom to top).\n\nA CTA of the chest for Pulmonary Embolism could include a Series that is used for timing how long intravenous contrast takes to get to the pulmonary arteries. In this contest we are given only one series per patient, so we don't have to figure out which series to use for a Study.\n\nFor a CT Scan, the images (SOPInstanceUID) can be stacked together to make up a 3D volume.\n\nFor DICOM, there is no requirement that images are organized in nested directories. Inside each DICOM file (in addition to image data) is MetaData that identifies the Patient, date of study, scanner used, position in a three-dimensional space that relates to other images, etc. There are typically dozens or hundreds of data items. Most have been anonymized for this study. Even if all the images were in one giant directory, you could use the metadata to reassemble the Studies, Series and Images.\n\nPydicom is a library that you can use in Python to inspect the metadata.\n\nA couple of Metadata items of interest:\n\nPixel size - unlike a photograph, CT images have the real size of the image recorded. So you can actually calculate true measurements of structures in each image.\n\nSlice Thickness - since the image is of a physical object, it has a thickness of how much of the chest was imaged in that slice. Comparable to a loaf of bread being sliced. Slices could be thick or thin. Typical is 1.25 mm. Some cases are up to 5 mm. Varies by machine and local imaging style.\n\nRows/Cols - all the Train images are 512 x 512. A typical value for a modern CT scanner. I wouldn't presume that the Test data is also 512 x 512, but if I had to guess, it would be. Best to write your code to handle other sizes.\n\nPatient position - FFP (Feet first prone), FFS (feet first supine), HFS (head first supine). Refers to the position of the patient in the machine. In the train dataset:\n\nFFP\t242\t\nFFS\t1732839\nHFS\t57513\t\n\nIn theory, this could affect the patient position. The FFP (prone) images suggest the patient is lying on their belly. Supposedly you would want to flip these images (but I haven't looked at the DICOM to see if this is true).\n\nInstanceNumber - a unique number for each image within a Series. In these series, usually starts at 1 and counts up in the physical order of the images. There are Studies that don't start at zero. I don't know if there are cases where numbers are skipped. Could count from top to bottom or bottom to top. I think the examples I checked are all top to bottom, but that might not be 100%.\n\nImage Patient Position - three floating point numbers that define a relative location in the real world. Image Patient Position [2] is the \"Z-axis\" and can be used to put the images in order.\n\nImage Orientation - another parameter that defines the orientation of the patient/images in space. Might be relevant for building three dimensional models and defining top/bottom of patient. I haven't checked.\n\nYou can use these parameters to reassemble the images in order and build a 3D model. Note that you are not guaranteed that the images give you a physically complete image. Nothing stops you from having gaps in the images. This would not be normal processing, but since we cannot see the real test data, we cannot check this.\n\nIn theory, if you put them in order by InstanceNumber or Image Patient Position [2], the change in position would equal the Slice Thickness. But in practice, you could have gaps or even overlaps (5 mm slices taken every 2.5 mm). How far you go to cover the possibilities depends on your model needs. If you look at the data in the Pulmonary Fibrosis competition, the CT images have missing images, gaps in images, repeated imaging through the patient in one Study, etc.\n\nLong answer to a short question.\n\n-Rich\n",
    "1009606": "May I know the relationship between\nRelationship between ****StudyInstanceUID, SeriesInstanceUID and SOPInstanceUID**** .\n\nI am looking for some explanation like :\n*One patient scan would be a StudyInstanceUID, and also a SeriesInstanceUID .\nMultiple Scan of Same Patient images collected in a sequence uniquely identified by SOPInstanceUID*\nOr could some one explain ",
    "1009827": "Well, in short terms:\n\nEach **StudyInstanceUID** contains one or more **SeriesInstanceUID**\nEach **SeriesInstanceUID** contrains one or more **SOPInstanceUID**\nEach **SOPInstanceUID** appended with `.dcm` is the filename for each DICOM file.\n\nOne Instance can be many series of scans. One scan will have many '.dcm' files."
  }
}