{
  "id": 109969,
  "title": "Metadata Usage Q&A",
  "url": "/competitions/rsna-intracranial-hemorrhage-detection/discussion/109969",
  "author_name": "Phil Culliton",
  "post_date": "2019-09-23T20:50:53.445000",
  "votes": 25,
  "comment_count": 24,
  "views": 0,
  "content": "<p>Hi all - feel free to ask questions about DICOM metadata usage in this thread, so we can have one central reference for everyone. I will confer with the host and get back to you with answers.</p>\n\n<p>We've already had <em>two cleared uses</em>, to my recollection:</p>\n\n<blockquote>\n  <p>You can read the DICOM metadata - especially \"Window Center\", \"Window Width\", \"Rescale Intercept\" and \"Rescale Slope\" - for data pre-processing.</p>\n</blockquote>\n\n<p>as answered <a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109479#latest-632585\">here</a> and <a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109479#latest-630263\">here</a>.</p>\n\n<blockquote>\n  <p>You can create a 3D construction using the series ID.</p>\n</blockquote>\n\n<p>as answered <a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109479#631888\">here</a>.</p>\n\n<p>Sorry if I've missed any questions that have been answered already, please feel free to link them here for everyone to check out. Of course, if you have posted questions that don't have answers yet, please link them here as well.</p>\n\n<p>Thanks for all of your questions so far! Thanks also to @wowfattie for the suggestion to have a central thread!</p>",
  "messages": [
    {
      "id": 632644,
      "postDate": "2019-09-23T20:50:53.447Z",
      "content": "<p>Hi all - feel free to ask questions about DICOM metadata usage in this thread, so we can have one central reference for everyone. I will confer with the host and get back to you with answers.</p>\n\n<p>We've already had <em>two cleared uses</em>, to my recollection:</p>\n\n<blockquote>\n  <p>You can read the DICOM metadata - especially \"Window Center\", \"Window Width\", \"Rescale Intercept\" and \"Rescale Slope\" - for data pre-processing.</p>\n</blockquote>\n\n<p>as answered <a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109479#latest-632585\">here</a> and <a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109479#latest-630263\">here</a>.</p>\n\n<blockquote>\n  <p>You can create a 3D construction using the series ID.</p>\n</blockquote>\n\n<p>as answered <a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109479#631888\">here</a>.</p>\n\n<p>Sorry if I've missed any questions that have been answered already, please feel free to link them here for everyone to check out. Of course, if you have posted questions that don't have answers yet, please link them here as well.</p>\n\n<p>Thanks for all of your questions so far! Thanks also to @wowfattie for the suggestion to have a central thread!</p>",
      "rawMarkdown": "Hi all - feel free to ask questions about DICOM metadata usage in this thread, so we can have one central reference for everyone. I will confer with the host and get back to you with answers.\n\nWe've already had *two cleared uses*, to my recollection:\n\n&gt; You can read the DICOM metadata - especially \"Window Center\", \"Window Width\", \"Rescale Intercept\" and \"Rescale Slope\" - for data pre-processing.\n\nas answered [here](https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109479#latest-632585) and [here](https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109479#latest-630263).\n\n&gt; You can create a 3D construction using the series ID.\n\nas answered [here](https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109479#631888).\n\nSorry if I've missed any questions that have been answered already, please feel free to link them here for everyone to check out. Of course, if you have posted questions that don't have answers yet, please link them here as well.\n\nThanks for all of your questions so far! Thanks also to @wowfattie for the suggestion to have a central thread!",
      "votes": 25
    },
    {
      "id": 641546,
      "postDate": "2019-10-04T19:33:03.570Z",
      "content": "<p>Recognizing the confusion generated by the statement “Submission predictions must be based entirely on the pixel data in the provided datasets” and the implications it has on metadata usage, the organizers have decided to retract this rule. The initial intent of the rule was for the algorithm not rely on metadata in order to limit over-fitting and maximize generalizability of the solution based on pixel data only. Given that (1) this generated confusion around metadata usage for preprocessing/model creation capabilities, (2) recognizing the metadata provided in the dataset is de-identified and the available fields do not contain information that can determine if an image contains intracranial hemorrhage, and (3) with the intent not to stifle creativity, the organizing committee has decided to retract this rule and allow all metadata to be used for model creation.</p>\n\n<p>Challenge Organizing Team</p>",
      "rawMarkdown": "Recognizing the confusion generated by the statement “Submission predictions must be based entirely on the pixel data in the provided datasets” and the implications it has on metadata usage, the organizers have decided to retract this rule. The initial intent of the rule was for the algorithm not rely on metadata in order to limit over-fitting and maximize generalizability of the solution based on pixel data only. Given that (1) this generated confusion around metadata usage for preprocessing/model creation capabilities, (2) recognizing the metadata provided in the dataset is de-identified and the available fields do not contain information that can determine if an image contains intracranial hemorrhage, and (3) with the intent not to stifle creativity, the organizing committee has decided to retract this rule and allow all metadata to be used for model creation.\n\nChallenge Organizing Team",
      "votes": 8
    },
    {
      "id": 633790,
      "postDate": "2019-09-25T12:26:35.657Z",
      "content": "<p>just to \"extend\" the question from Al-Khwârizmî - can we reconstruct the 3d image (using \"(0020, 0032) Image Position (Patient)\") and during <strong>prediction</strong> modify the slice prediction (for example by looking at nearby slices)?</p>",
      "rawMarkdown": "just to \"extend\" the question from Al-Khwârizmî - can we reconstruct the 3d image (using \"(0020, 0032) Image Position (Patient)\") and during **prediction** modify the slice prediction (for example by looking at nearby slices)?",
      "votes": 5,
      "replies": [
        {
          "id": 634809,
          "postDate": "2019-09-26T19:35:41.627Z",
          "content": "<p>Good question <a href=\"/steelrose\">@steelrose</a> - I'll check in with the host about it. My guess would be that anything after pre-processing will be off-limits, but I'll let you know when we have a ruling.</p>",
          "rawMarkdown": "Good question @steelrose - I'll check in with the host about it. My guess would be that anything after pre-processing will be off-limits, but I'll let you know when we have a ruling.",
          "votes": 3
        },
        {
          "id": 635112,
          "postDate": "2019-09-27T07:08:22.107Z",
          "content": "<p>It looks like the participants are inexorably refining the problem.</p>\n\n<p>I would like to add my two cents. If the goal of the RSNA is to advance the potential of machine learning relative to the state of the art or to human radiologists, then the machine should be given the same inputs that radiologists or prior work have. After all, what good is it to know that evaluating independent slices has 95% accuracy when in actual real-world practice a radiologist or a machine would have access to a complete scan? This would be just a waste of the inventive creativity of hundreds of participants for no useful result.</p>\n\n<p>At this early point, I suspect that some competitors are already using the 3D information from a complete scan while others are focusing on independent 2D slices. IMO, we need full clarity about what the problem statement is, and soon, so as not to go down false tracks.</p>\n\n<p>My suggestion is that <em>all</em> metadata that pertains to interpreting pixel data should be usable. That includes the scan id and order, the slice width, and the orientation: all and only the data produced by the CT scanner itself, and exactly the same playing field an agnostic human radiologist would have while viewing the slices. Externally derived metadata such as patient id and scanner id (if there is such) could legitimately be excluded, because there may be a correlation between hematomas and number of scans or the particular scanner, ones that the machine would use to \"cheat\".</p>\n\n<p>It also means that both the training and test sets should contain complete, typical scans. (I don't know whether this is the case for the test set.)</p>\n\n<p>Anyway, this is my assessment given that RSNA's goal is to move us closer to automated detection of hematomas from a cranial CT scan. If its goal is rather to find out what can be gleaned from a single 2D slice, then these opinions are moot. But I have a hard time imagining how producing an expert single slice classifier would be a useful outcome.</p>",
          "rawMarkdown": "It looks like the participants are inexorably refining the problem.\n\nI would like to add my two cents. If the goal of the RSNA is to advance the potential of machine learning relative to the state of the art or to human radiologists, then the machine should be given the same inputs that radiologists or prior work have. After all, what good is it to know that evaluating independent slices has 95% accuracy when in actual real-world practice a radiologist or a machine would have access to a complete scan? This would be just a waste of the inventive creativity of hundreds of participants for no useful result.\n\nAt this early point, I suspect that some competitors are already using the 3D information from a complete scan while others are focusing on independent 2D slices. IMO, we need full clarity about what the problem statement is, and soon, so as not to go down false tracks.\n\nMy suggestion is that *all* metadata that pertains to interpreting pixel data should be usable. That includes the scan id and order, the slice width, and the orientation: all and only the data produced by the CT scanner itself, and exactly the same playing field an agnostic human radiologist would have while viewing the slices. Externally derived metadata such as patient id and scanner id (if there is such) could legitimately be excluded, because there may be a correlation between hematomas and number of scans or the particular scanner, ones that the machine would use to \"cheat\".\n\nIt also means that both the training and test sets should contain complete, typical scans. (I don't know whether this is the case for the test set.)\n\nAnyway, this is my assessment given that RSNA's goal is to move us closer to automated detection of hematomas from a cranial CT scan. If its goal is rather to find out what can be gleaned from a single 2D slice, then these opinions are moot. But I have a hard time imagining how producing an expert single slice classifier would be a useful outcome.",
          "votes": 10
        },
        {
          "id": 635520,
          "postDate": "2019-09-27T18:48:33.290Z",
          "content": "<p>I agree completely.  After looking at the data, it would seem that there are 3D series for 2214 test and 19530 training series IDs (e.g. 3D scans).  In all unique series IDs, all the data either come from the training set or the test set (i.e. slice 5 is in the test set, but the rest are in training does not happen).  One of the issues is that the instance number is not in the DICOM metadata.  I've added that field after sorting on ImagePositionPatient to allow most converters (dcm2niix) to convert the image \"correctly\".   That said, keeping a 1-to-1 correspondence with the slices and the 3D volumes can be some tedious bookkeeping or difficult if the converter does anything like interpolation in the conversion process.   This bookkeeping/interpolation complicates many things likely due to the variable slice thickness within a series (2.5mm then 5mm) which is common.  </p>\n\n<p>Also, gantry tilt seems to be stripped off the header, which isn't the large of a problem, but may lead to changes in strategies (using an affine registration where a rigid would work if the data were tilt-corrected). </p>",
          "rawMarkdown": "I agree completely.  After looking at the data, it would seem that there are 3D series for 2214 test and 19530 training series IDs (e.g. 3D scans).  In all unique series IDs, all the data either come from the training set or the test set (i.e. slice 5 is in the test set, but the rest are in training does not happen).  One of the issues is that the instance number is not in the DICOM metadata.  I've added that field after sorting on ImagePositionPatient to allow most converters (dcm2niix) to convert the image \"correctly\".   That said, keeping a 1-to-1 correspondence with the slices and the 3D volumes can be some tedious bookkeeping or difficult if the converter does anything like interpolation in the conversion process.   This bookkeeping/interpolation complicates many things likely due to the variable slice thickness within a series (2.5mm then 5mm) which is common.  \n\nAlso, gantry tilt seems to be stripped off the header, which isn't the large of a problem, but may lead to changes in strategies (using an affine registration where a rigid would work if the data were tilt-corrected). ",
          "votes": 1
        },
        {
          "id": 636269,
          "postDate": "2019-09-29T06:48:14.277Z",
          "content": "<p>I think that for at least some cases 3D recon can be problematic. Look at PatientID  ID_beb49b44 for example. This case appears to consist of a mix of two scans (performed in quick succession) - this is easy to notice if you scroll through it in order of slice location (as defined by IPP). My immediate impression is that this is one of those cases where certain slices have artefact (streaks in lowest slices) and the techs rescan a partial volume. They then merge the originals and the rescanned into the same series.  As far as I can tell, a case like this cannot be reconstructed in 3D without manual removal of the duplicate slices. \nIt will be problematic to use a 3D approach if such cases will appear in the test set since the data is effectively 2 overlapping 3D volumes without enough information in the header to seperate them. Can anyone confirm?</p>",
          "rawMarkdown": "I think that for at least some cases 3D recon can be problematic. Look at PatientID  ID_beb49b44 for example. This case appears to consist of a mix of two scans (performed in quick succession) - this is easy to notice if you scroll through it in order of slice location (as defined by IPP). My immediate impression is that this is one of those cases where certain slices have artefact (streaks in lowest slices) and the techs rescan a partial volume. They then merge the originals and the rescanned into the same series.  As far as I can tell, a case like this cannot be reconstructed in 3D without manual removal of the duplicate slices. \nIt will be problematic to use a 3D approach if such cases will appear in the test set since the data is effectively 2 overlapping 3D volumes without enough information in the header to seperate them. Can anyone confirm?",
          "votes": 5
        },
        {
          "id": 640774,
          "postDate": "2019-10-04T08:11:54.733Z",
          "content": "<p>confirmed\nthere are about 300 of 19530 series in the training set with multiple scans mixed into the same series\nthere are about 5000 of 19530 series with varying slice thickness</p>",
          "rawMarkdown": "confirmed\nthere are about 300 of 19530 series in the training set with multiple scans mixed into the same series\nthere are about 5000 of 19530 series with varying slice thickness\n",
          "votes": 3
        }
      ]
    },
    {
      "id": 633378,
      "postDate": "2019-09-24T20:29:34.110Z",
      "content": "<p>Please, check my question here:\n<a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109953#latest-632564\">https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109953#latest-632564</a></p>",
      "rawMarkdown": "Please, check my question here:\nhttps://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109953#latest-632564",
      "votes": 3,
      "replies": [
        {
          "id": 634808,
          "postDate": "2019-09-26T19:34:23.683Z",
          "content": "<p>Thanks <a href=\"/amyaramine\">@amyaramine</a> - I put your questions to the host, and once we have a decision we'll post it here!</p>",
          "rawMarkdown": "Thanks @amyaramine - I put your questions to the host, and once we have a decision we'll post it here!",
          "votes": 1
        }
      ]
    },
    {
      "id": 644294,
      "postDate": "2019-10-08T15:37:59.043Z",
      "content": "<p>Suck it up :)\nReal life simulation. I'd say we should be the one used with the changes. \nregards,\nAdrian</p>",
      "rawMarkdown": "Suck it up :)\nReal life simulation. I'd say we should be the one used with the changes. \nregards,\nAdrian",
      "votes": 1
    },
    {
      "id": 634912,
      "postDate": "2019-09-27T00:08:24.787Z",
      "content": "<p><a href=\"/philculliton\">@philculliton</a>  Can you look at this as well: <a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109281#latest-633282\">https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109281#latest-633282</a> </p>\n\n<p>Are we able to reconstruct full 3D scans from the test set so to see which slices are adjacent to each other? </p>\n\n<p>Thanks.</p>",
      "rawMarkdown": "@philculliton  Can you look at this as well: https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109281#latest-633282 \n\nAre we able to reconstruct full 3D scans from the test set so to see which slices are adjacent to each other? \n\nThanks.",
      "votes": 2,
      "replies": [
        {
          "id": 639831,
          "postDate": "2019-10-03T15:59:21.223Z",
          "content": "<p>I have the same question.</p>",
          "rawMarkdown": "I have the same question.",
          "votes": 1
        }
      ]
    },
    {
      "id": 637525,
      "postDate": "2019-10-01T05:12:06.543Z",
      "content": "<p>We’ve confirmed with the host that <strong>metadata can be used for pre-processing of images, but they cannot be features in your model or used to change or label predictions post-modeling</strong>. This should help resolve many of the more specific questions that have been raised recently on what types of things can be done with the metadata.</p>",
      "rawMarkdown": "We’ve confirmed with the host that **metadata can be used for pre-processing of images, but they cannot be features in your model or used to change or label predictions post-modeling**. This should help resolve many of the more specific questions that have been raised recently on what types of things can be done with the metadata.",
      "replies": [
        {
          "id": 637842,
          "postDate": "2019-10-01T09:54:39.367Z",
          "content": "<p>I am still confused because if modelling is being done on reconstructed 3D scans then PatientID, study and slice number are necessarily being used in the model so why should these not be used explicitly as features in a model that uses 2D images?</p>",
          "rawMarkdown": "I am still confused because if modelling is being done on reconstructed 3D scans then PatientID, study and slice number are necessarily being used in the model so why should these not be used explicitly as features in a model that uses 2D images?",
          "votes": 3
        },
        {
          "id": 639830,
          "postDate": "2019-10-03T15:58:43.700Z",
          "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> I am also very confused by this clarification. According to your discussion post here \n<a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/110599\">https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/110599</a></p>\n\n<p>\"The clarification that we can provide here (on behalf of the host) is that the annotators were presented with a single image at a time to label. However, they did have access to the entire stack of images of the same study as they were annotating. So images were not necessarily annotated in isolation, but it is feasible or even expected (as is it is clinically) that the adjacent slices contribute to the interpretation process.\"</p>\n\n<p>The labeling process used the adjacent slices for annotation. Why can't we use post processing of the 2D images with the metadata, which is equivalent to 3D reconstructions of the scans? And this resembles clinical practice but the model can't be trained the same way?\nIf so, are 3D models allowed? if they are allowed can we use ImagePositionPatient[2] to get the order of the slides?</p>",
          "rawMarkdown": "@juliaelliott I am also very confused by this clarification. According to your discussion post here \nhttps://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/110599\n\n\"The clarification that we can provide here (on behalf of the host) is that the annotators were presented with a single image at a time to label. However, they did have access to the entire stack of images of the same study as they were annotating. So images were not necessarily annotated in isolation, but it is feasible or even expected (as is it is clinically) that the adjacent slices contribute to the interpretation process.\"\n\nThe labeling process used the adjacent slices for annotation. Why can't we use post processing of the 2D images with the metadata, which is equivalent to 3D reconstructions of the scans? And this resembles clinical practice but the model can't be trained the same way?\nIf so, are 3D models allowed? if they are allowed can we use ImagePositionPatient[2] to get the order of the slides?",
          "votes": 1
        },
        {
          "id": 640505,
          "postDate": "2019-10-04T04:33:15.957Z",
          "content": "<p>Can you confirm this:\n- You can build 3D-reconstructions for  e.g. brain extraction, in-slice translation and rotation, normalization as preprocessing \n- you are not allowed to build 3D-models or equivalently anything that uses more than one slice for prediction\n- essentially that means once the labels enter the training code you are not allowed to use anything other than the preprocessed single slice pixel data</p>",
          "rawMarkdown": "Can you confirm this:\n- You can build 3D-reconstructions for  e.g. brain extraction, in-slice translation and rotation, normalization as preprocessing \n- you are not allowed to build 3D-models or equivalently anything that uses more than one slice for prediction\n- essentially that means once the labels enter the training code you are not allowed to use anything other than the preprocessed single slice pixel data",
          "votes": 2
        },
        {
          "id": 640675,
          "postDate": "2019-10-04T06:52:03.553Z",
          "content": "<p>&gt; We’ve confirmed with the host that metadata can be used for pre-processing of images, but they cannot be features in your model or used to change or label predictions post-modeling.</p>\n\n<p>Like many others, I don't understand this one, please issue a clarification. Using metadata in pre-processing means using metadata. One can pre-process everything so it behaves as features or does label corrections. I believe nobody wants this competition to be about acrobatics on pre-processing so all valuable information is included. What does pre-processing even mean?</p>",
          "rawMarkdown": "&gt; We’ve confirmed with the host that metadata can be used for pre-processing of images, but they cannot be features in your model or used to change or label predictions post-modeling.\n\nLike many others, I don't understand this one, please issue a clarification. Using metadata in pre-processing means using metadata. One can pre-process everything so it behaves as features or does label corrections. I believe nobody wants this competition to be about acrobatics on pre-processing so all valuable information is included. What does pre-processing even mean?",
          "votes": 3
        },
        {
          "id": 640698,
          "postDate": "2019-10-04T07:13:11.200Z",
          "content": "<p>Very true, any preprocessing using metadata and multiple slices leads to leakage of the metadata and adjacent slices information into the model. I don't want to have to implicitly select for a preprocessing step that maximises this information leakage.\nI it is a bit late now to change the rules, a lot of people have started to invest time and money in preprocessing and metadata usage, it makes total sense from a medical perspective, so why don't you just allow using the data you provide?</p>",
          "rawMarkdown": "Very true, any preprocessing using metadata and multiple slices leads to leakage of the metadata and adjacent slices information into the model. I don't want to have to implicitly select for a preprocessing step that maximises this information leakage.\nI it is a bit late now to change the rules, a lot of people have started to invest time and money in preprocessing and metadata usage, it makes total sense from a medical perspective, so why don't you just allow using the data you provide?",
          "votes": 3
        },
        {
          "id": 640754,
          "postDate": "2019-10-04T07:55:00.717Z",
          "content": "<p>&gt; What does pre-processing even mean?</p>\n\n<p>+1</p>",
          "rawMarkdown": "&gt; What does pre-processing even mean?\n\n+1",
          "votes": 1
        },
        {
          "id": 641387,
          "postDate": "2019-10-04T16:36:22.860Z",
          "content": "<p>I think <a href=\"/philculliton\">@philculliton</a> mentioned here: <a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109281#635537\">https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109281#635537</a></p>\n\n<blockquote>\n  <p>As clarified in the Metadata Q&amp;A, creating a 3D image of a given study using the metadata is within the bounds of the competition.</p>\n</blockquote>\n\n<p>It looks like to me the the competition is about classifying hemorrhages based on CT scans. So as Phil mentioned above, in my opinion, yes 3D images can be used to help you label the particular slice of the scan. </p>\n\n<p>And as Julia mentioned, in my opinion, using metadata such as for example, patient ID number as one of your features to predict the hemorrhage because maybe a patient is in the train set and also in the test set, does not generalize to real world algorithms so that should not be a factor to build the model. </p>\n\n<p>Just my 2 cents. 👀 </p>",
          "rawMarkdown": "I think @philculliton mentioned here: https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109281#635537\n&gt; As clarified in the Metadata Q&amp;A, creating a 3D image of a given study using the metadata is within the bounds of the competition.\n\nIt looks like to me the the competition is about classifying hemorrhages based on CT scans. So as Phil mentioned above, in my opinion, yes 3D images can be used to help you label the particular slice of the scan. \n\nAnd as Julia mentioned, in my opinion, using metadata such as for example, patient ID number as one of your features to predict the hemorrhage because maybe a patient is in the train set and also in the test set, does not generalize to real world algorithms so that should not be a factor to build the model. \n \nJust my 2 cents. 👀 \n\n"
        },
        {
          "id": 641444,
          "postDate": "2019-10-04T17:38:15.100Z",
          "content": "<p><a href=\"/bopengiowa\">@bopengiowa</a> , please note a quote from the data description:</p>\n\n<p>&gt; You will notice some PatientIDs represented in both the stage 1 train and test sets. This is known and intentional. However, there will be no crossover of PatientIDs into stage 2 test.</p>\n\n<p>I also disagree with your statement in general, all the metadata is actually available to a doctor, and in real life application you will have it for your model. </p>\n\n<p>But my point is even beyond this. I don't mind if the rules do not match the real world, there can be a reason, may be some kind of hidden limitation, that forces them to prohibit the metadata. But the rules need to make sense. You can not allow using 3D scans and not allow taking information from other slices on later stages, such as merging features in the middle or post-processing. </p>",
          "rawMarkdown": "@bopengiowa , please note a quote from the data description:\n\n&gt; You will notice some PatientIDs represented in both the stage 1 train and test sets. This is known and intentional. However, there will be no crossover of PatientIDs into stage 2 test.\n\nI also disagree with your statement in general, all the metadata is actually available to a doctor, and in real life application you will have it for your model. \n\nBut my point is even beyond this. I don't mind if the rules do not match the real world, there can be a reason, may be some kind of hidden limitation, that forces them to prohibit the metadata. But the rules need to make sense. You can not allow using 3D scans and not allow taking information from other slices on later stages, such as merging features in the middle or post-processing. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 647651,
      "postDate": "2019-10-13T03:26:03.577Z",
      "content": "<p>How do we interpret the image orientation vector? I created this notebook to find the anomalous orientations however i could not find any correlation between the vector and the observed orientation</p>\n\n<p><a href=\"https://www.kaggle.com/nikperi/image-orientation?scriptVersionId=21878745\">https://www.kaggle.com/nikperi/image-orientation?scriptVersionId=21878745</a></p>",
      "rawMarkdown": "How do we interpret the image orientation vector? I created this notebook to find the anomalous orientations however i could not find any correlation between the vector and the observed orientation\n\nhttps://www.kaggle.com/nikperi/image-orientation?scriptVersionId=21878745"
    },
    {
      "id": 635526,
      "postDate": "2019-09-27T18:58:43.313Z",
      "content": "<p>To extend form <a href=\"/steelrose\">@steelrose</a> question about 3D models, I noted in this training/test set: \" In all unique series IDs, all the data either come from the training set or the test set (i.e. slice 5 is in the test set, but the rest are in training does not happen)\".  Can we assume the future set of series have the full data in there (not removing slices or if one slice is in validation all are in validation)?</p>",
      "rawMarkdown": "To extend form @steelrose question about 3D models, I noted in this training/test set: \" In all unique series IDs, all the data either come from the training set or the test set (i.e. slice 5 is in the test set, but the rest are in training does not happen)\".  Can we assume the future set of series have the full data in there (not removing slices or if one slice is in validation all are in validation)?",
      "replies": [
        {
          "id": 635533,
          "postDate": "2019-09-27T19:11:35.827Z",
          "content": "<p>Hi <a href=\"/muschellij2\">@muschellij2</a> - good question. You can assume that if one slice from a series is in validation, then all slices from that series are in validation.</p>",
          "rawMarkdown": "Hi @muschellij2 - good question. You can assume that if one slice from a series is in validation, then all slices from that series are in validation."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 641546,
      "author_name": "Luciano Prevedello",
      "author_url": "",
      "post_date": "2019-10-04T19:33:03.570000",
      "content": "<p>Recognizing the confusion generated by the statement “Submission predictions must be based entirely on the pixel data in the provided datasets” and the implications it has on metadata usage, the organizers have decided to retract this rule. The initial intent of the rule was for the algorithm not rely on metadata in order to limit over-fitting and maximize generalizability of the solution based on pixel data only. Given that (1) this generated confusion around metadata usage for preprocessing/model creation capabilities, (2) recognizing the metadata provided in the dataset is de-identified and the available fields do not contain information that can determine if an image contains intracranial hemorrhage, and (3) with the intent not to stifle creativity, the organizing committee has decided to retract this rule and allow all metadata to be used for model creation.</p>\n\n<p>Challenge Organizing Team</p>",
      "votes": 8,
      "replies": []
    },
    {
      "id": 633790,
      "author_name": "steelrose",
      "author_url": "",
      "post_date": "2019-09-25T12:26:35.657000",
      "content": "<p>just to \"extend\" the question from Al-Khwârizmî - can we reconstruct the 3d image (using \"(0020, 0032) Image Position (Patient)\") and during <strong>prediction</strong> modify the slice prediction (for example by looking at nearby slices)?</p>",
      "votes": 5,
      "replies": [
        {
          "id": 634809,
          "author_name": "Phil Culliton",
          "author_url": "",
          "post_date": "2019-09-26T19:35:41.627000",
          "content": "<p>Good question <a href=\"/steelrose\">@steelrose</a> - I'll check in with the host about it. My guess would be that anything after pre-processing will be off-limits, but I'll let you know when we have a ruling.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 635112,
          "author_name": "Malcolm McLean",
          "author_url": "",
          "post_date": "2019-09-27T07:08:22.107000",
          "content": "<p>It looks like the participants are inexorably refining the problem.</p>\n\n<p>I would like to add my two cents. If the goal of the RSNA is to advance the potential of machine learning relative to the state of the art or to human radiologists, then the machine should be given the same inputs that radiologists or prior work have. After all, what good is it to know that evaluating independent slices has 95% accuracy when in actual real-world practice a radiologist or a machine would have access to a complete scan? This would be just a waste of the inventive creativity of hundreds of participants for no useful result.</p>\n\n<p>At this early point, I suspect that some competitors are already using the 3D information from a complete scan while others are focusing on independent 2D slices. IMO, we need full clarity about what the problem statement is, and soon, so as not to go down false tracks.</p>\n\n<p>My suggestion is that <em>all</em> metadata that pertains to interpreting pixel data should be usable. That includes the scan id and order, the slice width, and the orientation: all and only the data produced by the CT scanner itself, and exactly the same playing field an agnostic human radiologist would have while viewing the slices. Externally derived metadata such as patient id and scanner id (if there is such) could legitimately be excluded, because there may be a correlation between hematomas and number of scans or the particular scanner, ones that the machine would use to \"cheat\".</p>\n\n<p>It also means that both the training and test sets should contain complete, typical scans. (I don't know whether this is the case for the test set.)</p>\n\n<p>Anyway, this is my assessment given that RSNA's goal is to move us closer to automated detection of hematomas from a cranial CT scan. If its goal is rather to find out what can be gleaned from a single 2D slice, then these opinions are moot. But I have a hard time imagining how producing an expert single slice classifier would be a useful outcome.</p>",
          "votes": 10,
          "replies": []
        },
        {
          "id": 635520,
          "author_name": "John M",
          "author_url": "",
          "post_date": "2019-09-27T18:48:33.290000",
          "content": "<p>I agree completely.  After looking at the data, it would seem that there are 3D series for 2214 test and 19530 training series IDs (e.g. 3D scans).  In all unique series IDs, all the data either come from the training set or the test set (i.e. slice 5 is in the test set, but the rest are in training does not happen).  One of the issues is that the instance number is not in the DICOM metadata.  I've added that field after sorting on ImagePositionPatient to allow most converters (dcm2niix) to convert the image \"correctly\".   That said, keeping a 1-to-1 correspondence with the slices and the 3D volumes can be some tedious bookkeeping or difficult if the converter does anything like interpolation in the conversion process.   This bookkeeping/interpolation complicates many things likely due to the variable slice thickness within a series (2.5mm then 5mm) which is common.  </p>\n\n<p>Also, gantry tilt seems to be stripped off the header, which isn't the large of a problem, but may lead to changes in strategies (using an affine registration where a rigid would work if the data were tilt-corrected). </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 636269,
          "author_name": "Soren",
          "author_url": "",
          "post_date": "2019-09-29T06:48:14.277000",
          "content": "<p>I think that for at least some cases 3D recon can be problematic. Look at PatientID  ID_beb49b44 for example. This case appears to consist of a mix of two scans (performed in quick succession) - this is easy to notice if you scroll through it in order of slice location (as defined by IPP). My immediate impression is that this is one of those cases where certain slices have artefact (streaks in lowest slices) and the techs rescan a partial volume. They then merge the originals and the rescanned into the same series.  As far as I can tell, a case like this cannot be reconstructed in 3D without manual removal of the duplicate slices. \nIt will be problematic to use a 3D approach if such cases will appear in the test set since the data is effectively 2 overlapping 3D volumes without enough information in the header to seperate them. Can anyone confirm?</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 640774,
          "author_name": "Victor",
          "author_url": "",
          "post_date": "2019-10-04T08:11:54.733000",
          "content": "<p>confirmed\nthere are about 300 of 19530 series in the training set with multiple scans mixed into the same series\nthere are about 5000 of 19530 series with varying slice thickness</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 633378,
      "author_name": "Al-Khwârizmî",
      "author_url": "",
      "post_date": "2019-09-24T20:29:34.110000",
      "content": "<p>Please, check my question here:\n<a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109953#latest-632564\">https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109953#latest-632564</a></p>",
      "votes": 3,
      "replies": [
        {
          "id": 634808,
          "author_name": "Phil Culliton",
          "author_url": "",
          "post_date": "2019-09-26T19:34:23.683000",
          "content": "<p>Thanks <a href=\"/amyaramine\">@amyaramine</a> - I put your questions to the host, and once we have a decision we'll post it here!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 644294,
      "author_name": "Adrian Zinovei",
      "author_url": "",
      "post_date": "2019-10-08T15:37:59.043000",
      "content": "<p>Suck it up :)\nReal life simulation. I'd say we should be the one used with the changes. \nregards,\nAdrian</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 634912,
      "author_name": "Bo Peng",
      "author_url": "",
      "post_date": "2019-09-27T00:08:24.787000",
      "content": "<p><a href=\"/philculliton\">@philculliton</a>  Can you look at this as well: <a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109281#latest-633282\">https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109281#latest-633282</a> </p>\n\n<p>Are we able to reconstruct full 3D scans from the test set so to see which slices are adjacent to each other? </p>\n\n<p>Thanks.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 639831,
          "author_name": "Maria Wellen",
          "author_url": "",
          "post_date": "2019-10-03T15:59:21.223000",
          "content": "<p>I have the same question.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 637525,
      "author_name": "Julia Elliott",
      "author_url": "",
      "post_date": "2019-10-01T05:12:06.543000",
      "content": "<p>We’ve confirmed with the host that <strong>metadata can be used for pre-processing of images, but they cannot be features in your model or used to change or label predictions post-modeling</strong>. This should help resolve many of the more specific questions that have been raised recently on what types of things can be done with the metadata.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 637842,
          "author_name": "Alison Davey",
          "author_url": "",
          "post_date": "2019-10-01T09:54:39.367000",
          "content": "<p>I am still confused because if modelling is being done on reconstructed 3D scans then PatientID, study and slice number are necessarily being used in the model so why should these not be used explicitly as features in a model that uses 2D images?</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 639830,
          "author_name": "Maria Wellen",
          "author_url": "",
          "post_date": "2019-10-03T15:58:43.700000",
          "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> I am also very confused by this clarification. According to your discussion post here \n<a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/110599\">https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/110599</a></p>\n\n<p>\"The clarification that we can provide here (on behalf of the host) is that the annotators were presented with a single image at a time to label. However, they did have access to the entire stack of images of the same study as they were annotating. So images were not necessarily annotated in isolation, but it is feasible or even expected (as is it is clinically) that the adjacent slices contribute to the interpretation process.\"</p>\n\n<p>The labeling process used the adjacent slices for annotation. Why can't we use post processing of the 2D images with the metadata, which is equivalent to 3D reconstructions of the scans? And this resembles clinical practice but the model can't be trained the same way?\nIf so, are 3D models allowed? if they are allowed can we use ImagePositionPatient[2] to get the order of the slides?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 640505,
          "author_name": "Victor",
          "author_url": "",
          "post_date": "2019-10-04T04:33:15.957000",
          "content": "<p>Can you confirm this:\n- You can build 3D-reconstructions for  e.g. brain extraction, in-slice translation and rotation, normalization as preprocessing \n- you are not allowed to build 3D-models or equivalently anything that uses more than one slice for prediction\n- essentially that means once the labels enter the training code you are not allowed to use anything other than the preprocessed single slice pixel data</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 640675,
          "author_name": "nosound",
          "author_url": "",
          "post_date": "2019-10-04T06:52:03.553000",
          "content": "<p>&gt; We’ve confirmed with the host that metadata can be used for pre-processing of images, but they cannot be features in your model or used to change or label predictions post-modeling.</p>\n\n<p>Like many others, I don't understand this one, please issue a clarification. Using metadata in pre-processing means using metadata. One can pre-process everything so it behaves as features or does label corrections. I believe nobody wants this competition to be about acrobatics on pre-processing so all valuable information is included. What does pre-processing even mean?</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 640698,
          "author_name": "Victor",
          "author_url": "",
          "post_date": "2019-10-04T07:13:11.200000",
          "content": "<p>Very true, any preprocessing using metadata and multiple slices leads to leakage of the metadata and adjacent slices information into the model. I don't want to have to implicitly select for a preprocessing step that maximises this information leakage.\nI it is a bit late now to change the rules, a lot of people have started to invest time and money in preprocessing and metadata usage, it makes total sense from a medical perspective, so why don't you just allow using the data you provide?</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 640754,
          "author_name": "cherring",
          "author_url": "",
          "post_date": "2019-10-04T07:55:00.717000",
          "content": "<p>&gt; What does pre-processing even mean?</p>\n\n<p>+1</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 641387,
          "author_name": "Bo Peng",
          "author_url": "",
          "post_date": "2019-10-04T16:36:22.860000",
          "content": "<p>I think <a href=\"/philculliton\">@philculliton</a> mentioned here: <a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109281#635537\">https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109281#635537</a></p>\n\n<blockquote>\n  <p>As clarified in the Metadata Q&amp;A, creating a 3D image of a given study using the metadata is within the bounds of the competition.</p>\n</blockquote>\n\n<p>It looks like to me the the competition is about classifying hemorrhages based on CT scans. So as Phil mentioned above, in my opinion, yes 3D images can be used to help you label the particular slice of the scan. </p>\n\n<p>And as Julia mentioned, in my opinion, using metadata such as for example, patient ID number as one of your features to predict the hemorrhage because maybe a patient is in the train set and also in the test set, does not generalize to real world algorithms so that should not be a factor to build the model. </p>\n\n<p>Just my 2 cents. 👀 </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 641444,
          "author_name": "nosound",
          "author_url": "",
          "post_date": "2019-10-04T17:38:15.100000",
          "content": "<p><a href=\"/bopengiowa\">@bopengiowa</a> , please note a quote from the data description:</p>\n\n<p>&gt; You will notice some PatientIDs represented in both the stage 1 train and test sets. This is known and intentional. However, there will be no crossover of PatientIDs into stage 2 test.</p>\n\n<p>I also disagree with your statement in general, all the metadata is actually available to a doctor, and in real life application you will have it for your model. </p>\n\n<p>But my point is even beyond this. I don't mind if the rules do not match the real world, there can be a reason, may be some kind of hidden limitation, that forces them to prohibit the metadata. But the rules need to make sense. You can not allow using 3D scans and not allow taking information from other slices on later stages, such as merging features in the middle or post-processing. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 647651,
      "author_name": "Nikhil Peri",
      "author_url": "",
      "post_date": "2019-10-13T03:26:03.577000",
      "content": "<p>How do we interpret the image orientation vector? I created this notebook to find the anomalous orientations however i could not find any correlation between the vector and the observed orientation</p>\n\n<p><a href=\"https://www.kaggle.com/nikperi/image-orientation?scriptVersionId=21878745\">https://www.kaggle.com/nikperi/image-orientation?scriptVersionId=21878745</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 635526,
      "author_name": "John M",
      "author_url": "",
      "post_date": "2019-09-27T18:58:43.313000",
      "content": "<p>To extend form <a href=\"/steelrose\">@steelrose</a> question about 3D models, I noted in this training/test set: \" In all unique series IDs, all the data either come from the training set or the test set (i.e. slice 5 is in the test set, but the rest are in training does not happen)\".  Can we assume the future set of series have the full data in there (not removing slices or if one slice is in validation all are in validation)?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 635533,
          "author_name": "Phil Culliton",
          "author_url": "",
          "post_date": "2019-09-27T19:11:35.827000",
          "content": "<p>Hi <a href=\"/muschellij2\">@muschellij2</a> - good question. You can assume that if one slice from a series is in validation, then all slices from that series are in validation.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "632644": "Hi all - feel free to ask questions about DICOM metadata usage in this thread, so we can have one central reference for everyone. I will confer with the host and get back to you with answers.\n\nWe've already had *two cleared uses*, to my recollection:\n\n&gt; You can read the DICOM metadata - especially \"Window Center\", \"Window Width\", \"Rescale Intercept\" and \"Rescale Slope\" - for data pre-processing.\n\nas answered [here](https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109479#latest-632585) and [here](https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109479#latest-630263).\n\n&gt; You can create a 3D construction using the series ID.\n\nas answered [here](https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109479#631888).\n\nSorry if I've missed any questions that have been answered already, please feel free to link them here for everyone to check out. Of course, if you have posted questions that don't have answers yet, please link them here as well.\n\nThanks for all of your questions so far! Thanks also to @wowfattie for the suggestion to have a central thread!",
    "641546": "Recognizing the confusion generated by the statement “Submission predictions must be based entirely on the pixel data in the provided datasets” and the implications it has on metadata usage, the organizers have decided to retract this rule. The initial intent of the rule was for the algorithm not rely on metadata in order to limit over-fitting and maximize generalizability of the solution based on pixel data only. Given that (1) this generated confusion around metadata usage for preprocessing/model creation capabilities, (2) recognizing the metadata provided in the dataset is de-identified and the available fields do not contain information that can determine if an image contains intracranial hemorrhage, and (3) with the intent not to stifle creativity, the organizing committee has decided to retract this rule and allow all metadata to be used for model creation.\n\nChallenge Organizing Team",
    "633790": "just to \"extend\" the question from Al-Khwârizmî - can we reconstruct the 3d image (using \"(0020, 0032) Image Position (Patient)\") and during **prediction** modify the slice prediction (for example by looking at nearby slices)?",
    "633378": "Please, check my question here:\nhttps://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109953#latest-632564",
    "644294": "Suck it up :)\nReal life simulation. I'd say we should be the one used with the changes. \nregards,\nAdrian",
    "634912": "@philculliton  Can you look at this as well: https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/109281#latest-633282 \n\nAre we able to reconstruct full 3D scans from the test set so to see which slices are adjacent to each other? \n\nThanks.",
    "637525": "We’ve confirmed with the host that **metadata can be used for pre-processing of images, but they cannot be features in your model or used to change or label predictions post-modeling**. This should help resolve many of the more specific questions that have been raised recently on what types of things can be done with the metadata.",
    "647651": "How do we interpret the image orientation vector? I created this notebook to find the anomalous orientations however i could not find any correlation between the vector and the observed orientation\n\nhttps://www.kaggle.com/nikperi/image-orientation?scriptVersionId=21878745",
    "635526": "To extend form @steelrose question about 3D models, I noted in this training/test set: \" In all unique series IDs, all the data either come from the training set or the test set (i.e. slice 5 is in the test set, but the rest are in training does not happen)\".  Can we assume the future set of series have the full data in there (not removing slices or if one slice is in validation all are in validation)?"
  }
}