{
  "id": 596296,
  "title": "Modality not provided in test.csv?",
  "url": "/competitions/rsna-intracranial-aneurysm-detection/discussion/596296",
  "author_name": "YYama",
  "post_date": "2025-08-02T18:11:30.071000",
  "votes": 11,
  "comment_count": 12,
  "views": 0,
  "content": "<p>I noticed that there is no Modality information in test.csv.<br>\nThe DICOM Modality contains only two types: \"CT\" or \"MR\". (It seems that the test DICOMs don't have Study Description either.)<br>\nHowever, MR has three types: MRA, T2WI, and T1CE, and the optimal conditions during inference may differ.</p>\n<p>Naturally, it's possible to distinguish these using various algorithms, but since sequence is metadata that can be easily obtained during imaging, it would be desirable to use it during inference without any special processing or restrictions.</p>\n<p>Would it be possible to provide Modality in test.csv similar to train.csv?<br>\nOr is there already some other way to use this information?</p>",
  "messages": [
    {
      "id": 3262020,
      "postDate": "2025-08-02T18:11:30.070Z",
      "content": "<p>I noticed that there is no Modality information in test.csv.<br>\nThe DICOM Modality contains only two types: \"CT\" or \"MR\". (It seems that the test DICOMs don't have Study Description either.)<br>\nHowever, MR has three types: MRA, T2WI, and T1CE, and the optimal conditions during inference may differ.</p>\n<p>Naturally, it's possible to distinguish these using various algorithms, but since sequence is metadata that can be easily obtained during imaging, it would be desirable to use it during inference without any special processing or restrictions.</p>\n<p>Would it be possible to provide Modality in test.csv similar to train.csv?<br>\nOr is there already some other way to use this information?</p>",
      "rawMarkdown": "I noticed that there is no Modality information in test.csv.\nThe DICOM Modality contains only two types: \"CT\" or \"MR\". (It seems that the test DICOMs don't have Study Description either.)\nHowever, MR has three types: MRA, T2WI, and T1CE, and the optimal conditions during inference may differ.\n\nNaturally, it's possible to distinguish these using various algorithms, but since sequence is metadata that can be easily obtained during imaging, it would be desirable to use it during inference without any special processing or restrictions.\n\nWould it be possible to provide Modality in test.csv similar to train.csv?\nOr is there already some other way to use this information?",
      "votes": 11
    },
    {
      "id": 3265689,
      "postDate": "2025-08-07T20:34:56.900Z",
      "content": "<p>Thanks for your comment. It may surprise you to know that the type of MR series cannot be easily inferred from DICOM header metadata in real clinical datasets. Often we rely on series description, which is a free text field that can and often is manually edited by MR technologists at the time of acquisition. For obvious reasons this field has to be removed to prevent PHI leak, but it’s not that reliable anyway. Trying to infer MR series type based on other scan parameters available in the DICOM header is also not straightforward since these parameters vary widely across scanner makes/models. </p>\n<p>If you believe that series type identification is important for your approach, then one option is to use a pixel data-based approach. See for example: <a href=\"https://pmc.ncbi.nlm.nih.gov/articles/PMC10831512/\" target=\"_blank\">https://pmc.ncbi.nlm.nih.gov/articles/PMC10831512/</a></p>",
      "rawMarkdown": "Thanks for your comment. It may surprise you to know that the type of MR series cannot be easily inferred from DICOM header metadata in real clinical datasets. Often we rely on series description, which is a free text field that can and often is manually edited by MR technologists at the time of acquisition. For obvious reasons this field has to be removed to prevent PHI leak, but it’s not that reliable anyway. Trying to infer MR series type based on other scan parameters available in the DICOM header is also not straightforward since these parameters vary widely across scanner makes/models. \n\nIf you believe that series type identification is important for your approach, then one option is to use a pixel data-based approach. See for example: https://pmc.ncbi.nlm.nih.gov/articles/PMC10831512/",
      "votes": 2,
      "replies": [
        {
          "id": 3265711,
          "postDate": "2025-08-07T21:22:16.373Z",
          "content": "<p>Thank you for your comment!</p>\n<p>As a radiologist in Japan, I’m familiar with that situation. However, I’m a bit surprised.<br>\nIn the hospitals I’ve worked at so far, while there were naturally some variations in notation within the series description, it wasn’t unreliable, and I was able to standardize it through rule-based NLP processing.<br>\nThis might be because the hospitals I’ve worked at had solid regulations and MR technologists diligently input the information. (There might also be differences between Japan and the United States.)</p>\n<p>And I’ve read that paper.<br>\nThe opening statement in the introduction that DICOM headers are unreliable didn’t quite resonate with me at the time, but now I understand the circumstances behind it. Very interesting.​​​​​​​​​​​​​​​​</p>",
          "rawMarkdown": "Thank you for your comment!\n\nAs a radiologist in Japan, I’m familiar with that situation. However, I’m a bit surprised.\nIn the hospitals I’ve worked at so far, while there were naturally some variations in notation within the series description, it wasn’t unreliable, and I was able to standardize it through rule-based NLP processing.\nThis might be because the hospitals I’ve worked at had solid regulations and MR technologists diligently input the information. (There might also be differences between Japan and the United States.)\n\nAnd I’ve read that paper.\nThe opening statement in the introduction that DICOM headers are unreliable didn’t quite resonate with me at the time, but now I understand the circumstances behind it. Very interesting.​​​​​​​​​​​​​​​​",
          "votes": 2,
          "replies": [
            {
              "id": 3265745,
              "postDate": "2025-08-07T22:41:31.820Z",
              "content": "<p>I think I would like to come work in Japan!!! 😂 Series descriptions can be quite unreliable here. Considering a multi-national dataset like this one, I think it would be extremely difficult to design an universal rules-based NLP for series description. Just to give you one example, some series in this dataset had the patients (or at least someone's) name in the series description before we wiped it 🤦.</p>",
              "rawMarkdown": "I think I would like to come work in Japan!!! 😂 Series descriptions can be quite unreliable here. Considering a multi-national dataset like this one, I think it would be extremely difficult to design an universal rules-based NLP for series description. Just to give you one example, some series in this dataset had the patients (or at least someone's) name in the series description before we wiped it 🤦.",
              "votes": 3
            },
            {
              "id": 3265750,
              "postDate": "2025-08-07T22:52:58.103Z",
              "content": "<p>I think that was just within my narrow range of observation, and it’s probably not like that throughout all of Japan lol. Actually, when I was doing sequence classification myself, there were differences in descriptions depending on the imaging equipment, so I implemented rule-based processing myself. If patient names are included in the descriptions, removing PHI must be quite challenging. I’m grateful that you overcame these difficulties and organized this competition.</p>",
              "rawMarkdown": "I think that was just within my narrow range of observation, and it’s probably not like that throughout all of Japan lol. Actually, when I was doing sequence classification myself, there were differences in descriptions depending on the imaging equipment, so I implemented rule-based processing myself. If patient names are included in the descriptions, removing PHI must be quite challenging. I’m grateful that you overcame these difficulties and organized this competition.",
              "votes": 2
            }
          ]
        },
        {
          "id": 3274179,
          "postDate": "2025-08-24T08:59:18.110Z",
          "content": "<p><a href=\"https://www.kaggle.com/evancalabrese\" target=\"_blank\">@evancalabrese</a> <a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> <br>\nHello, I have another question about the metadata in the hidden testset. When processing dicom, I perform a basic validation method by checking</p>\n<ul>\n<li>SeriesInstanceUID, InstanceNumber, Modality, ImagePositionPatient, ImageOrientationPatient in dicom</li>\n<li>length of ImagePositionPatient &gt;=3 and length of ImageOrientationPatient &gt;=6</li>\n</ul>\n<pre><code> ():\n        is_valid_dicom = \n         meta  [, ]:\n         ]:\n            is_valid_dicom = is_valid_dicom  meta  dicom_header\n\n            dicom_header  (dicom_header.ImageOrientationPatient) &lt; :\n            is_valid_dicom = \n\n            dicom_header  (dicom_header.ImagePositionPatient) &lt; :\n            is_valid_dicom = \n         is_valid_dicom\n</code></pre>\n<p>If the dicom file does not pass the check, it will be removed. This method works well on training data but I encountered a Series in the hidden testset that all of the dicoms are removed, which means all of the dicom of that series is invalid and lead to \"Notebook threw Exception\". It seems that this occurs after the dataset update. <a href=\"https://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/600658#3274261\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/600658#3274261</a></p>\n<p>Is it expected that some series does not contain ImagePositionPatient, InstanceNumber, ImageOrientationPatient or Modality in dicom metadata? I am not a radiologist but GPT5 tells me that these are must information. I don't think such a gap in dicom files between train and test data should occur and we cannot reconstruct the dicoms into ordered 3d volume without these information. Or is the submission system broken?</p>",
          "rawMarkdown": "@evancalabrese @ryanholbrook \nHello, I have another question about the metadata in the hidden testset. When processing dicom, I perform a basic validation method by checking\n- SeriesInstanceUID, InstanceNumber, Modality, ImagePositionPatient, ImageOrientationPatient in dicom\n- length of ImagePositionPatient >=3 and length of ImageOrientationPatient >=6\n\n```\ndef check_dicom(dicom_header):\n        is_valid_dicom = True\n        for meta in [\"SeriesInstanceUID\", \"InstanceNumber\"]:\n         ]:\n            is_valid_dicom = is_valid_dicom and meta in dicom_header\n\n        if \"ImageOrientationPatient\" not in dicom_header or len(dicom_header.ImageOrientationPatient) < 6:\n            is_valid_dicom = False\n\n        if \"ImagePositionPatient\" not in dicom_header or len(dicom_header.ImagePositionPatient) < 3:\n            is_valid_dicom = False\n        return is_valid_dicom\n\n```\n\nIf the dicom file does not pass the check, it will be removed. This method works well on training data but I encountered a Series in the hidden testset that all of the dicoms are removed, which means all of the dicom of that series is invalid and lead to \"Notebook threw Exception\". It seems that this occurs after the dataset update. https://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/600658#3274261\n\nIs it expected that some series does not contain ImagePositionPatient, InstanceNumber, ImageOrientationPatient or Modality in dicom metadata? I am not a radiologist but GPT5 tells me that these are must information. I don't think such a gap in dicom files between train and test data should occur and we cannot reconstruct the dicoms into ordered 3d volume without these information. Or is the submission system broken?",
          "votes": 2,
          "replies": [
            {
              "id": 3274834,
              "postDate": "2025-08-25T13:54:30.470Z",
              "content": "<p>These tags should all be present in the test set. I think there is currently an issue with the test set that may be causing this error. Hopefully we will have a fix soon.</p>",
              "rawMarkdown": "These tags should all be present in the test set. I think there is currently an issue with the test set that may be causing this error. Hopefully we will have a fix soon."
            }
          ]
        }
      ]
    },
    {
      "id": 3262349,
      "postDate": "2025-08-03T12:54:08.900Z",
      "content": "<p>Noting that modality is also available in the dcm files both for train and test sets, though I have not tried to use it that way yet. Modality is in the list of tags provided for the test set per the \"Data\" section of the competition.</p>",
      "rawMarkdown": "Noting that modality is also available in the dcm files both for train and test sets, though I have not tried to use it that way yet. Modality is in the list of tags provided for the test set per the \"Data\" section of the competition.",
      "replies": [
        {
          "id": 3262353,
          "postDate": "2025-08-03T13:01:33.697Z",
          "content": "<p>Thank you for your comment! As I mentioned in my post, I’m already aware of that. My question is whether it’s possible to obtain the type of MR sequence—of which there are three main kinds—as metadata. For more details, please refer to my original post.</p>",
          "rawMarkdown": "Thank you for your comment! As I mentioned in my post, I’m already aware of that. My question is whether it’s possible to obtain the type of MR sequence—of which there are three main kinds—as metadata. For more details, please refer to my original post."
        }
      ]
    },
    {
      "id": 3262058,
      "postDate": "2025-08-02T19:54:53.240Z",
      "content": "<p>You can find Modality in metadata.</p>\n<p>EDIT: dcm.Modality</p>",
      "rawMarkdown": "You can find Modality in metadata.\n\nEDIT: dcm.Modality",
      "replies": [
        {
          "id": 3262112,
          "postDate": "2025-08-02T23:53:57.393Z",
          "content": "<p>As I mentioned above, DICOM Modality tags don’t contain the three types of MR information, so my question is whether we can obtain information like what’s in train.csv during submission?</p>",
          "rawMarkdown": "As I mentioned above, DICOM Modality tags don’t contain the three types of MR information, so my question is whether we can obtain information like what’s in train.csv during submission?",
          "replies": [
            {
              "id": 3262123,
              "postDate": "2025-08-03T00:48:57.780Z",
              "content": "<p>My bad then, but I think since test.csv purpose is toy example you can spect same for hidden, so I guess no more explicit information but the one available at test, metadata or custom preprocessing.</p>",
              "rawMarkdown": "My bad then, but I think since test.csv purpose is toy example you can spect same for hidden, so I guess no more explicit information but the one available at test, metadata or custom preprocessing."
            },
            {
              "id": 3262126,
              "postDate": "2025-08-03T00:58:11.813Z",
              "content": "<p>To be honest, I agree as well. However, modality information (CTA, T2WI, T1CE, MRA) would normally be available beforehand in routine practice, so I'm a bit curious about why this might be hidden here.</p>",
              "rawMarkdown": "To be honest, I agree as well. However, modality information (CTA, T2WI, T1CE, MRA) would normally be available beforehand in routine practice, so I'm a bit curious about why this might be hidden here.",
              "votes": 4
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3265689,
      "author_name": "Evan Calabrese",
      "author_url": "",
      "post_date": "2025-08-07T20:34:56.900000",
      "content": "<p>Thanks for your comment. It may surprise you to know that the type of MR series cannot be easily inferred from DICOM header metadata in real clinical datasets. Often we rely on series description, which is a free text field that can and often is manually edited by MR technologists at the time of acquisition. For obvious reasons this field has to be removed to prevent PHI leak, but it’s not that reliable anyway. Trying to infer MR series type based on other scan parameters available in the DICOM header is also not straightforward since these parameters vary widely across scanner makes/models. </p>\n<p>If you believe that series type identification is important for your approach, then one option is to use a pixel data-based approach. See for example: <a href=\"https://pmc.ncbi.nlm.nih.gov/articles/PMC10831512/\" target=\"_blank\">https://pmc.ncbi.nlm.nih.gov/articles/PMC10831512/</a></p>",
      "votes": 2,
      "replies": [
        {
          "id": 3265711,
          "author_name": "YYama",
          "author_url": "",
          "post_date": "2025-08-07T21:22:16.373000",
          "content": "<p>Thank you for your comment!</p>\n<p>As a radiologist in Japan, I’m familiar with that situation. However, I’m a bit surprised.<br>\nIn the hospitals I’ve worked at so far, while there were naturally some variations in notation within the series description, it wasn’t unreliable, and I was able to standardize it through rule-based NLP processing.<br>\nThis might be because the hospitals I’ve worked at had solid regulations and MR technologists diligently input the information. (There might also be differences between Japan and the United States.)</p>\n<p>And I’ve read that paper.<br>\nThe opening statement in the introduction that DICOM headers are unreliable didn’t quite resonate with me at the time, but now I understand the circumstances behind it. Very interesting.​​​​​​​​​​​​​​​​</p>",
          "votes": 2,
          "replies": [
            {
              "id": 3265745,
              "author_name": "Evan Calabrese",
              "author_url": "",
              "post_date": "2025-08-07T22:41:31.820000",
              "content": "<p>I think I would like to come work in Japan!!! 😂 Series descriptions can be quite unreliable here. Considering a multi-national dataset like this one, I think it would be extremely difficult to design an universal rules-based NLP for series description. Just to give you one example, some series in this dataset had the patients (or at least someone's) name in the series description before we wiped it 🤦.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 3265750,
              "author_name": "YYama",
              "author_url": "",
              "post_date": "2025-08-07T22:52:58.103000",
              "content": "<p>I think that was just within my narrow range of observation, and it’s probably not like that throughout all of Japan lol. Actually, when I was doing sequence classification myself, there were differences in descriptions depending on the imaging equipment, so I implemented rule-based processing myself. If patient names are included in the descriptions, removing PHI must be quite challenging. I’m grateful that you overcame these difficulties and organized this competition.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        },
        {
          "id": 3274179,
          "author_name": "RihanPiggy",
          "author_url": "",
          "post_date": "2025-08-24T08:59:18.110000",
          "content": "<p><a href=\"https://www.kaggle.com/evancalabrese\" target=\"_blank\">@evancalabrese</a> <a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> <br>\nHello, I have another question about the metadata in the hidden testset. When processing dicom, I perform a basic validation method by checking</p>\n<ul>\n<li>SeriesInstanceUID, InstanceNumber, Modality, ImagePositionPatient, ImageOrientationPatient in dicom</li>\n<li>length of ImagePositionPatient &gt;=3 and length of ImageOrientationPatient &gt;=6</li>\n</ul>\n<pre><code> ():\n        is_valid_dicom = \n         meta  [, ]:\n         ]:\n            is_valid_dicom = is_valid_dicom  meta  dicom_header\n\n            dicom_header  (dicom_header.ImageOrientationPatient) &lt; :\n            is_valid_dicom = \n\n            dicom_header  (dicom_header.ImagePositionPatient) &lt; :\n            is_valid_dicom = \n         is_valid_dicom\n</code></pre>\n<p>If the dicom file does not pass the check, it will be removed. This method works well on training data but I encountered a Series in the hidden testset that all of the dicoms are removed, which means all of the dicom of that series is invalid and lead to \"Notebook threw Exception\". It seems that this occurs after the dataset update. <a href=\"https://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/600658#3274261\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/600658#3274261</a></p>\n<p>Is it expected that some series does not contain ImagePositionPatient, InstanceNumber, ImageOrientationPatient or Modality in dicom metadata? I am not a radiologist but GPT5 tells me that these are must information. I don't think such a gap in dicom files between train and test data should occur and we cannot reconstruct the dicoms into ordered 3d volume without these information. Or is the submission system broken?</p>",
          "votes": 2,
          "replies": [
            {
              "id": 3274834,
              "author_name": "Evan Calabrese",
              "author_url": "",
              "post_date": "2025-08-25T13:54:30.470000",
              "content": "<p>These tags should all be present in the test set. I think there is currently an issue with the test set that may be causing this error. Hopefully we will have a fix soon.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3262349,
      "author_name": "Vettejeep",
      "author_url": "",
      "post_date": "2025-08-03T12:54:08.900000",
      "content": "<p>Noting that modality is also available in the dcm files both for train and test sets, though I have not tried to use it that way yet. Modality is in the list of tags provided for the test set per the \"Data\" section of the competition.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3262353,
          "author_name": "YYama",
          "author_url": "",
          "post_date": "2025-08-03T13:01:33.697000",
          "content": "<p>Thank you for your comment! As I mentioned in my post, I’m already aware of that. My question is whether it’s possible to obtain the type of MR sequence—of which there are three main kinds—as metadata. For more details, please refer to my original post.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3262058,
      "author_name": "Ángel Jacinto Sánchez Ruiz",
      "author_url": "",
      "post_date": "2025-08-02T19:54:53.240000",
      "content": "<p>You can find Modality in metadata.</p>\n<p>EDIT: dcm.Modality</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3262112,
          "author_name": "YYama",
          "author_url": "",
          "post_date": "2025-08-02T23:53:57.393000",
          "content": "<p>As I mentioned above, DICOM Modality tags don’t contain the three types of MR information, so my question is whether we can obtain information like what’s in train.csv during submission?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3262123,
              "author_name": "Ángel Jacinto Sánchez Ruiz",
              "author_url": "",
              "post_date": "2025-08-03T00:48:57.780000",
              "content": "<p>My bad then, but I think since test.csv purpose is toy example you can spect same for hidden, so I guess no more explicit information but the one available at test, metadata or custom preprocessing.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3262126,
              "author_name": "YYama",
              "author_url": "",
              "post_date": "2025-08-03T00:58:11.813000",
              "content": "<p>To be honest, I agree as well. However, modality information (CTA, T2WI, T1CE, MRA) would normally be available beforehand in routine practice, so I'm a bit curious about why this might be hidden here.</p>",
              "votes": 4,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3262020": "I noticed that there is no Modality information in test.csv.\nThe DICOM Modality contains only two types: \"CT\" or \"MR\". (It seems that the test DICOMs don't have Study Description either.)\nHowever, MR has three types: MRA, T2WI, and T1CE, and the optimal conditions during inference may differ.\n\nNaturally, it's possible to distinguish these using various algorithms, but since sequence is metadata that can be easily obtained during imaging, it would be desirable to use it during inference without any special processing or restrictions.\n\nWould it be possible to provide Modality in test.csv similar to train.csv?\nOr is there already some other way to use this information?",
    "3265689": "Thanks for your comment. It may surprise you to know that the type of MR series cannot be easily inferred from DICOM header metadata in real clinical datasets. Often we rely on series description, which is a free text field that can and often is manually edited by MR technologists at the time of acquisition. For obvious reasons this field has to be removed to prevent PHI leak, but it’s not that reliable anyway. Trying to infer MR series type based on other scan parameters available in the DICOM header is also not straightforward since these parameters vary widely across scanner makes/models. \n\nIf you believe that series type identification is important for your approach, then one option is to use a pixel data-based approach. See for example: https://pmc.ncbi.nlm.nih.gov/articles/PMC10831512/",
    "3262349": "Noting that modality is also available in the dcm files both for train and test sets, though I have not tried to use it that way yet. Modality is in the list of tags provided for the test set per the \"Data\" section of the competition.",
    "3262058": "You can find Modality in metadata.\n\nEDIT: dcm.Modality"
  }
}