{
  "id": 386386,
  "title": "Question about StudyInstanceUID and PatientID",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/386386",
  "author_name": "Antti Isosalo",
  "post_date": "2023-02-12T19:31:34.378000",
  "votes": 0,
  "comment_count": 3,
  "views": 0,
  "content": "<p>👋 Greetings!</p>\n<p>I have question about StudyInstanceUID and patient_id.</p>\n<p>Usually, in <code>DICOM</code> files the <code>ds.StudyInstanceUID</code> gives an unique identifier for each examination. Patients on the other hand have a <code>ds.PatientID</code> which distinguishes them from other patients.</p>\n<p>So, <code>StudyInstanceUID</code> allows to distinguish one examination from another. (And then some time label will tell where the examination is positioned on the time axis.)</p>\n<p>If I am not mistaken, in this RSNA dataset for screening mammography, and especially the metadata, <code>patient_id</code> seems to equal <code>ds.StudyInstanceUID</code>? <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> </p>\n<p><strong>How are we able to stratify</strong> the data properly for deep learning, if there is no anonymous ID to tell if the same patient has several <strong>examinations</strong> in the dataset? Or should we conclude that each patient has only one examination (with several images) in the dataset?</p>\n<p><strong>EDIT</strong>: Here is my investigation regarding the <code>StudyInstanceUID</code>: <a href=\"https://www.kaggle.com/code/anttiisosalo/16-bit-post-processing-rsna-bc-detection?scriptVersionId=119182142&amp;cellId=28\" target=\"_blank\">🎗️ [16-bit post-processing] RSNA BC Detection</a></p>\n<p>Best,</p>\n<p>Antti 👍</p>",
  "messages": [
    {
      "id": 2145446,
      "postDate": "2023-02-15T04:45:08.480Z",
      "content": "<p>\"How are we able to stratify the data properly for deep learning, if there is no anonymous ID to tell if the same patient has several examinations in the dataset? Or should we conclude that each patient has only one examination (with several images) in the dataset?\"</p>\n<p>From my perspective, I assume that there are some duplicate DICOM images, but fewer. In addition, I normally distinguish them by implementing Pandas library (e.g. Series.unique). And match the number to the number of images in the competition dataset (54k).</p>\n<p>As you know, the process of training a Neural Network is a repeated job, not just a single process. That means, in each epoch, the patients' data is used to train our model repeatedly.</p>",
      "rawMarkdown": "\"How are we able to stratify the data properly for deep learning, if there is no anonymous ID to tell if the same patient has several examinations in the dataset? Or should we conclude that each patient has only one examination (with several images) in the dataset?\"\n\nFrom my perspective, I assume that there are some duplicate DICOM images, but fewer. In addition, I normally distinguish them by implementing Pandas library (e.g. Series.unique). And match the number to the number of images in the competition dataset (54k).\n\nAs you know, the process of training a Neural Network is a repeated job, not just a single process. That means, in each epoch, the patients' data is used to train our model repeatedly.",
      "replies": [
        {
          "id": 2146433,
          "postDate": "2023-02-15T21:36:16.557Z",
          "content": "<blockquote>\n  <p>That means, in each epoch, the patients' data is used to train our model repeatedly.</p>\n</blockquote>\n<p>Perhaps all the more reason to make sure that there is no patient-wise leakage between train and validation subsets? So that we don't accidentally learn (<em>i.e.</em>, fit to) <strong>the anatomy of some particular patient</strong>. Or that we don't get a model, which we assume to generalize well based on our validation set, but in truth is only good for breast cancer patterns distinctive (again) <strong>to some particular patient</strong>?</p>",
          "rawMarkdown": ">That means, in each epoch, the patients' data is used to train our model repeatedly.\n\nPerhaps all the more reason to make sure that there is no patient-wise leakage between train and validation subsets? So that we don't accidentally learn (*i.e.*, fit to) **the anatomy of some particular patient**. Or that we don't get a model, which we assume to generalize well based on our validation set, but in truth is only good for breast cancer patterns distinctive (again) **to some particular patient**?",
          "replies": [
            {
              "id": 2146638,
              "postDate": "2023-02-16T02:58:54.433Z",
              "content": "<p>Yes, the probability of a bias can occur in terms of the quality of a dataset. Assume that the 10,000 paitent images are distinguished well, but 2000 don't. It might follow the cases:</p>\n<ol>\n<li><p>Preprosessing is a good option to overcome the shortage of our data. Because it gives us a wide range of new data derived from the orginial dataset. It will be a good resource in terms of quantity and quality.</p></li>\n<li><p>Our model is difficult to train the medical images of DICOM. Because the patterns of tumers might be challenged to the model. (In this case, we have 2 options)<br>\nOp1 : change the model.<br>\nOp2 : have to think about making standard of factors, especially labels.</p></li>\n<li><p>Include outer dataset in order to fill the gap between biased data and non-biased one.</p></li>\n</ol>",
              "rawMarkdown": "Yes, the probability of a bias can occur in terms of the quality of a dataset. Assume that the 10,000 paitent images are distinguished well, but 2000 don't. It might follow the cases:\n\n1. Preprosessing is a good option to overcome the shortage of our data. Because it gives us a wide range of new data derived from the orginial dataset. It will be a good resource in terms of quantity and quality.\n\n2. Our model is difficult to train the medical images of DICOM. Because the patterns of tumers might be challenged to the model. (In this case, we have 2 options)\nOp1 : change the model.\nOp2 : have to think about making standard of factors, especially labels.\n\n3. Include outer dataset in order to fill the gap between biased data and non-biased one.\n\n\n",
              "votes": -1
            }
          ]
        }
      ]
    },
    {
      "id": 2141481,
      "postDate": "2023-02-12T19:31:34.380Z",
      "content": "<p>👋 Greetings!</p>\n<p>I have question about StudyInstanceUID and patient_id.</p>\n<p>Usually, in <code>DICOM</code> files the <code>ds.StudyInstanceUID</code> gives an unique identifier for each examination. Patients on the other hand have a <code>ds.PatientID</code> which distinguishes them from other patients.</p>\n<p>So, <code>StudyInstanceUID</code> allows to distinguish one examination from another. (And then some time label will tell where the examination is positioned on the time axis.)</p>\n<p>If I am not mistaken, in this RSNA dataset for screening mammography, and especially the metadata, <code>patient_id</code> seems to equal <code>ds.StudyInstanceUID</code>? <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> </p>\n<p><strong>How are we able to stratify</strong> the data properly for deep learning, if there is no anonymous ID to tell if the same patient has several <strong>examinations</strong> in the dataset? Or should we conclude that each patient has only one examination (with several images) in the dataset?</p>\n<p><strong>EDIT</strong>: Here is my investigation regarding the <code>StudyInstanceUID</code>: <a href=\"https://www.kaggle.com/code/anttiisosalo/16-bit-post-processing-rsna-bc-detection?scriptVersionId=119182142&amp;cellId=28\" target=\"_blank\">🎗️ [16-bit post-processing] RSNA BC Detection</a></p>\n<p>Best,</p>\n<p>Antti 👍</p>",
      "rawMarkdown": "👋 Greetings!\n\nI have question about StudyInstanceUID and patient_id.\n\nUsually, in `DICOM` files the `ds.StudyInstanceUID` gives an unique identifier for each examination. Patients on the other hand have a `ds.PatientID` which distinguishes them from other patients.\n\nSo, `StudyInstanceUID` allows to distinguish one examination from another. (And then some time label will tell where the examination is positioned on the time axis.)\n\nIf I am not mistaken, in this RSNA dataset for screening mammography, and especially the metadata, `patient_id` seems to equal `ds.StudyInstanceUID`? @sohier \n\n**How are we able to stratify** the data properly for deep learning, if there is no anonymous ID to tell if the same patient has several **examinations** in the dataset? Or should we conclude that each patient has only one examination (with several images) in the dataset?\n\n**EDIT**: Here is my investigation regarding the `StudyInstanceUID`: [🎗️ [16-bit post-processing] RSNA BC Detection](https://www.kaggle.com/code/anttiisosalo/16-bit-post-processing-rsna-bc-detection?scriptVersionId=119182142&cellId=28)\n\nBest,\n\nAntti 👍"
    }
  ],
  "comments": [
    {
      "id": 2145446,
      "author_name": "Hyunsoo Lee 1010",
      "author_url": "",
      "post_date": "2023-02-15T04:45:08.480000",
      "content": "<p>\"How are we able to stratify the data properly for deep learning, if there is no anonymous ID to tell if the same patient has several examinations in the dataset? Or should we conclude that each patient has only one examination (with several images) in the dataset?\"</p>\n<p>From my perspective, I assume that there are some duplicate DICOM images, but fewer. In addition, I normally distinguish them by implementing Pandas library (e.g. Series.unique). And match the number to the number of images in the competition dataset (54k).</p>\n<p>As you know, the process of training a Neural Network is a repeated job, not just a single process. That means, in each epoch, the patients' data is used to train our model repeatedly.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2146433,
          "author_name": "Antti Isosalo",
          "author_url": "",
          "post_date": "2023-02-15T21:36:16.557000",
          "content": "<blockquote>\n  <p>That means, in each epoch, the patients' data is used to train our model repeatedly.</p>\n</blockquote>\n<p>Perhaps all the more reason to make sure that there is no patient-wise leakage between train and validation subsets? So that we don't accidentally learn (<em>i.e.</em>, fit to) <strong>the anatomy of some particular patient</strong>. Or that we don't get a model, which we assume to generalize well based on our validation set, but in truth is only good for breast cancer patterns distinctive (again) <strong>to some particular patient</strong>?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2146638,
              "author_name": "Hyunsoo Lee 1010",
              "author_url": "",
              "post_date": "2023-02-16T02:58:54.433000",
              "content": "<p>Yes, the probability of a bias can occur in terms of the quality of a dataset. Assume that the 10,000 paitent images are distinguished well, but 2000 don't. It might follow the cases:</p>\n<ol>\n<li><p>Preprosessing is a good option to overcome the shortage of our data. Because it gives us a wide range of new data derived from the orginial dataset. It will be a good resource in terms of quantity and quality.</p></li>\n<li><p>Our model is difficult to train the medical images of DICOM. Because the patterns of tumers might be challenged to the model. (In this case, we have 2 options)<br>\nOp1 : change the model.<br>\nOp2 : have to think about making standard of factors, especially labels.</p></li>\n<li><p>Include outer dataset in order to fill the gap between biased data and non-biased one.</p></li>\n</ol>",
              "votes": -1,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2145446": "\"How are we able to stratify the data properly for deep learning, if there is no anonymous ID to tell if the same patient has several examinations in the dataset? Or should we conclude that each patient has only one examination (with several images) in the dataset?\"\n\nFrom my perspective, I assume that there are some duplicate DICOM images, but fewer. In addition, I normally distinguish them by implementing Pandas library (e.g. Series.unique). And match the number to the number of images in the competition dataset (54k).\n\nAs you know, the process of training a Neural Network is a repeated job, not just a single process. That means, in each epoch, the patients' data is used to train our model repeatedly.",
    "2141481": "👋 Greetings!\n\nI have question about StudyInstanceUID and patient_id.\n\nUsually, in `DICOM` files the `ds.StudyInstanceUID` gives an unique identifier for each examination. Patients on the other hand have a `ds.PatientID` which distinguishes them from other patients.\n\nSo, `StudyInstanceUID` allows to distinguish one examination from another. (And then some time label will tell where the examination is positioned on the time axis.)\n\nIf I am not mistaken, in this RSNA dataset for screening mammography, and especially the metadata, `patient_id` seems to equal `ds.StudyInstanceUID`? @sohier \n\n**How are we able to stratify** the data properly for deep learning, if there is no anonymous ID to tell if the same patient has several **examinations** in the dataset? Or should we conclude that each patient has only one examination (with several images) in the dataset?\n\n**EDIT**: Here is my investigation regarding the `StudyInstanceUID`: [🎗️ [16-bit post-processing] RSNA BC Detection](https://www.kaggle.com/code/anttiisosalo/16-bit-post-processing-rsna-bc-detection?scriptVersionId=119182142&cellId=28)\n\nBest,\n\nAntti 👍"
  }
}