{
  "id": 166441,
  "title": "Patient ID",
  "url": "/competitions/rsna-intracranial-hemorrhage-detection/discussion/166441",
  "author_name": "s nothing",
  "post_date": "2020-07-12T22:32:59.039000",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Is the patient ID attribute in the dicom the correct attribute to distinguish the patients? If that's correct, in the original train set, I find there are 18938 unique patients, and in the test set, 3518 unique patients. </p>\n\n<p>I'm trying to construct new train and validation sets from the original training sets, where the train and validation sets have images from different patients. </p>",
  "messages": [
    {
      "id": 926700,
      "postDate": "2020-07-12T22:32:59.040Z",
      "content": "<p>Is the patient ID attribute in the dicom the correct attribute to distinguish the patients? If that's correct, in the original train set, I find there are 18938 unique patients, and in the test set, 3518 unique patients. </p>\n\n<p>I'm trying to construct new train and validation sets from the original training sets, where the train and validation sets have images from different patients. </p>",
      "rawMarkdown": "Is the patient ID attribute in the dicom the correct attribute to distinguish the patients? If that's correct, in the original train set, I find there are 18938 unique patients, and in the test set, 3518 unique patients. \n\nI'm trying to construct new train and validation sets from the original training sets, where the train and validation sets have images from different patients. ",
      "votes": 1
    },
    {
      "id": 1324301,
      "postDate": "2021-05-26T19:25:17.457Z",
      "content": "<p>Hey!<br>\nI'm fairly new to python, Data Science and Kaggle in general and I'm trying to recreate and also \"solve\" this challenge but I'm having some issues and doubts and I would apreciate a lot if you could help me.</p>\n<p>So, for now, I'm able to get all slices from my directory:</p>\n<pre><code>#Set data directory\ndata_dir = 'D:\\RSNA Data Set\\\\rsna-intracranial-hemorrhage-detection\\stage_2_train\\'\n\n#Retrieve patients - List of dicom files\nprint(f\"Getting .dcm files from file path: {data_dir}\")\nslices_name = os.listdir(data_dir)\n\nslices = []\nfor slice_name in slices_name:\n    slice = dcmread(data_dir + slice_name)\n    slices.append(slice)\n</code></pre>\n<p>Is there any way that I can turn this list into a Data Frame? How can I create a <code>list</code> where the <code>len(list)</code> will be the number of distinct patients? Thank you!</p>",
      "rawMarkdown": "Hey!\nI'm fairly new to python, Data Science and Kaggle in general and I'm trying to recreate and also \"solve\" this challenge but I'm having some issues and doubts and I would apreciate a lot if you could help me.\n\nSo, for now, I'm able to get all slices from my directory:\n\n```\n#Set data directory\ndata_dir = 'D:\\RSNA Data Set\\\\rsna-intracranial-hemorrhage-detection\\stage_2_train\\'\n\n#Retrieve patients - List of dicom files\nprint(f\"Getting .dcm files from file path: {data_dir}\")\nslices_name = os.listdir(data_dir)\n\nslices = []\nfor slice_name in slices_name:\n    slice = dcmread(data_dir + slice_name)\n    slices.append(slice)\n```\n\nIs there any way that I can turn this list into a Data Frame? How can I create a `list ` where the `len(list)` will be the number of distinct patients? Thank you!"
    }
  ],
  "comments": [
    {
      "id": 1324301,
      "author_name": "JoseJoaoPimenta",
      "author_url": "",
      "post_date": "2021-05-26T19:25:17.457000",
      "content": "<p>Hey!<br>\nI'm fairly new to python, Data Science and Kaggle in general and I'm trying to recreate and also \"solve\" this challenge but I'm having some issues and doubts and I would apreciate a lot if you could help me.</p>\n<p>So, for now, I'm able to get all slices from my directory:</p>\n<pre><code>#Set data directory\ndata_dir = 'D:\\RSNA Data Set\\\\rsna-intracranial-hemorrhage-detection\\stage_2_train\\'\n\n#Retrieve patients - List of dicom files\nprint(f\"Getting .dcm files from file path: {data_dir}\")\nslices_name = os.listdir(data_dir)\n\nslices = []\nfor slice_name in slices_name:\n    slice = dcmread(data_dir + slice_name)\n    slices.append(slice)\n</code></pre>\n<p>Is there any way that I can turn this list into a Data Frame? How can I create a <code>list</code> where the <code>len(list)</code> will be the number of distinct patients? Thank you!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "926700": "Is the patient ID attribute in the dicom the correct attribute to distinguish the patients? If that's correct, in the original train set, I find there are 18938 unique patients, and in the test set, 3518 unique patients. \n\nI'm trying to construct new train and validation sets from the original training sets, where the train and validation sets have images from different patients. ",
    "1324301": "Hey!\nI'm fairly new to python, Data Science and Kaggle in general and I'm trying to recreate and also \"solve\" this challenge but I'm having some issues and doubts and I would apreciate a lot if you could help me.\n\nSo, for now, I'm able to get all slices from my directory:\n\n```\n#Set data directory\ndata_dir = 'D:\\RSNA Data Set\\\\rsna-intracranial-hemorrhage-detection\\stage_2_train\\'\n\n#Retrieve patients - List of dicom files\nprint(f\"Getting .dcm files from file path: {data_dir}\")\nslices_name = os.listdir(data_dir)\n\nslices = []\nfor slice_name in slices_name:\n    slice = dcmread(data_dir + slice_name)\n    slices.append(slice)\n```\n\nIs there any way that I can turn this list into a Data Frame? How can I create a `list ` where the `len(list)` will be the number of distinct patients? Thank you!"
  }
}