{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# RSNA Screening Mammography Breast Cancer Detection\nThis notebook explains how to transforms DICOM meta information to DataFrame.  \nThe DICOM file is prepared from the [RSNA Screening Mammography Breast Cancer Detection](https://www.kaggle.com/competitions/rsna-breast-cancer-detection) Competition.","metadata":{}},{"cell_type":"markdown","source":"The transformed csv files are stored in the **[RSNA-2022 Breast Cancer : DICOM Meta Information Dataset](https://www.kaggle.com/datasets/masatakaitakura/rsna2022-breast-cancer-dicom-meta-information) and in the [Input Data]** of this notebook. Please feel free to download and use this dataset, and if these datasets and this notebook help you, <font color=red>**please upvote both dataset and notebook!!**</font> ","metadata":{}},{"cell_type":"markdown","source":"The EDA for the transformed DICOM files are prepared in this notebook : [RSNA-2022 [EDA] DICOM Meta Information](https://www.kaggle.com/masatakaitakura/rsna-2022-eda-dicom-meta-information/edit)","metadata":{}},{"cell_type":"markdown","source":"## Contents\n- Overview of DICOM file\n- Transform to dataframe\n- [Tips] Handling DICOM meta information","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport os\nimport pydicom\nimport glob\nfrom tqdm import tqdm","metadata":{"execution":{"iopub.status.busy":"2022-12-17T23:26:08.710984Z","iopub.execute_input":"2022-12-17T23:26:08.71205Z","iopub.status.idle":"2022-12-17T23:26:08.889298Z","shell.execute_reply.started":"2022-12-17T23:26:08.7119Z","shell.execute_reply":"2022-12-17T23:26:08.887945Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# creating train and test folders\nos.makedirs('/kaggle/working/train')\nos.makedirs('/kaggle/working/test')","metadata":{"execution":{"iopub.status.busy":"2022-12-17T23:26:08.891783Z","iopub.execute_input":"2022-12-17T23:26:08.893824Z","iopub.status.idle":"2022-12-17T23:26:08.900271Z","shell.execute_reply.started":"2022-12-17T23:26:08.89377Z","shell.execute_reply":"2022-12-17T23:26:08.898772Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## train data","metadata":{}},{"cell_type":"code","source":"train = pd.read_csv('/kaggle/input/rsna-breast-cancer-detection/train.csv')\ntrain.head()","metadata":{"execution":{"iopub.status.busy":"2022-12-17T23:26:08.902236Z","iopub.execute_input":"2022-12-17T23:26:08.902777Z","iopub.status.idle":"2022-12-17T23:26:09.077966Z","shell.execute_reply.started":"2022-12-17T23:26:08.902726Z","shell.execute_reply":"2022-12-17T23:26:09.076422Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## test data","metadata":{}},{"cell_type":"code","source":"test = pd.read_csv('/kaggle/input/rsna-breast-cancer-detection/test.csv')\ntest.head()","metadata":{"execution":{"iopub.status.busy":"2022-12-17T05:37:40.608739Z","iopub.execute_input":"2022-12-17T05:37:40.609462Z","iopub.status.idle":"2022-12-17T05:37:40.629439Z","shell.execute_reply.started":"2022-12-17T05:37:40.609415Z","shell.execute_reply":"2022-12-17T05:37:40.628199Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### DICOM file meta information\nDICOM file contain some file information.\nEach file contains <font color=red>a rich information that might help us to understand the image deeply and how we split the data for training</font>. Some data might be provided as a train input directly to our model.","metadata":{}},{"cell_type":"markdown","source":"First, we see a single DICOM file's meta information.","metadata":{}},{"cell_type":"code","source":"dicom_path = \"/kaggle/input/rsna-breast-cancer-detection/train_images/10038/1967300488.dcm\"\ndcm_info = pydicom.dcmread(dicom_path)\ndcm_info","metadata":{"execution":{"iopub.status.busy":"2022-12-17T05:37:42.822282Z","iopub.execute_input":"2022-12-17T05:37:42.823127Z","iopub.status.idle":"2022-12-17T05:37:42.837912Z","shell.execute_reply.started":"2022-12-17T05:37:42.823061Z","shell.execute_reply":"2022-12-17T05:37:42.836714Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We have several way to confirm the DICOM meta information, and in this notebook, I will transform the information to DataFrame.","metadata":{}},{"cell_type":"markdown","source":"# Transform DICOM meta information to DataFrame\nNext, let's obtain these meta information and create the dataframe for some train data.","metadata":{}},{"cell_type":"markdown","source":"## Sample for 5 data","metadata":{}},{"cell_type":"code","source":"# sample for 5 data\n\n# prepare the blank dataframe\ndcm_df_5 = pd.DataFrame(columns=['name'])\n\nfor patient_id,image_id, machine_id in zip(tqdm(train['patient_id'].head()), train['image_id'].head(), train['machine_id'].head()):\n    # obtain the dicom_path\n    # obtain the dicom_path\n    dicom_path = \"/kaggle/input/rsna-breast-cancer-detection/train_images/\" + str(patient_id) + \"/\" + str(image_id) + \".dcm\"\n\n    # create df for each meta information\n    dataset = pydicom.dcmread(dicom_path)\n    dcm_list = []\n\n    for data_element in dataset:\n        dcm_list.append([data_element.name, data_element.value])\n\n    # delete the last meta information of [Pixel Data]\n    dcm_list = dcm_list[:-1]\n    \n    # insert `machine_id` into the list\n    dcm_list.append(['machine_id', machine_id])\n\n    # convert to df\n    dcm_df_ind = pd.DataFrame(dcm_list, columns = ['name', image_id])\n    \n    # combine each df\n    dcm_df_5 = pd.merge(dcm_df_5, dcm_df_ind, on='name', how='outer')\n\n# swap col and row\ndcm_df_5 = dcm_df_5.T\n\n# set row[0] as col name\ndcm_df_5.columns = dcm_df_5.iloc[0]\n\n# drop the unrequired first 'name' row\ndcm_df_5 = dcm_df_5.drop(index='name', axis=0)\n\ndcm_df_5","metadata":{"execution":{"iopub.status.busy":"2022-12-17T05:37:47.702026Z","iopub.execute_input":"2022-12-17T05:37:47.702512Z","iopub.status.idle":"2022-12-17T05:37:47.78422Z","shell.execute_reply.started":"2022-12-17T05:37:47.702473Z","shell.execute_reply":"2022-12-17T05:37:47.783135Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Save to csv file\ndcm_df_5.to_csv(\"/kaggle/working/dicom_meta_info_sample.csv\",index=True)","metadata":{"execution":{"iopub.status.busy":"2022-12-17T05:37:51.48816Z","iopub.execute_input":"2022-12-17T05:37:51.488568Z","iopub.status.idle":"2022-12-17T05:37:51.496126Z","shell.execute_reply.started":"2022-12-17T05:37:51.488533Z","shell.execute_reply":"2022-12-17T05:37:51.494523Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Confirmation\ndf_sample = pd.read_csv(\"/kaggle/working/dicom_meta_info_sample.csv\")\ndf_sample","metadata":{"execution":{"iopub.status.busy":"2022-12-17T05:37:56.194999Z","iopub.execute_input":"2022-12-17T05:37:56.195827Z","iopub.status.idle":"2022-12-17T05:37:56.227617Z","shell.execute_reply.started":"2022-12-17T05:37:56.195785Z","shell.execute_reply":"2022-12-17T05:37:56.225799Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Code for all train data\nThis code takes too much time to run, so be careful to run the following cells. <font color=red>I recommend you to find the result directly which I show at the beginning</font>.  \nThe train data has total 54,706 files, so the csv file is also devided into 6 files (~10,000 for each). The following cells are commented out to avoid the full run.","metadata":{}},{"cell_type":"markdown","source":"## 0 ~ 10,000 files","metadata":{"_kg_hide-input":true,"_kg_hide-output":true}},{"cell_type":"code","source":"# prepare the blank dataframe\ndcm_df = pd.DataFrame(columns=['name'])\n\nfor patient_id,image_id, machine_id in zip(tqdm(train['patient_id'][0:10000]), train['image_id'][0:10000], train['machine_id'][0:10000]):    # obtain the dicom_path\n    dicom_path = \"/kaggle/input/rsna-breast-cancer-detection/train_images/\" + str(patient_id) + \"/\" + str(image_id) + \".dcm\"\n\n    # create df for each meta information\n    dataset = pydicom.dcmread(dicom_path)\n    dcm_list = []\n\n    for data_element in dataset:\n        dcm_list.append([data_element.name, data_element.value])\n\n    # delete the last meta information of [Pixel Data]\n    dcm_list = dcm_list[:-1]\n    \n    # insert `machine_id` into the list\n    dcm_list.append(['machine_id', machine_id])\n\n    # convert to df\n    dcm_df_ind = pd.DataFrame(dcm_list, columns = ['name', image_id])\n    \n    # combine each df\n    dcm_df = pd.merge(dcm_df, dcm_df_ind, on='name', how='outer')\n    \n# swap col and row\ndcm_df = dcm_df.T\n\n# set row[0] as col name\ndcm_df.columns = dcm_df.iloc[0]\n\n# drop the unrequired first 'name' row\ndcm_df = dcm_df.drop(index='name', axis=0)\n\n# Save to csv file\ndcm_df.to_csv(\"/kaggle/working/train/dicom_meta_info_1.csv\",index=True)","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-12-17T21:19:23.14201Z","iopub.execute_input":"2022-12-17T21:19:23.142561Z","iopub.status.idle":"2022-12-17T21:19:23.242636Z","shell.execute_reply.started":"2022-12-17T21:19:23.142452Z","shell.execute_reply":"2022-12-17T21:19:23.241335Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 10,001 ~ 20,000 files","metadata":{"_kg_hide-input":true,"_kg_hide-output":true}},{"cell_type":"code","source":"# prepare the blank dataframe\ndcm_df = pd.DataFrame(columns=['name'])\n\nfor patient_id,image_id, machine_id in zip(tqdm(train['patient_id'][10000:20000]), train['image_id'][10000:20000], train['machine_id'][10000:20000]):    # obtain the dicom_path\n    dicom_path = \"/kaggle/input/rsna-breast-cancer-detection/train_images/\" + str(patient_id) + \"/\" + str(image_id) + \".dcm\"\n\n    # create df for each meta information\n    dataset = pydicom.dcmread(dicom_path)\n    dcm_list = []\n\n    for data_element in dataset:\n        dcm_list.append([data_element.name, data_element.value])\n\n    # delete the last meta information of [Pixel Data]\n    dcm_list = dcm_list[:-1]\n    \n    # insert `machine_id` into the list\n    dcm_list.append(['machine_id', machine_id])\n\n    # convert to df\n    dcm_df_ind = pd.DataFrame(dcm_list, columns = ['name', image_id])\n    \n    # combine each df\n    dcm_df = pd.merge(dcm_df, dcm_df_ind, on='name', how='outer')\n    \n# swap col and row\ndcm_df = dcm_df.T\n\n# set row[0] as col name\ndcm_df.columns = dcm_df.iloc[0]\n\n# drop the unrequired first 'name' row\ndcm_df = dcm_df.drop(index='name', axis=0)\n\n# Save to csv file\ndcm_df.to_csv(\"/kaggle/working/train/dicom_meta_info_2.csv\",index=True)","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-12-17T21:19:29.047937Z","iopub.execute_input":"2022-12-17T21:19:29.048341Z","iopub.status.idle":"2022-12-17T21:19:29.055181Z","shell.execute_reply.started":"2022-12-17T21:19:29.048299Z","shell.execute_reply":"2022-12-17T21:19:29.053896Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 20,001 ~ 30,000 files","metadata":{"_kg_hide-input":true,"_kg_hide-output":true}},{"cell_type":"code","source":"# prepare the blank dataframe\ndcm_df = pd.DataFrame(columns=['name'])\n\nfor patient_id,image_id, machine_id in zip(tqdm(train['patient_id'][20000:30000]), train['image_id'][20000:30000], train['machine_id'][20000:30000]):    # obtain the dicom_path\n    dicom_path = \"/kaggle/input/rsna-breast-cancer-detection/train_images/\" + str(patient_id) + \"/\" + str(image_id) + \".dcm\"\n\n    # create df for each meta information\n    dataset = pydicom.dcmread(dicom_path)\n    dcm_list = []\n\n    for data_element in dataset:\n        dcm_list.append([data_element.name, data_element.value])\n\n    # delete the last meta information of [Pixel Data]\n    dcm_list = dcm_list[:-1]\n    \n    # insert `machine_id` into the list\n    dcm_list.append(['machine_id', machine_id])\n\n    # convert to df\n    dcm_df_ind = pd.DataFrame(dcm_list, columns = ['name', image_id])\n    \n    # combine each df\n    dcm_df = pd.merge(dcm_df, dcm_df_ind, on='name', how='outer')\n    \n# swap col and row\ndcm_df = dcm_df.T\n\n# set row[0] as col name\ndcm_df.columns = dcm_df.iloc[0]\n\n# drop the unrequired first 'name' row\ndcm_df = dcm_df.drop(index='name', axis=0)\n\n# Save to csv file\ndcm_df.to_csv(\"/kaggle/working/train/dicom_meta_info_3.csv\",index=True)","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-12-17T05:37:06.46077Z","iopub.execute_input":"2022-12-17T05:37:06.461214Z","iopub.status.idle":"2022-12-17T05:37:08.239728Z","shell.execute_reply.started":"2022-12-17T05:37:06.461181Z","shell.execute_reply":"2022-12-17T05:37:08.238429Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 30,001 ~ 40,000 files","metadata":{"_kg_hide-input":true,"_kg_hide-output":true}},{"cell_type":"code","source":"# prepare the blank dataframe\ndcm_df = pd.DataFrame(columns=['name'])\n\nfor patient_id,image_id, machine_id in zip(tqdm(train['patient_id'][30000:40000]), train['image_id'][30000:40000], train['machine_id'][30000:40000]):    # obtain the dicom_path\n    dicom_path = \"/kaggle/input/rsna-breast-cancer-detection/train_images/\" + str(patient_id) + \"/\" + str(image_id) + \".dcm\"\n\n    # create df for each meta information\n    dataset = pydicom.dcmread(dicom_path)\n    dcm_list = []\n\n    for data_element in dataset:\n        dcm_list.append([data_element.name, data_element.value])\n\n    # delete the last meta information of [Pixel Data]\n    dcm_list = dcm_list[:-1]\n    \n    # insert `machine_id` into the list\n    dcm_list.append(['machine_id', machine_id])\n\n    # convert to df\n    dcm_df_ind = pd.DataFrame(dcm_list, columns = ['name', image_id])\n    \n    # combine each df\n    dcm_df = pd.merge(dcm_df, dcm_df_ind, on='name', how='outer')\n    \n# swap col and row\ndcm_df = dcm_df.T\n\n# set row[0] as col name\ndcm_df.columns = dcm_df.iloc[0]\n\n# drop the unrequired first 'name' row\ndcm_df = dcm_df.drop(index='name', axis=0)\n\n# Save to csv file\ndcm_df.to_csv(\"/kaggle/working/train/dicom_meta_info_4.csv\",index=True)","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-12-17T05:36:56.459047Z","iopub.execute_input":"2022-12-17T05:36:56.459534Z","iopub.status.idle":"2022-12-17T05:36:58.485781Z","shell.execute_reply.started":"2022-12-17T05:36:56.459498Z","shell.execute_reply":"2022-12-17T05:36:58.483983Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 40,001 ~ 50,000 files","metadata":{"_kg_hide-input":true,"_kg_hide-output":true}},{"cell_type":"code","source":"# prepare the blank dataframe\ndcm_df = pd.DataFrame(columns=['name'])\n\nfor patient_id,image_id, machine_id in zip(tqdm(train['patient_id'][40000:50000]), train['image_id'][40000:50000], train['machine_id'][40000:50000]):    # obtain the dicom_path\n    dicom_path = \"/kaggle/input/rsna-breast-cancer-detection/train_images/\" + str(patient_id) + \"/\" + str(image_id) + \".dcm\"\n\n    # create df for each meta information\n    dataset = pydicom.dcmread(dicom_path)\n    dcm_list = []\n\n    for data_element in dataset:\n        dcm_list.append([data_element.name, data_element.value])\n\n    # delete the last meta information of [Pixel Data]\n    dcm_list = dcm_list[:-1]\n    \n    # insert `machine_id` into the list\n    dcm_list.append(['machine_id', machine_id])\n\n    # convert to df\n    dcm_df_ind = pd.DataFrame(dcm_list, columns = ['name', image_id])\n    \n    # combine each df\n    dcm_df = pd.merge(dcm_df, dcm_df_ind, on='name', how='outer')\n    \n# swap col and row\ndcm_df = dcm_df.T\n\n# set row[0] as col name\ndcm_df.columns = dcm_df.iloc[0]\n\n# drop the unrequired first 'name' row\ndcm_df = dcm_df.drop(index='name', axis=0)\n\n# Save to csv file\ndcm_df.to_csv(\"/kaggle/working/train/dicom_meta_info_5.csv\",index=True)","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-12-17T05:36:47.646117Z","iopub.execute_input":"2022-12-17T05:36:47.64654Z","iopub.status.idle":"2022-12-17T05:36:51.055548Z","shell.execute_reply.started":"2022-12-17T05:36:47.646508Z","shell.execute_reply":"2022-12-17T05:36:51.053444Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 50,001 ~ 54,706 files","metadata":{"_kg_hide-input":true,"_kg_hide-output":true}},{"cell_type":"code","source":"# prepare the blank dataframe\ndcm_df = pd.DataFrame(columns=['name'])\n\nfor patient_id,image_id, machine_id in zip(tqdm(train['patient_id'][50000:54706]), train['image_id'][50000:54706], train['machine_id'][50000:54706]):    # obtain the dicom_path\n    dicom_path = \"/kaggle/input/rsna-breast-cancer-detection/train_images/\" + str(patient_id) + \"/\" + str(image_id) + \".dcm\"\n\n    # create df for each meta information\n    dataset = pydicom.dcmread(dicom_path)\n    dcm_list = []\n\n    for data_element in dataset:\n        dcm_list.append([data_element.name, data_element.value])\n\n    # delete the last meta information of [Pixel Data]\n    dcm_list = dcm_list[:-1]\n    \n    # insert `machine_id` into the list\n    dcm_list.append(['machine_id', machine_id])\n\n    # convert to df\n    dcm_df_ind = pd.DataFrame(dcm_list, columns = ['name', image_id])\n    \n    # combine each df\n    dcm_df = pd.merge(dcm_df, dcm_df_ind, on='name', how='outer')\n    \n# swap col and row\ndcm_df = dcm_df.T\n\n# set row[0] as col name\ndcm_df.columns = dcm_df.iloc[0]\n\n# drop the unrequired first 'name' row\ndcm_df = dcm_df.drop(index='name', axis=0)\n\n# Save to csv file\ndcm_df.to_csv(\"/kaggle/working/train/dicom_meta_info_6.csv\",index=True)","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-12-17T23:29:08.026781Z","iopub.execute_input":"2022-12-17T23:29:08.027945Z","iopub.status.idle":"2022-12-17T23:29:08.061199Z","shell.execute_reply.started":"2022-12-17T23:29:08.0279Z","shell.execute_reply":"2022-12-17T23:29:08.059958Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## test data","metadata":{"_kg_hide-input":true}},{"cell_type":"code","source":"# prepare the blank dataframe\ndcm_df = pd.DataFrame(columns=['name'])\n\nfor patient_id,image_id, machine_id in zip(tqdm(test['patient_id']), test['image_id'], test['machine_id']):    # obtain the dicom_path\n    dicom_path = \"/kaggle/input/rsna-breast-cancer-detection/test_images/\" + str(patient_id) + \"/\" + str(image_id) + \".dcm\"\n\n    # create df for each meta information\n    dataset = pydicom.dcmread(dicom_path)\n    dcm_list = []\n\n    for data_element in dataset:\n        dcm_list.append([data_element.name, data_element.value])\n\n    # delete the last meta information of [Pixel Data]\n    dcm_list = dcm_list[:-1]\n    \n    # insert `machine_id` into the list\n    dcm_list.append(['machine_id', machine_id])\n\n    # convert to df\n    dcm_df_ind = pd.DataFrame(dcm_list, columns = ['name', image_id])\n    \n    # combine each df\n    dcm_df = pd.merge(dcm_df, dcm_df_ind, on='name', how='outer')\n    \n# swap col and row\ndcm_df = dcm_df.T\n\n# set row[0] as col name\ndcm_df.columns = dcm_df.iloc[0]\n\n# drop the unrequired first 'name' row\ndcm_df = dcm_df.drop(index='name', axis=0)\n\n# Save to csv file\ndcm_df.to_csv(\"/kaggle/working/test/dicom_meta_info.csv\",index=True)","metadata":{"execution":{"iopub.status.busy":"2022-12-17T05:36:35.988541Z","iopub.status.idle":"2022-12-17T05:36:35.989813Z","shell.execute_reply.started":"2022-12-17T05:36:35.989512Z","shell.execute_reply":"2022-12-17T05:36:35.989539Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# # Confirmation\n# df = pd.read_csv(\"/kaggle/working/test/dicom_meta_info.csv\")\n# df","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-12-17T05:36:35.991476Z","iopub.status.idle":"2022-12-17T05:36:35.992037Z","shell.execute_reply.started":"2022-12-17T05:36:35.991747Z","shell.execute_reply":"2022-12-17T05:36:35.991773Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Thank you very much for reading, I hope this notebook helps you!\n\n**If you enjoyed the notebook, <font color=red>please upvote!</font> 🙏 Thank you, appreciate your support!**\n","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# [Tips] Handling DICOM meta information","metadata":{}},{"cell_type":"markdown","source":"We can obtain the 'keywords' for the dicom file by `.dir()` as shown below.","metadata":{"_kg_hide-input":false,"_kg_hide-output":false}},{"cell_type":"code","source":"dcm_info.dir()","metadata":{"_kg_hide-input":false,"_kg_hide-output":false,"execution":{"iopub.status.busy":"2022-12-17T05:36:35.993979Z","iopub.status.idle":"2022-12-17T05:36:35.994798Z","shell.execute_reply.started":"2022-12-17T05:36:35.994516Z","shell.execute_reply":"2022-12-17T05:36:35.994543Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can obtain the mete information by specifying the row title without space as show below.","metadata":{"_kg_hide-input":false,"_kg_hide-output":false}},{"cell_type":"code","source":"dcm_info_list = [\n    dcm_info.SOPInstanceUID,\n    dcm_info.ContentDate,\n    dcm_info.ContentTime,\n    dcm_info.PatientID, \n    dcm_info.BodyPartThickness,\n    dcm_info.CompressionForce,\n    dcm_info.ExposureControlMode,\n    dcm_info.ExposureControlModeDescription,\n    dcm_info.StudyInstanceUID,\n    dcm_info.SeriesInstanceUID,\n    dcm_info.InstanceNumber,\n    dcm_info.ImageLaterality,\n    dcm_info.SamplesPerPixel,\n    dcm_info.PhotometricInterpretation,\n    dcm_info.Rows,\n    dcm_info.Columns,\n    dcm_info.BitsAllocated,\n    dcm_info.BitsStored,\n    dcm_info.HighBit,\n    dcm_info.PixelRepresentation,\n    dcm_info.PixelIntensityRelationship,\n    dcm_info.PixelIntensityRelationshipSign,\n    dcm_info.WindowCenter,\n    dcm_info.WindowWidth,\n    dcm_info.RescaleIntercept,\n    dcm_info.RescaleSlope,\n    dcm_info.RescaleType,\n    dcm_info.VOILUTFunction,\n    dcm_info.LossyImageCompression,\n]\n\ndcm_info_list","metadata":{"_kg_hide-input":false,"_kg_hide-output":false,"execution":{"iopub.status.busy":"2022-12-17T05:36:35.996303Z","iopub.status.idle":"2022-12-17T05:36:35.99713Z","shell.execute_reply.started":"2022-12-17T05:36:35.996813Z","shell.execute_reply":"2022-12-17T05:36:35.996839Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"According to the [pydicom official document](https://pydicom.github.io/pydicom/stable/auto_examples/input_output/plot_printing_dataset.html#sphx-glr-auto-examples-input-output-plot-printing-dataset-py), we can obtain the meta information in my own format.","metadata":{"_kg_hide-input":false,"_kg_hide-output":false}},{"cell_type":"code","source":"# source : https://pydicom.github.io/pydicom/stable/auto_examples/input_output/plot_printing_dataset.html#sphx-glr-auto-examples-input-output-plot-printing-dataset-py\n\ndef myprint(dataset, indent=0):\n    \"\"\"Go through all items in the dataset and print them with custom format\n\n    Modelled after Dataset._pretty_str()\n    \"\"\"\n    dont_print = ['Pixel Data', 'File Meta Information Version']\n\n    indent_string = \"   \" * indent\n    next_indent_string = \"   \" * (indent + 1)\n\n    for data_element in dataset:\n        if data_element.VR == \"SQ\":   # a sequence\n            print(indent_string, data_element.name)\n            for sequence_item in data_element.value:\n                myprint(sequence_item, indent + 1)\n                print(next_indent_string + \"---------\")\n        else:\n            if data_element.name in dont_print:\n                print(\"\"\"<item not printed -- in the \"don't print\" list>\"\"\")\n            else:\n                repr_value = repr(data_element.value)\n                if len(repr_value) > 50:\n                    repr_value = repr_value[:50] + \"...\"\n                print(\"{0:s} {1:s} = {2:s}\".format(indent_string,\n                                                   data_element.name,\n                                                   repr_value))\n\n\n# Set the dicom file path\nds = pydicom.dcmread(dicom_path)\n\nmyprint(ds)","metadata":{"_kg_hide-input":false,"_kg_hide-output":false,"execution":{"iopub.status.busy":"2022-12-17T05:38:12.314047Z","iopub.execute_input":"2022-12-17T05:38:12.314476Z","iopub.status.idle":"2022-12-17T05:38:12.331318Z","shell.execute_reply.started":"2022-12-17T05:38:12.314445Z","shell.execute_reply":"2022-12-17T05:38:12.329886Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Other way to transform the DICOM file to df.  \nSorce: https://www.kaggle.com/code/servietsky/osic-transform-dicom-into-dataframe/notebook","metadata":{"_kg_hide-input":false,"_kg_hide-output":false}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport pydicom\nfrom pydicom import dcmread\nfrom pydicom.data import get_testdata_files\nimport glob, os\nfrom collections import defaultdict\nfrom tqdm import tqdm\nimport gc","metadata":{"_kg_hide-input":false,"_kg_hide-output":false,"execution":{"iopub.status.busy":"2022-12-17T05:38:13.389357Z","iopub.execute_input":"2022-12-17T05:38:13.390746Z","iopub.status.idle":"2022-12-17T05:38:13.398096Z","shell.execute_reply.started":"2022-12-17T05:38:13.390707Z","shell.execute_reply":"2022-12-17T05:38:13.396661Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"col_name = dcm_info.dir()\n\ndf = pd.DataFrame(columns=col_name)\nmy_dict = defaultdict(list)\n\nfor name in tqdm(glob.glob('/kaggle/input/rsna-breast-cancer-detection/train_images/10006/*')): # run for only patient_id = 10006\n# for name in tqdm.tqdm(glob.glob('/kaggle/input/rsna-breast-cancer-detection/train_images/*/*')): # run for all data\n    ds = pydicom.read_file(name)\n    for i in col_name :\n        if i in ds :\n            my_dict[i].append(str(ds[i].value))\n        else:\n            my_dict[i].append(np.nan)\n    df = pd.concat([df, pd.DataFrame(my_dict)], ignore_index = True)\n    del my_dict\n    my_dict = defaultdict(list)\ngc.collect()\n\ndf","metadata":{"_kg_hide-input":false,"_kg_hide-output":false,"execution":{"iopub.status.busy":"2022-12-17T05:38:13.652911Z","iopub.execute_input":"2022-12-17T05:38:13.653336Z","iopub.status.idle":"2022-12-17T05:38:14.096277Z","shell.execute_reply.started":"2022-12-17T05:38:13.653302Z","shell.execute_reply":"2022-12-17T05:38:14.094793Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"_kg_hide-input":false,"_kg_hide-output":false},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}