{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# [EDA] RSNA Screening Mammography Breast Cancer Detection\n","metadata":{}},{"cell_type":"markdown","source":"This notebook describes the [EDA] for the [RSNA Screening Mammography Breast Cancer Detection](https://www.kaggle.com/competitions/rsna-breast-cancer-detection)","metadata":{}},{"cell_type":"markdown","source":"## Contents\n- train.csv, test.csv, sample_sabmission.csv overview\n- Visualize the DICOM file images\n- Explore the DICOM file meta information","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport pandas_profiling\nimport numpy as np\nimport os\n\n\n# visualization\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n%matplotlib inline\n\n# install gdcm to open the DICOM file\n!pip install -qU python-gdcm pydicom pylibjpeg\n\nimport pydicom\nimport pylibjpeg\n","metadata":{"execution":{"iopub.status.busy":"2022-12-12T23:51:48.643029Z","iopub.execute_input":"2022-12-12T23:51:48.64345Z","iopub.status.idle":"2022-12-12T23:52:00.135494Z","shell.execute_reply.started":"2022-12-12T23:51:48.643421Z","shell.execute_reply":"2022-12-12T23:52:00.13384Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"section-one\"></a>\n# Overview of data","metadata":{}},{"cell_type":"markdown","source":"## train data","metadata":{}},{"cell_type":"code","source":"train = pd.read_csv('/kaggle/input/rsna-breast-cancer-detection/train.csv')\ntrain.head()","metadata":{"execution":{"iopub.status.busy":"2022-12-12T23:52:00.138136Z","iopub.execute_input":"2022-12-12T23:52:00.138546Z","iopub.status.idle":"2022-12-12T23:52:00.232967Z","shell.execute_reply.started":"2022-12-12T23:52:00.138505Z","shell.execute_reply":"2022-12-12T23:52:00.231769Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The following is the explanation of columns in [Kaggle-Data](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/data) \n\n`site_id` - ID code for the source hospital.  \n`patient_id` - ID code for the patient.  \n`image_id` - ID code for the image.  \n`laterality` - Whether the image is of the left or right breast.  \n`view` - The orientation of the image. The default for a screening exam is to capture two views per breast.  \n`age` - The patient's age in years.  \n`implant` - Whether or not the patient had breast implants. Site 1 only provides breast implant information at the patient level, not at the breast level.  \n`density` - A rating for how dense the breast tissue is, with A being the least dense and D being the most dense. Extremely dense tissue can make diagnosis more difficult. <font color=blue>Only provided for train.</font>  \n`machine_id` - An ID code for the imaging device.  \n`cancer` - Whether or not the breast was positive for cancer. The target value. <font color=blue>Only provided for train.  </font>  \n`biopsy` - Whether or not a follow-up biopsy was performed on the breast. <font color=blue>Only provided for train.  </font>  \n`invasive` - If the breast is positive for cancer, whether or not the cancer proved to be invasive. <font color=blue>Only provided for train.  </font>  \n`BIRADS` - 0 if the breast required follow-up, 1 if the breast was rated as negative for cancer, and 2 if the breast was rated as normal. <font color=blue>Only provided for train.  </font>  \n`prediction_id` - The ID for the matching submission row. Multiple images will share the same prediction ID. Test only.  \n`difficult_negative_case` - True if the case was unusually difficult. <font color=blue>Only provided for train.  </font>","metadata":{}},{"cell_type":"markdown","source":"Let's see the overview of the train data.","metadata":{}},{"cell_type":"markdown","source":"## Target columns and features","metadata":{}},{"cell_type":"code","source":"# 'cancer' and 'age'\nfig, ax1 = plt.subplots()\n\ncolor = 'tab:blue'\nax1.set_xlabel('age')\nax1.set_ylabel('cancer=0 count', color=color)\nax1.hist(train.loc[train['cancer'] == 0, 'age'].dropna(), bins=30, alpha=0.5, label='0', color=color)\nax1.tick_params(axis='y', labelcolor=color)\n\nax2 = ax1.twinx()  # instantiate a second axes that shares the same x-axis\n\ncolor = 'tab:orange'\nax2.set_ylabel('cancer=1 count', color=color)  # we already handled the x-label with ax1\nax2.hist(train.loc[train['cancer'] == 1, 'age'].dropna(), bins=30, alpha=0.8, label='1', color=color)\nax2.tick_params(axis='y', labelcolor=color)\n\nfig.tight_layout()  # otherwise the right y-label is slightly clipped\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-12-12T23:52:00.234694Z","iopub.execute_input":"2022-12-12T23:52:00.235139Z","iopub.status.idle":"2022-12-12T23:52:00.769976Z","shell.execute_reply.started":"2022-12-12T23:52:00.235095Z","shell.execute_reply":"2022-12-12T23:52:00.768669Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Statistical data for 'cancer' and 'age'\ntrain.loc[train['cancer'] == 1, 'age'].describe()","metadata":{"execution":{"iopub.status.busy":"2022-12-12T23:52:00.774053Z","iopub.execute_input":"2022-12-12T23:52:00.774555Z","iopub.status.idle":"2022-12-12T23:52:00.793397Z","shell.execute_reply.started":"2022-12-12T23:52:00.774509Z","shell.execute_reply":"2022-12-12T23:52:00.792058Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Findings:\n- cancers count vs age seems normal distribution.\n- the peak is appeared around age 60 to 75.\n- cancers appeared from age 40.","metadata":{}},{"cell_type":"code","source":"# 'cancer' and 'laterality'\nsns.countplot(x='laterality', hue='cancer', data=train)\nplt.legend(loc='upper right', title='cancer')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-12-12T23:52:00.795104Z","iopub.execute_input":"2022-12-12T23:52:00.796011Z","iopub.status.idle":"2022-12-12T23:52:01.047682Z","shell.execute_reply.started":"2022-12-12T23:52:00.795969Z","shell.execute_reply":"2022-12-12T23:52:01.046304Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Findings:\n- laterality data seems balanced.","metadata":{}},{"cell_type":"code","source":"# 'cancer' and 'view'\nsns.countplot(x='view', hue='cancer', data=train)\nplt.legend(loc='upper right', title='cancer')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-12-12T23:52:01.049589Z","iopub.execute_input":"2022-12-12T23:52:01.049947Z","iopub.status.idle":"2022-12-12T23:52:01.323241Z","shell.execute_reply.started":"2022-12-12T23:52:01.049917Z","shell.execute_reply":"2022-12-12T23:52:01.321667Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Findings\n- it seems there are no significand difference between 'cancer' and 'non-cancer'.\n- the view of 'CC' and 'MLO' are the most popular.","metadata":{}},{"cell_type":"code","source":"# 'cancer' and 'density' *ONLY available in TRAIN data\nsns.countplot(x='density', hue='cancer', data=train)\nplt.legend(loc='upper right', title='cancer')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-12-12T23:52:01.324959Z","iopub.execute_input":"2022-12-12T23:52:01.32604Z","iopub.status.idle":"2022-12-12T23:52:01.576622Z","shell.execute_reply.started":"2022-12-12T23:52:01.325999Z","shell.execute_reply":"2022-12-12T23:52:01.575256Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 'cancer' and 'BIRADS' *ONLY available in TRAIN data\nsns.countplot(x='BIRADS', hue='cancer', data=train)\nplt.legend(loc='upper right', title='cancer')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-12-12T23:52:01.578449Z","iopub.execute_input":"2022-12-12T23:52:01.579162Z","iopub.status.idle":"2022-12-12T23:52:01.822057Z","shell.execute_reply.started":"2022-12-12T23:52:01.579117Z","shell.execute_reply":"2022-12-12T23:52:01.820681Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"If you want to see more comprehensive overview, please see bellow.","metadata":{}},{"cell_type":"code","source":"train.profile_report()","metadata":{"execution":{"iopub.status.busy":"2022-12-12T23:52:01.823732Z","iopub.execute_input":"2022-12-12T23:52:01.825102Z","iopub.status.idle":"2022-12-12T23:52:21.669859Z","shell.execute_reply.started":"2022-12-12T23:52:01.825051Z","shell.execute_reply":"2022-12-12T23:52:21.668903Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Next, see the data of `cancer=1` to understand the target column features.","metadata":{}},{"cell_type":"code","source":"train_cancer = train[train['cancer']==1]\ntrain_cancer","metadata":{"execution":{"iopub.status.busy":"2022-12-12T23:52:21.673938Z","iopub.execute_input":"2022-12-12T23:52:21.675118Z","iopub.status.idle":"2022-12-12T23:52:21.703091Z","shell.execute_reply.started":"2022-12-12T23:52:21.675079Z","shell.execute_reply":"2022-12-12T23:52:21.701626Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# if you want to see the detail [profile_report], please run this cell.\n# train_cancer.profile_report()","metadata":{"execution":{"iopub.status.busy":"2022-12-12T23:52:21.704614Z","iopub.execute_input":"2022-12-12T23:52:21.704936Z","iopub.status.idle":"2022-12-12T23:52:21.71394Z","shell.execute_reply.started":"2022-12-12T23:52:21.704908Z","shell.execute_reply":"2022-12-12T23:52:21.712378Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## test data","metadata":{}},{"cell_type":"code","source":"test = pd.read_csv('/kaggle/input/rsna-breast-cancer-detection/test.csv')\ntest.head()","metadata":{"execution":{"iopub.status.busy":"2022-12-12T23:52:21.715606Z","iopub.execute_input":"2022-12-12T23:52:21.71659Z","iopub.status.idle":"2022-12-12T23:52:21.746498Z","shell.execute_reply.started":"2022-12-12T23:52:21.716543Z","shell.execute_reply":"2022-12-12T23:52:21.74501Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The test.csv only contains 4 rows, single person data.\n\"Only the first few rows of the test set are available for download.\"","metadata":{}},{"cell_type":"markdown","source":"## sample_submission","metadata":{}},{"cell_type":"code","source":"sample_submission = pd.read_csv('/kaggle/input/rsna-breast-cancer-detection/sample_submission.csv')\nsample_submission.head()","metadata":{"execution":{"iopub.status.busy":"2022-12-12T23:52:21.749006Z","iopub.execute_input":"2022-12-12T23:52:21.749918Z","iopub.status.idle":"2022-12-12T23:52:21.770271Z","shell.execute_reply.started":"2022-12-12T23:52:21.749869Z","shell.execute_reply":"2022-12-12T23:52:21.768861Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Visualization","metadata":{}},{"cell_type":"markdown","source":"To visualize the DICOM file (.dcm), the following function is defined. Note that the line #12 and #13 is commented out to see the difference between each `machine_id`. The further detail to be discussed later.\nIf you activate line #16 and comment out the line #15, you can obtain the original colored images.","metadata":{}},{"cell_type":"code","source":"# Function to visualize the DICOM file\n\ndef visualize(patient_id, figsize=(20, 20)):\n    image_num = list(train['patient_id']==patient_id).count(True)\n    fig, ax = plt.subplots(1, image_num, figsize=figsize)\n    ax = ax.flatten()\n    path_to_dcms = \"/kaggle/input/rsna-breast-cancer-detection/train_images\"\n    for i, dcm_path in enumerate(os.listdir(os.path.join(path_to_dcms, str(patient_id)))):\n        dcm = pydicom.dcmread(os.path.join(path_to_dcms, str(patient_id), dcm_path))\n        dcm = dcm.pixel_array\n        dcm = (dcm - dcm.min()) / (dcm.max() - dcm.min()) * 255\n#         if pydicom.dcmread(os.path.join(path_to_dcms, str(patient_id), dcm_path)).PhotometricInterpretation == \"MONOCHROME1\":\n#             dcm = 1 - dcm\n        ax[i].set_title(dcm_path)\n        ax[i].imshow(dcm, cmap=\"bone\") #cmap=\"bone\" for black-white image\n#         ax[i].imshow(dcm) #original color","metadata":{"execution":{"iopub.status.busy":"2022-12-12T23:55:09.235765Z","iopub.execute_input":"2022-12-12T23:55:09.236217Z","iopub.status.idle":"2022-12-12T23:55:09.247115Z","shell.execute_reply.started":"2022-12-12T23:55:09.236186Z","shell.execute_reply":"2022-12-12T23:55:09.245603Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# visualize the patient_id images\nvisualize(10006)\nvisualize(10011)\ntrain.head(8)","metadata":{"execution":{"iopub.status.busy":"2022-12-12T23:55:11.345279Z","iopub.execute_input":"2022-12-12T23:55:11.345728Z","iopub.status.idle":"2022-12-12T23:55:34.208847Z","shell.execute_reply.started":"2022-12-12T23:55:11.345683Z","shell.execute_reply":"2022-12-12T23:55:34.207914Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Basically, each `patient_id` has 4 mammography image, <font color=red>but some patient has more than 4 images</font>.   \nSo the `visualization` function consider this by checking an [image_num].","metadata":{}},{"cell_type":"code","source":"# Example: patient_id which has more than 4 images\npatient_id = 10049\nimage_num = list(train['patient_id']==patient_id).count(True)\nimage_num","metadata":{"execution":{"iopub.status.busy":"2022-12-12T23:57:24.637789Z","iopub.execute_input":"2022-12-12T23:57:24.63837Z","iopub.status.idle":"2022-12-12T23:57:24.658946Z","shell.execute_reply.started":"2022-12-12T23:57:24.638324Z","shell.execute_reply":"2022-12-12T23:57:24.65751Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train[train['patient_id']==10049]","metadata":{"execution":{"iopub.status.busy":"2022-12-12T23:52:58.78535Z","iopub.execute_input":"2022-12-12T23:52:58.786717Z","iopub.status.idle":"2022-12-12T23:52:58.807574Z","shell.execute_reply.started":"2022-12-12T23:52:58.786679Z","shell.execute_reply":"2022-12-12T23:52:58.806235Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Findings:\n- the mammography images contain some 'blank' area (ex. laterality=Right has blank in left area.) -> The detetion of breast area is discussed in this notebook : https://www.kaggle.com/code/masatakaitakura/rsna-2022-for-beginner-detect-breast-area\n- the background color seems different depende on the `machine_id` -> <font color=red>black and white contract should be aligned for each image??</font>","metadata":{}},{"cell_type":"markdown","source":"To consider the 2nd bullet of the above findings, we re-defined the function as bellow. The only difference is just removed the comment-out in line #12 and #13.","metadata":{}},{"cell_type":"code","source":"# Function to visualize the DICOM file\n\ndef visualize(patient_id, figsize=(20, 20)):\n    image_num = list(train['patient_id']==patient_id).count(True)\n    fig, ax = plt.subplots(1, image_num, figsize=figsize)\n    ax = ax.flatten()\n    path_to_dcms = \"/kaggle/input/rsna-breast-cancer-detection/train_images\"\n    for i, dcm_path in enumerate(os.listdir(os.path.join(path_to_dcms, str(patient_id)))):\n        dcm = pydicom.dcmread(os.path.join(path_to_dcms, str(patient_id), dcm_path))\n        dcm = dcm.pixel_array\n        dcm = (dcm - dcm.min()) / (dcm.max() - dcm.min()) * 255\n        if pydicom.dcmread(os.path.join(path_to_dcms, str(patient_id), dcm_path)).PhotometricInterpretation == \"MONOCHROME1\":\n            dcm = 1 - dcm\n        ax[i].set_title(dcm_path) #original color\n        ax[i].imshow(dcm, cmap=\"bone\") #cmap=\"bone\" for black-white image\n#         ax[i].imshow(dcm)","metadata":{"execution":{"iopub.status.busy":"2022-12-12T23:57:56.509632Z","iopub.execute_input":"2022-12-12T23:57:56.510086Z","iopub.status.idle":"2022-12-12T23:57:56.521331Z","shell.execute_reply.started":"2022-12-12T23:57:56.510051Z","shell.execute_reply":"2022-12-12T23:57:56.52031Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# visualize the patient_id images\nvisualize(10006)\nvisualize(10011)","metadata":{"execution":{"iopub.status.busy":"2022-12-12T23:57:58.026302Z","iopub.execute_input":"2022-12-12T23:57:58.026969Z","iopub.status.idle":"2022-12-12T23:58:20.916257Z","shell.execute_reply.started":"2022-12-12T23:57:58.026934Z","shell.execute_reply":"2022-12-12T23:58:20.914841Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we can obtain the visualization of remaining all DICOM file images.","metadata":{}},{"cell_type":"markdown","source":"## cancer mammography\nNext we check the cancer mammography images.","metadata":{}},{"cell_type":"code","source":"visualize(10130)\nvisualize(10226)\nvisualize(1025)\nvisualize(10432)\nvisualize(10589)\ntrain_cancer.head(12)","metadata":{"execution":{"iopub.status.busy":"2022-12-12T23:58:32.567209Z","iopub.execute_input":"2022-12-12T23:58:32.567665Z","iopub.status.idle":"2022-12-12T23:59:32.373615Z","shell.execute_reply.started":"2022-12-12T23:58:32.567633Z","shell.execute_reply":"2022-12-12T23:59:32.372189Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It is difficult to find out the cancer by visual judgement...however image_id=613462606 has the \"white block area\" so this seems a cancer.\n","metadata":{}},{"cell_type":"code","source":"def visualize_single(patient_id, image_id, figsize=(20, 20)):\n    fig, ax = plt.subplots(1, 1, figsize=figsize)\n    path_to_dcms = \"/kaggle/input/rsna-breast-cancer-detection/train_images\"\n    dcm = pydicom.dcmread(os.path.join(path_to_dcms, str(patient_id), str(image_id)+'.dcm'))\n    dcm = dcm.pixel_array\n    ax.imshow(dcm)","metadata":{"execution":{"iopub.status.busy":"2022-12-12T23:59:32.376099Z","iopub.execute_input":"2022-12-12T23:59:32.376635Z","iopub.status.idle":"2022-12-12T23:59:32.386305Z","shell.execute_reply.started":"2022-12-12T23:59:32.376587Z","shell.execute_reply":"2022-12-12T23:59:32.384748Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"visualize_single(patient_id=10130, image_id=613462606)","metadata":{"execution":{"iopub.status.busy":"2022-12-12T23:59:32.38826Z","iopub.execute_input":"2022-12-12T23:59:32.388799Z","iopub.status.idle":"2022-12-12T23:59:35.499861Z","shell.execute_reply.started":"2022-12-12T23:59:32.388754Z","shell.execute_reply":"2022-12-12T23:59:35.498114Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# machine_id","metadata":{}},{"cell_type":"markdown","source":"Now we check the `machine_id` information for more detail.\nThe machine_id has total 10 groups as below.","metadata":{}},{"cell_type":"code","source":"print(f\"machine_id of all train data: {train['machine_id'].unique()}\")\nprint(f\"machine_id of cancer=1 train data{train_cancer['machine_id'].unique()}\")","metadata":{"execution":{"iopub.status.busy":"2022-12-12T23:59:35.503731Z","iopub.execute_input":"2022-12-12T23:59:35.504278Z","iopub.status.idle":"2022-12-12T23:59:35.513147Z","shell.execute_reply.started":"2022-12-12T23:59:35.504211Z","shell.execute_reply":"2022-12-12T23:59:35.511812Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"`machine_id` is unbalanced. id=49 has the largest data and id=197 has the smallest number.","metadata":{}},{"cell_type":"code","source":"# 'machine_id'\nsns.countplot(x='machine_id', data=train)\nplt.legend(loc='upper right', title='machine_id')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-12-12T23:59:35.515305Z","iopub.execute_input":"2022-12-12T23:59:35.515943Z","iopub.status.idle":"2022-12-12T23:59:35.747252Z","shell.execute_reply.started":"2022-12-12T23:59:35.515902Z","shell.execute_reply":"2022-12-12T23:59:35.74599Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Next, let's see the mammography data for each `machine_id`.","metadata":{}},{"cell_type":"code","source":"train_dropduplicates = train.drop_duplicates(subset='machine_id')\ntrain_dropduplicates","metadata":{"execution":{"iopub.status.busy":"2022-12-12T23:59:35.748953Z","iopub.execute_input":"2022-12-12T23:59:35.749414Z","iopub.status.idle":"2022-12-12T23:59:35.774465Z","shell.execute_reply.started":"2022-12-12T23:59:35.749373Z","shell.execute_reply":"2022-12-12T23:59:35.773257Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# machine_id = 29\nvisualize(10006)\nvisualize(10025)\nvisualize(10048)\ntrain[train['machine_id']==29].head(12)","metadata":{"execution":{"iopub.status.busy":"2022-12-12T23:59:35.775856Z","iopub.execute_input":"2022-12-12T23:59:35.77727Z","iopub.status.idle":"2022-12-13T00:00:27.401723Z","shell.execute_reply.started":"2022-12-12T23:59:35.777234Z","shell.execute_reply":"2022-12-13T00:00:27.400281Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# machine_id = 21\nvisualize(10011)\nvisualize(10106)\nvisualize(10144)\ntrain[train['machine_id']==21].head(12)","metadata":{"execution":{"iopub.status.busy":"2022-12-13T00:00:27.403686Z","iopub.execute_input":"2022-12-13T00:00:27.404147Z","iopub.status.idle":"2022-12-13T00:00:44.237409Z","shell.execute_reply.started":"2022-12-13T00:00:27.404107Z","shell.execute_reply":"2022-12-13T00:00:44.236141Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# machine_id = 216\nvisualize(10038)\nvisualize(10285)\nvisualize(10302)\ntrain[train['machine_id']==216].head(12)","metadata":{"execution":{"iopub.status.busy":"2022-12-13T00:00:44.239361Z","iopub.execute_input":"2022-12-13T00:00:44.239708Z","iopub.status.idle":"2022-12-13T00:00:57.632387Z","shell.execute_reply.started":"2022-12-13T00:00:44.239678Z","shell.execute_reply":"2022-12-13T00:00:57.631145Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# machine_id = 93\nvisualize(10042)\nvisualize(10215)\nvisualize(10391)\ntrain[train['machine_id']==93].head(12)","metadata":{"execution":{"iopub.status.busy":"2022-12-13T00:00:57.637445Z","iopub.execute_input":"2022-12-13T00:00:57.638019Z","iopub.status.idle":"2022-12-13T00:01:18.048629Z","shell.execute_reply.started":"2022-12-13T00:00:57.637972Z","shell.execute_reply":"2022-12-13T00:01:18.047315Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# machine_id = 49\nvisualize(10049)\nvisualize(10095)\nvisualize(10097)\ntrain[train['machine_id']==49].head(12)","metadata":{"execution":{"iopub.status.busy":"2022-12-13T00:01:18.0507Z","iopub.execute_input":"2022-12-13T00:01:18.051116Z","iopub.status.idle":"2022-12-13T00:01:49.169072Z","shell.execute_reply.started":"2022-12-13T00:01:18.051077Z","shell.execute_reply":"2022-12-13T00:01:49.167467Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# machine_id = 48\nvisualize(10124)\nvisualize(10126)\nvisualize(10136)\ntrain[train['machine_id']==48].head(12)","metadata":{"execution":{"iopub.status.busy":"2022-12-13T00:01:49.171089Z","iopub.execute_input":"2022-12-13T00:01:49.171616Z","iopub.status.idle":"2022-12-13T00:02:27.021077Z","shell.execute_reply.started":"2022-12-13T00:01:49.171551Z","shell.execute_reply":"2022-12-13T00:02:27.019888Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# machine_id = 170\nvisualize(10243)\nvisualize(10589)\nvisualize(10668)\ntrain[train['machine_id']==170].head(12)","metadata":{"execution":{"iopub.status.busy":"2022-12-13T00:02:27.022797Z","iopub.execute_input":"2022-12-13T00:02:27.023178Z","iopub.status.idle":"2022-12-13T00:02:55.909892Z","shell.execute_reply.started":"2022-12-13T00:02:27.023145Z","shell.execute_reply":"2022-12-13T00:02:55.908415Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# machine_id = 210\nvisualize(10342)\nvisualize(10438)\nvisualize(10741)\ntrain[train['machine_id']==210].head(12)","metadata":{"execution":{"iopub.status.busy":"2022-12-13T00:02:55.912171Z","iopub.execute_input":"2022-12-13T00:02:55.913058Z","iopub.status.idle":"2022-12-13T00:03:52.11109Z","shell.execute_reply.started":"2022-12-13T00:02:55.913002Z","shell.execute_reply":"2022-12-13T00:03:52.109936Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# machine_id = 190\nvisualize(10577)\nvisualize(11664)\nvisualize(21809)\ntrain[train['machine_id']==190].head(12)","metadata":{"execution":{"iopub.status.busy":"2022-12-13T00:03:52.112593Z","iopub.execute_input":"2022-12-13T00:03:52.113532Z","iopub.status.idle":"2022-12-13T00:04:04.332172Z","shell.execute_reply.started":"2022-12-13T00:03:52.113499Z","shell.execute_reply":"2022-12-13T00:04:04.330641Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# machine_id = 197\nvisualize(13365)\nvisualize(17095)\ntrain[train['machine_id']==197].head(12)","metadata":{"execution":{"iopub.status.busy":"2022-12-13T00:04:04.334333Z","iopub.execute_input":"2022-12-13T00:04:04.334799Z","iopub.status.idle":"2022-12-13T00:04:19.944621Z","shell.execute_reply.started":"2022-12-13T00:04:04.33475Z","shell.execute_reply":"2022-12-13T00:04:19.943308Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The characteristics of the image may appear different for each machine. This may be coming from the setting of each machine, so next we see the <font color=red>DICOM file metadata</font>.","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# DICOM file meta information\nDICOM file contain some file information.\nEach image contains a rich information that might help us to understand the image deeply and how we split the data for training. Some data might be provided as a train input directly to our model.","metadata":{}},{"cell_type":"markdown","source":"The transformed csv files are stored in the **[RSNA-2022 Breast Cancer : DICOM Meta Information Dataset](https://www.kaggle.com/datasets/masatakaitakura/rsna2022-breast-cancer-dicom-meta-information) and in the [Input Data]** of this notebook. Please feel free to download and use this dataset, and if these datasets and this notebook help you, <font color=red>**please upvote both dataset and notebook!!**</font> ","metadata":{}},{"cell_type":"markdown","source":"The EDA for the transformed DICOM files are prepared in this notebook : [RSNA-2022 [EDA] DICOM Meta Information](https://www.kaggle.com/masatakaitakura/rsna-2022-eda-dicom-meta-information/edit)","metadata":{}},{"cell_type":"code","source":"dicom_path = \"/kaggle/input/rsna-breast-cancer-detection/train_images/10038/1967300488.dcm\"\ndcm_info = pydicom.dcmread(dicom_path)","metadata":{"execution":{"iopub.status.busy":"2022-12-13T00:04:19.946454Z","iopub.execute_input":"2022-12-13T00:04:19.94691Z","iopub.status.idle":"2022-12-13T00:04:19.958804Z","shell.execute_reply.started":"2022-12-13T00:04:19.946864Z","shell.execute_reply":"2022-12-13T00:04:19.957495Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dcm_info","metadata":{"execution":{"iopub.status.busy":"2022-12-13T00:04:19.960081Z","iopub.execute_input":"2022-12-13T00:04:19.960396Z","iopub.status.idle":"2022-12-13T00:04:19.97174Z","shell.execute_reply.started":"2022-12-13T00:04:19.960368Z","shell.execute_reply":"2022-12-13T00:04:19.970167Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can obtain the 'keywords' for the dicom file by `.dir()` as shown below.","metadata":{"_kg_hide-input":true}},{"cell_type":"code","source":"dcm_info.dir()","metadata":{"execution":{"iopub.status.busy":"2022-12-13T00:04:19.973522Z","iopub.execute_input":"2022-12-13T00:04:19.974791Z","iopub.status.idle":"2022-12-13T00:04:19.986036Z","shell.execute_reply.started":"2022-12-13T00:04:19.974744Z","shell.execute_reply":"2022-12-13T00:04:19.984587Z"},"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"You can obtain the mete information by specifying the row title without space as show below.","metadata":{"_kg_hide-input":true}},{"cell_type":"code","source":"dcm_info_list = [\n    dcm_info.SOPInstanceUID,\n    dcm_info.ContentDate,\n    dcm_info.ContentTime,\n    dcm_info.PatientID, \n    dcm_info.BodyPartThickness,\n    dcm_info.CompressionForce,\n    dcm_info.ExposureControlMode,\n    dcm_info.ExposureControlModeDescription,\n    dcm_info.StudyInstanceUID,\n    dcm_info.SeriesInstanceUID,\n    dcm_info.InstanceNumber,\n    dcm_info.ImageLaterality,\n    dcm_info.SamplesPerPixel,\n    dcm_info.PhotometricInterpretation,\n    dcm_info.Rows,\n    dcm_info.Columns,\n    dcm_info.BitsAllocated,\n    dcm_info.BitsStored,\n    dcm_info.HighBit,\n    dcm_info.PixelRepresentation,\n    dcm_info.PixelIntensityRelationship,\n    dcm_info.PixelIntensityRelationshipSign,\n    dcm_info.WindowCenter,\n    dcm_info.WindowWidth,\n    dcm_info.RescaleIntercept,\n    dcm_info.RescaleSlope,\n    dcm_info.RescaleType,\n    dcm_info.VOILUTFunction,\n    dcm_info.LossyImageCompression,\n]","metadata":{"execution":{"iopub.status.busy":"2022-12-13T00:04:19.98762Z","iopub.execute_input":"2022-12-13T00:04:19.988004Z","iopub.status.idle":"2022-12-13T00:04:19.999807Z","shell.execute_reply.started":"2022-12-13T00:04:19.987973Z","shell.execute_reply":"2022-12-13T00:04:19.998589Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dcm_info_list","metadata":{"execution":{"iopub.status.busy":"2022-12-13T00:04:20.001314Z","iopub.execute_input":"2022-12-13T00:04:20.003166Z","iopub.status.idle":"2022-12-13T00:04:20.017147Z","shell.execute_reply.started":"2022-12-13T00:04:20.00309Z","shell.execute_reply":"2022-12-13T00:04:20.015625Z"},"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"According to the [pydicom official document](https://pydicom.github.io/pydicom/stable/auto_examples/input_output/plot_printing_dataset.html#sphx-glr-auto-examples-input-output-plot-printing-dataset-py), we can obtain the meta information in my own format.","metadata":{"_kg_hide-input":true}},{"cell_type":"code","source":"# source : https://pydicom.github.io/pydicom/stable/auto_examples/input_output/plot_printing_dataset.html#sphx-glr-auto-examples-input-output-plot-printing-dataset-py\n\ndef myprint(dataset, indent=0):\n    \"\"\"Go through all items in the dataset and print them with custom format\n\n    Modelled after Dataset._pretty_str()\n    \"\"\"\n    dont_print = ['Pixel Data', 'File Meta Information Version']\n\n    indent_string = \"   \" * indent\n    next_indent_string = \"   \" * (indent + 1)\n\n    for data_element in dataset:\n        if data_element.VR == \"SQ\":   # a sequence\n            print(indent_string, data_element.name)\n            for sequence_item in data_element.value:\n                myprint(sequence_item, indent + 1)\n                print(next_indent_string + \"---------\")\n        else:\n            if data_element.name in dont_print:\n                print(\"\"\"<item not printed -- in the \"don't print\" list>\"\"\")\n            else:\n                repr_value = repr(data_element.value)\n                if len(repr_value) > 50:\n                    repr_value = repr_value[:50] + \"...\"\n                print(\"{0:s} {1:s} = {2:s}\".format(indent_string,\n                                                   data_element.name,\n                                                   repr_value))\n\n\n# Set the dicom file path\nds = pydicom.dcmread(dicom_path)\n\nmyprint(ds)","metadata":{"execution":{"iopub.status.busy":"2022-12-13T01:22:28.225621Z","iopub.execute_input":"2022-12-13T01:22:28.227375Z","iopub.status.idle":"2022-12-13T01:22:28.25037Z","shell.execute_reply.started":"2022-12-13T01:22:28.227307Z","shell.execute_reply":"2022-12-13T01:22:28.24866Z"},"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Next, we see the difference of meta information tags name between the `machine_id`.","metadata":{"execution":{"iopub.status.busy":"2022-12-12T03:28:07.888941Z","iopub.execute_input":"2022-12-12T03:28:07.88932Z","iopub.status.idle":"2022-12-12T03:28:07.896491Z","shell.execute_reply.started":"2022-12-12T03:28:07.88929Z","shell.execute_reply":"2022-12-12T03:28:07.89485Z"}}},{"cell_type":"code","source":"# pick-up `image_id`s for each `machine_id`s\ntrain_dropdup = train.drop_duplicates(subset='machine_id')\ntrain_dropdup","metadata":{"execution":{"iopub.status.busy":"2022-12-13T02:10:35.600811Z","iopub.execute_input":"2022-12-13T02:10:35.601311Z","iopub.status.idle":"2022-12-13T02:10:35.633664Z","shell.execute_reply.started":"2022-12-13T02:10:35.601273Z","shell.execute_reply":"2022-12-13T02:10:35.632535Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# difference of meta information tags name between the `machine_id`\n\n# prepare the blank list\ndcm_info_list = []\nmachine_id_list = []\n\nfor patient_id,image_id, machine_id in zip(train_dropdup['patient_id'], train_dropdup['image_id'], train_dropdup['machine_id']):\n    # obtain the keywords for each `machine_id`\n    dicom_path = \"/kaggle/input/rsna-breast-cancer-detection/train_images/\" + str(patient_id) + \"/\" + str(image_id) + \".dcm\"\n    keywords = pydicom.dcmread(dicom_path).dir()\n    dcm_info_list.append(keywords)\n    machine_id_list.append(machine_id)\n\n# convert to df\ndcm_info_df = pd.DataFrame(dcm_info_list, index = machine_id_list)\ndcm_info_df","metadata":{"execution":{"iopub.status.busy":"2022-12-13T02:11:12.268394Z","iopub.execute_input":"2022-12-13T02:11:12.26887Z","iopub.status.idle":"2022-12-13T02:11:12.405901Z","shell.execute_reply.started":"2022-12-13T02:11:12.268832Z","shell.execute_reply":"2022-12-13T02:11:12.405072Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"You can see the `machine_id`:170 has the most number of dicom meta data keywords. Please note this is the spot-check for each machine_id so this may not applicable for all other remaining data.  \nHowever, we can say <font color=red>each data have different number of dicom meta information</font>.","metadata":{}},{"cell_type":"markdown","source":"## Transform DICOM meta information to DataFrame\nNext, let's obtain these meta information and create the dataframe for some train data.","metadata":{}},{"cell_type":"code","source":"# difference of meta information tags name between the `machine_id`\n\n# prepare the blank dataframe\ndcm_df = pd.DataFrame(columns=['name'])\n\nfor patient_id,image_id, machine_id in zip(train_dropdup['patient_id'], train_dropdup['image_id'], train_dropdup['machine_id']):\n    # obtain the keywords for each `machine_id`\n    dicom_path = \"/kaggle/input/rsna-breast-cancer-detection/train_images/\" + str(patient_id) + \"/\" + str(image_id) + \".dcm\"\n\n    # create df for each meta information\n    dataset = pydicom.dcmread(dicom_path)\n    dcm_list = []\n\n    for data_element in dataset:\n        dcm_list.append([data_element.name, data_element.value])\n\n    # delete the last meta information of [Pixel Data]\n    dcm_list = dcm_list[:-1]\n    \n    # insert `machine_id` into the list\n    dcm_list.append(['machine_id', machine_id])\n\n    # convert to df\n    dcm_df_ind = pd.DataFrame(dcm_list, columns = ['name', image_id])\n    \n    # combine each df\n    dcm_df = pd.merge(dcm_df, dcm_df_ind, on='name', how='outer')\n\ndcm_df","metadata":{"execution":{"iopub.status.busy":"2022-12-13T02:35:18.606144Z","iopub.execute_input":"2022-12-13T02:35:18.606644Z","iopub.status.idle":"2022-12-13T02:35:18.743302Z","shell.execute_reply.started":"2022-12-13T02:35:18.606597Z","shell.execute_reply":"2022-12-13T02:35:18.741995Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Findings:\n- each data have different number of dicom meta information.\n- for further report, please see the following notebooks\n- [RSNA-2022 [EDA] DICOM Meta Information](https://www.kaggle.com/masatakaitakura/rsna-2022-eda-dicom-meta-information/edit)\n- [RSNA-2022 Breast Cancer : DICOM Meta Information Dataset](https://www.kaggle.com/datasets/masatakaitakura/rsna2022-breast-cancer-dicom-meta-information)","metadata":{}},{"cell_type":"markdown","source":"Separately, I transformed the DICOM file meta information to DataFrame. If you interested in that, please see - [RSNA-2022 [EDA] DICOM Meta Information](https://www.kaggle.com/masatakaitakura/rsna-2022-eda-dicom-meta-information/edit)\nand [RSNA-2022 Breast Cancer : DICOM Meta Information Dataset](https://www.kaggle.com/datasets/masatakaitakura/rsna2022-breast-cancer-dicom-meta-information).","metadata":{}},{"cell_type":"markdown","source":"Thank you very much for reading, I hope this notebook helps you!\n\n**If you enjoyed the notebook, please upvote! 🙏 Thank you, appreciate your support!**\n","metadata":{}}]}