{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"#  <center><font size = 3><span style=\"color:#38a3a5\"> <p style=\"background-color:#90e0ef;font-family:newtimeroman;color:#03045e;font-size:200%;text-align:center;border-radius:100px 10px;\">INTRODUCTION</p>   </span></font></center>\n \n<p style=\"font-famly:Roboto;font-size:130%;color:#a04070;margin-up:0\"> Before jumping into the modelling part of any problem it is necessary to understand the data first and when it is of size 460GB then it certainly takes times to uncover the different nuances of the dataset. In this notebook I have tried to perform and EDA on the dataset given so, that we can match the different aspects of the data given in the different csv files.</p>\n\n<!-- <a id='top'></a> -->\n<div class=\"list-group\" id=\"list-tab\" role=\"tablist\">\n <center><font size = 3><span style=\"color:#38a3a5\"> <p style=\"background-color:#90e0ef;font-family:newtimeroman;color:#03045e;font-size:200%;text-align:center;border-radius:100px 10px;\">Table of Contents</p>   </span></font></center> \n    \n* [1. Train](#1)\n    \n* [2. Train Series Meta](#2)\n    \n* [3. Train Dicom Tags](#3)\n    \n* [4. Segmentations](#4)\n    \n* [5. Train Images](#5)","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport seaborn as sb\nimport matplotlib.pyplot as plt\nimport nibabel as nib\n\nimport os\nimport pydicom\nfrom glob import glob\nfrom tqdm import tqdm, trange","metadata":{"execution":{"iopub.status.busy":"2023-08-01T10:36:45.685902Z","iopub.execute_input":"2023-08-01T10:36:45.686637Z","iopub.status.idle":"2023-08-01T10:36:46.348811Z","shell.execute_reply.started":"2023-08-01T10:36:45.686587Z","shell.execute_reply":"2023-08-01T10:36:46.347842Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id='1'></a>\n#  <center><font size = 3><span style=\"color:#a8dadc\"> <p style=\"background-color:#90e0ef;font-family:newtimeroman;color:#03045e;font-size:200%;text-align:center;border-radius:100px 10px;\">1. Train </p>   </span></font></center>\n\n<div class=\"alert alert-block alert-info\" style=\"font-size:14px; font-family:verdana; line-height: 1.7em;\">\n    * <b>{bowel/extravasation}_{healthy/injury}</b> - The two injury types with binary targets.<br>\n    * <b>{kidney/liver/spleen}_{healthy/low/high}</b> - The three injury types with three target levels based on the severity of the injury.<br>\n    * <b>any_injury</b> - If the value of binary target corresponding to any injury(no matter whether it is high or low) is 1 then this column will have value as equal to 1 showing that there is abnormality.<br>\n</div>","metadata":{}},{"cell_type":"code","source":"PATH = '/kaggle/input/rsna-2023-abdominal-trauma-detection'\ntrain = pd.read_csv(f'{PATH}/train.csv')\ntrain.head()","metadata":{"execution":{"iopub.status.busy":"2023-08-01T10:36:46.353568Z","iopub.execute_input":"2023-08-01T10:36:46.35435Z","iopub.status.idle":"2023-08-01T10:36:46.380045Z","shell.execute_reply.started":"2023-08-01T10:36:46.354321Z","shell.execute_reply":"2023-08-01T10:36:46.379016Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.shape","metadata":{"execution":{"iopub.status.busy":"2023-08-01T10:36:46.381357Z","iopub.execute_input":"2023-08-01T10:36:46.381863Z","iopub.status.idle":"2023-08-01T10:36:46.387441Z","shell.execute_reply.started":"2023-08-01T10:36:46.381834Z","shell.execute_reply":"2023-08-01T10:36:46.386507Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<p style=\"font-famly:Roboto;font-size:130%;color:#a04070;margin-up:0\"> What is the meaning of <b>Extravasation</b>?<br>\nThe leakage of blood from a vessel into tissues surrounding it. This can occur in injuries or burns or allergic reactions.</p>","metadata":{}},{"cell_type":"code","source":"temp = train['any_injury'].value_counts()\ntemp.plot(kind='pie', autopct='%1.1f%%')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-01T10:36:46.389668Z","iopub.execute_input":"2023-08-01T10:36:46.389966Z","iopub.status.idle":"2023-08-01T10:36:46.572638Z","shell.execute_reply.started":"2023-08-01T10:36:46.389941Z","shell.execute_reply":"2023-08-01T10:36:46.571184Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<p style=\"font-famly:Roboto;font-size:130%;color:#a04070;margin-up:0\">Let's find out which type of injury is more frequent in the dataset at hand. If there is a case of Data Imbalance then we will see some difference else none.</p>","metadata":{}},{"cell_type":"code","source":"_, ax = plt.subplots(2, 3, figsize=(10, 5))\nnum_image = 0\nfor col in train.columns:\n    if col.endswith('_healthy'):\n        row = num_image//3\n        cols = num_image%3\n        temp = train[col].value_counts()\n        temp.plot(kind='pie', autopct='%1.1f%%',\n                  ax=ax[row, cols])\n        num_image+=1\nax[1, 2].axis('off')\nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-01T10:36:47.031039Z","iopub.execute_input":"2023-08-01T10:36:47.031412Z","iopub.status.idle":"2023-08-01T10:36:47.660191Z","shell.execute_reply.started":"2023-08-01T10:36:47.031381Z","shell.execute_reply":"2023-08-01T10:36:47.659057Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<p style=\"font-famly:Roboto;font-size:130%;color:#a04070;margin-up:0\">Most of the cases of injury or unhealthy organs are found in <b>liver</b> and <b>spleen</b>(This contains WBC to fight germs in the blood) in the dataset.</p>","metadata":{}},{"cell_type":"markdown","source":"<a id='2'></a>\n#  <center><font size = 3><span style=\"color:#a8dadc\"> <p style=\"background-color:#90e0ef;font-family:newtimeroman;color:#03045e;font-size:200%;text-align:center;border-radius:100px 10px;\">2. Train Series Meta </p>   </span></font></center>","metadata":{}},{"cell_type":"code","source":"train_series_meta = pd.read_csv(f'{PATH}/train_series_meta.csv')\ntrain_series_meta.head()                                ","metadata":{"execution":{"iopub.status.busy":"2023-08-01T10:37:00.001174Z","iopub.execute_input":"2023-08-01T10:37:00.00161Z","iopub.status.idle":"2023-08-01T10:37:00.02285Z","shell.execute_reply.started":"2023-08-01T10:37:00.001574Z","shell.execute_reply":"2023-08-01T10:37:00.021717Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_series_meta.shape","metadata":{"execution":{"iopub.status.busy":"2023-08-01T10:37:00.220672Z","iopub.execute_input":"2023-08-01T10:37:00.221086Z","iopub.status.idle":"2023-08-01T10:37:00.227235Z","shell.execute_reply.started":"2023-08-01T10:37:00.22105Z","shell.execute_reply":"2023-08-01T10:37:00.22581Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_series_meta[train_series_meta['patient_id']==10004]","metadata":{"execution":{"iopub.status.busy":"2023-08-01T10:37:00.40113Z","iopub.execute_input":"2023-08-01T10:37:00.401519Z","iopub.status.idle":"2023-08-01T10:37:00.413403Z","shell.execute_reply.started":"2023-08-01T10:37:00.401468Z","shell.execute_reply":"2023-08-01T10:37:00.412253Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Background for Multiphasic CT Scans:\n<p style=\"font-famly:Roboto;font-size:130%;color:#a04070;margin-up:0\">\nIn Multiphasic Scans particular area is scanned in the different phases of contrast enhancement. For the contrast enhancement some contrast agents which are generally <b>Iodine-Based</b> are introduced in the Blood vessels to get better clarity of the vessels, organs and tissues.<br><br>\n    So, the patients with two scans can have different values of the <b>aortic_hu</b> but how is this possible if the person is same then the volume of aaorta should be same as well but the difference can be observed due to the introduction of the contrast enhancements for better scans and accurate measurements.<br><br>\nIt is also given that out of the two values higher one will show the late arterial phase that is when the <b>contrast enhancing agent</b> has reached the desired loaction in the body.\n</p>","metadata":{}},{"cell_type":"code","source":"temp = train_series_meta['patient_id'].value_counts()\ntemp.value_counts().plot.bar()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-01T10:37:07.761289Z","iopub.execute_input":"2023-08-01T10:37:07.761686Z","iopub.status.idle":"2023-08-01T10:37:07.963457Z","shell.execute_reply.started":"2023-08-01T10:37:07.761652Z","shell.execute_reply":"2023-08-01T10:37:07.962333Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<p style=\"font-famly:Roboto;font-size:130%;color:#a04070;margin-up:0\">So, every patient has at max two sessions of scans and we have equal number of patients with one and two sessions.\n</p>","metadata":{}},{"cell_type":"code","source":"sb.histplot(train_series_meta['aortic_hu'], kde=True)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-01T10:37:40.801251Z","iopub.execute_input":"2023-08-01T10:37:40.801684Z","iopub.status.idle":"2023-08-01T10:37:41.328612Z","shell.execute_reply.started":"2023-08-01T10:37:40.801646Z","shell.execute_reply":"2023-08-01T10:37:41.32751Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_series_meta[train_series_meta['aortic_hu']<0]","metadata":{"execution":{"iopub.status.busy":"2023-08-01T10:38:09.300721Z","iopub.execute_input":"2023-08-01T10:38:09.301121Z","iopub.status.idle":"2023-08-01T10:38:09.313247Z","shell.execute_reply.started":"2023-08-01T10:38:09.301091Z","shell.execute_reply":"2023-08-01T10:38:09.312007Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<p style=\"font-famly:Roboto;font-size:130%;color:#a04070;margin-up:0\">Seems like this is an outlier or by mistake there is a negative sign in the value.\n</p>","metadata":{}},{"cell_type":"code","source":"sb.histplot(np.abs(train_series_meta['aortic_hu']), \n            kde=True)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-01T10:38:26.671601Z","iopub.execute_input":"2023-08-01T10:38:26.672028Z","iopub.status.idle":"2023-08-01T10:38:27.076948Z","shell.execute_reply.started":"2023-08-01T10:38:26.671994Z","shell.execute_reply":"2023-08-01T10:38:27.075723Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_series_meta['incomplete_organ'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2023-08-01T10:38:27.380964Z","iopub.execute_input":"2023-08-01T10:38:27.381353Z","iopub.status.idle":"2023-08-01T10:38:27.390099Z","shell.execute_reply.started":"2023-08-01T10:38:27.381314Z","shell.execute_reply":"2023-08-01T10:38:27.388963Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<p style=\"font-famly:Roboto;font-size:130%;color:#a04070;margin-up:0\">For the most of the cases scans include complete organs.\n</p>","metadata":{}},{"cell_type":"markdown","source":"<a id='3'></a>\n#  <center><font size = 3><span style=\"color:#a8dadc\"> <p style=\"background-color:#90e0ef;font-family:newtimeroman;color:#03045e;font-size:200%;text-align:center;border-radius:100px 10px;\">3. Train DICOM Tags </p>   </span></font></center>","metadata":{}},{"cell_type":"code","source":"train_dicom_tags = pd.read_parquet(f'{PATH}/train_dicom_tags.parquet')\ntrain_dicom_tags.head()","metadata":{"execution":{"iopub.status.busy":"2023-08-01T10:38:34.50073Z","iopub.execute_input":"2023-08-01T10:38:34.501541Z","iopub.status.idle":"2023-08-01T10:38:40.267122Z","shell.execute_reply.started":"2023-08-01T10:38:34.50148Z","shell.execute_reply":"2023-08-01T10:38:40.265833Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_dicom_tags.shape","metadata":{"execution":{"iopub.status.busy":"2023-08-01T10:38:40.270972Z","iopub.execute_input":"2023-08-01T10:38:40.271355Z","iopub.status.idle":"2023-08-01T10:38:40.278067Z","shell.execute_reply.started":"2023-08-01T10:38:40.27132Z","shell.execute_reply":"2023-08-01T10:38:40.277031Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(glob(f'{PATH}/train_images/*/*/*.dcm'))","metadata":{"execution":{"iopub.status.busy":"2023-08-01T10:38:40.279405Z","iopub.execute_input":"2023-08-01T10:38:40.279795Z","iopub.status.idle":"2023-08-01T10:45:19.599682Z","shell.execute_reply.started":"2023-08-01T10:38:40.279763Z","shell.execute_reply":"2023-08-01T10:45:19.598754Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<p style=\"font-famly:Roboto;font-size:130%;color:#a04070;margin-up:0\">Here we can observe that the number of the DICOM Tags rows are approximately equal to the number of scans we have been provided with.\n</p>","metadata":{}},{"cell_type":"code","source":"train_dicom_tags.iloc[0]","metadata":{"execution":{"iopub.status.busy":"2023-08-01T10:45:19.601339Z","iopub.execute_input":"2023-08-01T10:45:19.601867Z","iopub.status.idle":"2023-08-01T10:45:19.610524Z","shell.execute_reply.started":"2023-08-01T10:45:19.601835Z","shell.execute_reply":"2023-08-01T10:45:19.609648Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<p style=\"font-famly:Roboto;font-size:130%;color:#a04070;margin-up:0\">Here we can extract the series number so, that we can get the unique dicom image data for each <b>patient_id</b> and <b>series_id</b>.\n</p>","metadata":{}},{"cell_type":"code","source":"temp = train_dicom_tags['SeriesInstanceUID'].str.split('.', expand=True)\ntrain_dicom_tags['series_id'] = temp[8]\ntrain_dicom_tags.head(2)","metadata":{"execution":{"iopub.status.busy":"2023-08-01T10:45:19.611643Z","iopub.execute_input":"2023-08-01T10:45:19.612523Z","iopub.status.idle":"2023-08-01T10:45:25.386147Z","shell.execute_reply.started":"2023-08-01T10:45:19.612471Z","shell.execute_reply":"2023-08-01T10:45:25.38499Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id='4'></a>\n#  <center><font size = 3><span style=\"color:#a8dadc\"> <p style=\"background-color:#90e0ef;font-family:newtimeroman;color:#03045e;font-size:200%;text-align:center;border-radius:100px 10px;\">4. Segmentations</p>   </span></font></center>\n### The NIfTI (Neuroimaging Informatics Technology Initiative) format, often denoted by the .nii extension, is a widely used file format for storing and sharing neuroimaging data. We can also find the extensive metadata that has been stored in this file format.","metadata":{}},{"cell_type":"markdown","source":"\n<p style=\"font-famly:Roboto;font-size:130%;color:#a04070;margin-up:0\"><b>TotalSegmentator: Robust Segmentation of 104 Anatomic Structures in CT Images</b><br>\nAll the segmentation that are provided here are extracted from another model that is trained to identify:</p>\n    <ul style=\"font-famly:Roboto;font-size:130%;color:#a04070;margin-up:0\">\n        <li>27 Organs</li>\n        <li>59 Bones</li>\n        <li>10 Muscles</li>\n        <li>8 Vessels</li>\n    </ul>\n<p style=\"font-famly:Roboto;font-size:130%;color:#a04070;margin-up:0\">That sums upto 104 different categories.</p>","metadata":{}},{"cell_type":"code","source":"segmentations = glob(f'{PATH}/segmentations/*')\nlen(segmentations)","metadata":{"execution":{"iopub.status.busy":"2023-08-01T10:46:37.940862Z","iopub.execute_input":"2023-08-01T10:46:37.941286Z","iopub.status.idle":"2023-08-01T10:46:38.000536Z","shell.execute_reply.started":"2023-08-01T10:46:37.941252Z","shell.execute_reply":"2023-08-01T10:46:37.999664Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.shape","metadata":{"execution":{"iopub.status.busy":"2023-08-01T10:46:39.210694Z","iopub.execute_input":"2023-08-01T10:46:39.211615Z","iopub.status.idle":"2023-08-01T10:46:39.219695Z","shell.execute_reply.started":"2023-08-01T10:46:39.211565Z","shell.execute_reply":"2023-08-01T10:46:39.218529Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<p style=\"font-famly:Roboto;font-size:130%;color:#a04070;margin-up:0\">So out of the 3147 patients we have been provided with the segmentations of the 206 patients only. But can we extract the segmentations for the remaining images by using the segmentor model by using which the provided segmentations has been extracted.\n</p>","metadata":{}},{"cell_type":"code","source":"seg_series_id = os.listdir(f'{PATH}/segmentations')\nseg_series_id = [int(val[:-4]) for val in seg_series_id]\nseg_series_id[:5]","metadata":{"execution":{"iopub.status.busy":"2023-08-01T10:46:43.260629Z","iopub.execute_input":"2023-08-01T10:46:43.261736Z","iopub.status.idle":"2023-08-01T10:46:43.269458Z","shell.execute_reply.started":"2023-08-01T10:46:43.261686Z","shell.execute_reply":"2023-08-01T10:46:43.268527Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = pd.merge(train_series_meta, train,\n              how='inner', on='patient_id')\ndf['has_seg'] = False\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2023-08-01T10:46:43.450607Z","iopub.execute_input":"2023-08-01T10:46:43.450973Z","iopub.status.idle":"2023-08-01T10:46:43.47574Z","shell.execute_reply.started":"2023-08-01T10:46:43.450945Z","shell.execute_reply":"2023-08-01T10:46:43.474537Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for val in tqdm(seg_series_id): df.loc[df['series_id']==val, 'has_seg'] = True\n\ndf.groupby(['any_injury', 'has_seg']).count()['patient_id']","metadata":{"execution":{"iopub.status.busy":"2023-08-01T10:47:29.920727Z","iopub.execute_input":"2023-08-01T10:47:29.921129Z","iopub.status.idle":"2023-08-01T10:47:29.934223Z","shell.execute_reply.started":"2023-08-01T10:47:29.921095Z","shell.execute_reply":"2023-08-01T10:47:29.933066Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sb.countplot(data=df, x='any_injury', hue='has_seg')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-01T10:47:48.741767Z","iopub.execute_input":"2023-08-01T10:47:48.742131Z","iopub.status.idle":"2023-08-01T10:47:48.96346Z","shell.execute_reply.started":"2023-08-01T10:47:48.742104Z","shell.execute_reply":"2023-08-01T10:47:48.962279Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<p style=\"font-famly:Roboto;font-size:130%;color:#a04070;margin-up:0\">If there is no injury then there is no menaing for getting the segmenattion of those cases because we would like to train a model which can predict which type of injury it is.\n</p>","metadata":{}},{"cell_type":"code","source":"# Check the shape\nnifti_file = nib.load(f'{PATH}/segmentations/10000.nii')\nmask = nifti_file.get_fdata()\nprint(mask.shape)","metadata":{"execution":{"iopub.status.busy":"2023-08-01T10:49:01.680974Z","iopub.execute_input":"2023-08-01T10:49:01.681407Z","iopub.status.idle":"2023-08-01T10:49:04.743055Z","shell.execute_reply.started":"2023-08-01T10:49:01.681371Z","shell.execute_reply":"2023-08-01T10:49:04.741857Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<p style=\"font-famly:Roboto;font-size:130%;color:#a04070;margin-up:0\">Here a point to observe is that for 483 scans we have 483 segmentations as well. But are they in order like for the first image in the .dcm files and the mask[ : , : , 0] belong to that image or not.\n</p>","metadata":{}},{"cell_type":"code","source":"print(nifti_file)","metadata":{"execution":{"iopub.status.busy":"2023-08-01T10:49:04.745978Z","iopub.execute_input":"2023-08-01T10:49:04.747031Z","iopub.status.idle":"2023-08-01T10:49:04.753626Z","shell.execute_reply.started":"2023-08-01T10:49:04.74698Z","shell.execute_reply":"2023-08-01T10:49:04.752561Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id='5'></a>\n#  <center><font size = 3><span style=\"color:#a8dadc\"> <p style=\"background-color:#90e0ef;font-family:newtimeroman;color:#03045e;font-size:200%;text-align:center;border-radius:100px 10px;\">5. Train Images</p></span></font></center>","metadata":{}},{"cell_type":"code","source":"df['num_scans'] = 0\nfor idx in tqdm(df.index):\n    temp = df.iloc[idx]\n    count = len(glob(f'{PATH}/train_images/{temp.patient_id}/{temp.series_id}/*.dcm'))\n    df.loc[idx, 'num_scans'] = count","metadata":{"execution":{"iopub.status.busy":"2023-08-01T10:49:53.800984Z","iopub.execute_input":"2023-08-01T10:49:53.801382Z","iopub.status.idle":"2023-08-01T10:50:03.402221Z","shell.execute_reply.started":"2023-08-01T10:49:53.80135Z","shell.execute_reply":"2023-08-01T10:50:03.400944Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df['num_scans'].mean(), df['num_scans'].mode()[0], df['num_scans'].median()","metadata":{"execution":{"iopub.status.busy":"2023-08-01T10:50:03.404537Z","iopub.execute_input":"2023-08-01T10:50:03.405851Z","iopub.status.idle":"2023-08-01T10:50:03.413454Z","shell.execute_reply.started":"2023-08-01T10:50:03.405814Z","shell.execute_reply":"2023-08-01T10:50:03.412558Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sb.histplot(df['num_scans'], kde=True)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-01T10:50:03.414778Z","iopub.execute_input":"2023-08-01T10:50:03.415865Z","iopub.status.idle":"2023-08-01T10:50:03.740398Z","shell.execute_reply.started":"2023-08-01T10:50:03.415831Z","shell.execute_reply":"2023-08-01T10:50:03.73878Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<p style=\"font-famly:Roboto;font-size:130%;color:#a04070;margin-up:0\">Bimodal distribution this means that generally images have around 200 ct_scans per patient or 700.\n</p>","metadata":{}},{"cell_type":"code","source":"df[df['series_id']==10000]","metadata":{"execution":{"iopub.status.busy":"2023-08-01T10:50:03.742543Z","iopub.execute_input":"2023-08-01T10:50:03.742895Z","iopub.status.idle":"2023-08-01T10:50:03.760694Z","shell.execute_reply.started":"2023-08-01T10:50:03.742864Z","shell.execute_reply":"2023-08-01T10:50:03.759562Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<p style=\"font-famly:Roboto;font-size:130%;color:#a04070;margin-up:0\">First segmentation which is shown in the segmentation directory so, we will try to explore this only.\n</p>","metadata":{}},{"cell_type":"code","source":"df[df['patient_id']==54722]","metadata":{"execution":{"iopub.status.busy":"2023-08-01T10:50:04.787162Z","iopub.execute_input":"2023-08-01T10:50:04.787831Z","iopub.status.idle":"2023-08-01T10:50:04.811473Z","shell.execute_reply.started":"2023-08-01T10:50:04.787778Z","shell.execute_reply":"2023-08-01T10:50:04.809772Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<p style=\"font-famly:Roboto;font-size:130%;color:#a04070;margin-up:0\">So, we will explore this 343 aortic_hu that is for the series_id 10000 scans as they will have better visibility than the other one.\n</p>","metadata":{}},{"cell_type":"code","source":"# idx = has_seg_idx[np.random.randint(len(has_seg_idx))]\nidx = 3602\ntemp = df.iloc[idx]\nimg_paths = glob(f'{PATH}/train_images/{temp.patient_id}/{temp.series_id}/*.dcm')\nimg_paths.sort()\nlen(img_paths)","metadata":{"execution":{"iopub.status.busy":"2023-08-01T10:50:05.870885Z","iopub.execute_input":"2023-08-01T10:50:05.871279Z","iopub.status.idle":"2023-08-01T10:50:05.881566Z","shell.execute_reply.started":"2023-08-01T10:50:05.871247Z","shell.execute_reply":"2023-08-01T10:50:05.880237Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(7, 7))\ndicom_file = pydicom.dcmread(img_paths[150])\nct_scan = dicom_file.pixel_array\nprint(np.min(ct_scan), np.max(ct_scan))\n\n# Bringing the image in scale of 0 to 255.\nct_scan = np.interp(ct_scan, [np.min(ct_scan), np.max(ct_scan)], [0,255])\nplt.imshow(ct_scan, cmap=plt.cm.gray)\nplt.axis('off')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-01T10:50:06.990929Z","iopub.execute_input":"2023-08-01T10:50:06.991344Z","iopub.status.idle":"2023-08-01T10:50:07.236607Z","shell.execute_reply.started":"2023-08-01T10:50:06.991311Z","shell.execute_reply":"2023-08-01T10:50:07.235493Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(dicom_file)","metadata":{"execution":{"iopub.status.busy":"2023-08-01T10:50:09.976041Z","iopub.execute_input":"2023-08-01T10:50:09.976438Z","iopub.status.idle":"2023-08-01T10:50:09.983723Z","shell.execute_reply.started":"2023-08-01T10:50:09.976406Z","shell.execute_reply":"2023-08-01T10:50:09.982539Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Why generally medical Imaging files are in the .dcm format?\n<p style=\"font-famly:Roboto;font-size:130%;color:#a04070;margin-up:0\">The answer to this question is that the images in the .dcm format contains metadata about the image related to the patient as well as the configuration that has been used while taking that particular image.<br><br>\n    Converting from the .dcm format to the common image formats like .png or .jpg help us save space by avoiding the metadata that has been there in the .dcm file. Also some times it is being done to introduce the anonymity in teh files so, that the patients personal details are not disclosed.\n</p>","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(7, 7))\nplt.imshow(mask[:,:,150])\nplt.axis('off')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-01T10:50:14.500638Z","iopub.execute_input":"2023-08-01T10:50:14.501017Z","iopub.status.idle":"2023-08-01T10:50:14.664798Z","shell.execute_reply.started":"2023-08-01T10:50:14.500987Z","shell.execute_reply":"2023-08-01T10:50:14.663633Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"_, ax = plt.subplots(1, 2, figsize=(10, 10))\n\nax[0].imshow(ct_scan, cmap=plt.cm.gray)\n\nax[1].imshow(mask[:,:,150])\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-01T10:50:14.681246Z","iopub.execute_input":"2023-08-01T10:50:14.682079Z","iopub.status.idle":"2023-08-01T10:50:15.117645Z","shell.execute_reply.started":"2023-08-01T10:50:14.682028Z","shell.execute_reply":"2023-08-01T10:50:15.116552Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<p style=\"font-famly:Roboto;font-size:130%;color:#a04070;margin-up:0\">As of now it doesn't seem like they are in order because when I tried to create overlay image using the scan and the mask they do not align properly.\n</p>","metadata":{}},{"cell_type":"markdown","source":"# Sample Submission","metadata":{}},{"cell_type":"code","source":"ss = pd.read_csv(f'{PATH}/sample_submission.csv')\nss.head()","metadata":{"execution":{"iopub.status.busy":"2023-08-01T10:50:56.086719Z","iopub.execute_input":"2023-08-01T10:50:56.087133Z","iopub.status.idle":"2023-08-01T10:50:56.112501Z","shell.execute_reply.started":"2023-08-01T10:50:56.087098Z","shell.execute_reply":"2023-08-01T10:50:56.11127Z"},"trusted":true},"execution_count":null,"outputs":[]}]}