{"cells":[{"metadata":{"_uuid":"ad9e7850-3e9d-4d00-a472-b3e0193c4e16","_cell_guid":"dc3ad3ec-f203-40dd-b953-59a091068207","trusted":true},"cell_type":"markdown","source":"This is my first experience in kaggle, and I am not native speaker of English (and python).\nI examined whether 'image position' in DICOM file has any information or predictive value for hemorhage detection. \nI thoght specific type of hemohage may occour more frequently in spcific Z position (hight) of the head.\nCode is heavily borrowed from https://www.kaggle.com/braquino/correct-images-sequece\n"},{"metadata":{"_uuid":"3de0722f-ee3d-457a-8777-8b50dbcf03e0","_cell_guid":"a65387c2-dde3-4c8e-a574-ee02d36ac463","trusted":true},"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport os\nimport torch, torch.nn as nn\nfrom torchvision import models, transforms, datasets\nimport matplotlib.pyplot as plt\nimport torch.nn.functional as F\nimport math\nimport pydicom\nfrom collections import Counter\nimport tqdm\nos.listdir('/kaggle/input/rsna-intracranial-hemorrhage-detection')\n\n# Any results you write to the current directory are saved as output.","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3171da5c-fe86-40cb-b206-957dffabb1b9","_cell_guid":"b1a6e392-7580-494e-8b3c-3de6c3d3ec58","trusted":true},"cell_type":"code","source":"input_path = '/kaggle/input/rsna-intracranial-hemorrhage-detection/'\ndf_train = pd.read_csv(input_path + 'stage_1_train.csv')\nprint(len(df_train))\ndf_train.head(10)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8831767e-de56-4b83-b16a-79adc90ac265","_cell_guid":"01dabe90-a499-48cd-83a7-4318aa860646","trusted":true},"cell_type":"code","source":"df_test = pd.read_csv(input_path + 'stage_1_sample_submission.csv')\nprint(len(df_test))\ndf_test.head(10)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"80ce89a7-4143-4367-8b6e-ae43823d3f7e","_cell_guid":"64a85455-fa7d-434e-aabc-37aca11c6c36","trusted":true},"cell_type":"code","source":"ds_columns = ['ID',\n              'PatientID',\n              'Modality',\n              'StudyInstance',\n              'SeriesInstance',\n                'PhotoInterpretation',\n              'Position0', 'Position1', 'Position2',\n              'Orientation0', 'Orientation1', 'Orientation2', 'Orientation3', 'Orientation4', 'Orientation5',\n              'PixelSpacing0', 'PixelSpacing1']\ndef extract_dicom_features(ds):\n    \n    ds_items = [ds.SOPInstanceUID,\n                ds.PatientID,\n                ds.Modality,\n                ds.StudyInstanceUID,\n                ds.SeriesInstanceUID,\n                ds.PhotometricInterpretation,\n                ds.ImagePositionPatient,\n                ds.ImageOrientationPatient,\n                ds.PixelSpacing]\n\n    line = []\n    for item in ds_items:\n        if type(item) is pydicom.multival.MultiValue:\n            line += [float(x) for x in item]\n        else:\n            line.append(item)\n\n    return line","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ee29912b-d0d1-4269-9f93-4eeac2ea4b55","_cell_guid":"4d243c0f-ce4f-4b59-bad9-0dfc85359229","trusted":true},"cell_type":"code","source":"list_img = os.listdir(input_path + 'stage_1_test_images')\nprint(len(list_img))\ndf_features = []\nfor img in tqdm.tqdm(list_img):\n    img_path = input_path + 'stage_1_test_images/' + img\n    ds = pydicom.read_file(img_path)\n    df_features.append(extract_dicom_features(ds))\ndf_features_test = pd.DataFrame(df_features, columns=ds_columns)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8a82973a-987f-4693-a08a-88d248627246","_cell_guid":"f80f8339-aa66-4e59-9143-c46747597104","trusted":true},"cell_type":"code","source":"list_img = os.listdir(input_path + 'stage_1_train_images')\nprint(len(list_img))\ndf_features = []\nfor img in tqdm.tqdm(list_img):\n    img_path = input_path + 'stage_1_train_images/' + img\n    ds = pydicom.read_file(img_path)\n    df_features.append(extract_dicom_features(ds))\ndf_features_train = pd.DataFrame(df_features, columns=ds_columns)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"00b1611d-6ac8-43b7-b0c3-e15fab537530","_cell_guid":"757c8d9d-02a4-49f1-9480-6a4e12609377","trusted":true},"cell_type":"code","source":"df_train[['ID', 'Subtype']] = df_train['ID'].str.rsplit(pat='_', n=1, expand=True)\ndf_train_new = df_train.pivot_table(index='ID', columns='Subtype').reset_index()\ndf_train_new.columns=['ID','any','epidural','intraparenchymal','intraventricular','subarachnoid','subdural']\ndf_train_merged = df_train_new.merge(df_features_train, how='right')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d39fed8a-4d23-4330-b471-ca5bd2bdeaf0","_cell_guid":"01538117-9521-43de-9102-eec82e9abb90","trusted":true},"cell_type":"code","source":"df_test[['ID', 'Subtype']] = df_test['ID'].str.rsplit(pat='_', n=1, expand=True)\ndf_test_new = df_test.pivot_table(index='ID', columns='Subtype').reset_index()\ndf_test_new.columns=['ID','any','epidural','intraparenchymal','intraventricular','subarachnoid','subdural']\ndf_test_merged = df_test_new.merge(df_features_test, how='right')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c7590d12-f864-49ea-92ca-297b8ad66b04","_cell_guid":"f68485e7-4496-4765-9666-c0a81251105e","trusted":true},"cell_type":"markdown","source":"Script to get z position relative to entire volume."},{"metadata":{"trusted":true},"cell_type":"code","source":"df_train_merged.to_csv('df_train_merged.csv', index=False)\ndf_test_merged.to_csv('df_test_merged.csv', index=False)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c7016dd5-b5e4-4d40-93e1-ac12040eea4e","_cell_guid":"ecd0c7dc-42cf-41ad-92b2-4de3e787e210","trusted":true},"cell_type":"code","source":"for df in [df_train_merged, df_test_merged]:\n    df['max_Position2']=0\n    df['min_Position2']=0\n    df['relative_Position2']=0\n    \n    for SeriesInstance in tqdm.tqdm(df[\"SeriesInstance\"].unique()):\n        sub_df=df[df['SeriesInstance']==SeriesInstance]\n        df.loc[df['SeriesInstance']==SeriesInstance, 'max_Position2'] = max(sub_df['Position2'])\n        df.loc[df['SeriesInstance']==SeriesInstance, 'min_Position2'] = min(sub_df['Position2'])\n\n    df['relative_Position2'] = (df['Position2']-df['min_Position2'])/(df['max_Position2']-df['min_Position2'])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5b36de5c-434c-4317-89af-d00a7179229a","_cell_guid":"de21062c-5429-47e2-a538-7b02fcb254ac","trusted":true},"cell_type":"markdown","source":"Sort and save as csv"},{"metadata":{"_uuid":"c402a89a-b547-4024-94db-d3a5246fc901","_cell_guid":"34d9a4f3-460d-4afb-9bf0-44f63121b442","trusted":true},"cell_type":"code","source":"df_train_merged=df_train_merged.sort_values(by=[\"PatientID\",\"SeriesInstance\",'Position2'])\ndf_train_merged.to_csv('df_train_merged.csv', index=False)\ndf_test_merged=df_test_merged.sort_values(by=[\"PatientID\",\"SeriesInstance\",'Position2'])\ndf_test_merged.to_csv('df_test_merged.csv', index=False)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"67478435-5edf-45ab-825d-e26a8e334143","_cell_guid":"e3375a2a-7696-4e15-a01f-58692de2788c","trusted":true},"cell_type":"markdown","source":"Show distribution of each type of hemorhage relative to Z position of slice"},{"metadata":{"_uuid":"0b044f47-8129-4e22-bfc2-fd341da6665b","_cell_guid":"508d4c3a-6aab-4264-8b0f-21f4dbc19252","trusted":true},"cell_type":"code","source":"z_normal = df_train_merged[df_train_merged[\"any\"] != 1]['relative_Position2']\nz_epidural = df_train_merged[df_train_merged[\"epidural\"] == 1]['relative_Position2']\nz_subdural = df_train_merged[df_train_merged[\"subdural\"] == 1]['relative_Position2']\nz_intraventricular = df_train_merged[df_train_merged[\"intraventricular\"] == 1]['relative_Position2']\nz_intraparenchymal = df_train_merged[df_train_merged[\"intraparenchymal\"] == 1]['relative_Position2']\nz_subarachnoid = df_train_merged[df_train_merged[\"subarachnoid\"] == 1]['relative_Position2']\nz_any = df_train_merged[df_train_merged[\"any\"] == 1]['relative_Position2']\n\nz_describe=pd.DataFrame({'normal': z_normal.describe(),\n                   'epidural': z_epidural.describe(),\n                   'subdural': z_subdural.describe(),\n                   'intraventricular': z_intraventricular.describe(),\n                   'intraparenchymal': z_intraparenchymal.describe(),\n                   'subarachnoid': z_subarachnoid.describe(),\n                   'any': z_any.describe()})\nprint(z_describe)\nbin_width=0.05\nkwargs = dict(histtype='stepfilled', normed=False, bins=np.arange(0, 1+bin_width, bin_width))\nfig, axes = plt.subplots(4, 2, figsize=(8, 16))\naxes[0, 0].hist(z_normal, **kwargs)\naxes[0, 0].set_title('normal')\naxes[0, 1].hist(z_epidural, **kwargs) \naxes[0, 1].set_title('epidural')\naxes[1, 0].hist(z_subdural, **kwargs) \naxes[1, 0].set_title('subdural')\naxes[1, 1].hist(z_intraventricular, **kwargs) \naxes[1, 1].set_title('intraventricular')\naxes[2, 0].hist(z_intraparenchymal, **kwargs) \naxes[2, 0].set_title('intraparenchymal')\naxes[2, 1].hist(z_subarachnoid, **kwargs) \naxes[2, 1].set_title('subarachnoid')\naxes[3, 0].hist(z_any, **kwargs) \naxes[3, 0].set_title('any')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"From these distributions, I can see..\n* Hemohages are NOT unifomly distributed, they often exist at the center (Z = 0.5) of each volume.\n* Uppermost (Z = 1.0) and lowermost (Z = 0) slices are likely to be normal.\n* Unfortunately, the difference among the distributions of each type of hemorhage is not so clear\n* Intraventricular hemorhage showed smaller variance, it occured more frequently in the center of the head.\n\nI am not medical doctor/radiotechnologist, but I guess the reason for these results are...\n* When doctors or radiotechnologists scan patients, they would postion the patients in a scanner so that the suspicious hemorhage point is at the center of field of view.\n* Intraventricular hemorhage exist only in the ventricule (ofcourse), and the ventricule exist in the center of the head."},{"metadata":{},"cell_type":"markdown","source":"I also noticed that  the slice is more likely to be abnormal if its adjacent slices are not normal.\nI don't know how to formulate such phenomenon (autoregression?), but I can show example.\nI can see slices with hemorhage in 'ID_2f48a87008' are aggregated in the center of the head."},{"metadata":{"trusted":true},"cell_type":"code","source":" print(df_train_merged[df_train_merged[\"SeriesInstance\"] == 'ID_2f48a87008'])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Maybe, 2.5D deep learning (a slice and adjacent ones as input) or 3D deeplearning can utilize such characteristics.\n\nAny comments are welcome"}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":1}