{"cells":[{"metadata":{},"cell_type":"markdown","source":"Here is some initial exploratory data analysis I did for [RSNA Intracranial Hemorrhage Detection competition](https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection).\n\nBits and pieces were taken from [Marco Vasquez's EDA kernel](https://www.kaggle.com/marcovasquez/basic-eda-data-visualization), so give him a shoutout! That said, a vast majority of the code is my own, simply to provide a better understanding of the data. I will likely update this notebook as more insights become apparent.\n\nAs I have no medical expertise, I will avoid discussing medical aspecs of the data. A vast majority of the medical imaging I am unable to read, and I will be using strictly ML methods to detect any anomolies."},{"metadata":{},"cell_type":"markdown","source":"# Training Labels\n\nLet us take a look at our training labels."},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import numpy as np\nimport pandas as pd\n\ndf = pd.read_csv('../input/rsna-intracranial-hemorrhage-detection/stage_1_train.csv')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"df.head(10)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"As described in the \"Data\" section of this competition, the column `ID` is broken up as follows:\n\n```\nID_(Patient ID)_(Hemorrhage Type)\n```\n\n* We want to get this in a slightly different form, for ease of lookup"},{"metadata":{"trusted":true},"cell_type":"code","source":"def get_id(s):\n    s = s.split('_')\n    return 'ID_' + s[1]\n\ndef get_hemorrhage_type(s):\n    s = s.split('_')\n    return s[2]\n    \ndf['Image_ID'] = df.ID.apply(get_id)\ndf['Hemhorrhage_Type'] = df.ID.apply(get_hemorrhage_type)\n\ndf.set_index('Image_ID')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df.head(10)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"All I did here is added two new columns. `Image_ID` (which makes loading a medical image in memory relatively easy) and `Hemhorrhage_Type`, which will tell us what we're looking for."},{"metadata":{},"cell_type":"markdown","source":"# Medical Images\n\nNow let us take a look at our medical images."},{"metadata":{"trusted":true},"cell_type":"code","source":"from os import listdir\nfrom os.path import isfile, join\nfrom pathlib import Path\n\n# To read medical images.\nimport pydicom\n\nimport matplotlib.pyplot as plt\n\ntrain_images_dir = Path('../input/rsna-intracranial-hemorrhage-detection/stage_1_test_images/')\ntrain_images = [str(train_images_dir / f) for f in listdir(train_images_dir) if isfile(train_images_dir / f)]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let us take a look at one image to see what we're dealing with."},{"metadata":{"trusted":true},"cell_type":"code","source":"train_images[0]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"As you can see, images are stored in the following format.\n\n```\nstage_1_test_images/ID_(Patient Id).dcm\n```\n\nThis is percisely the reason why we generated the new columns in the previous section. `.dcm` files, also called DICOM images, is a [specialized format for storing medical images](https://en.wikipedia.org/wiki/DICOM).\n\nWe can load them up using the [pydicom library](https://pydicom.github.io/pydicom/stable/index.html)."},{"metadata":{"trusted":true},"cell_type":"code","source":"ds = pydicom.dcmread(train_images[0])\nim = ds.pixel_array\n\nplt.imshow(im, cmap=plt.cm.gist_gray);","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"fig=plt.figure(figsize=(15, 10))\ncolumns = 5; rows = 4\nfor i in range(1, columns*rows +1):\n    ds = pydicom.dcmread(train_images[i])\n    fig.add_subplot(rows, columns, i)\n    plt.imshow(ds.pixel_array, cmap=plt.cm.gist_gray)\n    fig.add_subplot","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Statistics\n\nNow, let us get a better idea of some of the data we're working with"},{"metadata":{"trusted":true},"cell_type":"code","source":"import seaborn as sns\n\nsns.countplot(df.Label)","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":1}