{"cells":[{"metadata":{},"cell_type":"markdown","source":"# RSNA-STR Pulmonary Embolism Detection\n\n\n\n##### File descriptions\n\n- test - all test images\n- train - all train images (note that your submission kernels will NOT have access to this set of images, so you must build your models elsewhere and incorporate them into your submissions)\n- sample_submission.csv - contains rows for each UID+label combination that requires a prediction. Therefore it has a row for each image (for which you will be predicting the existence of a pulmonary embolism within the image) and row for each study+label that requires a study-level prediction.\n- train.csv - contains UIDs and all labels.\n- test.csv - contains UIDs.\n\n### Data fields\n\n- StudyInstanceUID - unique ID for each study (exam) in the data.\n- SeriesInstanceUID - unique ID for each series within the study.\n- SOPInstanceUID - unique ID for each image within the study (and data).\n- pe_present_on_image - image-level, notes whether any form of PE is present on the image.\n- negative_exam_for_pe - exam-level, whether there are any images in the study that have PE present.\n- qa_motion - informational, indicates whether radiologists noted an issue with motion in the study.\n- qa_contrast - informational, indicates whether radiologists noted an issue with contrast in the study.\n- flow_artifact - informational\n- rv_lv_ratio_gte_1 - exam-level, indicates whether the RV/LV ratio present in the study is >= 1\n- rv_lv_ratio_lt_1 - exam-level, indicates whether the RV/LV ratio present in the study is < 1\n- leftsided_pe - exam-level, indicates that there is PE present on the left side of the images in the study\n- chronic_pe - exam-level, indicates that the PE in the study is chronic\n- true_filling_defect_not_pe - informational, indicates a defect that is NOT PE\n- rightsided_pe - exam-level, indicates that there is PE present on the right side of the images in the study\n- acute_and_chronic_pe - exam-level, indicates that the PE present in the study is both acute AND chronic\n- central_pe - exam-level, indicates that there is PE present in the center of the images in the study\n- indeterminate -exam-level, indicates that while the study is not negative for PE, an ultimate set of exam-level labels could not be created, due to QA issues\n"},{"metadata":{"trusted":true},"cell_type":"code","source":"# let us install gdcm library \n!conda install -c conda-forge gdcm -y","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n\nimport os\nimport pydicom as dcm\nimport plotly.express as px\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport glob\nimport gdcm\nfrom matplotlib import animation, rc\n\nimport matplotlib\n%matplotlib inline\nmatplotlib.use(\"Agg\")\nimport matplotlib.pyplot as plt\nimport matplotlib.animation as animation\nTRAIN_DIR = \"../input/rsna-str-pulmonary-embolism-detection/train/\"\nfiles = glob.glob('../input/rsna-str-pulmonary-embolism-detection/train/*/*/*.dcm')\n\nrc('animation', html='jshtml')\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train = pd.read_csv('/kaggle/input/rsna-str-pulmonary-embolism-detection/train.csv')\ntest = pd.read_csv('/kaggle/input/rsna-str-pulmonary-embolism-detection/test.csv')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test.head()\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def bar_plot(column_name):\n    ds = train[column_name].value_counts().reset_index()\n    ds.columns = ['Values', 'Total Number']\n    fig = px.bar(\n        ds, \n        y='Values', \n        x=\"Total Number\", \n        orientation='h', \n        title='Bar plot of: ' + column_name,\n        width=600,\n        height=400\n    )\n    fig.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"col = train.columns\ncol","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"col[0+3]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# columns distribution "},{"metadata":{"trusted":true},"cell_type":"code","source":"len(col)-3","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"for i in range(len(col)-3):\n    bar_plot(col[i+3])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Columns and non-zero/zero samples"},{"metadata":{"trusted":true},"cell_type":"code","source":"# drop the first column ('sig_id'), and \ndf = train.drop(['StudyInstanceUID', 'SeriesInstanceUID', 'SOPInstanceUID'], axis=1).sum(axis=0).sort_values(ascending=False).reset_index()\ndf.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"\ndf.columns = ['column', 'nonzero_records']\nfig = px.bar(\n    df, \n    y='nonzero_records', \n    x='column', \n    orientation='v', \n    title='Columns and non zero samples', \n    height=500, \n    width=1000\n)\nfig.show()\n\n# drop the first column ('sig_id') and count the 0s in \ndf1 = train.drop(['StudyInstanceUID', 'SeriesInstanceUID', 'SOPInstanceUID'], axis=1).sum(axis=0).sort_values(ascending=False).reset_index()\ndf1.columns = ['column', 'zero_records']\ndf1['zero_records'] = len(train) -  df1['zero_records']\n# plot the bar \n\nfig = px.bar(\n    df1.head(50), \n    y='zero_records', \n    x='column', \n    orientation='v', \n    title='Columns with the zero samples ', \n    height=500, \n    width=1000\n)\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Let us check correlation: "},{"metadata":{"trusted":true},"cell_type":"code","source":"corr = train.corr()\ncorr.style.background_gradient(cmap='coolwarm')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Let us see the scans and save it as gif "},{"metadata":{},"cell_type":"markdown","source":"### Let's do animation (Inspired from https://www.kaggle.com/isaienkov/pulmonary-embolism-detection-eda)"},{"metadata":{},"cell_type":"markdown","source":"## glob module used here: \nThe glob module finds all the pathnames matching a specified pattern according to the rules used by the Unix shell, although results are returned in arbitrary order\n\n"},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"scans = glob.glob('/kaggle/input/rsna-str-pulmonary-embolism-detection/train/*/*/')\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def read_scan(path):\n    fragments = glob.glob(path + '/*')\n    \n    slices = []\n    for f in fragments:\n        img = dcm.dcmread(f)\n        img_data = img.pixel_array\n        length = int(img.InstanceNumber)\n        slices.append((length, img_data))\n    slices.sort()\n    return [s[1] for s in slices]\n\ndef animate(ims):\n    fig = plt.figure(figsize=(11,11))\n    plt.axis('off')\n    im = plt.imshow(ims[0], cmap='gray')\n\n    def animate_func(i):\n        im.set_array(ims[i])\n        return [im]\n\n    anim = animation.FuncAnimation(fig, animate_func, frames = len(ims), interval = 1000//24)\n    \n    return anim","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"movie = animate(read_scan(scans[1]))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"movie","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"movie.save('Test.gif', dpi=80, writer='imagemagick')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}