{"cells":[{"metadata":{},"cell_type":"markdown","source":"# * Pulmonary Embolism Detection - EDA [Beginner-Friendly] *\n![LEGO](https://a360-rtmagazine.s3.amazonaws.com/wp-content/uploads/2019/10/lung-pulmo-embolism-1500-1200x799.jpg)\n\n**If you like this kernel, an up-vote would be appreciated**"},{"metadata":{},"cell_type":"markdown","source":"# Introduction"},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","collapsed":true,"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":false},"cell_type":"markdown","source":"This is an EDA (exploratory data analysis) for the newly launched RNSA Embolism Detection competition on Kaggle that we're going to be working on today. An embolism is caused when your arteries are blocked off in your lung, preventing blood flow and stopping your lung from getting the oxygen it needs to carry out respiration. It is the most fatal cardiovascular disease in the United States of America (60,000 to 100,000 deaths per annum)."},{"metadata":{"trusted":true},"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\n%matplotlib inline\nimport seaborn as sns\nimport glob\nimport pydicom as dcm","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# 1- Loading the dataset"},{"metadata":{"trusted":true},"cell_type":"code","source":"train = pd.read_csv(\"../input/rsna-str-pulmonary-embolism-detection/train.csv\")\ntrain.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# 2- Exploring columns"},{"metadata":{"trusted":true},"cell_type":"code","source":"train.dtypes","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"| Field Name | Feature Type | Description |\n| :--- | :--- | :---- |\n| **StudyInstanceUID** | UID | unique ID for each study (exam) in the data. |\n| **SeriesInstanceUID** | UID | unique ID for each series within the study. |\n| **SOPInstanceUID** | UID |  unique ID for each image within the study (and data) |\n| **pe_present_on_image**|  image-level | notes whether any form of PE is present on the image.|\n| **negative_exam_for_pe**|  exam-level | whether there are any images in the study that have PE present.|\n| **qa_motion** |  informational | indicates whether radiologists noted an issue with motion in the study. | \n|**qa_contrast** |  informational | indicates whether radiologists noted an issue with contrast in the study.|\n| **flow_artifact** | informational | ---|\n| **rv_lv_ratio_gte_1** | exam-level| indicates whether the RV/LV ratio present in the study is >= 1|\n| **rv_lv_ratio_lt_1** | exam-level| indicates whether the RV/LV ratio present in the study is < 1|\n| **leftsided_pe** | exam-level | indicates that there is PE present on the left side of the images in the study| \n| **chronic_pe**  | exam-level | indicates that the PE in the study is chronic|\n| **true_filling_defect_not_pe** | informational | indicates a defect that is NOT PE|\n| **rightsided_pe** | exam-level | indicates that there is PE present on the right side of the images in the study|\n| **acute_and_chronic_pe** | exam-level| indicates that the PE present in the study is both acute AND chronic|\n| **central_pe** | exam-level| indicates that there is PE present in the center of the images in the study|\n| **indeterminate**  | exam-level| indicates that while the study is not negative for PE, an ultimate set of exam-level labels could not be created, due to QA issues|\n"},{"metadata":{},"cell_type":"markdown","source":"**So, What are we going to predict?**\n\n> Every study / exam has a row for each label that is scored (detailed in the Data page). It is uniquely indicated by the StudyInstanceUID. Every image, further, has a row for the PE Present on Image label and is uniquely indicated by the SOPInstanceUID. Your prediction file should have a number of rows equal to: (number of images) + (number of studies * number of scored labels).\n\nIn other words:\n- For each image, we're going to predict the column \"pe_present_on_image\"\n- For each \"StudyInstanceUID\" we're goin to predict:"},{"metadata":{},"cell_type":"markdown","source":"1. negative_exam_for_pe\n2. rv_lv_ratio_gte_1\n3. rv_lv_ratio_lt_1\n4. chronic_pe\n5. true_filling_defect_not_pe\n6. acute_and_chronic_pe\n7. rightsided_pe\n8. leftsided_pe\n9. central_pe"},{"metadata":{},"cell_type":"markdown","source":"# 3- Data visualization"},{"metadata":{"trusted":true},"cell_type":"code","source":"train_cols = ['pe_present_on_image', 'negative_exam_for_pe', 'qa_motion',\n       'qa_contrast', 'flow_artifact', 'rv_lv_ratio_gte_1', 'rv_lv_ratio_lt_1',\n       'leftsided_pe', 'chronic_pe', 'true_filling_defect_not_pe',\n       'rightsided_pe', 'acute_and_chronic_pe', 'central_pe', 'indeterminate']\n\ndef plot_grid(cols = train_cols):\n    fig=plt.figure(figsize=(12, 12))\n    columns = 3\n    rows = 5\n    for i in range(1, columns*rows):\n        col = cols[i-1]\n        fig.add_subplot(rows, columns, i)\n        train[col].value_counts().plot(kind = \"bar\")\n        indices = train[col].value_counts().index.tolist()\n        count_0 = train[col].value_counts()[0]\n        count_1 = train[col].value_counts()[1]\n        plt.xlabel(f\"{col}\\n {indices[0]}: {count_0}\\n {indices[1]}: {count_1}\")\n    plt.tight_layout()\n    plt.show()\n\nplot_grid()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We can note the following observations:\n\n- By looking at the images, only 5.4% of patients have active cases of PE\n- But exam shows that 32.4% of the images involve active cases of PE, which means that image-based prediction gives many false-negatives\n- There are almost no issues related to motion (0.8%) / contrast in the studies (1.6%)\n- RV/LV values are not consistent\n- Most defects are right-sided or left-sided. There are only few cases of central cases of PE"},{"metadata":{},"cell_type":"markdown","source":"![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F115173%2Fa2a5ee66b5799274141dd547cc3ea466%2FPE%20figure.jpg?generation=1599575183749576&alt=media)\n\nCredits for this schema goes to @redwankarimsony from his amazing notebook: [RSNA-STR Pulmonary Embolism](https://www.kaggle.com/redwankarimsony/rsna-str-pulmonary-embolism-eda)"},{"metadata":{},"cell_type":"markdown","source":"**Correlation matrix between the trainable columns**"},{"metadata":{"trusted":true},"cell_type":"code","source":"corr_mat = train[train_cols].corr()\nmask = np.triu(np.ones_like(corr_mat, dtype=bool))\nf, ax = plt.subplots(figsize=(14, 12))\nsns.heatmap(corr_mat, mask = mask, annot = True, vmax = 0.3, square = False, linewidths = 0.5, center = 0)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# 4- Image visualization and interpretation"},{"metadata":{"trusted":true},"cell_type":"code","source":"def read_dicom(file_path, show = False, cmap = 'gray'):\n    im = dcm.dcmread(file_path)\n    image = im.pixel_array\n    if show:\n        plt.imshow(image, cmap = 'gray')\n    return image","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def show_images_with_specific_condition(column_name):\n    \n    rows = 2\n    cols = 5\n    \n    train_with_condition = train[train[column_name] == 0]\n    train_with_condition = train_with_condition.sample(n = rows*cols) \n    train_image_file_paths = []\n    for _, entry in train_with_condition.iterrows():\n        train_image_file_paths.append('../input/rsna-str-pulmonary-embolism-detection/train/'+str(entry['StudyInstanceUID'])+'/'+str(entry['SeriesInstanceUID'])+'/'+str(entry['SOPInstanceUID'])+'.dcm')\n    counter  = 0\n    fig = plt.figure(figsize=(25,15))\n    fig.suptitle('Samples with ' + column_name + ' = 0', fontsize=40)\n    for path in train_image_file_paths:\n        fig.add_subplot(rows, cols, counter+1)\n        plt.imshow(read_dicom(path), cmap='gray')\n        plt.axis(False)\n        fig.add_subplot\n        counter += 1\n    \n    \n    train_with_condition = train[train[column_name] == 1]\n    train_with_condition = train_with_condition.sample(n = rows*cols) \n    train_image_file_paths = []\n    for _, entry in train_with_condition.iterrows():\n        train_image_file_paths.append('../input/rsna-str-pulmonary-embolism-detection/train/'+str(entry['StudyInstanceUID'])+'/'+str(entry['SeriesInstanceUID'])+'/'+str(entry['SOPInstanceUID'])+'.dcm')\n    counter  = 0\n    fig = plt.figure(figsize=(25,15))\n    fig.suptitle('Samples with ' + column_name + ' = 1', fontsize=40)\n    for path in train_image_file_paths:\n        fig.add_subplot(rows, cols, counter+1)\n        plt.imshow(read_dicom(path), cmap='gray')\n        plt.axis(False)\n        fig.add_subplot\n        counter += 1","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**4-1. Images with qa_motion**"},{"metadata":{"trusted":true},"cell_type":"code","source":"show_images_with_specific_condition('qa_motion')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**4-2. Images with qa_contrast**"},{"metadata":{"trusted":true},"cell_type":"code","source":"show_images_with_specific_condition('qa_contrast')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**4-3. Images with indeterminate**"},{"metadata":{"trusted":true},"cell_type":"code","source":"show_images_with_specific_condition('indeterminate')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**4-4. Images with rightsided_pe**"},{"metadata":{"trusted":true},"cell_type":"code","source":"show_images_with_specific_condition('rightsided_pe')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**4-5. Images with leftsided_pe**"},{"metadata":{"trusted":true},"cell_type":"code","source":"show_images_with_specific_condition('leftsided_pe')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**4-6. Images with central_pe**"},{"metadata":{"trusted":true},"cell_type":"code","source":"show_images_with_specific_condition('central_pe')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# **[KERNEL UNDER CONSTRUCTION]**"}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}