{"cells":[{"metadata":{},"cell_type":"markdown","source":"![image](https://prod-images-static.radiopaedia.org/images/17056972/2ef2ae41c8a0b5a070aa21140a14e0_gallery.jpeg)"},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","collapsed":true,"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":false},"cell_type":"markdown","source":"# Have a look\n\n* [Code Requirements](https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/overview/code-requirements) say that no internet access enabled.Let's turn off the Internet.\n\n# What is Pulmonary Embolism ?\n* Wikipedia says that [Pulmonary embolism (PE)](https://en.wikipedia.org/wiki/Pulmonary_embolism) is a blockage of an artery in the lungs by a substance that has moved from elsewhere in the body through the bloodstream.There is also a simple information about pulmonary thromboembolism posted in [this thread](https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/182376) as well."},{"metadata":{"trusted":true},"cell_type":"code","source":"from IPython.display import HTML","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"HTML('<iframe width=\"800\" height=\"500\" src=\"https://www.youtube.com/embed/4C6BB56fG1M\" frameborder=\"0\" allow=\"accelerometer; autoplay; encrypted-media; gyroscope; picture-in-picture\" allowfullscreen></iframe>')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Set up"},{"metadata":{"trusted":true},"cell_type":"code","source":"# linear algebra\nimport numpy as np\n# data processing, CSV file I/O (e.g. pd.read_csv)\nimport pandas as pd\n#Unix commands\nimport os\n\n# import useful tools\nfrom glob import glob\nfrom PIL import Image\nimport cv2\nimport pydicom\nimport scipy.ndimage\nfrom skimage import measure \nfrom mpl_toolkits.mplot3d.art3d import Poly3DCollection\nfrom skimage.morphology import disk, opening, closing\nfrom tqdm import tqdm\nfrom os import listdir, mkdir\n\nfrom PIL import Image\n\n\n# import data visualization\nimport matplotlib.pyplot as plt\nimport matplotlib.patches as patches\nimport seaborn as sns\n\nfrom bokeh.plotting import figure\nfrom bokeh.io import output_notebook, show, output_file\nfrom bokeh.models import ColumnDataSource, HoverTool, Panel\nfrom bokeh.models.widgets import Tabs\n\n# import data augmentation\nimport albumentations as albu\n\n# import math module\nimport math\n\n#Libraries\nimport pandas_profiling\nimport xgboost as xgb\nfrom sklearn.metrics import log_loss\nfrom sklearn.preprocessing import LabelEncoder\nfrom sklearn import preprocessing\nfrom sklearn.model_selection import KFold\nfrom sklearn.tree import DecisionTreeRegressor\n\n#used for changing color of text in print statement\nfrom colorama import Fore, Back, Style\ny_ = Fore.YELLOW\nr_ = Fore.RED\ng_ = Fore.GREEN\nb_ = Fore.BLUE\nm_ = Fore.MAGENTA\nsr_ = Style.RESET_ALL\n\n# One-hot encoding\nfrom sklearn.model_selection import cross_val_score\nfrom sklearn.ensemble import RandomForestRegressor","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Loading data"},{"metadata":{"trusted":true},"cell_type":"code","source":"# Setup the paths to train and test images\nDATASET = '../input/rsna-str-pulmonary-embolism-detection'\nTEST_DIR = '../input/rsna-str-pulmonary-embolism-detection/test/'\nTRAIN_DIR = '../input/rsna-str-pulmonary-embolism-detection/'\nSCAN_DIR = '../input/pulmonary-embolism-ct-data/'","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Display some of the training data\ntrain = pd.read_csv(TRAIN_DIR + \"train.csv\")\ntrain.head(10).style.applymap(lambda x: 'background-color:lightsteelblue')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print(f\"{b_}Number of rows in sample data: {r_}{train.shape[0]}\\n{b_}Number of columns in sample data: {r_}{train.shape[1]}\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Display some of the training data\ntrain.info()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Display some of the scan data\nscan = pd.read_csv(SCAN_DIR + \"Pulmonary_Embolism_CT_scans_data.csv\")\nscan.head(5).style.applymap(lambda x: 'background-color:lightsteelblue')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print(f\"{b_}Number of rows in sample data: {r_}{scan.shape[0]}\\n{b_}Number of columns in sample data: {r_}{scan.shape[1]}\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sample = pd.read_csv(TRAIN_DIR + \"sample_submission.csv\")\n# Confirmation of the format of samples for submission\nsample.head(3).style.applymap(lambda x: 'background-color:lightsteelblue')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* You'll find that the format of giving each ID a numerical value as a label is good.\n* [The page of this competition](https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/data) says that we have to predict a number of labels, at both the image and study level. "},{"metadata":{},"cell_type":"markdown","source":"# Checking data statistics"},{"metadata":{},"cell_type":"markdown","source":"* StudyInstanceUID - unique ID for each study (exam) in the data.\n* SeriesInstanceUID - unique ID for each series within the study.\n* SOPInstanceUID - unique ID for each image within the study (and data)."},{"metadata":{"trusted":true},"cell_type":"code","source":"print('The number of SOPInstanceUID is ' + str(len(train['SOPInstanceUID'].unique())))\nprint('The number of StudyInstanceUID is ' + str(len(train['StudyInstanceUID'].unique())))\nprint('The number of SeriesInstanceUID is ' + str(len(train['SeriesInstanceUID'].unique())))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The relationship between these IDs has been discussed in [this thread](https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/182736)."},{"metadata":{"trusted":true},"cell_type":"code","source":"# display IDs of the training data without duplicates\nprint(train['SOPInstanceUID'].drop_duplicates())\nprint(train['StudyInstanceUID'].drop_duplicates())\nprint(train['SeriesInstanceUID'].drop_duplicates())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Display of test data\ntest = pd.read_csv(TRAIN_DIR + \"test.csv\")\ntest.head(10).style.applymap(lambda x: 'background-color:lightsteelblue')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* You can see that there are 3 rows of IDs and 14 other rows of IDs."},{"metadata":{"trusted":true},"cell_type":"code","source":"# Display of test data\ntest.info(10)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Check for missing values in the training data\ntrain.isnull().sum()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* Therefore, we can conclude that there is no missing training data"},{"metadata":{"trusted":true},"cell_type":"code","source":"# coding: utf-8\nfrom tqdm import tqdm\nimport time\n\n# Set the total value \nbar = tqdm(total = 1000)\n# Add description\nbar.set_description('Progress rate')\nfor i in range(100):\n    # Set the progress\n    bar.update(25)\n    time.sleep(1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Let's check the max value and the max value for pe_present_on_image\nprint(\"Minimum number of value for pe_present_on_image is: {}\".format(train['pe_present_on_image'].min()), \"\\n\" +\n      \"Maximum number of value for pe_present_on_image is: {}\".format(train['pe_present_on_image'].max() ))\n# Let's check the max value and the max value for negative_exam_for_pe\nprint(\"Minimum number of value for negative_exam_for_pe is: {}\".format(train['negative_exam_for_pe'].min()), \"\\n\" +\n      \"Maximum number of value for negative_exam_for_pe is: {}\".format(train['negative_exam_for_pe'].max() ))\n# Let's check the max value and the max value for qa_motion\nprint(\"Minimum number of value for qa_motion is: {}\".format(train['qa_motion'].min()), \"\\n\" +\n      \"Maximum number of value for qa_motion is: {}\".format(train['qa_motion'].max() ))\n# Let's check the max value and the max value for qa_contrast\nprint(\"Minimum number of value for qa_contrast is: {}\".format(train['qa_contrast'].min()), \"\\n\" +\n      \"Maximum number of value for qa_contrast is: {}\".format(train['qa_contrast'].max() ))\n# Let's check the max value and the max value for flow_artifact\nprint(\"Minimum number of value for flow_artifact is: {}\".format(train['flow_artifact'].min()), \"\\n\" +\n      \"Maximum number of value for flow_artifact is: {}\".format(train['flow_artifact'].max() ))\n# Let's check the max value and the max value for rv_lv_ratio_gte_1\nprint(\"Minimum number of value for rv_lv_ratio_gte_1 is: {}\".format(train['rv_lv_ratio_gte_1'].min()), \"\\n\" +\n      \"Maximum number of value for rv_lv_ratio_gte_1 is: {}\".format(train['rv_lv_ratio_gte_1'].max() ))\n# Let's check the max value and the max value for rv_lv_ratio_lt_1\nprint(\"Minimum number of value for rv_lv_ratio_lt_1 is: {}\".format(train['rv_lv_ratio_lt_1'].min()), \"\\n\" +\n      \"Maximum number of value for rv_lv_ratio_lt_1 is: {}\".format(train['rv_lv_ratio_lt_1'].max() ))\n# Let's check the max value and the max value for leftsided_pe\nprint(\"Minimum number of value for leftsided_pe is: {}\".format(train['leftsided_pe'].min()), \"\\n\" +\n      \"Maximum number of value for leftsided_pe is: {}\".format(train['leftsided_pe'].max() ))\n# Let's check the max value and the max value for true_filling_defect_not_pe\nprint(\"Minimum number of value for true_filling_defect_not_pe is: {}\".format(train['true_filling_defect_not_pe'].min()), \"\\n\" +\n      \"Maximum number of value for true_filling_defect_not_pe is: {}\".format(train['true_filling_defect_not_pe'].max() ))\n# Let's check the max value and the max value for rightsided_pe\nprint(\"Minimum number of value for rightsided_pe is: {}\".format(train['rightsided_pe'].min()), \"\\n\" +\n      \"Maximum number of value for rightsided_pe is: {}\".format(train['rightsided_pe'].max() ))\n# Let's check the max value and the max value for central_pe\nprint(\"Minimum number of value for central_pe is: {}\".format(train['central_pe'].min()), \"\\n\" +\n      \"Maximum number of value for central_pe is: {}\".format(train['central_pe'].max() ))\n# Let's check the max value and the max value for indeterminate\nprint(\"Minimum number of value for rightsided_pe is: {}\".format(train['indeterminate'].min()), \"\\n\" +\n      \"Maximum number of value for rightsided_pe is: {}\".format(train['indeterminate'].max() ))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* We found that a minimum value of all column values except ID were 0.At the same time,we found that a maximum value of all column values except ID were 1."},{"metadata":{},"cell_type":"markdown","source":"* pe_present_on_image - image-level, notes whether any form of PE is present on the image."},{"metadata":{"trusted":true},"cell_type":"code","source":"# Draw a pie chart about pe_present_on_image.\nplt.pie(train[\"pe_present_on_image\"].value_counts(),labels=[\"0\",\"1\"],autopct=\"%.1f%%\")\nplt.title(\"Ratio of pe_present_on_image\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* qa_motion - informational, indicates whether radiologists noted an issue with motion in the study."},{"metadata":{"trusted":true},"cell_type":"code","source":"# Draw a pie chart about qa_motion.\nplt.pie(train[\"qa_motion\"].value_counts(),labels=[\"0\",\"1\"],autopct=\"%.1f%%\")\nplt.title(\"Ratio of qa_motion\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* negative_exam_for_pe - exam-level, whether there are any images in the study that have PE present."},{"metadata":{"trusted":true},"cell_type":"code","source":"# Draw a pie chart about negative_exam_for_pe.\nplt.pie(train[\"negative_exam_for_pe\"].value_counts(),labels=[\"0\",\"1\"],autopct=\"%.1f%%\")\nplt.title(\"Ratio of negative_exam_for_pe\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* qa_contrast - informational, indicates whether radiologists noted an issue with contrast in the study."},{"metadata":{"trusted":true},"cell_type":"code","source":"# Draw a pie chart about qa_contrast.\nplt.pie(train[\"qa_contrast\"].value_counts(),labels=[\"0\",\"1\"],autopct=\"%.1f%%\")\nplt.title(\"Ratio of qa_contrast\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* flow_artifact - informational"},{"metadata":{"trusted":true},"cell_type":"code","source":"# Draw a pie chart about flow_artifact.\nplt.pie(train[\"flow_artifact\"].value_counts(),labels=[\"0\",\"1\"],autopct=\"%.1f%%\")\nplt.title(\"Ratio of flow_artifact\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* rv_lv_ratio_gte_1 - exam-level, indicates whether the RV/LV ratio present in the study is >= 1"},{"metadata":{"trusted":true},"cell_type":"code","source":"# Draw a pie chart about rv_lv_ratio_gte_1.\nplt.pie(train[\"rv_lv_ratio_gte_1\"].value_counts(),labels=[\"0\",\"1\"],autopct=\"%.1f%%\")\nplt.title(\"Ratio of rv_lv_ratio_gte_1\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* rv_lv_ratio_lt_1 - exam-level, indicates whether the RV/LV ratio present in the study is < 1"},{"metadata":{"trusted":true},"cell_type":"code","source":"# Draw a pie chart about rv_lv_ratio_lt_1.\nplt.pie(train[\"rv_lv_ratio_lt_1\"].value_counts(),labels=[\"0\",\"1\"],autopct=\"%.1f%%\")\nplt.title(\"Ratio of rv_lv_ratio_lt_1\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* leftsided_pe - exam-level, indicates that there is PE present on the left side of the images in the study"},{"metadata":{"trusted":true},"cell_type":"code","source":"# Draw a pie chart about Ratio of leftsided_pe.\nplt.pie(train[\"leftsided_pe\"].value_counts(),labels=[\"0\",\"1\"],autopct=\"%.1f%%\")\nplt.title(\"Ratio of leftsided_pe\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* true_filling_defect_not_pe - informational, indicates a defect that is NOT PE"},{"metadata":{"trusted":true},"cell_type":"code","source":"# Draw a pie chart about Ratio of true_filling_defect_not_pe.\nplt.pie(train[\"true_filling_defect_not_pe\"].value_counts(),labels=[\"0\",\"1\"],autopct=\"%.1f%%\")\nplt.title(\"Ratio of true_filling_defect_not_pe\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* rightsided_pe - exam-level, indicates that there is PE present on the right side of the images in the study"},{"metadata":{"trusted":true},"cell_type":"code","source":"# Draw a pie chart about Ratio of rightsided_pe.\nplt.pie(train[\"rightsided_pe\"].value_counts(),labels=[\"0\",\"1\"],autopct=\"%.1f%%\")\nplt.title(\"Ratio of rightsided_pe\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* central_pe - exam-level, indicates that there is PE present in the center of the images in the study"},{"metadata":{"trusted":true},"cell_type":"code","source":"# Draw a pie chart about Ratio of central_pe.\nplt.pie(train[\"central_pe\"].value_counts(),labels=[\"0\",\"1\"],autopct=\"%.1f%%\")\nplt.title(\"Ratio of central_pe\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* indeterminate -exam-level, indicates that while the study is not negative for PE, an ultimate set of exam-level labels could not be created, due to QA issues"},{"metadata":{"trusted":true},"cell_type":"code","source":"# Draw a pie chart about Ratio of indeterminate.\nplt.pie(train[\"indeterminate\"].value_counts(),labels=[\"0\",\"1\"],autopct=\"%.1f%%\")\nplt.title(\"Ratio of indeterminate\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Find the unique number of pixelspacing_area. \nn = scan['pixelspacing_area'].nunique()\n# First, I'll use Sturgess's formula to find the appropriate number of classes in the histogram \nk = 1 + math.log2(n)\n# Display a histogram of the pixelspacing_area of the training data\nsns.distplot(scan['pixelspacing_area'], kde=True, rug=False, bins=int(k)) \n# Graph Title\nplt.title('pixelspacing_area')\n# Show Histogram\nplt.show() ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Find the unique number of pixelspacing_c. \nn = scan['pixelspacing_c'].nunique()\n# First, I'll use Sturgess's formula to find the appropriate number of classes in the histogram \nk = 1 + math.log2(n)\n# Display a histogram of the pixelspacing_c of the training data\nsns.distplot(scan['pixelspacing_c'], kde=True, rug=False, bins=int(k)) \n# Graph Title\nplt.title('pixelspacing_c')\n# Show Histogram\nplt.show() ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Find the unique number of pixelspacing_r. \nn = scan['pixelspacing_r'].nunique()\n# First, I'll use Sturgess's formula to find the appropriate number of classes in the histogram \nk = 1 + math.log2(n)\n# Display a histogram of the pixelspacing_c of the training data\nsns.distplot(scan['pixelspacing_r'], kde=True, rug=False, bins=int(k)) \n# Graph Title\nplt.title('pixelspacing_r')\n# Show Histogram\nplt.show() ","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Overview of Correlation"},{"metadata":{"trusted":true},"cell_type":"code","source":"# View the correlation heat map\ncorr_mat = train.corr(method='pearson')\nsns.heatmap(corr_mat,\n            vmin=-1.0,\n            vmax=1.0,\n            center=0,\n            annot=True, # True:Displays values in a grid\n            fmt='.1f',\n            xticklabels=corr_mat.columns.values,\n            yticklabels=corr_mat.columns.values\n           )\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* The correlation between qa_contrast and indeterminate appears to be high."},{"metadata":{},"cell_type":"markdown","source":"# Dicom Preprocessing"},{"metadata":{},"cell_type":"markdown","source":"* For dicom preprocessing, [Full Preprocessing Tutorial](https://www.kaggle.com/gzuidhof/full-preprocessing-tutorial) is usefull."},{"metadata":{"trusted":true},"cell_type":"code","source":"def extract_num(s, p, ret=0):\n    search = p.search(s)\n    if search:\n        return int(search.groups()[0])\n    else:\n        return ret","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"import pydicom\n\ndef plot_pixel_array(dataset, figsize=(5,5)):\n    plt.figure(figsize=figsize)\n    plt.imshow(dataset.pixel_array, cmap=plt.cm.bone)\n    plt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"file_path = \"../input/rsna-str-pulmonary-embolism-detection/train/0003b3d648eb/d2b2960c2bbf/00ac73cfc372.dcm\"\ndataset = pydicom.dcmread(file_path)\nplot_pixel_array(dataset)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Acknowledgements\n\n* [Pulmonary Embolism Dicom preprocessing & EDA](https://www.kaggle.com/nitindatta/pulmonary-embolism-dicom-preprocessing-eda)\n* [Pulmonary Dicom Preprocessing](https://www.kaggle.com/allunia/pulmonary-dicom-preprocessing)"},{"metadata":{},"cell_type":"markdown","source":"**Your upvote is the source of my motivation.**"}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}