{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Goal of the Competition\nThe goal of this competition is to classify the blood clot origins in ischemic stroke. Using whole slide digital pathology images, you'll build a model that differentiates between the two major acute ischemic stroke (AIS) etiology subtypes: cardiac and large artery atherosclerosis.\n\nYour work will enable healthcare providers to better identify the origins of blood clots in deadly strokes, making it easier for physicians to prescribe the best post-stroke therapeutic management and reducing the likelihood of a second stroke.\n\n# Context\nStroke remains the second-leading cause of death worldwide. Each year in the United States, over 700,000 individuals experience an ischemic stroke caused by a blood clot blocking an artery to the brain. A second stroke (23% of total events are recurrent) worsens the chances of the patient’s survival. However, subsequent strokes may be mitigated if physicians can determine stroke etiology, which influences the therapeutic management following stroke events.\n\nDuring the last decade, mechanical thrombectomy has become the standard of care treatment for acute ischemic stroke from large vessel occlusion. As a result, retrieved clots became amenable to analysis. Healthcare professionals are currently attempting to apply deep learning-based methods to predict ischemic stroke etiology and clot origin. However, unique data formats, image file sizes, as well as the number of available pathology slides create challenges you could lend a hand in solving.\n\nThe Mayo Clinic is a nonprofit American academic medical center focused on integrated health care, education, and research. Stroke Thromboembolism Registry of Imaging and Pathology (STRIP) is a uniquely large multicenter project led by Mayo Clinic Neurovascular Lab with the aim of histopathologic characterization of thromboemboli of various etiologies and examining clot composition and its relation to mechanical thrombectomy revascularization.\n\nTo decrease the chances of subsequent strokes, the Mayo Clinic Neurovascular Research Laboratory encourages data scientists to improve artificial intelligence-based etiology classification so that physicians are better equipped to prescribe the correct treatment. New computational and artificial intelligence approaches could help save the lives of stroke survivors and help us better understand the world's second-leading cause of death.","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"# Data Description\nThe dataset for this competition comprises over a thousand high-resolution whole-slide digital pathology images. Each slide depicts a blood clot from a patient that had experienced an acute ischemic stroke.\n\nThe slides comprising the training and test sets depict clots with an etiology (that is, origin) known to be either CE (Cardioembolic) or LAA (Large Artery Atherosclerosis). We include a set of supplemental slides with a either an unknown etiology or an etiology other than CE or LAA.\n\nYour task is to classify the etiology (CE or LAA) of the slides in the test set for each patient.\n\n# File and Data Field Descriptions\ntrain/ - A folder containing images in the TIFF format to be used as training data.\n\ntest/ - A folder containing images to be used as test data. The actual test data comprises about 280 images.\n\nother/ - A supplemental set of images with a either an unknown etiology or an etiology other than CE or LAA.\n\ntrain.csv Contains annotations for images in the train/ folder.\n\nimage_id - A unique identifier for this instance having the form {patient_id}_{image_num}. Corresponds to the image {image_id}.tif.\n\ncenter_id - Identifies the medical center where the slide was obtained.\n\npatient_id - Identifies the patient from whom the slide was obtained.\n\nimage_num - Enumerates images of clots obtained from the same patient.\n\nlabel - The etiology of the clot, either CE or LAA. This field is the classification target.\n\ntest.csv - Annotations for images in the test/ folder. Has the same fields as train.csv excluding label.\n\nother.csv - Annotations for images in the other/ folder. Has the same fields as train.csv. The center_id is unavailable for these images however.\n\nlabel - The etiology of the clot, either Unknown or Other.\nother_specified - The specific etiology, when known, in case the etiology is labeled as Other.\n\nsample_submission.csv - A sample submission file in the correct format. See the Evaluation page for more details. Note in particular that you should make one prediction per patient_id, not per image_id.","metadata":{}},{"cell_type":"markdown","source":"# Import Libraries","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom openslide import OpenSlide\nimport tifffile\norange_black = ['#fdc029', '#df861d', '#FF6347', '#aa3d01', '#a30e15', '#800000', '#171820']\n\n# Setting plot styling.\nplt.style.use('ggplot')\n\nplt.rcParams['figure.figsize'] = (18, 14)\nplt.rcParams['figure.dpi'] = 300\nplt.rcParams[\"axes.grid\"] = True\nplt.rcParams[\"grid.color\"] = orange_black[0]\nplt.rcParams[\"grid.alpha\"] = 0.5\nplt.rcParams[\"grid.linestyle\"] = '--'\nplt.rcParams[\"font.family\"] = \"monospace\"\n\nplt.rcParams['axes.edgecolor'] = 'black'\nplt.rcParams['figure.frameon'] = False\nplt.rcParams['axes.spines.left'] = True\nplt.rcParams['axes.spines.bottom'] = True\nplt.rcParams['axes.spines.top'] = False\nplt.rcParams['axes.spines.right'] = False\nplt.rcParams['axes.linewidth'] = 1.0\n\nimport warnings\nwarnings.filterwarnings(\"ignore\")","metadata":{"execution":{"iopub.status.busy":"2022-07-07T13:09:59.373912Z","iopub.execute_input":"2022-07-07T13:09:59.375106Z","iopub.status.idle":"2022-07-07T13:10:00.797015Z","shell.execute_reply.started":"2022-07-07T13:09:59.374964Z","shell.execute_reply":"2022-07-07T13:10:00.795457Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 2. Load Data","metadata":{}},{"cell_type":"code","source":"train = pd.read_csv(\"../input/mayo-clinic-strip-ai/train.csv\")\nprint(\"There are {} rows and {} colums in the training data\".format(train.shape[0],train.shape[1]))","metadata":{"execution":{"iopub.status.busy":"2022-07-07T13:10:00.799445Z","iopub.execute_input":"2022-07-07T13:10:00.799815Z","iopub.status.idle":"2022-07-07T13:10:00.824906Z","shell.execute_reply.started":"2022-07-07T13:10:00.799782Z","shell.execute_reply":"2022-07-07T13:10:00.823765Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 3. Take a look at the data","metadata":{}},{"cell_type":"code","source":"train.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-07T13:10:00.827087Z","iopub.execute_input":"2022-07-07T13:10:00.827995Z","iopub.status.idle":"2022-07-07T13:10:00.855387Z","shell.execute_reply.started":"2022-07-07T13:10:00.827952Z","shell.execute_reply":"2022-07-07T13:10:00.853896Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-07T13:10:00.859238Z","iopub.execute_input":"2022-07-07T13:10:00.861147Z","iopub.status.idle":"2022-07-07T13:10:00.893381Z","shell.execute_reply.started":"2022-07-07T13:10:00.861084Z","shell.execute_reply":"2022-07-07T13:10:00.891555Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# check for null values\ntrain.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-07T13:10:00.895764Z","iopub.execute_input":"2022-07-07T13:10:00.896645Z","iopub.status.idle":"2022-07-07T13:10:00.910164Z","shell.execute_reply.started":"2022-07-07T13:10:00.896592Z","shell.execute_reply":"2022-07-07T13:10:00.908756Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- There are no null values in the training data","metadata":{}},{"cell_type":"code","source":"# check for unique values for each column\nfor col in train.columns:\n    print(\"unique values for column {} : {}\".format(col,train[col].nunique()))","metadata":{"execution":{"iopub.status.busy":"2022-07-07T13:10:00.912838Z","iopub.execute_input":"2022-07-07T13:10:00.913751Z","iopub.status.idle":"2022-07-07T13:10:00.926681Z","shell.execute_reply.started":"2022-07-07T13:10:00.91368Z","shell.execute_reply":"2022-07-07T13:10:00.924711Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# check for no. of unique values for image id\nprint(\"We have {} unique image ids\".format(train.image_id.nunique()))","metadata":{"execution":{"iopub.status.busy":"2022-07-07T13:10:00.928917Z","iopub.execute_input":"2022-07-07T13:10:00.929787Z","iopub.status.idle":"2022-07-07T13:10:00.943776Z","shell.execute_reply.started":"2022-07-07T13:10:00.929735Z","shell.execute_reply":"2022-07-07T13:10:00.942228Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- We have no duplicates in the dataset ","metadata":{}},{"cell_type":"code","source":"# check for no. of unique patient ids\nprint(\"We have {} unique pateint ids\".format(train.patient_id.nunique()))","metadata":{"execution":{"iopub.status.busy":"2022-07-07T13:10:00.94558Z","iopub.execute_input":"2022-07-07T13:10:00.947278Z","iopub.status.idle":"2022-07-07T13:10:00.961396Z","shell.execute_reply.started":"2022-07-07T13:10:00.947224Z","shell.execute_reply":"2022-07-07T13:10:00.960188Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- We have more images than pateit ids, means we have few patients for which more than one image is available","metadata":{}},{"cell_type":"markdown","source":"# 4. EDA","metadata":{}},{"cell_type":"code","source":"print('Sample Image')\nimg = '../input/mayo-clinic-strip-ai/other/01f2b3_0.tif'\nregion = (0, 0)\nlevel = 0\nsize = (5000, 5000)\n\nplt.figure(figsize=(20,8))\nslide = OpenSlide(img)\nplt.imshow(slide.read_region(region, level, size))\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-07T13:10:00.962796Z","iopub.execute_input":"2022-07-07T13:10:00.963784Z","iopub.status.idle":"2022-07-07T13:10:07.020983Z","shell.execute_reply.started":"2022-07-07T13:10:00.963739Z","shell.execute_reply":"2022-07-07T13:10:07.019418Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Sample Images with label == CE')\nimgs = train.loc[train.label==\"CE\"].sample(1).image_id.values # more than 1 images can also be picked here\n\nregion = (0, 0)\nlevel = 0\nsize = (5000, 5000)\n\nplt.figure(figsize=(10,8))\nfor k in imgs:\n    slide = OpenSlide('../input/mayo-clinic-strip-ai/train/%s.tif'%k)\n    plt.axis('off')\n    plt.imshow(slide.read_region(region, level, size))\nplt.show()\n\nprint('Sample Images with label == LAA')\nimgs = train.loc[train.label==\"LAA\"].sample(1).image_id.values # more than 1 images can also be picked here\n\nplt.figure(figsize=(10,8))\nfor k in imgs:\n    slide = OpenSlide('../input/mayo-clinic-strip-ai/train/%s.tif'%k)\n    plt.axis('off')\n    plt.imshow(slide.read_region(region, level, size))\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-07T13:16:56.735118Z","iopub.execute_input":"2022-07-07T13:16:56.736214Z","iopub.status.idle":"2022-07-07T13:17:14.282521Z","shell.execute_reply.started":"2022-07-07T13:16:56.736174Z","shell.execute_reply":"2022-07-07T13:17:14.280268Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# check how many records are available for each patient\nsns.countplot(train.patient_id.value_counts(),palette=orange_black)\nplt.title('Patient Id Distribution')\nplt.ylabel('Count')\nplt.xlabel('Patient Id')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-07T13:10:12.510778Z","iopub.status.idle":"2022-07-07T13:10:12.51173Z","shell.execute_reply.started":"2022-07-07T13:10:12.511409Z","shell.execute_reply":"2022-07-07T13:10:12.511438Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- Most of the patients have only 1 image, maximum images for a patient is 5","metadata":{}},{"cell_type":"code","source":"# check for class distribution\nsns.countplot(train.label,palette=orange_black)\nplt.title('Class Distribution')\nplt.ylabel('Count')\nplt.xlabel('Label')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-07T13:10:12.513196Z","iopub.status.idle":"2022-07-07T13:10:12.514272Z","shell.execute_reply.started":"2022-07-07T13:10:12.514048Z","shell.execute_reply":"2022-07-07T13:10:12.514071Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- We have more observations for label \"CE\" than \"LAA\"","metadata":{}},{"cell_type":"code","source":"# center id distribution\nsns.countplot(train.center_id,hue=train.label,palette=orange_black)\nplt.title('Center Id Distribution')\nplt.ylabel('Count')\nplt.xlabel('Center Id')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-07T13:10:12.515893Z","iopub.status.idle":"2022-07-07T13:10:12.516289Z","shell.execute_reply.started":"2022-07-07T13:10:12.516101Z","shell.execute_reply":"2022-07-07T13:10:12.516119Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- Most of the images have center id as 11\n- center id count varies based on label, but that seems to because we have less data for the label = \"LAA\"","metadata":{}},{"cell_type":"code","source":"# no. of images distribution for patients\nsns.countplot(train.image_num,hue=train.label,palette=orange_black)\nplt.title('Image Number Distribution')\nplt.ylabel('Count')\nplt.xlabel('Image Number')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-07T13:10:12.517813Z","iopub.status.idle":"2022-07-07T13:10:12.518201Z","shell.execute_reply.started":"2022-07-07T13:10:12.518015Z","shell.execute_reply":"2022-07-07T13:10:12.518032Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- Most of the patients have only 1 image, max no. of images for a patient is 5\n- No. of images vary for both the labels, but thats seems to be because we have less data for the label \"LAA\"","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}