{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"\n\n<center><h2><b>  </b><a href=\"https://www.kaggle.com/competitions/mayo-clinic-strip-ai\"> Mayo Clinic - Strip AI</a><b> </b></h2></center>\n\n<br>\n\n<!--\nTRIED TO EMBED A YOUTUBE VIDEO UNSUCCESSFULLY\n\n <iframe width=\"853\" height=\"480\" src=\"https://www.youtube.com/embed/tMTsN8UtQIk\" title=\"Animation of Thrombectomy Using Aspiration\" frameborder=\"0\" allow=\"accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture\" allowfullscreen></iframe> \n-->\n\n<img src=\"https://storage.googleapis.com/kaggle-competitions/kaggle/37333/logos/header.png\">\n\n---\n\n<h3> Table Of Contents </h3>\n\n[...]\n\n---\n\nType of problem: **classification**\n\nType of data: **2-D images**\n\nMetadata:\n- **image_id** - A unique identifier for this instance having the form {patient_id}_{image_num}. Corresponds to the image {image_id}.tif.\n- **center_id** - Identifies the medical center where the slide was obtained.\n- **patient_id** - Identifies the patient from whom the slide was obtained.\n- **image_num** - Enumerates images of clots obtained from the same patient.\n- **label** - The etiology of the clot, either CE or LAA. This field is the classification target.\n\n<h3> Domain Description </h3>\n\nThis task will be mainly driven thanks to the emergence of **digital pathology** — an image-based environment for the acquisition, management and interpretation of pathology information supported by computational techniques for data extraction and analysis.\n\n<br><br>\n\n<center><img src=\"https://biolab.polito.it/app/uploads/2021/05/Digital-pathology-scaled.jpg\" width=\"800\" \n     height=\"960\"></center>\n\nThe digital pathology is now fully-integrated in new pathology workflow, based on Machine Learning and new ways to share pre-processed data (mainly made by high-resolution images) and results. Efforts in metadata standardization have also been made.\nAn intense use of ad-hoc databases is also included, as shown below:\n\n<img src=\"https://www.altexsoft.com/media/2022/04/word-image-26.png\">\n\n\n<h3> Competition Description </h3>\n\nThe goal of this competition is to classify the blood clot origins in ischemic stroke. Using whole slide digital pathology images, you'll build a model that differentiates between the two major acute ischemic stroke (AIS) etiology subtypes: cardiac and large artery atherosclerosis.\n\nThe two target variables are:\n\n- Cardioembolic (CE) clots\n- Large Artery Atherosclerosis (LAA) clots.\n\nThe **Acute ischemic stroke** (AIS) occurs when blood flow through a brain artery is blocked by a clot, a mass of thickened blood. Clots are either _thrombotic_ or _embolic_.\n\nAIS is responsible for almost 90% of all strokes.\n[...]\n\n---\n<h3> Data Description </h3>\n\nBefore performing any kind of pre-processing, data wrangling and model selection, it's pivotal to understand the data format, what they represent, how these images are taken from the patients and what could be the best strategy to analyze them.\n\n\nFirst, we know that the patients have enrolled in the Stroke Thromboembolism Registry of Imaging and Pathology (STRIP) and they underwent mechanical thrombectomy for retrieving clots, which then were sent to a central core lab for processing.\n\nThe **whole slide imaging** technique is performed, which permits to capture very high resolution images in a matter of minutes. After the mechanical thrombectomy, the tissues are put into small and very compact glass slides and these are put into special scanners, which then creates very high resolution digital images. <br>\nIn details, the **Martius Scarlet Blue** (MSB) staining technique is used. It is especially suitable for fibrin visualization and older clusters. This method is the most ideal for studying connective tissue and vascular patology (hence, vascular clots).\n\nIn most cases (we will give punctual statistics later during _Exploratory Data Analysis_), digital images are made of white pixels. These white spots can be removed to decrease the image size and resolution; in this case, to put it in more technical terms, we are performing a **region of interest** (ROI) analysis, by using trainable exclusion mappings.\n\n\n<h3> Computing Resources </h3>\n\nThe sheer amount of data do not permit to use classical computing tools and techniques. Hosting the files locally is too burdensome for us, hence _on-the-fly_ workflows are exploited.\n\nThe EDA is carried out from the Kaggle API.\n\n[...]\n\n---\n\n\n<h3> Relevant Literature </h3>\n\n- <a href=\"https://www.nature.com/articles/s41581-020-0321-6\"><i> Digital pathology and computational image analysis in nephropathology</i></a>, L. Barisoni et Al. (2020) \n- <a href=\"https://pubmed.ncbi.nlm.nih.gov/33722963/\"><i>Association between clot composition and stroke origin in mechanical thrombectomy patients: analysis of the Stroke Thromboembolism Registry of Imaging and Pathology</i></a>, W. Brinjikji et Al. (2021)\n- <a href=\"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8711251/\"><i>Digital Pathology and Artificial Intelligence</i></a>, M.K.K. Niazi et Al. (2019)\n- <a href=\"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3405921/\"><i>Managing and Querying Whole Slide Images</i></a>, F. Wang et Al. (2012)\n- <a href=\"https://pubmed.ncbi.nlm.nih.gov/30952688/\"><i>\" Platelet-rich clots as identified by Martius Scarlet Blue staining are isodense on NCCT \"</i></a> S. Fitzgerald et Al. (2019)\n- <a href=\"https://www.dovepress.com/whole-slide-imaging-in-pathology-advantages-limitations-and-emerging-p-peer-reviewed-fulltext-article-PLMI\"><i> Whole slide imaging in pathology: advantages, limitations, and emerging perspectives</i></a> N. Farahani et Al. (2015)\n\n---\n\n<h3>Exploratory Data Analysis: Metadata</h3>\n\nThe first part of EDA is dedicated to the _metadata_ of our images. The actual data on which the model is going to be trained are the images.","metadata":{}},{"cell_type":"code","source":"from openslide import OpenSlide    # Whole Slide Imagines Analysis tool\nfrom glob import glob\nfrom pprint import pprint\nfrom collections import defaultdict\nimport gc\nimport os\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sn\nimport cv2\nimport tifffile\nfrom PIL import Image\nfrom tqdm.auto import tqdm\nimport plotly.express as px\nimport colorama as cl\n\ntrain_df = pd.read_csv('../input/mayo-clinic-strip-ai/train.csv')\ntest_df = pd.read_csv('../input/mayo-clinic-strip-ai/test.csv')\nother_df = pd.read_csv('../input/mayo-clinic-strip-ai/other.csv')\n\ntrain_images = glob(\"/kaggle/input/mayo-clinic-strip-ai/train/*\")\ntest_images = glob(\"/kaggle/input/mayo-clinic-strip-ai/test/*\")\nother_images = glob(\"/kaggle/input/mayo-clinic-strip-ai/other/*\")\n","metadata":{"execution":{"iopub.status.busy":"2022-10-24T09:20:13.557751Z","iopub.execute_input":"2022-10-24T09:20:13.55827Z","iopub.status.idle":"2022-10-24T09:20:16.869272Z","shell.execute_reply.started":"2022-10-24T09:20:13.558175Z","shell.execute_reply":"2022-10-24T09:20:16.868138Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In this first part of EDA, we add as many metadata and information as possible, such as _size_, _width_ and _length_.","metadata":{}},{"cell_type":"code","source":"img_prop = defaultdict(list)\n\nfor i, path in enumerate(train_images):\n    img_path = train_images[i]\n    slide = OpenSlide(img_path)    \n    img_prop['image_id'].append(img_path[-12:-4])\n    img_prop['width (px)'].append(slide.dimensions[0])\n    img_prop['height (px)'].append(slide.dimensions[1])\n    img_prop['size'].append(str(round(os.path.getsize(img_path) / 1e6, 2)))\n    img_prop['path'].append(img_path)\n\nimage_data = pd.DataFrame(img_prop)\nimage_data['img_aspect_ratio'] = image_data['width (px)']/image_data['height (px)']\nimage_data.sort_values(by='image_id', inplace=True)\nimage_data.reset_index(inplace=True, drop=True)\n\nimage_data = image_data.merge(train_df, on='image_id')\nimage_data.head()","metadata":{"execution":{"iopub.status.busy":"2022-10-24T09:20:46.302677Z","iopub.execute_input":"2022-10-24T09:20:46.303116Z","iopub.status.idle":"2022-10-24T09:21:06.841208Z","shell.execute_reply.started":"2022-10-24T09:20:46.303074Z","shell.execute_reply":"2022-10-24T09:21:06.839753Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"By using **Python Image Library**, we can conveniently store and visualize large images. In our case, it's especially useful given their size. \n\nFor now we just make sure to visualize a small subset of training images, in order to understand the their nature: ","metadata":{}},{"cell_type":"code","source":"from PIL import Image\nImage.MAX_IMAGE_PIXELS = None  # you have to set this value to allow displaying high-resolution images\n\nCE_imgs = image_data.loc[image_data['label']=='CE','path']\nLAA_imgs = image_data.loc[image_data['label']=='LAA','path']","metadata":{"execution":{"iopub.status.busy":"2022-10-24T09:22:58.265573Z","iopub.execute_input":"2022-10-24T09:22:58.266053Z","iopub.status.idle":"2022-10-24T09:22:58.275703Z","shell.execute_reply.started":"2022-10-24T09:22:58.266018Z","shell.execute_reply":"2022-10-24T09:22:58.27429Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.style.use('default')\nfig, axes = plt.subplots(1,5, figsize=(16,16))\ntrain_images\nfor ax in axes.reshape(-1):\n    img_path = np.random.choice(CE_imgs)\n    img = Image.open(img_path)   \n    img.thumbnail((300,300), Image.Resampling.LANCZOS)\n    ax.imshow(img), ax.set_title(\"target: CE\")\nplt.axis(\"off\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-10-23T15:06:37.693885Z","iopub.execute_input":"2022-10-23T15:06:37.69482Z","iopub.status.idle":"2022-10-23T15:07:22.061022Z","shell.execute_reply.started":"2022-10-23T15:06:37.694785Z","shell.execute_reply":"2022-10-23T15:07:22.060041Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"image_data.head()","metadata":{"execution":{"iopub.status.busy":"2022-10-23T15:07:22.062415Z","iopub.execute_input":"2022-10-23T15:07:22.063066Z","iopub.status.idle":"2022-10-23T15:07:22.080381Z","shell.execute_reply.started":"2022-10-23T15:07:22.063033Z","shell.execute_reply":"2022-10-23T15:07:22.07905Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's see if the training dataset has **NaN**s:","metadata":{}},{"cell_type":"code","source":"train_df.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-10-23T15:07:22.081804Z","iopub.execute_input":"2022-10-23T15:07:22.082162Z","iopub.status.idle":"2022-10-23T15:07:22.098235Z","shell.execute_reply.started":"2022-10-23T15:07:22.08213Z","shell.execute_reply":"2022-10-23T15:07:22.097044Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The dataset have all values, this means no _missing data imputation_ is needed.","metadata":{}},{"cell_type":"code","source":"labels = train_df[\"label\"].value_counts()\nlabels\n\n# Percentage calculations\n\nTotal = labels[0]+labels[1]\nPercCE, PercLAA = round((labels[0]/(Total)),3)*100, round(round((labels[1]/(Total)),3)*100,2)\n\n\nplt.pie(labels, labels=[\"Cardioembolic clots \" + str(PercCE) +\"%\", \n                        \"Large Artery Atherosclerosis \" + str(PercLAA) +\"%\"], \n                        labeldistance=None)\n\nplt.legend(bbox_to_anchor=(1,0), loc=\"lower right\", \n                          bbox_transform=plt.gcf().transFigure)\nplt.title(\"Train %ages\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-10-23T15:07:22.100078Z","iopub.execute_input":"2022-10-23T15:07:22.100418Z","iopub.status.idle":"2022-10-23T15:07:22.251755Z","shell.execute_reply.started":"2022-10-23T15:07:22.100388Z","shell.execute_reply":"2022-10-23T15:07:22.249674Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The train dataset is _unbalanced_. In a classification context, this may lead to errors when only accounting model accuracy.  Better metrics will then used, such as **F1-Score** or **Recall**. \n\nThere are systematic algorithms that it is possible to use in order to generate synthetic samples. The most popular of such algorithms is called the **Synthetic Minority Over-sampling Technique** (SMOTE). We can also give a bigger weight to LAA, which is the under-represented class.\n\n<a href=\"https://www.mdpi.com/2075-4418/10/10/804/htm\">Harpaz et Al. (2020)</a> tell us \"<i> Ischemic stroke accounts for the majority of stroke cases, which makes ischemic stroke mechanism classification the second most important classification in stroke care. There are four main ischemic stroke etiologies: <b>20% cardioembolic (CE), 20% atherosclerosis (large artery disease, LAA)</b>, 25% lacunar (small vessel disease) and 5% other causes</i>\". <br>\nThis is important to know, because now we can impute the label imbalance to other factors. The training set distribution is not coherent with the population distribution.\n\nNow let's see how many patients are involved:","metadata":{}},{"cell_type":"code","source":"print(str(train_df[\"patient_id\"].value_counts()))\n\nn_pat = train_df[\"patient_id\"].unique().size\nprint(f\" \\n\\n - Number of Train images: {train_df.shape[0]}\")\nprint(\"\\n\\n - Number of patients in Train:\", n_pat)\n","metadata":{"execution":{"iopub.status.busy":"2022-10-23T15:07:22.254476Z","iopub.execute_input":"2022-10-23T15:07:22.255619Z","iopub.status.idle":"2022-10-23T15:07:22.27141Z","shell.execute_reply.started":"2022-10-23T15:07:22.25554Z","shell.execute_reply":"2022-10-23T15:07:22.269783Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can notice some patients got multiple strokes (!). <br>\nPatient # _91b9d3_ got five of them, or at least these are the ones that have been subject to whole slide imaging.\n\nA better visualization can be provided:\n","metadata":{}},{"cell_type":"code","source":"lb = list(range(5))\nbplt = sn.countplot(x=\"image_num\", data=train_df, dodge=False, palette=\"icefire\")\nbplt.bar_label(bplt.containers[0])\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-10-23T15:07:22.277152Z","iopub.execute_input":"2022-10-23T15:07:22.27903Z","iopub.status.idle":"2022-10-23T15:07:22.515554Z","shell.execute_reply.started":"2022-10-23T15:07:22.278945Z","shell.execute_reply":"2022-10-23T15:07:22.514622Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From this first plot we can notice that the image number distribution is highly skewed. This means the clots' slides are rarely from the same patient. <br>\nFor example, the <code>image_num = 1</code> means that there are 89 patients where the clots' slides are from.<br>\nThis observation may be useful to understand whether the clot severity and tissue degradation increase with the number of clots for each patient: <a href=\"https://n.neurology.org/content/54/3/674\">domain-knowledge</a> research hints us so.\n\nNow we can get, in details, how many WSI each patient underwent:","metadata":{}},{"cell_type":"code","source":"df = (train_df.groupby([\"patient_id\",\"label\"])[\"image_num\"].count().reset_index(name='image_count').sort_values(\"image_count\", ascending=False))\ndf[df[\"image_count\"]>1].set_index(\"patient_id\").style.background_gradient(cmap='Reds')\n","metadata":{"execution":{"iopub.status.busy":"2022-10-23T15:07:33.203671Z","iopub.execute_input":"2022-10-23T15:07:33.205223Z","iopub.status.idle":"2022-10-23T15:07:33.238971Z","shell.execute_reply.started":"2022-10-23T15:07:33.205177Z","shell.execute_reply":"2022-10-23T15:07:33.237781Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we plot the labels for number of images and their frequency:","metadata":{}},{"cell_type":"code","source":"df = pd.crosstab(index=train_df.image_num, columns=train_df.label).reset_index()\ndf[\"LAA/CE\"] = df.LAA/df.CE\nax = df.plot(x=\"image_num\", y=[\"CE\",\"LAA\"], kind=\"bar\", width=0.8, figsize=(14,7))\nax.bar_label(ax.containers[0])\nax.bar_label(ax.containers[1]);","metadata":{"execution":{"iopub.status.busy":"2022-10-23T15:07:22.553191Z","iopub.execute_input":"2022-10-23T15:07:22.553543Z","iopub.status.idle":"2022-10-23T15:07:22.910587Z","shell.execute_reply.started":"2022-10-23T15:07:22.553511Z","shell.execute_reply":"2022-10-23T15:07:22.909287Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the graph, we can notice that the vast majority of multiple strokes are associable with Cardioembolic strokes.. The ratio is consistent throughout the number of stroke events per patient. <br><br>\nThe next graph shows the distribution in greater details:","metadata":{}},{"cell_type":"code","source":"df = pd.crosstab(index=train_df.center_id, columns=train_df.label).reset_index()\ndf[\"LAA/CE\"] = df.LAA/df.CE\nax = df.plot(x=\"center_id\", y=[\"CE\",\"LAA\"], kind=\"bar\", width=0.8, figsize=(14,7))\nax.bar_label(ax.containers[0])\nax.bar_label(ax.containers[1])\ndf.plot(y=[\"LAA/CE\"], secondary_y=\"LAA/CE\", color=\"lightgreen\", linewidth=3, ax=ax);","metadata":{"execution":{"iopub.status.busy":"2022-10-23T15:07:22.912267Z","iopub.execute_input":"2022-10-23T15:07:22.913463Z","iopub.status.idle":"2022-10-23T15:07:23.493848Z","shell.execute_reply.started":"2022-10-23T15:07:22.913413Z","shell.execute_reply":"2022-10-23T15:07:23.492947Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The LAA/CE in the overall case is consistent with the ratio we have seen in the previous graph.\n\nNow can plot something about the Lab Centers (_center_id_) where the WSI come from.","metadata":{}},{"cell_type":"code","source":"df = df_train[[\"center_id\"]].value_counts().reset_index(name=\"# of images\")\nlb = list(df.center_id.values)\nod = df.sort_values(\"# of images\", ascending=False).center_id\n\nsn.barplot(x=\"center_id\",  y=\"# of images\", data=df, \n                    order=od, palette=\"icefire\")\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-10-23T15:07:23.495172Z","iopub.execute_input":"2022-10-23T15:07:23.4957Z","iopub.status.idle":"2022-10-23T15:07:23.763323Z","shell.execute_reply.started":"2022-10-23T15:07:23.495666Z","shell.execute_reply":"2022-10-23T15:07:23.76242Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can notice the WSI source is unbalanced. <br> Maybe Center_ID 11 is the most important hospital nationwide.   Even if there is a strong data skewness, this should not bring many problems, given that the medical and technical procedures are strongly consistent with national protocols.","metadata":{}},{"cell_type":"markdown","source":"<h4><b>Image Size</b></h4>\n\nAnother useful metadata to include in our inquiry is the image size.  All the images come in **Tag Image File Format** (_.tiff_), which is a computer file used to store raster graphics and image information.\n\nWe can start by plotting the size and aspect ratio distribution of our images:\n","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots(1,2, figsize=(16,5))\nsn.histplot(x='size', data = image_data, bins=100, ax=ax[0])\nax[0].set_title(\"Distribution of size\"), ax[0].set_ylabel(\"# of Images\")\nsn.histplot(x='img_aspect_ratio', data = image_data, bins=100, ax=ax[1])\nax[1].set_title(\"Image aspect ratio\"), ax[1].set_ylabel(\"%\")\nplt.show()\n\n","metadata":{"execution":{"iopub.status.busy":"2022-10-23T15:09:28.468539Z","iopub.execute_input":"2022-10-23T15:09:28.469035Z","iopub.status.idle":"2022-10-23T15:09:36.745704Z","shell.execute_reply.started":"2022-10-23T15:09:28.468968Z","shell.execute_reply":"2022-10-23T15:09:36.744436Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2> Exploratory Data Analysis: Images </h2>","metadata":{}},{"cell_type":"markdown","source":"**<h3>Model Brainstorming</h3>**\n\n\n---\n<h4> Relevant Lecture for Model Selection </h4>\n\n- <a href=\"https://www.frontiersin.org/articles/10.3389/fmed.2019.00264/full\"><i>Deep Learning for Whole Slide Image Analysis: An Overview</i></a> N. Dimitriou et Al. (2019)\n- <a href=\"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8401188/\"><i> Transfer Learning Approach for Classification of Histopathology Whole Slide Images</i></a> S. Ahmed et Al (2021)\n---\n\n\nAt first glance, given that we are dealing with images, the most suitable models could be **Convolutional AutoEncoders** (CAE) and **Convolutional Neural Networks** (CNN). But this may not be necessarily true, or at least these instruments may not be enough in their vanilla version.\n\nWhole Slide Images Analysis (from now WSIA) deal with images with billions of pixels. They suffer from high morphological heterogeneity as well as from different types of artifacts.\n\nThese conditions preclude the direct application of conventional deep learning techniques. Instead, practitioners are faced with **two non-trivial challenges**:\n- the visual understanding of the images, impeded by the morphological variance, artifacts, and typically small data sets;\n- the inability of the current state of the hardware to facilitate learning from images with such high resolution, thereby requiring some form of dimensionality reduction to the images.\n\nThe second point is a very impeding bottleneck: this is the reason why we are trying to adopt an online pipeline for on-the-fly image analysis or dig deeper in the WSI _pyramid approach_, which permits an intelligent resource exploitation by only loading and processing a subset image pixels when a zoom-in occurs in a specific region. The sub-region comes in a standard sizes and (e.g.: 256x256 pixels) are called _patches_.  <br>\nThe <code><a href=\"https://openslide.org/\">OpenSlide</a></code> Python Library permits to use these strategies efficiently.\n\n<center><img src=\"https://gut.bmj.com/content/gutjnl/70/6/1183/F1.large.jpg\" alt=\"pyramid\" width=\"800\" \n     height=\"500\"> </center>\n\n<br>\n\nEach single patch can carry metadata to help the researcher to understand some details, for example the nature and the type of cells in that patch.\n\n\n<h2><p> As suggested by Daniel, we can exploit <i>Transfer Learning</i>, we can write plenty about it</p></h2>\n<!-- The first one permits to get the latent variables and decrease the data dimensionality, which in our case is especially important to carry out. A single _.tif_ image may weigh up to 300 (!) MB. \n\nGiven that, analyzing image 008e5c_0 (100 MB), it's possible to notice that the image is mainly made by white pixels, which obviously do not carry meaningful information, apart from understanding the size of the clots.\n -->\nROI is carried out and all the images have to be centered accordingly.","metadata":{}},{"cell_type":"markdown","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}}]}