{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<h1><center><b>Mayo Clinic - STRIP AI - Exploratory Data Analysis</b></center></h1>","metadata":{"papermill":{"duration":0.042833,"end_time":"2020-11-30T22:55:26.624955","exception":false,"start_time":"2020-11-30T22:55:26.582122","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"![](https://storage.googleapis.com/kaggle-competitions/kaggle/37333/logos/header.png)","metadata":{"papermill":{"duration":0.041013,"end_time":"2020-11-30T22:55:26.708321","exception":false,"start_time":"2020-11-30T22:55:26.667308","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"<a id=\"top\"></a>\n<div class=\"list-group\" id=\"list-tab\" role=\"tablist\"><h2 class=\"list-group-item list-group-item-action active\" data-toggle=\"list\" style='color:white; background:red; border:0' role=\"tab\" aria-controls=\"home\"><center>Quick Navigation</center></h2>\n\n* [0. Goal and Context of Competition](#0)\n* [1. Basic Data Exploration](#1)\n* [2. Images Visualizations](#2)\n* [3. Baseline Submission](#3)","metadata":{"papermill":{"duration":0.04059,"end_time":"2020-11-30T22:55:26.790107","exception":false,"start_time":"2020-11-30T22:55:26.749517","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"<a id=\"0\"></a>\n<h2 style='background:red; border:0; color:white'><center>Goal of the Competition and Context</center></h2>","metadata":{"execution":{"iopub.status.busy":"2022-07-07T04:37:50.276569Z","iopub.execute_input":"2022-07-07T04:37:50.276946Z","iopub.status.idle":"2022-07-07T04:37:50.283917Z","shell.execute_reply.started":"2022-07-07T04:37:50.276915Z","shell.execute_reply":"2022-07-07T04:37:50.282509Z"}}},{"cell_type":"markdown","source":"\nThe goal of this competition is to classify the blood clot origins in ischemic stroke. Using whole slide digital pathology images, you'll build a model that differentiates between the two major acute ischemic stroke (AIS) etiology subtypes: cardiac and large artery atherosclerosis.\n\nYour work will enable healthcare providers to better identify the origins of blood clots in deadly strokes, making it easier for physicians to prescribe the best post-stroke therapeutic management and reducing the likelihood of a second stroke.\n\n<b>Context</b>\n \nStroke remains the second-leading cause of death worldwide. Each year in the United States, over 700,000 individuals experience an ischemic stroke caused by a blood clot blocking an artery to the brain. A second stroke (23% of total events are recurrent) worsens the chances of the patient’s survival. However, subsequent strokes may be mitigated if physicians can determine stroke etiology, which influences the therapeutic management following stroke events.\n\nDuring the last decade, mechanical thrombectomy has become the standard of care treatment for acute ischemic stroke from large vessel occlusion. As a result, retrieved clots became amenable to analysis. Healthcare professionals are currently attempting to apply deep learning-based methods to predict ischemic stroke etiology and clot origin. However, unique data formats, image file sizes, as well as the number of available pathology slides create challenges you could lend a hand in solving.\n\nThe Mayo Clinic is a nonprofit American academic medical center focused on integrated health care, education, and research. Stroke Thromboembolism Registry of Imaging and Pathology (STRIP) is a uniquely large multicenter project led by Mayo Clinic Neurovascular Lab with the aim of histopathologic characterization of thromboemboli of various etiologies and examining clot composition and its relation to mechanical thrombectomy revascularization.\n\nTo decrease the chances of subsequent strokes, the Mayo Clinic Neurovascular Research Laboratory encourages data scientists to improve artificial intelligence-based etiology classification so that physicians are better equipped to prescribe the correct treatment. New computational and artificial intelligence approaches could help save the lives of stroke survivors and help us better understand the world's second-leading cause of death.\n\nThis competition is about predicting the origins of blood clot resulting in ischemic stroke using Digital Pathology Slides (aka Whole Slide Images (WSI)) of the thrombotic material extracted mechanically during acute neurovascular procedures. Collaborative efforts across 18 institutions, led by Waleed Brinjikji, MD of the Mayo Clinic Rochester allowed the compilation of this unique dataset (https://pubmed.ncbi.nlm.nih.gov/33722963/).","metadata":{}},{"cell_type":"markdown","source":"<a id=\"1\"></a>\n<h2 style='background:red; border:0; color:white'><center>Basic Data Exploration<center></h2>","metadata":{"papermill":{"duration":0.041678,"end_time":"2020-11-30T22:55:52.319607","exception":false,"start_time":"2020-11-30T22:55:52.277929","status":"completed"},"tags":[]}},{"cell_type":"code","source":"import os\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sn\nimport cv2\nimport tifffile\nfrom PIL import Image\nfrom tqdm.auto import tqdm\nimport plotly.express as px","metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","papermill":{"duration":1.165167,"end_time":"2020-11-30T22:55:53.526664","exception":false,"start_time":"2020-11-30T22:55:52.361497","status":"completed"},"tags":[],"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-23T19:56:58.192984Z","iopub.execute_input":"2022-07-23T19:56:58.193859Z","iopub.status.idle":"2022-07-23T19:56:58.199478Z","shell.execute_reply.started":"2022-07-23T19:56:58.19382Z","shell.execute_reply":"2022-07-23T19:56:58.198774Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"BASE_PATH = \"../input/mayo-clinic-strip-ai/\"\nImage.MAX_IMAGE_PIXELS = 5_000_000_000","metadata":{"papermill":{"duration":0.052328,"end_time":"2020-11-30T22:55:53.62131","exception":false,"start_time":"2020-11-30T22:55:53.568982","status":"completed"},"tags":[],"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-23T19:56:58.200465Z","iopub.execute_input":"2022-07-23T19:56:58.201202Z","iopub.status.idle":"2022-07-23T19:56:58.221542Z","shell.execute_reply.started":"2022-07-23T19:56:58.20117Z","shell.execute_reply":"2022-07-23T19:56:58.220475Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Train dataset","metadata":{"papermill":{"duration":0.041995,"end_time":"2020-11-30T22:55:53.707282","exception":false,"start_time":"2020-11-30T22:55:53.665287","status":"completed"},"tags":[]}},{"cell_type":"code","source":"df_train = pd.read_csv(\n    os.path.join(BASE_PATH, \"train.csv\")\n)\ndf_train","metadata":{"papermill":{"duration":0.233896,"end_time":"2020-11-30T22:55:54.067575","exception":false,"start_time":"2020-11-30T22:55:53.833679","status":"completed"},"tags":[],"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-23T19:57:21.6839Z","iopub.execute_input":"2022-07-23T19:57:21.684238Z","iopub.status.idle":"2022-07-23T19:57:21.70455Z","shell.execute_reply.started":"2022-07-23T19:57:21.684209Z","shell.execute_reply":"2022-07-23T19:57:21.703481Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.info()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-23T19:57:23.348413Z","iopub.execute_input":"2022-07-23T19:57:23.349357Z","iopub.status.idle":"2022-07-23T19:57:23.374678Z","shell.execute_reply.started":"2022-07-23T19:57:23.349296Z","shell.execute_reply":"2022-07-23T19:57:23.373685Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Test dataset","metadata":{}},{"cell_type":"code","source":"df_test = pd.read_csv(\n    os.path.join(BASE_PATH, \"test.csv\")\n)\ndf_test","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-23T19:57:26.391267Z","iopub.execute_input":"2022-07-23T19:57:26.391635Z","iopub.status.idle":"2022-07-23T19:57:26.40886Z","shell.execute_reply.started":"2022-07-23T19:57:26.391602Z","shell.execute_reply":"2022-07-23T19:57:26.407303Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Other dataset","metadata":{}},{"cell_type":"code","source":"df_other = pd.read_csv(\n    os.path.join(BASE_PATH, \"other.csv\")\n)\ndf_other","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-23T19:57:38.186161Z","iopub.execute_input":"2022-07-23T19:57:38.186616Z","iopub.status.idle":"2022-07-23T19:57:38.208731Z","shell.execute_reply.started":"2022-07-23T19:57:38.186576Z","shell.execute_reply":"2022-07-23T19:57:38.207561Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Sample Submission","metadata":{"papermill":{"duration":0.043988,"end_time":"2020-11-30T22:55:54.154813","exception":false,"start_time":"2020-11-30T22:55:54.110825","status":"completed"},"tags":[]}},{"cell_type":"code","source":"df_sub = pd.read_csv(\n    os.path.join(BASE_PATH, \"sample_submission.csv\"))\n\nif len(df_sub)>4:\n    eda=False\nelse:\n    eda=True\n\ndf_sub","metadata":{"papermill":{"duration":0.061839,"end_time":"2020-11-30T22:55:54.261549","exception":false,"start_time":"2020-11-30T22:55:54.19971","status":"completed"},"tags":[],"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-23T19:58:10.186288Z","iopub.execute_input":"2022-07-23T19:58:10.186667Z","iopub.status.idle":"2022-07-23T19:58:10.20606Z","shell.execute_reply.started":"2022-07-23T19:58:10.186637Z","shell.execute_reply":"2022-07-23T19:58:10.204833Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Number of samples","metadata":{"papermill":{"duration":0.043697,"end_time":"2020-11-30T22:55:54.349366","exception":false,"start_time":"2020-11-30T22:55:54.305669","status":"completed"},"tags":[]}},{"cell_type":"code","source":"n_pat = df_train[\"patient_id\"].unique().size\nprint(f\"Number of Train images: {df_train.shape[0]}\")\nprint(\"Number of patients in Train:\", n_pat)\nprint(f\"Number of Test images: {df_test.shape[0]}\")\nprint(\"Numver of patients in Test:\", df_test[\"patient_id\"].unique().size)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-23T19:58:43.544934Z","iopub.execute_input":"2022-07-23T19:58:43.545337Z","iopub.status.idle":"2022-07-23T19:58:43.556765Z","shell.execute_reply.started":"2022-07-23T19:58:43.545289Z","shell.execute_reply":"2022-07-23T19:58:43.555508Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = pd.crosstab(index=df_train.patient_id, columns=df_train.label)\ndf.loc[df.CE>1,\"CE\"]=1\ndf.loc[df.LAA>1,\"LAA\"]=1\n\npCE = df.CE.sum()/n_pat\npLAA = df.LAA.sum()/n_pat\n\nif df.sum().sum() == n_pat:\n    print(\"Label is unique by Patient\")\nelse:\n    print(\"Label not is unique by Patient\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-23T19:58:45.268629Z","iopub.execute_input":"2022-07-23T19:58:45.269811Z","iopub.status.idle":"2022-07-23T19:58:45.310785Z","shell.execute_reply.started":"2022-07-23T19:58:45.269746Z","shell.execute_reply":"2022-07-23T19:58:45.309079Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, axs = plt.subplots(nrows=1, ncols=2, figsize=(14, 7))\naxs[0] = sn.countplot(x=\"label\", data=df_train, ax=axs[0])\naxs[0].bar_label(axs[0].containers[0])\naxs[1] = df_train[[\"label\"]].value_counts().plot.pie(autopct='%1.1f%%', \n                                                     ylabel=\"label\", \n                                                     labels = [\"CE\",\"LAA\"], \n                                                     shadow=True)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-10T20:59:51.613546Z","iopub.execute_input":"2022-07-10T20:59:51.613951Z","iopub.status.idle":"2022-07-10T20:59:51.838164Z","shell.execute_reply.started":"2022-07-10T20:59:51.613916Z","shell.execute_reply":"2022-07-10T20:59:51.837011Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = df_train[[\"center_id\"]].value_counts().reset_index(name=\"cnt\")\nlb = list(df.center_id.values)\nod = df.sort_values(\"cnt\", ascending=False).center_id\n\nfig, axs = plt.subplots(nrows=1, ncols=2, figsize=(14, 7))\ncl = sn.color_palette(\"Paired\")\naxs[0] = sn.barplot(x=\"center_id\",  y=\"cnt\", data=df, \n                    order=od, palette=cl, ax=axs[0])\naxs[0].bar_label(axs[0].containers[0])\naxs[1] = df[\"cnt\"].plot.pie(autopct='%1.1f%%', ylabel=\"center_id\", \n                            labels=lb, shadow=True, colors=cl)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-10T20:59:51.839985Z","iopub.execute_input":"2022-07-10T20:59:51.840364Z","iopub.status.idle":"2022-07-10T20:59:52.345466Z","shell.execute_reply.started":"2022-07-10T20:59:51.840332Z","shell.execute_reply":"2022-07-10T20:59:52.344296Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Image_num distribution","metadata":{}},{"cell_type":"code","source":"fig, axs = plt.subplots(nrows=1, ncols=2, figsize=(14, 7))\nlb = list(range(5))\nex = (0,0.1,0.1,0.4,0.6)\naxs[0] = sn.countplot(x=\"image_num\", data=df_train, dodge=False, ax=axs[0])\naxs[0].bar_label(axs[0].containers[0])\naxs[1] = df_train[[\"image_num\"]].value_counts().plot.pie(autopct='%1.1f%%', \n                                                         ylabel=\"image_num\", \n                                                         labels=lb, explode=ex)\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-10T20:59:52.34747Z","iopub.execute_input":"2022-07-10T20:59:52.347796Z","iopub.status.idle":"2022-07-10T20:59:52.664393Z","shell.execute_reply.started":"2022-07-10T20:59:52.347767Z","shell.execute_reply":"2022-07-10T20:59:52.663173Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Patients with more than two images","metadata":{}},{"cell_type":"code","source":"df = (df_train.groupby([\"patient_id\",\"label\"])[\"image_num\"].count().reset_index(name='image_count'))\ndf[df[\"image_count\"]>2].set_index(\"patient_id\").style.background_gradient(cmap='Reds')","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-10T20:59:52.666279Z","iopub.execute_input":"2022-07-10T20:59:52.666687Z","iopub.status.idle":"2022-07-10T20:59:52.689403Z","shell.execute_reply.started":"2022-07-10T20:59:52.666651Z","shell.execute_reply":"2022-07-10T20:59:52.688298Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Number of Centers by Patients","metadata":{}},{"cell_type":"code","source":"df = df_train.groupby([\"patient_id\"])[\"center_id\"].count().reset_index(name=\"center_count\")\nlb = list(range(1,6))\nfig, axs = plt.subplots(nrows=1, ncols=2, figsize=(14, 7))\naxs[0] = sn.countplot(x=\"center_count\", data=df, dodge=False, ax=axs[0])\naxs[0].bar_label(axs[0].containers[0])\naxs[1] = df[[\"center_count\"]].value_counts().plot.pie(autopct='%1.1f%%', \n                                                      ylabel=\"center_count\", \n                                                      labels=lb, explode=ex)\nplt.show()\n","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-10T20:59:52.692013Z","iopub.execute_input":"2022-07-10T20:59:52.692927Z","iopub.status.idle":"2022-07-10T20:59:53.011521Z","shell.execute_reply.started":"2022-07-10T20:59:52.69289Z","shell.execute_reply":"2022-07-10T20:59:53.010411Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Labels by Center","metadata":{}},{"cell_type":"code","source":"df = pd.crosstab(index=df_train.center_id, columns=df_train.label).reset_index()\ndf[\"LAA/CE\"] = df.LAA/df.CE\nax = df.plot(x=\"center_id\", y=[\"CE\",\"LAA\"], kind=\"bar\", width=0.8, figsize=(14,7))\nax.bar_label(ax.containers[0])\nax.bar_label(ax.containers[1])\ndf.plot(y=[\"LAA/CE\"], secondary_y=\"LAA/CE\", color=\"lightgreen\", linewidth=4, ax=ax);","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-10T20:59:53.013401Z","iopub.execute_input":"2022-07-10T20:59:53.013843Z","iopub.status.idle":"2022-07-10T20:59:53.47252Z","shell.execute_reply.started":"2022-07-10T20:59:53.0138Z","shell.execute_reply":"2022-07-10T20:59:53.471289Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Labels by Image_num","metadata":{"execution":{"iopub.status.busy":"2022-07-09T01:17:20.150707Z","iopub.execute_input":"2022-07-09T01:17:20.151179Z","iopub.status.idle":"2022-07-09T01:17:20.156957Z","shell.execute_reply.started":"2022-07-09T01:17:20.151144Z","shell.execute_reply":"2022-07-09T01:17:20.155605Z"}}},{"cell_type":"code","source":"df = pd.crosstab(index=df_train.image_num, columns=df_train.label).reset_index()\ndf[\"LAA/CE\"] = df.LAA/df.CE\nax = df.plot(x=\"image_num\", y=[\"CE\",\"LAA\"], kind=\"bar\", width=0.8, figsize=(14,7))\nax.bar_label(ax.containers[0])\nax.bar_label(ax.containers[1]);","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-10T20:59:53.473766Z","iopub.execute_input":"2022-07-10T20:59:53.474489Z","iopub.status.idle":"2022-07-10T20:59:53.737554Z","shell.execute_reply.started":"2022-07-10T20:59:53.474454Z","shell.execute_reply":"2022-07-10T20:59:53.736506Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Include Image Sizes in Train dataset","metadata":{}},{"cell_type":"code","source":"%%time\nsizes = []\nfor name in df_train[\"image_id\"]:\n    img = Image.open(os.path.join(BASE_PATH, \"train\", f\"{name}.tif\"))\n    sizes.append({\"img_height\": img.height, \n                  \"img_width\": img.width, \n                  \"img_size\": img.size[0]*img.size[1]/(1024**2)})\n\ndf_train = pd.concat([pd.DataFrame(sizes),df_train], axis=1)\ndel sizes\ndf_train","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-10T20:59:53.740232Z","iopub.execute_input":"2022-07-10T20:59:53.740592Z","iopub.status.idle":"2022-07-10T20:59:55.348556Z","shell.execute_reply.started":"2022-07-10T20:59:53.740562Z","shell.execute_reply":"2022-07-10T20:59:55.347262Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.describe()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-10T20:59:55.351274Z","iopub.execute_input":"2022-07-10T20:59:55.352217Z","iopub.status.idle":"2022-07-10T20:59:55.383053Z","shell.execute_reply.started":"2022-07-10T20:59:55.352172Z","shell.execute_reply":"2022-07-10T20:59:55.381752Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Image analysis size distribution","metadata":{}},{"cell_type":"code","source":"fig, axs = plt.subplots(nrows=1, ncols=2, figsize=(15, 5))\nfor i, col in enumerate([\"img_height\", \"img_width\"]):\n    _= sn.histplot(df_train[[col]], ax=axs[i], bins=40, kde=True)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-10T20:59:55.384298Z","iopub.execute_input":"2022-07-10T20:59:55.384611Z","iopub.status.idle":"2022-07-10T20:59:55.943841Z","shell.execute_reply.started":"2022-07-10T20:59:55.384582Z","shell.execute_reply":"2022-07-10T20:59:55.942496Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Image Sizes by Center","metadata":{}},{"cell_type":"code","source":"df = df_train.groupby([\"center_id\"])[[\"img_height\", \"img_width\"]].mean().reset_index()\nax = df.plot(x=\"center_id\", kind=\"bar\", width=0.8, figsize=(14,7))","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-10T20:59:55.945409Z","iopub.execute_input":"2022-07-10T20:59:55.946102Z","iopub.status.idle":"2022-07-10T20:59:56.218591Z","shell.execute_reply.started":"2022-07-10T20:59:55.946061Z","shell.execute_reply":"2022-07-10T20:59:56.217421Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"2\"></a>\n<h2 style='background:red; border:0; color:white'><center>Images Visualizations</center></h2>","metadata":{"papermill":{"duration":0.045958,"end_time":"2020-11-30T22:55:55.072135","exception":false,"start_time":"2020-11-30T22:55:55.026177","status":"completed"},"tags":[]}},{"cell_type":"code","source":"#Utility functions\n\ndef read_image(image_id, dset, scale=None, verbose=1):\n    with tifffile.TiffFile(os.path.join(BASE_PATH, dset, f\"{image_id}.tif\")) as tif:\n        tif_tags = {}\n        for tag in tif.pages[0].tags.values():\n            name, value = tag.name, tag.value\n            tif_tags[name] = value\n        del tif_tags[\"TileOffsets\"] \n        del tif_tags[\"TileByteCounts\"]\n        image = tif.pages[0].asarray()\n    \n    if verbose:\n        print(f\"[{image_id}] Image shape: {image.shape}\")\n    \n    if scale:\n        new_size = (image.shape[1] // scale, image.shape[0] // scale)\n        image = cv2.resize(image, new_size, interpolation=cv2.INTER_AREA)\n        if image.shape[1]>1.5*image.shape[0]:\n            out=cv2.transpose(image)\n            image=cv2.flip(out,flipCode=0)\n        \n        if verbose:\n            print(f\"[{image_id}] Resized Image shape: {image.shape}\")\n        \n    return image, tif_tags\n\ndef plot_image(image, image_id):\n    plt.figure(figsize=(16, 10))\n    plt.imshow(image)\n    plt.title(f\"Image {image_id}\", fontsize=18)  \n    plt.axis('off')\n    plt.show()\n    \ndef plot_list_img(sample_ids, dset, scale=20):\n    sample_images = []\n    for sample_id in sample_ids:\n        sample_images.append(read_image(sample_id, dset, scale=scale, verbose=0)[0])\n    plt.figure(figsize=(16, 16))\n    for ind, (tmp_id, tmp_image) in enumerate(zip(sample_ids, sample_images)):\n        plt.subplot(2, 5, ind + 1)\n        plt.imshow(tmp_image)\n        plt.title(f\"{tmp_id}\", fontsize=10) \n        plt.axis(\"off\")","metadata":{"_kg_hide-input":true,"papermill":{"duration":0.092112,"end_time":"2020-11-30T22:55:54.979734","exception":false,"start_time":"2020-11-30T22:55:54.887622","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-10T20:59:56.220158Z","iopub.execute_input":"2022-07-10T20:59:56.221296Z","iopub.status.idle":"2022-07-10T20:59:56.238188Z","shell.execute_reply.started":"2022-07-10T20:59:56.221228Z","shell.execute_reply":"2022-07-10T20:59:56.237001Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Iterative View","metadata":{}},{"cell_type":"code","source":"if eda:\n    img_id = \"026c97_0\"\n    img, tags = read_image(img_id, \"train\", scale=2)\n    print(tags)\n    fig = px.imshow(img)\n    fig.show()","metadata":{"_kg_hide-input":true,"_kg_hide-output":false,"execution":{"iopub.status.busy":"2022-07-10T20:59:56.239881Z","iopub.execute_input":"2022-07-10T20:59:56.240624Z","iopub.status.idle":"2022-07-10T21:00:03.204625Z","shell.execute_reply.started":"2022-07-10T20:59:56.240578Z","shell.execute_reply":"2022-07-10T21:00:03.200408Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Train images ","metadata":{"papermill":{"duration":0.045251,"end_time":"2020-11-30T23:00:01.062717","exception":false,"start_time":"2020-11-30T23:00:01.017466","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"### Images of patient id = 09644e (CE)","metadata":{}},{"cell_type":"code","source":"if eda:\n    sample_ids = [\"09644e_0\",\"09644e_1\",\"09644e_2\",\"09644e_3\",\"09644e_4\"]\n    plot_list_img(sample_ids, \"train\", scale=100)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-10T21:00:03.208674Z","iopub.execute_input":"2022-07-10T21:00:03.210396Z","iopub.status.idle":"2022-07-10T21:01:23.241551Z","shell.execute_reply.started":"2022-07-10T21:00:03.210197Z","shell.execute_reply":"2022-07-10T21:01:23.240563Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Images of patient id = 91b9d3 (LAA)","metadata":{}},{"cell_type":"code","source":"if eda:\n    sample_ids = [\"91b9d3_0\",\"91b9d3_1\",\"91b9d3_2\",\"91b9d3_3\",\"91b9d3_4\"]\n    plot_list_img(sample_ids, \"train\", scale=100)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-10T21:01:23.244886Z","iopub.execute_input":"2022-07-10T21:01:23.245255Z","iopub.status.idle":"2022-07-10T21:02:26.528075Z","shell.execute_reply.started":"2022-07-10T21:01:23.245206Z","shell.execute_reply":"2022-07-10T21:02:26.527016Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### CE class sample","metadata":{}},{"cell_type":"code","source":"if eda:\n    sample_ids = df_train[df_train.label==\"CE\"].image_id[:10].values\n    plot_list_img(sample_ids, \"train\", scale=100)","metadata":{"_kg_hide-input":true,"papermill":{"duration":245.851239,"end_time":"2020-11-30T23:00:00.969759","exception":false,"start_time":"2020-11-30T22:55:55.11852","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-10T21:02:26.52936Z","iopub.execute_input":"2022-07-10T21:02:26.529691Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### LAA class sample","metadata":{"execution":{"iopub.status.busy":"2022-07-07T00:03:41.436386Z","iopub.execute_input":"2022-07-07T00:03:41.437003Z","iopub.status.idle":"2022-07-07T00:03:41.441493Z","shell.execute_reply.started":"2022-07-07T00:03:41.436956Z","shell.execute_reply":"2022-07-07T00:03:41.440452Z"}}},{"cell_type":"code","source":"if eda:\n    sample_ids = df_train[df_train.label==\"LAA\"].image_id[:10].values\n    plot_list_img(sample_ids, \"train\", scale=100)","metadata":{"_kg_hide-output":false,"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Test Images","metadata":{}},{"cell_type":"code","source":"if eda:\n    sample_ids = df_test.image_id[:4].values\n    plot_list_img(sample_ids, \"train\", scale=100)","metadata":{"_kg_hide-input":true,"papermill":{"duration":34.477766,"end_time":"2020-11-30T23:02:52.107762","exception":false,"start_time":"2020-11-30T23:02:17.629996","status":"completed"},"tags":[],"_kg_hide-output":false,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Compare images (Test = Train ?)","metadata":{}},{"cell_type":"code","source":"if eda:\n    for image_id in list(df_test.image_id[:4].values):\n        image_train = read_image(image_id, \"train\", 100, verbose=0)[0]\n        image_test = read_image(image_id, \"test\", 100, verbose=0)[0]\n        diff = cv2.absdiff(image_train, image_test)\n        print(f\"Sum of DIFF of images {image_id} = \", diff.sum())","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"3\"></a>\n<h2 style='background:red; border:0; color:white'><center>Baseline Submission</center></h2>","metadata":{}},{"cell_type":"code","source":"df_sub.CE = 0.48\ndf_sub.LAA = 0.52\ndf_sub.to_csv(\"submission.csv\", index=False)\ndf_sub ","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### WORK IN PROGRESS...","metadata":{"papermill":{"duration":0.275493,"end_time":"2020-11-30T23:11:47.481592","exception":false,"start_time":"2020-11-30T23:11:47.206099","status":"completed"},"tags":[],"_kg_hide-input":true}}]}