{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.7.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":37333,"databundleVersionId":3949526,"sourceType":"competition"}],"dockerImageVersionId":30213,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<p id=\"toc\"></p>\n<h2 class=\"list-group-item list-group-item-action active\" data-toggle=\"list\" style=\"font-family: Verdana; font-size: 24px; font-style: normal; font-weight: bold; text-decoration: none; text-transform: none; letter-spacing: 3px; background-color: #CCCCFF; color: black;\" role=\"tab\" aria-controls=\"home\"><center><br>CONTENTS</center></h2>\n\n<h3 style=\"text-indent: 10vw; font-family: Verdana; font-size: 16px; font-style: normal; font-weight: normal; text-decoration: none; text-transform: none; letter-spacing: 2px; color: black; background-color: #ffffff;\"><a href=\"#background\">0&nbsp;&nbsp;&nbsp;&nbsp;BACKGROUND INFORMATION</a></h3>\n\n---\n\n<h3 style=\"text-indent: 10vw; font-family: Verdana; font-size: 16px; font-style: normal; font-weight: normal; text-decoration: none; text-transform: none; letter-spacing: 2px; color: black; background-color: #ffffff;\"><a href=\"#imports\">1&nbsp;&nbsp;&nbsp;&nbsp;IMPORTS</a></h3>\n\n---\n\n<h3 style=\"text-indent: 10vw; font-family: Verdana; font-size: 16px; font-style: normal; font-weight: normal; text-decoration: none; text-transform: none; letter-spacing: 2px; color: black; background-color: #ffffff;\"><a href=\"#metadata\">2&nbsp;&nbsp;&nbsp;&nbsp;METADATA ANALYSIS</a></h3>\n\n---\n\n<h3 style=\"text-indent: 10vw; font-family: Verdana; font-size: 16px; font-style: normal; font-weight: normal; text-decoration: none; text-transform: none; letter-spacing: 2px; color: black; background-color: #ffffff;\"><a href=\"#img\">3&nbsp;&nbsp;&nbsp;&nbsp;IMAGE ANALYSIS</a></h3>\n\n---\n\n\n<div class=\"list-group\" id=\"list-tab\" role=\"tablist\">\n<h3 class=\"list-group-item list-group-item-action active\" data-toggle=\"list\" style='background:#F08080; border:0; color:black' role=\"tab\" aria-controls=\"home\"><center><br>If you find this notebook useful, do give me an upvote, it motivates me a lot.<br><br> This notebook is still a work in progress. Keep checking for further developments!😊</center></h3>","metadata":{}},{"cell_type":"markdown","source":"<a id=\"background\"></a>\n\n<h2 style=\"font-family: Verdana; font-size: 24px; font-style: normal; font-weight: bold; text-decoration: none; text-transform: none; letter-spacing: 3px; background-color: #CCCCFF; color: black;\" id=\"background\"><left><br>&nbsp0. BACKGROUND INFORMATION <a href=\"#toc\">&#10514;</a><br></left> </h2>\n\n# Task Description:\n\nIn this competition, we'll classify the blood clot origins in acute ischemic stroke (AIS). There are **2** major AIS etiology subtypes:\n\n* **Cardioembolic (CE)**\n* **Large Artery Atherosclerosis (LAA)**\n\n=> This is a **classification** problem\n\n# File + Data Field Descriptions:\n\n1. **train/**: A folder containing images in the TIFF format to be used as training data.\n2. **test/**: A folder containing images to be used as test data. The actual test data comprises about 280 images.\n3. **other/**: A supplemental set of images with a either an unknown etiology or an etiology other than CE or LAA.\n4. **train.csv**: Contains annotations for images in the `train/` folder.\n    * `image_id`: A unique identifier for this instance having the form `{patient_id}_{image_num}`. Corresponds to the image `{image_id}.tif`.\n    * `center_id`: Identifies the medical center where the slide was obtained.\n    * `patient_id`: Identifies the patient from whom the slide was obtained.\n    * `image_num`: Enumerates images of clots obtained from the same patient.\n    * `label`: The etiology of the clot, either CE or LAA. This field is the classification target.\n5. **test.csv**: Annotations for images in the `test/` folder. Has the same fields as train.csv excluding label.\n6. **other.csv**: Annotations for images in the `other/` folder. Has the same fields as train.csv. The `center_id` is unavailable for these images however.\n    * `label`: The etiology of the clot, either `Unknown` or `Other`.\n    * `other_specified`: The specific etiology, when known, in case the etiology is labeled as `Other`.\n7. **sample_submission.csv**: A sample submission file in the correct format. Note in particular that you should make **one prediction** per `patient_id`, not per `image_id`.\n\n\n# Evaluation Metric:\n\nSubmissions are evaluated using a weighted multi-class logarithmic loss. The overall effect is such that each class is roughly equally important for the final score.\n\nEach image has been labeled with an etiology class, either CE or LAA. For each image, you must submit a probability for each class. The formula is then:\n\n$$Log Loss=−\\frac{\\sum_{i=1}^{M}w_i.\\sum_{j=1}^{N_i}\\frac{y_{ij}}{N_i}.ln p_{ij}}{\\sum_{i=1}^{M}w_i}$$\n\nwhere 𝑁 is the number of images in the class set, 𝑀 is the number of classes,  ln is the natural logarithm, $y_{ij}$ is 1 if observation 𝑖 belongs to class 𝑗 and 0 otherwise, $p_{ij}$ is the predicted probability that image 𝑖 belongs to class 𝑗.\n\nThe submitted probabilities for a given image are not required to sum to one because they are rescaled prior to being scored (each row is divided by the row sum). In order to avoid the extremes of the log function, each predicted probability 𝑝 is replaced with $max(min(p,1−10^{15}),10^{15})$.\n\n# Submission File Format:\n\nFor each `patient_id` in the test set, you must predict a probability for each of the two etiology classes. The file should contain a header and have the following format:\n\n```\npatient_id,CE,LAA\n01f2b3,0.5,0.5\n04de22,0.5,0.5\n0a47c9,0.5,0.5\n0af8b6,0.5,0.5\n...\n\n```","metadata":{}},{"cell_type":"markdown","source":"<a id=\"imports\"></a>\n\n<h2 style=\"font-family: Verdana; font-size: 24px; font-style: normal; font-weight: bold; text-decoration: none; text-transform: none; letter-spacing: 3px; background-color: #CCCCFF; color: black;\" id=\"imports\"><left><br>&nbsp1. IMPORTS <a href=\"#toc\">&#10514;</a><br></left> </h2>","metadata":{}},{"cell_type":"code","source":"import os\nimport sys\nimport pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport cv2\n\nfrom glob import glob\nfrom pprint import pprint\nfrom collections import defaultdict\nimport gc\n\nimport plotly\nfrom plotly import tools\nimport plotly.graph_objects as go\nfrom plotly.subplots import make_subplots\nimport plotly.express as px\nimport plotly.offline as pyo\nimport plotly.io as pio\nimport plotly.graph_objects as go\n#pio.templates.default = 'plotly_white'\nsns.set_theme(style=\"dark\")\nimport cv2\nimport tifffile as tiff\nfrom PIL import Image\n\nimport warnings\nwarnings.simplefilter(\"ignore\")","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2024-05-09T09:59:41.294513Z","iopub.execute_input":"2024-05-09T09:59:41.295091Z","iopub.status.idle":"2024-05-09T09:59:45.046818Z","shell.execute_reply.started":"2024-05-09T09:59:41.294977Z","shell.execute_reply":"2024-05-09T09:59:45.045311Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load Datasets","metadata":{}},{"cell_type":"code","source":"train_df = pd.read_csv('../input/mayo-clinic-strip-ai/train.csv')\ntest_df = pd.read_csv('../input/mayo-clinic-strip-ai/test.csv')\nother_df = pd.read_csv('../input/mayo-clinic-strip-ai/other.csv')","metadata":{"execution":{"iopub.status.busy":"2024-05-09T09:59:45.050416Z","iopub.execute_input":"2024-05-09T09:59:45.050939Z","iopub.status.idle":"2024-05-09T09:59:45.082671Z","shell.execute_reply.started":"2024-05-09T09:59:45.0509Z","shell.execute_reply":"2024-05-09T09:59:45.081467Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## *Train Dataset*","metadata":{}},{"cell_type":"code","source":"gmap = np.array([[1,2,3], [2,3,4], [1,2,3],[2,3,4],[1,2,3]])\ntrain_df.head(5).style.background_gradient(axis=None,gmap=gmap, cmap='Purples', \n                                            subset=['image_id','patient_id','label'])","metadata":{"execution":{"iopub.status.busy":"2024-05-09T09:59:45.087894Z","iopub.execute_input":"2024-05-09T09:59:45.088553Z","iopub.status.idle":"2024-05-09T09:59:45.211367Z","shell.execute_reply.started":"2024-05-09T09:59:45.088513Z","shell.execute_reply":"2024-05-09T09:59:45.209575Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## *Test Dataset*","metadata":{}},{"cell_type":"code","source":"test_df.head(5).style.background_gradient(axis=None,gmap=[[1,2],[2,3],[1,2],[2,3]], \n                                          cmap='Purples', \n                                          subset=['image_id','patient_id'])","metadata":{"execution":{"iopub.status.busy":"2024-05-09T09:59:45.2133Z","iopub.execute_input":"2024-05-09T09:59:45.213672Z","iopub.status.idle":"2024-05-09T09:59:45.232711Z","shell.execute_reply.started":"2024-05-09T09:59:45.213637Z","shell.execute_reply":"2024-05-09T09:59:45.231262Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## *Other Dataset*","metadata":{}},{"cell_type":"code","source":"other_df.head(5).style.background_gradient(axis=None,gmap=gmap, cmap='Purples', \n                                            subset=['image_id','image_num','label'])","metadata":{"execution":{"iopub.status.busy":"2024-05-09T09:59:45.234746Z","iopub.execute_input":"2024-05-09T09:59:45.235272Z","iopub.status.idle":"2024-05-09T09:59:45.259811Z","shell.execute_reply.started":"2024-05-09T09:59:45.23522Z","shell.execute_reply":"2024-05-09T09:59:45.257155Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"metadata\"></a>\n\n<h2 style=\"font-family: Verdana; font-size: 24px; font-style: normal; font-weight: bold; text-decoration: none; text-transform: none; letter-spacing: 3px; background-color: #CCCCFF; color: black;\" id=\"metadata\"><left><br>&nbsp2. METADATA ANALYSIS <a href=\"#toc\">&#10514;</a><br></left> </h2>","metadata":{}},{"cell_type":"markdown","source":"# Statistical Description","metadata":{}},{"cell_type":"code","source":"# https://www.kaggle.com/code/toomuchsauce/mental-health-plotly-interactive-viz\ndef EDA(df):\n    \n    print('\\033[1m' +'EXPLORATORY DATA ANALYSIS :'+ '\\033[0m\\n')\n    print('\\033[1m' + 'Shape of the data (rows, columns):' + '\\033[0m')\n    print(df.shape, \n          '\\n------------------------------------------------------------------------------------\\n')\n    \n    print('\\033[1m' + 'All columns from the dataframe :' + '\\033[0m')\n    print(df.columns, \n          '\\n------------------------------------------------------------------------------------\\n')\n    \n    print('\\033[1m' + 'Datatypes and Missing values:' + '\\033[0m')\n    print(df.info(), \n          '\\n------------------------------------------------------------------------------------\\n')\n    \n    for col in df.columns:\n        print('\\033[1m' + 'Unique values in {} :'.format(col) + '\\033[0m',len(df[col].unique()))\n    print('\\n------------------------------------------------------------------------------------\\n')\n    \n    print('\\033[1m' + 'Summary statistics for the data :' + '\\033[0m')\n    print(df.describe(include='all'), \n          '\\n------------------------------------------------------------------------------------\\n')\n    \n        \n    print('\\033[1m' + 'Memory used by the data :' + '\\033[0m')\n    print(df.memory_usage(), \n          '\\n------------------------------------------------------------------------------------\\n')\n    \n    print('\\033[1m' + 'Number of duplicate values :' + '\\033[0m')\n    print(df.duplicated().sum())\n          \nEDA(train_df)","metadata":{"execution":{"iopub.status.busy":"2024-05-09T09:59:45.262172Z","iopub.execute_input":"2024-05-09T09:59:45.262802Z","iopub.status.idle":"2024-05-09T09:59:45.31593Z","shell.execute_reply.started":"2024-05-09T09:59:45.262746Z","shell.execute_reply":"2024-05-09T09:59:45.314226Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data Visualization","metadata":{}},{"cell_type":"markdown","source":"## *Label Distribution*","metadata":{}},{"cell_type":"code","source":"labels_train = train_df.groupby('label')['label'].count()\nlabels_other = other_df.groupby('label')['label'].count()","metadata":{"execution":{"iopub.status.busy":"2024-05-09T09:59:45.317721Z","iopub.execute_input":"2024-05-09T09:59:45.318132Z","iopub.status.idle":"2024-05-09T09:59:45.327876Z","shell.execute_reply.started":"2024-05-09T09:59:45.318094Z","shell.execute_reply":"2024-05-09T09:59:45.326774Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# https://stackoverflow.com/questions/34888058/changing-width-of-bars-in-bar-chart-created-using-seaborn-factorplot\ndef change_width(ax, new_value) :\n    for patch in ax.patches :\n        current_width = patch.get_width()\n        diff = current_width - new_value\n\n        # we change the bar width\n        patch.set_width(new_value)\n\n        # we recenter the bar\n        patch.set_x(patch.get_x() + diff * .5)\n","metadata":{"execution":{"iopub.status.busy":"2024-05-09T09:59:45.329407Z","iopub.execute_input":"2024-05-09T09:59:45.329811Z","iopub.status.idle":"2024-05-09T09:59:45.343655Z","shell.execute_reply.started":"2024-05-09T09:59:45.32976Z","shell.execute_reply":"2024-05-09T09:59:45.342025Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.set_style(\"darkgrid\")\nfig, ax = plt.subplots(1,2, figsize=(16,6))\ng1 = sns.barplot(x=labels_train.index, y=labels_train.values, ax=ax[0])\nax[0].set_title(\"Distribution of Train Label\",fontsize=16)\nax[0].set_ylabel(\"Count\")\nax[0].set_xlabel(\"Label\")\ng1.grid(True)\nchange_width(ax[0],0.50)\nfor p in g1.patches:\n    x, w, h = p.get_x(), p.get_width(), p.get_height()\n    if h > 0:\n        g1.text(x + w / 2, h, f'{h:.0f}\\n', ha='center', va='center', size=11)\ng2 = sns.barplot(x=labels_other.index, y=labels_other.values, ax=ax[1])\nax[1].set_title(\"Distribution of Other Label\",fontsize=16)\nax[1].set_ylabel(\"Count\")\nax[1].set_xlabel(\"Label\")\ng2.grid(True)\nchange_width(ax[1],0.50)\nfor p in g2.patches:\n    x, w, h = p.get_x(), p.get_width(), p.get_height()\n    if h > 0:\n        g2.text(x + w / 2, h, f'{h:.0f}\\n', ha='center', va='center', size=11)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-05-09T09:59:45.348801Z","iopub.execute_input":"2024-05-09T09:59:45.349308Z","iopub.status.idle":"2024-05-09T09:59:45.831186Z","shell.execute_reply.started":"2024-05-09T09:59:45.349267Z","shell.execute_reply":"2024-05-09T09:59:45.830138Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## *Distribution of Labels by Center*","metadata":{}},{"cell_type":"code","source":"sns.set_style(\"dark\")\nplt.figure(figsize=(15,8))\ng = sns.histplot(data=train_df, x=\"center_id\", hue=\"label\", \n                 multiple=\"dodge\", color='label',discrete=True,\n                 kde=True, shrink=0.8)\ng.set(xlabel='Center ID', ylabel='Label Count')\ng.set(xlim=(0, 12), xticks=np.arange(0,12,1))\nfor p in g.patches:\n    x, w, h = p.get_x(), p.get_width(), p.get_height()\n    if h > 0:\n        g.text(x + w / 2, h, f'{h}\\n', ha='center', va='center', size=11)\ng.margins(y=0.07)\ng.grid(True)\ng.set_title('Distribution of Labels by Center', fontsize=16)","metadata":{"execution":{"iopub.status.busy":"2024-05-09T09:59:45.837231Z","iopub.execute_input":"2024-05-09T09:59:45.840749Z","iopub.status.idle":"2024-05-09T09:59:46.451128Z","shell.execute_reply.started":"2024-05-09T09:59:45.840618Z","shell.execute_reply":"2024-05-09T09:59:46.450183Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## *Distribution of Patients per Center*","metadata":{}},{"cell_type":"code","source":"labels = train_df.groupby('label')['label'].count()\ncenters = train_df.groupby(\"center_id\")['center_id'].count()","metadata":{"execution":{"iopub.status.busy":"2024-05-09T09:59:46.452478Z","iopub.execute_input":"2024-05-09T09:59:46.453888Z","iopub.status.idle":"2024-05-09T09:59:46.462619Z","shell.execute_reply.started":"2024-05-09T09:59:46.453845Z","shell.execute_reply":"2024-05-09T09:59:46.461187Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.set_style(\"darkgrid\")\nplt.figure(figsize=(15,8))\n\ng = sns.barplot(x=centers.index, y=centers.values, orient='v')\nfor p in g.patches:\n    x, w, h = p.get_x(), p.get_width(), p.get_height()\n    if h > 0:\n        g.text(x + w / 2, h, f'{h:.0f}\\n', ha='center', va='center', size=11)\ng.set_title('Distribution of Patients per Center', fontsize=16)\ng.set(ylabel=\"Patient Count\", xlabel='Center ID')\ng.grid(True)","metadata":{"execution":{"iopub.status.busy":"2024-05-09T09:59:46.464035Z","iopub.execute_input":"2024-05-09T09:59:46.464489Z","iopub.status.idle":"2024-05-09T09:59:46.848065Z","shell.execute_reply.started":"2024-05-09T09:59:46.464454Z","shell.execute_reply":"2024-05-09T09:59:46.846241Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"img\"></a>\n\n<h2 style=\"font-family: Verdana; font-size: 24px; font-style: normal; font-weight: bold; text-decoration: none; text-transform: none; letter-spacing: 3px; background-color: #CCCCFF; color: black;\" id=\"img\"><left><br>&nbsp3. IMAGE ANALYSIS <a href=\"#toc\">&#10514;</a><br></left> </h2>","metadata":{}},{"cell_type":"markdown","source":"# Image Visualization","metadata":{}},{"cell_type":"code","source":"# https://www.kaggle.com/code/datark1/eda-images-processing-and-exploration/notebook\n# ../input/mayo-clinic-strip-ai/train/006388_0.tif -> id=image_path[-12:-4]\n# add image path, width, height to metadata\nImage.MAX_IMAGE_PIXELS = None \ntrain_images = glob(\"/kaggle/input/mayo-clinic-strip-ai/train/*\")\n\nimg_features = defaultdict(list)\n\nfor img_path in train_images:\n    img = Image.open(img_path)\n    img_features['image_id'].append(img_path[-12:-4])\n    img_features['width'].append(img.size[0])\n    img_features['height'].append(img.size[1])\n    img_features['img_path'].append(img_path)\n    \nimg_data = pd.DataFrame(img_features)\nimg_data.sort_values(by='image_id', inplace=True)\nimg_data.reset_index(inplace=True, drop=True)\nimg_data = img_data.merge(train_df, on='image_id')","metadata":{"execution":{"iopub.status.busy":"2024-05-09T09:59:46.849916Z","iopub.execute_input":"2024-05-09T09:59:46.850354Z","iopub.status.idle":"2024-05-09T10:00:03.315414Z","shell.execute_reply.started":"2024-05-09T09:59:46.850315Z","shell.execute_reply":"2024-05-09T10:00:03.313848Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"img_data.head().style.background_gradient(axis=None,gmap=gmap, cmap='Purples', \n                                            subset=['image_id','img_path','label'])    ","metadata":{"execution":{"iopub.status.busy":"2024-05-09T10:00:03.318123Z","iopub.execute_input":"2024-05-09T10:00:03.318896Z","iopub.status.idle":"2024-05-09T10:00:03.340853Z","shell.execute_reply.started":"2024-05-09T10:00:03.318851Z","shell.execute_reply":"2024-05-09T10:00:03.338954Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"Image.MAX_IMAGE_PIXELS = None  # you have to set this value to allow displaying high-resolution images\n\n# paths to CE and LAA images\nCE_imgs = img_data.loc[img_data['label']=='CE','img_path']\nLAA_imgs = img_data.loc[img_data['label']=='LAA','img_path']","metadata":{"execution":{"iopub.status.busy":"2024-05-09T10:00:03.342841Z","iopub.execute_input":"2024-05-09T10:00:03.343294Z","iopub.status.idle":"2024-05-09T10:00:03.358684Z","shell.execute_reply.started":"2024-05-09T10:00:03.343256Z","shell.execute_reply":"2024-05-09T10:00:03.356511Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## *Batches of Patients with CE*","metadata":{}},{"cell_type":"code","source":"plt.style.use('default')\nfig = plt.figure(figsize=(16,16))\nfor i, path in enumerate(CE_imgs[:4]):\n    plt.subplot(2,2,i+1)\n    img = Image.open(path)\n    original_shape = img.size\n    file_name = path[-12:-4]\n    img.thumbnail((300,300), Image.Resampling.LANCZOS)\n    if img.height > img.width:\n        img = img.transpose(Image.Resampling.LANCZOS)\n    resampled_shape = img.size\n    plt.imshow(img)\n    plt.title(f'Image ID: {file_name}\\nOriginal shape: {original_shape}\\nResampled shape: {resampled_shape}')\nplt.suptitle(\"Etiology Type: CE\", fontsize=16)","metadata":{"execution":{"iopub.status.busy":"2024-05-09T10:00:03.361101Z","iopub.execute_input":"2024-05-09T10:00:03.361619Z","iopub.status.idle":"2024-05-09T10:01:00.721136Z","shell.execute_reply.started":"2024-05-09T10:00:03.361576Z","shell.execute_reply":"2024-05-09T10:01:00.719252Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## *Batches of Patients with LAA*","metadata":{}},{"cell_type":"code","source":"plt.style.use('default')\nfig = plt.figure(figsize=(16,16))\nfor i, path in enumerate(LAA_imgs[:4]):\n    plt.subplot(2,2,i+1)\n    img = Image.open(path)\n    original_shape = img.size\n    file_name = path[-12:-4]\n    img.thumbnail((300,300), Image.Resampling.LANCZOS)\n    if img.height > img.width:\n        img = img.transpose(Image.Resampling.LANCZOS)\n    resampled_shape = img.size\n    plt.imshow(img)\n    plt.title(f'Image ID: {file_name}\\nOriginal shape: {original_shape}\\nResampled shape: {resampled_shape}')\nplt.suptitle(\"Etiology Type: LAA\", fontsize=16)","metadata":{"execution":{"iopub.status.busy":"2024-05-09T10:01:00.72321Z","iopub.execute_input":"2024-05-09T10:01:00.724018Z","iopub.status.idle":"2024-05-09T10:02:35.582479Z","shell.execute_reply.started":"2024-05-09T10:01:00.723971Z","shell.execute_reply":"2024-05-09T10:02:35.581078Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Case Study Analysis","metadata":{}},{"cell_type":"markdown","source":"## *Helper Functions*","metadata":{}},{"cell_type":"code","source":"def highlight(rows):\n    df = lambda x: ['background: #CCCCFF' if x.name in rows\n                        else '' for i in x]\n    return df","metadata":{"execution":{"iopub.status.busy":"2024-05-09T10:02:35.584028Z","iopub.execute_input":"2024-05-09T10:02:35.585104Z","iopub.status.idle":"2024-05-09T10:02:35.592507Z","shell.execute_reply.started":"2024-05-09T10:02:35.585053Z","shell.execute_reply":"2024-05-09T10:02:35.591045Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def case_study(df, img_id):\n    plt.style.use('default')\n    \n    path = df.loc[df['image_id'] == img_id]['img_path'].values[0]\n    patient_id = df.loc[df['image_id'] == img_id]['patient_id'].values[0]\n    center_id = df.loc[df['image_id'] == img_id]['center_id'].values[0]\n    label = df.loc[df['image_id'] == img_id]['label'].values[0]\n    \n    # https://www.kaggle.com/code/datark1/eda-images-processing-and-exploration/notebook \n    img = cv2.imread(path, cv2.COLOR_BGR2RGB)\n    image_resized = cv2.resize(img, (0,0), fx=0.05, fy=0.05)  \n    \n    print('\\033[1m' + 'Case study : {}'.format(img_id) + '\\033[0m')\n    print('\\n------------------------------\\n')\n    print('\\033[1m' + 'General info: \\n' + '\\033[0m')\n    print('\\033[1m' + 'Patient ID: ' '\\033[0m' + f'{patient_id}')  \n    print('\\033[1m' + 'Center ID: ' '\\033[0m' + f'{center_id}')  \n    print('\\033[1m' + 'Etiology type: ' '\\033[0m' + f'{label}')\n    print('\\n------------------------------\\n')\n    print('\\033[1m' + 'Image info: \\n' + '\\033[0m')\n    print('\\033[1m' + 'Original shape: \\n' + '\\033[0m' + f'{img.shape}')\n    print('\\033[1m' + 'Resized shape: \\n' + '\\033[0m' + f'{image_resized.shape}')\n    \n    fig, ax = plt.subplots(1,2, figsize=(12,12))\n    ax[0].imshow(img)\n    ax[0].set_title('Original image')\n    \n    ax[1].imshow(image_resized)\n    ax[1].set_title('Resized image')\n    \n    # plt.suptitle('Patient ID: ' + f'{patient_id}', y=0.7)\n    plt.tight_layout()\n    plt.show()\n    return img, image_resized\n\n    ","metadata":{"execution":{"iopub.status.busy":"2024-05-09T10:02:35.594654Z","iopub.execute_input":"2024-05-09T10:02:35.595119Z","iopub.status.idle":"2024-05-09T10:02:35.612004Z","shell.execute_reply.started":"2024-05-09T10:02:35.595079Z","shell.execute_reply":"2024-05-09T10:02:35.610239Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def color_channel_analysis(resized_img):\n    colors = ['Red', 'Green', 'Blue']\n    plt.figure(figsize=(16,16))\n    for i in range(0,9,3):\n        color = int(i/3)\n\n        # color channel\n        plt.subplot(3, 3, i + 1)\n        plt.imshow(resized_img[:,:,color], aspect='auto') # align all subplots\n        plt.title(f'{colors[color]} Channel')\n\n        # histogram\n        ax = plt.subplot(3, 3, i + 2)\n        df = resized_img[:,:,color].ravel()\n        ax.axvline(150,c=\"red\",linestyle=\"--\")\n        sns.histplot(df, bins=np.arange(0,255), ax=ax)\n        plt.title(f'{colors[color]} Channel Histogram')\n\n        # threshold\n        plt.subplot(3, 3, i + 3) # threshold\n        thresh = resized_img.copy()\n        thresh[np.array(resized_img[:,:,color] < 150)] = 0\n        plt.imshow(thresh, aspect='auto')\n        plt.title(f'Threshold {colors[color]} Channel')\n    \n    #plt.suptitle('Color Channel Analysis', y = 1.0, fontsize=16)\n    plt.tight_layout()\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2024-05-09T10:02:35.614254Z","iopub.execute_input":"2024-05-09T10:02:35.614678Z","iopub.status.idle":"2024-05-09T10:02:35.631546Z","shell.execute_reply.started":"2024-05-09T10:02:35.614641Z","shell.execute_reply":"2024-05-09T10:02:35.630071Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## *Cardioembolic (CE)*","metadata":{}},{"cell_type":"code","source":"# Select img with width, height < 16k\nimg_data.loc[img_data['label']=='CE'].head(10).style.apply(highlight([4]), axis=1)","metadata":{"execution":{"iopub.status.busy":"2024-05-09T10:02:35.633313Z","iopub.execute_input":"2024-05-09T10:02:35.634569Z","iopub.status.idle":"2024-05-09T10:02:35.674758Z","shell.execute_reply.started":"2024-05-09T10:02:35.634516Z","shell.execute_reply":"2024-05-09T10:02:35.673261Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"img_CE, img_resized_CE = case_study(img_data, '026c97_0')","metadata":{"execution":{"iopub.status.busy":"2024-05-09T10:02:35.676183Z","iopub.execute_input":"2024-05-09T10:02:35.676545Z","iopub.status.idle":"2024-05-09T10:02:46.535547Z","shell.execute_reply.started":"2024-05-09T10:02:35.676512Z","shell.execute_reply":"2024-05-09T10:02:46.534121Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"color_channel_analysis(img_resized_CE)","metadata":{"execution":{"iopub.status.busy":"2024-05-09T10:02:46.537359Z","iopub.execute_input":"2024-05-09T10:02:46.537888Z","iopub.status.idle":"2024-05-09T10:02:51.381931Z","shell.execute_reply.started":"2024-05-09T10:02:46.53784Z","shell.execute_reply":"2024-05-09T10:02:51.370494Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## *Large Artery Atherosclerosis (LAA)*","metadata":{}},{"cell_type":"code","source":"# Select img with width, height < 16k\nimg_data.loc[img_data['label']=='LAA'].head(10).style.apply(highlight([24]), axis=1)","metadata":{"execution":{"iopub.status.busy":"2024-05-09T10:02:51.384112Z","iopub.execute_input":"2024-05-09T10:02:51.384831Z","iopub.status.idle":"2024-05-09T10:02:51.424603Z","shell.execute_reply.started":"2024-05-09T10:02:51.384778Z","shell.execute_reply":"2024-05-09T10:02:51.422103Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"img_LAA, img_resized_LAA = case_study(img_data, '08d3d8_0')","metadata":{"execution":{"iopub.status.busy":"2024-05-09T10:02:51.427131Z","iopub.execute_input":"2024-05-09T10:02:51.427609Z","iopub.status.idle":"2024-05-09T10:03:13.816757Z","shell.execute_reply.started":"2024-05-09T10:02:51.427567Z","shell.execute_reply":"2024-05-09T10:03:13.815741Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"color_channel_analysis(img_resized_LAA)","metadata":{"execution":{"iopub.status.busy":"2024-05-09T10:03:13.818196Z","iopub.execute_input":"2024-05-09T10:03:13.819142Z","iopub.status.idle":"2024-05-09T10:03:19.542786Z","shell.execute_reply.started":"2024-05-09T10:03:13.819093Z","shell.execute_reply":"2024-05-09T10:03:19.540514Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}