{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Mayo Clinic - STRIP AI \nImage Classification of Stroke Blood Clot Origin\n\n![banner](https://storage.googleapis.com/kaggle-competitions/kaggle/37333/logos/header.png?t=2022-06-29-00-47-20)\n\nThe goal of this competition is to classify the blood clot origins in ischemic stroke. Using whole slide digital pathology images, you'll build a model that differentiates between the two major acute ischemic stroke (AIS) etiology subtypes: cardiac and large artery atherosclerosis.","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"# Lets see what we've for dinner 🍴","metadata":{}},{"cell_type":"markdown","source":"The dataset for this competition comprises over a thousand high-resolution whole-slide digital pathology images. Each slide depicts a blood clot from a patient that had experienced an acute ischemic stroke.\n\nThe slides comprising the training and test sets depict clots with an etiology (that is, origin) known to be either `CE` (Cardioembolic) or `LAA` (Large Artery Atherosclerosis). We include a set of supplemental slides with a either an unknown etiology or an etiology other than CE or LAA.\n\nOur task is to classify the etiology (CE or LAA) of the slides in the test set for each patient.","metadata":{}},{"cell_type":"code","source":"!ls ../input/mayo-clinic-strip-ai","metadata":{"execution":{"iopub.status.busy":"2022-07-22T10:35:21.217985Z","iopub.execute_input":"2022-07-22T10:35:21.218428Z","iopub.status.idle":"2022-07-22T10:35:22.015387Z","shell.execute_reply.started":"2022-07-22T10:35:21.218389Z","shell.execute_reply":"2022-07-22T10:35:22.013677Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* `train/` - A folder containing images in the TIFF format to be used as training data.\n* `test/` - A folder containing images to be used as test data. The actual test data comprises about * 280 images.\n* `other/` - A supplemental set of images with a either an unknown etiology or an etiology other than CE or LAA.\n* `*.csv` files contains annotations for images in the `train/`, `test/` and `other/` folder. ","metadata":{}},{"cell_type":"markdown","source":"## Usual imports...🚶‍♀️🚶‍♂️","metadata":{}},{"cell_type":"code","source":"import cv2 \nimport glob\nimport pandas as pd \nimport numpy as np\nimport seaborn as sns\nimport matplotlib.pyplot as plt","metadata":{"execution":{"iopub.status.busy":"2022-07-22T18:08:50.959197Z","iopub.execute_input":"2022-07-22T18:08:50.959809Z","iopub.status.idle":"2022-07-22T18:08:51.812068Z","shell.execute_reply.started":"2022-07-22T18:08:50.959698Z","shell.execute_reply":"2022-07-22T18:08:51.810578Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Analyzing train.csv 🔍","metadata":{}},{"cell_type":"code","source":"train_df = pd.read_csv(\"../input/mayo-clinic-strip-ai/train.csv\")","metadata":{"execution":{"iopub.status.busy":"2022-07-22T11:08:28.987154Z","iopub.execute_input":"2022-07-22T11:08:28.988333Z","iopub.status.idle":"2022-07-22T11:08:29.006398Z","shell.execute_reply.started":"2022-07-22T11:08:28.988282Z","shell.execute_reply":"2022-07-22T11:08:29.005607Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-22T11:08:35.64252Z","iopub.execute_input":"2022-07-22T11:08:35.643236Z","iopub.status.idle":"2022-07-22T11:08:35.663274Z","shell.execute_reply.started":"2022-07-22T11:08:35.643192Z","shell.execute_reply":"2022-07-22T11:08:35.662101Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Description**\n\n* image_id : {patient_id}_{image_num}\n* patient_id : Unique ID of a patient\n* image_num : image number in slide for images w.r.t. patient (0:1st image, 1:2nd image, ... ) \n* label : type of stroke\n* center_id : id of center where sample was taken","metadata":{}},{"cell_type":"code","source":"train_df.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-22T11:10:12.89434Z","iopub.execute_input":"2022-07-22T11:10:12.894771Z","iopub.status.idle":"2022-07-22T11:10:12.901827Z","shell.execute_reply.started":"2022-07-22T11:10:12.894735Z","shell.execute_reply":"2022-07-22T11:10:12.901003Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## lets confirm absence of null values ","metadata":{}},{"cell_type":"code","source":"train_df.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-22T11:11:29.2715Z","iopub.execute_input":"2022-07-22T11:11:29.271884Z","iopub.status.idle":"2022-07-22T11:11:29.281099Z","shell.execute_reply.started":"2022-07-22T11:11:29.271854Z","shell.execute_reply":"2022-07-22T11:11:29.280361Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"good ","metadata":{}},{"cell_type":"markdown","source":"## Patients","metadata":{}},{"cell_type":"code","source":"train_df[\"patient_id\"].nunique()","metadata":{"execution":{"iopub.status.busy":"2022-07-22T11:21:00.361088Z","iopub.execute_input":"2022-07-22T11:21:00.361816Z","iopub.status.idle":"2022-07-22T11:21:00.370614Z","shell.execute_reply.started":"2022-07-22T11:21:00.36177Z","shell.execute_reply":"2022-07-22T11:21:00.369404Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**632 unique patients** but do we have 632 unique samples?  \ncuz we can also have a situation where single patient may've given sample multiple times?\n\nLet's find it out by checking unique values w.r.t. `patient_id` and`image_num` cuz if one patient has given samples twice then we might get duplicate values ","metadata":{}},{"cell_type":"code","source":"train_df[train_df[[\"patient_id\", \"image_num\"]].duplicated()]","metadata":{"execution":{"iopub.status.busy":"2022-07-22T11:24:55.751003Z","iopub.execute_input":"2022-07-22T11:24:55.751384Z","iopub.status.idle":"2022-07-22T11:24:55.764332Z","shell.execute_reply.started":"2022-07-22T11:24:55.751355Z","shell.execute_reply":"2022-07-22T11:24:55.763273Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Good!\nI am doing this cuz I just want to make sure we don't have any duplicate entries ","metadata":{}},{"cell_type":"markdown","source":"## Image Num\n\nFor each patient, we've multiple images in `whole-slide` of images and each image has some sequential weightage thats why `image_num` is a **nominal** attribute","metadata":{}},{"cell_type":"code","source":"ax = sns.countplot(\n    x = \"image_num\", \n    data = train_df,\n    palette = \"dark:salmon_r\", \n);","metadata":{"execution":{"iopub.status.busy":"2022-07-22T11:37:00.341102Z","iopub.execute_input":"2022-07-22T11:37:00.342231Z","iopub.status.idle":"2022-07-22T11:37:00.797649Z","shell.execute_reply.started":"2022-07-22T11:37:00.34219Z","shell.execute_reply":"2022-07-22T11:37:00.796484Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"so rarely we've `image_num` as 2,3 and 4\n\n**This we've to keep it in mind since while training model we will need our data to be consistent**","metadata":{}},{"cell_type":"markdown","source":"## Label","metadata":{}},{"cell_type":"markdown","source":"removing `image_num` other than 0 to avoid duplicate counts of each label","metadata":{}},{"cell_type":"code","source":"train_df_img_num_zero = train_df[train_df[\"image_num\"] == 0]","metadata":{"execution":{"iopub.status.busy":"2022-07-22T11:43:00.70466Z","iopub.execute_input":"2022-07-22T11:43:00.705127Z","iopub.status.idle":"2022-07-22T11:43:00.712016Z","shell.execute_reply.started":"2022-07-22T11:43:00.705088Z","shell.execute_reply":"2022-07-22T11:43:00.711005Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ax = sns.countplot(\n    x = \"label\", \n    data = train_df_img_num_zero,\n    palette = \"dark:salmon_r\", \n);\n\nfor p in ax.patches:\n    ax.annotate(f\"{p.get_height()/train_df.shape[0]*100:.2f}%\", (p.get_x()+0.15, p.get_height()+1))","metadata":{"execution":{"iopub.status.busy":"2022-07-22T11:43:15.497566Z","iopub.execute_input":"2022-07-22T11:43:15.497976Z","iopub.status.idle":"2022-07-22T11:43:15.651701Z","shell.execute_reply.started":"2022-07-22T11:43:15.49794Z","shell.execute_reply":"2022-07-22T11:43:15.650839Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Looks like there's an **imbalance** in our dataset\n\nThus, in later stage we need to do something while training our model(e.g. stratified k fold) or while preparing the dataset(e.g. undersampling)","metadata":{}},{"cell_type":"markdown","source":"# Analyzing train directory 🔬","metadata":{}},{"cell_type":"code","source":"!ls ../input/mayo-clinic-strip-ai/train/ | head -20","metadata":{"execution":{"iopub.status.busy":"2022-07-22T14:24:36.908445Z","iopub.execute_input":"2022-07-22T14:24:36.909018Z","iopub.status.idle":"2022-07-22T14:24:37.687094Z","shell.execute_reply.started":"2022-07-22T14:24:36.908967Z","shell.execute_reply":"2022-07-22T14:24:37.685251Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As mentioned before, we've images with filename as `image_id`i.e. `{patient_id}_{image_num}`\n\nThere's one problem tho...\n\n<a href=\"https://ibb.co/mG6JvCw\"><img src=\"https://i.ibb.co/L8nSQzW/image.png\" alt=\"image\" border=\"0\"></a>\n\n\n![sad memory](https://media.makeameme.org/created/when-you-remember-7d904e36a1.jpg)\n\n`whole-slide` images are high resolution images having memory in GBs which is just nearly impossible to process in our kaggle kernels directly but I found this [dataset](https://www.kaggle.com/datasets/yasufuminakama/mayo-train-images-size1024-n16?select=train_images_8) created by [@yasufuminakama](https://www.kaggle.com/yasufuminakama) where, using a `title` method shown in this [kernel](https://www.kaggle.com/code/analokamus/a-fast-tile-generation) by [@analokamus](https://www.kaggle.com/analokamus) GBs of .Tiff images or \"slides\" were converted into 16 patches of size `1024x1024`\n\nlet's start with a sample","metadata":{}},{"cell_type":"code","source":"def plot_tile(image_id, image_dir, tile_num, axis):\n    tile = cv2.imread(f\"../input/mayo-train-images-size1024-n16/{image_dir}/{image_id}_{tile_num}.jpg\")\n    tile = cv2.cvtColor(tile, cv2.COLOR_BGR2RGB)\n    axis.imshow(tile);\n    axis.axis('off')","metadata":{"execution":{"iopub.status.busy":"2022-07-22T18:29:45.703729Z","iopub.execute_input":"2022-07-22T18:29:45.704847Z","iopub.status.idle":"2022-07-22T18:29:45.711354Z","shell.execute_reply.started":"2022-07-22T18:29:45.704805Z","shell.execute_reply":"2022-07-22T18:29:45.710092Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"image_id = \"42529f_0\"\nimage_dir = \"train_images_3\"\n\nfig, axes = plt.subplots(4,4, figsize=(10,10))\naxes = axes.flatten()\n\nfor i, ax in enumerate(axes) :\n    plot_tile(image_id, image_dir, tile_num=i, axis=ax)","metadata":{"execution":{"iopub.status.busy":"2022-07-22T18:29:45.712733Z","iopub.execute_input":"2022-07-22T18:29:45.713112Z","iopub.status.idle":"2022-07-22T18:29:48.94078Z","shell.execute_reply.started":"2022-07-22T18:29:45.713079Z","shell.execute_reply":"2022-07-22T18:29:48.939692Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Refernces 🧾","metadata":{}},{"cell_type":"markdown","source":"Big thanks to [@yasufuminakama](https://www.kaggle.com/yasufuminakama) for providing instructions to create torch `Dataset` for processed dataset in the comment section of his [notebook](https://www.kaggle.com/code/yasufuminakama/mayo-train-images-size-1024-n-16-1/comments)\n\n```python\ntrain_df['image_dir'] = ''\ntrain_df.loc[:100,'image_dir'] = '../input/mayo-train-images-size1024-n16/train_images_1/'\ntrain_df.loc[100:200,'image_dir'] = '../input/mayo-train-images-size1024-n16/train_images_2/'\ntrain_df.loc[200:300,'image_dir'] = '../input/mayo-train-images-size1024-n16/train_images_3/'\ntrain_df.loc[300:400,'image_dir'] = '../input/mayo-train-images-size1024-n16/train_images_4/'\ntrain_df.loc[400:500,'image_dir'] = '../input/mayo-train-images-size1024-n16/train_images_5/'\ntrain_df.loc[500:600,'image_dir'] = '../input/mayo-train-images-size1024-n16/train_images_6/'\ntrain_df.loc[600:700,'image_dir'] = '../input/mayo-train-images-size1024-n16/train_images_7/'\ntrain_df.loc[700:,'image_dir'] = '../input/mayo-train-images-size1024-n16/train_images_8/'\n\nclass TrainDataset(Dataset):\n    def __init__(self, cfg, df, transform=None):\n        self.cfg = cfg\n        self.image_ids = df['image_id'].values\n        self.image_dirs = df['image_dir'].values\n        self.labels = df['target'].values\n        self.transform = transform\n\n    def __len__(self):\n        return len(self.image_ids)\n\n    def __getitem__(self, idx):\n        image_id = self.image_ids[idx]\n        image_dir = self.image_dirs[idx]\n        images = []\n        for i in range(self.cfg.n_images):\n            path = image_dir + image_id + f'_{i}.jpg'\n            image = cv2.imread(path)\n            image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\n            image = self.transform(image=image)[\"image\"]\n            images.append(image)\n        images = torch.stack(images, dim=0)\n        label = torch.tensor(self.labels[idx]).long()\n        return images, label\n    \nclass TestDataset(Dataset):\n    def __init__(self, cfg, df, transform=None):\n        self.cfg = cfg\n        self.crop_size = 1024\n        self.max_size = 20000 \n        self.file_paths = df['file_path'].values\n        self.transform = transform\n\n    def __len__(self):\n        return len(self.file_paths)\n\n    def __getitem__(self, idx):\n        file_path = self.file_paths[idx]\n        image = pyvips.Image.thumbnail(file_path, self.max_size)\n        image = vips2numpy(image) # you can find this in authors dataset overview page\n        images = []\n        for img in tile(image, sz=self.crop_size, N=self.cfg.n_images):\n            img = cv2.cvtColor(img, cv2.COLOR_RGB2BGR)\n            img = self.transform(image=img)[\"image\"]\n            images.append(img)\n        images = torch.stack(images, dim=0)\n        return images\n```\n\n\nWe need to install `pyvips` for creating `TestDataset`, refer this [notebook](https://www.kaggle.com/code/analokamus/how-to-use-pyvips-offline) for installation in kaggle kernels","metadata":{}},{"cell_type":"markdown","source":"In the discussion forum, I found one intersting discussion about handelling images taken from different centers\n> The slide preparation processes may differ across sites. Hence, you may want to look into stain normalization methodologies (I believe this was covered in other threads as well)   \n-Barbos (Competition Host)\n\n\nI was able to find this package [torchstain](https://github.com/EIDOSLAB/torchstain) which can be used for doing the same\n","metadata":{}},{"cell_type":"markdown","source":"Regarding which models we can build,   \nI found this [discussion](https://www.kaggle.com/competitions/mayo-clinic-strip-ai/discussion/337651) which suggests using `Multiple Instance Learning`\n\n![multiple instance learning](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1973217%2F1005d61e62b31ba4f5bac52a354fae79%2Fkaggle-mayo.001.png?generation=1658022399002616&alt=media)\n\nYou can also find implementation of the same in this [notebook](https://www.kaggle.com/code/analokamus/a-sample-of-multi-instance-learning-model/notebook)","metadata":{}},{"cell_type":"markdown","source":"That's pretty much it ! \n\n![bie bie](https://c.tenor.com/1EaGqSpMblYAAAAM/bye-okay.gif)","metadata":{}}]}