{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"**What are you trying to do in this notebook?**\n\nTo decrease the chances of subsequent strokes, With the help of Mayo Clinic Neurovascular Research Laboratory it encourages me to improve artificial intelligence-based etiology classification so that physicians can be better equipped to prescribe the correct treatment. New computational and artificial intelligence approaches could help save the lives of stroke survivors and help me better understand the world's second-leading cause of death.\n\nMy task is to classify the etiology (CE or LAA) of the slides in the test set for each patient. The slides comprising the training and test sets depict clots with an etiology (that is, origin) known to be either CE (Cardioembolic) or LAA (Large Artery Atherosclerosis). I include a set of supplemental slides with a either an unknown etiology or an etiology other than CE or LAA.\n\n**Why are you trying it?**\n\nIn this competition, I'm working on to classify the blood clot origins in ischemic stroke. Using whole slide digital pathology images, I'll build a model that differentiates between the two major acute ischemic stroke (AIS) etiology subtypes: cardiac and large artery atherosclerosis.\n\nMy work will enable healthcare providers to better identify the origins of blood clots in deadly strokes, making it easier for physicians to prescribe the best post-stroke therapeutic management and reducing the likelihood of a second stroke.\n\n**File and Data Field Descriptions** :-\n\n**train/** - A folder containing images in the TIFF format to be used as training data.\n\n**test/** - A folder containing images to be used as test data. The actual test data comprises about 280 images.\n\n**other/** - A supplemental set of images with a either an unknown etiology or an etiology other than CE or LAA. train.csv Contains annotations for images in the train/ folder.\n\n**image_id** - A unique identifier for this instance having the form {patientid}{image_num}. Corresponds to the image {image_id}.tif.\n\n**center_id** - Identifies the medical center where the slide was obtained.\n\n**patient_id** - Identifies the patient from whom the slide was obtained.\n\n**image_num** - Enumerates images of clots obtained from the same patient.\n\n**label** - The etiology of the clot, either CE or LAA. This field is the classification target.\n\n**test.csv** - Annotations for images in the test/ folder. Has the same fields as train.csv excluding label.\n\n**other.csv** - Annotations for images in the other/ folder. Has the same fields as train.csv. The center_id is unavailable for these images however.\n\n**label** - The etiology of the clot, either Unknown or Other.\n\n**other_specified** - The specific etiology, when known, in case the etiology is labeled as Other.\n\n**sample_submission.csv** - A sample submission file in the correct format. See the Evaluation page for more details. Note in particular that you should make one prediction per patient_id, not per image_id.","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"execution":{"iopub.status.busy":"2022-09-29T08:02:27.905794Z","iopub.execute_input":"2022-09-29T08:02:27.906705Z","iopub.status.idle":"2022-09-29T08:02:28.500535Z","shell.execute_reply.started":"2022-09-29T08:02:27.906555Z","shell.execute_reply":"2022-09-29T08:02:28.499232Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os\nimport gc\nimport cv2\nimport copy\nimport time\nimport random\nimport string\nimport joblib\nimport tifffile\nimport numpy as np \nimport pandas as pd \nimport torch","metadata":{"execution":{"iopub.status.busy":"2022-09-29T08:02:28.503487Z","iopub.execute_input":"2022-09-29T08:02:28.503996Z","iopub.status.idle":"2022-09-29T08:02:30.99193Z","shell.execute_reply.started":"2022-09-29T08:02:28.503937Z","shell.execute_reply":"2022-09-29T08:02:30.990589Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from torch import nn\nimport seaborn as sns\nfrom torchvision import models\nimport matplotlib.pyplot as plt\nfrom torch.utils.data import Dataset, DataLoader\nfrom sklearn.model_selection import train_test_split\nfrom tqdm.notebook import tqdm\nfrom torch.optim import lr_scheduler\nimport warnings\nwarnings.filterwarnings(\"ignore\")\ngc.enable()","metadata":{"execution":{"iopub.status.busy":"2022-09-29T08:02:30.993349Z","iopub.execute_input":"2022-09-29T08:02:30.994293Z","iopub.status.idle":"2022-09-29T08:02:32.387289Z","shell.execute_reply.started":"2022-09-29T08:02:30.994245Z","shell.execute_reply":"2022-09-29T08:02:32.385894Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"debug = False\ngenerate_new = True\ntest_df = pd.read_csv(\"../input/mayo-clinic-strip-ai/test.csv\")\nif(test_df.shape[0] == 4):\n    test_df = pd.concat([test_df for i in range(25)])\ndirs = [\"../input/mayo-clinic-strip-ai/train/\", \"../input/mayo-clinic-strip-ai/test/\"]","metadata":{"execution":{"iopub.status.busy":"2022-09-29T08:02:32.388909Z","iopub.execute_input":"2022-09-29T08:02:32.389385Z","iopub.status.idle":"2022-09-29T08:02:32.420107Z","shell.execute_reply.started":"2022-09-29T08:02:32.389325Z","shell.execute_reply":"2022-09-29T08:02:32.418834Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df","metadata":{"execution":{"iopub.status.busy":"2022-09-29T08:02:32.42328Z","iopub.execute_input":"2022-09-29T08:02:32.424239Z","iopub.status.idle":"2022-09-29T08:02:32.452697Z","shell.execute_reply.started":"2022-09-29T08:02:32.424188Z","shell.execute_reply":"2022-09-29T08:02:32.451211Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"try:\n    os.mkdir(\"../test/\")\nexcept:\n    pass\nfor i in tqdm(range(test_df.shape[0])):\n    img_id = test_df.iloc[i].image_id\n    try:\n        sz = os.path.getsize(dirs[1] + img_id + \".tif\")\n    except:\n        sz = 1000000000\n    if(sz > 8e8):\n        img = np.zeros((512,512,3), np.uint8)\n    else:\n        try:\n            img = cv2.resize(tifffile.imread(dirs[1] + img_id + \".tif\"), (512, 512))\n        except:\n            img = np.zeros((512,512,3), np.uint8)\n    cv2.imwrite(f\"../test/{img_id}.jpg\", img)\n    del img\n    gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-09-29T08:02:32.454934Z","iopub.execute_input":"2022-09-29T08:02:32.455725Z","iopub.status.idle":"2022-09-29T08:16:07.550391Z","shell.execute_reply.started":"2022-09-29T08:02:32.455658Z","shell.execute_reply":"2022-09-29T08:16:07.548952Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class ImgDataset(Dataset):\n    def __init__(self, df):\n        self.df = df \n        self.train = 'label' in df.columns\n    def __len__(self):\n        return len(self.df)\n    \n    def __getitem__(self, index):\n        if(generate_new):\n            paths = [\"../test/\", \"../train/\"]\n        else:\n            paths = [\"../input/jpg-images-strip-ai/test/\", \"../input/jpg-images-strip-ai/train/\"]\n        try:\n            image = cv2.imread(paths[self.train] + self.df.iloc[index].image_id + \".jpg\")\n        except:\n            image = np.zeros((512,512,3), np.uint8)\n        label = 0\n        try:\n            if len(image.shape) == 5:\n                image = image.squeeze().transpose(1, 2, 0)\n            image = cv2.resize(image, (512, 512)).transpose(2, 0, 1)\n        except:\n            image = np.zeros((3, 512, 512))\n        if(self.train):\n            label = {\"CE\" : 0, \"LAA\": 1}[self.df.iloc[index].label]\n        patient_id = self.df.iloc[index].patient_id\n        return image, label, patient_id","metadata":{"execution":{"iopub.status.busy":"2022-09-29T08:16:07.552132Z","iopub.execute_input":"2022-09-29T08:16:07.552509Z","iopub.status.idle":"2022-09-29T08:16:07.566384Z","shell.execute_reply.started":"2022-09-29T08:16:07.552476Z","shell.execute_reply":"2022-09-29T08:16:07.565036Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def predict(model, dataloader):\n    model.cuda()\n    model.eval()\n    dataloader = dataloader\n    outputs = []\n    s = nn.Softmax(dim=1)\n    ids = []\n    for item in tqdm(dataloader, leave=False):\n        patient_id = item[2][0]\n        try:\n            images = item[0].cuda().float()\n            ids.append(patient_id)\n            output = model(images)\n            outputs.append(s(output.cpu()[:,:2])[0].detach().numpy())\n        except:\n            ids.append(patient_id)\n            outputs.append(s(torch.tensor([[1, 1]]).float())[0].detach().numpy())\n    return np.array(outputs), ids","metadata":{"execution":{"iopub.status.busy":"2022-09-29T08:16:07.568399Z","iopub.execute_input":"2022-09-29T08:16:07.569109Z","iopub.status.idle":"2022-09-29T08:16:07.582075Z","shell.execute_reply.started":"2022-09-29T08:16:07.569058Z","shell.execute_reply":"2022-09-29T08:16:07.580918Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model = torch.jit.load('../input/coatnet-strip-ai-training/model.pth')\nbatch_size = 1\ntest_loader = DataLoader(\n    ImgDataset(test_df), \n    batch_size=batch_size, \n    shuffle=False, \n    num_workers=1\n)","metadata":{"execution":{"iopub.status.busy":"2022-09-29T08:16:07.583437Z","iopub.execute_input":"2022-09-29T08:16:07.5843Z","iopub.status.idle":"2022-09-29T08:16:11.759315Z","shell.execute_reply.started":"2022-09-29T08:16:07.584264Z","shell.execute_reply":"2022-09-29T08:16:11.758124Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission = pd.read_csv(\"../input/mayo-clinic-strip-ai/sample_submission.csv\")\nsubmission","metadata":{"execution":{"iopub.status.busy":"2022-09-29T08:16:11.760882Z","iopub.execute_input":"2022-09-29T08:16:11.761334Z","iopub.status.idle":"2022-09-29T08:16:11.781268Z","shell.execute_reply.started":"2022-09-29T08:16:11.76129Z","shell.execute_reply":"2022-09-29T08:16:11.780276Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.to_csv(\"submission.csv\", index = False)","metadata":{"execution":{"iopub.status.busy":"2022-09-29T08:16:11.782974Z","iopub.execute_input":"2022-09-29T08:16:11.783331Z","iopub.status.idle":"2022-09-29T08:16:11.795848Z","shell.execute_reply.started":"2022-09-29T08:16:11.783298Z","shell.execute_reply":"2022-09-29T08:16:11.794816Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Did it work?**\n\nBy Using whole slide digital pathology images, I'll build a model that differentiates between the two major acute ischemic stroke (AIS) etiology subtypes: cardiac and large artery atherosclerosis and to classify the blood clot origins in ischemic stroke.\n\nMy work will enable healthcare providers to better identify the origins of blood clots in deadly strokes, making it easier for physicians to prescribe the best post-stroke therapeutic management and reducing the likelihood of a second stroke.\n\n**What did you not understand about it?**\n\nWell, everything provides in the competition data page. I've no problem while working on it. The dataset for this competition comprises over a thousand high-resolution whole-slide digital pathology images. Each slide depicts a blood clot from a patient that had experienced an acute ischemic stroke.\n\n**I hope you find this notebook useful , Good Luck!**","metadata":{}}]}