{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Mayo Clinic - STRIP AI - Exploratory data analysis and Image processing with Pillow (PIL)\n\n## 1. Understand the train data\n\n### 1.1. descriptions of the train data\n\n- data file: train.csv\n- fields:\n    - features\n        -  image_id - A unique identifier for this instance having the form {patient_id}_{image_num}. Corresponds to the image {image_id}.tif.\n        -  center_id - Identifies the medical center where the slide was obtained.\n        -  patient_id - Identifies the patient from whom the slide was obtained.\n        -  image_num - Enumerates images of clots obtained from the same patient.\n    - target\n        -  label - The etiology of the clot, either CE or LAA. This field is the classification target.\n        \n### 1.2. summary of the train data\n\n- train data has 754 samples for 632 unique patients:\n    - the majority have only 1 image\n    - some patients have as many as 5 images thus 5 samples in the training data\n    - despite a patient may have more than 1 image, one patient has only one category of etiology (CE or LAA)  and only one center_id\n    \n- there are total of 11 centers\n    - most samples are from center 11: 257 out of 754\n    - center 4 has the 2nd most samples: 114 out of 754\n\n- label\n    - 72.5% samples are `CE` category; 27.5% are `LAA`\n    - 457 patients are `CE` category, accounts for 72% of all patients.   \n    - similar to the overall distribution, the majority are `CE` category per center_id except center_id = 3\n    - there are about equal number of samples in `CE` and `LAA`  categories   \n    \n### 1.3. understand the images\n\n- files sizes:\n\n    - most files are less than 500MB; however, there are a few files more than 2GB\n    - file sizes do not differ much between the 2 categories of clot `CE` and `LAA`\n    - large files are mostly from center 11\n    \n- images:\n    - explore if images for the same patient can be vastly different\n    - explore if images for different etiology of the clots look very different\n        - based on images of two patients, 2 different type of clots look different\n\n\n## 2. A bit exploration of the other data\n\n- other.csv - Annotations for images in the other/ folder. \n    - Has the same fields as train.csv. \n    - The center_id is unavailable for these images however.\n    - label - The etiology of the clot, either Unknown or Other.\n    - other_specified - The specific etiology, when known, in case the etiology is labeled as Other.\n\n\n## 3. Image processing with Pillow (PIL)\n\n- the `Pillow` package (`from PIL import Image`) offers to resize images in 2 ways: resize and thumbnail\n    - resize images: \n        - [https://pillow.readthedocs.io/en/stable/reference/Image.html?highlight=resize#PIL.Image.Image.resize](https://pillow.readthedocs.io/en/stable/reference/Image.html?highlight=resize#PIL.Image.Image.resize)\n        - when using resize, you need to calculate the original image height-width ratio and make sure the resized image retains the same ratio. \n    - Create thumbnails: \n        - [https://pillow.readthedocs.io/en/stable/reference/Image.html?highlight=thumbnail#create-thumbnails](https://pillow.readthedocs.io/en/stable/reference/Image.html?highlight=thumbnail#create-thumbnails)\n        - using thumbnail does not have to deal with the hassle of keeping original image ratio. \n    - addtional examples and explanations: [stackoverflow: How do I resize an image using PIL and maintain its aspect ratio?](https://stackoverflow.com/questions/273946/how-do-i-resize-an-image-using-pil-and-maintain-its-aspect-ratio)\n\n- Both thumnail and resize require defining the 'resample' filter. \n    - [https://pillow.readthedocs.io/en/stable/handbook/concepts.html#concept-filters](https://pillow.readthedocs.io/en/stable/handbook/concepts.html#concept-filters)\n    - for best quality: choose resampling filter `PIL.Image.LANCZOS`\n    - for fastest resizing: choose resampling filter `PIL.Image.NEAREST`\n- Addtional notes\n    - Rotate image: when the height of the image is much larger than the width of the image, it may be worthwhile to rotate the image in 90 degrees.\n        - the rotate method: `transpose`(https://pillow.readthedocs.io/en/stable/reference/Image.html?highlight=resize#PIL.Image.Image.resize)\n        - **!!! need to assign the transposed image to a new object to make the `transpose` work.**\n        - use `transpose(PIL.Image.Transpose.ROTATE_90)` to rotate image in 90 degrees\n    - close image to release memory\n        - use **`close()`** to destroy the image object and release memory: e.g. `img.close()`\n        - for the image object, using `del img` and `gc.collect()` to recycle memory do not help much in releasing the memory\n    - image attributes\n        - `img.size` returns *(width, height)* of the image, not *(height, width)*.\n        -  to get the height or width, use `img.height` and `img.width`\n    - max pixels to display:\n        - set `Image.MAX_IMAGE_PIXELS = None`  to disabled the upper limit of pixels to display.","metadata":{}},{"cell_type":"markdown","source":"#### Load packages","metadata":{}},{"cell_type":"code","source":"#basic libs\n\nimport pandas as pd\nimport numpy as np\nimport os\nfrom pathlib import Path\n\nfrom datetime import datetime, timedelta\nimport time\nfrom dateutil.relativedelta import relativedelta\n\nimport gc\nimport copy\n\n#additional data processing\n\nimport pyarrow.parquet as pq\nimport pyarrow as pa\n\nfrom sklearn.preprocessing import StandardScaler, MinMaxScaler\n\n\n#visualization\nimport seaborn as sns\nimport matplotlib.pyplot as plt\n\n#load images\nimport matplotlib.image as mpimg\nimport PIL\nfrom PIL import Image\n\n\n\n\n#settings\npd.options.display.max_rows = 100\npd.options.display.max_columns = 100\n\nImage.MAX_IMAGE_PIXELS = None\n\nimport warnings\nwarnings.filterwarnings(\"ignore\")\n\nimport pytorch_lightning as pl\nrandom_seed=1234\npl.seed_everything(random_seed)","metadata":{"execution":{"iopub.status.busy":"2022-07-26T08:06:53.373282Z","iopub.execute_input":"2022-07-26T08:06:53.373733Z","iopub.status.idle":"2022-07-26T08:06:58.019879Z","shell.execute_reply.started":"2022-07-26T08:06:53.373649Z","shell.execute_reply":"2022-07-26T08:06:58.018635Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os\n\nnext(os.walk('/kaggle/input/mayo-clinic-strip-ai'))","metadata":{"execution":{"iopub.status.busy":"2022-07-26T08:06:58.022374Z","iopub.execute_input":"2022-07-26T08:06:58.02362Z","iopub.status.idle":"2022-07-26T08:06:58.031104Z","shell.execute_reply.started":"2022-07-26T08:06:58.023575Z","shell.execute_reply":"2022-07-26T08:06:58.030008Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Understanding the train data","metadata":{}},{"cell_type":"code","source":"train_df = pd.read_csv('/kaggle/input/mayo-clinic-strip-ai/train.csv')\n\nprint(train_df.shape)\ntrain_df.head(2)","metadata":{"execution":{"iopub.status.busy":"2022-07-26T08:06:58.289923Z","iopub.execute_input":"2022-07-26T08:06:58.290502Z","iopub.status.idle":"2022-07-26T08:06:58.3209Z","shell.execute_reply.started":"2022-07-26T08:06:58.290474Z","shell.execute_reply":"2022-07-26T08:06:58.32013Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### patient_id\n\n\n- most patients have only one image;\n- some have as many as 5 iamges.","metadata":{}},{"cell_type":"code","source":"# check unique number of patients\nprint(train_df.shape, train_df['patient_id'].nunique())","metadata":{"execution":{"iopub.status.busy":"2022-07-26T08:07:00.215147Z","iopub.execute_input":"2022-07-26T08:07:00.215726Z","iopub.status.idle":"2022-07-26T08:07:00.225505Z","shell.execute_reply.started":"2022-07-26T08:07:00.215696Z","shell.execute_reply":"2022-07-26T08:07:00.224421Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['patient_id'].value_counts().hist(bins=10)","metadata":{"execution":{"iopub.status.busy":"2022-07-26T08:07:04.343528Z","iopub.execute_input":"2022-07-26T08:07:04.344492Z","iopub.status.idle":"2022-07-26T08:07:04.620467Z","shell.execute_reply.started":"2022-07-26T08:07:04.344456Z","shell.execute_reply":"2022-07-26T08:07:04.619172Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n##### image_num: \n\n- range from 0 to 4, indicating the 1st to the 5th image for a patient\n- nearly 84% of samples have image_num=0; less than 5% samples have image_num>=2     ","metadata":{}},{"cell_type":"code","source":"a = train_df['image_num'].value_counts().sort_index()\n\nt = pd.concat([a, 100*a/train_df.shape[0]], axis=1)\nt.columns = ['# samples', '% of samples']\nt.index.name = 'image_num'\ndisplay(t)\ndel a, t\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-07-26T08:07:07.251252Z","iopub.execute_input":"2022-07-26T08:07:07.252208Z","iopub.status.idle":"2022-07-26T08:07:07.453886Z","shell.execute_reply.started":"2022-07-26T08:07:07.252174Z","shell.execute_reply":"2022-07-26T08:07:07.452657Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['image_num'].value_counts().plot(kind='bar')","metadata":{"execution":{"iopub.status.busy":"2022-07-26T08:07:08.469311Z","iopub.execute_input":"2022-07-26T08:07:08.469992Z","iopub.status.idle":"2022-07-26T08:07:08.629014Z","shell.execute_reply.started":"2022-07-26T08:07:08.469959Z","shell.execute_reply":"2022-07-26T08:07:08.628033Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n##### center_id:\n\n- about 1/3 of samples from center 11","metadata":{}},{"cell_type":"code","source":"#the center id\n\na = train_df['center_id'].value_counts().sort_index()\n\nt = pd.concat([a, 100*a/train_df.shape[0]], axis=1)\nt.columns = ['# samples', '% of samples']\nt.index.name = 'center_id'\ndisplay(t)\ndel a, t\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-07-26T08:07:10.134415Z","iopub.execute_input":"2022-07-26T08:07:10.1348Z","iopub.status.idle":"2022-07-26T08:07:10.308258Z","shell.execute_reply.started":"2022-07-26T08:07:10.134768Z","shell.execute_reply":"2022-07-26T08:07:10.307118Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"p_center = pd.pivot_table(train_df, \n               index='patient_id', \n               columns='center_id', \n               values=['image_id'], \n               aggfunc={'image_id':[np.size]}\n              )\np_center.columns = [f'{c}' for _, _, c in p_center.columns]\np_center.isna().sum(axis=1).nunique()","metadata":{"execution":{"iopub.status.busy":"2022-07-26T08:07:11.944875Z","iopub.execute_input":"2022-07-26T08:07:11.945634Z","iopub.status.idle":"2022-07-26T08:07:11.978068Z","shell.execute_reply.started":"2022-07-26T08:07:11.945597Z","shell.execute_reply":"2022-07-26T08:07:11.97696Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"p_center.sort_values(by=['1', '2', '3', '4', '5', '6', '7', '8', '9', '10', '11'], inplace=True)\n\nplt.figure(figsize=(12,12))\n\nsns.heatmap(p_center, annot=False, cbar=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-26T08:07:12.870835Z","iopub.execute_input":"2022-07-26T08:07:12.871709Z","iopub.status.idle":"2022-07-26T08:07:13.765345Z","shell.execute_reply.started":"2022-07-26T08:07:12.87167Z","shell.execute_reply":"2022-07-26T08:07:13.76424Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### target varialbe: label\n\n- label\n    - 72.5% samples are `CE` category; 27.5% are `LAA`\n    - 457 patients are `CE` category, accounts for 72% of all patients.   \n- label v center_id\n    - similar to the overall distribution, the majority are `CE` category per center_id except center_id = 3\n    - there are about equal number of samples in `CE` and `LAA`  categories","metadata":{}},{"cell_type":"code","source":"#label\n\na = train_df['label'].value_counts().sort_index()\n\nt = pd.concat([a, 100*a/train_df.shape[0]], axis=1)\nt.columns = ['# samples', '% of samples']\nt.index.name = 'label'\ndisplay(t)\ndel a, t\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-07-26T08:07:16.234887Z","iopub.execute_input":"2022-07-26T08:07:16.235317Z","iopub.status.idle":"2022-07-26T08:07:16.414064Z","shell.execute_reply.started":"2022-07-26T08:07:16.235286Z","shell.execute_reply":"2022-07-26T08:07:16.412965Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['label'].value_counts().plot(kind='bar', figsize=(4,4))","metadata":{"execution":{"iopub.status.busy":"2022-07-26T08:07:17.51216Z","iopub.execute_input":"2022-07-26T08:07:17.512983Z","iopub.status.idle":"2022-07-26T08:07:17.622834Z","shell.execute_reply.started":"2022-07-26T08:07:17.512942Z","shell.execute_reply":"2022-07-26T08:07:17.621955Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#label\n\na = train_df[['patient_id', 'label']].drop_duplicates(keep='first')['label'].value_counts().sort_index()\n\nt = pd.concat([a, 100*a/train_df['patient_id'].nunique()], axis=1)\nt.columns = ['# unique patients', '% of unique patients']\nt.index.name = 'label'\ndisplay(t)\ndel a, t\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-07-26T08:07:18.965755Z","iopub.execute_input":"2022-07-26T08:07:18.966152Z","iopub.status.idle":"2022-07-26T08:07:19.148084Z","shell.execute_reply.started":"2022-07-26T08:07:18.966118Z","shell.execute_reply":"2022-07-26T08:07:19.147003Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"p_label = pd.pivot_table(train_df, \n               index='patient_id', \n               columns='label', \n               values=['image_id'], \n               aggfunc={'image_id':[np.size]}\n              )\np_label.columns = [f'{c}' for _, _, c in p_label.columns]\n# p_label.fillna(value = 0, inplace=True)\np_label['label_cnt'] =2-p_label.isna().sum(axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-07-26T08:07:20.832598Z","iopub.execute_input":"2022-07-26T08:07:20.833225Z","iopub.status.idle":"2022-07-26T08:07:20.857192Z","shell.execute_reply.started":"2022-07-26T08:07:20.833186Z","shell.execute_reply":"2022-07-26T08:07:20.856391Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"p_label['label_cnt'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-07-26T08:07:21.833408Z","iopub.execute_input":"2022-07-26T08:07:21.834204Z","iopub.status.idle":"2022-07-26T08:07:21.842027Z","shell.execute_reply.started":"2022-07-26T08:07:21.834166Z","shell.execute_reply":"2022-07-26T08:07:21.84097Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"p_label","metadata":{"execution":{"iopub.status.busy":"2022-07-26T08:07:22.837191Z","iopub.execute_input":"2022-07-26T08:07:22.838208Z","iopub.status.idle":"2022-07-26T08:07:22.854142Z","shell.execute_reply.started":"2022-07-26T08:07:22.838167Z","shell.execute_reply":"2022-07-26T08:07:22.853014Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"p_label['total']= p_label.sum(axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-07-26T08:07:23.833348Z","iopub.execute_input":"2022-07-26T08:07:23.834227Z","iopub.status.idle":"2022-07-26T08:07:23.840763Z","shell.execute_reply.started":"2022-07-26T08:07:23.834177Z","shell.execute_reply":"2022-07-26T08:07:23.839999Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"p_center = pd.pivot_table(train_df[['center_id', 'label', 'patient_id']].drop_duplicates(keep='first'), \n               index='center_id', \n               columns='label', \n               values=['patient_id'], \n               aggfunc={'patient_id':[np.size]}\n              )\np_center.columns = [f'{c}' for _, _, c in p_center.columns]\n","metadata":{"execution":{"iopub.status.busy":"2022-07-26T08:07:25.17555Z","iopub.execute_input":"2022-07-26T08:07:25.176246Z","iopub.status.idle":"2022-07-26T08:07:25.194034Z","shell.execute_reply.started":"2022-07-26T08:07:25.176208Z","shell.execute_reply":"2022-07-26T08:07:25.192967Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"p_center.plot(kind='bar', figsize=(12, 5), \n              title='unique num of patients by target category and center')","metadata":{"execution":{"iopub.status.busy":"2022-07-26T08:07:26.898689Z","iopub.execute_input":"2022-07-26T08:07:26.899092Z","iopub.status.idle":"2022-07-26T08:07:27.167166Z","shell.execute_reply.started":"2022-07-26T08:07:26.899059Z","shell.execute_reply":"2022-07-26T08:07:27.166087Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Explore the images\n\n- files sizes:\n\n    - most files are less than 500MB; however, there are a few files more than 2GB\n    - file sizes do not differ much between the 2 categories of clot `CE` and `LAA`\n    - large files are mostly from center 11\n    \n- image info:\n    - image width and height distribution\n    - image ratio (height/width) distribution\n    - image mode/compression/dpi\n    \n- images:\n    - explore if images for the same patient can be vastly different\n    - explore if images for different etiology of the clots look very different\n        - based on images of two patients, 2 different type of clots look different","metadata":{}},{"cell_type":"code","source":"%%time\n#train file sizes\n\ntrain_pic_folder = '/kaggle/input/mayo-clinic-strip-ai/train'\ntrain_pics = next(os.walk(train_pic_folder))[2]\n#\npic_stats = []\nfor pic in train_pics:\n    p = Path(f'{train_pic_folder}/{pic}')\n    img = Image.open(f'{train_pic_folder}/{pic}')\n    pic_stats.append([pic.split('.')[0], pic, p.stat().st_size/(1024**2), img.width, img.height, img.mode, img.info['compression'], img.info['dpi'] ])\n    img.close()\n    del img\n    gc.collect()\n    \npic_stats_df = pd.DataFrame(data = pic_stats, columns = ['image_id', 'image_name', 'size', 'width', 'height', 'mode', 'compression', 'dpi'])","metadata":{"execution":{"iopub.status.busy":"2022-07-26T08:19:51.772275Z","iopub.execute_input":"2022-07-26T08:19:51.773272Z","iopub.status.idle":"2022-07-26T08:24:01.304275Z","shell.execute_reply.started":"2022-07-26T08:19:51.773231Z","shell.execute_reply":"2022-07-26T08:24:01.302974Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pic_stats_df.head(2)","metadata":{"execution":{"iopub.status.busy":"2022-07-26T08:33:24.233257Z","iopub.execute_input":"2022-07-26T08:33:24.233696Z","iopub.status.idle":"2022-07-26T08:33:24.248119Z","shell.execute_reply.started":"2022-07-26T08:33:24.233658Z","shell.execute_reply":"2022-07-26T08:33:24.247Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pic_stats_df.sort_values(by='size', ascending=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-26T08:33:27.385888Z","iopub.execute_input":"2022-07-26T08:33:27.387047Z","iopub.status.idle":"2022-07-26T08:33:27.41289Z","shell.execute_reply.started":"2022-07-26T08:33:27.387008Z","shell.execute_reply":"2022-07-26T08:33:27.412027Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pic_stats_df['size'].plot(kind='hist', bins=50, figsize = (8, 5),\n                          title='distribution of images by file size (MB)')\n","metadata":{"execution":{"iopub.status.busy":"2022-07-26T08:33:38.764716Z","iopub.execute_input":"2022-07-26T08:33:38.765107Z","iopub.status.idle":"2022-07-26T08:33:39.033628Z","shell.execute_reply.started":"2022-07-26T08:33:38.765074Z","shell.execute_reply":"2022-07-26T08:33:39.032663Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pic_stats_df.shape, train_df.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-26T08:33:42.439236Z","iopub.execute_input":"2022-07-26T08:33:42.440238Z","iopub.status.idle":"2022-07-26T08:33:42.446422Z","shell.execute_reply.started":"2022-07-26T08:33:42.440199Z","shell.execute_reply":"2022-07-26T08:33:42.445374Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df = train_df.merge(pic_stats_df, on='image_id', how='left')\ntrain_df.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-26T08:33:43.557681Z","iopub.execute_input":"2022-07-26T08:33:43.55867Z","iopub.status.idle":"2022-07-26T08:33:43.57243Z","shell.execute_reply.started":"2022-07-26T08:33:43.558634Z","shell.execute_reply":"2022-07-26T08:33:43.571527Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.head(2)","metadata":{"execution":{"iopub.status.busy":"2022-07-26T08:33:44.844216Z","iopub.execute_input":"2022-07-26T08:33:44.84516Z","iopub.status.idle":"2022-07-26T08:33:44.861331Z","shell.execute_reply.started":"2022-07-26T08:33:44.845123Z","shell.execute_reply":"2022-07-26T08:33:44.86025Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df[['label', 'size']].groupby('label').plot(kind='hist', bins=50, figsize = (8, 5),\n                          title='distribution of images by file size (MB)')\n","metadata":{"execution":{"iopub.status.busy":"2022-07-26T08:33:45.483253Z","iopub.execute_input":"2022-07-26T08:33:45.483905Z","iopub.status.idle":"2022-07-26T08:33:46.020691Z","shell.execute_reply.started":"2022-07-26T08:33:45.48387Z","shell.execute_reply":"2022-07-26T08:33:46.01945Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['center_id'] = train_df['center_id'].astype('category')\ntrain_df[['center_id', 'size']].groupby('center_id').plot(kind='hist', bins=50, figsize = (8, 5),\n                          title='distribution of images by file size (MB)')\n\n","metadata":{"execution":{"iopub.status.busy":"2022-07-26T08:33:46.190629Z","iopub.execute_input":"2022-07-26T08:33:46.191503Z","iopub.status.idle":"2022-07-26T08:33:49.000342Z","shell.execute_reply.started":"2022-07-26T08:33:46.191459Z","shell.execute_reply":"2022-07-26T08:33:48.999368Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### image info\n\n- some images are extremely large (in width and height)\n- all images are in `RGB` mode and in `tiff_adobe_deflate` format\n- only 2 Dots per inches (dpi): (25.4, 25.4)  and (50497.015703125, 50497.015703125) \n","metadata":{}},{"cell_type":"code","source":"train_df['ratio'] = train_df['height']/train_df['width']","metadata":{"execution":{"iopub.status.busy":"2022-07-26T08:36:00.894861Z","iopub.execute_input":"2022-07-26T08:36:00.896019Z","iopub.status.idle":"2022-07-26T08:36:00.903192Z","shell.execute_reply.started":"2022-07-26T08:36:00.895964Z","shell.execute_reply":"2022-07-26T08:36:00.90181Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df[['height', 'width', 'ratio']].describe()","metadata":{"execution":{"iopub.status.busy":"2022-07-26T08:37:52.810897Z","iopub.execute_input":"2022-07-26T08:37:52.812115Z","iopub.status.idle":"2022-07-26T08:37:52.842243Z","shell.execute_reply.started":"2022-07-26T08:37:52.812063Z","shell.execute_reply":"2022-07-26T08:37:52.841008Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for c in ['height', 'width', 'ratio']:\n    train_df[[c]].plot(kind='hist', bins=50)","metadata":{"execution":{"iopub.status.busy":"2022-07-26T08:40:10.989439Z","iopub.execute_input":"2022-07-26T08:40:10.989803Z","iopub.status.idle":"2022-07-26T08:40:11.922697Z","shell.execute_reply.started":"2022-07-26T08:40:10.989776Z","shell.execute_reply":"2022-07-26T08:40:11.921623Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['mode'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-07-26T08:42:32.670915Z","iopub.execute_input":"2022-07-26T08:42:32.67142Z","iopub.status.idle":"2022-07-26T08:42:32.679417Z","shell.execute_reply.started":"2022-07-26T08:42:32.671389Z","shell.execute_reply":"2022-07-26T08:42:32.678275Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['compression'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-07-26T08:53:05.175287Z","iopub.execute_input":"2022-07-26T08:53:05.175692Z","iopub.status.idle":"2022-07-26T08:53:05.185291Z","shell.execute_reply.started":"2022-07-26T08:53:05.175661Z","shell.execute_reply":"2022-07-26T08:53:05.184179Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['dpi'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-07-26T08:53:16.108048Z","iopub.execute_input":"2022-07-26T08:53:16.108391Z","iopub.status.idle":"2022-07-26T08:53:16.117127Z","shell.execute_reply.started":"2022-07-26T08:53:16.108364Z","shell.execute_reply":"2022-07-26T08:53:16.116315Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### images\n\n- explore if images for the same patient can be vastly different\n- explore if images for different etiology of the clots look very different\n    - based on images of two patients, 2 different type of clots look different","metadata":{}},{"cell_type":"code","source":"#find a patient with at 5 small images\ntrain_df[(train_df['size']<500) & (train_df['image_num']==4)]","metadata":{"execution":{"iopub.status.busy":"2022-07-26T08:55:14.365587Z","iopub.execute_input":"2022-07-26T08:55:14.366347Z","iopub.status.idle":"2022-07-26T08:55:14.386496Z","shell.execute_reply.started":"2022-07-26T08:55:14.36631Z","shell.execute_reply":"2022-07-26T08:55:14.385121Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"patient_id='3d10be'\nlabel = 'CE'\ntrain_df[train_df['patient_id']==patient_id]","metadata":{"execution":{"iopub.status.busy":"2022-07-26T08:55:15.396678Z","iopub.execute_input":"2022-07-26T08:55:15.397069Z","iopub.status.idle":"2022-07-26T08:55:15.416507Z","shell.execute_reply.started":"2022-07-26T08:55:15.39704Z","shell.execute_reply":"2022-07-26T08:55:15.415379Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\nfig, axes = plt.subplots(nrows=1, ncols=5, figsize=(20, 6))\n\nprint(f'patient_id = {patient_id}, label={label}')\nfor i in range(5):\n    img_path = f'/kaggle/input/mayo-clinic-strip-ai/train/{patient_id}_{i}.tif'\n    img = Image.open(img_path)\n    fac = int(max(img.size)/224)\n    h, w = img.size\n    \n    axes[i].imshow(img.resize((int(h/fac), int(w/fac))))\n    axes[i].set_title(f'{patient_id}_{i}')\n    \n    img.close()\n    del img, fac\n    gc.collect()\n","metadata":{"execution":{"iopub.status.busy":"2022-07-26T08:55:16.999266Z","iopub.execute_input":"2022-07-26T08:55:16.999622Z","iopub.status.idle":"2022-07-26T08:55:43.395929Z","shell.execute_reply.started":"2022-07-26T08:55:16.999583Z","shell.execute_reply":"2022-07-26T08:55:43.395157Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"patient_id='91b9d3'\nlabel = 'LAA'\ntrain_df[train_df['patient_id']==patient_id]","metadata":{"execution":{"iopub.status.busy":"2022-07-26T08:55:57.395961Z","iopub.execute_input":"2022-07-26T08:55:57.397209Z","iopub.status.idle":"2022-07-26T08:55:57.42331Z","shell.execute_reply.started":"2022-07-26T08:55:57.397168Z","shell.execute_reply":"2022-07-26T08:55:57.422256Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\nfig, axes = plt.subplots(nrows=1, ncols=5, figsize=(20, 6))\n\nprint(f'patient_id = {patient_id}, label={label}')\nfor i in range(5):\n    img_path = f'/kaggle/input/mayo-clinic-strip-ai/train/{patient_id}_{i}.tif'\n    img = Image.open(img_path)\n    fac = int(max(img.size)/224)\n    h, w = img.size\n    \n    axes[i].imshow(img.resize((int(h/fac), int(w/fac))))\n    axes[i].set_title(f'{patient_id}_{i}')\n    img.close()\n    del img, fac\n    gc.collect()\n","metadata":{"execution":{"iopub.status.busy":"2022-07-26T08:56:12.748579Z","iopub.execute_input":"2022-07-26T08:56:12.748993Z","iopub.status.idle":"2022-07-26T08:57:30.35706Z","shell.execute_reply.started":"2022-07-26T08:56:12.748955Z","shell.execute_reply":"2022-07-26T08:57:30.35588Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfig, axes = plt.subplots(nrows=1, ncols=2, figsize=(20, 10))\npic_id = '037300_0.tif'\nimg = Image.open(f\"/kaggle/input/mayo-clinic-strip-ai/train/{pic_id}\")\nprint(img.height, img.width)\nimg.thumbnail((500, 500), resample=Image.Resampling.LANCZOS, reducing_gap=10)\nimg2 = img.transpose(PIL.Image.Transpose.ROTATE_90)\naxes[0].imshow(img)\naxes[0].set_title(f'{pic_id} - thumbnail (500, 500)')\naxes[1].imshow(img2)\naxes[1].set_title(f'{pic_id} - thumbnail (500, 500) rotate 90 degrees')\n\nimg.close()\nimg2.close()\n\ndel img, img2\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-07-26T09:12:02.521508Z","iopub.execute_input":"2022-07-26T09:12:02.521933Z","iopub.status.idle":"2022-07-26T09:12:28.130331Z","shell.execute_reply.started":"2022-07-26T09:12:02.521874Z","shell.execute_reply":"2022-07-26T09:12:28.129179Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Others","metadata":{}},{"cell_type":"code","source":"other_df = pd.read_csv('/kaggle/input/mayo-clinic-strip-ai/other.csv')\n\nprint(other_df.shape)\nother_df.head(2)","metadata":{"execution":{"iopub.status.busy":"2022-07-26T09:12:37.961809Z","iopub.execute_input":"2022-07-26T09:12:37.962222Z","iopub.status.idle":"2022-07-26T09:12:37.983416Z","shell.execute_reply.started":"2022-07-26T09:12:37.962189Z","shell.execute_reply":"2022-07-26T09:12:37.982241Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"other_df['label'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-07-26T09:12:38.755875Z","iopub.execute_input":"2022-07-26T09:12:38.756534Z","iopub.status.idle":"2022-07-26T09:12:38.765316Z","shell.execute_reply.started":"2022-07-26T09:12:38.756501Z","shell.execute_reply":"2022-07-26T09:12:38.764214Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#train file sizes\n\nother_pic_folder = '/kaggle/input/mayo-clinic-strip-ai/other'\nother_pics = next(os.walk(other_pic_folder))[2]\n#\npic_stats = []\nfor pic in other_pics:\n    p = Path(f'{other_pic_folder}/{pic}')\n    img = Image.open(f'{other_pic_folder}/{pic}')\n    pic_stats.append([pic.split('.')[0], pic, p.stat().st_size/(1024**2), img.width, img.height, img.mode, img.info['compression'], img.info['dpi'] ])\n    \n    img.close()\n    del img\n    gc.collect()\n    \nother_pic_stats_df = pd.DataFrame(data = pic_stats, columns = ['image_id', 'image_name', 'size', 'width', 'height', 'mode', 'compression', 'dpi'])","metadata":{"execution":{"iopub.status.busy":"2022-07-26T09:14:08.770396Z","iopub.execute_input":"2022-07-26T09:14:08.771432Z","iopub.status.idle":"2022-07-26T09:17:01.371914Z","shell.execute_reply.started":"2022-07-26T09:14:08.77139Z","shell.execute_reply":"2022-07-26T09:17:01.37075Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(other_df.shape, other_pic_stats_df.shape)\nother_df = other_df.merge(other_pic_stats_df, on='image_id', how='left')\nprint(other_df.shape, other_pic_stats_df.shape)\ndel other_pic_stats_df\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-07-26T09:17:01.37447Z","iopub.execute_input":"2022-07-26T09:17:01.374943Z","iopub.status.idle":"2022-07-26T09:17:01.795455Z","shell.execute_reply.started":"2022-07-26T09:17:01.374884Z","shell.execute_reply":"2022-07-26T09:17:01.794449Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"other_df['other_specified'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-07-26T09:17:01.796883Z","iopub.execute_input":"2022-07-26T09:17:01.797358Z","iopub.status.idle":"2022-07-26T09:17:01.807643Z","shell.execute_reply.started":"2022-07-26T09:17:01.797327Z","shell.execute_reply":"2022-07-26T09:17:01.806203Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"other_df[other_df['other_specified']=='Hypercoagulable']","metadata":{"execution":{"iopub.status.busy":"2022-07-26T09:17:01.810897Z","iopub.execute_input":"2022-07-26T09:17:01.811657Z","iopub.status.idle":"2022-07-26T09:17:01.837662Z","shell.execute_reply.started":"2022-07-26T09:17:01.811621Z","shell.execute_reply":"2022-07-26T09:17:01.836572Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfig, axes = plt.subplots(nrows=1, ncols=2, figsize=(20, 10))\npic_id = '54334d_0.tif'\nimg = Image.open(f\"/kaggle/input/mayo-clinic-strip-ai/other/{pic_id}\")\nprint(img.height, img.width)\nimg.thumbnail((500, 500), resample=Image.Resampling.LANCZOS, reducing_gap=10)\nimg2 = img.transpose(PIL.Image.Transpose.ROTATE_90)\naxes[0].imshow(img)\naxes[0].set_title(f'{pic_id} - thumbnail (500, 500)')\naxes[1].imshow(img2)\naxes[1].set_title(f'{pic_id} - thumbnail (500, 500) rotate 90 degrees')\n\nimg.close()\nimg2.close()\n\ndel img, img2\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-07-26T09:17:01.839012Z","iopub.execute_input":"2022-07-26T09:17:01.840107Z","iopub.status.idle":"2022-07-26T09:17:16.105547Z","shell.execute_reply.started":"2022-07-26T09:17:01.840068Z","shell.execute_reply":"2022-07-26T09:17:16.104413Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}