{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":45867,"databundleVersionId":6924515,"sourceType":"competition"},{"sourceId":6774400,"sourceType":"datasetVersion","datasetId":3895136},{"sourceId":6984590,"sourceType":"datasetVersion","datasetId":4014175},{"sourceId":147716186,"sourceType":"kernelVersion"}],"dockerImageVersionId":30558,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Cancer🔬 Classification: 🔎 Exploratory Data Analysis\n\nReading list:\n- https://www.intechopen.com/chapters/58601\n- https://www.mdpi.com/2072-6694/13/7/1512","metadata":{}},{"cell_type":"markdown","source":"The phenotypically informed histotype classification remains the mainstay of ovarian carcinoma subclassification. Histotypes of ovarian epithelial neoplasms have evolved with each edition of the WHO Classification of Female Genital Tumours. The current fifth edition (2020) lists five principal histotypes: high-grade serous carcinoma (HGSC), low-grade serous carcinoma (LGSC), mucinous carcinoma (MC), endometrioid carcinoma (EC) and clear cell carcinoma (CCC). Since histotypes arise from different cells of origin, cell lineage-specific diagnostic immunohistochemical markers and histotype-specific oncogenic alterations can confirm the morphological diagnosis. [REF](https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8774015/)\n\n[![cancers-14-00416-g001.jpg](https://i.postimg.cc/3xpdQ7Xm/cancers-14-00416-g001.jpg)](https://postimg.cc/75HxSFQZ)","metadata":{}},{"cell_type":"markdown","source":"The World Health Organization (WHO) recommends dividing ovarian carcinomas into five main epithelial types: High-grade serous carcinoma (HGSC), endometrioid (EN), clear cell (CC), mucinous (MC), and low-grade serous carcinomas (LGSC). These tumors not only differ at the molecular level but in many other aspects such as the response to treatment and aggressiveness. [REF](https://www.sciencedirect.com/science/article/pii/S2153353922005466)\n\n![image](https://ars.els-cdn.com/content/image/1-s2.0-S2153353922005466-gr1.jpg)","metadata":{}},{"cell_type":"code","source":"import os, glob\nimport pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\n\nDATASET_FOLDER = \"/kaggle/input/UBC-OCEAN/\"","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-12-15T18:31:30.469793Z","iopub.execute_input":"2023-12-15T18:31:30.470181Z","iopub.status.idle":"2023-12-15T18:31:30.912058Z","shell.execute_reply.started":"2023-12-15T18:31:30.47015Z","shell.execute_reply":"2023-12-15T18:31:30.910897Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Checkout some labels","metadata":{}},{"cell_type":"code","source":"df_train = pd.read_csv(os.path.join(DATASET_FOLDER, \"train.csv\"))\n# labels = list(df_train[\"label\"].unique())\nprint(f\"Dataset/train size: {len(df_train)}\")\ndisplay(df_train.head())","metadata":{"execution":{"iopub.status.busy":"2023-12-15T18:31:30.913954Z","iopub.execute_input":"2023-12-15T18:31:30.914422Z","iopub.status.idle":"2023-12-15T18:31:30.952559Z","shell.execute_reply.started":"2023-12-15T18:31:30.914388Z","shell.execute_reply":"2023-12-15T18:31:30.95149Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"_= df_train[[\"label\"]].value_counts().plot.pie(autopct='%1.1f%%', ylabel=\"label\")","metadata":{"execution":{"iopub.status.busy":"2023-12-15T18:31:30.953745Z","iopub.execute_input":"2023-12-15T18:31:30.954092Z","iopub.status.idle":"2023-12-15T18:31:31.216339Z","shell.execute_reply.started":"2023-12-15T18:31:30.954061Z","shell.execute_reply":"2023-12-15T18:31:31.214396Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_grouped = pd.DataFrame(df_train.groupby(\"is_tma\", as_index=False)[[\"label\"]].value_counts())\ndf_pivoted = df_grouped.pivot(index='label', columns='is_tma', values='count')\ndisplay(df_pivoted.T)\n_= df_pivoted.plot(kind=\"bar\", grid=True, figsize=(5, 2))","metadata":{"execution":{"iopub.status.busy":"2023-12-15T18:31:31.221484Z","iopub.execute_input":"2023-12-15T18:31:31.222666Z","iopub.status.idle":"2023-12-15T18:31:31.563088Z","shell.execute_reply.started":"2023-12-15T18:31:31.222602Z","shell.execute_reply":"2023-12-15T18:31:31.561887Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_test = pd.read_csv(os.path.join(DATASET_FOLDER, \"test.csv\"))\n# labels = list(df_train[\"label\"].unique())\nprint(f\"Dataset/test size: {len(df_test)}\")\ndisplay(df_test.head())","metadata":{"execution":{"iopub.status.busy":"2023-12-15T18:31:31.564478Z","iopub.execute_input":"2023-12-15T18:31:31.56481Z","iopub.status.idle":"2023-12-15T18:31:31.580745Z","shell.execute_reply.started":"2023-12-15T18:31:31.56478Z","shell.execute_reply":"2023-12-15T18:31:31.579651Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## How large are the images?","metadata":{}},{"cell_type":"code","source":"fig, axes = plt.subplots(nrows=2, ncols=1, figsize=(8, 5))\nfor i, col in enumerate([\"image_height\", \"image_width\"]):\n    _= df_train[[col]].hist(ax=axes[i], bins=35)","metadata":{"execution":{"iopub.status.busy":"2023-12-15T18:31:31.582258Z","iopub.execute_input":"2023-12-15T18:31:31.582585Z","iopub.status.idle":"2023-12-15T18:31:32.203964Z","shell.execute_reply.started":"2023-12-15T18:31:31.582554Z","shell.execute_reply":"2023-12-15T18:31:32.202523Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Visialize image sizes 📏 per class","metadata":{}},{"cell_type":"code","source":"import plotly.express as px\n\ndf_sizes = df_train[[\"image_width\", \"image_height\", \"label\"]]\nfor col in (\"image_width\", \"image_height\"):\n    df_sizes[col] = [round(i / 1000) * 1000 for i in df_sizes[col]]\n\ndf_sizes = df_sizes.groupby([\"image_width\", \"image_height\", \"label\"], as_index=False).size()\n# display(df_sizes.head())\nfig = px.scatter(\n    df_sizes, x=\"image_width\", y=\"image_height\", size=\"size\", color=\"label\",\n    height=450, width=900)\nfig.update_xaxes(range=[1_000, 120_000])\nfig.update_yaxes(scaleanchor=\"x\", scaleratio=1, range=[1_000, 60_000])\nfig.show() ","metadata":{"execution":{"iopub.status.busy":"2023-12-15T18:31:32.205765Z","iopub.execute_input":"2023-12-15T18:31:32.206171Z","iopub.status.idle":"2023-12-15T18:31:35.184566Z","shell.execute_reply.started":"2023-12-15T18:31:32.206138Z","shell.execute_reply":"2023-12-15T18:31:35.183395Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Show some samples 🖼️ per class\n\nNOte that not all images has thumbnails","metadata":{}},{"cell_type":"code","source":"for lb, dfg in df_train.groupby(\"label\"):\n    fig, axes = plt.subplots(ncols=4, figsize=(16, 4))\n    for i, name in enumerate(dfg[\"image_id\"].sample(4)):\n        # print(f\"{lb}: {name}\")\n        img_path = os.path.join(DATASET_FOLDER, \"train_thumbnails\", f\"{name}_thumbnail.png\")\n        if not os.path.isfile(img_path):\n            img_path = os.path.join(DATASET_FOLDER, \"train_images\", f\"{name}.png\")\n            print(f\"Missing thumbnail for {img_path} but img exists {os.path.isfile(img_path)}\")\n            continue\n        axes[i].imshow(plt.imread(img_path))\n        axes[i].set_title(f\"label: *{lb}* for img: {name}\")\n        axes[i].set_axis_off()\n    fig.show()","metadata":{"execution":{"iopub.status.busy":"2023-12-15T18:31:35.186166Z","iopub.execute_input":"2023-12-15T18:31:35.186682Z","iopub.status.idle":"2023-12-15T18:32:12.595089Z","shell.execute_reply.started":"2023-12-15T18:31:35.186648Z","shell.execute_reply":"2023-12-15T18:32:12.594003Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ls_images = glob.glob(os.path.join(DATASET_FOLDER, \"train_images\", \"*.png\"))\nprint(f\"number of images: {len(ls_images)}\")\nls_thumbnails = glob.glob(os.path.join(DATASET_FOLDER, \"train_thumbnails\", \"*.png\"))\nprint(f\"number of thumbnails: {len(ls_thumbnails)}\")","metadata":{"execution":{"iopub.status.busy":"2023-12-15T18:32:12.59673Z","iopub.execute_input":"2023-12-15T18:32:12.597143Z","iopub.status.idle":"2023-12-15T18:32:12.707531Z","shell.execute_reply.started":"2023-12-15T18:32:12.59711Z","shell.execute_reply":"2023-12-15T18:32:12.706227Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Prune uniform backround","metadata":{}},{"cell_type":"code","source":"def prune_image_rows_cols(img, thr=0.001):\n    # delete empty columns\n    for l in reversed(range(img.shape[1])):\n        if (np.sum(img[:, l]) / float(img.shape[0])) < thr:\n            img = np.delete(img, l, 1)\n    # delete empty rows\n    for l in reversed(range(img.shape[0])):\n        if (np.sum(img[l, :]) / float(img.shape[1])) < thr:\n            img = np.delete(img, l, 0)\n    return img","metadata":{"execution":{"iopub.status.busy":"2023-12-15T18:32:12.7113Z","iopub.execute_input":"2023-12-15T18:32:12.711955Z","iopub.status.idle":"2023-12-15T18:32:12.71879Z","shell.execute_reply.started":"2023-12-15T18:32:12.7119Z","shell.execute_reply":"2023-12-15T18:32:12.717924Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i, spl in df_train.sample(5).iterrows():\n    img_path = os.path.join(DATASET_FOLDER, \"train_thumbnails\", f\"{spl['image_id']}_thumbnail.png\")\n    if not os.path.isfile(img_path):\n        continue\n    img = plt.imread(img_path)\n    if img.shape[0] > img.shape[1]:\n        img = np.rollaxis(img, 0, 2)\n    img = prune_image_rows_cols(img)\n    \n    fig, axes = plt.subplots(nrows=1, ncols=2, figsize=(8, 4))\n    axes[0].imshow(img)\n    axes[0].set_title(f'{spl[\"label\"]}: {spl[\"image_id\"]}')\n    axes[1].imshow(img)\n    axes[1].set_title(\"pruned / cropped\")\n    fig.show()","metadata":{"execution":{"iopub.status.busy":"2023-12-15T18:32:12.720214Z","iopub.execute_input":"2023-12-15T18:32:12.720915Z","iopub.status.idle":"2023-12-15T18:32:54.887428Z","shell.execute_reply.started":"2023-12-15T18:32:12.720875Z","shell.execute_reply":"2023-12-15T18:32:54.886143Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Loading whole slices + prune\n\nThere is also shared datasets:\n- **[UBC-OCEAN: Tiles🖽 of Cancer🔬 2048px|scale*0.25](https://www.kaggle.com/datasets/jirkaborovec/tiles-of-cancer-2048px-scale-0-25)**\n- **[UBC-OCEAN: Tiles🖽 w/ masks🔬 2048px | scale 0.25](https://www.kaggle.com/datasets/jirkaborovec/ubc-ocean-tiles-w-masks-2048px-scale-0-25)** ","metadata":{}},{"cell_type":"code","source":"!ls /kaggle/input/pyvips-python-and-deb-package\n# intall the deb packages\n!dpkg -i --force-depends /kaggle/input/pyvips-python-and-deb-package/linux_packages/archives/*.deb\n# install the python wrapper\n!pip install pyvips -f /kaggle/input/pyvips-python-and-deb-package/python_packages/ --no-index","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2023-12-15T18:32:54.889154Z","iopub.execute_input":"2023-12-15T18:32:54.889506Z","iopub.status.idle":"2023-12-15T18:34:11.685242Z","shell.execute_reply.started":"2023-12-15T18:32:54.889474Z","shell.execute_reply":"2023-12-15T18:34:11.683749Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pyvips\nfrom tqdm.auto import tqdm\nfrom joblib import Parallel, delayed\n\nos.environ['VIPS_CONCURRENCY'] = '4'\n# set VIPS_DISC_THRESHOLD to a small number to prefer the use of disc, set it to a large number to prefer RAM\nos.environ['VIPS_DISC_THRESHOLD'] = '15gb'","metadata":{"execution":{"iopub.status.busy":"2023-12-15T18:34:11.687322Z","iopub.execute_input":"2023-12-15T18:34:11.687711Z","iopub.status.idle":"2023-12-15T18:34:12.445898Z","shell.execute_reply.started":"2023-12-15T18:34:11.687676Z","shell.execute_reply":"2023-12-15T18:34:12.444746Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def thumbnail_prune_scale(img_path: str, out_dir: str = \"train_thumbnails\") -> None:\n    name, _ = os.path.splitext(os.path.basename(img_path))\n    img = pyvips.Image.thumbnail(img_path, 3000).numpy()[..., :3]\n    plt.imsave(os.path.join(out_dir, f\"{name}_thumbnail.png\"), img)\n    img = prune_image_rows_cols(img)\n    plt.imsave(os.path.join(out_dir, f\"{name}_crop.png\"), img)\n\n\n!mkdir -p train_thumbnails\n\nls_imgs = sorted(glob.glob(\"/kaggle/input/UBC-OCEAN/train_images/*.png\"))\n\n# for p_img in tqdm(ls_imgs):\n#     thumbnail_prune_scale(p_img)\n    \n_= Parallel(n_jobs=4)(\n    delayed(thumbnail_prune_scale) (p_img) for p_img in tqdm(ls_imgs)\n)","metadata":{"execution":{"iopub.status.busy":"2023-12-15T18:34:12.447395Z","iopub.execute_input":"2023-12-15T18:34:12.447706Z","iopub.status.idle":"2023-12-15T19:14:29.166484Z","shell.execute_reply.started":"2023-12-15T18:34:12.447674Z","shell.execute_reply":"2023-12-15T19:14:29.163555Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ls_imgs = glob.glob(\"train_thumbnails/*_crop.png\")\n\nfig, axes = plt.subplots(nrows=4, figsize=(5, 18))\nfor i, img_path in enumerate(ls_imgs[:4]):\n    axes[i].imshow(plt.imread(img_path))\n    axes[i].set_axis_off()","metadata":{"execution":{"iopub.status.busy":"2023-12-15T19:14:36.062204Z","iopub.execute_input":"2023-12-15T19:14:36.062702Z","iopub.status.idle":"2023-12-15T19:14:41.958331Z","shell.execute_reply.started":"2023-12-15T19:14:36.062665Z","shell.execute_reply":"2023-12-15T19:14:41.956801Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Supplement with segm masks\n\nadjusted version for segmentations, see:\n- https://www.kaggle.com/competitions/UBC-OCEAN/discussion/455890\n- https://www.kaggle.com/datasets/sohier/ubc-ovarian-cancer-competition-supplemental-masks\n\nMask info - The masks use the following color codes:\n- Red: Tumor\n- Green: Stroma (healthy tissue)\n- Blue: Necrosis (dead or dying non-cancerous tissue)","metadata":{}},{"cell_type":"code","source":"ls_masks = sorted(glob.glob(\"/kaggle/input/ubc-ovarian-cancer-competition-supplemental-masks/*.png\"))\nprint(f\"found masks: {len(ls_masks)}\")","metadata":{"execution":{"iopub.status.busy":"2023-12-15T19:14:41.960881Z","iopub.execute_input":"2023-12-15T19:14:41.961322Z","iopub.status.idle":"2023-12-15T19:14:41.982387Z","shell.execute_reply.started":"2023-12-15T19:14:41.961278Z","shell.execute_reply":"2023-12-15T19:14:41.980938Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from PIL import Image\n\ndef thumbnail_scale(img_path: str, out_dir: str = \"train_masks\") -> None:\n    # RGB thumbnail\n    img = pyvips.Image.thumbnail(img_path, 3000).numpy()[..., :3]\n    name, _ = os.path.splitext(os.path.basename(img_path))\n    plt.imsave(os.path.join(out_dir, f\"{name}_mask.png\"), img)\n    # Conver RGB annotation to labels\n    bg = np.ones((img.shape[0], img.shape[1], 1)) * 128\n    stack = np.concatenate((bg, img), axis=2)\n    seg = np.argmax(stack, axis=2).astype(np.uint8)\n    Image.fromarray(seg).save(os.path.join(out_dir, f\"{name}_segm.png\")) \n\n\n!mkdir -p train_masks\n\n# for p_img in tqdm(ls_imgs):\n#     thumbnail_prune_scale(p_img)\n    \n_= Parallel(n_jobs=4)(\n    delayed(thumbnail_scale) (p_img) for p_img in tqdm(ls_masks)\n)","metadata":{"execution":{"iopub.status.busy":"2023-12-15T19:20:15.852197Z","iopub.execute_input":"2023-12-15T19:20:15.852683Z","iopub.status.idle":"2023-12-15T19:25:04.122778Z","shell.execute_reply.started":"2023-12-15T19:20:15.852649Z","shell.execute_reply":"2023-12-15T19:25:04.119419Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from skimage.color import label2rgb\n\nls_tnbs = glob.glob(\"train_thumbnails/*_thumbnail.png\")\nls_segs = glob.glob(\"train_masks/*.png\")\nnames = [os.path.basename(p).split(\"_\")[0] for p in ls_segs]\nnames = [n for n in names if any(f\"{n}_\" in p for p in ls_tnbs)]\n# print(names)\n\nfor i, idx in enumerate(names[:4]):\n    fig, axes = plt.subplots(ncols=3, figsize=(12, 5))\n    p_img = os.path.join(\"train_thumbnails\", f\"{idx}_thumbnail.png\")\n    p_mask = os.path.join(\"train_masks\", f\"{idx}_mask.png\")\n    p_segm = os.path.join(\"train_masks\", f\"{idx}_segm.png\")\n    img = plt.imread(p_img)[..., :3]\n    axes[0].imshow(img)\n    axes[0].set_axis_off()\n    axes[1].imshow(plt.imread(p_mask))\n    axes[1].set_axis_off()\n    seg = np.array(Image.open(p_segm))\n    try:\n        axes[2].imshow(label2rgb(\n            seg, img, colors=[(255,0,0), (0,0,255)],\n            alpha=0.01, bg_label=0, bg_color=None\n        ))\n    except:\n        print(f\"failing with image {img.shape} and segm {seg.shape}\")\n    axes[2].set_axis_off()","metadata":{"execution":{"iopub.status.busy":"2023-12-15T20:14:21.878898Z","iopub.execute_input":"2023-12-15T20:14:21.879354Z","iopub.status.idle":"2023-12-15T20:14:41.317489Z","shell.execute_reply.started":"2023-12-15T20:14:21.879319Z","shell.execute_reply":"2023-12-15T20:14:41.316054Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### prepared dataset\n\n**The tiled version of sumplementatry segmentation is shared in [UBC-OCEAN: Tiles🖽 w/ masks🔬 2048px | scale 0.25](https://www.kaggle.com/datasets/jirkaborovec/ubc-ocean-tiles-w-masks-2048px-scale-0-25)**\n\nwhich was generated by [Cancer🔬Subtype: cut 🖽 WSI -> tiles+mask | 0.25x](https://www.kaggle.com/code/jirkaborovec/cancer-subtype-cut-wsi-tiles-mask-0-25x/output)**","metadata":{}},{"cell_type":"code","source":"!cp /kaggle/input/UBC-OCEAN/train.csv .","metadata":{"execution":{"iopub.status.busy":"2023-12-15T19:25:04.126614Z","iopub.status.idle":"2023-12-15T19:25:04.12702Z","shell.execute_reply.started":"2023-12-15T19:25:04.126808Z","shell.execute_reply":"2023-12-15T19:25:04.126845Z"},"trusted":true},"execution_count":null,"outputs":[]}]}