{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"import warnings\nimport os\nwarnings.filterwarnings(\"ignore\")\n\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_kg_hide-output":true,"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-11-12T03:14:43.451505Z","iopub.execute_input":"2023-11-12T03:14:43.451886Z","iopub.status.idle":"2023-11-12T03:14:43.66171Z","shell.execute_reply.started":"2023-11-12T03:14:43.451861Z","shell.execute_reply":"2023-11-12T03:14:43.660646Z"},"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":" ## Introduction\n This notebooks mainly focuses on EDA(Exploratory Data Analysis), and we will expore what properties Ovarian Cancer Subtype(OCS) has.\n And I hope us can get some good properties to help us to better analyse **OCS**, and clean this dataset. \n here is a list to give you a quick view on this notebook.\n 1. Inital exploratory data analysis\n    * Overview on UBC-OCEAN\n    * The amount of data item\n    * Visualize data distribution\n 2. Visulize each ovarian cancer subtypes, and figure out how to classify them\n    * In this part, my goal is to extract some features to help following feature engineering.\n    * Also, I will quote some basic medical knowledge you have to know, I believe you will get benfit from my tile visulization\n \n**TODO List:**\n 1. HGSC(Finish)\n 2. LGSC(Finish)\n 3. EC(editing)\n 4. CC(TODO)\n 5. MC(TOD)","metadata":{}},{"cell_type":"markdown","source":"## 1.Inital exploratory data analysis","metadata":{}},{"cell_type":"code","source":"import numpy as np \nimport pandas as pd \n%matplotlib inline\nimport matplotlib.pyplot as plt\n\nDATASET_FOLDER = \"/kaggle/input/UBC-OCEAN/\"\n\nTRAIN_IMAGES = os.path.join(DATASET_FOLDER, \"train_images/\")\nTRAIN_THUMBNAILS_IMAGES = os.path.join(DATASET_FOLDER, \"train_thumbnails/\")\n\nTEST_IMAGES = os.path.join(DATASET_FOLDER, \"test_images/\")\nTEST_THUMBNAILS_IMAGES = os.path.join(DATASET_FOLDER, \"test_thumbnails/\")","metadata":{"execution":{"iopub.status.busy":"2023-11-12T03:14:43.662892Z","iopub.execute_input":"2023-11-12T03:14:43.663145Z","iopub.status.idle":"2023-11-12T03:14:43.988207Z","shell.execute_reply.started":"2023-11-12T03:14:43.663124Z","shell.execute_reply":"2023-11-12T03:14:43.987472Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train = pd.read_csv(os.path.join(DATASET_FOLDER, \"train.csv\"))\ndf_train.head() # show head data","metadata":{"execution":{"iopub.status.busy":"2023-11-12T03:14:43.989772Z","iopub.execute_input":"2023-11-12T03:14:43.99031Z","iopub.status.idle":"2023-11-12T03:14:44.030834Z","shell.execute_reply.started":"2023-11-12T03:14:43.99028Z","shell.execute_reply":"2023-11-12T03:14:44.02942Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.info() # non-null is good, we have no need to clean empty items","metadata":{"execution":{"iopub.status.busy":"2023-11-12T03:14:44.034537Z","iopub.execute_input":"2023-11-12T03:14:44.034987Z","iopub.status.idle":"2023-11-12T03:14:44.065056Z","shell.execute_reply.started":"2023-11-12T03:14:44.034955Z","shell.execute_reply":"2023-11-12T03:14:44.06411Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.describe()","metadata":{"execution":{"iopub.status.busy":"2023-11-12T03:14:44.066143Z","iopub.execute_input":"2023-11-12T03:14:44.066505Z","iopub.status.idle":"2023-11-12T03:14:44.091169Z","shell.execute_reply.started":"2023-11-12T03:14:44.066478Z","shell.execute_reply":"2023-11-12T03:14:44.090058Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Analyse components\n","metadata":{}},{"cell_type":"code","source":"label_list = list(df_train['label'].unique()) # show label types\nprint(f\"label list contains {len(label_list)} label types: {label_list}\\n\")\n\n# max & min image width,image height\nimage_width_max, image_width_min = df_train['image_width'].max(), df_train['image_width'].min()\nprint(f\"max image width: {image_width_max}\\nmin image width: {image_width_min}\")\n\nimage_height_max, image_height_min = df_train['image_height'].max(), df_train['image_height'].min()\nprint(f\"max image height: {image_height_max}\\nmin image height: {image_height_min}\")\n\ndf_train['label'].value_counts().plot(kind='pie',autopct='%.2f', legend=False)\nplt.plot()","metadata":{"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-11-12T03:14:44.092404Z","iopub.execute_input":"2023-11-12T03:14:44.092692Z","iopub.status.idle":"2023-11-12T03:14:44.28816Z","shell.execute_reply.started":"2023-11-12T03:14:44.092666Z","shell.execute_reply":"2023-11-12T03:14:44.287357Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Expore TMA and WSI data\nThere are two categories of images: whole slide images (WSI) and tissue microarray (TMA). Whole slide images are at 20x magnification and can be quite large. The TMAs are smaller (roughly 4,000x4,000 pixels) but at 40x magnification.  \n\n**Pay attention to this: is_tma - True if the slide is a tissue microarray. Only available for the train set.**","metadata":{}},{"cell_type":"code","source":"!ls /kaggle/input/pyvips-python-and-deb-package\n# intall the deb packages\n!dpkg -i --force-depends /kaggle/input/pyvips-python-and-deb-package/linux_packages/archives/*.deb\n# install the python wrapper\n!pip install pyvips -f /kaggle/input/pyvips-python-and-deb-package/python_packages/ --no-index","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-11-12T03:14:44.289367Z","iopub.execute_input":"2023-11-12T03:14:44.28983Z","iopub.status.idle":"2023-11-12T03:15:56.136587Z","shell.execute_reply.started":"2023-11-12T03:14:44.289804Z","shell.execute_reply":"2023-11-12T03:15:56.135462Z"},"_kg_hide-output":true,"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import glob\nimport PIL\nimport pyvips\n\ntrain_images = glob.glob(os.path.join(TRAIN_IMAGES,'*.png'))\nprint(f\"{len(train_images)} train items found under {TRAIN_IMAGES} folder.\")\n\ntrain_thumbnails_images = glob.glob(os.path.join(TRAIN_THUMBNAILS_IMAGES,'*.png'))\nprint(f\"{len(train_thumbnails_images)} train items found under {TRAIN_THUMBNAILS_IMAGES} folder.\")","metadata":{"execution":{"iopub.status.busy":"2023-11-12T03:15:56.137839Z","iopub.execute_input":"2023-11-12T03:15:56.13812Z","iopub.status.idle":"2023-11-12T03:15:56.502073Z","shell.execute_reply.started":"2023-11-12T03:15:56.138097Z","shell.execute_reply":"2023-11-12T03:15:56.500874Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!ls /kaggle/input/pyvips-python-and-deb-package\n# intall the deb packages\n!dpkg -i --force-depends /kaggle/input/pyvips-python-and-deb-package/linux_packages/archives/*.deb\n# install the python wrapper\n!pip install pyvips -f /kaggle/input/pyvips-python-and-deb-package/python_packages/ --no-index","metadata":{"scrolled":true,"_kg_hide-output":true,"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-11-12T03:15:56.503647Z","iopub.execute_input":"2023-11-12T03:15:56.503938Z","iopub.status.idle":"2023-11-12T03:17:12.564254Z","shell.execute_reply.started":"2023-11-12T03:15:56.503913Z","shell.execute_reply":"2023-11-12T03:17:12.562632Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Now we need some basic medical and histopathology knowledge about ovarian cancer\n* HGSC(High-Grade Serous Carcinoma)\n* LGSC(Low-Grade Serous Carcinoma)\n* EC(Endometrioid Carcinoma)\n* CC(Cell Carcinoma)\n* MC(Mucinous Carcinoma)\n\n### Classification rule\n\n>These are referred to as: serous, mucinous, clear-cell and endometrioid—appellations deriving from their morphology and tissue architecture as observed through microscopy. Furthermore, the assignment of a tumour grade, based on the apparent degree of cytological aberration, allows for an additional degree of stratification for serous and endometrioid EOCs\nfrom https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6412907/\n\n","metadata":{}},{"cell_type":"markdown","source":"## HGSC(High-Grade Serous Carcinoma)\n### Basic knowledge about HGSC\nIn general, HGSC accounts for 68% of ovarian cancer and have the worst prognosis, which is often caused by human's TP53 gene mutation.\n\nNow we need to know how it looks like in morphology.\n> From the perspective of a pathologist visualizing stained tissue sections under a microscope, HGSOC tumours typically feature solid masses of cells (Figure 1A) with slit-like fenestrations (Figure 1B) [9]. In some areas, the tumours often have a papillary (Figure 1C), glandular (Figure 1B) or cribriform (Figure 1D) architecture that is said to resemble the surface epithelium of the fallopian tube [9,11]. The regions of solid growth are frequently accompanied by areas of extensive necrosis (Figure 1E) [9]. In certain cases, HGSOC may present with areas displaying a solid growth pattern that simulates the appearance of endometrioid or transitional cell carcinoma (Figure 1D) [10]. Although morphologically distinct, these tumours show an immunoreactivity identical to typical HGSOC and are thus not considered as a separate entity [10]. Researchers have recently named a group of HGSOC as the SET (“Solid, pseudo-Endometrioid and/or Transitional cell carcinoma-like”) tumours [16]. It was found that SET tumours frequently associate with BRCA1 mutations and contain a greater number of tumour-infiltrating lymphocytes (Figure 1F) compared to typical HGSOC [16].\n\nif you are interested about it and want to go deeper, I recommend you to read this paper(section 3. Histopathology): https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6412907/. It lists some HGSC histopathology images, which can help us to understand better.\n","metadata":{}},{"cell_type":"code","source":"df_HGSC = df_train[(df_train['label'] == 'HGSC') & (df_train['is_tma'] == False)]\nplt.figure(figsize=(9,3),dpi=300)\nsample_HGSC = df_HGSC.iloc[4]\nsample_HGSC_path = os.path.join(TRAIN_THUMBNAILS_IMAGES,str(sample_HGSC['image_id'])+\"_thumbnail.png\")\nprint(f\"exist HGSC sample file: {os.path.exists(sample_HGSC_path)}\")\n\nHGSC_img = pyvips.Image.new_from_file(sample_HGSC_path)\n\nplt.title(sample_HGSC['label'],x = 0.5, y = -0.2)\nplt.subplot(1,2,1)\n_ = plt.imshow(HGSC_img.numpy())\n\nsample_HGSC = df_HGSC.iloc[1]\nsample_HGSC_path = os.path.join(TRAIN_THUMBNAILS_IMAGES,str(sample_HGSC['image_id'])+\"_thumbnail.png\")\nprint(f\"exist HGSC sample file: {os.path.exists(sample_HGSC_path)}\")\n\nHGSC_img = pyvips.Image.new_from_file(sample_HGSC_path)\n\nplt.subplot(1,2,2)\nplt.title(sample_HGSC['label'],x = 0.5, y = -0.2)\n_ = plt.imshow(HGSC_img.numpy())\nprint(f\"image information: width= {HGSC_img.width}, height= {HGSC_img.height}.\")","metadata":{"execution":{"iopub.status.busy":"2023-11-12T03:17:12.568334Z","iopub.execute_input":"2023-11-12T03:17:12.568681Z","iopub.status.idle":"2023-11-12T03:17:16.70298Z","shell.execute_reply.started":"2023-11-12T03:17:12.568652Z","shell.execute_reply":"2023-11-12T03:17:16.701082Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"***We found 3 tiles from this thumbnail image, now we will crop them***","metadata":{}},{"cell_type":"code","source":"from typing import Union\n\ndef img_tile(img: Union[str, pyvips.Image], tile_size: int, threshold:int = 0):\n    if type(img) == str:\n        assert os.path.exists(img)\n        img = pyvips.Image.new_from_file(img)\n    \n    W = img.width\n    H = img.height\n    \n    # crop image into tiles\n    w_tile_num = int(W/tile_size)\n    h_tile_num = int(H/tile_size)\n    \n    crop_list = []\n    for i in range(w_tile_num):\n        for j in range(h_tile_num):\n            img_crop = img.crop(i*tile_size, j*tile_size, tile_size , tile_size)\n            # if img_crop.numpy().sum() > threshold:\n            crop_list.append(img_crop)\n    \n    return crop_list\n\ndef tiles_plot(tiles: list[pyvips.Image]):\n    plt.figure(figsize=(10,10),dpi=200)\n    for i,img in enumerate(tiles):\n        plt.subplot(6, 6, i+1)\n        _ = plt.imshow(img)\n\nHGSC_tiles = img_tile(HGSC_img, 500, 50)\nprint(f\"Get {len(HGSC_tiles)} HGSC tiles\")\ntiles_plot(HGSC_tiles)","metadata":{"execution":{"iopub.status.busy":"2023-11-12T03:17:16.704657Z","iopub.execute_input":"2023-11-12T03:17:16.705225Z","iopub.status.idle":"2023-11-12T03:17:22.630431Z","shell.execute_reply.started":"2023-11-12T03:17:16.705185Z","shell.execute_reply":"2023-11-12T03:17:22.628307Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* tile(x,y) means 4th row and 5th col tile.\n\nwe can see papillary area on tile(4,5), (4,6). And that's the evidence of HGSC type ovarian cancer","metadata":{}},{"cell_type":"markdown","source":"## One more HGSC example","metadata":{}},{"cell_type":"code","source":"sample_HGSC = df_HGSC.iloc[3]\n\nsample_HGSC_path = os.path.join(TRAIN_THUMBNAILS_IMAGES,str(sample_HGSC['image_id'])+\"_thumbnail.png\")\nprint(f\"exist HGSC sample file: {os.path.exists(sample_HGSC_path)}\")\n\nHGSC_img = pyvips.Image.new_from_file(sample_HGSC_path)\n\nplt.figure(figsize=(9,3),dpi=300)\nplt.title(sample_HGSC['label'],x = 0.5, y = -0.2)\n_ = plt.imshow(HGSC_img.numpy())\nprint(f\"image information: width= {HGSC_img.width}, height= {HGSC_img.height}.\")","metadata":{"execution":{"iopub.status.busy":"2023-11-12T03:17:22.632189Z","iopub.execute_input":"2023-11-12T03:17:22.632706Z","iopub.status.idle":"2023-11-12T03:17:24.221681Z","shell.execute_reply.started":"2023-11-12T03:17:22.632659Z","shell.execute_reply":"2023-11-12T03:17:24.220489Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from typing import Union\n\ndef img_tile(img: Union[str, pyvips.Image], tile_size: int, threshold:int = 0):\n    if type(img) == str:\n        assert os.path.exists(img)\n        img = pyvips.Image.new_from_file(img)\n    \n    W = img.width\n    H = img.height\n    \n    # crop image into tiles\n    w_tile_num = int(W/tile_size)\n    h_tile_num = int(H/tile_size)\n    \n    crop_list = []\n    for i in range(w_tile_num):\n        for j in range(h_tile_num):\n            img_crop = img.crop(i*tile_size, j*tile_size, tile_size , tile_size)\n            # if img_crop.numpy().sum() > threshold:\n            crop_list.append(img_crop)\n    \n    return crop_list\n\ndef tiles_plot(tiles: list[pyvips.Image]):\n    plt.figure(figsize=(10,10),dpi=200)\n    for i,img in enumerate(tiles):\n        plt.subplot(6, 3, i+1)\n        _ = plt.imshow(img)\n\nHGSC_tiles = img_tile(HGSC_img, 500, 50)\nprint(f\"Get {len(HGSC_tiles)} HGSC tiles\")\ntiles_plot(HGSC_tiles)","metadata":{"execution":{"iopub.status.busy":"2023-11-12T03:17:24.222868Z","iopub.execute_input":"2023-11-12T03:17:24.224128Z","iopub.status.idle":"2023-11-12T03:17:27.399456Z","shell.execute_reply.started":"2023-11-12T03:17:24.224092Z","shell.execute_reply":"2023-11-12T03:17:27.398477Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Look at tiles in rol2 and rol3, that's a classical papillary HGSC case. Now we can get some conclusion from those tiles.\n1. Not all tiles are useful to classification result. some of tiles don't contain any information, they are just black images.\n2. Although HGSC is one of ovarian cancer subtype, it still has some different histopatholohy properties. We are lucky, since most of HGSC's tiles image will show some papillary area.\n3. Most of CNN model prefers to cut a whole image into severals tiles, and classify those tiles, finally some algorithm will be used to balance the classification scores. Maybe we can tune the balance score algorithm according to histopatholohy properties of tiles. ","metadata":{}},{"cell_type":"markdown","source":"## LGSC(Low-Grade Serous Carcinoma)\n\n**LGSC Microscopic Featrues**\n>Accurate pathological evaluation is crucial to managing LGSC properly. It typically shows uniform cells with frank destructive invasion associated with mild to moderate atypia and a low mitotic index.\nLGSC is characterized by a monotonous population of cuboidal, low columnar, and sometimes flattened cells with an amphophilic or lightly eosinophilic cytoplasm. The degree of atypia is mild to moderate, with evenly distributed chromatin; occasional cells with larger nuclei can be seen. As mentioned above, the number of mitoses is no more than 12 mitoses per 10 high-power fields. Destructive invasion is recognized by neoplastic cells in the tumor/ovarian stroma in an area that measures ≥3.0 mm in linear dimension or has desmoplasia. The invasive component may grow in various architectural patterns such as micropapillary, cribriform, elongated papillae, glandular, medium-sized papillae, nests, macro-papillae, cell clusters, and as single cells (Figure 1). Commonly, there is a mix of architectural patterns. Psammoma bodies are common in LGSC and may be numerous. Occasionally, extracellular and intracellular mucin can be observed. Necrosis or multinucleated tumor giant cells are not commonly seen [6,7,35,36]. Sporadically, LGSC is associated with high-grade serous carcinoma at presentation or recurrences [37].\n\nBabaier A, Mal H, Alselwi W, Ghatage P. Low-Grade Serous Carcinoma of the Ovary: The Current Status. Diagnostics (Basel). 2022 Feb 10;12(2):458. doi: 10.3390/diagnostics12020458. PMID: 35204549; PMCID: PMC8871133.\n\n![image.png](https://storage.googleapis.com/kagglesdsdata/datasets/3984006/6938165/LGSC_description.jpg?X-Goog-Algorithm=GOOG4-RSA-SHA256&X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20231110%2Fauto%2Fstorage%2Fgoog4_request&X-Goog-Date=20231110T142608Z&X-Goog-Expires=345600&X-Goog-SignedHeaders=host&X-Goog-Signature=54af2e09e5c78ee8989172c43e8266b9dda333994b16723748c37d20b18c53662b6ceae1ebc40633ca97897563faeb90ed5e910b88cb60c67300a5bd79ddc6e4bdfe5309d5853c2186f3dd1bea08d92dc5fc15f0036ef6cbbfdf4fa880b382b65115925f27d5149a51ceca9bbca57a640868a9bf551fd0e9eaf2be39ae66730396b144d50938e391e7eba3832c05d7805d832481ab41a38238894dbe34824356fdcdbfe7028011aadd7c22c2fc4f5eb8009c5376cfe2ddb8b4edd252794099dc2f039cc035733ffb6922ed0c6e7874bd2fd9aa34ba8a3d8c51bd03f8d4d0996fb5503a26afa5eb021772708f6ddf6bf7e4e60adf85e0351065b9672ceb6ab839)\n\n\nThis example images from: \nAmante S, Santos F, Cunha TM. Low-grade serous epithelial ovarian cancer: a comprehensive review and update for radiologists. Insights Imaging. 2021 May 11;12(1):60. doi: 10.1186/s13244-021-01004-7. PMID: 33974157; PMCID: PMC8113429.","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"df_LGSC = df_train[(df_train['label'] == 'LGSC') & (df_train['is_tma'] == False)]\nplt.figure(figsize=(10,10),dpi=200)\n\n\nsample_LGSC = df_LGSC.iloc[7]\nsample_LGSC_path = os.path.join(TRAIN_THUMBNAILS_IMAGES,str(sample_LGSC['image_id'])+\"_thumbnail.png\")\nprint(f\"exist LGSC sample file: {os.path.exists(sample_HGSC_path)}\")\n\nLGSC_img = pyvips.Image.new_from_file(sample_LGSC_path)\nplt.subplot(1,2,1)\nplt.title(sample_LGSC['label'],x = 0.5, y = -0.2)\n_ = plt.imshow(LGSC_img.numpy())\n\nprint(f\"image information: width= {LGSC_img.width}, height= {LGSC_img.height}.\")\n\nsample_LGSC = df_LGSC.iloc[0]\nsample_LGSC_path = os.path.join(TRAIN_THUMBNAILS_IMAGES,str(sample_LGSC['image_id'])+\"_thumbnail.png\")\nprint(f\"exist LGSC sample file: {os.path.exists(sample_HGSC_path)}\")\n\nLGSC_img = pyvips.Image.new_from_file(sample_LGSC_path)\nplt.subplot(1,2,2)\nplt.title(sample_LGSC['label'],x = 0.5, y = -0.3)\n_ = plt.imshow(LGSC_img.numpy())\n\nprint(f\"image information: width= {LGSC_img.width}, height= {LGSC_img.height}.\")","metadata":{"execution":{"iopub.status.busy":"2023-11-12T03:17:27.400835Z","iopub.execute_input":"2023-11-12T03:17:27.401312Z","iopub.status.idle":"2023-11-12T03:17:30.398707Z","shell.execute_reply.started":"2023-11-12T03:17:27.401283Z","shell.execute_reply":"2023-11-12T03:17:30.397213Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"LGSC_tiles = img_tile(LGSC_img, 500, 50)\nprint(f\"Get {len(LGSC_tiles)} HGSC tiles\")\n\nplt.figure(figsize=(10,10),dpi=200)\nfor i,img in enumerate(LGSC_tiles):\n    plt.subplot(6, 5, i+1)\n    _ = plt.imshow(img)","metadata":{"execution":{"iopub.status.busy":"2023-11-12T03:17:30.400243Z","iopub.execute_input":"2023-11-12T03:17:30.401349Z","iopub.status.idle":"2023-11-12T03:17:35.360562Z","shell.execute_reply.started":"2023-11-12T03:17:30.401299Z","shell.execute_reply":"2023-11-12T03:17:35.358494Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Classical LGSC example","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(5,5),dpi=200)\nimg = LGSC_tiles[6]\nplt.title(\"classical LGSC tile origin\")\n_ = plt.imshow(img)","metadata":{"execution":{"iopub.status.busy":"2023-11-12T03:17:35.362244Z","iopub.execute_input":"2023-11-12T03:17:35.363357Z","iopub.status.idle":"2023-11-12T03:17:36.118159Z","shell.execute_reply.started":"2023-11-12T03:17:35.363304Z","shell.execute_reply":"2023-11-12T03:17:36.116882Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## One classical LGSC tile example\n![LGSC_tile_example](https://storage.googleapis.com/kagglesdsdata/datasets/3984006/6938165/LGSC_example.jpg?X-Goog-Algorithm=GOOG4-RSA-SHA256&X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20231110%2Fauto%2Fstorage%2Fgoog4_request&X-Goog-Date=20231110T142755Z&X-Goog-Expires=345600&X-Goog-SignedHeaders=host&X-Goog-Signature=0c8556ece09a2532e05c807edfee694ab4cb7d7142900202dcc794966201210d34b3dd96b2aa7610e740166e12330c8ca4577913c9c45be9719ed39dfc057be574a850e11c65514f6640f1a7e9ed8e2a57b32158cdf764f35dd1cba27923b2342864b412eeedfd4340a4ced1e63ecbdb2323e53e059c63b6f0b855a4d404adcacc7010165eb7a57950d3163354c50a3d4c9129def620818dcab78aa37d023bf94d9449a681fb50e5181b3b03f09997c396bfaeb7b202496e0d7480f4ebb08eddb23a795920eb12b48b59a69e4f1d4dbf1320d9003bfa918ac2a84dcb2b74d69bc0e01cf1ef68c96b208633419c02f8d5ef604baafe0814135d4b0cd8d02787cd)","metadata":{}},{"cell_type":"markdown","source":"### Conclusion about LGSC tiles\n1. LGSC tiles are similar with HGSC tiles. They are all composed by papillaes. But most of LGSC tiles contains a large number of occasional cells with larger nuclei.(You can find some dark red spots from above tile images)\n2. We can find that most of tiles can show a large number of occasional cells with larger nuclei. That's a good criterion to extract useful tiles or patch. ","metadata":{}},{"cell_type":"markdown","source":"## EC(Endometrioid Carcinoma- Not Finish)\n\n","metadata":{}},{"cell_type":"code","source":"df_EC = df_train[(df_train['label'] == 'EC') & (df_train['is_tma'] == False)]\nplt.figure(figsize=(9,3),dpi=300)\nsample_EC = df_EC.iloc[4]\nsample_EC_path = os.path.join(TRAIN_THUMBNAILS_IMAGES,str(sample_EC['image_id'])+\"_thumbnail.png\")\nprint(f\"exist EC sample file: {os.path.exists(sample_HGSC_path)}\")\n\nEC_img = pyvips.Image.new_from_file(sample_EC_path)\n\nplt.title(sample_EC['label'],x = 0.5, y = -0.2)\nplt.subplot(1,2,1)\n_ = plt.imshow(EC_img.numpy())\n\nsample_EC = df_EC.iloc[1]\nsample_EC_path = os.path.join(TRAIN_THUMBNAILS_IMAGES,str(sample_EC['image_id'])+\"_thumbnail.png\")\nprint(f\"exist EC sample file: {os.path.exists(sample_EC_path)}\")\n\nEC_img = pyvips.Image.new_from_file(sample_EC_path)\n\nplt.subplot(1,2,2)\nplt.title(sample_EC['label'],x = 0.5, y = -0.2)\n_ = plt.imshow(EC_img.numpy())\nprint(f\"image information: width= {EC_img.width}, height= {EC_img.height}.\")","metadata":{"execution":{"iopub.status.busy":"2023-11-12T03:17:36.119485Z","iopub.execute_input":"2023-11-12T03:17:36.120345Z","iopub.status.idle":"2023-11-12T03:17:39.255637Z","shell.execute_reply.started":"2023-11-12T03:17:36.120313Z","shell.execute_reply":"2023-11-12T03:17:39.254602Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_EC = df_train[(df_train['label'] == 'EC') & (df_train['is_tma'] == False)]\nsample_EC = df_EC.iloc[4]\nsample_EC_path = os.path.join(TRAIN_THUMBNAILS_IMAGES,str(sample_EC['image_id'])+\"_thumbnail.png\")\nprint(f\"exist EC sample file: {os.path.exists(sample_HGSC_path)}\")\n\nEC_img = pyvips.Image.new_from_file(sample_EC_path)","metadata":{"execution":{"iopub.status.busy":"2023-11-12T03:17:39.256978Z","iopub.execute_input":"2023-11-12T03:17:39.258137Z","iopub.status.idle":"2023-11-12T03:17:39.268924Z","shell.execute_reply.started":"2023-11-12T03:17:39.258104Z","shell.execute_reply":"2023-11-12T03:17:39.267602Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"EC_tiles = img_tile(EC_img, 500, 50)\nprint(f\"Get {len(EC_tiles)} EC tiles\")\n\nplt.figure(figsize=(10,10),dpi=200)\nfor i,img in enumerate(EC_tiles):\n    plt.subplot(4, 5, i+1)\n    _ = plt.imshow(img)","metadata":{"execution":{"iopub.status.busy":"2023-11-12T03:17:39.270467Z","iopub.execute_input":"2023-11-12T03:17:39.270759Z","iopub.status.idle":"2023-11-12T03:17:42.763219Z","shell.execute_reply.started":"2023-11-12T03:17:39.270733Z","shell.execute_reply":"2023-11-12T03:17:42.761596Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(5,5),dpi=200)\nimg = EC_tiles[4]\nplt.title(\"classical EC tile origin\")\n_ = plt.imshow(img)","metadata":{"execution":{"iopub.status.busy":"2023-11-12T03:17:42.764653Z","iopub.execute_input":"2023-11-12T03:17:42.764985Z","iopub.status.idle":"2023-11-12T03:17:43.410089Z","shell.execute_reply.started":"2023-11-12T03:17:42.764954Z","shell.execute_reply":"2023-11-12T03:17:43.405842Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## CC(Cell Carcinoma-TODO)\n","metadata":{}},{"cell_type":"markdown","source":"## MC(Mucinous Carcinoma-TODO)","metadata":{}},{"cell_type":"markdown","source":"## GO deeper in Ovarian Cancer and its subtypes\n\n**GOAL list:**\n* figure out the principle of general ovarian cancer\n* figure out how to classify different types of ovarian cancer\n    * what properties they have\n    * and analyse how we could use them\n\n## References\n1. Zhu C, Xu Z, Zhang T, Qian L, Xiao W, Wei H, Jin T, Zhou Y. Updates of Pathogenesis, Diagnostic and Therapeutic Perspectives for Ovarian Clear Cell Carcinoma. J Cancer. 2021 Feb 22;12(8):2295-2316. doi: 10.7150/jca.53395. PMID: 33758607; PMCID: PMC7974897.\n2. Lisio MA, Fu L, Goyeneche A, Gao ZH, Telleria C. High-Grade Serous Ovarian Cancer: Basic Sciences, Clinical and Therapeutic Standpoints. Int J Mol Sci. 2019 Feb 22;20(4):952. doi: 10.3390/ijms20040952. PMID: 30813239; PMCID: PMC6412907.\n3. https://en.wikipedia.org/wiki/High-grade_serous_carcinoma\n4. Amante S, Santos F, Cunha TM. Low-grade serous epithelial ovarian cancer: a comprehensive review and update for radiologists. Insights Imaging. 2021 May 11;12(1):60. doi: 10.1186/s13244-021-01004-7. PMID: 33974157; PMCID: PMC8113429.\n5. Babaier A, Mal H, Alselwi W, Ghatage P. Low-Grade Serous Carcinoma of the Ovary: The Current Status. Diagnostics (Basel). 2022 Feb 10;12(2):458. doi: 10.3390/diagnostics12020458. PMID: 35204549; PMCID: PMC8871133.","metadata":{}},{"cell_type":"markdown","source":"**I will still update this notebook for looking for patterns between different subtypes of Ovarian carcinoma. Please upvote if you like this notebook and visualization content. I will be glad if this notebook can push your work.**","metadata":{}}]}