{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Cancer Classification with TensorFlow2 Keras.\n","metadata":{}},{"cell_type":"markdown","source":"## Introduction\nIn this data visualization notebook, I will explore an approach to classify cancer types in tissue samples using TensorFlow for image classification. The primary challenge is handling massive images, some of which can exceed gigabytes in size. Initially, I considered scaling down the images, If we simply scale down the images to feed the model, that is not training the model to look for cancer cells, it would train the model to look for paterns in the tissue sample shape. I believe that splitting the images into smaller segments and classifying those segments may provide a better solution.\n\n### Available Resources.\n\n#### A standard CPU\n\n- Versatile and suitable for a wide range of tasks.\n- General-purpose processing, ideal for everyday computing.\n- Good for sequential tasks and single-threaded applications.\n- Probably the best place to start for most people.\n\n#### NVIDIA Tensor Core GPU T4 x2\n\n- Excellent for parallel processing and handling complex graphics tasks.\n- Suitable for machine learning and deep learning workloads.\n- Provides high throughput and computational power for data-intensive tasks.\n- Good for image processing.\n\n#### NVIDIA Tesla GPU P100\n\n- Offers even greater computational power compared to the T4 GPUs.\n- Ideal for demanding scientific simulations and data analysis.\n- Excellent for deep learning and artificial intelligence tasks.\n- Faster memory and better performance for complex workloads.\n\n#### Google Cloud TensorFlow Processing Unit (TPU) VM v3-8\n\n- Designed specifically for machine learning workloads.\n- Offers exceptional performance for training deep neural networks.\n- Optimized for large-scale model training and inference.\n- Cost-effective choice for AI and ML tasks when compared to GPUs.\n","metadata":{}},{"cell_type":"markdown","source":"## Problem Statement\nThe main problem is to predict the types of cancer in tissue samples. We need to develop an efficient method for processing large images and classifying them accurately. ","metadata":{}},{"cell_type":"markdown","source":"## Proposed Approach\n### Image Segmentation with CUDA GPU acceleration\nTo address the issue of processing massive images, I propose the following approach:\n- **Image Segmentation**: Divide each large image into smaller segments or tiles, e.g., 225x225 pixel tiles. This allows us to focus on smaller sections of the tissue samples and may improve the accuracy of cancer detection.\n- **Currently only runs on CPU**: once I have the model trained on a small set of images and working effectively I plan to implement CUDA GPU acceleration to speed things up.","metadata":{}},{"cell_type":"markdown","source":"### Classification Model\n- **Neural Network Model**: Train a neural network model using TensorFlow Keras to classify each image segment. ","metadata":{}},{"cell_type":"markdown","source":"### Aggregation\n- **Majority Voting**: For each large tissue sample, I plan to collect predictions from all the individual segments. Use a majority voting mechanism to determine the most common classification of cancer type among the segments. This aggregated result will be used to classify the entire tissue sample.","metadata":{}},{"cell_type":"markdown","source":"## Benefits of the Proposed Approach\n- **Improved Accuracy**: By analyzing smaller segments, the model may be able to detect cancerous cells more accurately, as it focuses on finer details.\n- **Efficient Processing**: Processing smaller image segments is more memory-efficient and allows us to handle large images effectively.\n- **Interpretability**: We can visualize the predictions on individual segments, which may provide insights into why certain regions of tissue are classified as cancerous.\n","metadata":{}},{"cell_type":"markdown","source":"## Future Steps\n- **Data Preprocessing**: Explore different preprocessing techniques, such as data augmentation and normalization, to enhance model performance.\n- **Model Tuning**: Experiment with various neural network architectures and hyperparameters to optimize the classification model.\n- **Visualization**: Create data visualizations to understand the distribution of cancer types, performance metrics, and segmentation results.\n- **Evaluation**: Evaluate the model's performance using appropriate metrics, such as accuracy, precision, recall, and F1-score.\n- **Deployment**: Consider deploying the model for real-world cancer diagnosis, integrating it into a medical system.\n","metadata":{}},{"cell_type":"markdown","source":"## Introduction Conclusion\nIn this notebook, I have proposed an approach to classify cancer types in tissue samples using TensorFlow and image segmentation. By dividing large images into smaller segments and aggregating predictions, we aim to improve accuracy and efficiency in cancer diagnosis.\n\nStay tuned for the upcoming sections, where I will implement the proposed approach.\n\nLet's get started!","metadata":{}},{"cell_type":"markdown","source":"# Getting Started\n\nIn this section, I will guide you through the initial steps to get your cancer classification submission up and running. I will begin by importing the necessary libraries and data and then prepare it for image segmentation. Additionally, I will leverage CUDA GPU acceleration to optimize the segmentation process.\n\n## Prerequisites\nBefore proceeding, ensure you have a subtle understanding of the following prerequisites:\n- Python and basic programming.\n- CUDA-enabled GPU for accelerated image segmentation (optional but highly recommended for performance improvement).\n- Cancer tissue image dataset.\n\n\n","metadata":{}},{"cell_type":"markdown","source":"## Importing Required Libraries","metadata":{}},{"cell_type":"code","source":"import warnings\nwarnings.filterwarnings('ignore')\n\nimport os\nimport pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nfrom matplotlib.image import imread\n%matplotlib inline\nimport seaborn as sns\nfrom sklearn.metrics import classification_report , confusion_matrix , accuracy_score , auc\nfrom sklearn.model_selection import train_test_split\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator\n\nimport cv2\n#from google.colab.patches import cv2_imshow\nfrom PIL import Image \nimport tensorflow as tf\nfrom tensorflow import keras\nfrom keras import Sequential\nfrom keras.layers import Input, Dense,Conv2D , MaxPooling2D, Flatten,BatchNormalization,Dropout\nfrom tensorflow.keras.preprocessing import image_dataset_from_directory\nimport tensorflow_hub as hub \nfrom tqdm import tqdm","metadata":{"execution":{"iopub.status.busy":"2023-11-03T10:50:36.788253Z","iopub.execute_input":"2023-11-03T10:50:36.78866Z","iopub.status.idle":"2023-11-03T10:50:46.815991Z","shell.execute_reply.started":"2023-11-03T10:50:36.788628Z","shell.execute_reply":"2023-11-03T10:50:46.815188Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = pd.read_csv('/kaggle/input/UBC-OCEAN/train.csv')\ndf","metadata":{"execution":{"iopub.status.busy":"2023-11-03T09:57:19.969953Z","iopub.execute_input":"2023-11-03T09:57:19.970447Z","iopub.status.idle":"2023-11-03T09:57:20.008095Z","shell.execute_reply.started":"2023-11-03T09:57:19.970421Z","shell.execute_reply":"2023-11-03T09:57:20.007099Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Exploratory Data Analysis","metadata":{}},{"cell_type":"markdown","source":"Now here I show the 5 classes of cancer in our dataset.","metadata":{}},{"cell_type":"code","source":"class_labels = ['CC', 'EC', 'HGSC', 'LGSC', 'MC']\ndf['label'] = df['label'].replace({'CC':0, 'EC':1, 'HGSC':2, 'LGSC':3, 'MC':4})\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2023-11-03T09:57:20.009432Z","iopub.execute_input":"2023-11-03T09:57:20.009816Z","iopub.status.idle":"2023-11-03T09:57:20.025147Z","shell.execute_reply.started":"2023-11-03T09:57:20.00978Z","shell.execute_reply":"2023-11-03T09:57:20.024097Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.info()","metadata":{"execution":{"iopub.status.busy":"2023-11-03T09:57:20.028184Z","iopub.execute_input":"2023-11-03T09:57:20.028576Z","iopub.status.idle":"2023-11-03T09:57:20.052808Z","shell.execute_reply.started":"2023-11-03T09:57:20.028524Z","shell.execute_reply":"2023-11-03T09:57:20.051788Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.countplot(data=df, x=\"label\")\nplt.title(\"Ovarian Cancer Types Distributions\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-11-03T09:57:20.054173Z","iopub.execute_input":"2023-11-03T09:57:20.054499Z","iopub.status.idle":"2023-11-03T09:57:20.333888Z","shell.execute_reply.started":"2023-11-03T09:57:20.054465Z","shell.execute_reply":"2023-11-03T09:57:20.332915Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Types of Ovarian Cancer\n\n1. **High-Grade Serous Carcinoma (HGSC):**\n   - *Characteristics:* HGSC is the most common and aggressive type of ovarian cancer, often diagnosed at an advanced stage in older women.\n   - *Treatment:* Surgery followed by chemotherapy.\n\n2. **Low-Grade Serous Carcinoma (LGSC):**\n   - *Characteristics:* LGSC is less common and slower-growing, typically diagnosed at an earlier stage.\n   - *Treatment:* Surgery, sometimes with targeted therapies.\n\n3. **Endometrioid Carcinoma (EC):**\n   - *Characteristics:* EC resembles uterine lining tissue (endometrium), often diagnosed early with a better prognosis.\n   - *Treatment:* Surgery, with possible chemotherapy or hormonal therapy.\n\n4. **Clear Cell Carcinoma (CC):**\n   - *Characteristics:* CC is rare, often chemo-resistant, and associated with a poorer prognosis.\n   - *Treatment:* Surgery, possibly with chemotherapy or targeted therapies.\n\n5. **Mucinous Carcinoma (MC):**\n   - *Characteristics:* MC is less common, usually diagnosed early and localized.\n   - *Treatment:* Surgery, with consideration for chemotherapy based on disease extent.\n\nUnderstanding these ovarian cancer types is crucial for tailoring treatment plans and predicting outcomes.\n","metadata":{}},{"cell_type":"markdown","source":"### Getting a closer look.\n\nNow lets have a look each an example of each type.\nLet's create a new dataframe with the first 5 unqiue labels.","metadata":{}},{"cell_type":"code","source":"unique_labels_df = df.drop_duplicates(subset='label').head(5)\nunique_labels_df","metadata":{"execution":{"iopub.status.busy":"2023-11-03T09:57:20.33528Z","iopub.execute_input":"2023-11-03T09:57:20.335636Z","iopub.status.idle":"2023-11-03T09:57:20.348717Z","shell.execute_reply.started":"2023-11-03T09:57:20.335603Z","shell.execute_reply":"2023-11-03T09:57:20.347586Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Can you see the different types of cancer based on these images?","metadata":{}},{"cell_type":"code","source":"\n# Directory containing the images\nimage_dir = '/kaggle/input/UBC-OCEAN/train_thumbnails'\n\n# Create a figure with subplots for each image\nfig, axs = plt.subplots(1, 5, figsize=(15, 5))\n\n# Iterate over each row in the DataFrame with an index\nfor index, row in enumerate(unique_labels_df.iterrows()):\n    _, row_data = row  # Unpack the row data\n    image_id = row_data['image_id']\n    image_path = os.path.join(image_dir, f'{image_id}_thumbnail.png')\n    image = imread(image_path)\n    \n    # Display the image\n    axs[index].imshow(image)\n    axs[index].set_title(f'Label: {row_data[\"label\"]}')\n    axs[index].axis('off')\n\n\n# Adjust spacing between subplots\nplt.tight_layout()\n\n# Show the plot\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-11-03T09:57:20.35011Z","iopub.execute_input":"2023-11-03T09:57:20.350428Z","iopub.status.idle":"2023-11-03T09:57:28.217276Z","shell.execute_reply.started":"2023-11-03T09:57:20.350402Z","shell.execute_reply":"2023-11-03T09:57:28.216376Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## No, I don't think so. They all look pretty much the same.\n\n#### HINT: The Image Classification model won't see the different types of cancer either.","metadata":{}},{"cell_type":"code","source":"# Directory containing the images\nimage_dir = '/kaggle/input/UBC-OCEAN/train_thumbnails'\n\n# Directory to save the extracted tiles\nsamples_dir = '/kaggle/working/centermost_tiles'\nos.makedirs(samples_dir, exist_ok=True)\n\n# Create a figure with subplots for each image\nfig, axs = plt.subplots(1, 5, figsize=(15, 5))\n\n# Define the size of the centermost tile\ntile_size = (225, 225)\n\n# Iterate over each row in the DataFrame with an index\nfor index, row in enumerate(unique_labels_df.iterrows()):\n    _, row_data = row  # Unpack the row data\n    image_id = row_data['image_id']\n    image_path = os.path.join(image_dir, f'{image_id}_thumbnail.png')\n    image = imread(image_path)\n\n    # Calculate the center coordinates\n    center_x, center_y = image.shape[1] // 2, image.shape[0] // 2\n\n    # Calculate the coordinates for the top-left corner of the centermost tile\n    tile_x = center_x - tile_size[0] // 2\n    tile_y = center_y - tile_size[1] // 2\n\n    # Extract the centermost tile\n    centermost_tile = image[tile_y:tile_y + tile_size[1], tile_x:tile_x + tile_size[0]]\n\n    # Save the centermost tile to the samples directory\n    tile_filename = os.path.join(samples_dir, f'{image_id}_centermost_tile.png')\n    Image.fromarray(np.uint8(centermost_tile)).save(tile_filename)\n\n    # Display the image\n    axs[index].imshow(centermost_tile)\n    axs[index].set_title(f'Label: {row_data[\"label\"]}')\n    axs[index].axis('off')\n\n# Adjust spacing between subplots\nplt.tight_layout()\n\n# Show the plot\nplt.show()\n","metadata":{"_kg_hide-input":false,"execution":{"iopub.status.busy":"2023-11-03T09:57:28.218491Z","iopub.execute_input":"2023-11-03T09:57:28.218862Z","iopub.status.idle":"2023-11-03T09:57:30.177985Z","shell.execute_reply.started":"2023-11-03T09:57:28.218828Z","shell.execute_reply":"2023-11-03T09:57:30.177124Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Even after we zoom in on the thumbnails. At this scale the image classification model still won't be able to recognize the cancerous cells. \n\n### **It's like training a model to recognize people with satellite images of mountains. **","metadata":{}},{"cell_type":"markdown","source":"# Creating Centered 225x225 Tiles for TensorFlow Model Training\n\nTo optimize TensorFlow's performance, it's beneficial to work with smaller images. Here, we outline the process of creating centered 225x225 pixel tiles from full images for training our TensorFlow model.\n\n## Tile Generation Process\n\n1. **Centering the Tile:**\n   - To start, I extract a 225x225 pixel tile from each full image, ensuring that the selected tile is centered within the original image.\n\n2. **Handling Off-Center Content:**\n   - Keep in mind that not all images will have cell tissue in the direct center. In such cases, I will implement a color density algorithm to ensure the selection of relevant tiles.\n\n3. **Color Saturation Algorithm:**\n   - I use a color density algorithm that randomly selects tiles until the color density/variation exceeds a predefined threshold.\n   \n4. **OpenCV Structural Anomaly detection Algorithm:**\n   - As suggested by Ravi Varatha I use an OpenCV Structural Anomaly detection function operating on various parameters to get a probabiliy that image contains a cancer cell.\n   \n5. **Average of the color saturation score and the cancer probability score for a prediction of optimal training image.**\n\nBy following this procedure, I aim to create a standardized dataset of centered 225x225 pixel tiles that are ideal for training our TensorFlow model. This approach ensures that even images with off-center content contribute effectively to the training process.\n","metadata":{}},{"cell_type":"markdown","source":"## Data Preprocessing","metadata":{}},{"cell_type":"code","source":"\n# Set max image pixels to avoid memory errors\nImage.MAX_IMAGE_PIXELS = None  \n\ndef open_computer_vision_cancer_detection(region_array):\n    img = cv2.cvtColor(np.array(region_array), cv2.COLOR_RGB2BGR)\n    gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)\n    blurred = cv2.GaussianBlur(gray, (5, 5), 0)\n    thresh = cv2.threshold(blurred, 60, 255, cv2.THRESH_BINARY)[1]\n    params = cv2.SimpleBlobDetector_Params()\n    params.filterByArea = True\n    params.minArea = 10\n    params.maxArea = 1000\n    detector = cv2.SimpleBlobDetector_create(params)\n    keypoints = detector.detect(thresh)\n    num_blobs = len(keypoints)\n    if num_blobs > 5:\n        cancer_prob = 0.9\n    elif num_blobs > 2:\n        cancer_prob = 0.5\n    elif num_blobs > 1:\n        cancer_prob = 0.2\n    else:\n        cancer_prob = 0.1\n    return cancer_prob\n","metadata":{"_kg_hide-input":false,"execution":{"iopub.status.busy":"2023-11-03T09:57:30.17925Z","iopub.execute_input":"2023-11-03T09:57:30.179686Z","iopub.status.idle":"2023-11-03T09:57:30.190124Z","shell.execute_reply.started":"2023-11-03T09:57:30.179647Z","shell.execute_reply":"2023-11-03T09:57:30.18896Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def Percentage_non_zero(image):\n\n    # Calculate the total number of pixels in the image\n    total_pixels = np.prod(image.shape[:2])  # Assuming image is a NumPy array\n\n    # Count the number of non-zero pixels in the image\n    non_zero_pixels = np.count_nonzero(image != np.zeros((224, 224, 3)))\n\n    # Calculate the percentage of non-zero pixels\n    percentage_non_zero = (non_zero_pixels / total_pixels) * 100\n\n    # Check if at least 25% of the pixels are non-zero\n    \"\"\"if percentage_non_zero >= 25:\n        print(\"At least 25% of pixels are non-zero.\")\n    else:\n        print(\"Less than 25% of pixels are non-zero.\")\"\"\"\n    return percentage_non_zero\n","metadata":{"execution":{"iopub.status.busy":"2023-11-03T09:57:30.194231Z","iopub.execute_input":"2023-11-03T09:57:30.194561Z","iopub.status.idle":"2023-11-03T09:57:30.207918Z","shell.execute_reply.started":"2023-11-03T09:57:30.194535Z","shell.execute_reply":"2023-11-03T09:57:30.20696Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Data training","metadata":{}},{"cell_type":"code","source":"df[\"is_tma\"].value_counts()","metadata":{"execution":{"iopub.status.busy":"2023-11-03T09:57:30.209173Z","iopub.execute_input":"2023-11-03T09:57:30.20942Z","iopub.status.idle":"2023-11-03T09:57:30.221545Z","shell.execute_reply.started":"2023-11-03T09:57:30.209398Z","shell.execute_reply":"2023-11-03T09:57:30.220525Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Get a list of image IDs from the DataFrame\nimage_ids = df['image_id']\n\n# Create a list of image file names in the train_images directory\nimage_files = os.listdir(\"/kaggle/input/UBC-OCEAN/train_images\")\n\n# Create a list of thumbnail file names in the train_thumbnails directory\nthumbnail_files = os.listdir(\"/kaggle/input/UBC-OCEAN/train_thumbnails\")\n\n# Extract the image IDs from the file names (removing the extension)\n#image_ids_from_files = [t_id.split(\".\")[0] for t_id in image_files if df[\"is_tma\"] == True]\ndf_is_tma = df[df[\"is_tma\"] == True]\nimage_ids_from_is_tma = [id for id in df_is_tma['image_id']]\n\n# Extract the image IDs from the thumbnail file names (removing any additional parts)\nthumbnail_ids = [tmb_id.split(\"_\")[0] for tmb_id in thumbnail_files]\n\n# Find the image IDs that do not have corresponding thumbnail files\nmissing_ids = [int(id) for id in image_ids_from_is_tma if not id in thumbnail_ids]\nlen(missing_ids)","metadata":{"execution":{"iopub.status.busy":"2023-11-03T09:57:30.222937Z","iopub.execute_input":"2023-11-03T09:57:30.223936Z","iopub.status.idle":"2023-11-03T09:57:30.406946Z","shell.execute_reply.started":"2023-11-03T09:57:30.223893Z","shell.execute_reply":"2023-11-03T09:57:30.406091Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### For CPU","metadata":{}},{"cell_type":"code","source":"# Set max image pixels to avoid memory errors\n#Image.MAX_IMAGE_PIXELS = None \nImage.MAX_IMAGE_PIXELS = 10000000000\n\n# Define patch size and overlap (if needed)\npatch_size = (224,224)  # Adjust this according to your requirements\noverlap = 64  # Adjust this if you want overlapping patches\n\nimage_data = []\nimage_label = []\n\nfor img_id, label, tma in tqdm(zip(df['image_id'], df['label'], df['is_tma']), total=len(df)):\n                     \n    if tma==False:\n        img_name = str(img_id)+\"_thumbnail.png\"\n        large_image = Image.open(\"/kaggle/input/UBC-OCEAN/train_thumbnails/\"+img_name)\n        for y in range(0, large_image.height, patch_size[0] - overlap): # (0,2523,160) 224-64=160\n            for x in range(0, large_image.width, patch_size[1] - overlap):  # (0,3000,160) \n                patch = large_image.crop((x, y, x+patch_size[1], y+patch_size[0]))\n                image = np.array(patch)\n                openCV_score = open_computer_vision_cancer_detection(image)\n                percentage_non_zero = Percentage_non_zero(image)\n                if percentage_non_zero > 50 and openCV_score > 0.1:\n                    image_data.append(image)\n                    image_label.append(label)\n                    \n   \n    elif tma==True:\n        img_name = str(img_id)+\".png\"\n        large_image = Image.open(\"/kaggle/input/UBC-OCEAN/train_images/\"+img_name)\n        for y in range(0, large_image.height, patch_size[0] - overlap): # (0,2523,160)\n            for x in range(0, large_image.width, patch_size[1] - overlap):  # (0,3000,160)\n                patch = large_image.crop((x, y, x+patch_size[1], y+patch_size[0]))\n                image = np.array(patch)\n                openCV_score = open_computer_vision_cancer_detection(image)\n                percentage_non_zero = Percentage_non_zero(image)\n                if percentage_non_zero > 25 and openCV_score > 0.1:\n                    image_data.append(image)\n                    image_label.append(label)","metadata":{"execution":{"iopub.status.busy":"2023-11-03T09:57:30.408024Z","iopub.execute_input":"2023-11-03T09:57:30.408286Z","iopub.status.idle":"2023-11-03T10:04:00.924548Z","shell.execute_reply.started":"2023-11-03T09:57:30.408263Z","shell.execute_reply":"2023-11-03T10:04:00.923523Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_is_tma = df[df[\"is_tma\"] == True]\ndf_is_not_tma = df[df[\"is_tma\"] == False]\ndf_is_not_tma_first = df_is_not_tma.drop_duplicates(subset='label', keep='first')\ndf_is_not_tma_last = df_is_not_tma.drop_duplicates(subset='label', keep='last')","metadata":{"execution":{"iopub.status.busy":"2023-11-03T10:04:00.925791Z","iopub.execute_input":"2023-11-03T10:04:00.926079Z","iopub.status.idle":"2023-11-03T10:04:00.934217Z","shell.execute_reply.started":"2023-11-03T10:04:00.926056Z","shell.execute_reply":"2023-11-03T10:04:00.933305Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Concatenate df_is_not_tma_second to df_is_not_tma_first\ndf_is_not_tma_5 = pd.concat([df_is_not_tma_first, df_is_not_tma_last], ignore_index=True)\ndf_is_not_tma_5","metadata":{"execution":{"iopub.status.busy":"2023-11-03T10:04:00.935434Z","iopub.execute_input":"2023-11-03T10:04:00.935699Z","iopub.status.idle":"2023-11-03T10:04:00.955491Z","shell.execute_reply.started":"2023-11-03T10:04:00.935676Z","shell.execute_reply":"2023-11-03T10:04:00.954541Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for img_id, label, tma in tqdm(zip(df_is_not_tma_5['image_id'], df_is_not_tma_5['label'], df_is_not_tma_5['is_tma']), total=len(df_is_not_tma_5)):\n                     \n    img_name = str(img_id)+\".png\"\n    large_image = Image.open(\"/kaggle/input/UBC-OCEAN/train_images/\"+img_name)\n    for y in range(0, large_image.height, patch_size[0] - overlap): # (0,2523,160) 224-64=160\n        for x in range(0, large_image.width, patch_size[1] - overlap):  # (0,3000,160)  224-10=160\n            patch = large_image.crop((x, y, x+patch_size[1], y+patch_size[0]))\n            image = np.array(patch)\n            openCV_score = open_computer_vision_cancer_detection(image)\n            percentage_non_zero = Percentage_non_zero(image)\n            if percentage_non_zero > 50 and openCV_score > 0.1:\n                image_data.append(image)\n                image_label.append(label)","metadata":{"execution":{"iopub.status.busy":"2023-11-03T10:04:00.956653Z","iopub.execute_input":"2023-11-03T10:04:00.956971Z","iopub.status.idle":"2023-11-03T10:33:47.180494Z","shell.execute_reply.started":"2023-11-03T10:04:00.956948Z","shell.execute_reply":"2023-11-03T10:33:47.179544Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### For GPU","metadata":{}},{"cell_type":"code","source":"\"\"\"# Set max image pixels to avoid memory errors\n#Image.MAX_IMAGE_PIXELS = None \nImage.MAX_IMAGE_PIXELS = 10000000000\n\n# Define patch size and overlap (if needed)\npatch_size = (224,224)  # Adjust this according to your requirements\noverlap = 32  # Adjust this if you want overlapping patches\n\nimage_data = []\nimage_label = []\n\nfor img_id, label, tma in tqdm(zip(df['image_id'], df['label'], df['is_tma']), total=len(df)):\n    #print(img_id, label,  tma)\n    if tma==False:\n        img_name = str(img_id)+\".png\"\n        large_image = Image.open(\"/kaggle/input/UBC-OCEAN/train_images/\"+img_name)\n        for y in range(0, large_image.height, patch_size[0] - overlap): # (0,2523,192)\n            for x in range(0, large_image.width, patch_size[1] - overlap):  # (0,3000,192)  224-32=192\n                patch = large_image.crop((x, y, x+patch_size[1], y+patch_size[0]))\n                image = np.array(patch)\n                openCV_score = open_computer_vision_cancer_detection(image)\n                if not (np.all(image == np.zeros((224, 224, 3))) or openCV_score <= 0.1):\n                    image_data.append(image)\n                    image_label.append(label)\n                    \n   \n    elif tma==True and  img_id not in missing_ids:\n        img_name = str(img_id)+\"_thumbnail.png\"\n        large_image = Image.open(\"/kaggle/input/UBC-OCEAN/train_thumbnails/\"+img_name)\n        for y in range(0, large_image.height, patch_size[0] - overlap): # (0,2523,192)\n            for x in range(0, large_image.width, patch_size[1] - overlap):  # (0,3000,192)  224-32=192\n                patch = large_image.crop((x, y, x+patch_size[1], y+patch_size[0]))\n                image = np.array(patch)\n                openCV_score = open_computer_vision_cancer_detection(image)\n                percentage_non_zero = Percentage_non_zero(image)\n                if percentage_non_zero > 50 and openCV_score > 0.1:\n                    image_data.append(image)\n                    image_label.append(label)\"\"\"","metadata":{"execution":{"iopub.status.busy":"2023-11-03T10:33:47.181841Z","iopub.execute_input":"2023-11-03T10:33:47.182473Z","iopub.status.idle":"2023-11-03T10:33:47.191622Z","shell.execute_reply.started":"2023-11-03T10:33:47.182436Z","shell.execute_reply":"2023-11-03T10:33:47.190578Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(len(image_data))\nprint(len(image_label))","metadata":{"execution":{"iopub.status.busy":"2023-11-03T10:33:47.192878Z","iopub.execute_input":"2023-11-03T10:33:47.193179Z","iopub.status.idle":"2023-11-03T10:33:47.20545Z","shell.execute_reply.started":"2023-11-03T10:33:47.193148Z","shell.execute_reply":"2023-11-03T10:33:47.204369Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Image test","metadata":{}},{"cell_type":"code","source":"pd.read_csv('/kaggle/input/UBC-OCEAN/test.csv')","metadata":{"execution":{"iopub.status.busy":"2023-11-03T10:33:47.206783Z","iopub.execute_input":"2023-11-03T10:33:47.207418Z","iopub.status.idle":"2023-11-03T10:33:47.227908Z","shell.execute_reply.started":"2023-11-03T10:33:47.207383Z","shell.execute_reply":"2023-11-03T10:33:47.227003Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Set max image pixels to avoid memory errors\n#Image.MAX_IMAGE_PIXELS = None \nImage.MAX_IMAGE_PIXELS = 10000000000\n\n# Define patch size and overlap (if needed)\npatch_size = (224,224)  # Adjust this according to your requirements\noverlap = 32  # Adjust this if you want overlapping patches\n\nimage_test = []\nlarge_image = Image.open(\"/kaggle/input/UBC-OCEAN/test_images/41.png\")\nfor y in range(0, large_image.height, patch_size[0] - overlap): # (0,16987,192)\n    for x in range(0, large_image.width, patch_size[1] - overlap):  # (0,28469,192)  224-32=192\n        patch = large_image.crop((x, y, x+patch_size[1], y+patch_size[0]))\n        image = np.array(patch)\n        openCV_score = open_computer_vision_cancer_detection(image)\n        if not (np.any(image == np.zeros((224, 224, 3))) or openCV_score < 0.9):\n            image_test.append(image)","metadata":{"execution":{"iopub.status.busy":"2023-11-03T10:33:47.229001Z","iopub.execute_input":"2023-11-03T10:33:47.229295Z","iopub.status.idle":"2023-11-03T10:34:35.607788Z","shell.execute_reply.started":"2023-11-03T10:33:47.22927Z","shell.execute_reply":"2023-11-03T10:34:35.606962Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(image_test)","metadata":{"execution":{"iopub.status.busy":"2023-11-03T10:34:35.608874Z","iopub.execute_input":"2023-11-03T10:34:35.609132Z","iopub.status.idle":"2023-11-03T10:34:35.615562Z","shell.execute_reply.started":"2023-11-03T10:34:35.60911Z","shell.execute_reply":"2023-11-03T10:34:35.614501Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Mitigating Non-Tissue Background in Selected Samples: A Solution\n\nAvoiding the selection of the non-tissue background is a crucial step in this project. I plan to employ additional algorithms and image processing techniques to identify and select tissue samples while excluding non-tissue background. Here's some additional approaches I might try:\n\n1. **Color-Based Segmentation:**\n   - *Staining:* Tissue samples are often stained to highlight specific structures or features. Color-based segmentation techniques to identify and extract the stained regions that correspond to the tissue of interest might be helpful. I can finetune algorithms to recognize specific color thresholds associated with the stain, separating it from the background.\n\n2. **Thresholding:**\n   - *Brightness and Contrast:* By setting appropriate thresholds for brightness and contrast, I can enhance tissue structures while suppressing the background. Thresholding techniques help separate regions with high contrast (typically tissue) from those with low contrast (often background).\n\n3. **Texture Analysis:**\n   - *Texture Filters:* I plan to try some texture analysis algorithms, such as Gabor filters or wavelet transforms, to detect textural patterns that are characteristic of tissue structures. This can help differentiate tissue regions from non-tissue areas.\n\n4. **Edge Detection:**\n   - *Canny Edge Detection:* Algorithms like the Canny edge detector can identify abrupt changes in intensity or color, which often correspond to tissue boundaries. By detecting edges, I can outline the tissue regions.\n\nBy combining these techniques and customizing them for this specific purpose, I can create algorithms that effectively and accurately identify tissue regions while excluding non-tissue background in the images. I'm also open to any suggestions or feedback. Please feel free to share your ideas on how to better select cancer tissue samples from the full-resolution images.\n","metadata":{}},{"cell_type":"markdown","source":"**Cancer cells** can have various appearances depending on the type of cancer. However, some common characteristics of cancer cells may include:\n\n1. *Uncontrolled Growth:* Cancer cells often divide and multiply uncontrollably, leading to the formation of a tumor.\n\n2. *Variation in Size and Shape:* Cancer cells may vary in size and shape, and they can be larger or smaller than normal cells. Anisocytosis (variation in cell size) and pleomorphism (variation in cell shape) are common features.\n\n3. *Lack of Differentiation:* Cancer cells may be less specialized or differentiated compared to normal cells. This is called anaplasia, and it results in a loss of normal cell function.\n\n4. *Large Nuclei:* The nuclei of cancer cells can be larger and more irregularly shaped than those of normal cells. This is known as hyperchromasia.\n\n5. *Increased Nucleo-Cytoplasmic Ratio:* Cancer cells often have a higher ratio of nucleus to cytoplasm, which can be a sign of malignancy.\n\n6. *Abnormal Nucleoli:* The nucleoli (small structures in the nucleus) may be larger or more numerous in cancer cells.\n\n7. *Invasive Behavior:* Cancer cells can invade nearby tissues and may spread to distant organs through a process called metastasis.\n\n8. *Loss of Contact Inhibition:* Normal cells stop dividing when they come into contact with neighboring cells, a phenomenon known as contact inhibition. Cancer cells may continue to divide even when surrounded by other cells.\n\nIt's important to note that the appearance of cancer cells can vary greatly between different types of cancer. Cancer diagnosis often involves the examination of tissue samples under a microscope by a pathologist to identify these characteristics. The specific appearance of cancer cells is one of the factors used in diagnosing and classifying cancer.\n","metadata":{}},{"cell_type":"markdown","source":"### With the algorithm now capable of obtaining higher-quality tissue samples, the image classifier may finally have a fighting chance.","metadata":{}},{"cell_type":"markdown","source":"# Normalize images and generate batches of tensor image data for training, validation, and testing\n\nIn machine learning, when using Keras or any other framework, it's essential to understand the distinctions between train, validation, and test data. Here's a brief explanation of each:\n\n**1. Training Data:**\n- **Purpose:** Training data is used to train the model. This is where the model learns from patterns and features in the data.\n- **Quantity:** It typically comprises the largest portion of your dataset, often around 60-80% of the data.\n- **Role:** During training, the model's parameters are adjusted to minimize the difference between its predictions and the actual target values in the training data.\n\n**2. Validation Data:**\n- **Purpose:** Validation data is used during training to evaluate the model's performance and make adjustments. It helps in tuning hyperparameters and preventing overfitting.\n- **Quantity:** Typically a smaller portion of the dataset, usually around 10-20%.\n- **Role:** The model is not directly trained on validation data, but it's used to assess how well the model generalizes to data it hasn't seen before. This is crucial for choosing the best model and its hyperparameters.\n\n**3. Test Data:**\n- **Purpose:** Test data is used to evaluate the model's performance after training and hyperparameter tuning. It simulates how the model will perform in the real world with new, unseen data.\n- **Quantity:** Similar in size to the validation set, usually 10-20%.\n- **Role:** Test data provides an unbiased evaluation of the model's accuracy. It's essential to assess how well the model generalizes and whether it's suitable for its intended purpose.\n\nIn summary, training data is used to teach the model, validation data helps in refining the model, and test data serves as a final evaluation to ensure the model's generalization and performance on new, unseen data. Properly splitting and managing these datasets is crucial for developing robust and accurate machine learning models.","metadata":{}},{"cell_type":"markdown","source":"## Training Images'TILES' Visualization","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(40,60))\nk=1\nfor i in tqdm(range(650,750)):\n    plt.subplot(15,10,k)\n    plt.imshow(image_data[i])\n    plt.title(f\"Label:{class_labels[image_label[i]]}\")\n    k+=1\n   \n        ","metadata":{"execution":{"iopub.status.busy":"2023-11-03T10:34:35.616871Z","iopub.execute_input":"2023-11-03T10:34:35.617196Z","iopub.status.idle":"2023-11-03T10:35:04.322546Z","shell.execute_reply.started":"2023-11-03T10:34:35.617172Z","shell.execute_reply":"2023-11-03T10:35:04.321499Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Testing Images'TILES' Visualization","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(40,60))\nj=1\nfor i in range(500,650):\n    plt.subplot(15,10,j)\n    plt.imshow(image_test[i])\n    j+=1\n       ","metadata":{"execution":{"iopub.status.busy":"2023-11-03T10:35:04.323907Z","iopub.execute_input":"2023-11-03T10:35:04.324261Z","iopub.status.idle":"2023-11-03T10:35:39.526144Z","shell.execute_reply.started":"2023-11-03T10:35:04.324234Z","shell.execute_reply":"2023-11-03T10:35:39.525003Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Coverting image data into array for training","metadata":{}},{"cell_type":"code","source":"# Changing the list to numpy array\nimage_data = np.array(image_data)\nimage_label = np.array(image_label)\nimage_test = np.array(image_test)\nprint(image_data.shape, image_label.shape, image_test.shape)","metadata":{"execution":{"iopub.status.busy":"2023-11-03T10:35:39.527573Z","iopub.execute_input":"2023-11-03T10:35:39.527916Z","iopub.status.idle":"2023-11-03T10:35:39.921888Z","shell.execute_reply.started":"2023-11-03T10:35:39.527878Z","shell.execute_reply":"2023-11-03T10:35:39.920861Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Saving Image Data","metadata":{}},{"cell_type":"code","source":"import joblib\n\n# # Save the NumPy array to a file\n#joblib.dump(x3, 'x_24000.joblib')\njoblib.dump(image_data, 'image_data.joblib')\njoblib.dump(image_label, 'image_label.joblib')\njoblib.dump(image_test, 'image_test.joblib')","metadata":{"execution":{"iopub.status.busy":"2023-11-03T10:35:39.923431Z","iopub.execute_input":"2023-11-03T10:35:39.923716Z","iopub.status.idle":"2023-11-03T10:35:41.638847Z","shell.execute_reply.started":"2023-11-03T10:35:39.923692Z","shell.execute_reply":"2023-11-03T10:35:41.637871Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import joblib\nimage_data = joblib.load(\"/kaggle/working/image_data.joblib\")\nimage_label = joblib.load(\"/kaggle/working/image_label.joblib\")\nimage_test = joblib.load(\"/kaggle/working/image_test.joblib\")\nprint(image_data.shape)\nprint(image_label.shape)\nprint(image_test.shape)","metadata":{"execution":{"iopub.status.busy":"2023-11-03T10:50:13.814504Z","iopub.execute_input":"2023-11-03T10:50:13.814966Z","iopub.status.idle":"2023-11-03T10:50:20.780754Z","shell.execute_reply.started":"2023-11-03T10:50:13.814932Z","shell.execute_reply":"2023-11-03T10:50:20.779716Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Shuffling the training data","metadata":{}},{"cell_type":"code","source":"shuffle_indexes = np.arange(image_data.shape[0])\nnp.random.shuffle(shuffle_indexes)\nimage_data = image_data[shuffle_indexes]\nimage_label = image_label[shuffle_indexes]","metadata":{"execution":{"iopub.status.busy":"2023-11-03T10:50:58.844822Z","iopub.execute_input":"2023-11-03T10:50:58.84604Z","iopub.status.idle":"2023-11-03T10:50:59.279338Z","shell.execute_reply.started":"2023-11-03T10:50:58.846002Z","shell.execute_reply":"2023-11-03T10:50:59.278519Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Splitting the data into train and validation set test set","metadata":{}},{"cell_type":"code","source":"# Split your data into training and a temporary set (temp_set)\nx_train, x_temp, y_train, y_temp = train_test_split(image_data, image_label, test_size=0.3, random_state=42)\n\n# Split the temporary set into validation and test sets\nx_val, x_test, y_val, y_test = train_test_split(x_temp, y_temp, test_size=0.5, random_state=42)\n","metadata":{"execution":{"iopub.status.busy":"2023-11-03T10:51:04.172914Z","iopub.execute_input":"2023-11-03T10:51:04.173322Z","iopub.status.idle":"2023-11-03T10:51:04.72363Z","shell.execute_reply.started":"2023-11-03T10:51:04.173291Z","shell.execute_reply":"2023-11-03T10:51:04.722548Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"x_train.shape\", x_train.shape)\nprint(\"x_val.shape\", x_val.shape)\nprint(\"x_val.shape\", x_test.shape)\nprint(\"y_train.shape\", y_train.shape)\nprint(\"y_val.shape\", y_val.shape)\nprint(\"y_val.shape\", y_test.shape)","metadata":{"execution":{"iopub.status.busy":"2023-11-03T10:35:42.890745Z","iopub.execute_input":"2023-11-03T10:35:42.89101Z","iopub.status.idle":"2023-11-03T10:35:42.897116Z","shell.execute_reply.started":"2023-11-03T10:35:42.890986Z","shell.execute_reply":"2023-11-03T10:35:42.896128Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\"\"\"x_train, x_val, y_train, y_val = train_test_split(image_data, image_label, test_size=0.2, random_state=42, shuffle=True)\n\nprint(\"x_train.shape\", x_train.shape)\nprint(\"x_valid.shape\", x_val.shape)\nprint(\"y_train.shape\", y_train.shape)\nprint(\"y_valid.shape\", y_val.shape)\"\"\"","metadata":{"execution":{"iopub.status.busy":"2023-11-03T10:35:42.898232Z","iopub.execute_input":"2023-11-03T10:35:42.898527Z","iopub.status.idle":"2023-11-03T10:35:42.910203Z","shell.execute_reply.started":"2023-11-03T10:35:42.898503Z","shell.execute_reply":"2023-11-03T10:35:42.909287Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Data preprocessing before modeling","metadata":{}},{"cell_type":"code","source":"def preprocessing(img): \n    #img = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY) \n    #img = cv2.equalizeHist(img) \n    img = img / 255\n    return img \n  \nx_train = np.array(list(map(preprocessing, x_train))) \nx_val = np.array(list(map(preprocessing, x_val)))\nx_test = np.array(list(map(preprocessing, x_test)))\nimage_test = np.array(list(map(preprocessing, image_test)))\n\nprint(x_train.shape) \nprint(x_val.shape) \nprint(x_test.shape) ","metadata":{"execution":{"iopub.status.busy":"2023-11-03T10:51:08.580686Z","iopub.execute_input":"2023-11-03T10:51:08.581599Z","iopub.status.idle":"2023-11-03T10:51:24.364265Z","shell.execute_reply.started":"2023-11-03T10:51:08.58156Z","shell.execute_reply":"2023-11-03T10:51:24.363203Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## One hot encoding the labels","metadata":{}},{"cell_type":"code","source":"y_train = keras.utils.to_categorical(y_train, 5)\ny_val = keras.utils.to_categorical(y_val, 5)","metadata":{"execution":{"iopub.status.busy":"2023-11-03T10:51:24.365885Z","iopub.execute_input":"2023-11-03T10:51:24.366185Z","iopub.status.idle":"2023-11-03T10:51:24.371218Z","shell.execute_reply.started":"2023-11-03T10:51:24.36616Z","shell.execute_reply":"2023-11-03T10:51:24.370168Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Choosing a TensorFlow Image Classification Model for Ovarian Cancer Cell Tissue Samples\n\nSelecting the right TensorFlow image classification model is crucial.\n\nFirst, I want to explain a bit about my understanding of Convolutional Neural Networks:\n\nhttps://en.wikipedia.org/wiki/Convolutional_neural_network","metadata":{}},{"cell_type":"markdown","source":"## Convolutional Neural Networks (CNNs) for Cancer Research.\n\n\nIn the context of cell tissue cancer classification, Convolutional Neural Networks function as follows:\n\n1. **Convolutional Layers:** These layers detect important features in cell tissue images, such as the shapes and patterns that distinguish different cancer types.\n\n2. **Pooling Layers:** Pooling layers help identify these features regardless of their precise location within the tissue sample.\n\n3. **Flattening:** Information about the features detected is organized into a list of numbers (a vector), making it suitable for further analysis.\n\n4. **Fully Connected Layers:** These layers learn to recognize complex patterns indicative of specific cancer types, drawing from the detected features.\n\n5. **Output Layer:** Finally, the output layer makes predictions based on the patterns learned, providing classifications for different ovarian cancer types.\n","metadata":{}},{"cell_type":"markdown","source":"## What Makes Keras Special for Cancer Classification?\n\nKeras stands out for this project due to the following factors:\n\n1. **User-Friendly Interface:** It provides a straightforward interface, accommodating both beginners and experts in medical image classification.\n\n2. **Modularity:** Keras allows flexible model building, facilitating the integration of domain-specific knowledge about ovarian cancers.\n\n3. **Compatibility:** With multiple backends, Keras adapts well to various data formats and sources, including medical imaging data.\n\n4. **Community and Documentation:** Keras has a strong community and extensive documentation, this is great for a specialized task like cancer classification.\n\n5. **Flexibility:** Customization is possible, allowing for tailored models that focus on the unique characteristics of ovarian cancer cell tissue images.\n\n6. **Integration with TensorFlow:** As an integral part of TensorFlow, Keras combines deep learning capabilities with TensorFlow's versatility.\n\n7. **Transfer Learning:** Keras simplifies the use of pre-trained models, saving time and resources in adapting models for this specific task.\n\n8. **Abstraction from Low-Level Details:** It abstracts complex technicalities, freeing up time to concentrate on fine-tuning the model architecture for cancer classification.\n\n9. **Rapid Prototyping:** Ideal for quick experimentation, Keras facilitates the exploration of different network architectures to identify the best-performing model.\n\n10. **Scalability:** Keras is scalable, accommodating small-scale experiments as well as the processing of large-scale ovarian cancer image datasets.\n\n","metadata":{}},{"cell_type":"markdown","source":"## Choosing a Model with Keras and Sequential\n\nWhen selecting a TensorFlow image classification model in Keras, considering the characteristics of ovarian cancer cell tissue images, dataset size, task complexity, and available resources is crucial. For this specialized project, some pre-trained models particularly relevant to medical image classification include:\n\n- **ResNet:** Leveraging its depth and performance to identify subtle patterns in cancer cell images.\n\n- **Inception:** Efficiently utilizing computational resources to detect unique features indicative of different cancer types.\n\n- **MobileNet:** Balancing speed and accuracy, making it ideal for analyzing large volumes of cell tissue samples.\n\n- **EfficientNet:** Offering state-of-the-art performance, which could be beneficial in achieving high classification accuracy for ovarian cancer types.\n\n- **VGGNet:** Providing a simple yet effective approach, making it a good starting point for educational purposes or as a baseline model.\n\nIn Keras, utilizing the `Sequential` model, I can fine-tune the selected model on the ovarian cancer cell tissue images. Evaluate the model's performance on a validation dataset, and fine-tune hyperparameters within Keras to optimize results specifically for the classification of ovarian cancer types in cell tissue samples.","metadata":{}},{"cell_type":"markdown","source":"## Modeling part","metadata":{}},{"cell_type":"markdown","source":"### Data Augmentation","metadata":{}},{"cell_type":"code","source":"from tensorflow.keras.preprocessing.image import ImageDataGenerator\naug = ImageDataGenerator(\n    rotation_range=10,\n    zoom_range=0.15,\n    width_shift_range=0.1,\n    height_shift_range=0.1,\n    shear_range=0.15,\n    horizontal_flip=False,\n    vertical_flip=False,\n    fill_mode=\"nearest\")","metadata":{"execution":{"iopub.status.busy":"2023-11-03T10:35:58.625588Z","iopub.execute_input":"2023-11-03T10:35:58.626657Z","iopub.status.idle":"2023-11-03T10:35:58.638316Z","shell.execute_reply.started":"2023-11-03T10:35:58.626533Z","shell.execute_reply":"2023-11-03T10:35:58.637404Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mobilenetv2 = tf.keras.applications.MobileNetV2(weights='imagenet', include_top=False, input_shape=(224, 224, 3))","metadata":{"execution":{"iopub.status.busy":"2023-11-03T10:51:24.372428Z","iopub.execute_input":"2023-11-03T10:51:24.372827Z","iopub.status.idle":"2023-11-03T10:51:30.084721Z","shell.execute_reply.started":"2023-11-03T10:51:24.372792Z","shell.execute_reply":"2023-11-03T10:51:30.083789Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Freeze the base model layers\nmobilenetv2.trainable = False\n# Create the input tensor\ninputs = tf.keras.Input(shape=(224, 224, 3))\n\n# Apply the pre-trained base model to the inputs\nx = mobilenetv2(inputs, training=False)\n\n# Add a global average pooling layer\nx = tf.keras.layers.GlobalAveragePooling2D()(x)\n\n# Add the output layer with the desired number of classes\noutputs = Dense(units=5, activation='softmax')(x)\n\n# Create the final model\nmodel_ = tf.keras.Model(inputs, outputs)\n# get the new model summary\n\nmodel_.summary()","metadata":{"execution":{"iopub.status.busy":"2023-11-03T10:51:30.08639Z","iopub.execute_input":"2023-11-03T10:51:30.086696Z","iopub.status.idle":"2023-11-03T10:51:30.565121Z","shell.execute_reply.started":"2023-11-03T10:51:30.08667Z","shell.execute_reply":"2023-11-03T10:51:30.564116Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Train the model","metadata":{}},{"cell_type":"markdown","source":"#### with augmentation","metadata":{}},{"cell_type":"code","source":"# using a learning rate of 0.0001 and Adam optimizer\nmodel_.compile(optimizer=tf.keras.optimizers.Adam(learning_rate=0.0001), loss='categorical_crossentropy', metrics=['accuracy'])","metadata":{"execution":{"iopub.status.busy":"2023-11-03T10:51:33.552787Z","iopub.execute_input":"2023-11-03T10:51:33.553787Z","iopub.status.idle":"2023-11-03T10:51:33.57634Z","shell.execute_reply.started":"2023-11-03T10:51:33.553716Z","shell.execute_reply":"2023-11-03T10:51:33.575435Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"history = model_.fit_generator(aug.flow(x_train, y_train, batch_size=32), epochs=10, validation_data=aug.flow(x_val, y_val, batch_size=32))\n","metadata":{"execution":{"iopub.status.busy":"2023-11-03T10:36:04.257021Z","iopub.execute_input":"2023-11-03T10:36:04.257395Z","iopub.status.idle":"2023-11-03T10:49:57.858779Z","shell.execute_reply.started":"2023-11-03T10:36:04.257359Z","shell.execute_reply":"2023-11-03T10:49:57.857768Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.plot(history.history['loss']) \nplt.plot(history.history['val_loss']) \nplt.legend(['training_loss', 'validation_loss']) \nplt.title('Loss') \nplt.xlabel('epoch') ","metadata":{"execution":{"iopub.status.busy":"2023-11-03T10:49:57.860362Z","iopub.execute_input":"2023-11-03T10:49:57.860675Z","iopub.status.idle":"2023-11-03T10:49:58.216657Z","shell.execute_reply.started":"2023-11-03T10:49:57.860648Z","shell.execute_reply":"2023-11-03T10:49:58.215663Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.plot(history.history['accuracy']) \nplt.plot(history.history['val_accuracy']) \nplt.legend(['training_accuracy', 'validation_accuracy']) \nplt.title('accuracy') \nplt.xlabel('epoch') ","metadata":{"execution":{"iopub.status.busy":"2023-11-03T10:49:58.217835Z","iopub.execute_input":"2023-11-03T10:49:58.218149Z","iopub.status.idle":"2023-11-03T10:49:58.562656Z","shell.execute_reply.started":"2023-11-03T10:49:58.218122Z","shell.execute_reply":"2023-11-03T10:49:58.561686Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### without augmentation","metadata":{}},{"cell_type":"code","source":"# Running the model using 10 epochs and saving the history\n\nhistory = model_.fit(x_train, y_train, batch_size=32, epochs=10, validation_data=(x_val, y_val))","metadata":{"execution":{"iopub.status.busy":"2023-11-03T10:51:40.160642Z","iopub.execute_input":"2023-11-03T10:51:40.16179Z","iopub.status.idle":"2023-11-03T10:53:45.721602Z","shell.execute_reply.started":"2023-11-03T10:51:40.161722Z","shell.execute_reply":"2023-11-03T10:53:45.720694Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.plot(history.history['accuracy']) \nplt.plot(history.history['val_accuracy']) \nplt.legend(['training_accuracy', 'validation_accuracy']) \nplt.title('accuracy') \nplt.xlabel('epoch') ","metadata":{"execution":{"iopub.status.busy":"2023-11-03T10:53:45.72409Z","iopub.execute_input":"2023-11-03T10:53:45.731048Z","iopub.status.idle":"2023-11-03T10:53:46.099509Z","shell.execute_reply.started":"2023-11-03T10:53:45.731017Z","shell.execute_reply":"2023-11-03T10:53:46.098537Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.plot(history.history['loss']) \nplt.plot(history.history['val_loss']) \nplt.legend(['training_loss', 'validation_loss']) \nplt.title('Loss') \nplt.xlabel('epoch') ","metadata":{"execution":{"iopub.status.busy":"2023-11-03T10:53:46.101551Z","iopub.execute_input":"2023-11-03T10:53:46.10237Z","iopub.status.idle":"2023-11-03T10:53:46.431445Z","shell.execute_reply.started":"2023-11-03T10:53:46.102328Z","shell.execute_reply":"2023-11-03T10:53:46.430498Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Extract features for aggregation (majority voting)","metadata":{}},{"cell_type":"code","source":"# Extract features from your testing data (image_data)\nfeatures = model_.predict(x_test)\npred = features.argmax(axis=1)","metadata":{"execution":{"iopub.status.busy":"2023-11-03T10:53:46.433514Z","iopub.execute_input":"2023-11-03T10:53:46.433853Z","iopub.status.idle":"2023-11-03T10:53:52.703939Z","shell.execute_reply.started":"2023-11-03T10:53:46.433825Z","shell.execute_reply":"2023-11-03T10:53:52.703111Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Evaluation Model","metadata":{}},{"cell_type":"code","source":"model_.evaluate(x_test, keras.utils.to_categorical(y_test, 5))","metadata":{"execution":{"iopub.status.busy":"2023-11-03T10:53:52.705485Z","iopub.execute_input":"2023-11-03T10:53:52.706035Z","iopub.status.idle":"2023-11-03T10:53:56.59001Z","shell.execute_reply.started":"2023-11-03T10:53:52.705967Z","shell.execute_reply":"2023-11-03T10:53:56.588991Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.metrics import confusion_matrix\ncm = confusion_matrix(y_true=y_test, y_pred=pred)","metadata":{"execution":{"iopub.status.busy":"2023-11-03T10:53:56.591226Z","iopub.execute_input":"2023-11-03T10:53:56.591543Z","iopub.status.idle":"2023-11-03T10:53:56.599891Z","shell.execute_reply.started":"2023-11-03T10:53:56.591515Z","shell.execute_reply":"2023-11-03T10:53:56.598904Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import seaborn as sns\ndf_cm = pd.DataFrame(cm, index = [0,1,2,3,4],  columns = [0,1,2,3,4])\nplt.figure(figsize = (10,10))\nsns.heatmap(df_cm, annot=True)\nplt.xlabel('Predicted Label')\nplt.ylabel('True Label')","metadata":{"execution":{"iopub.status.busy":"2023-11-03T10:53:56.601107Z","iopub.execute_input":"2023-11-03T10:53:56.601378Z","iopub.status.idle":"2023-11-03T10:53:57.118699Z","shell.execute_reply.started":"2023-11-03T10:53:56.601355Z","shell.execute_reply":"2023-11-03T10:53:57.117773Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Classification report","metadata":{}},{"cell_type":"code","source":"from sklearn.metrics import classification_report\n\nprint(classification_report(y_test, pred))","metadata":{"execution":{"iopub.status.busy":"2023-11-03T10:53:57.120045Z","iopub.execute_input":"2023-11-03T10:53:57.120439Z","iopub.status.idle":"2023-11-03T10:53:57.137003Z","shell.execute_reply.started":"2023-11-03T10:53:57.120403Z","shell.execute_reply":"2023-11-03T10:53:57.135973Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Aggregation","metadata":{}},{"cell_type":"code","source":"from collections import Counter\n\n# Collect predictions from individual segments\nsegment_predictions = pred\n\n# Apply majority voting\nmajority_vote = Counter(segment_predictions).most_common(1)[0]\n\n# The most common classification is the final classification for the tissue sample\nfinal_classification = majority_vote[0]\nfinal_classification","metadata":{"execution":{"iopub.status.busy":"2023-11-03T10:53:57.138272Z","iopub.execute_input":"2023-11-03T10:53:57.138555Z","iopub.status.idle":"2023-11-03T10:53:57.14665Z","shell.execute_reply.started":"2023-11-03T10:53:57.138528Z","shell.execute_reply":"2023-11-03T10:53:57.145623Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Submission","metadata":{}},{"cell_type":"code","source":"df_submission = pd.read_csv(\"/kaggle/input/UBC-OCEAN/sample_submission.csv\")","metadata":{"execution":{"iopub.status.busy":"2023-11-03T10:53:57.155702Z","iopub.execute_input":"2023-11-03T10:53:57.155998Z","iopub.status.idle":"2023-11-03T10:53:57.17455Z","shell.execute_reply.started":"2023-11-03T10:53:57.155967Z","shell.execute_reply":"2023-11-03T10:53:57.173835Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_submission[\"label\"] = class_labels[final_classification]\ndf_submission","metadata":{"execution":{"iopub.status.busy":"2023-11-03T10:55:07.460212Z","iopub.execute_input":"2023-11-03T10:55:07.460891Z","iopub.status.idle":"2023-11-03T10:55:07.476309Z","shell.execute_reply.started":"2023-11-03T10:55:07.460856Z","shell.execute_reply":"2023-11-03T10:55:07.475123Z"},"trusted":true},"execution_count":null,"outputs":[]}]}