{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":45867,"databundleVersionId":6688004,"sourceType":"competition"}],"dockerImageVersionId":30558,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Cancer Classification with TensorFlow2 Keras.\n","metadata":{}},{"cell_type":"markdown","source":"## Author: Justin Newman\n\n\n### Computer Science and Data Visualization Graduate\n\n\nMore infomation about the author: https://newmanjustin.com\n\n\n\n**I am actively seeking additional team members.**\n\n**Contributions and constructive feedback is always welcome.**\n\n","metadata":{}},{"cell_type":"markdown","source":"## Introduction\nIn this data visualization notebook, I will explore an approach to classify cancer types in tissue samples using TensorFlow for image classification. The primary challenge is handling massive images, some of which can exceed gigabytes in size. Initially, I considered scaling down the images, If we simply scale down the images to feed the model, that is not training the model to look for cancer cells, it would train the model to look for paterns in the tissue sample shape. I believe that splitting the images into smaller segments and classifying those segments may provide a better solution.\n\n### Available Resources.\n\n#### A standard CPU\n\n- Versatile and suitable for a wide range of tasks.\n- General-purpose processing, ideal for everyday computing.\n- Good for sequential tasks and single-threaded applications.\n- Probably the best place to start for most people.\n\n#### NVIDIA Tensor Core GPU T4 x2\n\n- Excellent for parallel processing and handling complex graphics tasks.\n- Suitable for machine learning and deep learning workloads.\n- Provides high throughput and computational power for data-intensive tasks.\n- Good for image processing.\n\n#### NVIDIA Tesla GPU P100\n\n- Offers even greater computational power compared to the T4 GPUs.\n- Ideal for demanding scientific simulations and data analysis.\n- Excellent for deep learning and artificial intelligence tasks.\n- Faster memory and better performance for complex workloads.\n\n#### Google Cloud TensorFlow Processing Unit (TPU) VM v3-8\n\n- Designed specifically for machine learning workloads.\n- Offers exceptional performance for training deep neural networks.\n- Optimized for large-scale model training and inference.\n- Cost-effective choice for AI and ML tasks when compared to GPUs.\n","metadata":{}},{"cell_type":"markdown","source":"## Problem Statement\nThe main problem is to predict the types of cancer in tissue samples. We need to develop an efficient method for processing large images and classifying them accurately. ","metadata":{}},{"cell_type":"markdown","source":"## Proposed Approach\n### Image Segmentation with CUDA GPU acceleration\nTo address the issue of processing massive images, I propose the following approach:\n- **Image Segmentation**: Divide each large image into smaller segments or tiles, e.g., 225x225 pixel tiles. This allows us to focus on smaller sections of the tissue samples and may improve the accuracy of cancer detection.\n- **Currently only runs on CPU**: once I have the model trained on a small set of images and working effectively I plan to implement CUDA GPU acceleration to speed things up.","metadata":{}},{"cell_type":"markdown","source":"### Classification Model\n- **Neural Network Model**: Train a neural network model using TensorFlow Keras to classify each image segment. ","metadata":{}},{"cell_type":"markdown","source":"### Aggregation\n- **Majority Voting**: For each large tissue sample, I plan to collect predictions from all the individual segments. Use a majority voting mechanism to determine the most common classification of cancer type among the segments. This aggregated result will be used to classify the entire tissue sample.","metadata":{}},{"cell_type":"markdown","source":"## Benefits of the Proposed Approach\n- **Improved Accuracy**: By analyzing smaller segments, the model may be able to detect cancerous cells more accurately, as it focuses on finer details.\n- **Efficient Processing**: Processing smaller image segments is more memory-efficient and allows us to handle large images effectively.\n- **Interpretability**: We can visualize the predictions on individual segments, which may provide insights into why certain regions of tissue are classified as cancerous.\n","metadata":{}},{"cell_type":"markdown","source":"## Future Steps\n- **Data Preprocessing**: Explore different preprocessing techniques, such as data augmentation and normalization, to enhance model performance.\n- **Model Tuning**: Experiment with various neural network architectures and hyperparameters to optimize the classification model.\n- **Visualization**: Create data visualizations to understand the distribution of cancer types, performance metrics, and segmentation results.\n- **Evaluation**: Evaluate the model's performance using appropriate metrics, such as accuracy, precision, recall, and F1-score.\n- **Deployment**: Consider deploying the model for real-world cancer diagnosis, integrating it into a medical system.\n","metadata":{}},{"cell_type":"markdown","source":"## Introduction Conclusion\nIn this notebook, I have proposed an approach to classify cancer types in tissue samples using TensorFlow and image segmentation. By dividing large images into smaller segments and aggregating predictions, I aim to improve accuracy and efficiency in cancer diagnosis.\n\nStay tuned for the upcoming sections, where I will implement the proposed approach.\n\nLet's get started!","metadata":{}},{"cell_type":"markdown","source":"# Getting Started\n\nIn this section, I will guide you through the initial steps to get your cancer classification submission up and running. I will begin by importing the necessary libraries and data and then prepare it for image segmentation.\n\n## Prerequisites\nBefore proceeding, ensure you have a subtle understanding of the following prerequisites:\n- Python and basic programming.\n- CUDA-enabled GPU for accelerated image segmentation (optional but highly recommended for performance improvements).\n- Cancer tissue image dataset.\n\n\n","metadata":{}},{"cell_type":"code","source":"import pandas as pd\ndf = pd.read_csv('/kaggle/input/UBC-OCEAN/train.csv')\ndisplay(df)","metadata":{"execution":{"iopub.status.busy":"2024-01-01T20:11:01.381305Z","iopub.execute_input":"2024-01-01T20:11:01.381843Z","iopub.status.idle":"2024-01-01T20:11:01.40501Z","shell.execute_reply.started":"2024-01-01T20:11:01.381804Z","shell.execute_reply":"2024-01-01T20:11:01.403813Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now here I show the 5 classes of cancer in our dataset.","metadata":{}},{"cell_type":"code","source":"unique_labels = df['label'].unique()\nprint(unique_labels)","metadata":{"execution":{"iopub.status.busy":"2024-01-01T20:11:01.407777Z","iopub.execute_input":"2024-01-01T20:11:01.408717Z","iopub.status.idle":"2024-01-01T20:11:01.415948Z","shell.execute_reply.started":"2024-01-01T20:11:01.408668Z","shell.execute_reply":"2024-01-01T20:11:01.414591Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Types of Ovarian Cancer\n\n1. **High-Grade Serous Carcinoma (HGSC):**\n   - *Characteristics:* HGSC is the most common and aggressive type of ovarian cancer, often diagnosed at an advanced stage in older women.\n   - *Treatment:* Surgery followed by chemotherapy.\n\n2. **Low-Grade Serous Carcinoma (LGSC):**\n   - *Characteristics:* LGSC is less common and slower-growing, typically diagnosed at an earlier stage.\n   - *Treatment:* Surgery, sometimes with targeted therapies.\n\n3. **Endometrioid Carcinoma (EC):**\n   - *Characteristics:* EC resembles uterine lining tissue (endometrium), often diagnosed early with a better prognosis.\n   - *Treatment:* Surgery, with possible chemotherapy or hormonal therapy.\n\n4. **Clear Cell Carcinoma (CC):**\n   - *Characteristics:* CC is rare, often chemo-resistant, and associated with a poorer prognosis.\n   - *Treatment:* Surgery, possibly with chemotherapy or targeted therapies.\n\n5. **Mucinous Carcinoma (MC):**\n   - *Characteristics:* MC is less common, usually diagnosed early and localized.\n   - *Treatment:* Surgery, with consideration for chemotherapy based on disease extent.\n\nUnderstanding these ovarian cancer types is crucial for tailoring treatment plans and predicting outcomes.\n","metadata":{}},{"cell_type":"markdown","source":"### Getting a closer look.\n\nNow lets have a look each an example of each type.\nLet's create a new dataframe with the first 5 unqiue labels.","metadata":{}},{"cell_type":"code","source":"unique_labels_df = df.drop_duplicates(subset='label').head(5)\n\n# Display the new DataFrame\ndisplay(unique_labels_df)","metadata":{"execution":{"iopub.status.busy":"2024-01-01T20:11:01.417617Z","iopub.execute_input":"2024-01-01T20:11:01.418102Z","iopub.status.idle":"2024-01-01T20:11:01.437714Z","shell.execute_reply.started":"2024-01-01T20:11:01.418056Z","shell.execute_reply":"2024-01-01T20:11:01.436261Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Can you see the different types of cancer based on these images?","metadata":{}},{"cell_type":"code","source":"import os\nimport matplotlib.pyplot as plt\nfrom matplotlib.image import imread\n\n# Directory containing the images\nimage_dir = '/kaggle/input/UBC-OCEAN/train_thumbnails'\n\n# Create a figure with subplots for each image\nfig, axs = plt.subplots(1, 5, figsize=(15, 5))\n\n# Iterate over each row in the DataFrame with an index\nfor index, row in enumerate(unique_labels_df.iterrows()):\n    _, row_data = row  # Unpack the row data\n    image_id = row_data['image_id']\n    image_path = os.path.join(image_dir, f'{image_id}_thumbnail.png')\n    image = imread(image_path)\n    \n    # Display the image\n    axs[index].imshow(image)\n    axs[index].set_title(f'Label: {row_data[\"label\"]}')\n    axs[index].axis('off')\n\n\n# Adjust spacing between subplots\nplt.tight_layout()\n\n# Show the plot\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-01-01T20:14:46.530269Z","iopub.execute_input":"2024-01-01T20:14:46.530658Z","iopub.status.idle":"2024-01-01T20:14:48.870303Z","shell.execute_reply.started":"2024-01-01T20:14:46.530629Z","shell.execute_reply":"2024-01-01T20:14:48.868704Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## No, I don't think so. They all look pretty much the same.\n\n#### HINT: The Image Classification model won't see the different types of cancer either.","metadata":{}},{"cell_type":"code","source":"import os\nimport matplotlib.pyplot as plt\nfrom matplotlib.image import imread\nfrom PIL import Image\nimport numpy as np\n\n# Directory containing the images\nimage_dir = '/kaggle/input/UBC-OCEAN/train_thumbnails'\n\n# Directory to save the extracted tiles\nsamples_dir = '/kaggle/working/centermost_tiles'\nos.makedirs(samples_dir, exist_ok=True)\n\n# Create a figure with subplots for each image\nfig, axs = plt.subplots(1, 5, figsize=(15, 5))\n\n# Define the size of the centermost tile\ntile_size = (225, 225)\n\n# Iterate over each row in the DataFrame with an index\nfor index, row in enumerate(unique_labels_df.iterrows()):\n    _, row_data = row  # Unpack the row data\n    image_id = row_data['image_id']\n    image_path = os.path.join(image_dir, f'{image_id}_thumbnail.png')\n    image = imread(image_path)\n\n    # Calculate the center coordinates\n    center_x, center_y = image.shape[1] // 2, image.shape[0] // 2\n\n    # Calculate the coordinates for the top-left corner of the centermost tile\n    tile_x = center_x - tile_size[0] // 2\n    tile_y = center_y - tile_size[1] // 2\n\n    # Extract the centermost tile\n    centermost_tile = image[tile_y:tile_y + tile_size[1], tile_x:tile_x + tile_size[0]]\n\n    # Save the centermost tile to the samples directory\n    tile_filename = os.path.join(samples_dir, f'{image_id}_centermost_tile.png')\n    Image.fromarray(np.uint8(centermost_tile)).save(tile_filename)\n\n    # Display the image\n    axs[index].imshow(centermost_tile)\n    axs[index].set_title(f'Label: {row_data[\"label\"]}')\n    axs[index].axis('off')\n\n# Adjust spacing between subplots\nplt.tight_layout()\n\n# Show the plot\nplt.show()","metadata":{"_kg_hide-input":false,"execution":{"iopub.status.busy":"2024-01-01T20:15:05.378186Z","iopub.execute_input":"2024-01-01T20:15:05.378645Z","iopub.status.idle":"2024-01-01T20:15:07.982813Z","shell.execute_reply.started":"2024-01-01T20:15:05.378611Z","shell.execute_reply":"2024-01-01T20:15:07.981342Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Even after we zoom in on the thumbnails. At this scale the image classification model still won't be able to recognize the cancerous cells. \n\n### **It's like training a model to recognize people with satellite images of mountains. **","metadata":{}},{"cell_type":"markdown","source":"# Why Extracting Tiles is Required for Model Training\n\n- **Memory Efficiency**: Large images require significant memory, making tile extraction essential for working with manageable sections.\n\n- **Data Augmentation**: Extracting tiles expands the training dataset, improving model generalization.\n\n- **Feature Localization**: Focusing on image regions of interest helps models learn relevant patterns.\n\n- **Input Consistency**: Tile extraction enforces consistent input sizes, crucial for deep learning models.\n\n- **Batch Processing**: Organizing tiles into batches streamlines parallelized training.\n\n- **Resource Constraints**: Tile extraction is practical in resource-limited environments.\n\n- **Localization Tasks**: Enhances precision in object detection and classification.\n","metadata":{}},{"cell_type":"markdown","source":"# Experimenting with histolab for tile extraction.\n\nhttps://histolab.readthedocs.io/en/latest/readme.html#\n\nWhile the histolab library provides a convenient way to extract tiles from large images, it's important to note that very large images can sometimes run into memory issues when using the library's built-in functions. In such cases, a custom implementation or pull request may be necessary to optimize memory consumption.\n\n## Community Input Needed\n\nIf you've encountered success in processing tiles for Image 51346 using Histolab, I'd greatly appreciate your insights. Currently, working with this particular image seems to pose memory challenges, even with dedicated efforts to mitigate memory consumption.\n\nPerhaps a collaborative approach, combining custom tile extraction methods with the capabilities of Histolab, could lead to optimal results. \n\nPlease consider sharing your feedback or any successful strategies in the comments.","metadata":{}},{"cell_type":"code","source":"#pip install histolab\n\n# from PIL import Image\n# # Set max image pixels to avoid memory errors\n# Image.MAX_IMAGE_PIXELS = None  \n\n# import os\n# from histolab.slide import Slide\n# from histolab.tiler import RandomTiler\n\n# import os\n# import matplotlib.pyplot as plt\n# import gc\n\n\n# def generate_tiles(image_id):\n#     output_dir = '/kaggle/working/tiles'\n#     os.makedirs(output_dir, exist_ok=True)\n#     image_path = f'/kaggle/input/UBC-OCEAN/train_images/{image_id}.png'\n#     image_slide = Slide(image_path, processed_path=output_dir)\n#     random_tiles_extractor = RandomTiler(\n#         tile_size=(225, 225),\n#         n_tiles=5,\n#         seed=7,\n#         level=0,  # (0 for the native resolution)\n#         check_tissue=True,  # Check if tiles contain tissue\n#         prefix=f'{image_id}_',  # Prefix for tile filenames\n#         suffix=\".png\"  # File format for the tiles\n#     )\n    \n#     random_tiles_extractor.locate_tiles(slide=image_slide)\n#     random_tiles_extractor.extract(image_slide)\n    \n#     del random_tiles_extractor\n#     del image_slide\n    \n# gc.collect()\n# # for image_id in data_frame['image_id']:\n# #     generate_tiles(image_id)\n# generate_tiles(51346)","metadata":{"execution":{"iopub.status.busy":"2024-01-01T20:11:11.473283Z","iopub.status.idle":"2024-01-01T20:11:11.474069Z","shell.execute_reply.started":"2024-01-01T20:11:11.473827Z","shell.execute_reply":"2024-01-01T20:11:11.473849Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# New tile extraction process.\n\nTo optimize TensorFlow's performance, it's beneficial to work with smaller images. Here, we outline the process of creating 225x225 pixel tiles from full images for training our TensorFlow model.\n\n## Tile Generation Process\n\n1. **Centering the Tile:**\n   - To start, I extract a 225x225 pixel tile from each full image, ensuring that the selected tile is centered within the original image.\n\n2. **Handling Off-Center Content:**\n   - Keep in mind that not all images will have cell tissue in the direct center. In such cases, I will implement an algorithm to ensure the selection of relevant tiles.\n\n3. **Color Saturation Algorithm:**\n   - I use a saturation algorithm that randomly selects tiles until the color density/variation exceeds a predefined threshold.\n   \n4. **OpenCV Structural Anomaly detection Algorithm:**\n   - As suggested by Ravi Varatha I implemented an OpenCV Structural Anomaly detection function operating on various parameters to get a probabiliy that image contains a cancer cell.\n   \n5. **Average of the color saturation score and the cancer probability score for a prediction of optimal training image.**\n\nBy following this procedure, I aim to create a standardized dataset of 225x225 pixel tiles that are ideal for training our TensorFlow model. This approach ensures that even images with off-center content contribute effectively to the training process.\n","metadata":{}},{"cell_type":"code","source":"import os\nimport matplotlib.pyplot as plt\nfrom matplotlib.image import imread\nimport pandas as pd\nfrom PIL import Image\nimport numpy as np\nimport cv2\n\n# Set max image pixels to avoid memory errors\nImage.MAX_IMAGE_PIXELS = None  \n\n# Parameters\n\ndef generate_tiles(image_id_for_debugging):\n    output_dir = '/kaggle/working/tiles'\n    #os.rmdir(output_dir)\n    os.makedirs(output_dir, exist_ok=True)\n    tissue_density_threshold = 0.5\n    zoom_width = 225\n    zoom_height = 225\n    shift_amount = 3000\n    \n#     def percent_color(image):\n#         # allow RGB values to differ by\n#         threshold = 5\n#         # target the background color\n#         target_color = (235, 229, 233, 255)\n#         if image.shape[2] == 4:\n#             color_difference = np.abs(image - target_color)\n#             matching_pixels = np.sum(np.all(color_difference <= threshold, axis=2))\n#         else:\n#             color_difference = np.abs(image[:, :, :3] - target_color[:3])\n#             matching_pixels = np.sum(np.all(color_difference <= threshold, axis=2))\n#         total_pixels = image.shape[0] * image.shape[1]\n#         percentage = (matching_pixels / total_pixels) * 100\n#         return 1-percentage\n\n\n    def open_computer_vision_cancer_detection(region_array):\n        img = cv2.cvtColor(np.array(region_array), cv2.COLOR_RGB2BGR)\n        gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)\n        blurred = cv2.GaussianBlur(gray, (5, 5), 0)\n        thresh = cv2.threshold(blurred, 60, 255, cv2.THRESH_BINARY)[1]\n        params = cv2.SimpleBlobDetector_Params()\n        params.filterByArea = True\n        params.minArea = 10\n        params.maxArea = 1000\n        detector = cv2.SimpleBlobDetector_create(params)\n        keypoints = detector.detect(thresh)\n        num_blobs = len(keypoints)\n        if num_blobs > 5:\n            cancer_prob = 0.9\n        elif num_blobs > 2:\n            cancer_prob = 0.5\n        else:\n            cancer_prob = 0.1\n        return cancer_prob\n\n\n    def calculate_saturation(img_array):\n        img_hsv = cv2.cvtColor(img_array, cv2.COLOR_RGB2HSV)\n        sat = img_hsv[:,:,1]\n        min_sat = 50\n        colorful_pixels = sat > min_sat\n        num_colorful_pixels = np.count_nonzero(colorful_pixels)\n        total_pixels = img_array.shape[0] * img_array.shape[1]\n        pct_colorful = num_colorful_pixels / total_pixels\n        return pct_colorful\n\n\n    # Get the specific image information for debugging\n    row = df[df['image_id'] == image_id_for_debugging].iloc[0]\n    img_id = row['image_id']\n    image_filename = f'/kaggle/input/UBC-OCEAN/train_images/{img_id}.png'\n\n    img = Image.open(image_filename)\n\n    # Calculate the center of the image\n    center_x = row['image_width'] // 2\n    center_y = row['image_height'] // 2\n\n    # Define the desired dimensions for zooming\n    zoom_width = 225\n    zoom_height = 225\n\n    # Initialize variables for shifting the center\n    center_x_shifted = center_x\n    center_y_shifted = center_y\n\n    # Calculate the coordinates for cropping\n    x_min = center_x - zoom_width // 2\n    x_max = center_x + zoom_width // 2\n    y_min = center_y - zoom_height // 2\n    y_max = center_y + zoom_height // 2\n\n    try_iteration = 0\n    top_tissue_densities = []\n    \n    sample_locations = {}\n    sample_locations[img_id] = {}\n    \n    while try_iteration < 50:\n        \n        x_min = center_x_shifted - zoom_width // 2\n        x_max = center_x_shifted + zoom_width // 2\n        y_min = center_y_shifted - zoom_height // 2\n        y_max = center_y_shifted + zoom_height // 2\n        location = (x_min, y_min, x_max, y_max)\n        sample_locations[img_id][try_iteration] = location\n        \n        cropped_img = img.crop(location)\n        \n        sat_score = calculate_saturation(np.array(cropped_img))\n        openCV_score = open_computer_vision_cancer_detection(np.array(cropped_img))\n#         pc = percent_color(np.array(cropped_img))\n        top_tissue_densities.append(((sat_score + openCV_score)/2, try_iteration))\n        \n        shift_amount = 3000\n        center_x_shifted += np.random.randint(-shift_amount, shift_amount)\n        center_y_shifted += np.random.randint(-shift_amount, shift_amount)\n        \n        try_iteration += 1\n\n    \n    top_tissue_densities.sort(reverse=True, key=lambda x: x[0])\n    top_5_tissue_densities = top_tissue_densities[:5]\n    keepers = [index for _, index in top_5_tissue_densities]\n    \n#     print(\"Best iterations: \", keepers)\n#     print(\"Locations: \")\n#     print(sample_locations)\n    \n    for i, index in enumerate(keepers):\n        cropped_img = img.crop(sample_locations[img_id][index])\n        output_filename = os.path.join(output_dir, f'tile_{img_id}_{i}.png')\n        cropped_img.save(output_filename)\n\n    \n    cropped_img.close()\n    img.close()\n    del cropped_img\n    del img\n\n\n    # Create a figure with subplots for each of the top 5 images\n    fig, axs = plt.subplots(1, 5, figsize=(15, 5))\n    for i, index in enumerate(keepers):\n        image_path = os.path.join(output_dir, f'tile_{img_id}_{i}.png')\n        image = imread(image_path)\n\n        axs[i].imshow(image)\n        axs[i].set_title(f\" Image : {img_id}, {df.loc[df['image_id'] == img_id, 'label'].values[0]}\")\n        axs[i].axis('off')\n    plt.tight_layout()\n    plt.show()\n","metadata":{"_kg_hide-input":false,"execution":{"iopub.status.busy":"2024-01-01T20:15:11.048138Z","iopub.execute_input":"2024-01-01T20:15:11.048602Z","iopub.status.idle":"2024-01-01T20:15:11.079238Z","shell.execute_reply.started":"2024-01-01T20:15:11.048564Z","shell.execute_reply":"2024-01-01T20:15:11.077927Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for image_id in df['image_id']:\n    generate_tiles(image_id)","metadata":{"execution":{"iopub.status.busy":"2024-01-01T20:15:29.880486Z","iopub.execute_input":"2024-01-01T20:15:29.88093Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Mitigating Non-Tissue Background in Selected Samples: A Solution\n\nAvoiding the selection of the non-tissue background is a crucial step in this project. I plan to employ additional algorithms and image processing techniques to identify and select tissue samples while excluding non-tissue background. Here's some additional approaches I might try:\n\n1. **Color-Based Segmentation:**\n   - *Staining:* Tissue samples are often stained to highlight specific structures or features. Color-based segmentation techniques to identify and extract the stained regions that correspond to the tissue of interest might be helpful. I can finetune algorithms to recognize specific color thresholds associated with the stain, separating it from the background.\n\n2. **Thresholding:**\n   - *Brightness and Contrast:* By setting appropriate thresholds for brightness and contrast, I can enhance tissue structures while suppressing the background. Thresholding techniques help separate regions with high contrast (typically tissue) from those with low contrast (often background).\n\n3. **Texture Analysis:**\n   - *Texture Filters:* I plan to try some texture analysis algorithms, such as Gabor filters or wavelet transforms, to detect textural patterns that are characteristic of tissue structures. This can help differentiate tissue regions from non-tissue areas.\n\n4. **Edge Detection:**\n   - *Canny Edge Detection:* Algorithms like the Canny edge detector can identify abrupt changes in intensity or color, which often correspond to tissue boundaries. By detecting edges, I can outline the tissue regions.\n\nBy combining these techniques and customizing them for this specific purpose, I can create algorithms that effectively and accurately identify tissue regions while excluding non-tissue background in the images. I'm also open to any suggestions or feedback. Please feel free to share your ideas on how to better select cancer tissue samples from the full-resolution images.\n","metadata":{}},{"cell_type":"markdown","source":"**Cancer cells** can have various appearances depending on the type of cancer. However, some common characteristics of cancer cells may include:\n\n1. *Uncontrolled Growth:* Cancer cells often divide and multiply uncontrollably, leading to the formation of a tumor.\n\n2. *Variation in Size and Shape:* Cancer cells may vary in size and shape, and they can be larger or smaller than normal cells. Anisocytosis (variation in cell size) and pleomorphism (variation in cell shape) are common features.\n\n3. *Lack of Differentiation:* Cancer cells may be less specialized or differentiated compared to normal cells. This is called anaplasia, and it results in a loss of normal cell function.\n\n4. *Large Nuclei:* The nuclei of cancer cells can be larger and more irregularly shaped than those of normal cells. This is known as hyperchromasia.\n\n5. *Increased Nucleo-Cytoplasmic Ratio:* Cancer cells often have a higher ratio of nucleus to cytoplasm, which can be a sign of malignancy.\n\n6. *Abnormal Nucleoli:* The nucleoli (small structures in the nucleus) may be larger or more numerous in cancer cells.\n\n7. *Invasive Behavior:* Cancer cells can invade nearby tissues and may spread to distant organs through a process called metastasis.\n\n8. *Loss of Contact Inhibition:* Normal cells stop dividing when they come into contact with neighboring cells, a phenomenon known as contact inhibition. Cancer cells may continue to divide even when surrounded by other cells.\n\nIt's important to note that the appearance of cancer cells can vary greatly between different types of cancer. Cancer diagnosis often involves the examination of tissue samples under a microscope by a pathologist to identify these characteristics. The specific appearance of cancer cells is one of the factors used in diagnosing and classifying cancer.\n","metadata":{}},{"cell_type":"markdown","source":"### With the algorithm now capable of obtaining higher-quality tissue samples, the image classifier may finally have a fighting chance.","metadata":{}},{"cell_type":"markdown","source":"# Normalize images and generate batches of tensor image data for training, validation, and testing\n\nIn machine learning, when using Keras or any other framework, it's essential to understand the distinctions between train, validation, and test data. Here's a brief explanation of each:\n\n**1. Training Data:**\n- **Purpose:** Training data is used to train the model. This is where the model learns from patterns and features in the data.\n- **Quantity:** It typically comprises the largest portion of your dataset, often around 60-80% of the data.\n- **Role:** During training, the model's parameters are adjusted to minimize the difference between its predictions and the actual target values in the training data.\n\n**2. Validation Data:**\n- **Purpose:** Validation data is used during training to evaluate the model's performance and make adjustments. It helps in tuning hyperparameters and preventing overfitting.\n- **Quantity:** Typically a smaller portion of the dataset, usually around 10-20%.\n- **Role:** The model is not directly trained on validation data, but it's used to assess how well the model generalizes to data it hasn't seen before. This is crucial for choosing the best model and its hyperparameters.\n\n**3. Test Data:**\n- **Purpose:** Test data is used to evaluate the model's performance after training and hyperparameter tuning. It simulates how the model will perform in the real world with new, unseen data.\n- **Quantity:** Similar in size to the validation set, usually 10-20%.\n- **Role:** Test data provides an unbiased evaluation of the model's accuracy. It's essential to assess how well the model generalizes and whether it's suitable for its intended purpose.\n\nIn summary, training data is used to teach the model, validation data helps in refining the model, and test data serves as a final evaluation to ensure the model's generalization and performance on new, unseen data. Properly splitting and managing these datasets is crucial for developing robust and accurate machine learning models.","metadata":{}},{"cell_type":"markdown","source":"# Choosing a TensorFlow Image Classification Model for Ovarian Cancer Cell Tissue Samples\n\nSelecting the right TensorFlow image classification model is crucial.\n\nFirst, I want to explain a bit about my understanding of Convolutional Neural Networks:\n\nhttps://en.wikipedia.org/wiki/Convolutional_neural_network","metadata":{}},{"cell_type":"markdown","source":"## Convolutional Neural Networks (CNNs) for Cancer Research.\n\n\nIn the context of cell tissue cancer classification, Convolutional Neural Networks function as follows:\n\n1. **Convolutional Layers:** These layers detect important features in cell tissue images, such as the shapes and patterns that distinguish different cancer types.\n\n2. **Pooling Layers:** Pooling layers help identify these features regardless of their precise location within the tissue sample.\n\n3. **Flattening:** Information about the features detected is organized into a list of numbers (a vector), making it suitable for further analysis.\n\n4. **Fully Connected Layers:** These layers learn to recognize complex patterns indicative of specific cancer types, drawing from the detected features.\n\n5. **Output Layer:** Finally, the output layer makes predictions based on the patterns learned, providing classifications for different ovarian cancer types.\n\nNow, I want to explain why I am using Keras for this specific task:\n\nhttps://www.tensorflow.org/guide/keras#who_should_use_keras\n\nhttps://medium.com/analytics-vidhya/sub-classifying-lung-cancer-with-tensorflow-2-and-keras-616353e59e5e","metadata":{}},{"cell_type":"markdown","source":"## What Makes Keras Special for Cancer Classification?\n\nKeras stands out for this project due to the following factors:\n\n1. **User-Friendly Interface:** It provides a straightforward interface, accommodating both beginners and experts in medical image classification.\n\n2. **Modularity:** Keras allows flexible model building, facilitating the integration of domain-specific knowledge about ovarian cancers.\n\n3. **Compatibility:** With multiple backends, Keras adapts well to various data formats and sources, including medical imaging data.\n\n4. **Community and Documentation:** Keras has a strong community and extensive documentation, this is great for a specialized task like cancer classification.\n\n5. **Flexibility:** Customization is possible, allowing for tailored models that focus on the unique characteristics of ovarian cancer cell tissue images.\n\n6. **Integration with TensorFlow:** As an integral part of TensorFlow, Keras combines deep learning capabilities with TensorFlow's versatility.\n\n7. **Transfer Learning:** Keras simplifies the use of pre-trained models, saving time and resources in adapting models for this specific task.\n\n8. **Abstraction from Low-Level Details:** It abstracts complex technicalities, freeing up time to concentrate on fine-tuning the model architecture for cancer classification.\n\n9. **Rapid Prototyping:** Ideal for quick experimentation, Keras facilitates the exploration of different network architectures to identify the best-performing model.\n\n10. **Scalability:** Keras is scalable, accommodating small-scale experiments as well as the processing of large-scale ovarian cancer image datasets.\n\n","metadata":{}},{"cell_type":"markdown","source":"## Choosing a Model with Keras and Sequential\n\nWhen selecting a TensorFlow image classification model in Keras, considering the characteristics of ovarian cancer cell tissue images, dataset size, task complexity, and available resources is crucial. For this specialized project, some pre-trained models particularly relevant to medical image classification include:\n\n- **ResNet:** Leveraging its depth and performance to identify subtle patterns in cancer cell images.\n\n- **Inception:** Efficiently utilizing computational resources to detect unique features indicative of different cancer types.\n\n- **MobileNet:** Balancing speed and accuracy, making it ideal for analyzing large volumes of cell tissue samples.\n\n- **EfficientNet:** Offering state-of-the-art performance, which could be beneficial in achieving high classification accuracy for ovarian cancer types.\n\n- **VGGNet:** Providing a simple yet effective approach, making it a good starting point for educational purposes or as a baseline model.\n\nIn Keras, utilizing the `Sequential` model, I can fine-tune the selected model on the ovarian cancer cell tissue images. Evaluate the model's performance on a validation dataset, and fine-tune hyperparameters within Keras to optimize results specifically for the classification of ovarian cancer types in cell tissue samples.","metadata":{}},{"cell_type":"code","source":"# Once the tile processing algorithm is consistently yeilding optimal tiles for each image I will implement this step.","metadata":{"execution":{"iopub.status.busy":"2024-01-01T20:11:11.480074Z","iopub.status.idle":"2024-01-01T20:11:11.480866Z","shell.execute_reply.started":"2024-01-01T20:11:11.480645Z","shell.execute_reply":"2024-01-01T20:11:11.480667Z"},"trusted":true},"execution_count":null,"outputs":[]}]}