{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"**Introduction Evaluation:**\n\n1. **Clarity and Informative Content**: Your introduction provides a clear and informative overview of your project, including the problem statement, proposed approach, and the benefits of your approach. It's well-structured and outlines your plan comprehensively.\n\n**Suggestions for Improvement:**\n\n1. **Clear Objectives**: It would be beneficial to explicitly state the objectives of your project right at the beginning. For example, you can mention that the primary goal is to develop a cancer type classification system using TensorFlow and image segmentation.\n\n2. **Bullet Points**: Consider presenting the information about available resources and benefits of the proposed approach as bullet points for better readability and organization.\n\n3. **Data Source Information**: It would be helpful to mention the source of your cancer tissue image dataset and provide specific details about the dataset. Is it publicly available, or did you collect it yourself?\n\n4. **Future Steps Plan**: In the \"Future Steps\" section, you can provide a rough timeline or plan for when you expect to complete each step. This can help readers understand your project's progress and anticipated milestones.\n\n5. **Visual Aids**: Since you plan to create data visualizations, consider adding a visualization or a sample image to give readers an early glimpse of what to expect in your project.\n\n","metadata":{}},{"cell_type":"code","source":"# Import necessary libraries for image processing and visualization\nimport os  # Library for interacting with the operating system\nimport matplotlib.pyplot as plt  # Library for creating visualizations\nfrom matplotlib.image import imread  # Function for reading images\nimport pandas as pd  # Import the pandas library and alias it as 'pd'\nfrom PIL import Image  # Python Imaging Library for image processing\nimport numpy as np  # Library for numerical operations\nimport cv2  # OpenCV library for computer vision and image processing\n\n# Note: The following code repeats the same import statements multiple times,\n# which is unnecessary and can be avoided in practice. It's a good practice to\n# include import statements at the beginning of your script, and you don't need\n# to repeat them multiple times. Keeping your imports organized makes your code\n# more readable and efficient.\n\n# Additional unnecessary import statements\n# Repeating import statements can clutter your code.\n# The following imports are already included above.\n\n# os\n# matplotlib\n# Image\n# numpy\n\n","metadata":{"execution":{"iopub.status.busy":"2023-10-16T01:26:10.943976Z","iopub.execute_input":"2023-10-16T01:26:10.944449Z","iopub.status.idle":"2023-10-16T01:26:10.951308Z","shell.execute_reply.started":"2023-10-16T01:26:10.944417Z","shell.execute_reply":"2023-10-16T01:26:10.950095Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Read a CSV file using pandas and store the data in a DataFrame named 'df'\ndf = pd.read_csv('/kaggle/input/UBC-OCEAN/train.csv')\n\n# Display the contents of the DataFrame 'df'\ndisplay(df)\n","metadata":{"execution":{"iopub.status.busy":"2023-10-16T01:26:10.978493Z","iopub.execute_input":"2023-10-16T01:26:10.979101Z","iopub.status.idle":"2023-10-16T01:26:10.998761Z","shell.execute_reply.started":"2023-10-16T01:26:10.979065Z","shell.execute_reply":"2023-10-16T01:26:10.997886Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\"Now, I will display the five classes of cancer present in our dataset.\"","metadata":{}},{"cell_type":"code","source":"# Get unique values from the 'label' column in the DataFrame 'df'\nunique_labels = df['label'].unique()\n\n# Print the unique labels\nprint(unique_labels)\n","metadata":{"execution":{"iopub.status.busy":"2023-10-16T01:26:11.015497Z","iopub.execute_input":"2023-10-16T01:26:11.016206Z","iopub.status.idle":"2023-10-16T01:26:11.023181Z","shell.execute_reply.started":"2023-10-16T01:26:11.016162Z","shell.execute_reply":"2023-10-16T01:26:11.022038Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Types of Ovarian Cancer**\n\n1. **High-Grade Serous Carcinoma (HGSC)**\n   - **Characteristics:** HGSC is the most common and aggressive type of ovarian cancer, often diagnosed at an advanced stage in older women.\n   - **Treatment:** Surgery followed by chemotherapy.\n\n2. **Low-Grade Serous Carcinoma (LGSC)**\n   - **Characteristics:** LGSC is less common and slower-growing, typically diagnosed at an earlier stage.\n   - **Treatment:** Surgery, sometimes with targeted therapies.\n\n3. **Endometrioid Carcinoma (EC)**\n   - **Characteristics:** EC resembles uterine lining tissue (endometrium), often diagnosed early with a better prognosis.\n   - **Treatment:** Surgery, with possible chemotherapy or hormonal therapy.\n\n4. **Clear Cell Carcinoma (CC)**\n   - **Characteristics:** CC is rare, often chemo-resistant, and associated with a poorer prognosis.\n   - **Treatment:** Surgery, possibly with chemotherapy or targeted therapies.\n\n5. **Mucinous Carcinoma (MC)**\n   - **Characteristics:** MC is less common, usually diagnosed early and localized.\n   - **Treatment:** Surgery, with consideration for chemotherapy based on disease extent.\n\nUnderstanding these ovarian cancer types is crucial for tailoring treatment plans and predicting outcomes.\n\n**Getting a Closer Look**\nNow, let's examine an example of each type. We'll create a new dataframe with the first 5 unique labels.\n\n","metadata":{}},{"cell_type":"code","source":"# Create a new DataFrame 'unique_labels_df' by dropping duplicates in the 'label' column and selecting the first 5 unique rows\nunique_labels_df = df.drop_duplicates(subset='label').head(5)\n\n# Display the new DataFrame 'unique_labels_df'\ndisplay(unique_labels_df)\n","metadata":{"execution":{"iopub.status.busy":"2023-10-16T01:26:11.046396Z","iopub.execute_input":"2023-10-16T01:26:11.047657Z","iopub.status.idle":"2023-10-16T01:26:11.064542Z","shell.execute_reply.started":"2023-10-16T01:26:11.047611Z","shell.execute_reply":"2023-10-16T01:26:11.063256Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Can you distinguish between different cancer types based on these images?**","metadata":{}},{"cell_type":"code","source":"# Directory containing the images\nimage_dir = '/kaggle/input/UBC-OCEAN/train_thumbnails'\n\n# Create a figure with subplots for each image\nfig, axs = plt.subplots(1, 5, figsize=(15, 5))\n\n# Iterate over each row in the DataFrame with an index\nfor index, (_, row_data) in enumerate(unique_labels_df.iterrows()):\n    image_id = row_data['image_id']\n    image_path = os.path.join(image_dir, f'{image_id}_thumbnail.png')\n    image = imread(image_path)\n    \n    # Display the image\n    axs[index].imshow(image)\n    axs[index].set_title(f'Label: {row_data[\"label\"]}')\n    axs[index].axis('off')\n\n# Adjust spacing between subplots\nplt.tight_layout()\n\n# Show the plot\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2023-10-16T01:26:11.07653Z","iopub.execute_input":"2023-10-16T01:26:11.077191Z","iopub.status.idle":"2023-10-16T01:26:20.3912Z","shell.execute_reply.started":"2023-10-16T01:26:11.077153Z","shell.execute_reply":"2023-10-16T01:26:20.389759Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**No, I don't think so. They all look pretty much the same.**\n  \n**HINT:** The Image Classification model won't differentiate between the different types of cancer either.","metadata":{}},{"cell_type":"code","source":"#!pip install scikit-image\n# sample training image\nimport pandas as pd\nimport matplotlib.pyplot as plt\nfrom skimage import io\nimport os\nimport seaborn as sns\nimport cv2\nimport random\nimport os\nimport glob\nio.imshow('/kaggle/input/UBC-OCEAN/train_thumbnails/10642_thumbnail.png')","metadata":{"execution":{"iopub.status.busy":"2023-10-16T01:26:20.394246Z","iopub.execute_input":"2023-10-16T01:26:20.395574Z","iopub.status.idle":"2023-10-16T01:26:21.568818Z","shell.execute_reply.started":"2023-10-16T01:26:20.395524Z","shell.execute_reply":"2023-10-16T01:26:21.568034Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load train data\ntrain_df = pd.read_csv('/kaggle/input/UBC-OCEAN/train.csv')\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-10-16T01:26:21.570006Z","iopub.execute_input":"2023-10-16T01:26:21.570765Z","iopub.status.idle":"2023-10-16T01:26:21.584515Z","shell.execute_reply.started":"2023-10-16T01:26:21.570731Z","shell.execute_reply":"2023-10-16T01:26:21.583112Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(train_df.shape)","metadata":{"execution":{"iopub.status.busy":"2023-10-16T01:26:21.58787Z","iopub.execute_input":"2023-10-16T01:26:21.588887Z","iopub.status.idle":"2023-10-16T01:26:21.594679Z","shell.execute_reply.started":"2023-10-16T01:26:21.588848Z","shell.execute_reply":"2023-10-16T01:26:21.593277Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.isna().sum().sum()","metadata":{"execution":{"iopub.status.busy":"2023-10-16T01:26:21.596102Z","iopub.execute_input":"2023-10-16T01:26:21.596468Z","iopub.status.idle":"2023-10-16T01:26:21.610689Z","shell.execute_reply.started":"2023-10-16T01:26:21.596436Z","shell.execute_reply":"2023-10-16T01:26:21.609523Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Import necessary libraries\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n# List of columns for which we want to visualize the distributions\ncolumns = [col for col in train_df.columns if col != 'label']\n\n# Loop to iterate over each column\nfor col in columns:\n    \n    # Create a subplot for 3 columns (3 plots)\n    fig, axs = plt.subplots(figsize=(15, 5), ncols=3)\n    \n    # 1st plot - Distribution of the sample dataset\n    sns.histplot(data=train_df, x=col, kde=True, ax=axs[0])\n    axs[0].set_title('Sample Distribution')\n    \n    # 2nd plot - Distribution of the selected column where the outcome is 1 (has diabetes)\n    sns.histplot(data=train_df[train_df['label'] == \"HGSC\"], x=col, kde=True, ax=axs[1], color='orange')\n    axs[1].set_title('label - HGSC')\n    \n    # 3rd plot - Distribution of the selected column where the outcome is 0 (doesn't have diabetes)\n    sns.histplot(data=train_df[train_df['label'] == \"LGSC\"], x=col, kde=True, ax=axs[2], color='green')\n    axs[2].set_title('label - LGSC')\n    \n    # Show the plots\n    plt.tight_layout()\n    plt.show()\n","metadata":{"execution":{"iopub.status.busy":"2023-10-16T01:26:21.612294Z","iopub.execute_input":"2023-10-16T01:26:21.612969Z","iopub.status.idle":"2023-10-16T01:26:25.573182Z","shell.execute_reply.started":"2023-10-16T01:26:21.612916Z","shell.execute_reply":"2023-10-16T01:26:25.57183Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display information about the DataFrame\ntrain_df.info()\n","metadata":{"execution":{"iopub.status.busy":"2023-10-16T01:26:25.574607Z","iopub.execute_input":"2023-10-16T01:26:25.574993Z","iopub.status.idle":"2023-10-16T01:26:25.589884Z","shell.execute_reply.started":"2023-10-16T01:26:25.574962Z","shell.execute_reply":"2023-10-16T01:26:25.587955Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Check if the 'image_id' column has unique values\nprint(train_df['image_id'].is_unique)\n\n# Check if the 'label' column has unique values\nprint(train_df['label'].is_unique)\n","metadata":{"execution":{"iopub.status.busy":"2023-10-16T01:26:25.591588Z","iopub.execute_input":"2023-10-16T01:26:25.59207Z","iopub.status.idle":"2023-10-16T01:26:25.609202Z","shell.execute_reply.started":"2023-10-16T01:26:25.592033Z","shell.execute_reply":"2023-10-16T01:26:25.607601Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Obtain descriptive statistics for the 'image_height' and 'image_width' columns\n# in the train_df DataFrame.\n\n# Importing necessary libraries\nimport pandas as pd\n\n# Select the 'image_height' and 'image_width' columns and calculate statistics\nstatistics = train_df[['image_height', 'image_width']].describe()\n\n# Display the descriptive statistics\nprint(statistics)\n","metadata":{"execution":{"iopub.status.busy":"2023-10-16T01:26:25.611152Z","iopub.execute_input":"2023-10-16T01:26:25.61155Z","iopub.status.idle":"2023-10-16T01:26:25.633709Z","shell.execute_reply.started":"2023-10-16T01:26:25.611515Z","shell.execute_reply":"2023-10-16T01:26:25.632276Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Print the number of rows in the 'image_id' column of the DataFrame\nprint(train_df['image_id'].shape[0])\n\n# Count the number of files in the 'train_images' directory\nimport os\nimage_files = os.listdir('/kaggle/input/UBC-OCEAN/train_images')\nprint(len(image_files))\n\n# Count the number of files in the 'train_thumbnails' directory\nthumbnail_files = os.listdir('/kaggle/input/UBC-OCEAN/train_thumbnails')\nprint(len(thumbnail_files))\n","metadata":{"execution":{"iopub.status.busy":"2023-10-16T01:26:25.63787Z","iopub.execute_input":"2023-10-16T01:26:25.638267Z","iopub.status.idle":"2023-10-16T01:26:25.648277Z","shell.execute_reply.started":"2023-10-16T01:26:25.638233Z","shell.execute_reply":"2023-10-16T01:26:25.646772Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Calculate and display the value counts of the 'label' column in train_df\nlabel_counts = train_df['label'].value_counts()\nprint(label_counts)\n","metadata":{"execution":{"iopub.status.busy":"2023-10-16T01:26:25.65044Z","iopub.execute_input":"2023-10-16T01:26:25.650876Z","iopub.status.idle":"2023-10-16T01:26:25.668332Z","shell.execute_reply.started":"2023-10-16T01:26:25.650835Z","shell.execute_reply":"2023-10-16T01:26:25.667399Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Import the warnings module to manage warning messages\nimport warnings\n\n# Use the filterwarnings method to ignore warning messages\nwarnings.filterwarnings('ignore')\n","metadata":{"execution":{"iopub.status.busy":"2023-10-16T01:26:25.669984Z","iopub.execute_input":"2023-10-16T01:26:25.670392Z","iopub.status.idle":"2023-10-16T01:26:25.680329Z","shell.execute_reply.started":"2023-10-16T01:26:25.670326Z","shell.execute_reply":"2023-10-16T01:26:25.678826Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display the value counts for the 'label' column\nlabel_value_counts = train_df['label'].value_counts()\nprint(label_value_counts)\n","metadata":{"execution":{"iopub.status.busy":"2023-10-16T01:26:25.682035Z","iopub.execute_input":"2023-10-16T01:26:25.683107Z","iopub.status.idle":"2023-10-16T01:26:25.699307Z","shell.execute_reply.started":"2023-10-16T01:26:25.683063Z","shell.execute_reply":"2023-10-16T01:26:25.697917Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"HGSC = train_df[train_df['label']==\"HGSC\"]\nEC = train_df[train_df['label']==\"EC\"]\nCC = train_df[train_df['label']==\"CC\"]\nLGSC = train_df[train_df['label']==\"LGSC\"]\nMC = train_df[train_df['label']==\"MC\"]","metadata":{"execution":{"iopub.status.busy":"2023-10-16T01:26:25.701316Z","iopub.execute_input":"2023-10-16T01:26:25.701741Z","iopub.status.idle":"2023-10-16T01:26:25.717793Z","shell.execute_reply.started":"2023-10-16T01:26:25.701704Z","shell.execute_reply":"2023-10-16T01:26:25.716266Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Import necessary libraries\nimport matplotlib.pyplot as plt\n\n# Set the figure size\nplt.figure(figsize=(20, 6))\n\n# Set the font size\nplt.rcParams['font.size'] = 14\n\n# Set the colors\ncolors = ['lightgreen', 'lightblue', 'purple', 'blue', 'yellow']\n\n# Plot the pie chart for the training set\nplt.subplot(1, 1, 1)\nplt.pie([len(HGSC), len(EC), len(CC), len(LGSC), len(MC)],\n        labels=['HGSC', 'EC', 'CC', 'LGSC', 'MC'],\n        autopct='%1.1f%%',\n        colors=colors)\nplt.title('Training Set')\n\n# Add a main title to the figure\nplt.suptitle('Distribution of HGSC, EC, CC, LGSC and MC Images in the Training data', fontsize=20, y=1.05)\n\n# Show the plot\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2023-10-16T01:26:25.719304Z","iopub.execute_input":"2023-10-16T01:26:25.720138Z","iopub.status.idle":"2023-10-16T01:26:25.936022Z","shell.execute_reply.started":"2023-10-16T01:26:25.720098Z","shell.execute_reply":"2023-10-16T01:26:25.934686Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Define the file paths for training and testing data\ntrain_data = glob.glob('/kaggle/input/UBC-OCEAN/train_images/*.png')\ntest_data = glob.glob('/kaggle/input/UBC-OCEAN/test_images/*.png')\n\n# Print the number of images in each dataset\nprint(f\"The Training Set contains: {len(train_data)} images\")\nprint(f\"The Testing Set contains: {len(test_data)} images\")\n","metadata":{"execution":{"iopub.status.busy":"2023-10-16T01:26:25.938224Z","iopub.execute_input":"2023-10-16T01:26:25.938936Z","iopub.status.idle":"2023-10-16T01:26:25.949912Z","shell.execute_reply.started":"2023-10-16T01:26:25.938899Z","shell.execute_reply":"2023-10-16T01:26:25.94832Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Calculate total counts\ntotal_train = len(train_data)\ntotal_test = len(test_data)\n\n# Set the figure size\nplt.figure(figsize=(4, 4))\n\n# Set the font size\nplt.rcParams['font.size'] = 12\n\n# Set the colors\ncolors = ['lightgreen', 'red']\n\n# Plot the pie chart for the total set\nplt.pie([total_train, total_test], labels=['Training Set', 'Testing Set'], autopct='%1.1f%%', colors=colors)\nplt.title('Distribution of Images in Training and Testing Sets')\n\n# Show the plot\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2023-10-16T01:26:25.951455Z","iopub.execute_input":"2023-10-16T01:26:25.952703Z","iopub.status.idle":"2023-10-16T01:26:26.104929Z","shell.execute_reply.started":"2023-10-16T01:26:25.952648Z","shell.execute_reply":"2023-10-16T01:26:26.103171Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\n\n# Read the CSV file into a DataFrame\ntraindf = pd.read_csv('/kaggle/input/UBC-OCEAN/train.csv', dtype=str)\n\n# Display the DataFrame\ntraindf\n","metadata":{"execution":{"iopub.status.busy":"2023-10-16T01:26:26.10709Z","iopub.execute_input":"2023-10-16T01:26:26.108725Z","iopub.status.idle":"2023-10-16T01:26:26.135461Z","shell.execute_reply.started":"2023-10-16T01:26:26.10866Z","shell.execute_reply":"2023-10-16T01:26:26.134128Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"testdf=pd.read_csv('/kaggle/input/UBC-OCEAN/test.csv',dtype=str)\ntestdf","metadata":{"execution":{"iopub.status.busy":"2023-10-16T01:26:26.137476Z","iopub.execute_input":"2023-10-16T01:26:26.139213Z","iopub.status.idle":"2023-10-16T01:26:26.161015Z","shell.execute_reply.started":"2023-10-16T01:26:26.139142Z","shell.execute_reply":"2023-10-16T01:26:26.159751Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"traindf","metadata":{"execution":{"iopub.status.busy":"2023-10-16T01:26:26.163284Z","iopub.execute_input":"2023-10-16T01:26:26.164328Z","iopub.status.idle":"2023-10-16T01:26:26.181959Z","shell.execute_reply.started":"2023-10-16T01:26:26.16427Z","shell.execute_reply":"2023-10-16T01:26:26.180684Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Directory containing the images\nimage_dir = '/kaggle/input/UBC-OCEAN/train_thumbnails'\n\n# Directory to save the extracted tiles\nsamples_dir = '/kaggle/working/centermost_tiles'\nos.makedirs(samples_dir, exist_ok=True)\n\n# Create a figure with subplots for each image\nfig, axs = plt.subplots(1, 5, figsize=(15, 5))\n\n# Define the size of the centermost tile\ntile_size = (225, 225)\n\n# Iterate over each row in the DataFrame with an index\nfor index, (_, row_data) in enumerate(unique_labels_df.iterrows()):\n    image_id = row_data['image_id']\n    image_path = os.path.join(image_dir, f'{image_id}_thumbnail.png')\n    image = imread(image_path)\n\n    # Calculate the center coordinates\n    center_x, center_y = image.shape[1] // 2, image.shape[0] // 2\n\n    # Calculate the coordinates for the top-left corner of the centermost tile\n    tile_x = center_x - tile_size[0] // 2\n    tile_y = center_y - tile_size[1] // 2\n\n    # Extract the centermost tile\n    centermost_tile = image[tile_y:tile_y + tile_size[1], tile_x:tile_x + tile_size[0]]\n\n    # Save the centermost tile to the samples directory\n    tile_filename = os.path.join(samples_dir, f'{image_id}_centermost_tile.png')\n    Image.fromarray(np.uint8(centermost_tile)).save(tile_filename)\n\n    # Display the image\n    axs[index].imshow(centermost_tile)\n    axs[index].set_title(f'Label: {row_data[\"label\"]}')\n    axs[index].axis('off')\n\n# Adjust spacing between subplots\nplt.tight_layout()\n\n# Show the plot\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2023-10-16T01:26:26.183239Z","iopub.execute_input":"2023-10-16T01:26:26.183991Z","iopub.status.idle":"2023-10-16T01:26:28.190994Z","shell.execute_reply.started":"2023-10-16T01:26:26.183957Z","shell.execute_reply":"2023-10-16T01:26:28.189817Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Challenges in Recognizing Cancerous Cells**\n\n- **Even after we zoom in on the thumbnails**, at this scale, the image classification model still won't be able to recognize the cancerous cells.\n\n- It's like training a model to recognize people with satellite images of mountains.\n\n**Why Extracting Tiles is Required for Model Training**\n\n- **Memory Efficiency**: Large images require significant memory, making tile extraction essential for working with manageable sections.\n\n- **Data Augmentation**: Extracting tiles expands the training dataset, improving model generalization.\n\n- **Feature Localization**: Focusing on image regions of interest helps models learn relevant patterns.\n\n- **Input Consistency**: Tile extraction enforces consistent input sizes, crucial for deep learning models.\n\n- **Batch Processing**: Organizing tiles into batches streamlines parallelized training.\n\n- **Resource Constraints**: Tile extraction is practical in resource-limited environments.\n\n- **Localization Tasks**: Enhances precision in object detection and classification.\n\n**Experimenting with histolab for tile extraction.**\n\n[Histolab Documentation](https://histolab.readthedocs.io/en/latest/readme.html)\n\n- While the histolab library provides a convenient way to extract tiles from large images, it's important to note that very large images can sometimes run into memory issues when using the library's built-in functions. In such cases, a custom implementation or pull request may be necessary to optimize memory consumption.\n\n**Community Input Needed**\n\n- If you've encountered success in processing tiles for Image 51346 using Histolab, I'd greatly appreciate your insights. Currently, working with this particular image seems to pose memory challenges, even with dedicated efforts to mitigate memory consumption.\n\n- Perhaps a collaborative approach, combining custom tile extraction methods with the capabilities of Histolab, could lead to optimal results.\n\n- Please consider sharing your feedback or any successful strategies in the comments.\n\n","metadata":{}},{"cell_type":"markdown","source":"**New Tile Extraction Process**\n\n*Optimizing TensorFlow's Performance*\n\nTo optimize TensorFlow's performance, it's beneficial to work with smaller images. Here, we outline the process of creating 225x225 pixel tiles from full images for training our TensorFlow model.\n\n**Tile Generation Process**\n\n1. **Centering the Tile:**\n\n   - To start, I extract a 225x225 pixel tile from each full image, ensuring that the selected tile is centered within the original image.\n\n2. **Handling Off-Center Content:**\n\n   - Keep in mind that not all images will have cell tissue in the direct center. In such cases, I will implement an algorithm to ensure the selection of relevant tiles.\n\n3. **Color Saturation Algorithm:**\n\n   - I use a saturation algorithm that randomly selects tiles until the color density/variation exceeds a predefined threshold.\n\n4. **OpenCV Structural Anomaly Detection Algorithm:**\n\n   - As suggested by Ravi Varatha, I implemented an OpenCV Structural Anomaly detection function operating on various parameters to get a probability that the image contains a cancer cell.\n\n5. **Average of the Color Saturation Score and the Cancer Probability Score for a Prediction of Optimal Training Image.**\n\nBy following this procedure, I aim to create a standardized dataset of 225x225 pixel tiles that are ideal for training our TensorFlow model. This approach ensures that even images with off-center content contribute effectively to the training process.\n\n","metadata":{}},{"cell_type":"code","source":"# Set max image pixels to avoid memory errors\nImage.MAX_IMAGE_PIXELS = None  \n\n# Parameters\n\ndef generate_tiles(image_id_for_debugging):\n    output_dir = '/kaggle/working/tiles'\n    os.makedirs(output_dir, exist_ok=True)\n    tissue_density_threshold = 0.5\n    zoom_width = 225\n    zoom_height = 225\n    shift_amount = 3000\n\n    def open_computer_vision_cancer_detection(region_array):\n        img = cv2.cvtColor(np.array(region_array), cv2.COLOR_RGB2BGR)\n        gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)\n        blurred = cv2.GaussianBlur(gray, (5, 5), 0)\n        thresh = cv2.threshold(blurred, 60, 255, cv2.THRESH_BINARY)[1]\n        params = cv2.SimpleBlobDetector_Params()\n        params.filterByArea = True\n        params.minArea = 10\n        params.maxArea = 1000\n        detector = cv2.SimpleBlobDetector_create(params)\n        keypoints = detector.detect(thresh)\n        num_blobs = len(keypoints)\n        if num_blobs > 5:\n            cancer_prob = 0.9\n        elif num_blobs > 2:\n            cancer_prob = 0.5\n        else:\n            cancer_prob = 0.1\n        return cancer_prob\n\n    def calculate_saturation(img_array):\n        img_hsv = cv2.cvtColor(img_array, cv2.COLOR_RGB2HSV)\n        sat = img_hsv[:,:,1]\n        min_sat = 50\n        colorful_pixels = sat > min_sat\n        num_colorful_pixels = np.count_nonzero(colorful_pixels)\n        total_pixels = img_array.shape[0] * img_array.shape[1]\n        pct_colorful = num_colorful_pixels / total_pixels\n        return pct_colorful\n\n    # Get the specific image information for debugging\n    row = unique_labels_df[unique_labels_df['image_id'] == image_id_for_debugging].iloc[0]\n    img_id = row['image_id']\n    image_filename = f'/kaggle/input/UBC-OCEAN/train_images/{img_id}.png'\n\n    img = Image.open(image_filename)\n\n    # Calculate the center of the image\n    center_x = row['image_width'] // 2\n    center_y = row['image_height'] // 2\n\n    # Define the desired dimensions for zooming\n    zoom_width = 225\n    zoom_height = 225\n\n    # Initialize variables for shifting the center\n    center_x_shifted = center_x\n    center_y_shifted = center_y\n\n    # Calculate the coordinates for cropping\n    x_min = center_x - zoom_width // 2\n    x_max = center_x + zoom_width // 2\n    y_min = center_y - zoom_height // 2\n    y_max = center_y + zoom_height // 2\n\n    try_iteration = 0\n    top_tissue_densities = []\n    \n    sample_locations = {}\n    sample_locations[img_id] = {}\n    \n    while try_iteration < 50:\n        \n        x_min = center_x_shifted - zoom_width // 2\n        x_max = center_x_shifted + zoom_width // 2\n        y_min = center_y_shifted - zoom_height // 2\n        y_max = center_y_shifted + zoom_height // 2\n        location = (x_min, y_min, x_max, y_max)\n        sample_locations[img_id][try_iteration] = location\n        \n        cropped_img = img.crop(location)\n        \n        sat_score = calculate_saturation(np.array(cropped_img))\n        openCV_score = open_computer_vision_cancer_detection(np.array(cropped_img))\n        top_tissue_densities.append(((sat_score + openCV_score)/2, try_iteration))\n        \n        shift_amount = 3000\n        center_x_shifted += np.random.randint(-shift_amount, shift_amount)\n        center_y_shifted += np.random.randint(-shift_amount, shift_amount)\n        \n        try_iteration += 1\n\n    top_tissue_densities.sort(reverse=True, key=lambda x: x[0])\n    top_5_tissue_densities = top_tissue_densities[:5]\n    keepers = [index for _, index in top_5_tissue_densities]\n    \n    for i, index in enumerate(keepers):\n        cropped_img = img.crop(sample_locations[img_id][index])\n        output_filename = os.path.join(output_dir, f'tile_{img_id}_{i}.png')\n        cropped_img.save(output_filename)\n\n    cropped_img.close()\n    img.close()\n    del cropped_img\n    del img\n\n    # Create a figure with subplots for each of the top 5 images\n    fig, axs = plt.subplots(1, 5, figsize=(15, 5))\n    for i, index in enumerate(keepers):\n        image_path = os.path.join(output_dir, f'tile_{img_id}_{i}.png')\n        image = imread(image_path)\n\n        axs[i].imshow(image)\n        axs[i].set_title(f\" Image: {img_id}, {unique_labels_df.loc[unique_labels_df['image_id'] == img_id, 'label'].values[0]}\")\n        axs[i].axis('off')\n    plt.tight_layout()\n    plt.show()\n\nfor image_id in unique_labels_df['image_id']:\n    generate_tiles(image_id)\n","metadata":{"execution":{"iopub.status.busy":"2023-10-16T01:26:28.192447Z","iopub.execute_input":"2023-10-16T01:26:28.1928Z","iopub.status.idle":"2023-10-16T01:32:03.817066Z","shell.execute_reply.started":"2023-10-16T01:26:28.192769Z","shell.execute_reply":"2023-10-16T01:32:03.814869Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Mitigating Non-Tissue Background in Selected Samples: A Solution**\n\nAvoiding the selection of non-tissue backgrounds is a fundamental and critical step within the scope of this project. To accomplish this, a combination of advanced algorithms and image processing techniques will be employed, with the primary objective of identifying and selecting tissue samples while efficiently excluding non-tissue backgrounds. The following strategies are being considered:\n\n1. **Color-Based Segmentation:**\n\n   *Staining:* A prevalent technique in the preparation of tissue samples is staining, which highlights specific structural features. Leveraging color-based segmentation techniques, we aim to discern and extract the stained regions corresponding to the tissue of interest. To achieve this, our approach involves fine-tuning algorithms to recognize specific color thresholds associated with the stain, effectively isolating it from the background.\n\n2. **Thresholding:**\n\n   *Brightness and Contrast:* Utilizing appropriate thresholds for brightness and contrast, we intend to emphasize tissue structures while effectively suppressing background elements. Thresholding techniques offer the capability to distinguish regions with high contrast (typically representing tissue) from those with low contrast (often indicative of background).\n\n3. **Texture Analysis:**\n\n   *Texture Filters:* The project's strategy includes exploring various texture analysis algorithms, such as Gabor filters and wavelet transforms. These algorithms are designed to identify distinctive textural patterns inherent to tissue structures. The utilization of texture analysis contributes to the differentiation of tissue regions from non-tissue areas.\n\n4. **Edge Detection:**\n\n   *Canny Edge Detection:* The application of algorithms like the Canny edge detector plays a significant role in identifying abrupt changes in intensity or color. These changes frequently correspond to the boundaries of tissue structures. By effectively detecting edges, we can outline and demarcate the tissue regions with precision.\n\nThrough the strategic amalgamation and customization of these techniques tailored to the specific requirements of our project, the objective is to develop algorithms that proficiently and accurately identify tissue regions while meticulously excluding non-tissue backgrounds in the images. Furthermore, the project is receptive to suggestions and feedback from experts and peers, with an open invitation for any novel ideas that can further enhance the selection of cancer tissue samples from high-resolution images.\n\n**Characteristics of Cancer Cells:**\n\nCancer cells manifest a diverse array of features contingent upon the specific type of cancer. Nevertheless, several common characteristics can be delineated, which may include:\n\n- **Uncontrolled Growth:** A prominent hallmark of cancer cells is their propensity to divide and multiply uncontrollably. This unbridled proliferation often culminates in the formation of tumors.\n\n- **Variation in Size and Shape:** Cancer cells frequently exhibit variations in both size and shape, occasionally surpassing or undershooting the dimensions of normal cells. Anisocytosis, denoting irregular cell size, and pleomorphism, which denotes irregular cell shape, are characteristic features.\n\n- **Lack of Differentiation:** Cancer cells may demonstrate reduced specialization or differentiation in comparison to their normal counterparts. This lack of differentiation, known as anaplasia, results in a loss of typical cell function.\n\n- **Large Nuclei:** Cancer cell nuclei tend to be larger and exhibit irregular shapes compared to those of normal cells, a condition referred to as hyperchromasia.\n\n- **Increased Nucleo-Cytoplasmic Ratio:** Cancer cells often present with a higher ratio of nucleus to cytoplasm, a phenomenon indicative of malignancy.\n\n- **Abnormal Nucleoli:** The nucleoli, small structures within the nucleus, may be larger or more numerous in cancer cells, further contributing to their atypical appearance.\n\n- **Invasive Behavior:** A defining characteristic of cancer cells is their capability to invade nearby tissues, and in certain cases, disseminate to distant organs via metastasis.\n\n- **Loss of Contact Inhibition:** Unlike normal cells that halt their division when neighboring cells are encountered, cancer cells frequently fail to exhibit contact inhibition, leading to uncontrolled proliferation even in crowded cellular environments.\n\nIt is essential to emphasize that the specific attributes of cancer cells can significantly diverge based on the type of cancer under consideration. The visual inspection of tissue samples by a pathologist under a microscope, to identify these distinctive characteristics, remains a key component in the diagnosis and classification of cancer.\n\nWith the current improvements in the algorithm's capability to procure high-quality tissue samples, the image classifier is now better equipped for enhanced performance.\n\n**Normalizing Images and Generating Batches of Tensor Image Data for Training, Validation, and Testing:**\n\nIn the domain of machine learning, whether utilizing Keras or any other framework, it is imperative to comprehend the role of and distinctions between training, validation, and test data. Each serves a distinct purpose in the model development process, and here's a succinct elucidation of their roles:\n\n1. **Training Data:**\n\n   - **Purpose:** Training data is fundamentally employed to instruct the model. It is within this dataset that the model learns the intricate patterns and distinctive features present in the data.\n   \n   - **Quantity:** Training data typically constitutes the most extensive portion of the dataset, often encompassing approximately 60-80% of the entire dataset.\n   \n   - **Role:** During the training phase, the model adjusts its parameters and fine-tunes its internal features to minimize the disparities between its predictions and the actual target values present in the training data.\n\n2. **Validation Data:**\n\n   - **Purpose:** Validation data plays an essential role during the training process. Its primary function is to gauge the model's performance and to facilitate adjustments. This dataset aids in the fine-tuning of hyperparameters and the prevention of overfitting.\n   \n   - **Quantity:** It typically represents a smaller subset of the dataset, commonly comprising approximately 10-20%.\n   \n   - **Role:** The model is not directly trained on validation data. Instead, it serves as a benchmark for evaluating how well the model generalizes to data it has not encountered during training. It is indispensable for selecting the optimal model and its corresponding hyperparameters.\n\n3. **Test Data:**\n\n   - **Purpose:** Test data, reserved for the final evaluation stage, is instrumental in assessing the model's real-world performance after training and hyperparameter optimization. This dataset simulates the conditions in which the model will operate with novel, unseen data.\n   \n   - **Quantity:** Comparable in size to the validation set, typically representing 10-20% of the entire dataset.\n   \n   - **Role:** The test data furnishes an impartial and unbiased evaluation of the model's accuracy, thereby serving as a crucial determinant of how effectively the model generalizes to fulfill its intended purpose.\n\nIn summation, the purpose of the training data is to teach the model, validation data assists in fine-tuning it, and test data substantiates its real-world performance and suitability for the task at hand. The judicious management and appropriate partitioning of these datasets are pivotal in the development of robust and accurate machine learning models.\n\n**Choosing a TensorFlow Image Classification Model for Ovarian Cancer Cell Tissue Samples:**\n\nThe selection of an apt TensorFlow image classification model is of paramount importance in this endeavor. To arrive at an informed decision, it is prudent to comprehend the underlying principles of Convolutional Neural Networks (CNNs) and their significance in the realm of cancer research. In the context of classifying cell tissue cancer, CNNs operate in the following manner:\n\n- **Convolutional Layers:** These layers are designed to detect critical features within cell tissue images, such as shapes and patterns that differentiate varying cancer types.\n\n- **Pooling Layers:** The purpose of pooling layers is to recognize these important features, regardless of their precise location within the tissue sample. This enhances the model's ability to generalize effectively.\n\n- **Flattening:** Information pertaining to the features detected is collated and organized into a numerical vector, thus rendering it amenable to further analysis.\n\n- **Fully Connected Layers:** These layers specialize in the recognition of intricate patterns that are indicative of specific cancer types. This recognition is based on the information gleaned from the detected features.\n\n- **Output Layer:** The final layer is responsible for making predictions based on the patterns learned by the model. It provides classifications for different types of ovarian cancer based on the features identified.\n\nNow, let's delve into the rationale behind the choice of Keras for this specific task and why it is a fitting tool:\n\n1. **User-Friendly Interface:** Keras offers a user-friendly and intuitive interface, making it accessible to individuals ranging from beginners to experts in medical image classification. Its straightforward design streamlines the model development process.\n\n2. **Modularity:** Keras is renowned for its modularity, allowing for flexible model construction. This adaptability enables the incorporation of domain-specific knowledge specific to ovarian cancers.\n\n3. **Compatibility:** Keras seamlessly adapts to various data formats and sources, including medical imaging data. It provides an interface that can readily handle the complexities of medical image analysis.\n\n4. **Community and Documentation:** Keras boasts a robust and supportive community, which is particularly advantageous for a specialized task like cancer classification. Furthermore, extensive documentation is available, ensuring that developers have access to valuable resources and expertise.\n\n5. **Flexibility:** Keras is highly flexible, allowing for the customization of models to focus on the unique characteristics of ovarian cancer cell tissue images. This adaptability is indispensable in accommodating the idiosyncrasies of the dataset.\n\n6. **Integration with TensorFlow:** As an integral component of the TensorFlow ecosystem, Keras combines the prowess of deep learning with the versatility of TensorFlow. This synergy offers a potent platform for developing and deploying deep learning models.\n\n7. **Transfer Learning:** Keras simplifies the integration of pre-trained models into the workflow. This is a valuable time-saving feature, as it obviates the need to build models from scratch, allowing for rapid deployment.\n\n8. **Abstraction from Low-Level Details:** Keras abstracts the complexity of low-level technicalities, enabling developers to focus on fine-tuning model architecture and achieving desired classification outcomes.\n\n9. **Rapid Prototyping:** Keras is ideal for rapid model experimentation, offering an environment that encourages the exploration of various network architectures. This promotes the discovery of the best-performing model configuration.\n\n10. **Scalability:** Keras is highly scalable and can accommodate experiments of varying scales. This includes small-scale projects as well as large-scale analysis of extensive ovarian cancer image datasets.\n\nWith this understanding of Keras' merits, the choice of the Keras framework for this specific task is clear. It aligns with the requirements of the project, which demand a flexible, accessible, and highly customizable platform for the development of deep learning models tailored to the challenges of ovarian cancer cell tissue image classification.\n\nNow, let's discuss the selection of an appropriate model within Keras and the utilization of the Sequential model. Choosing the right TensorFlow image classification model is pivotal for the success of this project. The following pre-trained models are particularly relevant to medical image classification and warrant consideration:\n\n1. **ResNet:** ResNet is distinguished by its depth and exceptional performance. It excels in identifying subtle patterns within cancer cell images, a feature that could prove invaluable in our context.\n\n2. **Inception:** The Inception model is renowned for its efficient use of computational resources. It excels in detecting unique features that are indicative of different cancer types, a quality that aligns with the project's objectives.\n\n3. **MobileNet:** Balancing speed and accuracy, MobileNet is an excellent choice for analyzing large volumes of cell tissue samples. Its ability to achieve a favorable trade-off between computational efficiency and classification accuracy is highly beneficial.\n\n4. **EfficientNet:** As a state-of-the-art performer, EfficientNet offers the promise of achieving high classification accuracy for various ovarian cancer types. Its proficiency in extracting subtle features is advantageous.\n\n5. **VGGNet:** VGGNet provides a straightforward yet effective approach. It is an excellent starting point, especially for educational purposes or as a baseline model, serving as a foundation for more advanced models.\n\nBy utilizing the Keras Sequential model, we can fine-tune the selected model on ovarian cancer cell tissue images. The next steps include evaluating the model's performance on a validation dataset and fine-tuning hyperparameters within Keras. This optimization process is specifically geared toward enhancing the model's efficacy in classifying ovarian cancer types based on cell tissue samples.\n\nIn conclusion, the combination of advanced image processing techniques, profound knowledge of cancer cell characteristics, strategic data partitioning, and the choice of an appropriate TensorFlow image classification model within the Keras framework will be instrumental in achieving the project's goals and advancing the diagnosis and classification of ovarian cancer.","metadata":{}},{"cell_type":"code","source":"df.to_csv(\"submission.csv\", index=False)","metadata":{"execution":{"iopub.status.busy":"2023-10-16T01:32:03.81959Z","iopub.execute_input":"2023-10-16T01:32:03.820251Z","iopub.status.idle":"2023-10-16T01:32:03.837493Z","shell.execute_reply.started":"2023-10-16T01:32:03.820186Z","shell.execute_reply":"2023-10-16T01:32:03.836299Z"},"trusted":true},"execution_count":null,"outputs":[]}]}