{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# <mark>UBC Ovarian Cancer Subtype Classification and Outlier Detection (UBC-OCEAN) - EDA</mark>\n<span style=\"font-size:22px;color:purple\"> Thank you for having a look at my notebook - advice and feedback always welcomed!</span>\n\n\n<div class=\"alert alert-block alert-info\" style=\"font-size:14px; font-family:verdana;\">\n    📌 Dataset Link: <a href=\"https://www.kaggle.com/competitions/UBC-OCEAN/data\">https://www.kaggle.com/competitions/UBC-OCEAN/data</a>\n</div>\n\n\n## **Exploratory Data Analysis (EDA)** \nEDA is a crucial step in understanding and preparing your data for any data analysis or machine learning project, including UBC Ovarian Cancer Subtype Classification and Outlier Detection (UBC-OCEAN). Here's a step-by-step guide on how to perform an EDA for this dataset:\n\n**Overview**\n\nThe goal of the UBC Ovarian Cancer subtypE clAssification and outlier detectioN (UBC-OCEAN) competition is to classify ovarian cancer subtypes. You will build a model trained on the world's most extensive ovarian cancer dataset of histopathology images obtained from more than 20 medical centers.\n\n\n**Data Collection:**\n\nBegin by obtaining the UBC-OCEAN dataset, which should include information on ovarian cancer subtypes and possibly outlier detection data. Ensure you have a clear understanding of the dataset's structure and the meaning of each variable.\nData Loading:\n\nImport the dataset into your preferred data analysis environment, such as Python with libraries like pandas, numpy, and matplotlib/seaborn for visualization.\n\n**Data Loading:**\n\nImport the dataset into your preferred data analysis environment, such as Python with libraries like pandas, numpy, and matplotlib/seaborn for visualization.\n\n\n\n\n\n### **Initial Exploration:**\n\n**1 - Start by examining the basic characteristics of the data:**\n\n    Check the first few rows using df.head().\n    Check the data types and missing values using df.info().\n    Calculate basic statistics using df.describe().\n    Data Cleaning:\n\n**2 - Handle missing values, outliers, and duplicates:**\n\n    Use techniques like imputation for missing values.\n    Identify and deal with outliers appropriately.\n    Remove duplicate rows if necessary.\n    \n**3 - Data Visualization:**\n\n    Create visualizations to gain insights into the data:\n    Histograms and box plots for numerical features.\n    Bar plots for categorical features.\n    Correlation matrix and scatter plots to understand relationships between variables.\n    \n**4 - Feature Analysis:**\n\n    Explore relationships between features and the target variable(s) for classification and outlier detection.\n    Visualize how different features vary across different subtypes or classes.\n    Use box plots, violin plots, or swarm plots to compare feature distributions.\n\n**5 - Outlier Detection:**\n\n    If your dataset contains information related to outlier detection, perform a dedicated EDA for this aspect:\n    Visualize outliers using scatter plots or box plots.\n    Apply statistical methods or machine learning techniques to identify outliers.\n\n**6 - Dimensionality Reduction (optional):**\n\n    If the dataset has many features, consider dimensionality reduction techniques like Principal Component Analysis (PCA) to reduce the number of variables while preserving important information.\n\n**8 - Summary and Insights:**\n\n    Summarize your findings from the EDA, including any patterns, trends, or anomalies observed.\n    Document any data preprocessing steps applied.\n\n**7 - Next Steps:**\n\n    Based on your EDA findings, plan your next steps, which may include feature engineering, model selection, and further data preprocessing.\n\n\nRemember that EDA is an iterative process, and you may need to revisit these steps as you delve deeper into the dataset and develop your machine learning or data analysis models.","metadata":{}},{"cell_type":"code","source":"%%capture \n# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-10-07T04:28:57.866043Z","iopub.execute_input":"2023-10-07T04:28:57.866399Z","iopub.status.idle":"2023-10-07T04:28:58.184115Z","shell.execute_reply.started":"2023-10-07T04:28:57.866368Z","shell.execute_reply":"2023-10-07T04:28:58.18322Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#!pip install scikit-image","metadata":{"execution":{"iopub.status.busy":"2023-10-07T04:28:58.718303Z","iopub.execute_input":"2023-10-07T04:28:58.719155Z","iopub.status.idle":"2023-10-07T04:28:58.72344Z","shell.execute_reply.started":"2023-10-07T04:28:58.719122Z","shell.execute_reply":"2023-10-07T04:28:58.722185Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nimport matplotlib.pyplot as plt\nfrom skimage import io\nimport os\nimport seaborn as sns\nimport cv2\nimport random\nimport os\nimport glob","metadata":{"execution":{"iopub.status.busy":"2023-10-07T04:28:59.19012Z","iopub.execute_input":"2023-10-07T04:28:59.192727Z","iopub.status.idle":"2023-10-07T04:29:00.049547Z","shell.execute_reply.started":"2023-10-07T04:28:59.192691Z","shell.execute_reply":"2023-10-07T04:29:00.048626Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load train data\ntrain_df = pd.read_csv('/kaggle/input/UBC-OCEAN/train.csv')\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-10-07T04:29:00.384184Z","iopub.execute_input":"2023-10-07T04:29:00.384497Z","iopub.status.idle":"2023-10-07T04:29:00.402419Z","shell.execute_reply.started":"2023-10-07T04:29:00.384471Z","shell.execute_reply":"2023-10-07T04:29:00.401498Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['label'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2023-10-07T04:29:01.540818Z","iopub.execute_input":"2023-10-07T04:29:01.541157Z","iopub.status.idle":"2023-10-07T04:29:01.549162Z","shell.execute_reply.started":"2023-10-07T04:29:01.54113Z","shell.execute_reply":"2023-10-07T04:29:01.548206Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(train_df.shape)","metadata":{"execution":{"iopub.status.busy":"2023-10-07T04:29:01.986545Z","iopub.execute_input":"2023-10-07T04:29:01.987058Z","iopub.status.idle":"2023-10-07T04:29:01.992905Z","shell.execute_reply.started":"2023-10-07T04:29:01.98702Z","shell.execute_reply":"2023-10-07T04:29:01.991966Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.isna().sum().sum()","metadata":{"execution":{"iopub.status.busy":"2023-10-07T04:29:02.309464Z","iopub.execute_input":"2023-10-07T04:29:02.310132Z","iopub.status.idle":"2023-10-07T04:29:02.317113Z","shell.execute_reply.started":"2023-10-07T04:29:02.310103Z","shell.execute_reply":"2023-10-07T04:29:02.316097Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <span style=\"color:purple\">No null values present</span>\n​\n<div class=\"alert alert-block alert-info\">\n<b>Good:</b> now to move onto the next steps </div>","metadata":{}},{"cell_type":"code","source":"# list of columns we want the distributions for\ncolumns = [col for col in train_df.columns if col!='label']\n\n# loop to iterate over each column\nfor col in columns:\n    \n    # subplot for 3 columns (3 plots)\n    fig, axs = plt.subplots(figsize=(15,5), ncols=3)\n    \n    # 1st plot - distribution of the sample dataset\n    sns.histplot(data=train_df, x=col, kde=True, ax=axs[0])\n    axs[0].set_title('Sample Distribution')\n    \n    # 2nd plot - distribution of the selected column where the outcome is 1 (has diabetes)\n    sns.histplot(data=train_df[train_df['label']==\"HGSC\"], x=col, kde=True, ax=axs[1], color='orange')\n    axs[1].set_title('label - HGSC')\n    \n    # 3rd plot - distribution of the selected column where the outcome is 0 (doesn't have diabetes)\n    sns.histplot(data=train_df[train_df['label']==\"LGSC\"], x=col, kde=True, ax=axs[2], color='green')\n    axs[2].set_title('label - LGSC')\n    \n    # showing the plots\n    plt.tight_layout()\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-07T04:29:03.220734Z","iopub.execute_input":"2023-10-07T04:29:03.221105Z","iopub.status.idle":"2023-10-07T04:29:06.235538Z","shell.execute_reply.started":"2023-10-07T04:29:03.221079Z","shell.execute_reply":"2023-10-07T04:29:06.234531Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.info()","metadata":{"execution":{"iopub.status.busy":"2023-10-07T04:29:06.237446Z","iopub.execute_input":"2023-10-07T04:29:06.238389Z","iopub.status.idle":"2023-10-07T04:29:06.251253Z","shell.execute_reply.started":"2023-10-07T04:29:06.23835Z","shell.execute_reply":"2023-10-07T04:29:06.249786Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(train_df.image_id.is_unique)\nprint(train_df.label.is_unique)","metadata":{"execution":{"iopub.status.busy":"2023-10-07T04:29:06.252693Z","iopub.execute_input":"2023-10-07T04:29:06.254058Z","iopub.status.idle":"2023-10-07T04:29:06.264466Z","shell.execute_reply.started":"2023-10-07T04:29:06.25402Z","shell.execute_reply":"2023-10-07T04:29:06.263317Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Exploring image dimensions refers to understanding and working with the various aspects of an image's size and resolution. In the context of digital images, there are several key dimensions and attributes to consider:\n\n1. **Resolution**: Resolution refers to the number of pixels (individual points of color) contained in an image. It is usually expressed in terms of pixels per inch (PPI) or dots per inch (DPI). Higher resolution images have more detail and are suitable for printing, while lower resolution images may be used for web display or on-screen viewing.\n\n2. **Pixel Dimensions**: Pixel dimensions specify the width and height of an image in pixels. For example, an image might have dimensions of 1920x1080 pixels, which means it is 1920 pixels wide and 1080 pixels tall. This is commonly used for specifying the size of digital images.\n\n3. **Aspect Ratio**: The aspect ratio is the ratio of an image's width to its height. Common aspect ratios include 4:3 (standard television), 16:9 (widescreen television), and 1:1 (square). Maintaining the correct aspect ratio is important to prevent image distortion.\n\n4. **Physical Dimensions**: Physical dimensions refer to the size of an image when printed or displayed in the physical world. This is determined by both the pixel dimensions and the resolution. For example, an image with dimensions of 3000x2000 pixels at 300 DPI will be 10x6.67 inches when printed.\n\n5. **File Size**: The file size of an image is measured in bytes or kilobytes (KB), megabytes (MB), etc. It depends on factors such as the color depth, compression, and pixel dimensions. Larger images with more detail tend to have larger file sizes.\n\n6. **Color Depth**: Color depth, also known as bit depth, determines the number of colors a pixel can represent. Common color depths include 8-bit (256 colors), 24-bit (true color), and 32-bit (true color with alpha channel for transparency).\n\n7. **DPI vs. PPI**: DPI (dots per inch) is often used in the context of printing, indicating how many ink dots a printer can produce in a linear inch. PPI (pixels per inch) is used for screen displays, representing the number of pixels in an inch of screen space.\n\n8. **Scaling**: Scaling an image involves resizing it to different dimensions. You can scale an image up (enlargement) or down (reduction). Be aware that scaling too much can lead to a loss of image quality, especially when making an image larger.\n\n9. **Cropping**: Cropping involves cutting out a portion of an image to focus on a specific area. This changes the pixel dimensions and aspect ratio of the image.\n\n10. **Compression**: Image compression reduces file size by removing redundant or less important data. It can be lossless (no quality loss) or lossy (some quality loss). Common image formats like JPEG use lossy compression.\n\nUnderstanding and managing these image dimensions and attributes is essential for various purposes, including graphic design, photography, web development, and printing. Depending on your specific needs, you may need to adjust these dimensions and attributes accordingly to achieve the desired result.","metadata":{}},{"cell_type":"code","source":"train_df[['image_height', 'image_width']].describe()","metadata":{"execution":{"iopub.status.busy":"2023-10-07T04:29:06.267177Z","iopub.execute_input":"2023-10-07T04:29:06.267845Z","iopub.status.idle":"2023-10-07T04:29:06.292419Z","shell.execute_reply.started":"2023-10-07T04:29:06.267809Z","shell.execute_reply":"2023-10-07T04:29:06.291064Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(train_df.image_id.shape[0])\nprint(len(os.listdir('/kaggle/input/UBC-OCEAN/train_images')))\nprint(len(os.listdir('/kaggle/input/UBC-OCEAN/train_thumbnails')))","metadata":{"execution":{"iopub.status.busy":"2023-10-07T04:29:06.294074Z","iopub.execute_input":"2023-10-07T04:29:06.294467Z","iopub.status.idle":"2023-10-07T04:29:06.302622Z","shell.execute_reply.started":"2023-10-07T04:29:06.29443Z","shell.execute_reply":"2023-10-07T04:29:06.301508Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import imageio\n","metadata":{"execution":{"iopub.status.busy":"2023-10-07T04:29:06.305455Z","iopub.execute_input":"2023-10-07T04:29:06.305813Z","iopub.status.idle":"2023-10-07T04:29:06.310238Z","shell.execute_reply.started":"2023-10-07T04:29:06.30578Z","shell.execute_reply":"2023-10-07T04:29:06.309269Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\"\"\"import numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nfrom tqdm.notebook import tqdm\n\nimport albumentations as A\nimport matplotlib.image as mpimg\nimport imageio\nimport scipy.ndimage as ndi\"\"\"","metadata":{"execution":{"iopub.status.busy":"2023-10-07T04:29:06.923547Z","iopub.execute_input":"2023-10-07T04:29:06.923902Z","iopub.status.idle":"2023-10-07T04:29:06.930568Z","shell.execute_reply.started":"2023-10-07T04:29:06.923872Z","shell.execute_reply":"2023-10-07T04:29:06.929575Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import warnings\nwarnings.filterwarnings('ignore')","metadata":{"execution":{"iopub.status.busy":"2023-10-07T04:29:08.879672Z","iopub.execute_input":"2023-10-07T04:29:08.880834Z","iopub.status.idle":"2023-10-07T04:29:08.885526Z","shell.execute_reply.started":"2023-10-07T04:29:08.880791Z","shell.execute_reply":"2023-10-07T04:29:08.884567Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.label.value_counts()","metadata":{"execution":{"iopub.status.busy":"2023-10-07T04:29:09.420348Z","iopub.execute_input":"2023-10-07T04:29:09.421455Z","iopub.status.idle":"2023-10-07T04:29:09.429531Z","shell.execute_reply.started":"2023-10-07T04:29:09.421415Z","shell.execute_reply":"2023-10-07T04:29:09.428525Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"HGSC = train_df[train_df['label']==\"HGSC\"]\nEC = train_df[train_df['label']==\"EC\"]\nCC = train_df[train_df['label']==\"CC\"]\nLGSC = train_df[train_df['label']==\"LGSC\"]\nMC = train_df[train_df['label']==\"MC\"]","metadata":{"execution":{"iopub.status.busy":"2023-10-07T04:29:09.988545Z","iopub.execute_input":"2023-10-07T04:29:09.989669Z","iopub.status.idle":"2023-10-07T04:29:09.99842Z","shell.execute_reply.started":"2023-10-07T04:29:09.989615Z","shell.execute_reply":"2023-10-07T04:29:09.99745Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Set the figure size\nplt.figure(figsize=(20, 6))\n\n# Set the font size\nplt.rcParams['font.size'] = 14\n\n# Set the colors\ncolors = ['lightgreen', 'lightblue', 'purple', 'blue', 'yellow']\n\n# Plot the pie chart for the training set\nplt.subplot(1, 1, 1)\nplt.pie([len(HGSC), len(EC), len(CC), len(LGSC), len(MC)], labels=['HGSC', 'EC', 'CC', 'LGSC', 'MC'], autopct='%1.1f%%', colors=colors)\nplt.title('Training Set')\n\n\n# Add a main title to the figure\nplt.suptitle('Distribution of HGSC, EC, CC, LGSC and MC Images in the Training data', fontsize=20, y=1.05)\n\n# Show the plot\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-07T04:29:10.615834Z","iopub.execute_input":"2023-10-07T04:29:10.616888Z","iopub.status.idle":"2023-10-07T04:29:10.808398Z","shell.execute_reply.started":"2023-10-07T04:29:10.616847Z","shell.execute_reply":"2023-10-07T04:29:10.807548Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Examine the Data Split of training and testing data\ntrain_data = glob.glob('/kaggle/input/UBC-OCEAN/train_images/*.png')\ntest_data = glob.glob('/kaggle/input/UBC-OCEAN/test_images/*.png')\n\nprint(f\"The Training Set contains: {len(train_data)} images\")\nprint(f\"The Testing Set contains: {len(test_data)} images\")","metadata":{"execution":{"iopub.status.busy":"2023-10-07T04:29:11.322692Z","iopub.execute_input":"2023-10-07T04:29:11.323062Z","iopub.status.idle":"2023-10-07T04:29:11.333594Z","shell.execute_reply.started":"2023-10-07T04:29:11.323035Z","shell.execute_reply":"2023-10-07T04:29:11.332508Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Calculate total counts\ntotal_train = len(train_data)\ntotal_test = len(test_data)\n\n# Set the figure size\nplt.figure(figsize=(4, 4))\n\n# Set the font size\nplt.rcParams['font.size'] = 12\n\n# Set the colors\ncolors = ['lightgreen', 'red']\n\n# Plot the pie chart for the total set\nplt.pie([total_train, total_test], labels=['Training Set', 'Testing Set'], autopct='%1.1f%%', colors=colors)\nplt.title('Distribution of Images in Training and Testing Sets')\n\n# Show the plot\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-07T04:29:12.536237Z","iopub.execute_input":"2023-10-07T04:29:12.536587Z","iopub.status.idle":"2023-10-07T04:29:12.663023Z","shell.execute_reply.started":"2023-10-07T04:29:12.536556Z","shell.execute_reply":"2023-10-07T04:29:12.662058Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Data Visualization 📈 \n\n### Randomly Visualize Images","metadata":{}},{"cell_type":"code","source":"# sample training image\n\nio.imshow('/kaggle/input/UBC-OCEAN/train_thumbnails/10642_thumbnail.png')","metadata":{"execution":{"iopub.status.busy":"2023-10-07T04:29:13.936443Z","iopub.execute_input":"2023-10-07T04:29:13.936803Z","iopub.status.idle":"2023-10-07T04:29:15.004396Z","shell.execute_reply.started":"2023-10-07T04:29:13.936767Z","shell.execute_reply":"2023-10-07T04:29:15.00352Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"HGSC = train_df[train_df['label']==\"HGSC\"]\nEC = train_df[train_df['label']==\"EC\"]\nCC = train_df[train_df['label']==\"CC\"]\nLGSC = train_df[train_df['label']==\"LGSC\"]\nMC = train_df[train_df['label']==\"MC\"]","metadata":{"execution":{"iopub.status.busy":"2023-10-07T04:29:16.291608Z","iopub.execute_input":"2023-10-07T04:29:16.292722Z","iopub.status.idle":"2023-10-07T04:29:16.301382Z","shell.execute_reply.started":"2023-10-07T04:29:16.29268Z","shell.execute_reply":"2023-10-07T04:29:16.300336Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!pip install torch-summary\n!pip install torch-lr-finder","metadata":{"execution":{"iopub.status.busy":"2023-10-07T04:29:17.608888Z","iopub.execute_input":"2023-10-07T04:29:17.609355Z","iopub.status.idle":"2023-10-07T04:29:17.615608Z","shell.execute_reply.started":"2023-10-07T04:29:17.609316Z","shell.execute_reply":"2023-10-07T04:29:17.614602Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os\nimport numpy as np\nimport cv2\nimport matplotlib.pyplot as plt\nimport matplotlib.image as mpimg\n%matplotlib inline\nfrom PIL import Image\nfrom IPython.display import display\nimport torch\nimport torch.nn as nn\nfrom torch.utils.data import DataLoader\nimport torch.nn.functional as F\nfrom torchvision import datasets, transforms, models\nfrom torch.optim.lr_scheduler import StepLR\nfrom torchsummary import summary\nfrom tqdm import tqdm\nimport torchvision.models as models\nimport PIL\nfrom torchvision.models import resnet50, ResNet50_Weights\nimport timm\nfrom torch.utils.data import Dataset\nimport glob","metadata":{"execution":{"iopub.status.busy":"2023-10-07T04:29:18.094058Z","iopub.execute_input":"2023-10-07T04:29:18.094416Z","iopub.status.idle":"2023-10-07T04:29:18.100919Z","shell.execute_reply.started":"2023-10-07T04:29:18.094385Z","shell.execute_reply":"2023-10-07T04:29:18.099896Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"torch.manual_seed(0)","metadata":{"execution":{"iopub.status.busy":"2023-10-07T04:29:18.832141Z","iopub.execute_input":"2023-10-07T04:29:18.833315Z","iopub.status.idle":"2023-10-07T04:29:18.840417Z","shell.execute_reply.started":"2023-10-07T04:29:18.833269Z","shell.execute_reply":"2023-10-07T04:29:18.839281Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_path = '/kaggle/input/UBC-OCEAN/train_images","metadata":{"execution":{"iopub.status.busy":"2023-10-07T04:29:19.323933Z","iopub.execute_input":"2023-10-07T04:29:19.324275Z","iopub.status.idle":"2023-10-07T04:29:19.330141Z","shell.execute_reply.started":"2023-10-07T04:29:19.324248Z","shell.execute_reply":"2023-10-07T04:29:19.329106Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class_name = os.listdir(data_path)\nprint(class_name)","metadata":{"execution":{"iopub.status.busy":"2023-10-07T04:29:19.594894Z","iopub.execute_input":"2023-10-07T04:29:19.595233Z","iopub.status.idle":"2023-10-07T04:29:19.601301Z","shell.execute_reply.started":"2023-10-07T04:29:19.595204Z","shell.execute_reply":"2023-10-07T04:29:19.600409Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class_map = {\"HGSC\" : 0, \"EC\": 1, \"CC\": 2, \"LGSC\": 3, \"MC\": 4}\nrev_class_map = {0 : \"HGSC\", 1 : \"EC\", 2 : \"CC\", 3 : \"LGSC\", 4 : \"MC\"}","metadata":{"execution":{"iopub.status.busy":"2023-10-07T04:29:20.313582Z","iopub.execute_input":"2023-10-07T04:29:20.314561Z","iopub.status.idle":"2023-10-07T04:29:20.320682Z","shell.execute_reply.started":"2023-10-07T04:29:20.314531Z","shell.execute_reply":"2023-10-07T04:29:20.319571Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from tensorflow.keras.models import Sequential\nfrom tensorflow.keras.layers import Dense, Activation,Conv2D, Flatten, Dropout, MaxPooling2D, BatchNormalization\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator\nfrom keras import regularizers, optimizers\nimport os\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport pandas as pd","metadata":{"execution":{"iopub.status.busy":"2023-10-07T04:29:22.032649Z","iopub.execute_input":"2023-10-07T04:29:22.033374Z","iopub.status.idle":"2023-10-07T04:29:24.66109Z","shell.execute_reply.started":"2023-10-07T04:29:22.033342Z","shell.execute_reply":"2023-10-07T04:29:24.660094Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"traindf=pd.read_csv('/kaggle/input/UBC-OCEAN/train.csv',dtype=str)\ntraindf","metadata":{"execution":{"iopub.status.busy":"2023-10-07T04:29:24.662786Z","iopub.execute_input":"2023-10-07T04:29:24.664007Z","iopub.status.idle":"2023-10-07T04:29:24.681771Z","shell.execute_reply.started":"2023-10-07T04:29:24.663969Z","shell.execute_reply":"2023-10-07T04:29:24.680813Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"testdf=pd.read_csv('/kaggle/input/UBC-OCEAN/test.csv',dtype=str)\ntestdf","metadata":{"execution":{"iopub.status.busy":"2023-10-07T04:29:27.900465Z","iopub.execute_input":"2023-10-07T04:29:27.901205Z","iopub.status.idle":"2023-10-07T04:29:27.91325Z","shell.execute_reply.started":"2023-10-07T04:29:27.901168Z","shell.execute_reply":"2023-10-07T04:29:27.912185Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"datagen=ImageDataGenerator(rescale=1./255.,validation_split=0.25)\ntrain_generator=datagen.flow_from_dataframe(\ndataframe=traindf,\ndirectory=\"/kaggle/input/UBC-OCEAN/train_images/\",\nx_col=\"image_id\",\ny_col=\"label\",\nsubset=\"training\",\nbatch_size=32,\nseed=42,\nshuffle=True,\nclass_mode=\"categorical\",\ntarget_size=(100,100))","metadata":{"execution":{"iopub.status.busy":"2023-10-07T04:29:28.335171Z","iopub.execute_input":"2023-10-07T04:29:28.335517Z","iopub.status.idle":"2023-10-07T04:29:28.347479Z","shell.execute_reply.started":"2023-10-07T04:29:28.335488Z","shell.execute_reply":"2023-10-07T04:29:28.346542Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def append_ext(fn):\n    return fn+\".png\"\n\ndef append_ext_thum(fn):\n    return fn+\"_thumbnail.png\"\n\n\ntraindf[\"image_id_path\"]=traindf[\"image_id\"].apply(append_ext)\ntraindf[\"image_id_path_thum\"]=traindf[\"image_id\"].apply(append_ext_thum)\n\n\ntestdf[\"image_id_path\"]=testdf[\"image_id\"].apply(append_ext)\ntestdf[\"image_id_path_thum\"]=testdf[\"image_id\"].apply(append_ext_thum)\n","metadata":{"execution":{"iopub.status.busy":"2023-10-07T04:29:29.388375Z","iopub.execute_input":"2023-10-07T04:29:29.388727Z","iopub.status.idle":"2023-10-07T04:29:29.398291Z","shell.execute_reply.started":"2023-10-07T04:29:29.388696Z","shell.execute_reply":"2023-10-07T04:29:29.396963Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"traindf","metadata":{"execution":{"iopub.status.busy":"2023-10-07T04:29:30.504204Z","iopub.execute_input":"2023-10-07T04:29:30.504853Z","iopub.status.idle":"2023-10-07T04:29:30.53068Z","shell.execute_reply.started":"2023-10-07T04:29:30.504812Z","shell.execute_reply":"2023-10-07T04:29:30.529801Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"testdf","metadata":{"execution":{"iopub.status.busy":"2023-10-07T04:29:30.898884Z","iopub.execute_input":"2023-10-07T04:29:30.899564Z","iopub.status.idle":"2023-10-07T04:29:30.911358Z","shell.execute_reply.started":"2023-10-07T04:29:30.899529Z","shell.execute_reply":"2023-10-07T04:29:30.910509Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"testdf['Image_path'] = [os.path.join('/kaggle/input/UBC-OCEAN/test_images', image) for image in testdf['image_id_path']]\ntestdf['Image_path_thumbnails'] = [os.path.join('/kaggle/input/UBC-OCEAN/test_thumbnails', image) for image in testdf['image_id_path_thum']]\ntestdf","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"traindf['Image_path'] = [os.path.join('/kaggle/input/UBC-OCEAN/train_images', image) for image in traindf['image_id_path']]\ntraindf['Image_path_thumbnails'] = [os.path.join('/kaggle/input/UBC-OCEAN/train_thumbnails', image) for image in traindf['image_id_path_thum']]\ntraindf","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"full_path_random = np.random.choice(traindf['Image_path_thumbnails'],5)\nfull_path_random","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def image_viewer(dataset, index, ax):\n    image_path =  dataset['Image_path_thumbnails'][index]\n    image      =  Image.open(image_path)\n    ax.imshow(image)\n    \ndef plot_some_images(dataset, title):\n    fig, axs = plt.subplots(nrows = 1,ncols = 2,figsize=(20,8))\n    for ind, ax in enumerate(axs.flat):\n            index = random.randrange(len(dataset))\n            image_viewer(dataset, index, ax)\n            ax.set_title(dataset['label'][index], fontsize = 8)\n            ax.axis('off')\n            fig.suptitle(title, fontsize = 15)\n    plt.show()\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_some_images(traindf, 'Trainig Images')\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\nfrom sklearn.metrics import confusion_matrix, classification_report\nfrom sklearn.utils.class_weight import compute_class_weight","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class_weights = compute_class_weight(class_weight = \"balanced\",\n                                     classes= np.unique(traindf['label']),\n                                     y= traindf['label'])\n\nclasses = (np.unique(traindf['label']))\nclass_weights_forplot = dict(zip(classes, class_weights))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"classes","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class_weights_forplot","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class_weights = dict(zip(range(43), class_weights))\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **Model Building**\n","metadata":{}},{"cell_type":"code","source":"import tensorflow as tf\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_generator = tf.keras.preprocessing.image.ImageDataGenerator(\n    preprocessing_function=tf.keras.applications.efficientnet.preprocess_input,\n)\ntest_generator = tf.keras.preprocessing.image.ImageDataGenerator(\n    preprocessing_function=tf.keras.applications.efficientnet.preprocess_input\n)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_images = train_generator.flow_from_dataframe(\n    dataframe=traindf,\n    x_col='Image_path_thumbnails',\n    y_col= 'label',\n    target_size=(256, 256),\n    color_mode='grayscale',\n    class_mode=\"categorical\",\n    batch_size=64,\n    shuffle=True,\n    seed=210,\n)\n\ntest_images = test_generator.flow_from_dataframe(\n    dataframe=traindf[0:10],\n    x_col='Image_path_thumbnails',\n    y_col= 'label',\n    target_size=(256, 256),\n    class_mode=\"categorical\",\n    color_mode='grayscale',\n    batch_size=64,\n    shuffle=False\n)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"trans_arc = tf.keras.applications.EfficientNetB0(weights = \"imagenet\", include_top = False,\n                         input_shape=(256, 256,3), pooling='max')\nfor l in trans_arc.layers:\n    l.trainable = False\ninputs = trans_arc.input\nflatten = trans_arc.output\n\nx = tf.keras.layers.Dense(256, activation='relu')(flatten)\nx = tf.keras.layers.BatchNormalization()(x)\nx = tf.keras.layers.Dropout(0.3)(x)\n\nx = tf.keras.layers.Dense(128, activation='relu')(x)\nx = tf.keras.layers.BatchNormalization()(x)\n\noutputs = tf.keras.layers.Dense(1, activation='sigmoid')(x)\n\n\nmodel = tf.keras.Model(inputs=inputs, outputs=outputs)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.summary()\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"os.mkdir(\"/kaggle/working/checkpoints\")\ncb_csvlogger = tf.keras.callbacks.CSVLogger(\n                                            filename='/content/checkpoints/training_log.csv',\n                                            separator=',',\n                                            append=False)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"loss = [tf.keras.losses.binary_crossentropy]\n\ninitial_learning_rate = 0.005\n\nlr_schedule = tf.keras.optimizers.schedules.ExponentialDecay(\n    initial_learning_rate,\n    decay_steps=82,\n    decay_rate=0.9,\n    staircase=True)\n\noptimizer = tf.keras.optimizers.Adam(\n    learning_rate= lr_schedule,\n    beta_1=0.9,\n    beta_2=0.999,\n    epsilon=1e-07,\n)\nmetrics= ['accuracy']\n\nmodel.compile(\n    optimizer=optimizer,\n    loss= loss,\n    metrics=metrics\n    )","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_images","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from tensorflow.python.client import device_lib\nprint(device_lib.list_local_devices())","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# https://www.tensorflow.org/guide/gpu?hl=pt-br\nwith tf.device(\"/device:GPU:0\"):\n\n    history = model.fit(\n      train_images,\n      validation_data=test_images,\n      verbose = True,\n      epochs=5,\n      class_weight = class_weights,\n      callbacks=[\n          tf.keras.callbacks.LearningRateScheduler(lr_schedule),\n          tf.keras.callbacks.EarlyStopping(\n              monitor='val_loss',\n              patience=7,\n              restore_best_weights=True),\n          #cb_csvlogger\n      ]\n  )","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# list all data in history\nprint(history.history.keys())\n# summarize history for accuracy\nplt.plot(history.history['accuracy'])\nplt.plot(history.history['val_accuracy'])  # RAISE ERROR\nplt.title('model accuracy')\nplt.ylabel('accuracy')\nplt.xlabel('epoch')\nplt.legend(['train', 'validation'], loc='upper left')\nplt.show()\n# summarize history for loss\nplt.plot(history.history['loss'])\nplt.plot(history.history['val_loss']) #RAISE ERROR\nplt.title('model loss')\nplt.ylabel('loss')\nplt.xlabel('epoch')\nplt.legend(['train', 'validation'], loc='upper left')\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **Model Evaluation**","metadata":{}},{"cell_type":"code","source":"def plot_model_evaluation(model, test_data, n_classes, target_labels):\n\n    results = model.evaluate(test_data, verbose=0)\n    loss = results[0]\n    acc = results[1]\n\n    print(\"    Test Loss: {:.5f}\".format(loss))\n    print(\"Test Accuracy: {:.2f}%\".format(acc * 100))\n\n    y_pred = np.squeeze((model.predict(test_data) >= 0.5).astype(int))\n    cm = confusion_matrix(test_data.labels, y_pred)\n    clr = classification_report(test_data.labels, y_pred, target_names=target_labels)\n\n    plt.figure(figsize=(15, 15))\n    sns.heatmap(cm, annot=True, fmt='g', vmin=0, cmap='Blues', cbar=False)\n    plt.xticks(ticks=np.arange(n_classes) + 0.5, labels=list(test_data.class_indices.keys()), rotation=90)\n    plt.yticks(ticks=np.arange(n_classes) + 0.5, labels=list(test_data.class_indices.keys()), rotation=0)\n    plt.xlabel(\"Predicted\")\n    plt.ylabel(\"Actual\")\n    plt.title(\"Confusion Matrix\")\n    plt.show()\n\n    print(\"Classification Report:\\n----------------------\\n\", clr)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_model_evaluation(model, test_images, 3, traindf[0:10]['label'].unique())","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Dears,**\n\n**I hope this message finds you well. I am excited to share that I am participating in the UBC Ovarian Cancer Subtype Classification and Outlier Detection (UBC-OCEAN) competition, and I need your support!**\n\n**As part of this competition, I have conducted an in-depth Exploratory Data Analysis (EDA) to gain crucial insights into the dataset, which is a fundamental step in developing effective solutions for cancer subtype classification and outlier detection. Now, I am reaching out to request your vote and support for my EDA submission.**\n\n**Your vote can make a significant difference in this competition and help me advance to the next stages. Here's how you can support me:**\n\n\n\n**1. Cast Your Vote:**\n\nVisit the competition platform and find my EDA submission.\nClick on the \"Vote\" or \"Support\" button to cast your vote.\n\n**2. Share with Your Network:**\n\nSpread the word among your friends, family, and colleagues who may be interested in supporting my work.\n\n**3. Provide Feedback:**\n\nIf you have any feedback or suggestions on my EDA, please feel free to share them with me. Your input is valuable and can help me improve.\nI am committed to making a positive impact in the field of cancer research, and your support will bring me one step closer to achieving that goal.\n\nThank you for taking the time to read this message, and I genuinely appreciate your support in this competition. Together, we can contribute to the fight against ovarian cancer and advance the field of data-driven healthcare.\n\nIf you have any questions or need more information about my EDA, please don't hesitate to reach out to me. Your support means the world to me!\n\nWarm regards,\nJeferson S. Pazze","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}}]}