{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# <mark>UBC Ovarian Cancer Subtype Classification and Outlier Detection (UBC-OCEAN) - EDA</mark>\n<span style=\"font-size:22px;color:purple\"> Thank you for having a look at my notebook - advice and feedback always welcomed!</span>\n\n\n<div class=\"alert alert-block alert-info\" style=\"font-size:14px; font-family:verdana;\">\n    📌 Dataset Link: <a href=\"https://www.kaggle.com/competitions/UBC-OCEAN/data\">https://www.kaggle.com/competitions/UBC-OCEAN/data</a>\n</div>\n\n\n## **Exploratory Data Analysis (EDA)** \nEDA is a crucial step in understanding and preparing your data for any data analysis or machine learning project, including UBC Ovarian Cancer Subtype Classification and Outlier Detection (UBC-OCEAN). Here's a step-by-step guide on how to perform an EDA for this dataset:\n\n**Overview**\n\nThe goal of the UBC Ovarian Cancer subtypE clAssification and outlier detectioN (UBC-OCEAN) competition is to classify ovarian cancer subtypes. You will build a model trained on the world's most extensive ovarian cancer dataset of histopathology images obtained from more than 20 medical centers.\n\n\n**Data Collection:**\n\nBegin by obtaining the UBC-OCEAN dataset, which should include information on ovarian cancer subtypes and possibly outlier detection data. Ensure you have a clear understanding of the dataset's structure and the meaning of each variable.\nData Loading:\n\nImport the dataset into your preferred data analysis environment, such as Python with libraries like pandas, numpy, and matplotlib/seaborn for visualization.\n\n**Data Loading:**\n\nImport the dataset into your preferred data analysis environment, such as Python with libraries like pandas, numpy, and matplotlib/seaborn for visualization.\n\n\n\n\n\n### **Initial Exploration:**\n\n**1 - Start by examining the basic characteristics of the data:**\n\n    Check the first few rows using df.head().\n    Check the data types and missing values using df.info().\n    Calculate basic statistics using df.describe().\n    Data Cleaning:\n\n**2 - Handle missing values, outliers, and duplicates:**\n\n    Use techniques like imputation for missing values.\n    Identify and deal with outliers appropriately.\n    Remove duplicate rows if necessary.\n    \n**3 - Data Visualization:**\n\n    Create visualizations to gain insights into the data:\n    Histograms and box plots for numerical features.\n    Bar plots for categorical features.\n    Correlation matrix and scatter plots to understand relationships between variables.\n    \n**4 - Feature Analysis:**\n\n    Explore relationships between features and the target variable(s) for classification and outlier detection.\n    Visualize how different features vary across different subtypes or classes.\n    Use box plots, violin plots, or swarm plots to compare feature distributions.\n\n**5 - Outlier Detection:**\n\n    If your dataset contains information related to outlier detection, perform a dedicated EDA for this aspect:\n    Visualize outliers using scatter plots or box plots.\n    Apply statistical methods or machine learning techniques to identify outliers.\n\n**6 - Dimensionality Reduction (optional):**\n\n    If the dataset has many features, consider dimensionality reduction techniques like Principal Component Analysis (PCA) to reduce the number of variables while preserving important information.\n\n**8 - Summary and Insights:**\n\n    Summarize your findings from the EDA, including any patterns, trends, or anomalies observed.\n    Document any data preprocessing steps applied.\n\n**7 - Next Steps:**\n\n    Based on your EDA findings, plan your next steps, which may include feature engineering, model selection, and further data preprocessing.\n\n\nRemember that EDA is an iterative process, and you may need to revisit these steps as you delve deeper into the dataset and develop your machine learning or data analysis models.","metadata":{}},{"cell_type":"code","source":"import os\nimport gc\nimport cv2\nimport math\nimport copy\nimport time\nimport random\nimport glob\nfrom matplotlib import pyplot as plt\nimport seaborn as sns\n\n# For data manipulation\nimport numpy as np\nimport pandas as pd\n\n# Pytorch Imports\nimport torch\nimport torch.nn as nn\nimport torch.optim as optim\nimport torch.nn.functional as F\nfrom torch.optim import lr_scheduler\nfrom torch.utils.data import Dataset, DataLoader\nfrom torch.cuda import amp\nimport torchvision\n\n# Utils\nimport joblib\nfrom tqdm import tqdm\nfrom collections import defaultdict\n\n# Sklearn Imports\nfrom sklearn.preprocessing import LabelEncoder\nfrom sklearn.model_selection import StratifiedKFold\n\n# For Image Models\nimport timm\n\n# Albumentations for augmentations\nimport albumentations as A\nfrom albumentations.pytorch import ToTensorV2\n\n# For colored terminal text\nfrom colorama import Fore, Back, Style\nb_ = Fore.BLUE\nsr_ = Style.RESET_ALL\n\nimport warnings\nwarnings.filterwarnings(\"ignore\")\n\n# For descriptive error messages\nos.environ['CUDA_LAUNCH_BLOCKING'] = \"1\"","metadata":{"execution":{"iopub.status.busy":"2023-10-09T18:48:24.756746Z","iopub.execute_input":"2023-10-09T18:48:24.757423Z","iopub.status.idle":"2023-10-09T18:48:24.764748Z","shell.execute_reply.started":"2023-10-09T18:48:24.757389Z","shell.execute_reply":"2023-10-09T18:48:24.76376Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Model Building**\n","metadata":{}},{"cell_type":"code","source":"CONFIG = {\n    \"seed\": 42,\n    \"img_size\": 512,\n    \"model_name\": \"tf_efficientnet_b0_ns\",\n    \"num_classes\": 5,\n    \"valid_batch_size\": 64,\n    \"device\": torch.device(\"cuda:0\" if torch.cuda.is_available() else \"cpu\"),\n}","metadata":{"execution":{"iopub.status.busy":"2023-10-09T18:48:24.769782Z","iopub.execute_input":"2023-10-09T18:48:24.770371Z","iopub.status.idle":"2023-10-09T18:48:24.780599Z","shell.execute_reply.started":"2023-10-09T18:48:24.770344Z","shell.execute_reply":"2023-10-09T18:48:24.779616Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"CONFIG","metadata":{"execution":{"iopub.status.busy":"2023-10-09T18:48:24.78283Z","iopub.execute_input":"2023-10-09T18:48:24.78386Z","iopub.status.idle":"2023-10-09T18:48:24.792349Z","shell.execute_reply.started":"2023-10-09T18:48:24.783828Z","shell.execute_reply":"2023-10-09T18:48:24.791167Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def set_seed(seed=42):\n    '''Sets the seed of the entire notebook so results are the same every time we run.\n    This is for REPRODUCIBILITY.'''\n    np.random.seed(seed)\n    torch.manual_seed(seed)\n    torch.cuda.manual_seed(seed)\n    # When running on the CuDNN backend, two further options must be set\n    torch.backends.cudnn.deterministic = True\n    torch.backends.cudnn.benchmark = False\n    # Set a fixed value for the hash seed\n    os.environ['PYTHONHASHSEED'] = str(seed)\n    \nset_seed(CONFIG['seed'])","metadata":{"execution":{"iopub.status.busy":"2023-10-09T18:48:24.793626Z","iopub.execute_input":"2023-10-09T18:48:24.794564Z","iopub.status.idle":"2023-10-09T18:48:24.801225Z","shell.execute_reply.started":"2023-10-09T18:48:24.794533Z","shell.execute_reply":"2023-10-09T18:48:24.800356Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ROOT_DIR = '/kaggle/input/UBC-OCEAN'\nTEST_DIR = '/kaggle/input/UBC-OCEAN/test_thumbnails'\nALT_TEST_DIR = '/kaggle/input/UBC-OCEAN/test_images'\nLABEL_ENCODER_BIN = \"/kaggle/input/ubcpytorchwith-classweights-training-fold1of5/label_encoder.pkl\"\nBEST_WEIGHT = \"/kaggle/input/ubcpytorchwith-classweights-training-fold1of5/Acc0.66_Loss1.0244_epoch16.bin\"","metadata":{"execution":{"iopub.status.busy":"2023-10-09T18:48:24.803749Z","iopub.execute_input":"2023-10-09T18:48:24.804117Z","iopub.status.idle":"2023-10-09T18:48:24.810058Z","shell.execute_reply.started":"2023-10-09T18:48:24.804082Z","shell.execute_reply":"2023-10-09T18:48:24.809179Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def get_test_file_path(image_id):\n    if os.path.exists(f\"{TEST_DIR}/{image_id}_thumbnail.png\"):\n        return f\"{TEST_DIR}/{image_id}_thumbnail.png\"\n    else:\n        return f\"{ALT_TEST_DIR}/{image_id}.png\"","metadata":{"execution":{"iopub.status.busy":"2023-10-09T18:48:24.811807Z","iopub.execute_input":"2023-10-09T18:48:24.812382Z","iopub.status.idle":"2023-10-09T18:48:24.820629Z","shell.execute_reply.started":"2023-10-09T18:48:24.812351Z","shell.execute_reply":"2023-10-09T18:48:24.819799Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = pd.read_csv(f\"{ROOT_DIR}/test.csv\")\ndf['file_path'] = df['image_id'].apply(get_test_file_path)\ndf['label'] = 0 # dummy\ndf","metadata":{"execution":{"iopub.status.busy":"2023-10-09T18:48:24.821976Z","iopub.execute_input":"2023-10-09T18:48:24.822871Z","iopub.status.idle":"2023-10-09T18:48:24.840624Z","shell.execute_reply.started":"2023-10-09T18:48:24.822838Z","shell.execute_reply":"2023-10-09T18:48:24.83974Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load train data\ntraindf = pd.read_csv('/kaggle/input/UBC-OCEAN/train.csv')\ntraindf.head()","metadata":{"execution":{"iopub.status.busy":"2023-10-09T18:48:24.842699Z","iopub.execute_input":"2023-10-09T18:48:24.843295Z","iopub.status.idle":"2023-10-09T18:48:24.856209Z","shell.execute_reply.started":"2023-10-09T18:48:24.843263Z","shell.execute_reply":"2023-10-09T18:48:24.855299Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"traindf['label'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2023-10-09T18:48:24.857283Z","iopub.execute_input":"2023-10-09T18:48:24.858124Z","iopub.status.idle":"2023-10-09T18:48:24.8657Z","shell.execute_reply.started":"2023-10-09T18:48:24.85809Z","shell.execute_reply":"2023-10-09T18:48:24.864662Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(traindf.shape)","metadata":{"execution":{"iopub.status.busy":"2023-10-09T18:48:24.954243Z","iopub.execute_input":"2023-10-09T18:48:24.954759Z","iopub.status.idle":"2023-10-09T18:48:24.959943Z","shell.execute_reply.started":"2023-10-09T18:48:24.954727Z","shell.execute_reply":"2023-10-09T18:48:24.958718Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"traindf.isna().sum().sum()\n","metadata":{"execution":{"iopub.status.busy":"2023-10-09T18:48:24.962018Z","iopub.execute_input":"2023-10-09T18:48:24.962588Z","iopub.status.idle":"2023-10-09T18:48:24.971822Z","shell.execute_reply.started":"2023-10-09T18:48:24.962557Z","shell.execute_reply":"2023-10-09T18:48:24.970666Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <span style=\"color:purple\">No null values present</span>\n​\n<div class=\"alert alert-block alert-info\">\n<b>Good:</b> now to move onto the next steps </div>","metadata":{}},{"cell_type":"code","source":"# list of columns we want the distributions for\ncolumns = [col for col in traindf.columns if col!='label']\n\n# loop to iterate over each column\nfor col in columns:\n    \n    # subplot for 3 columns (3 plots)\n    fig, axs = plt.subplots(figsize=(15,5), ncols=3)\n    \n    # 1st plot - distribution of the sample dataset\n    sns.histplot(data=traindf, x=col, kde=True, ax=axs[0])\n    axs[0].set_title('Sample Distribution')\n    \n    # 2nd plot - distribution of the selected column where the outcome is 1 (has diabetes)\n    sns.histplot(data=traindf[traindf['label']==\"HGSC\"], x=col, kde=True, ax=axs[1], color='orange')\n    axs[1].set_title('label - HGSC')\n    \n    # 3rd plot - distribution of the selected column where the outcome is 0 (doesn't have diabetes)\n    sns.histplot(data=traindf[traindf['label']==\"LGSC\"], x=col, kde=True, ax=axs[2], color='green')\n    axs[2].set_title('label - LGSC')\n    \n    # showing the plots\n    plt.tight_layout()\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-09T18:48:24.973085Z","iopub.execute_input":"2023-10-09T18:48:24.973881Z","iopub.status.idle":"2023-10-09T18:48:27.951059Z","shell.execute_reply.started":"2023-10-09T18:48:24.973851Z","shell.execute_reply":"2023-10-09T18:48:27.950175Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"traindf.info()","metadata":{"execution":{"iopub.status.busy":"2023-10-09T18:48:27.952288Z","iopub.execute_input":"2023-10-09T18:48:27.953085Z","iopub.status.idle":"2023-10-09T18:48:27.96426Z","shell.execute_reply.started":"2023-10-09T18:48:27.953049Z","shell.execute_reply":"2023-10-09T18:48:27.963351Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(traindf.image_id.is_unique)\nprint(traindf.label.is_unique)","metadata":{"execution":{"iopub.status.busy":"2023-10-09T18:48:27.968112Z","iopub.execute_input":"2023-10-09T18:48:27.968376Z","iopub.status.idle":"2023-10-09T18:48:27.984329Z","shell.execute_reply.started":"2023-10-09T18:48:27.96835Z","shell.execute_reply":"2023-10-09T18:48:27.983254Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Exploring image dimensions refers to understanding and working with the various aspects of an image's size and resolution. In the context of digital images, there are several key dimensions and attributes to consider:\n​\n1. **Resolution**: Resolution refers to the number of pixels (individual points of color) contained in an image. It is usually expressed in terms of pixels per inch (PPI) or dots per inch (DPI). Higher resolution images have more detail and are suitable for printing, while lower resolution images may be used for web display or on-screen viewing.\n​\n2. **Pixel Dimensions**: Pixel dimensions specify the width and height of an image in pixels. For example, an image might have dimensions of 1920x1080 pixels, which means it is 1920 pixels wide and 1080 pixels tall. This is commonly used for specifying the size of digital images.\n​\n3. **Aspect Ratio**: The aspect ratio is the ratio of an image's width to its height. Common aspect ratios include 4:3 (standard television), 16:9 (widescreen television), and 1:1 (square). Maintaining the correct aspect ratio is important to prevent image distortion.\n​\n4. **Physical Dimensions**: Physical dimensions refer to the size of an image when printed or displayed in the physical world. This is determined by both the pixel dimensions and the resolution. For example, an image with dimensions of 3000x2000 pixels at 300 DPI will be 10x6.67 inches when printed.\n​\n5. **File Size**: The file size of an image is measured in bytes or kilobytes (KB), megabytes (MB), etc. It depends on factors such as the color depth, compression, and pixel dimensions. Larger images with more detail tend to have larger file sizes.\n​\n6. **Color Depth**: Color depth, also known as bit depth, determines the number of colors a pixel can represent. Common color depths include 8-bit (256 colors), 24-bit (true color), and 32-bit (true color with alpha channel for transparency).\n​\n7. **DPI vs. PPI**: DPI (dots per inch) is often used in the context of printing, indicating how many ink dots a printer can produce in a linear inch. PPI (pixels per inch) is used for screen displays, representing the number of pixels in an inch of screen space.\n​\n8. **Scaling**: Scaling an image involves resizing it to different dimensions. You can scale an image up (enlargement) or down (reduction). Be aware that scaling too much can lead to a loss of image quality, especially when making an image larger.\n​\n9. **Cropping**: Cropping involves cutting out a portion of an image to focus on a specific area. This changes the pixel dimensions and aspect ratio of the image.\n​\n10. **Compression**: Image compression reduces file size by removing redundant or less important data. It can be lossless (no quality loss) or lossy (some quality loss). Common image formats like JPEG use lossy compression.\n​\nUnderstanding and managing these image dimensions and attributes is essential for various purposes, including graphic design, photography, web development, and printing. Depending on your specific needs, you may need to adjust these dimensions and attributes accordingly to achieve the desired result.","metadata":{}},{"cell_type":"code","source":"traindf[['image_height', 'image_width']].describe()","metadata":{"execution":{"iopub.status.busy":"2023-10-09T18:48:27.985487Z","iopub.execute_input":"2023-10-09T18:48:27.986538Z","iopub.status.idle":"2023-10-09T18:48:28.003256Z","shell.execute_reply.started":"2023-10-09T18:48:27.9865Z","shell.execute_reply":"2023-10-09T18:48:28.002202Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(traindf.image_id.shape[0])\nprint(len(os.listdir('/kaggle/input/UBC-OCEAN/train_images')))\nprint(len(os.listdir('/kaggle/input/UBC-OCEAN/train_thumbnails')))","metadata":{"execution":{"iopub.status.busy":"2023-10-09T18:48:28.004662Z","iopub.execute_input":"2023-10-09T18:48:28.005199Z","iopub.status.idle":"2023-10-09T18:48:28.012724Z","shell.execute_reply.started":"2023-10-09T18:48:28.005168Z","shell.execute_reply":"2023-10-09T18:48:28.011632Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"traindf.label.value_counts()","metadata":{"execution":{"iopub.status.busy":"2023-10-09T18:48:28.014154Z","iopub.execute_input":"2023-10-09T18:48:28.014728Z","iopub.status.idle":"2023-10-09T18:48:28.023783Z","shell.execute_reply.started":"2023-10-09T18:48:28.01469Z","shell.execute_reply":"2023-10-09T18:48:28.022743Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"HGSC = traindf[traindf['label']==\"HGSC\"]\nEC = traindf[traindf['label']==\"EC\"]\nCC = traindf[traindf['label']==\"CC\"]\nLGSC = traindf[traindf['label']==\"LGSC\"]\nMC = traindf[traindf['label']==\"MC\"]","metadata":{"execution":{"iopub.status.busy":"2023-10-09T18:48:28.024975Z","iopub.execute_input":"2023-10-09T18:48:28.025981Z","iopub.status.idle":"2023-10-09T18:48:28.034969Z","shell.execute_reply.started":"2023-10-09T18:48:28.02595Z","shell.execute_reply":"2023-10-09T18:48:28.034098Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Data Visualization 📈 \n\n### Randomly Visualize Images","metadata":{}},{"cell_type":"code","source":"# Set the figure size\nplt.figure(figsize=(20, 6))\n\n# Set the font size\nplt.rcParams['font.size'] = 14\n\n# Set the colors\ncolors = ['lightgreen', 'lightblue', 'purple', 'blue', 'yellow']\n\n# Plot the pie chart for the training set\nplt.subplot(1, 1, 1)\nplt.pie([len(HGSC), len(EC), len(CC), len(LGSC), len(MC)], labels=['HGSC', 'EC', 'CC', 'LGSC', 'MC'], autopct='%1.1f%%', colors=colors)\nplt.title('Training Set')\n\n\n# Add a main title to the figure\nplt.suptitle('Distribution of HGSC, EC, CC, LGSC and MC Images in the Training data', fontsize=20, y=1.05)\n\n# Show the plot\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-09T18:48:28.036317Z","iopub.execute_input":"2023-10-09T18:48:28.037261Z","iopub.status.idle":"2023-10-09T18:48:28.227973Z","shell.execute_reply.started":"2023-10-09T18:48:28.037228Z","shell.execute_reply":"2023-10-09T18:48:28.226956Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Examine the Data Split of training and testing data\ntrain_data = glob.glob('/kaggle/input/UBC-OCEAN/train_images/*.png')\ntest_data = glob.glob('/kaggle/input/UBC-OCEAN/test_images/*.png')\n\nprint(f\"The Training Set contains: {len(train_data)} images\")\nprint(f\"The Testing Set contains: {len(test_data)} images\")","metadata":{"execution":{"iopub.status.busy":"2023-10-09T18:48:28.229234Z","iopub.execute_input":"2023-10-09T18:48:28.229983Z","iopub.status.idle":"2023-10-09T18:48:28.243044Z","shell.execute_reply.started":"2023-10-09T18:48:28.229946Z","shell.execute_reply":"2023-10-09T18:48:28.242012Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Calculate total counts\ntotal_train = len(train_data)\ntotal_test = len(test_data)\n\n# Set the figure size\nplt.figure(figsize=(4, 4))\n\n# Set the font size\nplt.rcParams['font.size'] = 12\n\n# Set the colors\ncolors = ['lightgreen', 'red']\n\n# Plot the pie chart for the total set\nplt.pie([total_train, total_test], labels=['Training Set', 'Testing Set'], autopct='%1.1f%%', colors=colors)\nplt.title('Distribution of Images in Training and Testing Sets')\n\n# Show the plot\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-09T18:48:28.24438Z","iopub.execute_input":"2023-10-09T18:48:28.245272Z","iopub.status.idle":"2023-10-09T18:48:28.354627Z","shell.execute_reply.started":"2023-10-09T18:48:28.245241Z","shell.execute_reply":"2023-10-09T18:48:28.353854Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_sub = pd.read_csv(f\"{ROOT_DIR}/sample_submission.csv\")\ndf_sub","metadata":{"execution":{"iopub.status.busy":"2023-10-09T18:48:28.355869Z","iopub.execute_input":"2023-10-09T18:48:28.3566Z","iopub.status.idle":"2023-10-09T18:48:28.373534Z","shell.execute_reply.started":"2023-10-09T18:48:28.356569Z","shell.execute_reply":"2023-10-09T18:48:28.372276Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"encoder = joblib.load( LABEL_ENCODER_BIN )","metadata":{"execution":{"iopub.status.busy":"2023-10-09T18:48:28.378341Z","iopub.execute_input":"2023-10-09T18:48:28.37929Z","iopub.status.idle":"2023-10-09T18:48:28.385218Z","shell.execute_reply.started":"2023-10-09T18:48:28.37926Z","shell.execute_reply":"2023-10-09T18:48:28.384003Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Training Configuration**  ⚙️\n\nA training configuration is a vital component of any machine learning or deep learning project, serving as the blueprint that outlines the parameters, settings, and conditions under which a model is trained. This essential document encapsulates the entire training process, providing clarity and reproducibility to the development and deployment of AI systems. Here, we delve into the key elements and considerations that constitute a comprehensive training configuration:\n\n**Model Architecture:** The configuration specifies the architecture of the neural network or machine learning model being trained. It outlines the layers, nodes, and connections that define the model's structure. This includes details like the type of layers (e.g., convolutional, recurrent), activation functions, and any custom layers or modifications.\n\n**Data Preparation:** It outlines the methods and procedures for data preprocessing, augmentation, and normalization. This may include data scaling, one-hot encoding, or image augmentation techniques. Proper data preparation is critical for model convergence and performance.\n\n**Hyperparameters:** The training configuration specifies hyperparameters, which are settings that control the learning process. This includes parameters like learning rate, batch size, epochs, and optimization algorithms (e.g., Adam, SGD). Tinkering with these hyperparameters can significantly impact model training outcomes.\n\n**Loss Function:** The choice of the loss function is pivotal to training. This component of the configuration details the specific loss metric that the model optimizes during training, aligning it with the objectives of the project (e.g., mean squared error for regression, cross-entropy for classification).\n\n**Metrics for Evaluation:** The configuration lists the evaluation metrics used to assess the model's performance during and after training. Common metrics include accuracy, F1-score, mean absolute error (MAE), and mean squared error (MSE).\n\n**Regularization Techniques:** If applicable, regularization techniques such as dropout, L1, or L2 regularization are specified in the configuration to prevent overfitting.\n\n**Checkpointing:** Configuration may include settings for model checkpointing, which periodically saves the model's weights and progress during training. This is essential for resuming training or selecting the best model for deployment.\n\n**Early Stopping:** Parameters for early stopping, based on validation metrics, are often included. This helps prevent overtraining by halting training when the model's performance on validation data plateaus or deteriorates.\n\n**Hardware and Environment:** It mentions the hardware resources utilized during training, including CPU, GPU, or TPUs, as well as the software environment, such as the version of deep learning frameworks (e.g., TensorFlow, PyTorch) and the operating system.\n\n**Batch Processing:** Configuration can also include information on distributed training, specifying whether training is done in a single batch or in mini-batches, and whether it spans multiple GPUs or nodes.\n\n**Data Splits:** The division of data into training, validation, and test sets is outlined in the configuration to ensure proper model assessment and generalization.\n\n**Documentation:** A well-documented training configuration is crucial for reproducibility. It should contain comments and explanations for each parameter and decision made, enabling easy sharing and replication of the training process.\n\nIn summary, a training configuration is the comprehensive roadmap that guides the development of machine learning and deep learning models. It plays a pivotal role in achieving reproducibility, scalability, and the successful deployment of AI systems, ensuring that the model can be trained consistently and effectively for various applications.","metadata":{}},{"cell_type":"code","source":"class UBCDataset(Dataset):\n    def __init__(self, df, transforms=None):\n        self.df = df\n        self.file_names = df['file_path'].values\n        self.labels = df['label'].values\n        self.transforms = transforms\n        \n    def __len__(self):\n        return len(self.df)\n    \n    def __getitem__(self, index):\n        img_path = self.file_names[index]\n        img = cv2.imread(img_path)\n        img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)\n        label = self.labels[index]\n        \n        if self.transforms:\n            img = self.transforms(image=img)[\"image\"]\n            \n        return {\n            'image': img,\n            'label': torch.tensor(label, dtype=torch.long)\n        }","metadata":{"execution":{"iopub.status.busy":"2023-10-09T18:48:28.386948Z","iopub.execute_input":"2023-10-09T18:48:28.387988Z","iopub.status.idle":"2023-10-09T18:48:28.399248Z","shell.execute_reply.started":"2023-10-09T18:48:28.387951Z","shell.execute_reply":"2023-10-09T18:48:28.39789Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Augmentations**","metadata":{}},{"cell_type":"code","source":"data_transforms = {\n    \"valid\": A.Compose([\n        A.Resize(CONFIG['img_size'], CONFIG['img_size']),\n        A.Normalize(\n                mean=[0.485, 0.456, 0.406], \n                std=[0.229, 0.224, 0.225], \n                max_pixel_value=255.0, \n                p=1.0\n            ),\n        ToTensorV2()], p=1.)\n}","metadata":{"execution":{"iopub.status.busy":"2023-10-09T18:48:28.400706Z","iopub.execute_input":"2023-10-09T18:48:28.401319Z","iopub.status.idle":"2023-10-09T18:48:28.412195Z","shell.execute_reply.started":"2023-10-09T18:48:28.401288Z","shell.execute_reply":"2023-10-09T18:48:28.411016Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **GeM Pooling**","metadata":{}},{"cell_type":"code","source":"class GeM(nn.Module):\n    def __init__(self, p=3, eps=1e-6):\n        super(GeM, self).__init__()\n        self.p = nn.Parameter(torch.ones(1)*p)\n        self.eps = eps\n\n    def forward(self, x):\n        return self.gem(x, p=self.p, eps=self.eps)\n        \n    def gem(self, x, p=3, eps=1e-6):\n        return F.avg_pool2d(x.clamp(min=eps).pow(p), (x.size(-2), x.size(-1))).pow(1./p)\n        \n    def __repr__(self):\n        return self.__class__.__name__ + \\\n                '(' + 'p=' + '{:.4f}'.format(self.p.data.tolist()[0]) + \\\n                ', ' + 'eps=' + str(self.eps) + ')'","metadata":{"execution":{"iopub.status.busy":"2023-10-09T18:48:28.414333Z","iopub.execute_input":"2023-10-09T18:48:28.416616Z","iopub.status.idle":"2023-10-09T18:48:28.432969Z","shell.execute_reply.started":"2023-10-09T18:48:28.416571Z","shell.execute_reply":"2023-10-09T18:48:28.431926Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Create Model**\n\n<!DOCTYPE html>\n<html>\n<head>\n</head>\n<body>\n    <h1>EfficientNet</h1>\n    <p>\n        EfficientNet is a family of convolutional neural network (CNN) architectures designed for efficient and effective deep learning in computer vision tasks. It was introduced in 2019 by researchers at Google AI. EfficientNet models are known for their exceptional performance in image classification tasks while being computationally efficient, making them suitable for a wide range of applications.\n    </p>\n    <p>\n        Key features of EfficientNet include:\n    </p>\n    <ul>\n        <li>Compound Scaling: EfficientNet employs a novel compound scaling method that balances the network's depth, width, and resolution to achieve optimal performance without a significant increase in computational cost.</li>\n        <li>Efficient Building Blocks: The architecture incorporates efficient building blocks like depthwise separable convolutions and squeeze-and-excitation blocks to reduce the number of parameters and computational overhead.</li>\n        <li>Variants: EfficientNet comes in various variants (e.g., EfficientNet-B0, B1, B2, ..., B7) that offer different trade-offs between model size and accuracy, allowing users to choose the one that suits their specific requirements.</li>\n        <li>State-of-the-Art Performance: EfficientNet models have achieved top performance in benchmark datasets such as ImageNet, outperforming many previous CNN architectures with smaller model sizes.</li>\n    </ul>\n    <p>\n        EfficientNet has become a popular choice in the field of computer vision due to its ability to achieve impressive results with fewer parameters, making it practical for deployment on resource-constrained devices and applications.\n    </p>\n    <p>\n        If you are working on image classification or related tasks, considering EfficientNet as part of your deep learning architecture can lead to efficient and accurate results.\n    </p>\n</body>\n</html>\n","metadata":{}},{"cell_type":"code","source":"class EfficientNetB0(nn.Module):\n    '''\n    EfficientNet B0 fine-tune.\n    '''\n    def __init__(self, model_name, num_classes, pretrained=False, checkpoint_path=None):\n        '''\n        Fine tune for EfficientNetB0\n        Args\n            n_classes : int - Number of classification categories.\n            learnable_modules : tuple - Names of the modules to fine-tune.\n        Return\n            \n        '''\n        super(EfficientNetB0, self).__init__()\n        self.model = timm.create_model(model_name, pretrained=pretrained)\n\n        in_features = self.model.classifier.in_features\n        self.model.classifier = nn.Identity()\n        self.model.global_pool = nn.Identity()\n        self.pooling = GeM()\n        self.linear = nn.Linear(in_features, num_classes)\n        self.softmax = nn.Softmax(dim=1)\n\n    def forward(self, images):\n        \"\"\"\n        Forward function for the fine-tuned model\n        Args\n            x: \n        Return\n            result\n        \"\"\"\n        features = self.model(images)\n        pooled_features = self.pooling(features).flatten(1)\n        output = self.linear(pooled_features)\n        return output\n\n    \nmodel = EfficientNetB0(CONFIG['model_name'], CONFIG['num_classes'])\nmodel.load_state_dict(torch.load( BEST_WEIGHT ))\nmodel.to(CONFIG['device']);","metadata":{"execution":{"iopub.status.busy":"2023-10-09T18:48:28.437835Z","iopub.execute_input":"2023-10-09T18:48:28.438476Z","iopub.status.idle":"2023-10-09T18:48:28.634174Z","shell.execute_reply.started":"2023-10-09T18:48:28.438433Z","shell.execute_reply":"2023-10-09T18:48:28.63326Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#model","metadata":{"execution":{"iopub.status.busy":"2023-10-09T18:48:28.635544Z","iopub.execute_input":"2023-10-09T18:48:28.636121Z","iopub.status.idle":"2023-10-09T18:48:28.640161Z","shell.execute_reply.started":"2023-10-09T18:48:28.63609Z","shell.execute_reply":"2023-10-09T18:48:28.639263Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_dataset = UBCDataset(df, transforms=data_transforms[\"valid\"])\ntest_loader = DataLoader(test_dataset, batch_size=CONFIG['valid_batch_size'], \n                          num_workers=2, shuffle=False, pin_memory=True)","metadata":{"execution":{"iopub.status.busy":"2023-10-09T18:48:28.641462Z","iopub.execute_input":"2023-10-09T18:48:28.64202Z","iopub.status.idle":"2023-10-09T18:48:28.649492Z","shell.execute_reply.started":"2023-10-09T18:48:28.641988Z","shell.execute_reply":"2023-10-09T18:48:28.648555Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"preds = []\nwith torch.no_grad():\n    bar = tqdm(enumerate(test_loader), total=len(test_loader))\n    for step, data in bar:        \n        images = data['image'].to(CONFIG[\"device\"], dtype=torch.float)        \n        batch_size = images.size(0)\n        outputs = model(images)\n        _, predicted = torch.max(model.softmax(outputs), 1)\n        preds.append( predicted.detach().cpu().numpy() )\npreds = np.concatenate(preds).flatten()\npred_labels = encoder.inverse_transform( preds )","metadata":{"execution":{"iopub.status.busy":"2023-10-09T18:48:28.650785Z","iopub.execute_input":"2023-10-09T18:48:28.651305Z","iopub.status.idle":"2023-10-09T18:48:29.01387Z","shell.execute_reply.started":"2023-10-09T18:48:28.651275Z","shell.execute_reply":"2023-10-09T18:48:29.012593Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pred_labels","metadata":{"execution":{"iopub.status.busy":"2023-10-09T18:48:29.01546Z","iopub.execute_input":"2023-10-09T18:48:29.015838Z","iopub.status.idle":"2023-10-09T18:48:29.022815Z","shell.execute_reply.started":"2023-10-09T18:48:29.015803Z","shell.execute_reply":"2023-10-09T18:48:29.021881Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_sub[\"label\"] = pred_labels\ndf_sub.to_csv(\"submission.csv\", index=False)","metadata":{"execution":{"iopub.status.busy":"2023-10-09T18:48:29.024341Z","iopub.execute_input":"2023-10-09T18:48:29.024942Z","iopub.status.idle":"2023-10-09T18:48:29.036889Z","shell.execute_reply.started":"2023-10-09T18:48:29.024909Z","shell.execute_reply":"2023-10-09T18:48:29.035987Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_sub","metadata":{"execution":{"iopub.status.busy":"2023-10-09T18:48:29.038366Z","iopub.execute_input":"2023-10-09T18:48:29.038725Z","iopub.status.idle":"2023-10-09T18:48:29.050203Z","shell.execute_reply.started":"2023-10-09T18:48:29.038694Z","shell.execute_reply":"2023-10-09T18:48:29.049142Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Dears,**\n\n**I hope this message finds you well. I am excited to share that I am participating in the UBC Ovarian Cancer Subtype Classification and Outlier Detection (UBC-OCEAN) competition, and I need your support!**\n\n**As part of this competition, I have conducted an in-depth Exploratory Data Analysis (EDA) to gain crucial insights into the dataset, which is a fundamental step in developing effective solutions for cancer subtype classification and outlier detection. Now, I am reaching out to request your vote and support for my EDA submission.**\n\n**Your vote can make a significant difference in this competition and help me advance to the next stages. Here's how you can support me:**\n\n\n\n**1. Cast Your Vote:**\n\nVisit the competition platform and find my EDA submission.\nClick on the \"Vote\" or \"Support\" button to cast your vote.\n\n**2. Share with Your Network:**\n\nSpread the word among your friends, family, and colleagues who may be interested in supporting my work.\n\n**3. Provide Feedback:**\n\nIf you have any feedback or suggestions on my EDA, please feel free to share them with me. Your input is valuable and can help me improve.\nI am committed to making a positive impact in the field of cancer research, and your support will bring me one step closer to achieving that goal.\n\nThank you for taking the time to read this message, and I genuinely appreciate your support in this competition. Together, we can contribute to the fight against ovarian cancer and advance the field of data-driven healthcare.\n\nIf you have any questions or need more information about my EDA, please don't hesitate to reach out to me. Your support means the world to me!\n\nWarm regards,\nJeferson S. Pazze\n\nLinkedin: https://www.linkedin.com/in/jeferson-souza-pazze-53806469/","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}