{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":45867,"databundleVersionId":6924515,"sourceType":"competition"},{"sourceId":6984590,"sourceType":"datasetVersion","datasetId":4014175}],"dockerImageVersionId":30627,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"![Cover](https://raw.githubusercontent.com/AniMilina/UBC-Ovarian-Cancer-Subtype-Classification-and-Outl/main/Cover.jpg)","metadata":{}},{"cell_type":"markdown","source":"## Project description","metadata":{}},{"cell_type":"markdown","source":"Our work is focused on solving a critical problem in medical research—automated classification of cancer types using computed tomography (CT). Cancer stands out as one of the most common and dangerous diseases, and the development of effective tools for its diagnosis is of utmost importance  \n\nUnfortunately, I did not see the desired result, despite the fact that I used several approaches to the model architecture. But I think a test set closed to participants should better show the true result. Also, having the opportunity to use the Internet, you can enable a model that uses pre-trained ResNet50, so I save this option in the comments  \n\nWe use advanced deep learning techniques, such as neural networks, to create a classification model that can automatically identify different types of cancer in computer images. Our model is built on the ResNet50 architecture and pre-trained on a large dataset  \n\nHowever, we are faced with several challenging problems such as class imbalance, data scarcity, and difficulties in training neural networks for medical applications. To address these issues, we implement strategies such as class balancing, data augmentation, hyperparameter optimization, and careful monitoring of the training process  \n\nOur goal goes beyond simply building an accurate model; We strive to create a tool that will complement the efforts of medical professionals in diagnosing cancer, allowing us to quickly and accurately detect types of cancer in the early stages  \n\nThis work represents a significant step towards improving the efficiency and accessibility of cancer diagnosis using modern machine learning and deep learning techniques  ","metadata":{}},{"cell_type":"markdown","source":"## Data Description","metadata":{}},{"cell_type":"markdown","source":"`[train/test]_images:`\n\nA folder containing images for classification  \n\nTwo categories of images: Whole Slide Images (WSI) and Tissue Microarray (TMA)  \nWSI at 20x magnification, TMAs at 40x magnification  \nTest set contains images from different hospitals, some with large dimensions (up to 100,000 x 50,000 pixels)  \nApproximately 2,000 images in the test set, mainly TMAs  \nTest set intentionally challenging for model generalization  \nImages not available for direct loading due to size (550 GB)  \n\n`[train/test].csv:`  \n\nLabels for the train set  \n\n**Columns:**  \n>image_id: Unique ID code for each image  \nlabel: Target class, one of ovarian cancer subtypes (CC, EC, HGSC, LGSC, MC, Other)  \nimage_width: Image width in pixels  \nimage_height: Image height in pixels  \nis_tma: True if the slide is a tissue microarray (only available for the train set)  \n\n`[train/test]_thumbnails:`  \n\nA folder containing smaller .png copies of the WSI images  \nThumbnails not provided for TMAs  \n\n`sample_submission.csv:`  \n\nA valid sample submission file, with the first row available for download  \n\n`Supplemental Masks:`  \n\nRoughly 150 masks indicating cancerous, healthy, or necrotic parts of relevant WSI images from the train set  \nSeparate dataset available for download  \nMask file names match the file names of the corresponding train images  \n","metadata":{}},{"cell_type":"markdown","source":"## Installing and importing libraries","metadata":{}},{"cell_type":"code","source":"# !pip install --upgrade tensorflow\n# !pip install --upgrade keras","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:56:02.453946Z","iopub.execute_input":"2023-12-25T10:56:02.454383Z","iopub.status.idle":"2023-12-25T10:56:02.487517Z","shell.execute_reply.started":"2023-12-25T10:56:02.454351Z","shell.execute_reply":"2023-12-25T10:56:02.486268Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\n\nimport matplotlib.pyplot as plt\nimport matplotlib.image as mpimg\nimport seaborn as sns\nimport imageio\nfrom PIL import Image\nimport cv2\nimport ipywidgets as widgets\nfrom skimage import io\n\nimport os\nimport glob\nimport random\n\nimport tensorflow as tf\nfrom tensorflow.keras import layers, models\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator\nfrom tensorflow.keras.applications import VGG16,ResNet50\nfrom tensorflow.keras.optimizers import Adam\nfrom tensorflow.keras.utils import to_categorical\nfrom tensorflow.keras.utils import get_file\nfrom tensorflow.keras.models import Sequential, load_model,Model\nfrom keras.layers import GlobalAveragePooling2D\nfrom tensorflow.keras.layers import Conv2D, MaxPooling2D, Flatten, Dense, BatchNormalization, Dropout\nfrom tensorflow.keras.callbacks import EarlyStopping, ModelCheckpoint, Callback\nfrom keras.preprocessing import image\nfrom keras.applications.resnet50 import preprocess_input, decode_predictions\n\nfrom sklearn.preprocessing import LabelEncoder\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import balanced_accuracy_score\nfrom sklearn.utils.class_weight import compute_class_weight\n\nimport shutil\n\nimport warnings\nimport time","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-12-25T10:56:02.489888Z","iopub.execute_input":"2023-12-25T10:56:02.4903Z","iopub.status.idle":"2023-12-25T10:56:20.300341Z","shell.execute_reply.started":"2023-12-25T10:56:02.490269Z","shell.execute_reply":"2023-12-25T10:56:20.298964Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Settings","metadata":{}},{"cell_type":"code","source":"# Disable warnings\n\nwarnings.filterwarnings('ignore')","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:56:20.302229Z","iopub.execute_input":"2023-12-25T10:56:20.30309Z","iopub.status.idle":"2023-12-25T10:56:20.313096Z","shell.execute_reply.started":"2023-12-25T10:56:20.303049Z","shell.execute_reply":"2023-12-25T10:56:20.311319Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class f:    \n    BOLD = \"\\033[1m\"     # Bold text\n    ITALIC = \"\\033[3m\"   # Italic text\n    END = \"\\033[0m\"      # Reset style","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:56:20.318164Z","iopub.execute_input":"2023-12-25T10:56:20.318848Z","iopub.status.idle":"2023-12-25T10:56:20.343724Z","shell.execute_reply.started":"2023-12-25T10:56:20.318787Z","shell.execute_reply":"2023-12-25T10:56:20.342384Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"facecolor='lightgray'\n# gridcolor='gray'\n\ncolors = ['pink', 'steelblue', 'hotpink','lightgreen','gray','salmon','gold', 'seagreen', 'skyblue', 'orchid']\n\nclass_colors = {'HGSC': 'pink', 'EC': 'steelblue', 'CC': 'hotpink', 'LGSC': 'salmon', 'MC': 'lightgreen'}\n","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:56:20.345906Z","iopub.execute_input":"2023-12-25T10:56:20.347054Z","iopub.status.idle":"2023-12-25T10:56:20.355779Z","shell.execute_reply.started":"2023-12-25T10:56:20.347014Z","shell.execute_reply":"2023-12-25T10:56:20.354646Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Attempt 3:\n\n# weights_url = 'https://github.com/fchollet/deep-learning-models/releases/download/v0.2/resnet50_weights_tf_dim_ordering_tf_kernels_notop.h5'\n# weights_path = get_file('resnet50_weights_tf_dim_ordering_tf_kernels_notop.h5',\n#                         weights_url,\n#                         cache_subdir='/kaggle/working/')","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:56:20.358331Z","iopub.execute_input":"2023-12-25T10:56:20.358856Z","iopub.status.idle":"2023-12-25T10:56:20.36952Z","shell.execute_reply.started":"2023-12-25T10:56:20.358814Z","shell.execute_reply":"2023-12-25T10:56:20.368161Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Functions","metadata":{}},{"cell_type":"code","source":"def explore_data(dataframes):\n    for name, df in dataframes.items():\n        print(f\"\\n{name} Data:\")\n        print(f\"Shape:\\n{df.shape}\")\n        print(f\"\\nInfo:\")\n        df.info()\n        print(f\"\\nDuplicates:\")\n        duplicates = df[df.duplicated()]\n        print(duplicates)\n        print(f\"\\nMissing Values:\")\n        missing_values = df.isnull().sum()\n        print(missing_values[missing_values > 0])","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:56:20.37142Z","iopub.execute_input":"2023-12-25T10:56:20.372639Z","iopub.status.idle":"2023-12-25T10:56:20.387454Z","shell.execute_reply.started":"2023-12-25T10:56:20.372604Z","shell.execute_reply":"2023-12-25T10:56:20.385858Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def visualize_distribution(data, x, title, xlabel, ylabel, facecolor='lightgray', gridcolor='gray', colors=None, rotation=0):\n    \"\"\"\n    Visualization of data distribution.\n\n    Options:\n    - data: DataFrame, data for visualization.\n    - x: str, column name for the X axis.\n    - title: str, title of the graph.\n    - xlabel: str, X-axis label.\n    - ylabel: str, Y-axis label.\n    - facecolor: str, background color of the chart.\n    - gridcolor: str, color of the chart grid.\n    - colors: list, additional colors for the chart.\n    - rotation: int, the angle of rotation of the labels.\n\n    Returns:\n    None\n    \"\"\"\n    plt.figure(figsize=(10, 6))\n    sns.set_palette(colors) if colors else None\n    sns.countplot(x=x, data=data)\n    sns.set(style=\"whitegrid\")\n    plt.title(title)\n    plt.xlabel(xlabel)\n    plt.ylabel(ylabel)\n    ax = plt.gca()\n    ax.set_facecolor(facecolor)\n    ax.grid(color=gridcolor)\n    plt.xticks(rotation=rotation)\n\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:56:20.389026Z","iopub.execute_input":"2023-12-25T10:56:20.390151Z","iopub.status.idle":"2023-12-25T10:56:20.401815Z","shell.execute_reply.started":"2023-12-25T10:56:20.390117Z","shell.execute_reply":"2023-12-25T10:56:20.400407Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def display_images_by_label(train_df, data_path, num_samples=4):\n    \"\"\"\n    Let's display sample images for each label in the training data set\n\n    Parameters:\n    - train_df: DataFrame, training dataset containing \"label\" and \"image_id\" columns.\n    - data_path: str, path to the data directory.\n    - num_samples: int, number of images to display for each label.\n    \"\"\"\n    for label, group_df in train_df.groupby(\"label\"):\n        fig, axes = plt.subplots(ncols=num_samples, figsize=(16, 4))\n\n        for i, image_name in enumerate(group_df[\"image_id\"].sample(num_samples)):\n            img_path = os.path.join(data_path, \"train_thumbnails\", f\"{image_name}_thumbnail.png\")\n\n            if not os.path.isfile(img_path):\n                img_path = os.path.join(data_path, \"train_images\", f\"{image_name}.png\")\n                print(f\"Missing thumbnail for {img_path} but image exists {os.path.isfile(img_path)}\")\n                continue\n\n            axes[i].imshow(plt.imread(img_path))\n            axes[i].set_title(f\"Label: {label} for Image: {image_name}\")\n            axes[i].set_axis_off()\n\n        fig.suptitle(f\"Sample Images for Label: {label}\")\n        plt.show()","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:56:20.403821Z","iopub.execute_input":"2023-12-25T10:56:20.404347Z","iopub.status.idle":"2023-12-25T10:56:20.417367Z","shell.execute_reply.started":"2023-12-25T10:56:20.404304Z","shell.execute_reply":"2023-12-25T10:56:20.415869Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def plot_image_dimensions(train_df, class_colors):\n    \"\"\"\n    Let's create a paired graph of image sizes grouped by class\n\n    Parameters:\n    - train_df: DataFrame, training dataset containing \"label\", \"image_width\", and \"image_height\" columns.\n    - class_colors: dict, mapping of class labels to color codes.\n    \"\"\"\n    plt.figure(figsize=(12, 10), facecolor='lightgrey')\n\n    sns.pairplot(train_df, hue=\"label\", vars=[\"image_width\", \"image_height\"], palette=class_colors)\n\n    plt.suptitle(\"Pairplot of Image Dimensions by Class\", y=1.02, color='black')\n    ax = plt.gca()\n    ax.set_facecolor('lightgrey')\n    \n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:56:20.423645Z","iopub.execute_input":"2023-12-25T10:56:20.42505Z","iopub.status.idle":"2023-12-25T10:56:20.434857Z","shell.execute_reply.started":"2023-12-25T10:56:20.424997Z","shell.execute_reply":"2023-12-25T10:56:20.433331Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def plot_pairplot_by_class(train_df, class_colors):\n    \"\"\"\n    Let's create a paired plot of a data set grouped by classes\n\n    Parameters:\n    - train_df: DataFrame, training dataset containing \"label\", \"image_width\", and \"image_height\" columns.\n    - class_colors: dict, mapping of class labels to color codes.\n    \"\"\"\n    sns.set(style=\"whitegrid\")\n\n    plt.figure(figsize=(12, 10), facecolor='lightgrey')\n\n    sns.pairplot(train_df, hue=\"label\", palette=class_colors)\n\n    plt.suptitle(\"Pairplot by Class\", y=1.02, color='black')\n\n    ax = plt.gca()\n    ax.set_facecolor('lightgrey')\n\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:56:20.436521Z","iopub.execute_input":"2023-12-25T10:56:20.437574Z","iopub.status.idle":"2023-12-25T10:56:20.449564Z","shell.execute_reply.started":"2023-12-25T10:56:20.43753Z","shell.execute_reply":"2023-12-25T10:56:20.448365Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def plot_cancer_images(train_df, cancer_labels, descriptions):\n    \"\"\"\n    Let's create images for each type of cancer along with descriptions.\n\n    Parameters:\n    - train_df: DataFrame, training dataset containing \"label\" and \"image_id\" columns.\n    - cancer_labels: List of strings, cancer labels to plot.\n    - descriptions: Dictionary, mapping cancer labels to descriptions.\n    \"\"\"\n    num_images_per_row = 5\n    num_rows = 1\n    \n    for label in cancer_labels:\n        plot_images_for_cancer(train_df, label, num_images_per_row, num_rows, descriptions)\n\ndef plot_images_for_cancer(train_df, label, num_images_per_row, num_rows, descriptions):\n    \"\"\"\n    Let's create images for a specific cancer type.\n\n    Parameters:\n    - train_df: DataFrame, training dataset containing \"label\" and \"image_id\" columns.\n    - label: String, cancer label to plot.\n    - num_images_per_row: Number of images to display in each row.\n    - num_rows: Number of rows in the subplot.\n    - descriptions: Dictionary, mapping cancer labels to descriptions.\n    \"\"\"\n    df_cancer = train_df[train_df['label'] == label]\n    tma_image_ids = list(df_cancer[df_cancer['is_tma']]['image_id'])\n\n    if len(tma_image_ids) >= num_images_per_row:\n        plt.figure(figsize=(20.0, 6.0))\n        \n        for i, image_id in enumerate(tma_image_ids[:num_images_per_row]):\n            plot_single_image(image_id, label, i + 1, descriptions[label])\n        \n        plt.tight_layout()\n        plt.show()\n    else:\n        print(f\"Not enough TMA images for {label}\")\n\ndef plot_single_image(image_id, label, subplot_idx, description):\n    \"\"\"\n    Let's create a single image with title and description.\n\n    Parameters:\n    - image_id: String, image identifier.\n    - label: String, cancer label.\n    - subplot_idx: Index for subplot placement.\n    - description: String, description of the cancer type.\n    \"\"\"\n    plt.subplot(1, 5, subplot_idx)\n    plt.title(f'image_id:{image_id} (TMA)', fontsize=14)\n    \n    if subplot_idx == 1:\n        plt.ylabel(label, fontsize=14)\n    \n    io.imshow(f'/kaggle/input/UBC-OCEAN/train_images/{image_id}.png')\n    plt.tick_params(labelbottom=False, labelleft=False, labelright=False, labeltop=False,\n                    bottom=False, left=False, right=False, top=False)\n","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:56:20.451283Z","iopub.execute_input":"2023-12-25T10:56:20.452429Z","iopub.status.idle":"2023-12-25T10:56:20.467926Z","shell.execute_reply.started":"2023-12-25T10:56:20.452394Z","shell.execute_reply":"2023-12-25T10:56:20.466771Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class BalancedAccuracyCallback(Callback):\n    def __init__(self, validation_data):\n        super().__init__()\n        self.validation_data = validation_data\n\n    def on_epoch_end(self, epoch, logs=None):\n        y_true = self.validation_data[1]\n        y_pred = self.model.predict(self.validation_data[0])\n        y_true = np.argmax(y_true, axis=1)\n        y_pred = np.argmax(y_pred, axis=1)\n        balanced_acc = balanced_accuracy_score(y_true, y_pred)\n        print(f'Balanced Accuracy: {balanced_acc}')","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:56:20.4693Z","iopub.execute_input":"2023-12-25T10:56:20.470376Z","iopub.status.idle":"2023-12-25T10:56:20.485237Z","shell.execute_reply.started":"2023-12-25T10:56:20.47034Z","shell.execute_reply":"2023-12-25T10:56:20.483658Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def plot_wsi_images(train_df, label, num_images=5):\n    \"\"\"\n    Plot whole-slide images (WSI) for a specific cancer type along with descriptions.\n\n    Parameters:\n    - train_df: DataFrame, training dataset containing \"label\" and \"image_id\" columns.\n    - label: str, cancer label to plot.\n    - num_images: int, number of images to plot.\n    \"\"\"\n    df_tmp = train_df[train_df['label'] == label]\n    image_id_list = list(df_tmp[~df_tmp['is_tma']]['image_id'].sample(num_images))\n\n    plt.figure(figsize=(20.0, 6.0))\n    \n    for i, image_id in enumerate(image_id_list):\n        plt.subplot(1, num_images, i+1)\n        \n        if i == 0:\n            plt.title(f'image_id:{image_id} (WSI)', fontsize=14)\n            plt.ylabel(label, fontsize=14)\n        else:\n            plt.title(f'image_id:{image_id} (WSI)', fontsize=14)\n            \n        img_path = f'/kaggle/input/UBC-OCEAN/train_thumbnails/{image_id}_thumbnail.png'\n        io.imshow(img_path)\n        \n        plt.tick_params(labelbottom=False, labelleft=False, labelright=False, labeltop=False, \n                        bottom=False, left=False, right=False, top=False)\n    \n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:56:20.487131Z","iopub.execute_input":"2023-12-25T10:56:20.487666Z","iopub.status.idle":"2023-12-25T10:56:20.499836Z","shell.execute_reply.started":"2023-12-25T10:56:20.487625Z","shell.execute_reply":"2023-12-25T10:56:20.498628Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Read data","metadata":{}},{"cell_type":"code","source":"data_path = '/kaggle/input/UBC-OCEAN/'\n\ntrain_csv_path = os.path.join(data_path, 'train.csv')\ntrain_df = pd.read_csv(train_csv_path)\n\ntrain_images_path = os.path.join(data_path, 'train_images')\ntrain_thumbnails_path = os.path.join(data_path, 'train_thumbnails')\n\ntest_csv_path = os.path.join(data_path, 'test.csv')\ntest_df = pd.read_csv(test_csv_path)\n\ntest_images_path = os.path.join(data_path, 'test_images')\ntest_thumbnails_path = os.path.join(data_path, 'test_thumbnails')\n\nupdated_image_ids_path = os.path.join(data_path, 'updated_image_ids.json')","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:56:20.501268Z","iopub.execute_input":"2023-12-25T10:56:20.502447Z","iopub.status.idle":"2023-12-25T10:56:20.537838Z","shell.execute_reply.started":"2023-12-25T10:56:20.502401Z","shell.execute_reply":"2023-12-25T10:56:20.536823Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Train Data:\")\ndisplay(train_df.head())\n\nprint(\"\\nTest Data:\")\ndisplay(test_df.head())","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:56:20.539354Z","iopub.execute_input":"2023-12-25T10:56:20.540083Z","iopub.status.idle":"2023-12-25T10:56:20.576324Z","shell.execute_reply.started":"2023-12-25T10:56:20.540041Z","shell.execute_reply":"2023-12-25T10:56:20.575212Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In this case, we can use an interactive widget to realize more efficient display of images without loading them entirely into memory. This will allow us to view images interactively without a large memory load. This can be used if necessary","metadata":{}},{"cell_type":"code","source":"# image_width_cm = 4\n# image_height_cm = 4 ","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:56:20.578041Z","iopub.execute_input":"2023-12-25T10:56:20.578754Z","iopub.status.idle":"2023-12-25T10:56:20.58327Z","shell.execute_reply.started":"2023-12-25T10:56:20.578713Z","shell.execute_reply":"2023-12-25T10:56:20.582239Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# image_widget = widgets.Image(format='png', width=int(image_width_cm * 100), height=int(image_height_cm * 100))\n\n# def update_image_widget(selected_image_id):\n#     image_path = os.path.join(train_images_path, f'{selected_image_id}.png')\n#     image = Image.open(image_path)\n#     image_widget.value = open(image_path, 'rb').read()\n\n# image_selector = widgets.Dropdown(options=train_df['image_id'], description='Image ID:')\n\n# def on_image_select(change):\n#     selected_image_id = change['new']\n#     update_image_widget(selected_image_id)\n\n# image_selector.observe(on_image_select, names='value')\n\n# display(widgets.HBox([image_selector, image_widget]))","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:56:20.584698Z","iopub.execute_input":"2023-12-25T10:56:20.585434Z","iopub.status.idle":"2023-12-25T10:56:20.59536Z","shell.execute_reply.started":"2023-12-25T10:56:20.585373Z","shell.execute_reply":"2023-12-25T10:56:20.594047Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# image_widget = widgets.Image(format='png', width=int(image_width_cm * 100), height=int(image_height_cm * 100))\n\n# def update_image_widget(selected_image_id):\n#     image_path = os.path.join(test_images_path, f'{selected_image_id}.png')\n#     image = Image.open(image_path)\n#     image_widget.value = open(image_path, 'rb').read()\n\n# image_selector = widgets.Dropdown(options=test_df['image_id'], description='Image ID:')\n\n# def on_image_select(change):\n#     selected_image_id = change['new']\n#     update_image_widget(selected_image_id)\n\n# image_selector.observe(on_image_select, names='value')\n\n# display(widgets.HBox([image_selector, image_widget]))","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:56:20.597096Z","iopub.execute_input":"2023-12-25T10:56:20.597959Z","iopub.status.idle":"2023-12-25T10:56:20.611605Z","shell.execute_reply.started":"2023-12-25T10:56:20.597914Z","shell.execute_reply":"2023-12-25T10:56:20.610056Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Let's check the received data","metadata":{}},{"cell_type":"code","source":"dataframes = {\n    'Train Data': train_df,\n    'Test Data': test_df\n}","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:56:20.613311Z","iopub.execute_input":"2023-12-25T10:56:20.614427Z","iopub.status.idle":"2023-12-25T10:56:20.623192Z","shell.execute_reply.started":"2023-12-25T10:56:20.614384Z","shell.execute_reply":"2023-12-25T10:56:20.622299Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"explore_data(dataframes)","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:56:20.624734Z","iopub.execute_input":"2023-12-25T10:56:20.625397Z","iopub.status.idle":"2023-12-25T10:56:20.679534Z","shell.execute_reply.started":"2023-12-25T10:56:20.625359Z","shell.execute_reply":"2023-12-25T10:56:20.678542Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Train Data:\n\n`Shape:`\n\n>There are 538 records in the data\nThere are 5 columns\n\n`Info:`\n\n>image_id, label, image_width, image_height and is_tma - five columns\nlabel is of type object (string type)\nis_tma is of type bool (Boolean type)\nNo missing values in the data\n\n`Duplicates:`\n\n>No duplicates\n\n`Missing Values:`\n\n>No missing values\n\n### Test Data:\n\n`Shape:`\n\n>There is 1 record in the data\nThere are 3 columns\n\n`Info:`\n\n>image_id, image_width and image_height - three columns\nAll columns are of type int64\nNo missing values in the data\n\n`Duplicates:`\n\n>No duplicates\n\n`Missing Values:`\n\n>No missing values\n\n### Task:\n\nBased on the data provided, the task is to classify the types of ovarian cancer subtypes. Key columns for the classification task:\n\n>image_id: Unique image identifier\nlabel: Type of ovarian cancer (target variable)\n\nImportant additional columns:\n\n>image_width and image_height: Image sizes\nis_tma: Binary attribute indicating whether the image is a tissue microarray (TMA)","metadata":{}},{"cell_type":"markdown","source":"## Exploratory Data Analysis","metadata":{}},{"cell_type":"markdown","source":"### Let's display sample images for each label in the training data set","metadata":{}},{"cell_type":"code","source":"display_images_by_label(train_df, data_path)","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:56:20.68091Z","iopub.execute_input":"2023-12-25T10:56:20.681908Z","iopub.status.idle":"2023-12-25T10:56:54.465763Z","shell.execute_reply.started":"2023-12-25T10:56:20.681871Z","shell.execute_reply":"2023-12-25T10:56:54.464355Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Let's Visualize Image Sizes","metadata":{}},{"cell_type":"code","source":"# Statistics calculation\n\nwidth_stats = train_df['image_width'].describe()\nheight_stats = train_df['image_height'].describe()\n\n# Display statistics\n\nprint(\"Statistics for Image Widths:\")\nprint(width_stats)\n\nprint(\"\\nStatistics for Image Heights:\")\nprint(height_stats)","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:56:54.467575Z","iopub.execute_input":"2023-12-25T10:56:54.467913Z","iopub.status.idle":"2023-12-25T10:56:54.491505Z","shell.execute_reply.started":"2023-12-25T10:56:54.467884Z","shell.execute_reply":"2023-12-25T10:56:54.49025Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Visualizing image sizes\n\nplt.figure(figsize=(12, 6))\nplt.subplot(1, 2, 1)\nsns.histplot(train_df['image_width'], bins=20, kde=True, color='pink')\nplt.title('Distribution of Image Widths')\nax = plt.gca()\nax.set_facecolor(facecolor)\n\nplt.subplot(1, 2, 2)\nsns.histplot(train_df['image_height'], bins=20, kde=True)\nplt.title('Distribution of Image Heights')\nax = plt.gca()\nax.set_facecolor(facecolor)\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:56:54.49366Z","iopub.execute_input":"2023-12-25T10:56:54.494479Z","iopub.status.idle":"2023-12-25T10:56:55.594211Z","shell.execute_reply.started":"2023-12-25T10:56:54.494418Z","shell.execute_reply":"2023-12-25T10:56:55.593068Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display statistics\n\nstatistics_by_class = train_df.groupby(\"label\")[[\"image_width\", \"image_height\"]].describe().transpose()\nprint(\"Statistics by Class:\")\nprint(statistics_by_class)","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:56:55.595627Z","iopub.execute_input":"2023-12-25T10:56:55.595951Z","iopub.status.idle":"2023-12-25T10:56:55.643503Z","shell.execute_reply.started":"2023-12-25T10:56:55.595922Z","shell.execute_reply":"2023-12-25T10:56:55.642529Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_image_dimensions(train_df, class_colors)","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:56:55.644955Z","iopub.execute_input":"2023-12-25T10:56:55.645955Z","iopub.status.idle":"2023-12-25T10:56:57.958747Z","shell.execute_reply.started":"2023-12-25T10:56:55.645913Z","shell.execute_reply":"2023-12-25T10:56:57.957588Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_pairplot_by_class(train_df, class_colors)","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:56:57.96019Z","iopub.execute_input":"2023-12-25T10:56:57.960594Z","iopub.status.idle":"2023-12-25T10:57:06.768763Z","shell.execute_reply.started":"2023-12-25T10:56:57.960563Z","shell.execute_reply":"2023-12-25T10:57:06.767304Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### An analysis of image sizes depending on cancer types was carried out. Below are statistics for each type of cancer:  \n\n`Ovarian Cellular Cancer - Clear Cell (CC):`  \n\n>Average image width: 52,992 pixels  \nAverage image height: 31,205 pixels  \nWidth standard deviation: 21.069 pixels  \nHeight standard deviation: 10.439 pixels  \n\n`Endometrioid Cancer (EC):`  \n\n>Average image width: 47.486 pixels  \nAverage image height: 29.935 pixels  \nWidth standard deviation: 19.315 pixels  \nHeight standard deviation: 10.485 pixels  \n\n`High Grade Serous Carcinoma (HGSC):`  \n\n>Average image width: 48.637 pixels  \nAverage image height: 28,939 pixels  \nWidth standard deviation: 19,250 pixels  \nHeight standard deviation: 9.605 pixels  \n\n`Low Grade Serous Carcinoma (LGSC):`  \n\n>Average image width: 43,519 pixels  \nAverage image height: 24,774 pixels  \nWidth standard deviation: 20.967 pixels  \nHeight standard deviation: 11.954 pixels  \n\n`Mucous Carcinoma (MC):`  \n\n>Average image width: 50,193 pixels  \nAverage image height: 34,872 pixels  \nWidth standard deviation: 21,504 pixels  \nHeight standard deviation: 13.589 pixels  \n\n#### Unique characteristics of each type of cancer:  \nDifferent types of cancer exhibit unique image sizes, both in width and height. This may be due to the morphological characteristics of tumors of each specific type  \n\n#### Variety of sizes within each category:  \nThere is significant variation in image sizes within each cancer category. This indicates differences in the size and shape of tumors even within the same disease group  \n\n#### Visual interpretation strategies:  \nThe presented graphs and statistics provide clinicians and researchers with additional information to develop strategies for visual interpretation of medical images. Understanding the variability in tumor size across cancer types can help guide diagnosis and guide treatment approaches  ","metadata":{}},{"cell_type":"markdown","source":"### Let's look at the distribution of classes (cancer types) in the training data set","metadata":{}},{"cell_type":"code","source":"class_distribution = train_df['label'].value_counts()\nprint(\"\\nDistribution of Cancer Types in Train Data:\\n\")\nprint(class_distribution)","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:57:06.781337Z","iopub.execute_input":"2023-12-25T10:57:06.781953Z","iopub.status.idle":"2023-12-25T10:57:06.791161Z","shell.execute_reply.started":"2023-12-25T10:57:06.781916Z","shell.execute_reply":"2023-12-25T10:57:06.789658Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There is an imbalance of classes, where some classes are represented by many more examples than others. This can impact model performance, especially if the model cannot learn effectively on less represented classes  \n\nIn the case of class imbalance, it is important to pay attention to choosing an appropriate evaluation metric and taking appropriate measures to combat the imbalance  ","metadata":{}},{"cell_type":"code","source":"# Distribution of classes in train_df\n\nplt.figure(figsize=(10, 6))\nsns.countplot(x='label', data=train_df,palette=colors )\nplt.title('Distribution of Cancer Types in Train Data')\nplt.xlabel('Cancer Type')\nplt.ylabel('Count')\nax = plt.gca()\nax.set_facecolor(facecolor)\n# ax.grid(color=gridcolor)\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:57:06.79282Z","iopub.execute_input":"2023-12-25T10:57:06.793405Z","iopub.status.idle":"2023-12-25T10:57:07.136214Z","shell.execute_reply.started":"2023-12-25T10:57:06.79337Z","shell.execute_reply":"2023-12-25T10:57:07.134654Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Analysis of Class Distribution in Training Data\n\n#### 1. Class Distribution\n\nThe distribution of cancer types in the training data is as follows:\n\n>HGSC (High Grade Serous Carcinoma): 222 images\nEC (Endometrioid): 124 images\nCC (Ovarian Cellular Cancer - Clear Cell): 99 images\nLGSC (Low Grade Serous): 47 images\nMC (Mucous Carcinoma): 46 images\n\n#### 2. Conclusions and Scientific Notes\n\nCancer Subtypes Presented:\n\n`HGSC (High Grade Serous Carcinoma):`\n\n>The most common type of cancer in the training data\nThe species is associated with aggressive growth and is often diagnosed in advanced stages\nOften presents a challenge for accurate diagnosis\n\n`EC (Endometrioid):`\n\n>Second most common type of cancer in the data\nUsually associated with a hyperestrogenic state and often detected in early stages\n\n`CC (Ovarian Cellular Cancer - Clear Cell):`\n\n>Represents a comparatively smaller group\nCharacterized by a clear cellular structure and may have unique clinical characteristics\n\n`LGSC (Low Grade Serous):`\n\n>Represented in data, but comparatively less frequently\nLow grade types often grow slower but may be more difficult to treat\n\n`MC (Mucous Carcinoma):`\n\n>Another rare form of cancer represented in the data\nOften associated with mucus formation in tissues\n\n\n### HGSC Dominance\n\nWhy is HGSC Predominant?\n\n`Frequency in Real Clinical Cases:`\n\n>HGSC is often encountered in clinical practice due to its high frequency in the general structure of malignant ovarian tumors\nAccording to medical research reports, high-grade serous carcinoma accounts for a significant proportion of ovarian cancer cases\n\n`Diagnosis at Late Stages:`\n\n>One of the reasons for the predominance of HGSC in the training data set is its frequent diagnosis at later stages of development\nDelay in diagnosis may be due to the absence of specific clinical symptoms in the early stages or their similarity with other diseases","metadata":{}},{"cell_type":"code","source":"# Visualization of the presence of tissue microarrays (TMA)\n\nplt.figure(figsize=(6, 4))\nsns.countplot(x='is_tma', data=train_df,palette=colors)\nplt.title('Distribution of TMA (Tissue Microarray)')\nplt.xlabel('Is TMA')\nplt.ylabel('Count')\nax = plt.gca()\nax.set_facecolor(facecolor)\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:57:07.138654Z","iopub.execute_input":"2023-12-25T10:57:07.14006Z","iopub.status.idle":"2023-12-25T10:57:07.425144Z","shell.execute_reply.started":"2023-12-25T10:57:07.139994Z","shell.execute_reply":"2023-12-25T10:57:07.423701Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Based on the graph, which shows the majority of samples without tissue microarrays (TMAs), the following preliminary conclusions can be drawn:\n\n`Predominance of Samples without TMA:`\n\nThe large number of samples lacking TMA indicates that the study favors analysis of individual tissue samples rather than bulk analysis using TMA\n\n`Heterogeneity and Diversity:`\n\nThe absence of TMA may be due to the heterogeneity and diversity of tissue samples. Each sample is likely to have unique characteristics that require individual attention and analysis\n\n`More Careful Approach:`\n\nThe predominance of samples without TMA may indicate the need for a more thorough approach to the analysis of each sample. This may include additional preparation methods, detailed microscopic examination, and other analytical steps.","metadata":{}},{"cell_type":"markdown","source":"### Samples: train_images","metadata":{}},{"cell_type":"code","source":"# Descriptions for each cancer type\n\ndescriptions = {\n    'HGSC': 'The most common type of cancer in the training data.\\nThe species is associated with aggressive growth and is often diagnosed in advanced stages.\\nOften presents a challenge for accurate diagnosis.',\n    'EC': 'Second most common type of cancer in the data.\\nUsually associated with a hyperestrogenic state and often detected in early stages.',\n    'CC': 'Represents a comparatively smaller group.\\nCharacterized by a clear cellular structure and may have unique clinical characteristics.',\n    'LGSC': 'Represented in data, but comparatively less frequently.\\nLow grade types often grow slower but may be more difficult to treat.',\n    'MC': 'Another rare form of cancer represented in the data.\\nOften associated with mucus formation in tissues.'\n}\n\ncancer_labels_to_plot = ['HGSC', 'EC', 'CC', 'LGSC', 'MC']\n\nplot_cancer_images(train_df, cancer_labels_to_plot, descriptions)","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:57:07.426771Z","iopub.execute_input":"2023-12-25T10:57:07.42712Z","iopub.status.idle":"2023-12-25T10:58:06.465249Z","shell.execute_reply.started":"2023-12-25T10:57:07.42709Z","shell.execute_reply":"2023-12-25T10:58:06.46294Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Samples: train_thumbnails","metadata":{}},{"cell_type":"code","source":"descriptions = {\n    'HGSC': 'The most common type of cancer in the training data.\\nThe species is associated with aggressive growth and is often diagnosed in advanced stages.\\nOften presents a challenge for accurate diagnosis.',\n    'EC': 'Second most common type of cancer in the data.\\nUsually associated with a hyperestrogenic state and often detected in early stages.',\n    'CC': 'Represents a comparatively smaller group.\\nCharacterized by a clear cellular structure and may have unique clinical characteristics.',\n    'LGSC': 'Represented in data, but comparatively less frequently.\\nLow grade types often grow slower but may be more difficult to treat.',\n    'MC': 'Another rare form of cancer represented in the data.\\nOften associated with mucus formation in tissues.'\n}\n\ncancer_labels_to_plot = ['HGSC', 'CC', 'EC', 'LGSC', 'MC']\n\nfor label in cancer_labels_to_plot:\n    plot_wsi_images(train_df, label)\n    print(f'{label} ({descriptions[label]}):\\n')","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:58:06.467426Z","iopub.execute_input":"2023-12-25T10:58:06.467938Z","iopub.status.idle":"2023-12-25T10:58:37.573883Z","shell.execute_reply.started":"2023-12-25T10:58:06.467897Z","shell.execute_reply":"2023-12-25T10:58:37.5729Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Even a non-medical analyst can see how different the image structures are for each type of cancer.  \n\nFor example:  \n\n`\"HGSC\":` The most common type of cancer in the training data. This species is associated with aggressive growth and is often diagnosed in advanced stages. Often poses a problem for accurate diagnosis.  \n`\"EC\":` The second most common type of cancer in the data. Typically associated with a hyperestrogenic state and often detected in the early stages.  \n`\"CC\":` Represents a comparatively smaller group. It is characterized by a distinct cellular structure and may have unique clinical characteristics.  \n`\"LGSC\":` Present in the data, but comparatively less frequently. Low-grade types often grow slower but are more difficult to treat.  \n`\"MC\":` Another rare form of cancer represented in the data. Often associated with the formation of mucus in tissues. And there is obviously a lot of liquid in the pictures.  ","metadata":{}},{"cell_type":"markdown","source":"## Data preparation","metadata":{}},{"cell_type":"code","source":"# # Determining the number of samples for training and test\n\nnum_samples = 3000 # Set the number of samples for training\nnum_samples_test = 600 # Set the number of samples for the test\n\n# # Determine the number of classes and image sizes\n\n# num_classes = 10 # Set the number of classes\n# image_height = 224 # Set the height of the image\n# image_width = 224 # Set the width of the image\n# num_channels = 3 # Set the number of color channels (RGB)","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:58:37.57515Z","iopub.execute_input":"2023-12-25T10:58:37.576135Z","iopub.status.idle":"2023-12-25T10:58:37.582776Z","shell.execute_reply.started":"2023-12-25T10:58:37.576098Z","shell.execute_reply":"2023-12-25T10:58:37.581352Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Attempt 2: \n\n# Determine the number of classes and image sizes\n\nnum_classes = 10 # Set the number of classes\nimage_height = 256 # Set the height of the image\nimage_width = 256 # Set the width of the image\nnum_channels = 3 # Set the number of color channels (RGB)","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:58:37.584417Z","iopub.execute_input":"2023-12-25T10:58:37.584842Z","iopub.status.idle":"2023-12-25T10:58:37.606318Z","shell.execute_reply.started":"2023-12-25T10:58:37.584799Z","shell.execute_reply":"2023-12-25T10:58:37.604811Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create random images for training\n\nX_train = np.random.rand(num_samples, image_height, image_width, num_channels)\ny_train = np.random.randint(0, num_classes, size=num_samples)\n\n# Create random images for the test\n\nX_test = np.random.rand(num_samples_test, image_height, image_width, num_channels)\ny_test = np.random.randint(0, num_classes, size=num_samples_test)","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:58:37.60823Z","iopub.execute_input":"2023-12-25T10:58:37.608982Z","iopub.status.idle":"2023-12-25T10:58:46.320718Z","shell.execute_reply.started":"2023-12-25T10:58:37.608718Z","shell.execute_reply":"2023-12-25T10:58:46.319444Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Generate random labels for testing\n\ny_train_encoded = to_categorical(y_train)\ny_test_encoded = to_categorical(y_test)","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:58:46.322224Z","iopub.execute_input":"2023-12-25T10:58:46.322604Z","iopub.status.idle":"2023-12-25T10:58:46.329502Z","shell.execute_reply.started":"2023-12-25T10:58:46.322573Z","shell.execute_reply":"2023-12-25T10:58:46.328075Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# y_train_encoded","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:58:46.331631Z","iopub.execute_input":"2023-12-25T10:58:46.332074Z","iopub.status.idle":"2023-12-25T10:58:46.341031Z","shell.execute_reply.started":"2023-12-25T10:58:46.33204Z","shell.execute_reply":"2023-12-25T10:58:46.339608Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train, X_val, y_train_encoded, y_val_encoded = train_test_split(\n    X_train, y_train_encoded, test_size=0.2, random_state=1234567, stratify=y_train_encoded\n)","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:58:46.342731Z","iopub.execute_input":"2023-12-25T10:58:46.343806Z","iopub.status.idle":"2023-12-25T10:58:49.270681Z","shell.execute_reply.started":"2023-12-25T10:58:46.343771Z","shell.execute_reply":"2023-12-25T10:58:49.269363Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train.shape, X_val.shape, y_train_encoded.shape, y_val_encoded.shape","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:58:49.272443Z","iopub.execute_input":"2023-12-25T10:58:49.272953Z","iopub.status.idle":"2023-12-25T10:58:49.282122Z","shell.execute_reply.started":"2023-12-25T10:58:49.272921Z","shell.execute_reply":"2023-12-25T10:58:49.280575Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Unique classes in the training set:\", np.unique(np.argmax(y_train_encoded, axis=1)))\nprint(\"Unique classes in the validation set:\", np.unique(np.argmax(y_val_encoded, axis=1)))","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:58:49.284235Z","iopub.execute_input":"2023-12-25T10:58:49.284708Z","iopub.status.idle":"2023-12-25T10:58:49.293888Z","shell.execute_reply.started":"2023-12-25T10:58:49.284666Z","shell.execute_reply":"2023-12-25T10:58:49.292408Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# datagen = ImageDataGenerator(\n#     rotation_range=20,\n#     width_shift_range=0.2,\n#     height_shift_range=0.2,\n#     shear_range=0.2,\n#     zoom_range=0.2,\n#     horizontal_flip=True,\n#     fill_mode='nearest'\n# )","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:58:49.295681Z","iopub.execute_input":"2023-12-25T10:58:49.296778Z","iopub.status.idle":"2023-12-25T10:58:49.30171Z","shell.execute_reply.started":"2023-12-25T10:58:49.29673Z","shell.execute_reply":"2023-12-25T10:58:49.300774Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Attempt 2 \n\nbatch_size = 20\nsteps_per_epoch = len(X_train) // batch_size","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:58:49.303371Z","iopub.execute_input":"2023-12-25T10:58:49.303865Z","iopub.status.idle":"2023-12-25T10:58:49.312889Z","shell.execute_reply.started":"2023-12-25T10:58:49.303834Z","shell.execute_reply":"2023-12-25T10:58:49.31122Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Attempt 2:\n\ndatagen = ImageDataGenerator(\n    rotation_range=50,\n    width_shift_range=0.2,\n    height_shift_range=0.2,\n    shear_range=0.2,\n    zoom_range=0.5,\n    horizontal_flip=True,\n    fill_mode='nearest'\n)","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:58:49.314917Z","iopub.execute_input":"2023-12-25T10:58:49.31538Z","iopub.status.idle":"2023-12-25T10:58:49.327452Z","shell.execute_reply.started":"2023-12-25T10:58:49.315347Z","shell.execute_reply":"2023-12-25T10:58:49.326184Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_generator = datagen.flow(X_train, y_train_encoded, batch_size=batch_size)","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:58:49.329273Z","iopub.execute_input":"2023-12-25T10:58:49.329649Z","iopub.status.idle":"2023-12-25T10:58:50.531804Z","shell.execute_reply.started":"2023-12-25T10:58:49.329619Z","shell.execute_reply":"2023-12-25T10:58:50.530298Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Model training","metadata":{}},{"cell_type":"code","source":"model = Sequential()\nmodel.add(Conv2D(32, (3, 3), activation='relu', input_shape=(image_height, image_width, num_channels)))\nmodel.add(MaxPooling2D((2, 2)))\nmodel.add(Conv2D(64, (3, 3), activation='relu'))\nmodel.add(MaxPooling2D((2, 2)))\nmodel.add(Conv2D(128, (3, 3), activation='relu'))\nmodel.add(MaxPooling2D((2, 2)))\nmodel.add(Flatten())\nmodel.add(Dense(128, activation='relu'))\nmodel.add(Dropout(0.5))\nmodel.add(Dense(num_classes, activation='softmax'))","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:58:50.533483Z","iopub.execute_input":"2023-12-25T10:58:50.533937Z","iopub.status.idle":"2023-12-25T10:58:50.908513Z","shell.execute_reply.started":"2023-12-25T10:58:50.5339Z","shell.execute_reply":"2023-12-25T10:58:50.907014Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# # Attempt 2:\n\n# base_model = ResNet50(weights='imagenet', include_top=False, input_shape=(image_height, image_width, num_channels))","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:58:50.910315Z","iopub.execute_input":"2023-12-25T10:58:50.910756Z","iopub.status.idle":"2023-12-25T10:58:50.916687Z","shell.execute_reply.started":"2023-12-25T10:58:50.91072Z","shell.execute_reply":"2023-12-25T10:58:50.915184Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# # Attempt 2:\n\n# base_model.trainable = False","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:58:50.91842Z","iopub.execute_input":"2023-12-25T10:58:50.918862Z","iopub.status.idle":"2023-12-25T10:58:50.92902Z","shell.execute_reply.started":"2023-12-25T10:58:50.918828Z","shell.execute_reply":"2023-12-25T10:58:50.927358Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# # Attempt 2:\n\n# model = Sequential()\n# model.add(base_model)\n# model.add(layers.GlobalAveragePooling2D())\n# model.add(layers.Dense(256, activation='relu'))\n# model.add(layers.Dropout(0.5))\n# model.add(layers.Dense(num_classes, activation='softmax'))","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:58:50.930725Z","iopub.execute_input":"2023-12-25T10:58:50.931089Z","iopub.status.idle":"2023-12-25T10:58:50.939456Z","shell.execute_reply.started":"2023-12-25T10:58:50.931058Z","shell.execute_reply":"2023-12-25T10:58:50.938329Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# # Attempt 3:\n\n# My kernel crashed, but you can try to balance the classes\n\n# class_weights = compute_class_weight('balanced', np.unique(y_train), y_train)\n\n# class_weight_dict = dict(enumerate(class_weights))\n\n# model.compile(optimizer=Adam(learning_rate=0.001),\n#               loss='categorical_crossentropy',\n#               metrics=['accuracy'],\n#               class_weight=class_weight_dict)","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:58:50.940989Z","iopub.execute_input":"2023-12-25T10:58:50.94162Z","iopub.status.idle":"2023-12-25T10:58:50.950964Z","shell.execute_reply.started":"2023-12-25T10:58:50.941581Z","shell.execute_reply":"2023-12-25T10:58:50.949364Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.compile(optimizer=Adam(learning_rate=0.0001),\n              loss='categorical_crossentropy',\n              metrics=['accuracy'])","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:58:50.952648Z","iopub.execute_input":"2023-12-25T10:58:50.95384Z","iopub.status.idle":"2023-12-25T10:58:50.982155Z","shell.execute_reply.started":"2023-12-25T10:58:50.953805Z","shell.execute_reply":"2023-12-25T10:58:50.9808Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# print(\"Prediction shape before training:\", model.predict(X_val).shape)","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:58:50.984883Z","iopub.execute_input":"2023-12-25T10:58:50.986089Z","iopub.status.idle":"2023-12-25T10:58:50.990765Z","shell.execute_reply.started":"2023-12-25T10:58:50.986038Z","shell.execute_reply":"2023-12-25T10:58:50.989508Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# weights_path = '/kaggle/working/resnet50_weights_tf_dim_ordering_tf_kernels_notop.h5'","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:58:50.992689Z","iopub.execute_input":"2023-12-25T10:58:50.993251Z","iopub.status.idle":"2023-12-25T10:58:51.006156Z","shell.execute_reply.started":"2023-12-25T10:58:50.993205Z","shell.execute_reply":"2023-12-25T10:58:51.004425Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Attempt 3:\n\n# base_model = ResNet50(weights=None, include_top=False, input_shape=(image_height, image_width, num_channels))","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:58:51.008076Z","iopub.execute_input":"2023-12-25T10:58:51.008555Z","iopub.status.idle":"2023-12-25T10:58:51.014813Z","shell.execute_reply.started":"2023-12-25T10:58:51.008514Z","shell.execute_reply":"2023-12-25T10:58:51.013566Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Attempt 3:\n\n# base_model.load_weights(weights_path, by_name=True)","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:58:51.016737Z","iopub.execute_input":"2023-12-25T10:58:51.017165Z","iopub.status.idle":"2023-12-25T10:58:51.031752Z","shell.execute_reply.started":"2023-12-25T10:58:51.017125Z","shell.execute_reply":"2023-12-25T10:58:51.030183Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Attempt 3:\n\n# model = Sequential()\n# model.add(base_model)\n# model.add(GlobalAveragePooling2D())\n# model.add(Dense(128, activation='relu'))\n# model.add(Dense(num_classes, activation='softmax'))","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:58:51.033702Z","iopub.execute_input":"2023-12-25T10:58:51.034154Z","iopub.status.idle":"2023-12-25T10:58:51.044114Z","shell.execute_reply.started":"2023-12-25T10:58:51.034112Z","shell.execute_reply":"2023-12-25T10:58:51.042748Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Attempt 3:\n\n# model.compile(optimizer='adam', loss='categorical_crossentropy', metrics=['accuracy'])","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:58:51.045829Z","iopub.execute_input":"2023-12-25T10:58:51.046274Z","iopub.status.idle":"2023-12-25T10:58:51.054909Z","shell.execute_reply.started":"2023-12-25T10:58:51.046232Z","shell.execute_reply":"2023-12-25T10:58:51.053917Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.summary()","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:58:51.056069Z","iopub.execute_input":"2023-12-25T10:58:51.056537Z","iopub.status.idle":"2023-12-25T10:58:51.09691Z","shell.execute_reply.started":"2023-12-25T10:58:51.056495Z","shell.execute_reply":"2023-12-25T10:58:51.09096Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"early_stopping = EarlyStopping(monitor='val_loss', patience=3, restore_best_weights=True)\nmodel_checkpoint = ModelCheckpoint('/kaggle/working/best_model.h5', monitor='val_loss', save_best_only=True)","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:58:51.098555Z","iopub.execute_input":"2023-12-25T10:58:51.098937Z","iopub.status.idle":"2023-12-25T10:58:51.10557Z","shell.execute_reply.started":"2023-12-25T10:58:51.098907Z","shell.execute_reply":"2023-12-25T10:58:51.104164Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"balanced_accuracy_callback = BalancedAccuracyCallback(validation_data=(X_val, y_val_encoded))","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:58:51.107254Z","iopub.execute_input":"2023-12-25T10:58:51.107719Z","iopub.status.idle":"2023-12-25T10:58:51.116683Z","shell.execute_reply.started":"2023-12-25T10:58:51.107686Z","shell.execute_reply":"2023-12-25T10:58:51.115185Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# history_gen = model.fit(\n#     train_generator,\n#     epochs=10,\n#     validation_data=(X_val, y_val_encoded),\n#     callbacks=[early_stopping, model_checkpoint]\n# )","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:58:51.118798Z","iopub.execute_input":"2023-12-25T10:58:51.119234Z","iopub.status.idle":"2023-12-25T10:58:51.127688Z","shell.execute_reply.started":"2023-12-25T10:58:51.1192Z","shell.execute_reply":"2023-12-25T10:58:51.126449Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# y_train_indices","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:58:51.128999Z","iopub.execute_input":"2023-12-25T10:58:51.129364Z","iopub.status.idle":"2023-12-25T10:58:51.140294Z","shell.execute_reply.started":"2023-12-25T10:58:51.129335Z","shell.execute_reply":"2023-12-25T10:58:51.138845Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Attempt 2: \n\nepochs = 5\n\nhistory_gen = model.fit(\n    train_generator,\n    epochs=epochs,\n    validation_data=(X_val, y_val_encoded),\n    callbacks=[early_stopping, model_checkpoint, balanced_accuracy_callback]\n)","metadata":{"execution":{"iopub.status.busy":"2023-12-25T10:58:51.142575Z","iopub.execute_input":"2023-12-25T10:58:51.143217Z","iopub.status.idle":"2023-12-25T11:22:47.119326Z","shell.execute_reply.started":"2023-12-25T10:58:51.14318Z","shell.execute_reply":"2023-12-25T11:22:47.116546Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(history_gen.history)","metadata":{"execution":{"iopub.status.busy":"2023-12-25T11:22:47.123749Z","iopub.execute_input":"2023-12-25T11:22:47.124272Z","iopub.status.idle":"2023-12-25T11:22:47.135122Z","shell.execute_reply.started":"2023-12-25T11:22:47.124232Z","shell.execute_reply":"2023-12-25T11:22:47.133485Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.save('/kaggle/working/model_UBC_OCSCO.keras')","metadata":{"execution":{"iopub.status.busy":"2023-12-25T11:22:47.137014Z","iopub.execute_input":"2023-12-25T11:22:47.138356Z","iopub.status.idle":"2023-12-25T11:22:47.61651Z","shell.execute_reply.started":"2023-12-25T11:22:47.138316Z","shell.execute_reply":"2023-12-25T11:22:47.615094Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Verifying the model on test data","metadata":{}},{"cell_type":"code","source":"sample_indices = random.sample(range(len(X_test)), 3)  \nsample_images = X_test[sample_indices]\nsample_true_labels = y_test[sample_indices]\n\nsample_pred_probs = model.predict(sample_images)\nsample_pred_labels = np.argmax(sample_pred_probs, axis=1)","metadata":{"execution":{"iopub.status.busy":"2023-12-25T11:22:47.618669Z","iopub.execute_input":"2023-12-25T11:22:47.619048Z","iopub.status.idle":"2023-12-25T11:22:47.766755Z","shell.execute_reply.started":"2023-12-25T11:22:47.619017Z","shell.execute_reply":"2023-12-25T11:22:47.765315Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"True Labels:\")\nprint(sample_true_labels)\n\nprint(\"\\nPredicted Labels:\")\nprint(sample_pred_labels)","metadata":{"execution":{"iopub.status.busy":"2023-12-25T11:22:47.768539Z","iopub.execute_input":"2023-12-25T11:22:47.769142Z","iopub.status.idle":"2023-12-25T11:22:47.777313Z","shell.execute_reply.started":"2023-12-25T11:22:47.769102Z","shell.execute_reply":"2023-12-25T11:22:47.775998Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission = pd.read_csv('/kaggle/input/UBC-OCEAN/sample_submission.csv')\nsubmission.head()","metadata":{"execution":{"iopub.status.busy":"2023-12-25T11:22:47.779233Z","iopub.execute_input":"2023-12-25T11:22:47.779712Z","iopub.status.idle":"2023-12-25T11:22:47.831632Z","shell.execute_reply.started":"2023-12-25T11:22:47.779672Z","shell.execute_reply":"2023-12-25T11:22:47.830429Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_pred_labels = np.argmax(sample_pred_probs, axis=1)\n\nsubmission = pd.read_csv('/kaggle/input/UBC-OCEAN/sample_submission.csv')\n\nsubmission['label'] = sample_pred_labels[:len(submission)]","metadata":{"execution":{"iopub.status.busy":"2023-12-25T11:22:47.833361Z","iopub.execute_input":"2023-12-25T11:22:47.834489Z","iopub.status.idle":"2023-12-25T11:22:47.844941Z","shell.execute_reply.started":"2023-12-25T11:22:47.83441Z","shell.execute_reply":"2023-12-25T11:22:47.843073Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.to_csv('/kaggle/working/submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2023-12-25T11:22:47.846908Z","iopub.execute_input":"2023-12-25T11:22:47.847725Z","iopub.status.idle":"2023-12-25T11:22:47.898628Z","shell.execute_reply.started":"2023-12-25T11:22:47.847686Z","shell.execute_reply":"2023-12-25T11:22:47.897031Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission","metadata":{"execution":{"iopub.status.busy":"2023-12-25T11:22:47.901067Z","iopub.execute_input":"2023-12-25T11:22:47.901493Z","iopub.status.idle":"2023-12-25T11:22:47.913535Z","shell.execute_reply.started":"2023-12-25T11:22:47.901441Z","shell.execute_reply":"2023-12-25T11:22:47.912205Z"},"trusted":true},"execution_count":null,"outputs":[]}]}