{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# UBC-OCEAN Exploratory Data Analysis\n\nIn this notebook, we will conduct an exploratory data analysis (EDA) on the UBC-OCEAN dataset for the task of classifying the type of ovarian cancer from microscopy scans of biopsy samples. The goal is to understand the distribution of the data, and identify any potential challenges or patterns that may be relevant for the classification problem.","metadata":{}},{"cell_type":"code","source":"# Import packages\nimport numpy as pd\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport cv2","metadata":{"execution":{"iopub.status.busy":"2023-11-05T06:11:13.968646Z","iopub.execute_input":"2023-11-05T06:11:13.969295Z","iopub.status.idle":"2023-11-05T06:11:16.528234Z","shell.execute_reply.started":"2023-11-05T06:11:13.969256Z","shell.execute_reply":"2023-11-05T06:11:16.525382Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data Description\nLet's read in the data for the training set and start exploring it. These are the field definitions provided by the competition Dataset Description:\n- **image_id**: A unique ID code for each image.\n- **label** - The target class. One of these subtypes of ovarian cancer: CC, EC, HGSC, LGSC, MC, Other. The Other class is not present in the training set; identifying outliers is one of the challenges of this competition. Only available for the train set.\n- **image_width** - The image width in pixels.\n- **image_height** - The image height in pixels.\n- **is_tma** - True if the slide is a tissue microarray. Only available for the train set.","metadata":{}},{"cell_type":"code","source":"train_df = pd.read_csv('/kaggle/input/UBC-OCEAN/train.csv')\ntrain_df.info()","metadata":{"execution":{"iopub.status.busy":"2023-11-05T06:11:16.530602Z","iopub.execute_input":"2023-11-05T06:11:16.531397Z","iopub.status.idle":"2023-11-05T06:11:16.595409Z","shell.execute_reply.started":"2023-11-05T06:11:16.531356Z","shell.execute_reply":"2023-11-05T06:11:16.59431Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The training dataset consists of 538 images. The data type of the **image_id** column is int64 when it should be object (string) as it represents an identifier. This is addressed in the next code cell:","metadata":{}},{"cell_type":"code","source":"train_df['image_id'] = train_df['image_id'].astype(str)\ntrain_df.info()","metadata":{"execution":{"iopub.status.busy":"2023-11-05T06:11:16.597182Z","iopub.execute_input":"2023-11-05T06:11:16.597856Z","iopub.status.idle":"2023-11-05T06:11:16.613448Z","shell.execute_reply.started":"2023-11-05T06:11:16.597818Z","shell.execute_reply":"2023-11-05T06:11:16.612205Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now let's call df.describe() to gain high-level insight on our data.","metadata":{}},{"cell_type":"code","source":"train_df.describe(include='all')","metadata":{"execution":{"iopub.status.busy":"2023-11-05T06:11:16.616762Z","iopub.execute_input":"2023-11-05T06:11:16.617622Z","iopub.status.idle":"2023-11-05T06:11:16.659199Z","shell.execute_reply.started":"2023-11-05T06:11:16.617577Z","shell.execute_reply":"2023-11-05T06:11:16.65786Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can make a several observations from the DataFrame description:\n- The most prevalent label is **HGSC (n=222)**.\n- image_width and image_height range from 2,964-10,5763 pixels and 2,964-50,155 pixels, respectively.\n- There are many more **Whole Slide Images (n=513)** than there are **Tissue Microarrays (n=25)**.\n\n----","metadata":{}},{"cell_type":"code","source":"train_df['image_type'] = train_df['is_tma'].apply(lambda x: 'TMA' if x else 'WSI')\ntrain_df['image_type'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2023-11-05T06:11:16.660559Z","iopub.execute_input":"2023-11-05T06:11:16.660858Z","iopub.status.idle":"2023-11-05T06:11:16.672867Z","shell.execute_reply.started":"2023-11-05T06:11:16.660832Z","shell.execute_reply":"2023-11-05T06:11:16.672012Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Label Distribution Analysis\n\nNext, let's examine the distribution of the class labels.","metadata":{}},{"cell_type":"code","source":"# Visualize number of samples of each label\nplt.figure(figsize=(6, 3))\nsns.countplot(\n    data=train_df, \n    y='label', \n    order=train_df['label'].value_counts().index,\n    palette='viridis')\n\nplt.xlabel('Number of Samples')\nplt.ylabel('Cancer Subtype')\nplt.grid(True)\n\nplt.show()\n\n# Show proportions of labels\nlabel_proportions = train_df['label'].value_counts(normalize=True)*100\nprint(label_proportions)","metadata":{"execution":{"iopub.status.busy":"2023-11-05T06:11:16.674454Z","iopub.execute_input":"2023-11-05T06:11:16.67501Z","iopub.status.idle":"2023-11-05T06:11:16.946323Z","shell.execute_reply.started":"2023-11-05T06:11:16.674982Z","shell.execute_reply":"2023-11-05T06:11:16.94521Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The above shows an unequal distribution of labels, indicating the presence of **Class Imbalance**. Class imbalance poses the potential of a model to be biased towards the more common classes relative to the less common classes. In the context of our classification task, a model may not learn sufficiently to accurately predict the less common labels, LGSC and MC, as a consequence.\n\nThere are several strategies we can consider to tackle class imbalance:\n- **Oversampling:** Generate additional copies of images in the minority classes.\n- **Class Weighting:** Assign weights to classes that are proportional to their frequency in the dataset during training so that the model is encouraged to pay more attention to the minority classes.\n- **Data Augmentation:** Generate additional samples of the minority classes by randomly applying some transformation to images such as rotating and flipping.","metadata":{}},{"cell_type":"markdown","source":"Earlier we noted that there are only 25 TMAs - we should check whether all of the class labels are present in this subset. If there are only one or two TMA samples of a certain label, or if a label is missing entirely within the TMAs, this would spell considerable challenge in our model being able to predict that label on TMA samples.","metadata":{}},{"cell_type":"code","source":"# Visualize number of samples of each label within the TMA images\ntma_df = train_df[train_df['is_tma']]\nplt.figure(figsize=(6, 3))\nsns.countplot(\n    data=tma_df, \n    y='label', \n    order=tma_df['label'].value_counts().index,\n    palette='viridis')\n\nplt.xlabel('Number of Samples')\nplt.ylabel('Cancer Subtype')\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-11-05T06:11:16.948044Z","iopub.execute_input":"2023-11-05T06:11:16.948876Z","iopub.status.idle":"2023-11-05T06:11:17.181427Z","shell.execute_reply.started":"2023-11-05T06:11:16.948836Z","shell.execute_reply":"2023-11-05T06:11:17.179916Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"All of the class labels are present at equal distribution within the 25 TMA samples.","metadata":{}},{"cell_type":"markdown","source":"## Image Size Analysis\n\nNext, let's plot the image_width and image_height in a scatteplot to understand the distribution of image sizes. We also color the points to examine the relationships between image size vs. image type (left) and vs. label (right).","metadata":{}},{"cell_type":"code","source":"fig, axes = plt.subplots(1, 2, figsize=(12, 4))\n\n# Scatterplot of image_height vs image_width, colored by image_type\nscatter1 = sns.scatterplot(data=train_df, x='image_width', y='image_height', hue='is_tma', ax=axes[0])\naxes[1].set_xlabel('Image Width')\naxes[1].set_ylabel('Image Height')\naxes[0].set_aspect('equal', adjustable='box')\naxes[0].set_title('Image Sizes Colored by Image Type')\n\nlegend1 = scatter1.legend(loc='center left', bbox_to_anchor=(1, 0.5))\n\n# Scatterplot of image_height vs image_width, colored by label\nscatter2 = sns.scatterplot(data=train_df, x='image_width', y='image_height', hue='label', ax=axes[1])\naxes[1].set_xlabel('Image Width')\naxes[1].set_ylabel('Image Height')\naxes[1].set_aspect('equal', adjustable='box')\naxes[1].set_title('Image Sizes Colored by Label')\n\nlegend2 = scatter2.legend(loc='center left', bbox_to_anchor=(1, 0.5))\n\nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-11-05T06:11:17.18278Z","iopub.execute_input":"2023-11-05T06:11:17.183149Z","iopub.status.idle":"2023-11-05T06:11:18.163486Z","shell.execute_reply.started":"2023-11-05T06:11:17.183104Z","shell.execute_reply":"2023-11-05T06:11:18.162567Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- The left scatterplot confirms that the images for TMAs range around 4000x4000 pixels, while the WSIs are typically much larger and vary widely in their dimensions. Notably, there are WSIs whose widths are much greater than their heights and vice versa.\n- The right scatterplot shows that the labels are randomly distributed in their image sizes. ","metadata":{"execution":{"iopub.status.busy":"2023-11-05T05:18:27.920118Z","iopub.execute_input":"2023-11-05T05:18:27.920451Z","iopub.status.idle":"2023-11-05T05:18:27.926678Z","shell.execute_reply.started":"2023-11-05T05:18:27.920423Z","shell.execute_reply":"2023-11-05T05:18:27.92553Z"}}},{"cell_type":"markdown","source":"## Image Type Analysis\n\nThe dataset is comprised of two types of images.\n- **Whole Slide Image (WSI):** High resolution images produced from the digitization of conventional glass slides. WSIs can be considerably larger than TMAs and are at 20x magnification \n- **Tissue Microarrays (TMA):** Images of small tissue cores extracted in a paraffin block. TMAs are roughly 4,000 x 4,000 pixels and are at 40x magnification\n\nLet's look at some examples of each.","metadata":{}},{"cell_type":"code","source":"# Show 4 examples of TMAs\ntma_examples = train_df[train_df['is_tma']]['image_id'].sample(n=4, random_state=1)\n\nfig, axes = plt.subplots(1, 4, figsize=(12, 3))\nfor i, id in enumerate(tma_examples):\n    img = cv2.imread(f'/kaggle/input/UBC-OCEAN/train_images/{id}.png')\n    \n    axes[i].imshow(img)\n    axes[i].set_title(id)\n    axes[i].axis('off')\n\nplt.suptitle('Example TMAs')\nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-11-05T06:11:18.164439Z","iopub.execute_input":"2023-11-05T06:11:18.164723Z","iopub.status.idle":"2023-11-05T06:11:26.316405Z","shell.execute_reply.started":"2023-11-05T06:11:18.164696Z","shell.execute_reply":"2023-11-05T06:11:26.315322Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Show 4 examples of WSIs\n# Thumbnails are used here since they are much faster to load\nwsi_examples = train_df[~train_df['is_tma']]['image_id'].sample(n=4, random_state=1)\n\nfig, axes = plt.subplots(1, 4, figsize=(12, 3))\nfor i, id in enumerate(wsi_examples):\n    img = cv2.imread(f'/kaggle/input/UBC-OCEAN/train_thumbnails/{id}_thumbnail.png')\n    \n    axes[i].imshow(img)\n    axes[i].set_title(id)\n    axes[i].axis('off')\n    \nplt.suptitle('Example WSIs')\nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-11-05T06:11:26.319314Z","iopub.execute_input":"2023-11-05T06:11:26.319716Z","iopub.status.idle":"2023-11-05T06:11:30.044332Z","shell.execute_reply.started":"2023-11-05T06:11:26.319687Z","shell.execute_reply":"2023-11-05T06:11:30.042888Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Visually, TMAs and WSIs are clearly distinct from each other. Below is a summary of the visual differences, based on the above examples.\n\n| **TMAs**                           | **WSIs**                                         |\n|-----------------------------------|--------------------------------------------------|\n| Single tissue specimen per image  | One or two tissue specimens per image           |\n| Circular boundary of the tissue   | Irregular boundary of the tissue specimen(s)    |\n| Darker color profile (purple)     | Lighter color profile (pink)                     |\n| No pre-processing of the background is apparent | The background outside of the specimen borders appear to have been pre-processed to be black |\n\nThe presence of multiple specimens being present in some of the WSIs is worth exploring further. It is noticeable that these samples have an image width that is roughly twice the image height. But as shown earlier in the image size scatterplots, there are samples where image width is more than double the height, with one example being more than five times as large. There are also samples whose heights are considerably greater than their widths.\n\nWe can selectively view the images with the greatest width and height disparity (in both directions) to investigate further.","metadata":{}},{"cell_type":"code","source":"# Calculate the width-height difference for each sample and sort on the resulting values among WSIs\ntrain_df['width_height_difference'] = train_df['image_width'] - train_df['image_height']\ntrain_df_sorted = train_df[~train_df['is_tma']].sort_values('width_height_difference')\n\n# Select the top 4 images with the largest width-height difference and the top 4 with the smallest\nlongest_images = train_df_sorted.head(4)['image_id']\nwidest_images = train_df_sorted.tail(4)['image_id']\n\n# Display the images\nfig, axes = plt.subplots(2, 4, figsize=(12, 6))\nfor i, id in enumerate(longest_images):\n    img = cv2.imread(f'/kaggle/input/UBC-OCEAN/train_thumbnails/{id}_thumbnail.png')\n    \n    axes[0, i].imshow(img)\n    axes[0, i].set_title(id)\n    axes[0, i].axis('off')\n    \nfor i, id in enumerate(widest_images):\n    img = cv2.imread(f'/kaggle/input/UBC-OCEAN/train_thumbnails/{id}_thumbnail.png')\n    \n    axes[1, i].imshow(img)\n    axes[1, i].set_title(id)\n    axes[1, i].axis('off')\n    \nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-11-05T06:11:30.045816Z","iopub.execute_input":"2023-11-05T06:11:30.046207Z","iopub.status.idle":"2023-11-05T06:11:42.515021Z","shell.execute_reply.started":"2023-11-05T06:11:30.04617Z","shell.execute_reply":"2023-11-05T06:11:42.513965Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- The longest images on the vertical axis feature single tissue specimens\n- The longest images on the horizontal axis feature two to four tissue specimens","metadata":{}},{"cell_type":"markdown","source":"It is typical for images to be processed so that all of them are of the same image size for image classification. One simple approach would be to pad the images, but this may not be a good approach for images with multiple tissue specimens. One way to handle multi-specimen images could be to separate the specimens into their own images as a way to increase the number of training samples while reducing the potential complexity introduced by having multiple specimens within the same image.","metadata":{}},{"cell_type":"markdown","source":"# Conclusion\n- The dataset is imbalanced with respect to the label and appropriate measures such as class weighting during training should be considered.\n- The dataset is also imbalanced with respect to image type, with only 25 out of 538 samples consisting of TMAs. Considering that the test set is comprised mostly of TMAs, we should apply augmentation to the TMAs in the training set so that our model can learn more from the type of data we expect it to receive in the test set.\n- There are instances of WSIs that contain more than one tissue specimen. We may want to consider processing these cases so that all images contain just one specimen.","metadata":{}}]}