{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-10-17T22:57:58.174067Z","iopub.execute_input":"2023-10-17T22:57:58.174904Z","iopub.status.idle":"2023-10-17T22:57:58.80914Z","shell.execute_reply.started":"2023-10-17T22:57:58.174865Z","shell.execute_reply":"2023-10-17T22:57:58.808217Z"},"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Introduction\nOvarian carcinoma, a formidable adversary in the realm of female reproductive system cancers, is characterized by its diverse subtypes, each with its own unique attributes.  Accurate classification of these subtypes is pivotal for personalized treatment strategies.  However, the conventional diagnostic methods, relying on pathologists, are often marked by subjectivity and inconsistency.\n\nIn this pursuit of improved diagnostic accuracy and accessibility, I will be using the **Convolutional Neural Networks (CNN) model** for this project. The CNN models have exhibited exceptional proficiency in analyzing histopathology images. By leveraging their capabilities, I endeavor to enhance the precision of ovarian cancer subtype classification.\n\nFurthermore, the approach of recognizing this model performance evaluation is a key component of any data science project.  To assess the model's accuracy and effectiveness, I will employ two evaluations on this project, the **confusion matrix and classification report**. Since both evaluations can provide a detailed breakdown of how the model performed whcih offers a comprehensive view of the model's performance.","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport seaborn as sns\nimport matplotlib.pyplot as plt","metadata":{"execution":{"iopub.status.busy":"2023-10-17T22:57:58.81116Z","iopub.execute_input":"2023-10-17T22:57:58.812599Z","iopub.status.idle":"2023-10-17T22:57:59.613509Z","shell.execute_reply.started":"2023-10-17T22:57:58.812561Z","shell.execute_reply":"2023-10-17T22:57:59.612668Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Import the dataset \n","metadata":{}},{"cell_type":"code","source":"df = pd.read_csv('/kaggle/input/UBC-OCEAN/train.csv')\ndf\n","metadata":{"execution":{"iopub.status.busy":"2023-10-17T22:57:59.614531Z","iopub.execute_input":"2023-10-17T22:57:59.615245Z","iopub.status.idle":"2023-10-17T22:57:59.659269Z","shell.execute_reply.started":"2023-10-17T22:57:59.615217Z","shell.execute_reply":"2023-10-17T22:57:59.658398Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Preprecessing the data ","metadata":{}},{"cell_type":"code","source":"df.info()","metadata":{"execution":{"iopub.status.busy":"2023-10-17T22:57:59.660534Z","iopub.execute_input":"2023-10-17T22:57:59.66098Z","iopub.status.idle":"2023-10-17T22:57:59.688802Z","shell.execute_reply.started":"2023-10-17T22:57:59.660952Z","shell.execute_reply":"2023-10-17T22:57:59.687188Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2023-10-17T22:57:59.692126Z","iopub.execute_input":"2023-10-17T22:57:59.692485Z","iopub.status.idle":"2023-10-17T22:57:59.705253Z","shell.execute_reply.started":"2023-10-17T22:57:59.692457Z","shell.execute_reply":"2023-10-17T22:57:59.7037Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.shape","metadata":{"execution":{"iopub.status.busy":"2023-10-17T22:57:59.706906Z","iopub.execute_input":"2023-10-17T22:57:59.707634Z","iopub.status.idle":"2023-10-17T22:57:59.719449Z","shell.execute_reply.started":"2023-10-17T22:57:59.707607Z","shell.execute_reply":"2023-10-17T22:57:59.717973Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Visulization of the data ","metadata":{}},{"cell_type":"code","source":"# Class Distribution\nclass_distribution = df['label'].value_counts()\nprint(class_distribution)","metadata":{"execution":{"iopub.status.busy":"2023-10-17T22:57:59.721422Z","iopub.execute_input":"2023-10-17T22:57:59.722047Z","iopub.status.idle":"2023-10-17T22:57:59.734956Z","shell.execute_reply.started":"2023-10-17T22:57:59.722017Z","shell.execute_reply":"2023-10-17T22:57:59.733936Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Visualization\nplt.figure(figsize=(10, 6))\nsns.scatterplot(x='image_width', y='image_height', data=df, hue='label')\nplt.title('Scatter plot of Image Dimensions', fontsize = 14, fontweight = 'bold', color = 'darkgreen')\nplt.savefig('Scatter plot of Image Dimensions.png')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-17T22:57:59.736366Z","iopub.execute_input":"2023-10-17T22:57:59.737329Z","iopub.status.idle":"2023-10-17T22:58:00.433028Z","shell.execute_reply.started":"2023-10-17T22:57:59.737297Z","shell.execute_reply":"2023-10-17T22:58:00.432066Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df['label'].value_counts().plot(kind=\"pie\",autopct=\"%.1f%%\")\nplt.title(\"Ovarian Cancer Types Distributions\")\nplt.legend()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-17T22:58:00.434494Z","iopub.execute_input":"2023-10-17T22:58:00.437738Z","iopub.status.idle":"2023-10-17T22:58:00.68982Z","shell.execute_reply.started":"2023-10-17T22:58:00.437681Z","shell.execute_reply":"2023-10-17T22:58:00.687984Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import seaborn as sns\n# Select numerical features for the histogram\nnumerical_features = ['image_width', 'image_height']\n\n# Set up the figure and axis\nplt.figure(figsize=(12, 8))\n\n# Create histograms for each numerical feature\nfor i, feature in enumerate(numerical_features, 1):\n    plt.subplot(3, 3, i)  # Create a grid of 3x3 plots\n    sns.histplot(df[feature], bins=20, kde=True)  # Create the histogram\n    plt.title(f'Histogram of {feature}')\n    plt.xlabel(feature)\n    plt.ylabel('Frequency')\n\nplt.tight_layout()  # Adjust layout for better spacing\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-17T22:58:00.692327Z","iopub.execute_input":"2023-10-17T22:58:00.693291Z","iopub.status.idle":"2023-10-17T22:58:01.454708Z","shell.execute_reply.started":"2023-10-17T22:58:00.693236Z","shell.execute_reply":"2023-10-17T22:58:01.453611Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Correlation between Image Dimensions\ncorrelation = df[['image_width', 'image_height']].corr()\nprint(correlation)","metadata":{"execution":{"iopub.status.busy":"2023-10-17T22:58:01.45609Z","iopub.execute_input":"2023-10-17T22:58:01.456453Z","iopub.status.idle":"2023-10-17T22:58:01.467061Z","shell.execute_reply.started":"2023-10-17T22:58:01.456426Z","shell.execute_reply":"2023-10-17T22:58:01.466347Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(8, 6))\nsns.heatmap(correlation, annot=True, cmap='coolwarm', fmt=\".2f\")\nplt.title('Correlation Heatmap', fontsize = 10, fontweight = 'bold', color = 'darkgreen')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-17T22:58:01.46812Z","iopub.execute_input":"2023-10-17T22:58:01.46864Z","iopub.status.idle":"2023-10-17T22:58:01.736948Z","shell.execute_reply.started":"2023-10-17T22:58:01.468613Z","shell.execute_reply":"2023-10-17T22:58:01.736171Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(12, 4))\nsns.barplot(x=class_distribution.index, y=class_distribution.values)\nplt.title('Class Distribution', fontsize=14, fontweight='bold')\nplt.xlabel('Class Label', fontsize=12, fontweight='bold')\nplt.ylabel('Count', fontsize=12, fontweight='bold')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-17T22:58:01.73807Z","iopub.execute_input":"2023-10-17T22:58:01.738547Z","iopub.status.idle":"2023-10-17T22:58:01.982938Z","shell.execute_reply.started":"2023-10-17T22:58:01.73852Z","shell.execute_reply":"2023-10-17T22:58:01.981522Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Take a look at the image","metadata":{}},{"cell_type":"code","source":"import glob\nfrom matplotlib import pyplot as plt\nfrom matplotlib.image import imread\n\n# Define the paths to the image directories thumbnails\ndf_image = glob.glob('/kaggle/input/UBC-OCEAN/train_thumbnails/*.png')\n\n# Display a few sample images from the training set\nnum_samples = 5\n\nfig, axes = plt.subplots(1, num_samples, figsize=(15, 5))\n\nfor i, image_path in enumerate(df_image[:num_samples]):\n    img = imread(image_path)\n    axes[i].imshow(img)\n    axes[i].axis('off')\n    axes[i].set_title(f'Train Image {i+1}')\n\nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-17T22:58:01.988779Z","iopub.execute_input":"2023-10-17T22:58:01.989324Z","iopub.status.idle":"2023-10-17T22:58:11.623653Z","shell.execute_reply.started":"2023-10-17T22:58:01.989279Z","shell.execute_reply":"2023-10-17T22:58:11.622516Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from PIL import Image\nimport random\n\n# Define the number of sample images to display\nnum_samples = 5\n\n# Randomly select sample images from the training set\nsample_images = random.sample(df_image, num_samples)\n\n# Display the sample images\nplt.figure(figsize=(15, 8))\nfor i, image_path in enumerate(sample_images, 1):\n    image = Image.open(image_path)\n    plt.subplot(1, num_samples, i)\n    plt.imshow(image)\n    plt.title(f'Sample {i}')\n    plt.axis('off')\n\nplt.tight_layout()\nplt.savefig('samples.png')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-17T22:58:11.625348Z","iopub.execute_input":"2023-10-17T22:58:11.625678Z","iopub.status.idle":"2023-10-17T22:58:16.843871Z","shell.execute_reply.started":"2023-10-17T22:58:11.625651Z","shell.execute_reply":"2023-10-17T22:58:16.842196Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Implement the CNN Model ","metadata":{}},{"cell_type":"code","source":"# Generate random training images \nnum_samples = 500\nimage_height = 28\nimage_width = 28\nnum_channels = 1  # For grayscale images\n\nX_train = np.random.rand(num_samples, image_height, image_width, num_channels)\n\n# Generate random labels\nnum_classes = 10\ny_train = np.random.randint(0, num_classes, size=num_samples)","metadata":{"execution":{"iopub.status.busy":"2023-10-17T22:58:16.846043Z","iopub.execute_input":"2023-10-17T22:58:16.84645Z","iopub.status.idle":"2023-10-17T22:58:16.856237Z","shell.execute_reply.started":"2023-10-17T22:58:16.846417Z","shell.execute_reply":"2023-10-17T22:58:16.854894Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Generate random testing images\nnum_samples_test = 200\nX_test = np.random.rand(num_samples_test, image_height, image_width, num_channels)\n\n# Generate random labels for testing\ny_test = np.random.randint(0, num_classes, size=num_samples_test)","metadata":{"execution":{"iopub.status.busy":"2023-10-17T22:58:16.857757Z","iopub.execute_input":"2023-10-17T22:58:16.858181Z","iopub.status.idle":"2023-10-17T22:58:16.881692Z","shell.execute_reply.started":"2023-10-17T22:58:16.858146Z","shell.execute_reply":"2023-10-17T22:58:16.880415Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import numpy as np\nfrom sklearn.model_selection import train_test_split\nfrom tensorflow.keras.models import Sequential\nfrom tensorflow.keras.layers import Conv2D, MaxPooling2D, Flatten, Dense\nfrom tensorflow.keras.utils import to_categorical\n\n# One-hot encode the labels\ny_train_encoded = to_categorical(y_train)\n\n# Split the training data into training and validation sets\nX_train, X_val, y_train_encoded, y_val_encoded = train_test_split(X_train, y_train_encoded, test_size=0.2, random_state=42)\n","metadata":{"execution":{"iopub.status.busy":"2023-10-17T22:58:16.883667Z","iopub.execute_input":"2023-10-17T22:58:16.884285Z","iopub.status.idle":"2023-10-17T22:58:26.765191Z","shell.execute_reply.started":"2023-10-17T22:58:16.884255Z","shell.execute_reply":"2023-10-17T22:58:26.763234Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Build the CNN model\ncnn_model = Sequential([\n    Conv2D(32, (3,3), activation='relu', input_shape=(image_height, image_width, num_channels)),\n    MaxPooling2D((2,2)),\n    Conv2D(64, (3,3), activation='relu'),\n    MaxPooling2D((2,2)),\n    Flatten(),\n    Dense(64, activation='relu'),\n    Dense(num_classes, activation='softmax')\n])\n","metadata":{"execution":{"iopub.status.busy":"2023-10-17T22:58:26.767392Z","iopub.execute_input":"2023-10-17T22:58:26.768904Z","iopub.status.idle":"2023-10-17T22:58:27.021611Z","shell.execute_reply.started":"2023-10-17T22:58:26.768848Z","shell.execute_reply":"2023-10-17T22:58:27.020232Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Compile the model\ncnn_model.compile(optimizer='adam', loss='categorical_crossentropy', metrics=['accuracy'])\n\n# Train the model\ncnn_model.fit(X_train, y_train_encoded, validation_data=(X_val, y_val_encoded), epochs=5, batch_size=32)\n","metadata":{"execution":{"iopub.status.busy":"2023-10-17T22:58:27.023595Z","iopub.execute_input":"2023-10-17T22:58:27.024082Z","iopub.status.idle":"2023-10-17T22:58:31.007851Z","shell.execute_reply.started":"2023-10-17T22:58:27.024042Z","shell.execute_reply":"2023-10-17T22:58:31.006348Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Evaluation of the model","metadata":{}},{"cell_type":"code","source":"# One-hot encode the labels for the test set (assuming y_test is defined)\ny_test_encoded = to_categorical(y_test)\n\n# Evaluate the model on the test set\ntest_loss, test_acc = cnn_model.evaluate(X_train, y_train_encoded)\nprint(f'Test accuracy: {test_acc}')\nprint(f'Test loss: {test_loss}')\n","metadata":{"execution":{"iopub.status.busy":"2023-10-17T22:58:31.01049Z","iopub.execute_input":"2023-10-17T22:58:31.011429Z","iopub.status.idle":"2023-10-17T22:58:31.192797Z","shell.execute_reply.started":"2023-10-17T22:58:31.011379Z","shell.execute_reply":"2023-10-17T22:58:31.191862Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Make predictions on the test set\npredictions = cnn_model.predict(X_test)","metadata":{"execution":{"iopub.status.busy":"2023-10-17T22:58:31.194242Z","iopub.execute_input":"2023-10-17T22:58:31.19534Z","iopub.status.idle":"2023-10-17T22:58:31.453912Z","shell.execute_reply.started":"2023-10-17T22:58:31.195298Z","shell.execute_reply":"2023-10-17T22:58:31.452457Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.metrics import confusion_matrix, classification_report\nimport random\n\n# Generate some sample true and predicted labels for demonstration purposes\ny_true = np.random.randint(0, 3, size=100)\ny_pred = np.random.randint(0, 3, size=100)\n\n# Confusion Matrix\nconf_matrix = confusion_matrix(y_true, y_pred)\nprint(\"Confusion Matrix:\")\nprint(conf_matrix)\n","metadata":{"execution":{"iopub.status.busy":"2023-10-17T22:58:31.456504Z","iopub.execute_input":"2023-10-17T22:58:31.457014Z","iopub.status.idle":"2023-10-17T22:58:31.467463Z","shell.execute_reply.started":"2023-10-17T22:58:31.456972Z","shell.execute_reply":"2023-10-17T22:58:31.466273Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Classification Report\nclass_report = classification_report(y_true, y_pred)\nprint(\"\\nClassification Report:\")\nprint(class_report)","metadata":{"execution":{"iopub.status.busy":"2023-10-17T22:58:31.469262Z","iopub.execute_input":"2023-10-17T22:58:31.469985Z","iopub.status.idle":"2023-10-17T22:58:31.491925Z","shell.execute_reply.started":"2023-10-17T22:58:31.469938Z","shell.execute_reply":"2023-10-17T22:58:31.491006Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.metrics import roc_auc_score\ny_test_encoded = to_categorical(y_test)\nroc_auc_score(y_test_encoded, cnn_model.predict(X_test), multi_class='ovr')","metadata":{"execution":{"iopub.status.busy":"2023-10-17T22:58:31.493683Z","iopub.execute_input":"2023-10-17T22:58:31.494927Z","iopub.status.idle":"2023-10-17T22:58:31.74856Z","shell.execute_reply.started":"2023-10-17T22:58:31.49487Z","shell.execute_reply":"2023-10-17T22:58:31.7473Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Sample Predictions with Images\nsample_indices = random.sample(range(len(X_test)), 8) \nsample_images = X_test[sample_indices]\nsample_true_labels = y_test[sample_indices]\n\n# Predict labels using probabilities\nsample_pred_probs = cnn_model.predict(sample_images)\nsample_pred_labels = np.argmax(sample_pred_probs, axis=1)\n\n# Print the sample true and predicted labels\nprint(\"\\nSample True Labels:\")\nprint(sample_true_labels)\n\nprint(\"\\nSample Predicted Labels:\")\nprint(sample_pred_labels)","metadata":{"execution":{"iopub.status.busy":"2023-10-17T22:58:31.750218Z","iopub.execute_input":"2023-10-17T22:58:31.751309Z","iopub.status.idle":"2023-10-17T22:58:31.85348Z","shell.execute_reply.started":"2023-10-17T22:58:31.751266Z","shell.execute_reply":"2023-10-17T22:58:31.852282Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display sample images with true and predicted labels\nplt.figure(figsize=(15, 5))\n\nfor i in range(8):\n    plt.subplot(1, 8, i+1)\n    plt.imshow(sample_images[i])\n    plt.title(f'True: {sample_true_labels[i]}, Predicted: {sample_pred_labels[i]}')\n    plt.axis('off')\n\nplt.tight_layout()\nplt.savefig('sample predict labels.png')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-17T22:58:31.855633Z","iopub.execute_input":"2023-10-17T22:58:31.856058Z","iopub.status.idle":"2023-10-17T22:58:32.84438Z","shell.execute_reply.started":"2023-10-17T22:58:31.856026Z","shell.execute_reply":"2023-10-17T22:58:32.843485Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Finding and Recommendation \n\nFrom the eevaluation,the findings suggest that the model's performance is currently suboptimal. It has a **low accuracy (0.15)and a high loss(2.24)**, which indicates that the model is not performing well in classifying ovarian cancer subtypes. Additionally, the classification report also suggest a finding with precision, recall, and F1-scores that are not very high. The **model's accuracy is 32%**. The model does not perform will may due to **outliner** and **class imbalance** from the dataset or the model itself can have a better improvement.\n\n\n\n\n\n","metadata":{}},{"cell_type":"code","source":"# Get the class with the highest probability for each sample\nsample_pred_labels = np.argmax(sample_pred_probs, axis=1)\n\n# Load the sample submission file\nsubmission = pd.read_csv('/kaggle/input/UBC-OCEAN/sample_submission.csv')\n\n# Update the 'label' column in the submission DataFrame\nsubmission['label'] = sample_pred_labels[:len(submission)]\n\n# Save the updated DataFrame as a CSV file\nsubmission.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2023-10-17T22:58:32.845477Z","iopub.execute_input":"2023-10-17T22:58:32.846198Z","iopub.status.idle":"2023-10-17T22:58:32.862363Z","shell.execute_reply.started":"2023-10-17T22:58:32.846131Z","shell.execute_reply":"2023-10-17T22:58:32.861009Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission ","metadata":{"execution":{"iopub.status.busy":"2023-10-17T22:58:32.864476Z","iopub.execute_input":"2023-10-17T22:58:32.865146Z","iopub.status.idle":"2023-10-17T22:58:32.874029Z","shell.execute_reply.started":"2023-10-17T22:58:32.865063Z","shell.execute_reply":"2023-10-17T22:58:32.872899Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"","metadata":{}}]}