{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"Certainly, I will provide a detailed explanation of the code :\n\n1. **First Cell: Importing Libraries and Reading Data**\n   - In this part, various necessary libraries are imported, such as Pandas, Matplotlib, Scikit-learn, and others.\n   - After that, a CSV file containing training data is read using Pandas. The data is stored in a DataFrame.\n\n2. **Second Cell: Data Analysis and Visualization**\n   - Here, the data is displayed and analyzed in the DataFrame. Data is plotted using charts to understand its distribution.\n   - The data is also visualized based on categories (HGSC and LGSC) using multiple charts.\n\n3. **Third Cell: Further Data Analysis**\n   - Data analysis continues with additional information about the data, such as image dimensions and class distribution.\n   - A pie chart is drawn to illustrate the class distribution within the data.\n\n4. **Fourth Cell: Loading Data from Image Files**\n   - In this section, the number of images available in image files is determined.\n   - A pie chart is plotted to show the distribution of images between the training and test sets.\n\n5. **Fifth Cell: Displaying a Random Image from the Training Set**\n   - A random image from the training set is displayed in this section using the `skimage` library.\n\n6. **Sixth Cell: Preparing Data for Training**\n   - Data preparation steps are performed here to make it ready for training a machine learning model.\n   - The `ImageDataGenerator` from TensorFlow is used to preprocess the data and set some properties such as image resizing.\n\n7. **Seventh Cell: Defining and Configuring the Machine Learning Model**\n   - The machine learning model is defined using TensorFlow and Keras.\n   - Layers, units, and other model properties are configured.\n   - Information about important aspects like the learning rate and used optimizations during training is specified.\n\n8. **Eighth Cell: Training the Model**\n   - This part involves training the machine learning model.\n   - It's trained using data prepared earlier, and crucial information is provided, such as the learning rate schedule and early stopping.\n\n9. **Ninth Cell: Displaying Training Progress**\n   - The training progress is displayed in this cell using line plots for model accuracy and loss over epochs.\n\n10. **Tenth Cell: Evaluating the Model**\n   - The model's performance is evaluated using test data.\n   - The code calculates and displays various metrics such as loss, accuracy, confusion matrix, and a classification report.\n\nPlease note that specific details might vary depending on the data and the exact problem the code is addressing. These explanations provide a general overview of the code sections and their functionality.","metadata":{}},{"cell_type":"code","source":"# Import libraries\nimport pandas as pd\nimport matplotlib.pyplot as plt\nfrom skimage import io\nimport os\nimport seaborn as sns\nimport cv2\nimport random\nimport glob\nimport imageio\nfrom sklearn.metrics import confusion_matrix, classification_report\nfrom sklearn.utils.class_weight import compute_class_weight\nimport tensorflow as tf\nfrom tensorflow.keras.models import Sequential, Model\nfrom tensorflow.keras.layers import Dense, Activation, Conv2D, Flatten, Dropout, MaxPooling2D, BatchNormalization\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator\nfrom tensorflow.keras.optimizers import Adam\nfrom tensorflow.keras.applications import EfficientNetB0\nimport numpy as np\n\n# Load the training file\ntrain_df = pd.read_csv('/kaggle/input/UBC-OCEAN/train.csv')\nprint(train_df.head())\nprint(train_df['label'].value_counts())\nprint(train_df.shape)\nprint(train_df.isna().sum().sum())\n\n# Create a list of column names we want to distribute\ncolumns = [col for col in train_df.columns if col != 'label']\n\n# Loop to check the data distribution\nfor col in columns:\n    fig, axs = plt.subplots(figsize=(15, 5), ncols=3)\n    \n    # Sample data distribution\n    sns.histplot(data=train_df, x=col, kde=True, ax=axs[0])\n    axs[0].set_title('Sample Distribution')\n    \n    # Data distribution for label 1\n    sns.histplot(data=train_df[train_df['label'] == \"HGSC\"], x=col, kde=True, ax=axs[1], color='orange')\n    axs[1].set_title('label - HGSC')\n    \n    # Data distribution for label 0\n    sns.histplot(data=train_df[train_df['label'] == \"LGSC\"], x=col, kde=True, ax=axs[2], color='green')\n    axs[2].set_title('label - LGSC')\n    \n    plt.tight_layout()\n    plt.show()\n\n# Information about the data\ntrain_df.info()\nprint(train_df['image_id'].is_unique)\nprint(train_df['label'].is_unique)\nprint(train_df[['image_height', 'image_width']].describe())\n\n# Number of images in the training set\nprint(train_df['image_id'].shape[0])\nprint(len(os.listdir('/kaggle/input/UBC-OCEAN/train_images')))\nprint(len(os.listdir('/kaggle/input/UBC-OCEAN/train_thumbnails')))\n\n# Class distribution in the training set\ntrain_df['label'].value_counts().plot(kind='pie', autopct='%1.1f%%')\nplt.title('Distribution of Training Set')\nplt.show()\n\n# Load the data\ntrain_data = glob.glob('/kaggle/input/UBC-OCEAN/train_images/*.png')\ntest_data = glob.glob('/kaggle/input/UBC-OCEAN/test_images/*.png')\nprint(f\"The Training Set contains: {len(train_data)} images\")\nprint(f\"The Testing Set contains: {len(test_data)} images\")\n\n# Distribution of images between training and testing sets\ntotal_train = len(train_data)\ntotal_test = len(test_data)\nplt.pie([total_train, total_test], labels=['Training Set', 'Testing Set'], autopct='%1.1f%%', colors=['lightgreen', 'red'])\nplt.title('Distribution of Images in Training and Testing Sets')\nplt.show()\n\n# Display a random image from the training set\nimage_path = '/kaggle/input/UBC-OCEAN/train_thumbnails/10642_thumbnail.png'\nimage = io.imread(image_path)\nplt.imshow(image)\nplt.show()\n\n# Read data sets\ntraindf = pd.read_csv('/kaggle/input/UBC-OCEAN/train.csv', dtype=str)\ntestdf = pd.read_csv('/kaggle/input/UBC-OCEAN/test.csv', dtype=str)\n\n# Prepare the data\ndatagen = ImageDataGenerator(rescale=1./255., validation_split=0.25)\ntrain_generator = datagen.flow_from_dataframe(\n    dataframe=traindf,\n    directory=\"/kaggle/input/UBC-OCEAN/train_images/\",\n    x_col=\"image_id\",\n    y_col=\"label\",\n    subset=\"training\",\n    batch_size=32,\n    seed=42,\n    shuffle=True,\n    class_mode=\"categorical\",\n    target_size=(100, 100)\n)\n\n# Functions to convert file names\ndef append_ext(fn):\n    return fn + \".png\"\n\ndef append_ext_thum(fn):\n    return fn + \"_thumbnail.png\"\n\n# Disable the warning about converting from bool to numpy.uint8\nimport warnings\nwarnings.filterwarnings(\"ignore\", category=np.VisibleDeprecationWarning)\n","metadata":{"execution":{"iopub.status.busy":"2023-10-10T23:30:35.011204Z","iopub.execute_input":"2023-10-10T23:30:35.011607Z","iopub.status.idle":"2023-10-10T23:30:39.470178Z","shell.execute_reply.started":"2023-10-10T23:30:35.011577Z","shell.execute_reply":"2023-10-10T23:30:39.469227Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(testdf.columns)\n","metadata":{"execution":{"iopub.status.busy":"2023-10-10T23:30:39.471917Z","iopub.execute_input":"2023-10-10T23:30:39.472191Z","iopub.status.idle":"2023-10-10T23:30:39.477179Z","shell.execute_reply.started":"2023-10-10T23:30:39.47217Z","shell.execute_reply":"2023-10-10T23:30:39.476136Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Transform file names for the test set\ntestdf[\"image_id_path\"] = testdf[\"image_id\"].apply(append_ext)\ntestdf[\"image_id_path_thum\"] = testdf[\"image_id\"].apply(append_ext_thum)\n\n# Create paths for images in the test set\ntestdf['Image_path'] = [os.path.join('/kaggle/input/UBC-OCEAN/test_images', image) for image in testdf['image_id_path']]\ntestdf['Image_path_thumbnails'] = [os.path.join('/kaggle/input/UBC-OCEAN/test_thumbnails', image) for image in testdf['image_id_path_thum']]\n\n# Create paths for images in the training set\ntraindf[\"image_id_path\"] = traindf[\"image_id\"].apply(append_ext)\ntraindf[\"image_id_path_thum\"] = traindf[\"image_id\"].apply(append_ext_thum)\ntraindf['Image_path'] = [os.path.join('/kaggle/input/UBC-OCEAN/train_images', image) for image in traindf['image_id_path']]\ntraindf['Image_path_thumbnails'] = [os.path.join('/kaggle/input/UBC-OCEAN/train_thumbnails', image) for image in traindf['image_id_path_thum']]\n\n# Display random images from the training set\nfull_path_random = np.random.choice(traindf['Image_path_thumbnails'], 5)\n\ndef image_viewer(dataset, index, ax):\n    image_path = dataset['Image_path_thumbnails'][index]\n    image = io.imread(image_path)\n    ax.imshow(image)\n\ndef plot_some_images(dataset, title):\n    fig, axs = plt.subplots(nrows=1, ncols=2, figsize=(20, 8))\n    for ind, ax in enumerate(axs.flat):\n        index = random.randrange(len(dataset))\n        image_viewer(dataset, index, ax)\n        ax.set_title(dataset['label'][index], fontsize=8)\n        ax.axis('off')\n    fig.suptitle(title, fontsize=15)\n    plt.show()\n\nplot_some_images(traindf, 'Training Images')\n\n# Calculate class weights\nclass_weights = compute_class_weight(class_weight=\"balanced\",\n                                     classes=np.unique(traindf['label']),\n                                     y=traindf['label'])\nclasses = (np.unique(traindf['label']))\nclass_weights_forplot = dict(zip(classes, class_weights))\nclasses\nclass_weights_forplot\nclass_weights = dict(zip(range(43), class_weights))\n\n# Prepare data generators for the training and testing sets\ntrain_generator = tf.keras.preprocessing.image.ImageDataGenerator(\n    preprocessing_function=tf.keras.applications.efficientnet.preprocess_input\n)\ntest_generator = tf.keras.preprocessing.image.ImageDataGenerator(\n    preprocessing_function=tf.keras.applications.efficientnet.preprocess_input\n)\n\ntrain_images = train_generator.flow_from_dataframe(\n    dataframe=traindf,\n    x_col='Image_path_thumbnails',\n    y_col='label',\n    target_size=(256, 256),\n    color_mode='grayscale',\n    class_mode=\"categorical\",\n    batch_size=64,\n    shuffle=True,\n    seed=210\n)\n\ntest_images = test_generator.flow_from_dataframe(\n    dataframe=traindf[0:10],\n    x_col='Image_path_thumbnails',\n    y_col='label',\n    target_size=(256, 256),\n    class_mode=\"categorical\",\n    color_mode='grayscale',\n    batch_size=64,\n    shuffle=False\n)\n\n# Create the neural network model\ntrans_arc = EfficientNetB0(weights=\"imagenet\", include_top=False,\n                            input_shape=(256, 256, 3), pooling='max')\n\nfor layer in trans_arc.layers:\n    layer.trainable = False\n\ninputs = trans_arc.input\nflatten = trans_arc.output\n\nx = Dense(256, activation='relu')(flatten)\nx = BatchNormalization()(x)\nx = Dropout(0.3)(x)\n\nx = Dense(128, activation='relu')(x)\nx = BatchNormalization()(x)\n\noutputs = Dense(1, activation='sigmoid')(x)\n\nmodel = Model(inputs=inputs, outputs=outputs)\nmodel.summary()\n\n# Create a folder to save the model if it doesn't exist\nif not os.path.exists(\"/kaggle/working/checkpoints\"):\n    os.mkdir(\"/kaggle/working/checkpoints\")\n\n# Specify and configure the loss, learning rate, and metrics\nloss = [tf.keras.losses.binary_crossentropy]\n\ninitial_learning_rate = 0.005\n\nlr_schedule = tf.keras.optimizers.schedules.ExponentialDecay(\n    initial_learning_rate,\n    decay_steps=82,\n    decay_rate=0.9,\n    staircase=True\n)\n\noptimizer = Adam(\n    learning_rate=lr_schedule,\n    beta_1=0.9,\n    beta_2=0.999,\n    epsilon=1e-07\n)\n\nmetrics = ['accuracy']\n\nmodel.compile(\n    optimizer=optimizer,\n    loss=loss,\n    metrics=metrics\n)\n","metadata":{"execution":{"iopub.status.busy":"2023-10-10T23:30:39.478449Z","iopub.execute_input":"2023-10-10T23:30:39.47906Z","iopub.status.idle":"2023-10-10T23:30:44.478374Z","shell.execute_reply.started":"2023-10-10T23:30:39.479036Z","shell.execute_reply":"2023-10-10T23:30:44.477303Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Print information about available devices for running TensorFlow on GPUs\nfrom tensorflow.python.client import device_lib\nprint(device_lib.list_local_devices())\n\n# Train the model using GPU if available\nwith tf.device(\"/device:GPU:0\"):\n    history = model.fit(\n        train_images,\n        validation_data=test_images,\n        verbose=True,\n        epochs=5,\n        class_weight=class_weights,\n        callbacks=[\n            tf.keras.callbacks.LearningRateScheduler(lr_schedule),\n            tf.keras.callbacks.EarlyStopping(\n                monitor='val_loss',\n                patience=7,\n                restore_best_weights=True\n            )\n        ]\n    )\n\n# Display training progress\nprint(history.history.keys())\nplt.plot(history.history['accuracy'])\nplt.plot(history.history['val_accuracy'])\nplt.title('Model Accuracy')\nplt.ylabel('Accuracy')\nplt.xlabel('Epoch')\nplt.legend(['Train', 'Validation'], loc='upper left')\nplt.show()\n\nplt.plot(history.history['val_loss'])\nplt.title('Model Loss')\nplt.ylabel('Loss')\nplt.xlabel('Epoch')\nplt.legend(['Train', 'Validation'], loc='upper left')\nplt.show()\n\n# Evaluate the model\ndef plot_model_evaluation(model, test_data, n_classes, target_labels):\n    results = model.evaluate(test_data, verbose=0)\n    loss = results[0]\n    acc = results[1]\n\n    print(\"Test Loss: {:.5f}\".format(loss))\n    print(\"Test Accuracy: {:.2f}%\".format(acc * 100))\n\n    y_pred = np.squeeze((model.predict(test_data) >= 0.5).astype(int))\n    cm = confusion_matrix(test_data.labels, y_pred)\n    clr = classification_report(test_data.labels, y_pred, target_names=target_labels)\n\n    plt.figure(figsize=(15, 15))\n    sns.heatmap(cm, annot=True, fmt='g', vmin=0, cmap='Blues', cbar=False)\n    plt.xticks(ticks=np.arange(n_classes) + 0.5, labels=list(test_data.class_indices.keys()), rotation=90)\n    plt.yticks(ticks=np.arange(n_classes) + 0.5, labels=list(test_data.class_indices.keys()), rotation=0)\n    plt.xlabel(\"Predicted\")\n    plt.ylabel(\"Actual\")\n    plt.title(\"Confusion Matrix\")\n    plt.show()\n\n    print(\"Classification Report:\\n\", clr)\n\n# Evaluate the model using test data\nplot_model_evaluation(model, test_images, 3, traindf[0:10]['label'].unique())\n","metadata":{"execution":{"iopub.status.busy":"2023-10-10T23:30:44.479881Z","iopub.execute_input":"2023-10-10T23:30:44.480186Z","iopub.status.idle":"2023-10-10T23:41:16.912909Z","shell.execute_reply.started":"2023-10-10T23:30:44.480163Z","shell.execute_reply":"2023-10-10T23:41:16.911938Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load the sample submission file (modify the file path accordingly)\nsubmission = pd.read_csv('/kaggle/input/UBC-OCEAN/sample_submission.csv')\n# Save the updated DataFrame as a CSV file\nsubmission.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2023-10-10T23:41:16.915379Z","iopub.execute_input":"2023-10-10T23:41:16.916345Z","iopub.status.idle":"2023-10-10T23:41:16.934315Z","shell.execute_reply.started":"2023-10-10T23:41:16.916306Z","shell.execute_reply":"2023-10-10T23:41:16.933344Z"},"trusted":true},"execution_count":null,"outputs":[]}]}