{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.12.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":18647,"databundleVersionId":1126921,"sourceType":"competition"},{"sourceId":14427011,"sourceType":"datasetVersion","datasetId":9214943}],"dockerImageVersionId":31234,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## CMM Summer Course\n\nHello! Welcome to the lecture. I'm going to introduce you to the world of Deep Learning applied to Histopathology","metadata":{}},{"cell_type":"markdown","source":"In this notebook, we will use a famous dataset called PANDA. This dataset contains **~10k WSIs** of prostate cancer, classified at different grades. We will use it to estimate the severity of each case based on ISUP grade.\n\nFirst of all, we will learn to preprocess a slide, reduce noise, segment and extract features from image.\n\nAfter that, we will visualize the different features using some dimensionality reduction techniques and unsupervised learning. \n\nFinally, we will train and optimize a supervised model to predict ISUP grade based only in the WSI.","metadata":{}},{"cell_type":"markdown","source":"Let's install the requirements of the lecture.\n\nWe will use [TRIDENT](https://github.com/mahmoodlab/TRIDENT) package. This is an open source library to preprocess WSIs.\n\nIn the repository you will find the different pretrained models to use. In this case, we will use a very successful one called Virchow2. If you want to use it, please contact the owners and ask for a permision in [Huggingface](https://huggingface.co/paige-ai/Virchow2). ","metadata":{}},{"cell_type":"code","source":"%%bash\ngit clone https://github.com/mahmoodlab/TRIDENT.git\npip install --upgrade pip setuptools wheel\ncd TRIDENT && pip install -e . --no-deps\npip install timm==0.9.16 segmentation-models-pytorch einops_exts --no-deps\npip install -U imagecodecs","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-08T14:48:13.947023Z","iopub.execute_input":"2026-01-08T14:48:13.947265Z","iopub.status.idle":"2026-01-08T14:48:43.158797Z","shell.execute_reply.started":"2026-01-08T14:48:13.947237Z","shell.execute_reply":"2026-01-08T14:48:43.158251Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport os\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nfrom kaggle_secrets import UserSecretsClient\nfrom huggingface_hub import login\nimport h5py\nimport openslide\nfrom sklearn.cluster import KMeans\nfrom sklearn.mixture import GaussianMixture\nfrom sklearn.decomposition import PCA\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.neural_network import MLPClassifier\nfrom sklearn.model_selection import cross_val_score, train_test_split, GridSearchCV\nfrom sklearn.pipeline import Pipeline\nfrom sklearn.metrics import classification_report, confusion_matrix\nimport warnings\nfrom sklearn.preprocessing import label_binarize\nfrom sklearn.metrics import roc_curve, auc\nwarnings.filterwarnings('ignore')","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2026-01-08T15:57:44.126965Z","iopub.execute_input":"2026-01-08T15:57:44.12724Z","iopub.status.idle":"2026-01-08T15:57:44.132667Z","shell.execute_reply.started":"2026-01-08T15:57:44.127218Z","shell.execute_reply":"2026-01-08T15:57:44.132028Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"You need to use your huggingface token to download the pretrained models.","metadata":{}},{"cell_type":"code","source":"token = UserSecretsClient().get_secret(\"CMM_Summer_Course\")\nlogin(token=token)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-08T14:48:45.238052Z","iopub.execute_input":"2026-01-08T14:48:45.238481Z","iopub.status.idle":"2026-01-08T14:48:45.879777Z","shell.execute_reply.started":"2026-01-08T14:48:45.238458Z","shell.execute_reply":"2026-01-08T14:48:45.879207Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"This is the dataset. We will use image_id and isup_grade columns. Gleason is a derivate of gleason_score, so we will not worry about it. ","metadata":{}},{"cell_type":"code","source":"train = pd.read_csv(\"/kaggle/input/prostate-cancer-grade-assessment/train.csv\")\ntrain","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-08T14:48:45.880556Z","iopub.execute_input":"2026-01-08T14:48:45.880814Z","iopub.status.idle":"2026-01-08T14:48:45.937675Z","shell.execute_reply.started":"2026-01-08T14:48:45.880784Z","shell.execute_reply":"2026-01-08T14:48:45.937106Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plt.subplot(2,2,1)\nsns.countplot(train, x=\"data_provider\")\nplt.subplot(2,2,2)\nsns.countplot(train, x=\"isup_grade\")\nplt.suptitle(\"Some counts\")\nplt.tight_layout()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-08T14:48:45.938408Z","iopub.execute_input":"2026-01-08T14:48:45.938643Z","iopub.status.idle":"2026-01-08T14:48:46.228556Z","shell.execute_reply.started":"2026-01-08T14:48:45.938622Z","shell.execute_reply":"2026-01-08T14:48:46.227964Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Preprocessing and Feature extraction","metadata":{}},{"cell_type":"markdown","source":"We will preprocess a single slide.","metadata":{}},{"cell_type":"code","source":"sample = train.sample(1, random_state=42)\ndisplay(sample)\nsample[\"image_id\"] = sample[\"image_id\"] + \".tiff\"\nsample = sample.rename(columns={\"image_id\": \"wsi\"})[[\"wsi\"]]\nsample.to_csv(\"sample.csv\", index=False)\ntiff_file = sample.wsi.values[0]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-08T14:48:46.229502Z","iopub.execute_input":"2026-01-08T14:48:46.229873Z","iopub.status.idle":"2026-01-08T14:48:46.244138Z","shell.execute_reply.started":"2026-01-08T14:48:46.229844Z","shell.execute_reply":"2026-01-08T14:48:46.24355Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"slide = openslide.OpenSlide(f\"/kaggle/input/prostate-cancer-grade-assessment/train_images/{tiff_file}\")\nthumbnail = slide.get_thumbnail((1024, 1024))\nthumbnail","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-08T14:48:46.246503Z","iopub.execute_input":"2026-01-08T14:48:46.246791Z","iopub.status.idle":"2026-01-08T14:48:46.572279Z","shell.execute_reply.started":"2026-01-08T14:48:46.246764Z","shell.execute_reply":"2026-01-08T14:48:46.571524Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"%%bash\npython TRIDENT/run_batch_of_slides.py \\\n--wsi_dir /kaggle/input/prostate-cancer-grade-assessment/train_images/ \\\n--task seg \\\n--job_dir ./trident_processed \\\n--gpu 0 \\\n--segmenter hest \\\n--custom_list_of_wsis ./sample.csv","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-08T14:48:46.573101Z","iopub.execute_input":"2026-01-08T14:48:46.573351Z","iopub.status.idle":"2026-01-08T14:49:58.312616Z","shell.execute_reply.started":"2026-01-08T14:48:46.573331Z","shell.execute_reply":"2026-01-08T14:49:58.312063Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"%%bash\npython TRIDENT/run_batch_of_slides.py \\\n--wsi_dir /kaggle/input/prostate-cancer-grade-assessment/train_images/ \\\n--task coords \\\n--job_dir ./trident_processed \\\n--custom_list_of_wsis ./sample.csv \\\n--patch_size 224 \\\n--mag 20","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-08T14:49:58.313502Z","iopub.execute_input":"2026-01-08T14:49:58.313734Z","iopub.status.idle":"2026-01-08T14:50:04.811764Z","shell.execute_reply.started":"2026-01-08T14:49:58.313714Z","shell.execute_reply":"2026-01-08T14:50:04.811231Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"%%bash\npython TRIDENT/run_batch_of_slides.py \\\n--wsi_dir /kaggle/input/prostate-cancer-grade-assessment/train_images/ \\\n--task feat \\\n--job_dir ./trident_processed \\\n--gpu 0 \\\n--custom_list_of_wsis ./sample.csv \\\n--patch_size 224 \\\n--mag 20 \\\n--patch_encoder virchow2","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-08T14:50:04.812708Z","iopub.execute_input":"2026-01-08T14:50:04.812977Z","iopub.status.idle":"2026-01-08T14:51:22.128485Z","shell.execute_reply.started":"2026-01-08T14:50:04.812952Z","shell.execute_reply":"2026-01-08T14:51:22.127937Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"%%bash\npython TRIDENT/run_batch_of_slides.py \\\n--wsi_dir /kaggle/input/prostate-cancer-grade-assessment/train_images/ \\\n--task feat \\\n--job_dir ./trident_processed \\\n--gpu 0 \\\n--custom_list_of_wsis ./sample.csv \\\n--patch_size 224 \\\n--mag 20 \\\n--slide_encoder mean-virchow2","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-08T14:51:22.129357Z","iopub.execute_input":"2026-01-08T14:51:22.129592Z","iopub.status.idle":"2026-01-08T14:51:28.359666Z","shell.execute_reply.started":"2026-01-08T14:51:22.12957Z","shell.execute_reply":"2026-01-08T14:51:28.359111Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Features visualization and unsupervised learning","metadata":{}},{"cell_type":"code","source":"h5_file = tiff_file.replace(\"tiff\", \"h5\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-08T14:51:28.3606Z","iopub.execute_input":"2026-01-08T14:51:28.360963Z","iopub.status.idle":"2026-01-08T14:51:28.364569Z","shell.execute_reply.started":"2026-01-08T14:51:28.360874Z","shell.execute_reply":"2026-01-08T14:51:28.364022Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"We will open our .h5 file using h5py. The name is exactly the same, but with a .h5 as a suffix.","metadata":{}},{"cell_type":"code","source":"file = h5py.File(f\"/kaggle/working/trident_processed/20x_224px_0px_overlap/features_virchow2/{h5_file}\")\nfile.keys()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-08T14:51:28.36529Z","iopub.execute_input":"2026-01-08T14:51:28.365519Z","iopub.status.idle":"2026-01-08T14:51:28.384712Z","shell.execute_reply.started":"2026-01-08T14:51:28.3655Z","shell.execute_reply":"2026-01-08T14:51:28.384195Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"As you can see, this file contains coords and features of each patch.","metadata":{}},{"cell_type":"code","source":"coords = file[\"coords\"][()]\ncoords.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-08T14:51:28.385551Z","iopub.execute_input":"2026-01-08T14:51:28.385769Z","iopub.status.idle":"2026-01-08T14:51:28.403458Z","shell.execute_reply.started":"2026-01-08T14:51:28.38575Z","shell.execute_reply":"2026-01-08T14:51:28.402797Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"It has 747 patches, with 2 coordinates (x and y). ","metadata":{}},{"cell_type":"code","source":"features = file[\"features\"][()]\nfeatures.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-08T14:51:28.404378Z","iopub.execute_input":"2026-01-08T14:51:28.404717Z","iopub.status.idle":"2026-01-08T14:51:28.428025Z","shell.execute_reply.started":"2026-01-08T14:51:28.404687Z","shell.execute_reply":"2026-01-08T14:51:28.427381Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"And each patch has a 2560 dimentional embedding. ","metadata":{}},{"cell_type":"markdown","source":"But 2560 looks like a huge information, can we summary it in a lower dimensional space?","metadata":{}},{"cell_type":"code","source":"pca = PCA(0.8)\ntransformed = pca.fit_transform(features)\ntransformed = pd.DataFrame(transformed)\ntransformed","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-08T14:51:28.428753Z","iopub.execute_input":"2026-01-08T14:51:28.429029Z","iopub.status.idle":"2026-01-08T14:51:28.820907Z","shell.execute_reply.started":"2026-01-08T14:51:28.429002Z","shell.execute_reply":"2026-01-08T14:51:28.820204Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Let's visualize just the first components","metadata":{}},{"cell_type":"code","source":"sns.pairplot(transformed[transformed.columns[0:3]])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-08T14:51:28.821558Z","iopub.execute_input":"2026-01-08T14:51:28.82178Z","iopub.status.idle":"2026-01-08T14:51:29.912933Z","shell.execute_reply.started":"2026-01-08T14:51:28.821757Z","shell.execute_reply":"2026-01-08T14:51:29.912253Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Can you show the explained variance of your PCA? Plot values and cumulated.\n\n*hint: look for [PCA](https://scikit-learn.org/stable/modules/generated/sklearn.decomposition.PCA.html) attributes*","metadata":{}},{"cell_type":"code","source":"#Code here:","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-08T14:51:29.913956Z","iopub.execute_input":"2026-01-08T14:51:29.914327Z","iopub.status.idle":"2026-01-08T14:51:29.91745Z","shell.execute_reply.started":"2026-01-08T14:51:29.914298Z","shell.execute_reply":"2026-01-08T14:51:29.916823Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Can you cluster patches to check different groups of them?\n\n*hint: You can use [KMeans](http://https://scikit-learn.org/stable/modules/generated/sklearn.cluster.KMeans.html) or [GaussianMixtures](https://scikit-learn.org/stable/modules/generated/sklearn.mixture.GaussianMixture.html). Try with 4 groups.*","metadata":{}},{"cell_type":"code","source":"#Code here:","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-08T14:51:29.918247Z","iopub.execute_input":"2026-01-08T14:51:29.918462Z","iopub.status.idle":"2026-01-08T14:51:29.930921Z","shell.execute_reply.started":"2026-01-08T14:51:29.918442Z","shell.execute_reply":"2026-01-08T14:51:29.930432Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Plot the first 2 components of your PCA painted by cluster.\n\n*hint: Use hue=cluster as parameter to paint with [seaborn.scatterplot](https://seaborn.pydata.org/generated/seaborn.scatterplot.html)*","metadata":{}},{"cell_type":"code","source":"def plot_patches_sample(df, title):\n    n = 10\n    rows, cols = 2, 5\n    fig, axes = plt.subplots(rows, cols, figsize=(15, 6))\n    axes = axes.flatten()\n    for ax, (_, row) in zip(axes, df.sample(n, random_state=42).iterrows()):\n        patch = slide.read_region(\n            (row[0], row[1]),\n            level=0,\n            size=(224, 224)\n        ).convert(\"RGB\")\n        ax.imshow(patch)\n        ax.axis(\"off\")\n    for ax in axes[n:]:\n        ax.axis(\"off\")\n    plt.suptitle(title)\n    plt.tight_layout()\n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-08T14:51:29.931923Z","iopub.execute_input":"2026-01-08T14:51:29.932256Z","iopub.status.idle":"2026-01-08T14:51:29.945454Z","shell.execute_reply.started":"2026-01-08T14:51:29.932235Z","shell.execute_reply":"2026-01-08T14:51:29.944948Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#Code here:\n#Example: plot_patches_sample(cluster_0, \"Cluster 0\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-08T14:51:29.946232Z","iopub.execute_input":"2026-01-08T14:51:29.94646Z","iopub.status.idle":"2026-01-08T14:51:29.962569Z","shell.execute_reply.started":"2026-01-08T14:51:29.946442Z","shell.execute_reply":"2026-01-08T14:51:29.961938Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Separate your different clusters in different datasets.\n\n*hint: name your dfs cluster_0, cluster_1, and so on*\n\nUse the following function to plot some patches of your dataset.","metadata":{}},{"cell_type":"code","source":"#Code here:","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-08T14:51:29.965721Z","iopub.execute_input":"2026-01-08T14:51:29.965951Z","iopub.status.idle":"2026-01-08T14:51:29.975699Z","shell.execute_reply.started":"2026-01-08T14:51:29.965931Z","shell.execute_reply":"2026-01-08T14:51:29.975154Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Using the full dataset\n\nOkay, we've preprocessed and analyzed a single WSI. What about the remaining 99.99%?\n\nAt HPC UOH, we preprocessed the 10K slides to accelerate (a little bit) this process.\n\nLike in a cooking TV-show, we will import the full dataset already preprocessed :P ","metadata":{}},{"cell_type":"markdown","source":"We will train a supervised classification algorithm with our features:\n\nFor mathematicians, we will create a situation like *Xw = y*\n\nWhere *X* is a rectangular matrix with our features (one per slide). *w* will be a vector with our coeficients/weights (our goal is to obtain it), and *y* is our target vector.\n\nIf *X* is a tight matrix (N > M), this problem is overdetermined, and we can obtain a certain solution, that fits \"best\" than the rest.","metadata":{}},{"cell_type":"markdown","source":"But we have a problem! D: How do we use our patches to train a model? We already have a matrix of PatchesxEmbeddings per slide. What can we do?\n\nThe simplest way to obtain our matrix *A* is to mean-pool the patch features by slide. By luck, we already did this with TRIDENT.\n\nLet's import our slide-pooled dataset","metadata":{}},{"cell_type":"code","source":"wsi_id = []\nslide_dataset = []\nfor a in os.listdir(f\"/kaggle/input/panda-preprocessed/20x_224px_0px_overlap/slide_features_mean-virchow2\"):\n    file = h5py.File(f\"/kaggle/input/panda-preprocessed/20x_224px_0px_overlap/slide_features_mean-virchow2/{a}\")\n    slide_dataset.append(file[\"features\"])\n    wsi_id.append(a.split(\".\")[0])\nslide_dataset = np.array(slide_dataset)\nwsi_id = pd.DataFrame(wsi_id, columns=[\"image_id\"])\nwsi_id","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-08T15:43:16.432948Z","iopub.execute_input":"2026-01-08T15:43:16.433233Z","iopub.status.idle":"2026-01-08T15:43:36.01941Z","shell.execute_reply.started":"2026-01-08T15:43:16.433212Z","shell.execute_reply":"2026-01-08T15:43:36.018638Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"y = wsi_id.merge(train, on=\"image_id\", how=\"left\")[\"isup_grade\"]\nX = slide_dataset\nX.shape, y.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-08T15:43:36.020952Z","iopub.execute_input":"2026-01-08T15:43:36.021251Z","iopub.status.idle":"2026-01-08T15:43:36.0316Z","shell.execute_reply.started":"2026-01-08T15:43:36.021232Z","shell.execute_reply":"2026-01-08T15:43:36.031031Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"sns.countplot(y = y)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-08T15:43:36.032313Z","iopub.execute_input":"2026-01-08T15:43:36.03264Z","iopub.status.idle":"2026-01-08T15:43:36.149554Z","shell.execute_reply.started":"2026-01-08T15:43:36.032619Z","shell.execute_reply":"2026-01-08T15:43:36.149054Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Use a PCA + scatterplot (as before) to plot the different slides embeddings and paint it by label.","metadata":{}},{"cell_type":"code","source":"#Code here:","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-08T16:03:55.724283Z","iopub.execute_input":"2026-01-08T16:03:55.72511Z","iopub.status.idle":"2026-01-08T16:03:55.728117Z","shell.execute_reply.started":"2026-01-08T16:03:55.72508Z","shell.execute_reply":"2026-01-08T16:03:55.727392Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"We will build a pipeline with a PCA for dimensionality reducion + a linear model as our supervised learning method.","metadata":{}},{"cell_type":"code","source":"pipe = Pipeline([\n    (\"pca\", PCA(0.9)),\n    (\"lr\", LogisticRegression())\n])\nscores = cross_val_score(pipe, X, y, scoring = \"roc_auc_ovr\")\nscores.mean(), scores.std()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-08T15:43:36.157681Z","iopub.execute_input":"2026-01-08T15:43:36.158017Z","iopub.status.idle":"2026-01-08T15:43:38.786977Z","shell.execute_reply.started":"2026-01-08T15:43:36.157998Z","shell.execute_reply":"2026-01-08T15:43:38.78642Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Can you fit your model with a subset of your data and optimize parameters?\n\n*hint: Use [train_test_split](https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.train_test_split.html) to split your dataset. Use [GridSearchCV](https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.GridSearchCV.html) to optimize hyperparams. Just optimize l1_ratio of [LogisticRegression](https://scikit-learn.org/stable/modules/generated/sklearn.linear_model.LogisticRegression.html).*","metadata":{}},{"cell_type":"code","source":"#Code here:","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-08T15:43:38.787711Z","iopub.execute_input":"2026-01-08T15:43:38.788108Z","iopub.status.idle":"2026-01-08T15:43:38.792645Z","shell.execute_reply.started":"2026-01-08T15:43:38.788085Z","shell.execute_reply":"2026-01-08T15:43:38.792094Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Predict in your test dataset. Show metrics like accuracy, precision, recall and f1.\n\n*hint: Predict and then use [classification_report](https://scikit-learn.org/stable/modules/generated/sklearn.metrics.classification_report.html)*","metadata":{}},{"cell_type":"code","source":"#Code here:","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-08T15:43:38.793532Z","iopub.execute_input":"2026-01-08T15:43:38.794527Z","iopub.status.idle":"2026-01-08T15:43:38.815734Z","shell.execute_reply.started":"2026-01-08T15:43:38.794502Z","shell.execute_reply":"2026-01-08T15:43:38.815051Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Show a confussion matrix:\n\n*hint: Use y_pred and [confusion_matrix](https://scikit-learn.org/stable/modules/generated/sklearn.metrics.confusion_matrix.html)*","metadata":{}},{"cell_type":"code","source":"#Code here:","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-08T16:01:33.579283Z","iopub.execute_input":"2026-01-08T16:01:33.579598Z","iopub.status.idle":"2026-01-08T16:01:33.583121Z","shell.execute_reply.started":"2026-01-08T16:01:33.579575Z","shell.execute_reply":"2026-01-08T16:01:33.5825Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Show a ROC curve.\n\n*hint: For ROC curve you will need to predict probability of each slide, and to plot a line per label. You can use the function I coded below.*","metadata":{}},{"cell_type":"code","source":"def plot_roc_curve(y_test, y_prob):\n    n_classes = y_prob.shape[1]\n    classes = np.arange(n_classes)\n    y_test_bin = label_binarize(y_test, classes=classes)\n    plt.figure(figsize=(7, 6))\n    for i in range(n_classes):\n        fpr, tpr, _ = roc_curve(y_test_bin[:, i], y_prob[:, i])\n        roc_auc = auc(fpr, tpr)\n    \n        plt.plot(\n            fpr, tpr,\n            label=f\"Class {i} (AUC = {roc_auc:.3f})\"\n        )\n    plt.plot([0, 1], [0, 1], linestyle=\"--\", color=\"gray\")\n    plt.xlabel(\"False Positive Rate\")\n    plt.ylabel(\"True Positive Rate\")\n    plt.title(\"ROC Curve One-vs-Rest (Multiclass)\")\n    plt.legend()\n    plt.grid(alpha=0.3)\n    plt.tight_layout()\n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-08T16:05:55.119507Z","iopub.execute_input":"2026-01-08T16:05:55.119774Z","iopub.status.idle":"2026-01-08T16:05:55.125466Z","shell.execute_reply.started":"2026-01-08T16:05:55.119754Z","shell.execute_reply":"2026-01-08T16:05:55.124813Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#Code here:","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-08T15:43:38.816611Z","iopub.execute_input":"2026-01-08T15:43:38.816919Z","iopub.status.idle":"2026-01-08T15:43:38.835634Z","shell.execute_reply.started":"2026-01-08T15:43:38.816867Z","shell.execute_reply":"2026-01-08T15:43:38.835105Z"}},"outputs":[],"execution_count":null}]}