{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Introduction","metadata":{}},{"cell_type":"markdown","source":"This notebook represents one of my first attempts to face a real machine learning challenge. Out of the toy datasets that I have been using before, something fundamentally different happens here with the trainning and validation accuracies both staying close to 0.5 or quickly reaching close to 1. Everything here seems okay to me, so **I would kindly appreciate any tip that someone more experience can offer me.**\n\nAs this is an initial stage of the proccess to build my final model, this DNN will only use images as input. Once I see some kind of learning, I will proceed to include metadata like patient age.\n\nTo write this notebook I have read and adapted code from different sources, specially from Andrada Olteanu: https://www.kaggle.com/code/andradaolteanu/rsna-breast-cancer-eda-pytorch-baseline.\n\nThe preprocessed dataset provided by Paul Bacher is great in my opinion by offering images well cropped, with the same size and same orientation. ","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport matplotlib.image as mpimg\nimport gc\nfrom time import time\nimport datetime as dtime\nfrom datetime import datetime\nimport json\nimport os\nfrom tqdm import tqdm\n\n# torch and torchvision\nimport torch\nfrom torch import nn\nfrom torch.utils.data import Dataset\nfrom torch.utils.data import DataLoader\nfrom torch.optim.lr_scheduler import ReduceLROnPlateau\nimport torchvision\nfrom torchvision import transforms\nfrom torchvision.transforms import RandomRotation\nimport torchvision.models as models\n\n# albumentation for image augmentation\nfrom albumentations import (ToFloat, Normalize, VerticalFlip, HorizontalFlip, Compose, Resize,\n                            RandomBrightnessContrast, HueSaturationValue, Blur, GaussNoise,\n                            Rotate, RandomResizedCrop, Cutout, ShiftScaleRotate, ToGray, FancyPCA, Sharpen)\nfrom albumentations.pytorch import ToTensorV2\n\n# SKlearn\nfrom sklearn.model_selection import StratifiedKFold, GroupKFold\nfrom sklearn.metrics import accuracy_score, roc_auc_score, confusion_matrix\nfrom sklearn.preprocessing import LabelEncoder, OneHotEncoder","metadata":{"execution":{"iopub.status.busy":"2023-01-29T19:04:41.864992Z","iopub.execute_input":"2023-01-29T19:04:41.865658Z","iopub.status.idle":"2023-01-29T19:04:46.59421Z","shell.execute_reply.started":"2023-01-29T19:04:41.865551Z","shell.execute_reply":"2023-01-29T19:04:46.593203Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"os.chdir('/kaggle/input/rsna-breast-cancer-detection')\ntrain = pd.read_csv(\"train.csv\")\nbase_path = \"/kaggle/input/rsna-bcd-1024x512-preprocessed/train_images\"\ntrain['path'] = base_path + \"/\" + train.patient_id.astype(str) + \"/\" + train.image_id.astype(str) + \".png\"","metadata":{"execution":{"iopub.status.busy":"2023-01-29T19:04:46.597032Z","iopub.execute_input":"2023-01-29T19:04:46.597602Z","iopub.status.idle":"2023-01-29T19:04:46.79185Z","shell.execute_reply.started":"2023-01-29T19:04:46.597563Z","shell.execute_reply":"2023-01-29T19:04:46.790798Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Encode categorical variables\nle_1 = LabelEncoder()\nle_2 = LabelEncoder()\n\ntrain['laterality'] = le_1.fit_transform(train['laterality'])\ntrain['view'] = le_2.fit_transform(train['view'])\n\ntrain.head()","metadata":{"execution":{"iopub.status.busy":"2023-01-29T19:04:46.793286Z","iopub.execute_input":"2023-01-29T19:04:46.793887Z","iopub.status.idle":"2023-01-29T19:04:46.844591Z","shell.execute_reply.started":"2023-01-29T19:04:46.793846Z","shell.execute_reply":"2023-01-29T19:04:46.843407Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"a = train[train.cancer == 1]\n\nprint(\"The number of cancer positive images is {} from a total of {} images. This makes for a ratio of {}.\".format(len(a), len(train), len(a)/len(train)))","metadata":{"execution":{"iopub.status.busy":"2023-01-29T19:04:46.847504Z","iopub.execute_input":"2023-01-29T19:04:46.848563Z","iopub.status.idle":"2023-01-29T19:04:46.866445Z","shell.execute_reply.started":"2023-01-29T19:04:46.848522Z","shell.execute_reply":"2023-01-29T19:04:46.865346Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The original dataframe is very inbalanced with a really low number of cancer positive images. Ideally, this should be solved using image augmentation techniques. To keep this notebook simple I decided just to those patients that have a cancer so there is close to 50% cancer positive and cancer negative images. These images are likely to be similar.\n\nThis subset is located in the dataframe \"train2\".","metadata":{}},{"cell_type":"code","source":"## Balanced sample:\n\nids_cancer = train[train.cancer == 1].patient_id\ntrain2=train[train.patient_id.isin(ids_cancer)].reset_index(drop=True)\n\na = train2[train2.cancer == 1]\nprint(\"The number of cancer positive images is {} from a total of {} images. This makes for a ratio of {}.\".format(len(a), len(train2), len(a)/len(train2)))","metadata":{"execution":{"iopub.status.busy":"2023-01-29T19:04:46.86768Z","iopub.execute_input":"2023-01-29T19:04:46.86866Z","iopub.status.idle":"2023-01-29T19:04:46.883703Z","shell.execute_reply.started":"2023-01-29T19:04:46.868613Z","shell.execute_reply":"2023-01-29T19:04:46.882576Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train2.head()","metadata":{"execution":{"iopub.status.busy":"2023-01-29T19:04:46.885224Z","iopub.execute_input":"2023-01-29T19:04:46.885865Z","iopub.status.idle":"2023-01-29T19:04:46.907142Z","shell.execute_reply.started":"2023-01-29T19:04:46.885823Z","shell.execute_reply":"2023-01-29T19:04:46.906406Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"N_imgs = 6\nfig, ax = plt.subplots(1,N_imgs, figsize=(15, 15))\nfor n,i in train2.iloc[0:N_imgs].iterrows():\n    img = mpimg.imread(i.path)\n    ax[n].imshow(img)\n    ax[n].set_title(i.cancer)\n    ax[n].axis('off')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-01-29T19:04:46.908547Z","iopub.execute_input":"2023-01-29T19:04:46.909201Z","iopub.status.idle":"2023-01-29T19:04:47.847287Z","shell.execute_reply.started":"2023-01-29T19:04:46.909162Z","shell.execute_reply":"2023-01-29T19:04:47.846399Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Dataset","metadata":{}},{"cell_type":"code","source":"img.shape","metadata":{"execution":{"iopub.status.busy":"2023-01-29T19:04:47.848309Z","iopub.execute_input":"2023-01-29T19:04:47.848665Z","iopub.status.idle":"2023-01-29T19:04:47.856013Z","shell.execute_reply.started":"2023-01-29T19:04:47.848634Z","shell.execute_reply":"2023-01-29T19:04:47.85491Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class img_dataset(Dataset):\n    def __init__(self, df, show_original=False):\n        self.df = df\n        self.show_original = show_original\n        self.transform = Compose([\n            #Resize(height=512, width=256), \n            Resize(height=256, width=128), \n            #RandomResizedCrop(height=256, width=128, scale=(0.995,0.999)),\n            Blur(blur_limit=3),\n            #GaussNoise(var_limit=(0.00002,0.00003), mean=0.1),\n            Sharpen(alpha=(0.025,0.05)),\n            #RandomRotation(5),\n            ToTensorV2()])\n        self.transform2 = Compose([\n            Resize(height=256, width=128), \n            ToTensorV2()\n        ])\n\n    def __len__(self):\n        return len(self.df)\n\n    def __getitem__(self, idx):\n        img = mpimg.imread(self.df.iloc[idx].path)/255\n        img = img.astype('float32')\n        #label = OneHotEncoder(categories=[[0,1]]).fit_transform(np.array([self.df.iloc[idx].cancer]).reshape(-1,1)).toarray()\n        label = np.array(self.df.iloc[idx].cancer, dtype=np.float32)\n\n        # Apply transforms\n        if self.show_original:\n            origial_img = self.transform2(image=img)['image']\n            origial_img = np.concatenate([origial_img,origial_img,origial_img], axis = 0) \n        transf_image = self.transform(image=img)['image']\n        transf_image = np.concatenate([transf_image,transf_image,transf_image], axis = 0) \n        #rotater = RandomRotation(degrees=(0, 5))\n        #transf_image = rotater(transf_image)\n        \n        if self.show_original:\n            return origial_img, transf_image, label\n        else:\n            return transf_image, label","metadata":{"execution":{"iopub.status.busy":"2023-01-29T19:04:47.857026Z","iopub.execute_input":"2023-01-29T19:04:47.858195Z","iopub.status.idle":"2023-01-29T19:04:47.878157Z","shell.execute_reply.started":"2023-01-29T19:04:47.858161Z","shell.execute_reply":"2023-01-29T19:04:47.876851Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dataset = img_dataset(train2, show_original=True)\ntrain_dataloader = DataLoader(dataset, batch_size=2, shuffle=True)","metadata":{"execution":{"iopub.status.busy":"2023-01-29T19:04:47.888886Z","iopub.execute_input":"2023-01-29T19:04:47.889449Z","iopub.status.idle":"2023-01-29T19:04:47.897071Z","shell.execute_reply.started":"2023-01-29T19:04:47.889415Z","shell.execute_reply":"2023-01-29T19:04:47.896191Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## check that it works\nfor k, data in enumerate(train_dataloader):\n    original_img, image, targets = data\n    print(\n          \"Image:\", image.shape, \"\\n\" +\n          \"=\"*40)\n\n    # Plot augmented images\n    fig, ax = plt.subplots(1,image.shape[0],figsize=(5,5))\n    for n,img in enumerate(image):\n        ax[n].imshow(img.cpu().numpy().squeeze()[0])\n        ax[n].axis('off')\n        ax[n].set_title(int(targets[n].item()))\n      \n    plt.show()\n\n    # Plot original images\n    fig, ax = plt.subplots(1,original_img.shape[0],figsize=(5,5))\n    for n,img in enumerate(original_img):\n      ax[n].imshow(img.cpu().numpy().squeeze()[0])\n      ax[n].axis('off')\n      \n    plt.show()\n\n    if k>4:\n        break\n    ","metadata":{"execution":{"iopub.status.busy":"2023-01-29T19:04:47.900537Z","iopub.execute_input":"2023-01-29T19:04:47.901668Z","iopub.status.idle":"2023-01-29T19:04:50.651354Z","shell.execute_reply.started":"2023-01-29T19:04:47.901627Z","shell.execute_reply":"2023-01-29T19:04:50.65031Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"DEVICE = torch.device('cuda' if torch.cuda.is_available() else 'cpu')\nprint('Device available now:', DEVICE)\n\ndef data_to_device(data):\n    image, targets = data\n    return image.to(DEVICE), targets.to(DEVICE)","metadata":{"execution":{"iopub.status.busy":"2023-01-29T19:04:50.653047Z","iopub.execute_input":"2023-01-29T19:04:50.653762Z","iopub.status.idle":"2023-01-29T19:04:50.771822Z","shell.execute_reply.started":"2023-01-29T19:04:50.653722Z","shell.execute_reply":"2023-01-29T19:04:50.77072Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# CNN","metadata":{}},{"cell_type":"markdown","source":"To keep it simple and fast, I will use the first version of efficientnet.\n\nThe idea is to train all layers of this model instead of just using it as a feature extractor. ","metadata":{}},{"cell_type":"code","source":"model = models.efficientnet_b0(weights='IMAGENET1K_V1')\nfor params in model.parameters():\n    params.requires_grad = True\nmodel.classifier[1] = nn.Linear(in_features=1280, out_features=1)\n\nmodel = model.to(DEVICE)","metadata":{"execution":{"iopub.status.busy":"2023-01-29T19:04:59.304204Z","iopub.execute_input":"2023-01-29T19:04:59.304615Z","iopub.status.idle":"2023-01-29T19:04:59.312435Z","shell.execute_reply.started":"2023-01-29T19:04:59.304583Z","shell.execute_reply":"2023-01-29T19:04:59.311424Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def plot_images(x):\n    for n,a in enumerate(x):\n        plt.imshow(img.cpu()[n].numpy().squeeze()[0])\n        plt.axis('off')\n        plt.show()\n\n        print(torch.sigmoid(x)[n].item())\n        print(label[n].item())\n        print('======================================')\n\n    del x\n    gc.collect()","metadata":{"execution":{"iopub.status.busy":"2023-01-29T19:05:03.068496Z","iopub.execute_input":"2023-01-29T19:05:03.068849Z","iopub.status.idle":"2023-01-29T19:05:03.075712Z","shell.execute_reply.started":"2023-01-29T19:05:03.068814Z","shell.execute_reply":"2023-01-29T19:05:03.074718Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## Evaluate model before training.\n\ndataset = img_dataset(train2, show_original=False)\ntrain_dataloader = DataLoader(dataset, batch_size=2, shuffle=True)\n\nimg, label = data_to_device(next(iter(train_dataloader)))\nmodel.eval()\nwith torch.no_grad():\n    outputs = model(img)\n\n\nplot_images(x=outputs)","metadata":{"execution":{"iopub.status.busy":"2023-01-29T19:05:03.07704Z","iopub.execute_input":"2023-01-29T19:05:03.077824Z","iopub.status.idle":"2023-01-29T19:05:11.847627Z","shell.execute_reply.started":"2023-01-29T19:05:03.077789Z","shell.execute_reply":"2023-01-29T19:05:11.846567Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Trainning:","metadata":{}},{"cell_type":"markdown","source":"My trainning function will receive the dataframe and model as input, build dataloaders and train the model. The following block of code is heavily based on the one written by Andrada Olteanu.\n\nI am using crossvalidation, so the data is divided into multiple folds.","metadata":{}},{"cell_type":"code","source":"def train_loop(model, df):\n    # Split in folds\n    group_fold = GroupKFold(n_splits = FOLDS)\n\n    # Generate indices to split data into training and test set.\n    k_folds = group_fold.split(X = np.zeros(len(df)), \n                               y = df['cancer'],\n                               groups=df['image_id'])\n\n    # For each fold\n    for i, (train_index, valid_index) in enumerate(k_folds):\n        print('FOLD: {}'.format(i+1))\n\n        # --- Create Instances ---\n        # Best ROC score in this fold\n        best_roc = None\n        # Reset patience before every fold\n        patience_f = PATIENCE\n\n        # Optimizer/ Scheduler/ Criterion\n        optimizer = torch.optim.Adam(model.parameters(), lr = LR, \n                                     weight_decay=WD)\n        scheduler = ReduceLROnPlateau(optimizer=optimizer, mode='max', \n                                      patience=LR_PATIENCE, verbose=True, factor=LR_FACTOR)\n        criterion = nn.BCEWithLogitsLoss()\n\n\n        # --- Read in Data ---\n        train_data = df.iloc[train_index].reset_index(drop=True)\n        valid_data = df.iloc[valid_index].reset_index(drop=True)\n\n        # Create Data instances\n        train = img_dataset(train_data)\n        valid = img_dataset(valid_data)\n\n        # Dataloaders\n        train_loader = DataLoader(train, batch_size=BATCH_SIZE1, \n                                  shuffle=True, num_workers=WORKERS)\n        valid_loader = DataLoader(valid, batch_size=BATCH_SIZE2, \n                                  shuffle=False, num_workers=WORKERS)\n\n        # === EPOCHS ===\n        for epoch in range(EPOCHS):\n            print('EPOCH: {}'.format(epoch+1))\n            start_time = time()\n            correct = 0\n            train_losses = 0\n\n            # === TRAIN ===\n            # Sets the module in training mode.\n            model.train()\n\n            # For each batch\n            for k, data in tqdm(enumerate(train_loader)):\n                # Save them to device\n                image, targets = data_to_device(data)\n\n                # Clear gradients first; very important\n                # usually done BEFORE prediction\n                optimizer.zero_grad()\n\n                # Log Probabilities & Backpropagation\n                out = model(image)\n                loss = criterion(out, targets.unsqueeze(1).float())\n                loss.backward()\n                optimizer.step()\n\n                # --- Save information after this batch ---\n                # Save loss\n                train_losses += loss.item()\n                # From log probabilities to actual probabilities\n                train_preds = torch.round(torch.sigmoid(out)) # 0 and 1\n                # Number of correct predictions\n                correct += (train_preds.cpu() == targets.cpu().unsqueeze(1)).sum().item()\n\n            # Compute Train Accuracy\n            train_acc = correct / len(train_index)\n\n            # === EVAL ===\n            # Sets the model in evaluation mode.\n            model.eval()\n\n            # Create matrix to store evaluation predictions (for accuracy)\n            valid_preds = torch.zeros(size = (len(valid_index), 1), \n                                      device=DEVICE, dtype=torch.float32)\n\n            # Disables gradients (we need to be sure no optimization happens)\n            with torch.no_grad():\n                for k, data in tqdm(enumerate(valid_loader)):\n                    # Save them to device\n                    image, targets = data_to_device(data)\n\n                    out = model(image)\n                    pred = torch.sigmoid(out)\n                    valid_preds[k*image.shape[0] : k*image.shape[0] + image.shape[0]] = pred\n\n                # Calculate accuracy\n                valid_acc = accuracy_score(valid_data['cancer'].values, \n                                           torch.round(valid_preds.cpu()))\n                # Calculate ROC\n                valid_roc = roc_auc_score(valid_data['cancer'].values, \n                                          valid_preds.cpu())\n\n                # Calculate time on Train + Eval\n                duration = str(dtime.timedelta(seconds=time() - start_time))[:7]\n\n                # PRINT INFO\n                final_logs = '{} | Epoch: {}/{} | Loss: {:.4} | Acc_tr: {:.3} | Acc_vd: {:.3} | ROC: {:.3}'.\\\n                                format(duration, epoch+1, EPOCHS, \n                                       train_losses, train_acc, valid_acc, valid_roc)\n                print(final_logs)\n\n                # Update scheduler (for learning_rate)\n                scheduler.step(valid_roc)\n        # === CLEANING ===\n        # Clear memory\n        del train, valid, train_loader, valid_loader, image, targets\n        gc.collect()","metadata":{"execution":{"iopub.status.busy":"2023-01-29T19:05:49.327905Z","iopub.execute_input":"2023-01-29T19:05:49.328258Z","iopub.status.idle":"2023-01-29T19:05:49.346782Z","shell.execute_reply.started":"2023-01-29T19:05:49.328226Z","shell.execute_reply":"2023-01-29T19:05:49.345284Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"FOLDS = 3\nEPOCHS = 3\nPATIENCE = 3\nWORKERS = 2\nLR = 0.0005\nWD = 0.0\nLR_PATIENCE = 1            # 1 model not improving until lr is decreasing\nLR_FACTOR = 0.4            # by how much the lr is decreasing\n\nBATCH_SIZE1 = 8           # for train\nBATCH_SIZE2 = 4           # for valid","metadata":{"execution":{"iopub.status.busy":"2023-01-29T19:05:50.28817Z","iopub.execute_input":"2023-01-29T19:05:50.288562Z","iopub.status.idle":"2023-01-29T19:05:50.295035Z","shell.execute_reply.started":"2023-01-29T19:05:50.28853Z","shell.execute_reply":"2023-01-29T19:05:50.293566Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The dataframe train2 only contains images from patients with a cancer. That way, the classes cancer==1 and ==0 are balanced while images are relatively similar\n","metadata":{}},{"cell_type":"code","source":"train_loop(model=model, df=train2)","metadata":{"execution":{"iopub.status.busy":"2023-01-29T19:06:00.609417Z","iopub.execute_input":"2023-01-29T19:06:00.610146Z","iopub.status.idle":"2023-01-29T19:10:05.147811Z","shell.execute_reply.started":"2023-01-29T19:06:00.610109Z","shell.execute_reply":"2023-01-29T19:10:05.144589Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"img, label = data_to_device(next(iter(train_dataloader)))\nmodel.eval()\nwith torch.no_grad():\n    outputs = model(img)\n\n\nplot_images(x=outputs)","metadata":{"execution":{"iopub.status.busy":"2023-01-29T19:10:05.150584Z","iopub.execute_input":"2023-01-29T19:10:05.151154Z","iopub.status.idle":"2023-01-29T19:10:05.571436Z","shell.execute_reply.started":"2023-01-29T19:10:05.151114Z","shell.execute_reply":"2023-01-29T19:10:05.570326Z"},"trusted":true},"execution_count":null,"outputs":[]}]}