{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Autoencoder for anomaly detection\nWe saw in exploratory data analysis that there is a huge class imbalance in the training data, cancer images consists of only ~2% of all the images (you can find my EDA notebook [here](https://www.kaggle.com/code/leventelippenszky/eda-baseline-submission-age-normalized/notebook)). As I wanted to play around with autoencoders for an other project, I implemented an approach where I train an **autoencoder** on non-cancer images and use the reconstruction error to perform classification. \nAn autoencoder is an unsupervised model, which learns efficient representation of the data. It contains two components:\n* An **encoder** that takes an image as input, and outputs a low-dimensional embedding (representation) of the image.\n* A **decoder** that takes the low-dimensional embedding, and reconstructs the image.\n\nThe idead behind my approach is pretty straightforward:\n* The autoencoder is forced to learn a **compressed representation** of the image due to the *bottleneck architecture*\n* **Only non-cancer images** are used to **train** the network, and it learns how to reconstruct them with high fidelity\n* As a consequence, if a breast **cancer image** is fed into the network, the autoencoder will have a **harder time to reproduce** that image as it has never seen any before\n* We can use the **reconstruction errors scaled into [0,1] as predictions**, where the larger the error, the more probable it is a breast cancer image","metadata":{}},{"cell_type":"code","source":"try:\n    import pylibjpeg\nexcept:\n    !pip install /kaggle/input/rsna-2022-whl/{pydicom-2.3.0-py3-none-any.whl,pylibjpeg-1.4.0-py3-none-any.whl,python_gdcm-3.0.15-cp37-cp37m-manylinux_2_17_x86_64.manylinux2014_x86_64.whl}","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-12-30T14:09:57.359318Z","iopub.execute_input":"2022-12-30T14:09:57.359922Z","iopub.status.idle":"2022-12-30T14:10:12.971911Z","shell.execute_reply.started":"2022-12-30T14:09:57.359817Z","shell.execute_reply":"2022-12-30T14:10:12.970676Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os\nimport cv2\nimport glob\nimport pydicom\nimport numpy as np\nimport pandas as pd\nfrom typing import List\nimport matplotlib.pyplot as plt\nfrom sklearn.model_selection import StratifiedGroupKFold\nfrom sklearn.decomposition import PCA\n\nimport torch\nfrom torch import nn\nfrom torch.utils.data import Dataset, DataLoader\nfrom torch.optim import Adam\nfrom torchvision import transforms","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-12-30T14:10:12.974952Z","iopub.execute_input":"2022-12-30T14:10:12.975701Z","iopub.status.idle":"2022-12-30T14:10:17.844999Z","shell.execute_reply.started":"2022-12-30T14:10:12.975656Z","shell.execute_reply":"2022-12-30T14:10:17.843684Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"DATA_DIR = \"/kaggle/input/rsna-breast-cancer-detection\"\nTRAIN_IMG_DIR = \"/kaggle/input/rsna-breast-cancer-512-pngs\"\nDEVICE = \"cuda\" if torch.cuda.is_available() else \"cpu\"\n\nclass config:    \n    num_folds = 5\n    fold = 0\n    batch_size = 128\n    debug = False\n    num_epochs = 1 if debug else 10","metadata":{"execution":{"iopub.status.busy":"2022-12-30T14:10:17.84659Z","iopub.execute_input":"2022-12-30T14:10:17.847221Z","iopub.status.idle":"2022-12-30T14:10:17.974633Z","shell.execute_reply.started":"2022-12-30T14:10:17.84719Z","shell.execute_reply":"2022-12-30T14:10:17.973122Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df = pd.read_csv(os.path.join(DATA_DIR, \"train.csv\"))\nprint(train_df.shape)\ntrain_df","metadata":{"execution":{"iopub.status.busy":"2022-12-30T14:10:17.977932Z","iopub.execute_input":"2022-12-30T14:10:17.978641Z","iopub.status.idle":"2022-12-30T14:10:18.152071Z","shell.execute_reply.started":"2022-12-30T14:10:17.978593Z","shell.execute_reply":"2022-12-30T14:10:18.151047Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Dataset","metadata":{}},{"cell_type":"code","source":"class BreastCancerDataset(Dataset):\n    def __init__(self, df: pd.DataFrame, img_dir: str, transform=None):\n        self.df = df\n        self.img_dir = img_dir\n        self.transform = transform\n        \n    def __len__(self):\n        return self.df.shape[0]\n    \n    def __getitem__(self, idx):\n        image_path = os.path.join(TRAIN_IMG_DIR, f\"{self.df.loc[self.df.index[idx], 'patient_id']}_{self.df.loc[self.df.index[idx], 'image_id']}.png\")\n        image = cv2.imread(image_path)\n        \n        if self.transform:\n            image = self.transform(image)\n        \n        cancer = self.df.loc[self.df.index[idx], \"cancer\"]\n        return image, cancer\n\n    \ndef show_batch(batch: List[torch.Tensor]):\n    fig, axs = plt.subplots(nrows=2, ncols=np.ceil(len(batch[0])/2).astype(int), figsize=(15,15))\n    axs = axs.flatten()\n    for idx in range(len(batch[0])):\n        image = batch[0][idx].permute(1, 2, 0)\n        axs[idx].imshow(image, cmap=\"gray\")\n        axs[idx].axis(\"off\")\n        axs[idx].set_title(f\"Cancer label: {batch[1][idx]}\")\n        \n\ntransforms = transforms.Compose([transforms.ToTensor()])","metadata":{"execution":{"iopub.status.busy":"2022-12-30T14:10:18.153796Z","iopub.execute_input":"2022-12-30T14:10:18.154556Z","iopub.status.idle":"2022-12-30T14:10:18.16676Z","shell.execute_reply.started":"2022-12-30T14:10:18.154516Z","shell.execute_reply":"2022-12-30T14:10:18.165557Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ds = BreastCancerDataset(train_df, TRAIN_IMG_DIR, transform=transforms)\ntrain_loader = DataLoader(ds, batch_size=4)","metadata":{"execution":{"iopub.status.busy":"2022-12-30T14:10:18.168519Z","iopub.execute_input":"2022-12-30T14:10:18.168932Z","iopub.status.idle":"2022-12-30T14:10:18.181224Z","shell.execute_reply.started":"2022-12-30T14:10:18.168895Z","shell.execute_reply":"2022-12-30T14:10:18.180212Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"batch = next(iter(train_loader))\nshow_batch(batch)","metadata":{"execution":{"iopub.status.busy":"2022-12-30T14:10:18.182945Z","iopub.execute_input":"2022-12-30T14:10:18.183406Z","iopub.status.idle":"2022-12-30T14:10:19.196223Z","shell.execute_reply.started":"2022-12-30T14:10:18.183372Z","shell.execute_reply":"2022-12-30T14:10:19.195139Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Build autoencoder\n\nThe original image of size 3x512x512 is encoded into a 128x32x32 representation, which is a **6x compression**.\nI experimented with different network designs, these are the **takeaways** from the experiments:\n* decreasing the bottleneck layer by around an order of magnitude increases the loss by an order of magnitude\n* max pooling and max unpooling results in less detailed reconstructions, therere are small patches in the images with similar pixel intensities","metadata":{}},{"cell_type":"code","source":"def compute_conv_output_size(n: int, f: int, s: int, p: int, op: int=0, transpose: bool=False) -> int:\n    \"\"\"Computes the output size after a convolutional operation.\n    \n    Args:\n        n: Input size, h = w = n.\n        f: Filter size.\n        s: Stride.\n        p: Padding.\n        op: Output padding.\n        transpose: Boolean indicating transpose convolution.\n    \"\"\"\n    if transpose:\n        return (n - 1)*s - 2*p + f + op\n    else:    \n        return np.floor((n + 2*p - f)/s + 1).astype(int)\n\n\nclass Autoencoder(nn.Module):\n    def __init__(self):\n        super(Autoencoder, self).__init__()\n        # N, 3, 512, 512\n        self.encoder = nn.Sequential(\n            nn.Conv2d(3, 16, kernel_size=3, stride=2, padding=1), # N, 16, 256, 256\n            nn.ReLU(),\n            nn.Conv2d(16, 32, kernel_size=3, stride=2, padding=1), # N, 32, 128, 128\n            nn.ReLU(),\n            nn.Conv2d(32, 64, kernel_size=3, stride=2, padding=1), # N, 64, 64, 64\n            nn.ReLU(),\n            nn.Conv2d(64, 128, kernel_size=3, stride=2, padding=1), # N, 128, 32, 32\n        )\n        # N, 128, 32, 32\n        self.decoder = nn.Sequential(\n            nn.ConvTranspose2d(128, 64, kernel_size=3, stride=2, padding=1, output_padding=1), # N, 64, 64, 64\n            nn.ReLU(),\n            nn.ConvTranspose2d(64, 32, kernel_size=3, stride=2, padding=1, output_padding=1), # N, 32, 128, 128\n            nn.ReLU(),\n            nn.ConvTranspose2d(32, 16, kernel_size=3, stride=2, padding=1, output_padding=1), # N, 16, 256, 256\n            nn.ReLU(),\n            nn.ConvTranspose2d(16, 3, kernel_size=3, stride=2, padding=1, output_padding=1), # N, 3, 512, 512\n            nn.Sigmoid()\n        )\n\n    def forward(self, x):\n        x = self.encoder(x)\n        x = self.decoder(x)\n        return x\n    \n    def encode(self, x):\n        return self.encoder(x)","metadata":{"execution":{"iopub.status.busy":"2022-12-30T14:10:19.197346Z","iopub.execute_input":"2022-12-30T14:10:19.197691Z","iopub.status.idle":"2022-12-30T14:10:19.223578Z","shell.execute_reply.started":"2022-12-30T14:10:19.197656Z","shell.execute_reply":"2022-12-30T14:10:19.222765Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Train autoencoder\n\nThe autoencoder was trained on only easy negative images.","metadata":{}},{"cell_type":"code","source":"def create_folds(df: pd.DataFrame, num_folds: int):\n    df_proc = df.copy()\n    sgkf = StratifiedGroupKFold(num_folds)\n    for i, (_, val_index) in enumerate(sgkf.split(df_proc, df_proc[\"cancer\"], groups=df_proc[\"patient_id\"])):\n        df_proc.loc[df_proc.index[val_index], \"fold\"] = int(i)\n    return df_proc  \n    \n\ndef get_loaders_one_fold(fold: int, df: pd.DataFrame, batch_size: int):\n    # train on easy negative images only\n    df_train = df[(df[\"fold\"] != fold) & (df[\"cancer\"] == 0) & (~df[\"difficult_negative_case\"])].reset_index(drop=True)\n    train_ds = BreastCancerDataset(df_train, TRAIN_IMG_DIR, transform=transforms)\n    train_loader = DataLoader(train_ds, batch_size, shuffle=True, num_workers=os.cpu_count())\n    \n    df_val = df[df[\"fold\"] == fold].reset_index(drop=True)\n    val_ds = BreastCancerDataset(df_val, TRAIN_IMG_DIR, transform=transforms)\n    val_loader = DataLoader(val_ds, batch_size, shuffle=False, num_workers=os.cpu_count())\n    return train_loader, val_loader\n\n\nclass AverageCalc:\n    '''\n    Calculates and stores the average and current value.\n    Used to update the loss.\n    '''\n    def __init__(self):\n        self.reset()\n    \n    def reset(self):\n        self.value = 0\n        self.avg = 0\n        self.sum = 0\n        self.count = 0\n    \n    def update(self, value, size=1):\n        self.value = value\n        self.sum += value * size\n        self.count += size\n        self.avg = self.sum/self.count\n        \n\ndef train_fn(train_loader, model, optimizer, batch_size, debug):\n    loss_fn = nn.MSELoss(reduction=\"mean\")\n    run_loss = AverageCalc()\n    model.train()\n    for i, (x, _) in enumerate(train_loader):\n        if debug and (i > 10):\n            break\n        x = x.to(DEVICE)\n        xhat = model(x)\n\n        loss = loss_fn(xhat, x)\n        loss.backward()\n        run_loss.update(loss.item(), batch_size)\n\n        optimizer.step()\n        optimizer.zero_grad()\n    return run_loss.avg\n\n\n@torch.no_grad()\ndef val_fn(val_loader, model, batch_size, debug):\n    model.eval()\n    loss_fn = nn.MSELoss(reduction=\"none\")\n    losses = []\n    for i, (x, _) in enumerate(val_loader):\n        if debug and (i > 10):\n            break\n        x = x.to(DEVICE)\n        xhat = model(x)\n        loss = loss_fn(xhat, x)\n        losses.append(loss.mean(dim=(1,2,3)))\n    return torch.cat(losses, dim=0)\n\n\ndef run_training_one_fold(fold: int, df: pd.DataFrame, batch_size: int, num_epochs: int, debug: bool=False):\n    train_loader, val_loader = get_loaders_one_fold(fold, df, batch_size)\n    model = Autoencoder().to(DEVICE)\n    \n    optimizer = Adam(model.parameters())\n    for epoch in range(num_epochs):\n        train_loss = train_fn(train_loader, model, optimizer, batch_size, debug)\n        print(f\"Train loss epoch {epoch}: {train_loss:.4f}\")\n        val_losses = val_fn(val_loader, model, batch_size, debug)\n        print(f\"Val loss epoch {epoch}: {val_losses.mean():.4f}\")\n    torch.save(model.state_dict(), f\"autoencoder_fold{fold}.pt\")\n    return val_losses.detach().cpu().numpy()","metadata":{"execution":{"iopub.status.busy":"2022-12-30T14:10:19.225214Z","iopub.execute_input":"2022-12-30T14:10:19.225936Z","iopub.status.idle":"2022-12-30T14:10:19.259284Z","shell.execute_reply.started":"2022-12-30T14:10:19.225887Z","shell.execute_reply":"2022-12-30T14:10:19.257973Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df = create_folds(train_df, config.num_folds)\nval_losses = run_training_one_fold(fold=config.fold, df=train_df, batch_size=config.batch_size,\n                                   num_epochs=config.num_epochs, debug=config.debug)","metadata":{"execution":{"iopub.status.busy":"2022-12-30T14:10:19.266775Z","iopub.execute_input":"2022-12-30T14:10:19.267438Z","iopub.status.idle":"2022-12-30T14:11:21.117362Z","shell.execute_reply.started":"2022-12-30T14:10:19.267393Z","shell.execute_reply":"2022-12-30T14:11:21.115296Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## View reconstruction","metadata":{}},{"cell_type":"code","source":"def load_autoencoder(path: str):\n    model = Autoencoder()\n    model.load_state_dict(torch.load(path))\n    model.to(DEVICE)\n    return model\n\n\n@torch.no_grad()\ndef show_reconstruction(batch: List[torch.Tensor], model, num_images: int):\n    model.eval()\n    images = batch[0][:num_images, :, :, :]\n    cancer_labels = batch[1][:num_images]\n    fig, axs = plt.subplots(nrows=images.shape[0], ncols=3, figsize=(20, 30))\n    for idx in range(images.shape[0]):\n        image = images[idx, :, :, :].to(DEVICE)\n        encoded = model.encode(image)\n        decoded = model(image)\n        \n        axs[idx, 0].imshow(image.permute(1, 2, 0).to(\"cpu\"), cmap=\"gray\")\n        axs[idx, 0].set_title(f\"Original image, cancer label: {cancer_labels[idx]}\")\n        axs[idx, 1].imshow(torch.reshape(encoded.to(\"cpu\"), (512, 256)))\n        axs[idx, 1].set_title(\"Latent space\")\n        axs[idx, 2].imshow(decoded.permute(1, 2, 0).to(\"cpu\"), cmap=\"gray\")\n        axs[idx, 2].set_title(\"Reconstructed image\")\n    fig.tight_layout()","metadata":{"execution":{"iopub.status.busy":"2022-12-30T14:11:21.119238Z","iopub.execute_input":"2022-12-30T14:11:21.120396Z","iopub.status.idle":"2022-12-30T14:11:21.131588Z","shell.execute_reply.started":"2022-12-30T14:11:21.120344Z","shell.execute_reply":"2022-12-30T14:11:21.130479Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model = load_autoencoder(f\"autoencoder_fold{config.fold}.pt\")\n_, val_loader = get_loaders_one_fold(fold=0, df=train_df, batch_size=config.batch_size)\nbatch = next(iter(val_loader))\nshow_reconstruction(batch, model, num_images=4)","metadata":{"execution":{"iopub.status.busy":"2022-12-30T14:11:21.132938Z","iopub.execute_input":"2022-12-30T14:11:21.133546Z","iopub.status.idle":"2022-12-30T14:11:28.972288Z","shell.execute_reply.started":"2022-12-30T14:11:21.133437Z","shell.execute_reply":"2022-12-30T14:11:28.968955Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Reconstruction loss","metadata":{}},{"cell_type":"code","source":"def plot_val_loss_hist(df: pd.DataFrame, val_losses: np.ndarray, fold: int):\n    val_df = df[df[\"fold\"] == fold]\n    val_df[\"loss\"] = val_losses\n    \n    plt.figure(figsize=(12, 8))\n    val_df.loc[val_df[\"cancer\"] == 0, \"loss\"].plot.hist(bins=30, alpha=0.8, label=\"non-cancer\")\n    val_df.loc[val_df[\"cancer\"] == 1, \"loss\"].plot.hist(bins=30, alpha=0.8, label=\"cancer\")\n    plt.legend()\n    return val_df","metadata":{"execution":{"iopub.status.busy":"2022-12-30T14:11:28.974214Z","iopub.execute_input":"2022-12-30T14:11:28.974831Z","iopub.status.idle":"2022-12-30T14:11:28.982246Z","shell.execute_reply.started":"2022-12-30T14:11:28.974777Z","shell.execute_reply":"2022-12-30T14:11:28.980994Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if not config.debug:\n    val_df = plot_val_loss_hist(train_df, val_losses, config.fold)","metadata":{"execution":{"iopub.status.busy":"2022-12-30T14:11:28.983956Z","iopub.execute_input":"2022-12-30T14:11:28.984954Z","iopub.status.idle":"2022-12-30T14:11:28.996595Z","shell.execute_reply.started":"2022-12-30T14:11:28.984918Z","shell.execute_reply":"2022-12-30T14:11:28.995505Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Classification using reconstruction loss\n\nAs expected, unsupervised learning performs way worse that supervised in this problem.","metadata":{}},{"cell_type":"code","source":"# https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369267\ndef pfbeta_torch(labels, preds, beta=1):\n    preds = preds.clip(0, 1)\n    y_true_count = labels.sum()\n    ctp = preds[labels==1].sum()\n    cfp = preds[labels==0].sum()\n    beta_squared = beta * beta\n    c_precision = ctp / (ctp + cfp)\n    c_recall = ctp / y_true_count\n    if (c_precision > 0 and c_recall > 0):\n        result = (1 + beta_squared) * (c_precision * c_recall) / (beta_squared * c_precision + c_recall)\n        return result\n    else:\n        return 0.0","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-12-30T14:11:28.998067Z","iopub.execute_input":"2022-12-30T14:11:28.999068Z","iopub.status.idle":"2022-12-30T14:11:29.008642Z","shell.execute_reply.started":"2022-12-30T14:11:28.999009Z","shell.execute_reply":"2022-12-30T14:11:29.007785Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if not config.debug:\n    val_df[\"pred\"] = val_df[\"loss\"] / val_df[\"loss\"].max()\n    print(f\"Probabilistic F-score on validation set: {pfbeta_torch(val_df['cancer'].to_numpy(), val_df['pred'].to_numpy()):.4f}\")","metadata":{"execution":{"iopub.status.busy":"2022-12-30T14:11:29.010114Z","iopub.execute_input":"2022-12-30T14:11:29.011171Z","iopub.status.idle":"2022-12-30T14:11:29.026304Z","shell.execute_reply.started":"2022-12-30T14:11:29.011113Z","shell.execute_reply":"2022-12-30T14:11:29.024993Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Visualize latent space in 2 dimensions\n\nI visualized the first two principal components of the latent space. We saw from the results earlier that the autoencoder cannot really distinguish between the two classes. Therefore, the points of the two classes do not form separate clusters in the latent space either. ","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:00:28.928915Z","iopub.execute_input":"2022-12-28T09:00:28.929808Z","iopub.status.idle":"2022-12-28T09:00:28.935048Z","shell.execute_reply.started":"2022-12-28T09:00:28.929767Z","shell.execute_reply":"2022-12-28T09:00:28.933835Z"}}},{"cell_type":"code","source":"@torch.no_grad()\ndef get_encodings(loader, num_iter, model):\n    model.eval()\n    encodings_all = []\n    cancer_labels_all = []\n    for i, (images, cancer_labels) in enumerate(loader):\n        if i < num_iter:\n            images = images.to(DEVICE)\n            encodings = model.encode(images)\n            encodings_all.append(encodings.flatten(start_dim=1))\n            cancer_labels_all.append(cancer_labels)\n        else:\n            break\n    encodings_all = torch.cat(encodings_all, dim=0)\n    cancer_labels_all = torch.cat(cancer_labels_all, dim=0)\n    return encodings_all.detach().cpu().numpy(), cancer_labels_all.detach().cpu().numpy()\n\n\ndef pca_on_encodings(encodings: np.ndarray, cancer_labels: np.ndarray):\n    # center features\n    encodings = encodings - encodings.mean(axis=0)\n    pca = PCA(n_components=2)\n    encodings_2d = pca.fit_transform(encodings)\n    print(f\"Explained variance ratio: {pca.explained_variance_ratio_.sum():.2f}\")\n    \n    non_cancer_encodings_2d = encodings_2d[~cancer_labels.astype(bool), :]\n    cancer_encodings_2d = encodings_2d[cancer_labels.astype(bool), :]\n    plt.figure(figsize=(15, 15))\n    plt.plot(non_cancer_encodings_2d[:, 0], non_cancer_encodings_2d[:, 1], \"bo\", label=\"non-cancer\")\n    plt.plot(cancer_encodings_2d[:, 0], cancer_encodings_2d[:, 1], \"go\", label=\"cancer\")\n    plt.legend()","metadata":{"execution":{"iopub.status.busy":"2022-12-30T14:11:29.02797Z","iopub.execute_input":"2022-12-30T14:11:29.028436Z","iopub.status.idle":"2022-12-30T14:11:29.042936Z","shell.execute_reply.started":"2022-12-30T14:11:29.028401Z","shell.execute_reply":"2022-12-30T14:11:29.041683Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"encodings, cancer_labels = get_encodings(loader=val_loader, num_iter=30, model=model)\npca_on_encodings(encodings, cancer_labels)","metadata":{"execution":{"iopub.status.busy":"2022-12-30T14:11:29.044721Z","iopub.execute_input":"2022-12-30T14:11:29.04516Z","iopub.status.idle":"2022-12-30T14:12:37.157695Z","shell.execute_reply.started":"2022-12-30T14:11:29.045127Z","shell.execute_reply":"2022-12-30T14:12:37.156334Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## References\n* https://www.cs.toronto.edu/~lczhang/360/lec/w05/autoencoder.html\n* https://www.kaggle.com/code/cdeotte/dog-autoencoder\n* https://www.kaggle.com/code/robinteuwens/anomaly-detection-with-auto-encoders\n* https://en.wikipedia.org/wiki/Autoencoder\n* https://arxiv.org/pdf/1603.07285.pdf\n* https://github.com/vdumoulin/conv_arithmetic/blob/master/README.md\n","metadata":{}}]}