{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Clinical Background\nThe main subtypes present in the training data are:\n- CC - clear cell carcinoma\n- EC - endometrial carcinoma\n- HGSC - high grade serous carcinoma\n- LGSC - low grade serous carcinoma\n- MC - mucinous carcinoma\n\nAdditional rare subtypes that could account for the outlier Other label are:\n- Transitional cell carcinomas\n- Mixed carcinomas\n- Undifferentiated carcinomas\n- Squamous cell carcinoma\n- Papillary thyroid carcinoma\n- Sebaceous carcinoma\n- Carcinoid tumors associated with mature cystic teratomas\n\nHGSC is the most common type, followed by LGSC. CCC and EC follow with similar frequency, while primary MC is uncommon (majority are actually metastases from other areas).","metadata":{}},{"cell_type":"markdown","source":"# Histology\n\nImages in this dataset appear to use standard brightfield illumination with H&E stain. This stain is a good general purpose stain for morphology characterization. It lacks more detailed molecular data, but should be able to distinguish between the different cancer subtypes. See figures in [^1] and [^2] and ignore the other stains used.\n\n## Serous Carcinomas\n![](https://cdn.sanity.io/images/0vv8moc6/cancernetwork/fb08321ede2cf8cacb03004797203909fea9bb1a-481x483.png?fit=crop&auto=format)\n> HGSCs are characterized by papillary (Figure 1A), glandular, solid, and transitional (Figure 1B) patterns. The diagnosis of HGSC is typically straightforward, especially when there is a predominantly papillary pattern with associated psammoma bodies. The solid pattern can cause difficulty in distinguishing HGSC from EC. However, morphologic features such as glands that are slit-like rather than smooth/round, with prominent cellular budding and bizarre nuclei, are more typical of serous carcinoma. ECs, on the other hand, are associated with squamous differentiation, adenofibromatous background, and endometriosis.[^1]\n\nHGSC presents in older women and has a worse overall prognosis, while LGSC presents in younger women with a better prognosis. Unfortunately, we do not have age or other demographic data to help with classifications.\n\n## Endometrial Carcinoma\n> EC of the ovary is usually associated with endometriosis or an endometrioid adenofibroma/borderline tumor. It typically presents as a unilateral mass confined to the ovary-that is, as stage I disease. The diagnosis of EC in the majority of cases is based on the presence of back-to-back glands rather than infiltrative invasion (Figure 1D). Several characteristic histologic features are seen in EC, including squamous morules, mucinous differentiation, clear cell change, spindle morphology, and secretory change, with the latter sometimes mimicking CCC.[^1]\n\n> Lastly, it is important to mention the concept of de-differentiated carcinoma, which is the presence of an undifferentiated carcinoma adjacent to a low-grade EC. These tumors are often misdiagnosed as EC, FIGO grade 3. The entity was first described by Silva and colleagues, who found that the presence of the undifferentiated carcinoma component in ovarian and endometrial endometrioid carcinoma, regardless of the percentage, conferred a much worse prognosis. In the ovary, such tumors may represent the far end of the spectrum of HGSC carcinoma or can be associated with low-grade EC. Undifferentiated carcinoma is characterized by the presence of discohesive, monotonous cells resembling lymphoma.[^1]\n\nIdentifying and differentiating between EC and undifferentiated carcinoma may be an important outlier result to classify.\n\n## Clear Cell Carcinoma\n![](https://www.cancernetwork.com/_next/image?url=https%3A%2F%2Fcdn.sanity.io%2Fimages%2F0vv8moc6%2Fcancernetwork%2F7a794c25a823fc0919ad97215baae31bacfa57be-495x487.png%3Ffit%3Dcrop%26auto%3Dformat&w=1080&q=75)\n> CCC comprises approximately 10% of ovarian carcinomas. CCCs are often associated with endometriosis, similar to EC. They are rarely bilateral, and patients usually present with stage I or II disease. Histologically, CCCs are characterized by solid, papillary, and tubulocystic patterns (Figure 2A). The papillae are typically short with hyalinized fibrovascular cores and are lined by hobnail cells with clear cytoplasm (Figure 2B). The nuclei of CCC are usually uniformly atypical, with large nucleoli; notably, bizarre atypia is not frequently identified in HGSC. The diagnosis of CCC is challenging because both EC and HGSC can have prominent clear cell change, particularly in the solid and papillary areas. Studies have shown that tumors with clear cell change without the characteristic architectural pattern of CCC likely represent either EC or HGSC.[^1]\n\nThese are more difficult to classify as well, since clear cell changes can occur in several cancer types.\n\n## Mucinous Carcinoma\n![](https://www.cancernetwork.com/_next/image?url=https%3A%2F%2Fcdn.sanity.io%2Fimages%2F0vv8moc6%2Fcancernetwork%2F0ae51d1f348a8c8360db565b43ebc4b602ccaf35-498x500.png%3Ffit%3Dcrop%26auto%3Dformat&w=1080&q=75)\n> Ovarian MUCs often show a spectrum of changes that include areas resembling cystadenoma and borderline tumor. MUCs can show two patterns of invasion, namely expansile type (Figure 3A) and infiltrative type (Figure 3B). The expansile pattern is characterized by the presence of back-to-back glands, with minimal to no intervening stroma measuring at least 5 mm in one dimension. In addition, there is no evidence of stromal invasion or desmoplastic reaction. Infiltrative invasion is characterized by the presence of infiltrating glands with associated desmoplasia.\n\nThis class is likely one of the more difficult ones due to its relative rarity, and the difficulty of definitively identifying it versus cancers spreading from other parts of the body. It may also be difficult to distinguish between MC and EC as both can have mucinous differentiation.\n\n## Transitional Cell Tumors\n> Transitional cell neoplasms comprise approximately 1% to 2% of ovarian tumors. They are broadly classified as benign, borderline, and malignant Brenner tumor; and transitional cell carcinomas (TCCs). Malignant Brenner tumor is characterized by the presence of epithelium resembling urothelial neoplasms and shows the presence of stromal invasion. Most importantly, when diagnosing this subtype, areas of benign Brenner tumor must be identified and metastasis from the urinary tract should be excluded. TCC, on the other hand, is now widely considered to be a variant of HGSC and is typically positive for WT1, further supporting this impression.[^1]\n\nA rarer, potential outlier class to detect.\n\n[^1]: Morphologic, Immunophenotypic, and Molecular Features of Epithelial Ovarian Cancer. Cancer Network. Published February 16, 2016. Accessed October 7, 2023. [https://www.cancernetwork.com/view/morphologic-immunophenotypic-and-molecular-features-epithelial-ovarian-cancer](https://www.cancernetwork.com/view/morphologic-immunophenotypic-and-molecular-features-epithelial-ovarian-cancer)","metadata":{}},{"cell_type":"markdown","source":"## Challenges\nWhole slide images are an unreasonable resolution to train on. Luckily, [pjmathematician](https://www.kaggle.com/code/pjmathematician/ucbo-256-tiles-loading) has created a useful dataset of 256x256 tiles from the whole slide images. This zoomed in view is much easier to train and gives us access to smaller image details compared to using thumbnails, but may lose some of the overall context of the slide.\n\nFor this baseline, I am simply using a random patch per slide during training. For future work, it would improve performance to make use of all the tiles for each slide when training.\n\n## Pytorch Approach\nThis notebook uses a standard Pytorch pipeline, combined with Weights & Biases experiment tracking. To log your own runs, place your API key into Add-ons > Secrets, and add it with the \"wandb_api\" label. To run hyperparameter sweeps, configure and create the sweep on W&B to get the sweep ID and replace mine. Then, run the agent.\n\nA custom Dataset is created to simply retrieve random patches and the corresponding labels. A good first step is to replace the random image tile selection with some other method to use all tiles. Examples are split into training and validation sets, and are stratified by label to have a similar distribution. A standard cross entropy loss is used as the loss function, and some simple optimizer and scheduler presets are available as well. Finally, images are transformed using the standard Pytorch ResNet transforms, but additional transforms could be helpful to augment the dataset.","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport os\nimport seaborn as sns\nimport torch\nimport glob\nimport numpy as np\nimport wandb\nimport gc\nimport random\nimport copy\n\nimport torch.nn as nn\nimport torch.nn.functional as F\n\nfrom skimage import io\nimport torchvision\nimport torchvision.transforms as transforms\nfrom torchvision.models import ResNet50_Weights\nfrom torch.utils.data import Dataset, DataLoader, Subset\nfrom sklearn.model_selection import train_test_split\n\nfrom kaggle_secrets import UserSecretsClient\nuser_secrets = UserSecretsClient()\nwandb_api = user_secrets.get_secret(\"wandb_api\") \nwandb.login(key=wandb_api)\n\ninput_path = \"/kaggle/input/\"\ndevice = torch.device(\"cuda\" if (torch.cuda.is_available()) else \"cpu\")\n\nclasses = ['HGSC', 'LGSC', 'EC', 'CC', 'MC']\ncdict = {c:idx for (idx, c) in enumerate(classes)}\nprint(cdict)","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-10-26T05:12:39.418048Z","iopub.execute_input":"2023-10-26T05:12:39.418788Z","iopub.status.idle":"2023-10-26T05:12:39.813542Z","shell.execute_reply.started":"2023-10-26T05:12:39.418748Z","shell.execute_reply":"2023-10-26T05:12:39.812586Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class UBCDataset(Dataset):\n    def __init__(self, df, transform=None):\n        self.df = df\n        self.transform = transform\n    def __len__(self):\n        return len(self.df)\n    def __getitem__(self, idx):\n        label = cdict[self.df.loc[df.index[idx], 'label']]\n        path = self.df.loc[df.index[idx], 'tile_path']\n        img_paths = glob.glob(path + '/*.png')\n#         temp solution to just pick a random image\n        ip = img_paths[np.random.randint(len(img_paths))]\n        \n        img = read_img(ip)\n        \n        if self.transform:\n            img = self.transform(img)\n        \n        return img, label\n    \ndef get_image_path(image_id:int):\n    if 4 <= image_id <= 15188:\n        path = \"/kaggle/input/ucbo-tiles-256-1\"\n    elif 15209 <= image_id <= 30515:\n        path = \"/kaggle/input/ucbo-tiles-256-2\"\n    elif 30539 <= image_id <= 38687:\n        path = \"/kaggle/input/ucbo-tiles-256-3\"\n    elif 38849 <= image_id <= 65300:\n        path = \"/kaggle/input/ucbo-tiles-256-4\"\n    elif 65371 <= image_id <= 65533:\n        path = \"/kaggle/input/ucbo-tiles-256-5\"\n    return os.path.join(path, \"256_\"+str(image_id))\n\ndef read_img(path):\n    img = io.imread(path)\n    return img\n\ndef buildDataframe(path=input_path + \"UBC-OCEAN/train.csv\"):\n    train = pd.read_csv(path)\n    train['tile_path'] = train['image_id'].apply(lambda x: get_image_path(x))\n    train['ntiles'] = train['tile_path'].apply(lambda x: len(os.listdir(x)))\n    return train\n\ndef buildDatasets(df):\n    train_idx, val_idx = train_test_split(df.index, test_size=0.2, \n                                          shuffle=True, stratify=df.label)\n    transform = transforms.Compose([transforms.ToTensor(),\n                                    transforms.Normalize((0.5), (0.5))])\n    \n#     if config.model == \"PretrainedResNet50\":\n    transform = transforms.Compose([transforms.ToTensor(),\n                                    ResNet50_Weights.DEFAULT.transforms(antialias=True)])\n    \n    dataset = UBCDataset(df=df, transform=transform)\n\n    train_dataset = Subset(dataset, train_idx)\n    validation_dataset = Subset(dataset, val_idx)\n\n    return train_idx, train_dataset, validation_dataset\n\ndef buildDataloaders(df, train_idx, train_dataset, validation_dataset, \n                     batch_size, num_workers):\n    train_loader = DataLoader(train_dataset, batch_size=batch_size, num_workers=num_workers, shuffle=True, pin_memory=True, persistent_workers=True)\n    val_loader = DataLoader(validation_dataset, batch_size=batch_size, num_workers=num_workers, pin_memory=True, persistent_workers=True)\n\n    return train_loader, val_loader\n\ndef make(config=None):\n    df = buildDataframe()\n    train_idx, train_dataset, validation_dataset = buildDatasets(df)\n    train_loader, val_loader = buildDataloaders(df, train_idx, train_dataset, \n                                                validation_dataset, 16, 4)\n    criterion = nn.CrossEntropyLoss()\n    \n    if config.model == \"ResNet34\":\n        net = ResNet34().to(device)\n    elif config.model == \"PretrainedResNet50\":\n        net = PretrainedResNet50().to(device)\n    \n    if config.optimizer == \"SGD\":\n        optimizer = torch.optim.SGD(net.parameters(), lr=config.lr, momentum=config.momentum, weight_decay=config.weight_decay)\n    elif config.optimizer == \"AdamW\":\n        optimizer = torch.optim.AdamW(net.parameters(), lr=config.lr, weight_decay=config.weight_decay, betas=(config.beta1, config.beta2), fused=True)\n    \n    if config.scheduler == \"None\":\n        sched = None\n    elif config.scheduler == \"OneCycleLR\":\n        sched = torch.optim.lr_scheduler.OneCycleLR(optimizer, max_lr=config.max_lr, steps_per_epoch=len(train_loader), epochs=config.epochs)\n    \n    wandb.watch(net, criterion, log=\"all\", log_freq=26)\n    return net, train_loader, val_loader, optimizer, criterion, sched\n\ndef train_epoch(epoch, length, net, train_loader, optimizer, criterion, sched):\n    running_loss = 0.\n    avg_loss = 0.\n    \n    for i, data in enumerate(train_loader):\n        inputs, labels = data[0].to(device), data[1].to(device)\n\n        optimizer.zero_grad()\n        outputs = net(inputs)\n        loss = criterion(outputs, labels)\n        loss.backward()\n        optimizer.step()\n        \n        if sched is not None:\n            sched.step()\n        \n        running_loss += loss.item()\n        if i % length == length-1:\n            avg_loss = running_loss / len(train_loader)\n            n_examples  = (epoch*length) + i + 1\n            \n            print(f'[{epoch + 1:4d}, {n_examples:5d}] loss: {avg_loss:.3f}')\n            wandb.log({\"train\": {\"epoch\":epoch, \"avg_loss\":avg_loss}},\n                     step=n_examples)\n\ndef validate(epoch, length, net, val_loader, optimizer, criterion):\n    running_loss = 0.\n    avg_loss = 0.\n    \n    with torch.no_grad():\n        for i, data in enumerate(val_loader):\n            inputs, labels = data[0].to(device), data[1].to(device)\n            outputs = net(inputs)\n            loss = criterion(outputs, labels)\n\n            running_loss += loss.item()\n        \n        avg_loss = running_loss / (i+1)\n        \n        if np.isnan(avg_loss):\n            raise Exception(f'val avg_loss is NaN')\n        \n        print(f'             vloss: {avg_loss:.3f}')\n        wandb.log({\"val\": {\"epoch\":epoch, \"avg_loss\":avg_loss}},\n                 step=(epoch*length) + length)\n        \n        return avg_loss\n    \ndef train(config, net, train_loader, val_loader, optimizer, criterion, sched):\n    length = len(train_loader)\n    min_vloss = 999\n    best_net = None\n    \n    for epoch in range(config.epochs):\n        net.train()\n        train_epoch(epoch, length, net, train_loader, optimizer, criterion, sched)\n        \n        net.eval()\n        vloss = validate(epoch, length, net, val_loader, optimizer, criterion)\n        \n        if vloss < min_vloss:\n            wandb.run.summary[\"best_loss\"] = vloss\n            min_vloss = vloss\n            best_net = copy.deepcopy(net)\n    \n    return best_net\n\ndef run(config=None, mode=\"online\"):\n    with wandb.init(project=\"ubc-ocean\", config=config, mode=mode):\n        config = wandb.config\n        params = make(config)\n        gc.collect()\n        net = train(config, *params)\n        wandb.finish()\n        return net","metadata":{"execution":{"iopub.status.busy":"2023-10-26T05:12:39.815719Z","iopub.execute_input":"2023-10-26T05:12:39.816545Z","iopub.status.idle":"2023-10-26T05:12:39.847095Z","shell.execute_reply.started":"2023-10-26T05:12:39.816518Z","shell.execute_reply":"2023-10-26T05:12:39.846174Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class ResNet34(nn.Module):\n    def __init__(self, inc=3):\n        super(ResNet34, self).__init__()\n        self.in_channels = 64\n\n        self.conv1 = self.conv_block(inc, 64, kernel_size=7, stride=2, padding=3)\n        self.conv2_x = self.make_layer(64, 3)\n        self.conv3_x = self.make_layer(128, 4, stride=2)\n        self.conv4_x = self.make_layer(256, 6, stride=2)\n        self.conv5_x = self.make_layer(512, 3, stride=2)\n\n        self.avgpool = nn.AdaptiveAvgPool2d((1, 1))\n        self.fc1 = nn.Linear(512, 64)\n        self.fc2 = nn.Linear(64, 5)\n\n    def conv_block(self, in_channels, out_channels, kernel_size, stride=1, padding=0):\n        layers = [\n            nn.Conv2d(in_channels, out_channels, kernel_size, stride=stride, padding=padding, bias=False),\n            nn.BatchNorm2d(out_channels),\n            nn.ReLU()\n        ]\n        return nn.Sequential(*layers)\n\n    def make_layer(self, out_channels, num_blocks, stride=1):\n        layers = []\n        for _ in range(num_blocks):\n            layers.append(self.conv_block(self.in_channels, out_channels, kernel_size=3, stride=stride, padding=1))\n            self.in_channels = out_channels\n        return nn.Sequential(*layers)\n\n    def forward(self, x):\n        x = self.conv1(x)\n        x = self.conv2_x(x)\n        x = self.conv3_x(x)\n        x = self.conv4_x(x)\n        x = self.conv5_x(x)\n\n        x = self.avgpool(x)\n        x = torch.flatten(x, 1)\n        x = F.relu(self.fc1(x))\n        x = self.fc2(x)\n\n        return x\n\nclass PretrainedResNet50(nn.Module):\n    def __init__(self):\n        super().__init__()\n        self.rnet = torchvision.models.resnet50(weights=ResNet50_Weights.DEFAULT)\n        self.rnet.fc = torch.nn.Linear(in_features=2048, out_features=5)\n        \n    def forward(self, x):\n        x = self.rnet(x)\n        \n        return x","metadata":{"execution":{"iopub.status.busy":"2023-10-26T05:12:39.848398Z","iopub.execute_input":"2023-10-26T05:12:39.848743Z","iopub.status.idle":"2023-10-26T05:12:39.86577Z","shell.execute_reply.started":"2023-10-26T05:12:39.848716Z","shell.execute_reply":"2023-10-26T05:12:39.864997Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = buildDataframe()\ntrain_idx, train_dataset, validation_dataset = buildDatasets(df)\ntrain_loader, val_loader = buildDataloaders(df, train_idx, train_dataset, \n                                            validation_dataset, 16, 1)\n\nimg, labels = next(iter(train_loader))\nprint(img.shape)\nprint(labels.shape)\nio.imshow(img[0].numpy().transpose(1,2,0))\nprint(classes[labels[0]])","metadata":{"execution":{"iopub.status.busy":"2023-10-26T05:12:39.86806Z","iopub.execute_input":"2023-10-26T05:12:39.868435Z","iopub.status.idle":"2023-10-26T05:12:41.110086Z","shell.execute_reply.started":"2023-10-26T05:12:39.868402Z","shell.execute_reply":"2023-10-26T05:12:41.109024Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Ensure deterministic behavior\ntorch.backends.cudnn.deterministic = True\nrandom.seed(42)\nnp.random.seed(42)\ntorch.manual_seed(42)\ntorch.cuda.manual_seed_all(42)\n\nconfig = dict(\n    epochs=64,\n    optimizer=\"AdamW\",\n    lr=0.001,\n    momentum=0.9,\n    model=\"PretrainedResNet50\",\n    beta1=0.9,\n    beta2=0.999,\n    weight_decay=0.01,\n    max_lr=0.01,\n    scheduler=\"OneCycleLR\"\n)\n\n# sweep_id = \"\"\n# net = wandb.agent(sweep_id, run)\nnet = run(config, \"online\")","metadata":{"execution":{"iopub.status.busy":"2023-10-26T05:12:41.111744Z","iopub.execute_input":"2023-10-26T05:12:41.112084Z","iopub.status.idle":"2023-10-26T05:15:33.228551Z","shell.execute_reply.started":"2023-10-26T05:12:41.112048Z","shell.execute_reply":"2023-10-26T05:15:33.227202Z"},"trusted":true},"execution_count":null,"outputs":[]}]}