{
  "id": 465358,
  "title": "13th Place Solution for the UBC-OCEAN Competition",
  "url": "/competitions/UBC-OCEAN/discussion/465358",
  "author_name": "MPWARE",
  "post_date": "2024-01-04T00:16:15.002000",
  "votes": 48,
  "comment_count": 16,
  "views": 0,
  "content": "<p>First of all we would like to thank Kaggle and the sponsors for this interesting research competition.</p>\n<h1>Context</h1>\n<ul>\n<li>Business context: <a href=\"https://www.kaggle.com/competitions/UBC-OCEAN/overview\" target=\"_blank\">https://www.kaggle.com/competitions/UBC-OCEAN/overview</a></li>\n<li>Data context: <a href=\"https://www.kaggle.com/competitions/UBC-OCEAN/data\" target=\"_blank\">https://www.kaggle.com/competitions/UBC-OCEAN/data</a></li>\n</ul>\n<h1>Overview of the approach</h1>\n<p>Our final solution was a combination of separated high scoring TMA and WSI models. For WSI models we’ve used some self-supervised learning (SSL) pretrained features extractor executed on 224 pixels tiles followed by Multiple Instance Learning (MIL) models. Same for TMA model with additional regular Transformer and CNN backbones. Outliers detection takes place in postprocessing and relies on both embedding distance and probabilities distributions.</p>\n<h1>Detail of the submission</h1>\n<p><strong>WSI models:</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Fe18dd035e408c85875e7b74d4862bc2e%2Fwsi_model.png?generation=1704326923139467&amp;alt=media\" alt=\"\"><br>\n<strong>TMA models:</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2F5ba1a216d5328e17c20cc55e866fcbb5%2Ftma_model.png?generation=1704326948322245&amp;alt=media\" alt=\"\"><br>\n<strong>Post processing:</strong><br>\nTMA models have been trained with ArcFace loss and an ArcMarginProduct sub center module to try to separate embeddings space as much as possible per class. It allows detecting outliers based on a similarity distance. Our distance threshold has been fine tuned on public LB as we didn’t have any solid sample of real outlier/rare class. <a href=\"https://github.com/facebookresearch/faiss\" target=\"_blank\">Faiss</a> package has been used to find nearest neighbors.<br>\nWe also performed another thresholding based on probabilities distributions. When all probabilities are low enough we switch the predicted class as outlier. The threshold has been calibrated on CV to avoid more than 10% outliers and checked on public LB.</p>\n<h1>WSI model training</h1>\n<p>Before training the WSI model, our medical intuition supported by this article led us to hypothesize that the relevant information for the subtype prediction was probably more located at low level. All the 513 available WSI in the training set were thus downscaled to 10x magnification. Since all these WSI had a black unicolor background, we then performed an otsu thresholding, in order to discard all the background tiles. We then tiled all the detected tissue into N non-overlapping 224px tiles. Insofar as it was not possible to infer tumor segmentation during submission because of the time limitation, we decided to keep all the tumor and the non-tumor tiles for training.<br>\nAll these N tiles were then encoded with CTransPath and Lunit-DINO, 2 features extractors trained using self-supervised learning on diverse pathology dataset. According to this article about the robustness of these models to stain variations, we did not perform any kind of augmentation or normalization preprocessing.<br>\nWe then trained and evaluated several MIL architecture into a weighted ensemble. The best CV were obtained by combining three of them : </p>\n<ul>\n<li>Clustering-constrained attention MIL (<a href=\"https://github.com/mahmoodlab/CLAM\" target=\"_blank\">https://github.com/mahmoodlab/CLAM</a>)</li>\n<li>Dual-stream MIL (<a href=\"https://github.com/binli123/dsmil-wsi\" target=\"_blank\">https://github.com/binli123/dsmil-wsi</a>)</li>\n<li>A weighted sum of the embeddings.<br>\nMIL training procedure and parameters:</li>\n<li>CV4, Stratified Group KFold</li>\n<li>No augmentation</li>\n<li>No normalization</li>\n<li>Batch size =1, epochs = 32</li>\n<li>AdamW optimizer, CosineAnnealingLR, LR=5e-3</li>\n<li>Cross-Entropy Loss</li>\n</ul>\n<h1>TMA model training</h1>\n<p>The UBC training dataset was coming with only 25 TMA samples and we know, according to the description, that TMA in the test set are the majority. We’ve detected around 65% to 70% of images with sides less than 6000 pixels. We’ve decided to generate some TMA based on the WSI provided in the training set. We’ve developed a custom augmentation that is detailed in this notebook: <a href=\"https://www.kaggle.com/code/mpware/ubc-tma-generator-from-wsi\" target=\"_blank\">https://www.kaggle.com/code/mpware/ubc-tma-generator-from-wsi</a></p>\n<pre><code> ():\n    A.Compose([\n       SimulateTMA((-, -), radius_ratio=(, ), ellipse_ratio=(, ), angle=(-, ), \n                   background_color=(-, -, -), background_color_ratio=(, ), \n                   noise_level=(/, /), black_replacement_color=, p=, always_apply=),\n       A.OneOf([\n           Stainer(ref_images=tma_images, method=, luminosity=, p=),\n           Stainer(ref_images=, method=, luminosity=, p=),\n           Stainer(ref_images=tma_images, method=, luminosity=, p=),\n       ], p=),        \n   ], p=p)\n</code></pre>\n<p>The idea is to identify tiles with tumoral tissue and crop an ellipse shape as could be a real TMA picked by an operator.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Fd37aaf25fc882b4f0c477f975514b388%2Ftma_v1.png?generation=1704327065491575&amp;alt=media\" alt=\"\"><br>\nThe crops are then augmented with stains based on the 25 TMA as references. As the WSI magnification is mainly x20 the generated TMA are also x20. Here are some generated samples:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2F7d4e508fdbfd96b7f04f8356408803d2%2Ftma_v2.png?generation=1704327082748102&amp;alt=media\" alt=\"\"><br>\nA final step was to review them manually to drop bad generated TMA (especially when tumor mask was not available / complete). Our best private LB (0.58) was with such validated TMA. Unfortunately we did not select it as final submission.<br>\n<strong>MIL</strong> training procedure and parameters:</p>\n<ul>\n<li>CV4, Stratified Group KFold</li>\n<li>Random batch sampler to balance samples</li>\n<li>Augmentations: Stain: Vahadane, Macenko, Reinhard</li>\n<li>Mask on attention, batch size = 32, epochs = 32</li>\n<li>EMA</li>\n<li>AdamW optimizer, CosineAnnealingLR, LR=1e-3</li>\n<li>Cross Entropy Loss<br>\n<br><br>\n<strong>Transformer/CNN</strong> training procedure and parameters:</li>\n<li>ImageNet pretrained backbones (<a href=\"https://github.com/huggingface/pytorch-image-models\" target=\"_blank\">Timm</a>):<ul>\n<li>tiny_vit_21m_512.dist_in22k_ft_in1k</li>\n<li>tf_efficientnetv2_s_in21ft1k</li></ul></li>\n<li>Augmentations: <ul>\n<li>H/V flips, Rot90</li>\n<li>Stain: Vahadane, Macenko, Reinhard</li>\n<li>Random BrightnessContrast/Gamma, HueSaturationValue, ColorJitter, CLAHE</li>\n<li>GaussianBlur, MotionBlur, GaussNoise</li>\n<li>Cut Mix, DropOut</li></ul></li>\n<li>EMA</li>\n<li>Batch size = 32, Epochs = 32</li>\n<li>AdamW optimizer, CosineAnnealingLR, LR=1e-4</li>\n<li>Cross Entropy Loss<br>\nModels have been trained with full data after checking stability on cross validation.<br>\nHere is a 2D t-SNE projection of TinyVit trained embeddings on generated 23k TMA:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Fe0859f7f5adffe00c3e30cb246b40101%2Ftma_training_embeddings.png?generation=1704327121454397&amp;alt=media\" alt=\"\"></li>\n</ul>\n<h1>Other useful strategies or approaches</h1>\n<p>Validation was quite difficult, MIL models were overfitting quite fast. Using EMA helped to limit it. </p>\n<h1>Model inference</h1>\n<p>Most of the inference time was lost in image loading/tiling. We’ve implemented a multiprocess inference to benefit from all CPUs but optimized to balance the memory issues due to concurrent large images loading. It reduced the loading + tiling of all images to around 5h30-6h. Features exaction was the most time consuming task that is why we’ve limited ourselves to the two best ones. We’ve limited the number of tiles to 350 max and at the end our inference ran in around 11h15-30min.</p>\n<h1>What did not work or improve?</h1>\n<p>A quick sum up of what did not work or not improve:<br>\nResnet50-based features extractors such as RetCCL and Lunit-BT.<br>\nExternal data:</p>\n<ul>\n<li><a href=\"https://portal.gdc.cancer.gov/repository?facetTab=cases&amp;filters=%7B%22op%22%3A%22and%22%2C%22content%22%3A%5B%7B%22op%22%3A%22in%22%2C%22content%22%3A%7B%22field%22%3A%22cases.primary_site%22%2C%22value%22%3A%5B%22ovary%22%5D%7D%7D%2C%7B%22op%22%3A%22in%22%2C%22content%22%3A%7B%22field%22%3A%22files.data_type%22%2C%22value%22%3A%5B%22Slide%20Image%22%5D%7D%7D%5D%7D\" target=\"_blank\">TCGA</a></li>\n<li><a href=\"https://www.cancerimagingarchive.net/collection/ovarian-bevacizumab-response/\" target=\"_blank\">ATEC</a><br>\n<br><br>\nUsually adding more data is always better but here it did not help on both CV and LB. However the quality of many slides was very bad and could explain it. Also the labeling of some was not obvious.<br>\nTrain a WSI model based on ImageNet pre-trained backbone. It worked but Ctranspath and LunitDINO outperformed it.<br>\nDown scale to x5 for WSI models (instead of x10)<br>\nTrain a tumor segmentation model in order to sample tumor TMA, but since subtypes have significant morphological variations, we preferred to train a stroma segmentation model and predict the carcinoma mask by complementarity. Finally it was impossible to set up a unique threshold for TMA selection because of high variation in epithelial surface area between solid and mucinous architecture. <br>\nPseudo labeling has not been tried.</li>\n</ul>\n<h1>Sources</h1>\n<ul>\n<li>1) Deep Learning for Detecting BRCA Mutations in High-Grade Ovarian Cancer Based on an Innovative Tumor Segmentation Method From Whole Slide Images: <a href=\"https://www.modernpathology.org/article/S0893-3952(23)00209-0/fulltext\" target=\"_blank\">https://www.modernpathology.org/article/S0893-3952(23)00209-0/fulltext</a></li>\n<li>2) A Good Feature Extractor Is All You Need for Weakly Supervised Learning in Histopathology: <a href=\"https://arxiv.org/pdf/2311.11772.pdf\" target=\"_blank\">https://arxiv.org/pdf/2311.11772.pdf</a><br>\n<br><br>\nWe had a lot of fun solving this kaggle. It was a lot of data to handle, in addition to the ML challenge it was an optimization challenge to make the inference fast.<br>\n<br><br>\nRaphaël Bourgade and MPWARE</li>\n</ul>",
  "messages": [
    {
      "id": 2586060,
      "postDate": "2024-01-04T00:16:15.003Z",
      "content": "<p>First of all we would like to thank Kaggle and the sponsors for this interesting research competition.</p>\n<h1>Context</h1>\n<ul>\n<li>Business context: <a href=\"https://www.kaggle.com/competitions/UBC-OCEAN/overview\" target=\"_blank\">https://www.kaggle.com/competitions/UBC-OCEAN/overview</a></li>\n<li>Data context: <a href=\"https://www.kaggle.com/competitions/UBC-OCEAN/data\" target=\"_blank\">https://www.kaggle.com/competitions/UBC-OCEAN/data</a></li>\n</ul>\n<h1>Overview of the approach</h1>\n<p>Our final solution was a combination of separated high scoring TMA and WSI models. For WSI models we’ve used some self-supervised learning (SSL) pretrained features extractor executed on 224 pixels tiles followed by Multiple Instance Learning (MIL) models. Same for TMA model with additional regular Transformer and CNN backbones. Outliers detection takes place in postprocessing and relies on both embedding distance and probabilities distributions.</p>\n<h1>Detail of the submission</h1>\n<p><strong>WSI models:</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Fe18dd035e408c85875e7b74d4862bc2e%2Fwsi_model.png?generation=1704326923139467&amp;alt=media\" alt=\"\"><br>\n<strong>TMA models:</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2F5ba1a216d5328e17c20cc55e866fcbb5%2Ftma_model.png?generation=1704326948322245&amp;alt=media\" alt=\"\"><br>\n<strong>Post processing:</strong><br>\nTMA models have been trained with ArcFace loss and an ArcMarginProduct sub center module to try to separate embeddings space as much as possible per class. It allows detecting outliers based on a similarity distance. Our distance threshold has been fine tuned on public LB as we didn’t have any solid sample of real outlier/rare class. <a href=\"https://github.com/facebookresearch/faiss\" target=\"_blank\">Faiss</a> package has been used to find nearest neighbors.<br>\nWe also performed another thresholding based on probabilities distributions. When all probabilities are low enough we switch the predicted class as outlier. The threshold has been calibrated on CV to avoid more than 10% outliers and checked on public LB.</p>\n<h1>WSI model training</h1>\n<p>Before training the WSI model, our medical intuition supported by this article led us to hypothesize that the relevant information for the subtype prediction was probably more located at low level. All the 513 available WSI in the training set were thus downscaled to 10x magnification. Since all these WSI had a black unicolor background, we then performed an otsu thresholding, in order to discard all the background tiles. We then tiled all the detected tissue into N non-overlapping 224px tiles. Insofar as it was not possible to infer tumor segmentation during submission because of the time limitation, we decided to keep all the tumor and the non-tumor tiles for training.<br>\nAll these N tiles were then encoded with CTransPath and Lunit-DINO, 2 features extractors trained using self-supervised learning on diverse pathology dataset. According to this article about the robustness of these models to stain variations, we did not perform any kind of augmentation or normalization preprocessing.<br>\nWe then trained and evaluated several MIL architecture into a weighted ensemble. The best CV were obtained by combining three of them : </p>\n<ul>\n<li>Clustering-constrained attention MIL (<a href=\"https://github.com/mahmoodlab/CLAM\" target=\"_blank\">https://github.com/mahmoodlab/CLAM</a>)</li>\n<li>Dual-stream MIL (<a href=\"https://github.com/binli123/dsmil-wsi\" target=\"_blank\">https://github.com/binli123/dsmil-wsi</a>)</li>\n<li>A weighted sum of the embeddings.<br>\nMIL training procedure and parameters:</li>\n<li>CV4, Stratified Group KFold</li>\n<li>No augmentation</li>\n<li>No normalization</li>\n<li>Batch size =1, epochs = 32</li>\n<li>AdamW optimizer, CosineAnnealingLR, LR=5e-3</li>\n<li>Cross-Entropy Loss</li>\n</ul>\n<h1>TMA model training</h1>\n<p>The UBC training dataset was coming with only 25 TMA samples and we know, according to the description, that TMA in the test set are the majority. We’ve detected around 65% to 70% of images with sides less than 6000 pixels. We’ve decided to generate some TMA based on the WSI provided in the training set. We’ve developed a custom augmentation that is detailed in this notebook: <a href=\"https://www.kaggle.com/code/mpware/ubc-tma-generator-from-wsi\" target=\"_blank\">https://www.kaggle.com/code/mpware/ubc-tma-generator-from-wsi</a></p>\n<pre><code> ():\n    A.Compose([\n       SimulateTMA((-, -), radius_ratio=(, ), ellipse_ratio=(, ), angle=(-, ), \n                   background_color=(-, -, -), background_color_ratio=(, ), \n                   noise_level=(/, /), black_replacement_color=, p=, always_apply=),\n       A.OneOf([\n           Stainer(ref_images=tma_images, method=, luminosity=, p=),\n           Stainer(ref_images=, method=, luminosity=, p=),\n           Stainer(ref_images=tma_images, method=, luminosity=, p=),\n       ], p=),        \n   ], p=p)\n</code></pre>\n<p>The idea is to identify tiles with tumoral tissue and crop an ellipse shape as could be a real TMA picked by an operator.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Fd37aaf25fc882b4f0c477f975514b388%2Ftma_v1.png?generation=1704327065491575&amp;alt=media\" alt=\"\"><br>\nThe crops are then augmented with stains based on the 25 TMA as references. As the WSI magnification is mainly x20 the generated TMA are also x20. Here are some generated samples:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2F7d4e508fdbfd96b7f04f8356408803d2%2Ftma_v2.png?generation=1704327082748102&amp;alt=media\" alt=\"\"><br>\nA final step was to review them manually to drop bad generated TMA (especially when tumor mask was not available / complete). Our best private LB (0.58) was with such validated TMA. Unfortunately we did not select it as final submission.<br>\n<strong>MIL</strong> training procedure and parameters:</p>\n<ul>\n<li>CV4, Stratified Group KFold</li>\n<li>Random batch sampler to balance samples</li>\n<li>Augmentations: Stain: Vahadane, Macenko, Reinhard</li>\n<li>Mask on attention, batch size = 32, epochs = 32</li>\n<li>EMA</li>\n<li>AdamW optimizer, CosineAnnealingLR, LR=1e-3</li>\n<li>Cross Entropy Loss<br>\n<br><br>\n<strong>Transformer/CNN</strong> training procedure and parameters:</li>\n<li>ImageNet pretrained backbones (<a href=\"https://github.com/huggingface/pytorch-image-models\" target=\"_blank\">Timm</a>):<ul>\n<li>tiny_vit_21m_512.dist_in22k_ft_in1k</li>\n<li>tf_efficientnetv2_s_in21ft1k</li></ul></li>\n<li>Augmentations: <ul>\n<li>H/V flips, Rot90</li>\n<li>Stain: Vahadane, Macenko, Reinhard</li>\n<li>Random BrightnessContrast/Gamma, HueSaturationValue, ColorJitter, CLAHE</li>\n<li>GaussianBlur, MotionBlur, GaussNoise</li>\n<li>Cut Mix, DropOut</li></ul></li>\n<li>EMA</li>\n<li>Batch size = 32, Epochs = 32</li>\n<li>AdamW optimizer, CosineAnnealingLR, LR=1e-4</li>\n<li>Cross Entropy Loss<br>\nModels have been trained with full data after checking stability on cross validation.<br>\nHere is a 2D t-SNE projection of TinyVit trained embeddings on generated 23k TMA:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Fe0859f7f5adffe00c3e30cb246b40101%2Ftma_training_embeddings.png?generation=1704327121454397&amp;alt=media\" alt=\"\"></li>\n</ul>\n<h1>Other useful strategies or approaches</h1>\n<p>Validation was quite difficult, MIL models were overfitting quite fast. Using EMA helped to limit it. </p>\n<h1>Model inference</h1>\n<p>Most of the inference time was lost in image loading/tiling. We’ve implemented a multiprocess inference to benefit from all CPUs but optimized to balance the memory issues due to concurrent large images loading. It reduced the loading + tiling of all images to around 5h30-6h. Features exaction was the most time consuming task that is why we’ve limited ourselves to the two best ones. We’ve limited the number of tiles to 350 max and at the end our inference ran in around 11h15-30min.</p>\n<h1>What did not work or improve?</h1>\n<p>A quick sum up of what did not work or not improve:<br>\nResnet50-based features extractors such as RetCCL and Lunit-BT.<br>\nExternal data:</p>\n<ul>\n<li><a href=\"https://portal.gdc.cancer.gov/repository?facetTab=cases&amp;filters=%7B%22op%22%3A%22and%22%2C%22content%22%3A%5B%7B%22op%22%3A%22in%22%2C%22content%22%3A%7B%22field%22%3A%22cases.primary_site%22%2C%22value%22%3A%5B%22ovary%22%5D%7D%7D%2C%7B%22op%22%3A%22in%22%2C%22content%22%3A%7B%22field%22%3A%22files.data_type%22%2C%22value%22%3A%5B%22Slide%20Image%22%5D%7D%7D%5D%7D\" target=\"_blank\">TCGA</a></li>\n<li><a href=\"https://www.cancerimagingarchive.net/collection/ovarian-bevacizumab-response/\" target=\"_blank\">ATEC</a><br>\n<br><br>\nUsually adding more data is always better but here it did not help on both CV and LB. However the quality of many slides was very bad and could explain it. Also the labeling of some was not obvious.<br>\nTrain a WSI model based on ImageNet pre-trained backbone. It worked but Ctranspath and LunitDINO outperformed it.<br>\nDown scale to x5 for WSI models (instead of x10)<br>\nTrain a tumor segmentation model in order to sample tumor TMA, but since subtypes have significant morphological variations, we preferred to train a stroma segmentation model and predict the carcinoma mask by complementarity. Finally it was impossible to set up a unique threshold for TMA selection because of high variation in epithelial surface area between solid and mucinous architecture. <br>\nPseudo labeling has not been tried.</li>\n</ul>\n<h1>Sources</h1>\n<ul>\n<li>1) Deep Learning for Detecting BRCA Mutations in High-Grade Ovarian Cancer Based on an Innovative Tumor Segmentation Method From Whole Slide Images: <a href=\"https://www.modernpathology.org/article/S0893-3952(23)00209-0/fulltext\" target=\"_blank\">https://www.modernpathology.org/article/S0893-3952(23)00209-0/fulltext</a></li>\n<li>2) A Good Feature Extractor Is All You Need for Weakly Supervised Learning in Histopathology: <a href=\"https://arxiv.org/pdf/2311.11772.pdf\" target=\"_blank\">https://arxiv.org/pdf/2311.11772.pdf</a><br>\n<br><br>\nWe had a lot of fun solving this kaggle. It was a lot of data to handle, in addition to the ML challenge it was an optimization challenge to make the inference fast.<br>\n<br><br>\nRaphaël Bourgade and MPWARE</li>\n</ul>",
      "rawMarkdown": "\nFirst of all we would like to thank Kaggle and the sponsors for this interesting research competition.\n\n\n# Context\n\n\n\n* Business context: [https://www.kaggle.com/competitions/UBC-OCEAN/overview](https://www.kaggle.com/competitions/UBC-OCEAN/overview)\n* Data context: [https://www.kaggle.com/competitions/UBC-OCEAN/data](https://www.kaggle.com/competitions/UBC-OCEAN/data)\n\n\n# Overview of the approach\n\nOur final solution was a combination of separated high scoring TMA and WSI models. For WSI models we’ve used some self-supervised learning (SSL) pretrained features extractor executed on 224 pixels tiles followed by Multiple Instance Learning (MIL) models. Same for TMA model with additional regular Transformer and CNN backbones. Outliers detection takes place in postprocessing and relies on both embedding distance and probabilities distributions.\n\n\n# Detail of the submission\n\n**WSI models:**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Fe18dd035e408c85875e7b74d4862bc2e%2Fwsi_model.png?generation=1704326923139467&alt=media)\n\n\n**TMA models:**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2F5ba1a216d5328e17c20cc55e866fcbb5%2Ftma_model.png?generation=1704326948322245&alt=media)\n\n**Post processing:**\n\nTMA models have been trained with ArcFace loss and an ArcMarginProduct sub center module to try to separate embeddings space as much as possible per class. It allows detecting outliers based on a similarity distance. Our distance threshold has been fine tuned on public LB as we didn’t have any solid sample of real outlier/rare class. [Faiss](https://github.com/facebookresearch/faiss) package has been used to find nearest neighbors.\n\nWe also performed another thresholding based on probabilities distributions. When all probabilities are low enough we switch the predicted class as outlier. The threshold has been calibrated on CV to avoid more than 10% outliers and checked on public LB.\n\n\n# WSI model training\n\nBefore training the WSI model, our medical intuition supported by this article<sup>1</sup> led us to hypothesize that the relevant information for the subtype prediction was probably more located at low level. All the 513 available WSI in the training set were thus downscaled to 10x magnification. Since all these WSI had a black unicolor background, we then performed an otsu thresholding, in order to discard all the background tiles. We then tiled all the detected tissue into N non-overlapping 224px tiles. Insofar as it was not possible to infer tumor segmentation during submission because of the time limitation, we decided to keep all the tumor and the non-tumor tiles for training.\n\nAll these N tiles were then encoded with CTransPath and Lunit-DINO, 2 features extractors trained using self-supervised learning on diverse pathology dataset. According to this article<sup>2</sup> about the robustness of these models to stain variations, we did not perform any kind of augmentation or normalization preprocessing.\n\nWe then trained and evaluated several MIL architecture into a weighted ensemble. The best CV were obtained by combining three of them : \n\n\n\n* Clustering-constrained attention MIL ([https://github.com/mahmoodlab/CLAM](https://github.com/mahmoodlab/CLAM))\n* Dual-stream MIL ([https://github.com/binli123/dsmil-wsi](https://github.com/binli123/dsmil-wsi))\n* A weighted sum of the embeddings.\n\nMIL training procedure and parameters:\n\n\n\n* CV4, Stratified Group KFold\n* No augmentation\n* No normalization\n* Batch size =1, epochs = 32\n* AdamW optimizer, CosineAnnealingLR, LR=5e-3\n* Cross-Entropy Loss\n\n\n# TMA model training\n\nThe UBC training dataset was coming with only 25 TMA samples and we know, according to the description, that TMA in the test set are the majority. We’ve detected around 65% to 70% of images with sides less than 6000 pixels. We’ve decided to generate some TMA based on the WSI provided in the training set. We’ve developed a custom augmentation that is detailed in this notebook: [https://www.kaggle.com/code/mpware/ubc-tma-generator-from-wsi](https://www.kaggle.com/code/mpware/ubc-tma-generator-from-wsi)\n\n```python\ndef tma_augmentation(p=1.0):\n    return A.Compose([\n        SimulateTMA((-1, -1), radius_ratio=(0.6, 1.0), ellipse_ratio=(0.85, 1.15), angle=(-90., 90.), \n                    background_color=(-1, -1, -1), background_color_ratio=(0.80, 1.0), \n                    noise_level=(20./5, 100./5), black_replacement_color=None, p=1.0, always_apply=True),\n        A.OneOf([\n            Stainer(ref_images=tma_images, method='vahadane', luminosity=True, p=0.34),\n            Stainer(ref_images=None, method='macenko', luminosity=False, p=0.33),\n            Stainer(ref_images=tma_images, method='reinhard', luminosity=False, p=0.33),\n        ], p=0.60),        \n    ], p=p)\n```\n\nThe idea is to identify tiles with tumoral tissue and crop an ellipse shape as could be a real TMA picked by an operator.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Fd37aaf25fc882b4f0c477f975514b388%2Ftma_v1.png?generation=1704327065491575&alt=media)\n\n\nThe crops are then augmented with stains based on the 25 TMA as references. As the WSI magnification is mainly x20 the generated TMA are also x20. Here are some generated samples:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2F7d4e508fdbfd96b7f04f8356408803d2%2Ftma_v2.png?generation=1704327082748102&alt=media)\n\nA final step was to review them manually to drop bad generated TMA (especially when tumor mask was not available / complete). Our best private LB (0.58) was with such validated TMA. Unfortunately we did not select it as final submission.\n\n**MIL** training procedure and parameters:\n\n\n* CV4, Stratified Group KFold\n* Random batch sampler to balance samples\n* Augmentations: Stain: Vahadane, Macenko, Reinhard\n* Mask on attention, batch size = 32, epochs = 32\n* EMA\n* AdamW optimizer, CosineAnnealingLR, LR=1e-3\n* Cross Entropy Loss\n\n<br/>\n**Transformer/CNN** training procedure and parameters:\n\n\n* ImageNet pretrained backbones ([Timm](https://github.com/huggingface/pytorch-image-models)):\n    * tiny_vit_21m_512.dist_in22k_ft_in1k\n    * tf_efficientnetv2_s_in21ft1k\n* Augmentations: \n    * H/V flips, Rot90\n    * Stain: Vahadane, Macenko, Reinhard\n    * Random BrightnessContrast/Gamma, HueSaturationValue, ColorJitter, CLAHE\n    * GaussianBlur, MotionBlur, GaussNoise\n    * Cut Mix, DropOut\n* EMA\n* Batch size = 32, Epochs = 32\n* AdamW optimizer, CosineAnnealingLR, LR=1e-4\n* Cross Entropy Loss\n\nModels have been trained with full data after checking stability on cross validation.\n\nHere is a 2D t-SNE projection of TinyVit trained embeddings on generated 23k TMA:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Fe0859f7f5adffe00c3e30cb246b40101%2Ftma_training_embeddings.png?generation=1704327121454397&alt=media)\n\n\n\n# Other useful strategies or approaches\n\nValidation was quite difficult, MIL models were overfitting quite fast. Using EMA helped to limit it. \n\n\n# Model inference\n\nMost of the inference time was lost in image loading/tiling. We’ve implemented a multiprocess inference to benefit from all CPUs but optimized to balance the memory issues due to concurrent large images loading. It reduced the loading + tiling of all images to around 5h30-6h. Features exaction was the most time consuming task that is why we’ve limited ourselves to the two best ones. We’ve limited the number of tiles to 350 max and at the end our inference ran in around 11h15-30min.\n\n\n# What did not work or improve?\n\nA quick sum up of what did not work or not improve:\n\nResnet50-based features extractors such as RetCCL and Lunit-BT.\n\nExternal data:\n\n\n\n* [TCGA](https://portal.gdc.cancer.gov/repository?facetTab=cases&filters=%7B%22op%22%3A%22and%22%2C%22content%22%3A%5B%7B%22op%22%3A%22in%22%2C%22content%22%3A%7B%22field%22%3A%22cases.primary_site%22%2C%22value%22%3A%5B%22ovary%22%5D%7D%7D%2C%7B%22op%22%3A%22in%22%2C%22content%22%3A%7B%22field%22%3A%22files.data_type%22%2C%22value%22%3A%5B%22Slide%20Image%22%5D%7D%7D%5D%7D)\n* [ATEC](https://www.cancerimagingarchive.net/collection/ovarian-bevacizumab-response/)\n\n<br/>\nUsually adding more data is always better but here it did not help on both CV and LB. However the quality of many slides was very bad and could explain it. Also the labeling of some was not obvious.\n\nTrain a WSI model based on ImageNet pre-trained backbone. It worked but Ctranspath and LunitDINO outperformed it.\n\nDown scale to x5 for WSI models (instead of x10)\n\nTrain a tumor segmentation model in order to sample tumor TMA, but since subtypes have significant morphological variations, we preferred to train a stroma segmentation model and predict the carcinoma mask by complementarity. Finally it was impossible to set up a unique threshold for TMA selection because of high variation in epithelial surface area between solid and mucinous architecture. \n\nPseudo labeling has not been tried.\n\n\n# Sources\n\n\n* 1) Deep Learning for Detecting BRCA Mutations in High-Grade Ovarian Cancer Based on an Innovative Tumor Segmentation Method From Whole Slide Images: [https://www.modernpathology.org/article/S0893-3952(23)00209-0/fulltext](https://www.modernpathology.org/article/S0893-3952(23)00209-0/fulltext)\n* 2) A Good Feature Extractor Is All You Need for Weakly Supervised Learning in Histopathology: [https://arxiv.org/pdf/2311.11772.pdf](https://arxiv.org/pdf/2311.11772.pdf)\n\n<br/>\nWe had a lot of fun solving this kaggle. It was a lot of data to handle, in addition to the ML challenge it was an optimization challenge to make the inference fast.\n\n<br/>\nRaphaël Bourgade and MPWARE",
      "votes": 48
    },
    {
      "id": 2586517,
      "postDate": "2024-01-04T09:13:57.120Z",
      "content": "<p>The generation of TMA is a great idea; I would be curious if you observed any significant improvement with these rounded patches compared to using the easy squared tiles?</p>",
      "rawMarkdown": "The generation of TMA is a great idea; I would be curious if you observed any significant improvement with these rounded patches compared to using the easy squared tiles?",
      "votes": 1,
      "replies": [
        {
          "id": 2587434,
          "postDate": "2024-01-04T18:34:17.033Z",
          "content": "<p>That's a good question, our first TMAs generated were squared (small dataset) then we had the idea to be closer to real TMA like in inference. Indeed, we've not evaluted this specific impact because we've also improved models at the same time and it worked.</p>",
          "rawMarkdown": "That's a good question, our first TMAs generated were squared (small dataset) then we had the idea to be closer to real TMA like in inference. Indeed, we've not evaluted this specific impact because we've also improved models at the same time and it worked.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2586223,
      "postDate": "2024-01-04T04:01:55.513Z",
      "content": "<p>Congratulations on getting to the 15th position in this competition. Thanks for sharing the details of your model with colorful diagrams. </p>",
      "rawMarkdown": "Congratulations on getting to the 15th position in this competition. Thanks for sharing the details of your model with colorful diagrams. ",
      "votes": 1
    },
    {
      "id": 2586119,
      "postDate": "2024-01-04T01:43:52.157Z",
      "content": "<p>That’s a very solid solution. How much time did you spend (approximately) to implement tma generation part? How did you come up with the staining techniques? That is awesome.</p>",
      "rawMarkdown": "That’s a very solid solution. How much time did you spend (approximately) to implement tma generation part? How did you come up with the staining techniques? That is awesome.",
      "votes": 1,
      "replies": [
        {
          "id": 2586452,
          "postDate": "2024-01-04T08:25:01.780Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">@sergiosaharovskiy</a> , I would say it took around a few days to implement and a few hours and discussions with my teammate <a href=\"https://www.kaggle.com/raphaelbourgade\" target=\"_blank\">@raphaelbourgade</a> to find the correct settings to be close to real TMAs. Once the TMAs (around 23k) were generated we got better public LB results. To go beyond, we decided to review manually them to drop bad ones, it took (so) many hours for no benefit on public LB but +0.02 on private LB. </p>\n<p>For the stain techniques it's something regular in histopatology domain. We've just combined existing known packages in an <code>Albumentations</code> class and selected the 25 TMAs in train dataset as reference. In a paper we've referenced in \"Sources\" section, it is reported that stain augmentation is not so important and that features extractor is key. That's true, but based on our scores I can say stain augmentation helped to get around +0.01.</p>",
          "rawMarkdown": "Thanks @sergiosaharovskiy , I would say it took around a few days to implement and a few hours and discussions with my teammate @raphaelbourgade to find the correct settings to be close to real TMAs. Once the TMAs (around 23k) were generated we got better public LB results. To go beyond, we decided to review manually them to drop bad ones, it took (so) many hours for no benefit on public LB but +0.02 on private LB. \n\nFor the stain techniques it's something regular in histopatology domain. We've just combined existing known packages in an `Albumentations` class and selected the 25 TMAs in train dataset as reference. In a paper we've referenced in \"Sources\" section, it is reported that stain augmentation is not so important and that features extractor is key. That's true, but based on our scores I can say stain augmentation helped to get around +0.01.",
          "votes": 2,
          "replies": [
            {
              "id": 2586832,
              "postDate": "2024-01-04T12:31:42.183Z",
              "content": "<p>I am really impressed. It seems for me that it was a competition where the teamwork was essential. I appreciate your time for putting it all together.</p>",
              "rawMarkdown": "I am really impressed. It seems for me that it was a competition where the teamwork was essential. I appreciate your time for putting it all together.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2586082,
      "postDate": "2024-01-04T00:54:05.893Z",
      "content": "<p>What is your only TMA private score? My only private TMA is 0.33. My only private WSI is low 🥲</p>",
      "rawMarkdown": "What is your only TMA private score? My only private TMA is 0.33. My only private WSI is low 🥲",
      "votes": 1,
      "replies": [
        {
          "id": 2586092,
          "postDate": "2024-01-04T01:03:42.613Z",
          "content": "<p>Let me a few days to send a late submission to answer you</p>",
          "rawMarkdown": "Let me a few days to send a late submission to answer you",
          "replies": [
            {
              "id": 2588766,
              "postDate": "2024-01-05T17:59:32.243Z",
              "content": "<p><a href=\"https://www.kaggle.com/quan0095\" target=\"_blank\">@quan0095</a> I've submited and replaced all WSI predictions by HGSC predictions:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Fade4320e88b0a8fc9d5896d4b48537b3%2Ftma_ony.png?generation=1704477551260196&amp;alt=media\" alt=\"\"></p>",
              "rawMarkdown": "@quan0095 I've submited and replaced all WSI predictions by HGSC predictions:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Fade4320e88b0a8fc9d5896d4b48537b3%2Ftma_ony.png?generation=1704477551260196&alt=media)"
            },
            {
              "id": 2588816,
              "postDate": "2024-01-05T18:37:39.803Z",
              "content": "<p>I replace all WSI by \"AAA\" predictions. </p>",
              "rawMarkdown": "I replace all WSI by \"AAA\" predictions. ",
              "votes": 1
            },
            {
              "id": 2589312,
              "postDate": "2024-01-06T09:12:15.540Z",
              "content": "<p>Let me try that</p>",
              "rawMarkdown": "Let me try that"
            },
            {
              "id": 2591750,
              "postDate": "2024-01-08T07:12:27.470Z",
              "content": "<p><a href=\"https://www.kaggle.com/quan0095\" target=\"_blank\">@quan0095</a> 0.42 with AAA for WSI<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Fc504aaa8ac357370602af8e8d0c2fe88%2Ftma_only.png?generation=1704697928748728&amp;alt=media\" alt=\"\"></p>",
              "rawMarkdown": "@quan0095 0.42 with AAA for WSI\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Fc504aaa8ac357370602af8e8d0c2fe88%2Ftma_only.png?generation=1704697928748728&alt=media)",
              "votes": 2
            },
            {
              "id": 2591774,
              "postDate": "2024-01-08T07:31:49.440Z",
              "content": "<p>Your private only TMA is impressive</p>",
              "rawMarkdown": "Your private only TMA is impressive",
              "votes": 1
            },
            {
              "id": 2592746,
              "postDate": "2024-01-08T18:57:03.297Z",
              "content": "<p>We've put most of our efforts on TMA models.</p>",
              "rawMarkdown": "We've put most of our efforts on TMA models."
            }
          ]
        }
      ]
    },
    {
      "id": 2586884,
      "postDate": "2024-01-04T13:13:15.347Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> , I think generation TMA on WSI is a good idea. But I think use the stain of the TMA in training set is very dangerous in this competition. Since the test set is get from over 20 medical centers, which is very diverse, while the training set only has 25 TMAs. So if you use the stain on the TMA in training set, I think you suffer a risk of overfitting to the TMAs in training set. </p>",
      "rawMarkdown": "Hi @mpware , I think generation TMA on WSI is a good idea. But I think use the stain of the TMA in training set is very dangerous in this competition. Since the test set is get from over 20 medical centers, which is very diverse, while the training set only has 25 TMAs. So if you use the stain on the TMA in training set, I think you suffer a risk of overfitting to the TMAs in training set. ",
      "votes": 2,
      "replies": [
        {
          "id": 2587429,
          "postDate": "2024-01-04T18:28:55.123Z",
          "content": "<p>True, but for MIL model stain was applied on 60% of generated TMA, 2/3 with 25 TMA references and 1/3 with a generic one. For Transformer/CNN models additional color augmentation was added too.</p>",
          "rawMarkdown": "True, but for MIL model stain was applied on 60% of generated TMA, 2/3 with 25 TMA references and 1/3 with a generic one. For Transformer/CNN models additional color augmentation was added too.",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2586517,
      "author_name": "Jirka",
      "author_url": "",
      "post_date": "2024-01-04T09:13:57.120000",
      "content": "<p>The generation of TMA is a great idea; I would be curious if you observed any significant improvement with these rounded patches compared to using the easy squared tiles?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2587434,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2024-01-04T18:34:17.033000",
          "content": "<p>That's a good question, our first TMAs generated were squared (small dataset) then we had the idea to be closer to real TMA like in inference. Indeed, we've not evaluted this specific impact because we've also improved models at the same time and it worked.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2586223,
      "author_name": "C R Suthikshn Kumar",
      "author_url": "",
      "post_date": "2024-01-04T04:01:55.513000",
      "content": "<p>Congratulations on getting to the 15th position in this competition. Thanks for sharing the details of your model with colorful diagrams. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2586119,
      "author_name": "SSS",
      "author_url": "",
      "post_date": "2024-01-04T01:43:52.157000",
      "content": "<p>That’s a very solid solution. How much time did you spend (approximately) to implement tma generation part? How did you come up with the staining techniques? That is awesome.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2586452,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2024-01-04T08:25:01.780000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">@sergiosaharovskiy</a> , I would say it took around a few days to implement and a few hours and discussions with my teammate <a href=\"https://www.kaggle.com/raphaelbourgade\" target=\"_blank\">@raphaelbourgade</a> to find the correct settings to be close to real TMAs. Once the TMAs (around 23k) were generated we got better public LB results. To go beyond, we decided to review manually them to drop bad ones, it took (so) many hours for no benefit on public LB but +0.02 on private LB. </p>\n<p>For the stain techniques it's something regular in histopatology domain. We've just combined existing known packages in an <code>Albumentations</code> class and selected the 25 TMAs in train dataset as reference. In a paper we've referenced in \"Sources\" section, it is reported that stain augmentation is not so important and that features extractor is key. That's true, but based on our scores I can say stain augmentation helped to get around +0.01.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2586832,
              "author_name": "SSS",
              "author_url": "",
              "post_date": "2024-01-04T12:31:42.183000",
              "content": "<p>I am really impressed. It seems for me that it was a competition where the teamwork was essential. I appreciate your time for putting it all together.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2586082,
      "author_name": "Quan Vu",
      "author_url": "",
      "post_date": "2024-01-04T00:54:05.893000",
      "content": "<p>What is your only TMA private score? My only private TMA is 0.33. My only private WSI is low 🥲</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2586092,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2024-01-04T01:03:42.613000",
          "content": "<p>Let me a few days to send a late submission to answer you</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2588766,
              "author_name": "MPWARE",
              "author_url": "",
              "post_date": "2024-01-05T17:59:32.243000",
              "content": "<p><a href=\"https://www.kaggle.com/quan0095\" target=\"_blank\">@quan0095</a> I've submited and replaced all WSI predictions by HGSC predictions:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Fade4320e88b0a8fc9d5896d4b48537b3%2Ftma_ony.png?generation=1704477551260196&amp;alt=media\" alt=\"\"></p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2588816,
              "author_name": "Quan Vu",
              "author_url": "",
              "post_date": "2024-01-05T18:37:39.803000",
              "content": "<p>I replace all WSI by \"AAA\" predictions. </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2589312,
              "author_name": "MPWARE",
              "author_url": "",
              "post_date": "2024-01-06T09:12:15.540000",
              "content": "<p>Let me try that</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2591750,
              "author_name": "MPWARE",
              "author_url": "",
              "post_date": "2024-01-08T07:12:27.470000",
              "content": "<p><a href=\"https://www.kaggle.com/quan0095\" target=\"_blank\">@quan0095</a> 0.42 with AAA for WSI<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Fc504aaa8ac357370602af8e8d0c2fe88%2Ftma_only.png?generation=1704697928748728&amp;alt=media\" alt=\"\"></p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2591774,
              "author_name": "Quan Vu",
              "author_url": "",
              "post_date": "2024-01-08T07:31:49.440000",
              "content": "<p>Your private only TMA is impressive</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2592746,
              "author_name": "MPWARE",
              "author_url": "",
              "post_date": "2024-01-08T18:57:03.297000",
              "content": "<p>We've put most of our efforts on TMA models.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2586884,
      "author_name": "ForcewithMe",
      "author_url": "",
      "post_date": "2024-01-04T13:13:15.347000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> , I think generation TMA on WSI is a good idea. But I think use the stain of the TMA in training set is very dangerous in this competition. Since the test set is get from over 20 medical centers, which is very diverse, while the training set only has 25 TMAs. So if you use the stain on the TMA in training set, I think you suffer a risk of overfitting to the TMAs in training set. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 2587429,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2024-01-04T18:28:55.123000",
          "content": "<p>True, but for MIL model stain was applied on 60% of generated TMA, 2/3 with 25 TMA references and 1/3 with a generic one. For Transformer/CNN models additional color augmentation was added too.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2586060": "\nFirst of all we would like to thank Kaggle and the sponsors for this interesting research competition.\n\n\n# Context\n\n\n\n* Business context: [https://www.kaggle.com/competitions/UBC-OCEAN/overview](https://www.kaggle.com/competitions/UBC-OCEAN/overview)\n* Data context: [https://www.kaggle.com/competitions/UBC-OCEAN/data](https://www.kaggle.com/competitions/UBC-OCEAN/data)\n\n\n# Overview of the approach\n\nOur final solution was a combination of separated high scoring TMA and WSI models. For WSI models we’ve used some self-supervised learning (SSL) pretrained features extractor executed on 224 pixels tiles followed by Multiple Instance Learning (MIL) models. Same for TMA model with additional regular Transformer and CNN backbones. Outliers detection takes place in postprocessing and relies on both embedding distance and probabilities distributions.\n\n\n# Detail of the submission\n\n**WSI models:**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Fe18dd035e408c85875e7b74d4862bc2e%2Fwsi_model.png?generation=1704326923139467&alt=media)\n\n\n**TMA models:**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2F5ba1a216d5328e17c20cc55e866fcbb5%2Ftma_model.png?generation=1704326948322245&alt=media)\n\n**Post processing:**\n\nTMA models have been trained with ArcFace loss and an ArcMarginProduct sub center module to try to separate embeddings space as much as possible per class. It allows detecting outliers based on a similarity distance. Our distance threshold has been fine tuned on public LB as we didn’t have any solid sample of real outlier/rare class. [Faiss](https://github.com/facebookresearch/faiss) package has been used to find nearest neighbors.\n\nWe also performed another thresholding based on probabilities distributions. When all probabilities are low enough we switch the predicted class as outlier. The threshold has been calibrated on CV to avoid more than 10% outliers and checked on public LB.\n\n\n# WSI model training\n\nBefore training the WSI model, our medical intuition supported by this article<sup>1</sup> led us to hypothesize that the relevant information for the subtype prediction was probably more located at low level. All the 513 available WSI in the training set were thus downscaled to 10x magnification. Since all these WSI had a black unicolor background, we then performed an otsu thresholding, in order to discard all the background tiles. We then tiled all the detected tissue into N non-overlapping 224px tiles. Insofar as it was not possible to infer tumor segmentation during submission because of the time limitation, we decided to keep all the tumor and the non-tumor tiles for training.\n\nAll these N tiles were then encoded with CTransPath and Lunit-DINO, 2 features extractors trained using self-supervised learning on diverse pathology dataset. According to this article<sup>2</sup> about the robustness of these models to stain variations, we did not perform any kind of augmentation or normalization preprocessing.\n\nWe then trained and evaluated several MIL architecture into a weighted ensemble. The best CV were obtained by combining three of them : \n\n\n\n* Clustering-constrained attention MIL ([https://github.com/mahmoodlab/CLAM](https://github.com/mahmoodlab/CLAM))\n* Dual-stream MIL ([https://github.com/binli123/dsmil-wsi](https://github.com/binli123/dsmil-wsi))\n* A weighted sum of the embeddings.\n\nMIL training procedure and parameters:\n\n\n\n* CV4, Stratified Group KFold\n* No augmentation\n* No normalization\n* Batch size =1, epochs = 32\n* AdamW optimizer, CosineAnnealingLR, LR=5e-3\n* Cross-Entropy Loss\n\n\n# TMA model training\n\nThe UBC training dataset was coming with only 25 TMA samples and we know, according to the description, that TMA in the test set are the majority. We’ve detected around 65% to 70% of images with sides less than 6000 pixels. We’ve decided to generate some TMA based on the WSI provided in the training set. We’ve developed a custom augmentation that is detailed in this notebook: [https://www.kaggle.com/code/mpware/ubc-tma-generator-from-wsi](https://www.kaggle.com/code/mpware/ubc-tma-generator-from-wsi)\n\n```python\ndef tma_augmentation(p=1.0):\n    return A.Compose([\n        SimulateTMA((-1, -1), radius_ratio=(0.6, 1.0), ellipse_ratio=(0.85, 1.15), angle=(-90., 90.), \n                    background_color=(-1, -1, -1), background_color_ratio=(0.80, 1.0), \n                    noise_level=(20./5, 100./5), black_replacement_color=None, p=1.0, always_apply=True),\n        A.OneOf([\n            Stainer(ref_images=tma_images, method='vahadane', luminosity=True, p=0.34),\n            Stainer(ref_images=None, method='macenko', luminosity=False, p=0.33),\n            Stainer(ref_images=tma_images, method='reinhard', luminosity=False, p=0.33),\n        ], p=0.60),        \n    ], p=p)\n```\n\nThe idea is to identify tiles with tumoral tissue and crop an ellipse shape as could be a real TMA picked by an operator.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Fd37aaf25fc882b4f0c477f975514b388%2Ftma_v1.png?generation=1704327065491575&alt=media)\n\n\nThe crops are then augmented with stains based on the 25 TMA as references. As the WSI magnification is mainly x20 the generated TMA are also x20. Here are some generated samples:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2F7d4e508fdbfd96b7f04f8356408803d2%2Ftma_v2.png?generation=1704327082748102&alt=media)\n\nA final step was to review them manually to drop bad generated TMA (especially when tumor mask was not available / complete). Our best private LB (0.58) was with such validated TMA. Unfortunately we did not select it as final submission.\n\n**MIL** training procedure and parameters:\n\n\n* CV4, Stratified Group KFold\n* Random batch sampler to balance samples\n* Augmentations: Stain: Vahadane, Macenko, Reinhard\n* Mask on attention, batch size = 32, epochs = 32\n* EMA\n* AdamW optimizer, CosineAnnealingLR, LR=1e-3\n* Cross Entropy Loss\n\n<br/>\n**Transformer/CNN** training procedure and parameters:\n\n\n* ImageNet pretrained backbones ([Timm](https://github.com/huggingface/pytorch-image-models)):\n    * tiny_vit_21m_512.dist_in22k_ft_in1k\n    * tf_efficientnetv2_s_in21ft1k\n* Augmentations: \n    * H/V flips, Rot90\n    * Stain: Vahadane, Macenko, Reinhard\n    * Random BrightnessContrast/Gamma, HueSaturationValue, ColorJitter, CLAHE\n    * GaussianBlur, MotionBlur, GaussNoise\n    * Cut Mix, DropOut\n* EMA\n* Batch size = 32, Epochs = 32\n* AdamW optimizer, CosineAnnealingLR, LR=1e-4\n* Cross Entropy Loss\n\nModels have been trained with full data after checking stability on cross validation.\n\nHere is a 2D t-SNE projection of TinyVit trained embeddings on generated 23k TMA:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Fe0859f7f5adffe00c3e30cb246b40101%2Ftma_training_embeddings.png?generation=1704327121454397&alt=media)\n\n\n\n# Other useful strategies or approaches\n\nValidation was quite difficult, MIL models were overfitting quite fast. Using EMA helped to limit it. \n\n\n# Model inference\n\nMost of the inference time was lost in image loading/tiling. We’ve implemented a multiprocess inference to benefit from all CPUs but optimized to balance the memory issues due to concurrent large images loading. It reduced the loading + tiling of all images to around 5h30-6h. Features exaction was the most time consuming task that is why we’ve limited ourselves to the two best ones. We’ve limited the number of tiles to 350 max and at the end our inference ran in around 11h15-30min.\n\n\n# What did not work or improve?\n\nA quick sum up of what did not work or not improve:\n\nResnet50-based features extractors such as RetCCL and Lunit-BT.\n\nExternal data:\n\n\n\n* [TCGA](https://portal.gdc.cancer.gov/repository?facetTab=cases&filters=%7B%22op%22%3A%22and%22%2C%22content%22%3A%5B%7B%22op%22%3A%22in%22%2C%22content%22%3A%7B%22field%22%3A%22cases.primary_site%22%2C%22value%22%3A%5B%22ovary%22%5D%7D%7D%2C%7B%22op%22%3A%22in%22%2C%22content%22%3A%7B%22field%22%3A%22files.data_type%22%2C%22value%22%3A%5B%22Slide%20Image%22%5D%7D%7D%5D%7D)\n* [ATEC](https://www.cancerimagingarchive.net/collection/ovarian-bevacizumab-response/)\n\n<br/>\nUsually adding more data is always better but here it did not help on both CV and LB. However the quality of many slides was very bad and could explain it. Also the labeling of some was not obvious.\n\nTrain a WSI model based on ImageNet pre-trained backbone. It worked but Ctranspath and LunitDINO outperformed it.\n\nDown scale to x5 for WSI models (instead of x10)\n\nTrain a tumor segmentation model in order to sample tumor TMA, but since subtypes have significant morphological variations, we preferred to train a stroma segmentation model and predict the carcinoma mask by complementarity. Finally it was impossible to set up a unique threshold for TMA selection because of high variation in epithelial surface area between solid and mucinous architecture. \n\nPseudo labeling has not been tried.\n\n\n# Sources\n\n\n* 1) Deep Learning for Detecting BRCA Mutations in High-Grade Ovarian Cancer Based on an Innovative Tumor Segmentation Method From Whole Slide Images: [https://www.modernpathology.org/article/S0893-3952(23)00209-0/fulltext](https://www.modernpathology.org/article/S0893-3952(23)00209-0/fulltext)\n* 2) A Good Feature Extractor Is All You Need for Weakly Supervised Learning in Histopathology: [https://arxiv.org/pdf/2311.11772.pdf](https://arxiv.org/pdf/2311.11772.pdf)\n\n<br/>\nWe had a lot of fun solving this kaggle. It was a lot of data to handle, in addition to the ML challenge it was an optimization challenge to make the inference fast.\n\n<br/>\nRaphaël Bourgade and MPWARE",
    "2586517": "The generation of TMA is a great idea; I would be curious if you observed any significant improvement with these rounded patches compared to using the easy squared tiles?",
    "2586223": "Congratulations on getting to the 15th position in this competition. Thanks for sharing the details of your model with colorful diagrams. ",
    "2586119": "That’s a very solid solution. How much time did you spend (approximately) to implement tma generation part? How did you come up with the staining techniques? That is awesome.",
    "2586082": "What is your only TMA private score? My only private TMA is 0.33. My only private WSI is low 🥲",
    "2586884": "Hi @mpware , I think generation TMA on WSI is a good idea. But I think use the stain of the TMA in training set is very dangerous in this competition. Since the test set is get from over 20 medical centers, which is very diverse, while the training set only has 25 TMAs. So if you use the stain on the TMA in training set, I think you suffer a risk of overfitting to the TMAs in training set. "
  }
}