{
  "id": 466455,
  "title": "1st Place Solution 🥇 [Owkin]",
  "url": "/competitions/UBC-OCEAN/discussion/466455",
  "author_name": "jibounet",
  "post_date": "2024-01-08T17:22:46.265000",
  "votes": 65,
  "comment_count": 37,
  "views": 0,
  "content": "<h1>1st Place Solution 🥇 [Owkin] -- Phikon &amp; Chowder</h1>\n<h2>Introduction</h2>\n<p>First of all, we would like to thank the University of British Columbia (UBC) for this exceptional multi-centric cohort and the Kaggle staff for organizing this competition. We got into this competition to showcase the efficiency and robustness of Phikon (<a href=\"https://huggingface.co/owkin/phikon\" target=\"_blank\">model card</a>, <a href=\"https://www.medrxiv.org/content/10.1101/2023.07.21.23292757v1\" target=\"_blank\">paper</a>, <a href=\"https://huggingface.co/blog/EazyAl/phikon\" target=\"_blank\">blog post</a>), the foundation model (FM) for digital pathology made available to the community by <a href=\"https://www.owkin.com/\" target=\"_blank\">Owkin</a> last November. We are very pleased with the outcome and really enjoyed participating in this competition.</p>\n<p>Our solution is straightforward: we trained an ensemble of <a href=\"https://arxiv.org/pdf/1802.02212.pdf\" target=\"_blank\">Chowder</a> models on top of Phikon tile embeddings. We used high entropy predictions to detect outliers. We did not use extra training data nor annotations (other than the ones provided by the organizers). Our <em>winning submission</em> submission scored 0.64/<strong>0.66</strong> (public/private) and our <em>top submission</em> scored 0.62/<strong>0.68</strong>. These submissions run in approximately 6 hours 🚀</p>\n<p>Our code is available <a href=\"https://www.kaggle.com/code/jbschiratti/winning-submission\" target=\"_blank\">here</a>. We cleaned our code (removed comments, unused code, added sections…) and created the <code>winning_submission</code> notebook. After a late submission, this notebook scores 0.63/<strong>0.66</strong>. The slight difference on the public LB is likely due to differences in the sampling of patches.</p>\n<p>Our solution write-up is structured as follows:</p>\n<ol>\n<li><a href=\"#1-our-main-takeaways\">Main takeaways</a></li>\n<li><a href=\"#2-matter-detection-and-tiling\">Matter detection and tiling</a></li>\n<li><a href=\"#3-feature-extraction\">Feature extraction</a></li>\n<li><a href=\"#4-subtypes-classification\">Subtypes classification</a></li>\n<li><a href=\"#5-outlier-detection\">Outlier detection</a></li>\n<li><a href=\"#6-some-(un)successful-ideas\">Some (un)successful ideas</a></li>\n<li><a href=\"#7-notes\">Notes</a></li>\n</ol>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F705762%2Fe5d2f490c74b954bf94ce12871659276%2Fovr_pipeline.png?generation=1704733650042787&amp;alt=media\" alt=\"Our pipeline\"></p>\n<h2>1. Our main takeaways</h2>\n<ul>\n<li><p>Foundation models (or domain-specific large vision models) are the way of the future. Our results further validate the effectiveness of <a href=\"https://huggingface.co/owkin/phikon\" target=\"_blank\">Phikon</a>, Owkin's foundation model for digital pathology. The next frontier is multimodality: combining spatial-omics, imaging, clinical data - from genotype to phenotype - and blending it with medical knowledge and reasoning powered by Large Language Models (LLMs).</p></li>\n<li><p>Occam's razor: our simple and efficient pipeline outperformed more complex approaches. In particular, <a href=\"https://arxiv.org/pdf/1802.02212.pdf\" target=\"_blank\">Chowder</a> - a Multiple Instance Learning (MIL) model - is still on par with more recent MIL models (e.g. <a href=\"https://arxiv.org/abs/2106.00908\" target=\"_blank\">TransMIL</a>, <a href=\"https://arxiv.org/abs/2203.12081\" target=\"_blank\">DTFD-MIL</a>), especially when combined with ensembling techniques. As opposed to more elaborate MIL models, Chowder is also interpretable.</p></li>\n<li><p>This competition was not easy! The PNG image format posed a major challenge: loading large PNG images with standard libraries (<a href=\"https://pillow.readthedocs.io/en/stable/\" target=\"_blank\">Pillow</a>, <a href=\"https://opencv.org/\" target=\"_blank\">OpenCV</a>) was time consuming and used a lot of RAM. Standard formats in digital pathology (SVS, TIFF, NDPI) store data pyramidally to prevent the need for loading the entire images into RAM. Although Kaggle staff members <a href=\"https://www.kaggle.com/competitions/UBC-OCEAN/discussion/446688\" target=\"_blank\">acknowledged that PNG format was a bad choice</a>, pyramidal images were not made available to the participants. Furthermore, either by design or as a result of converting pyramidal images to PNG, useful metadata such as mpp (image resolution in microns per pixels) and ICC profile (if available) were stripped from the images. In addition to this, working with images at different resolutions and dealing with outliers (rare variants and normal cases) in the test set made this competition challenging (and quite interesting!).</p></li>\n<li><p>How well do our models generalize? Locally, in cross-validation (CV), the balanced accuracy scores were in the (0.8, 0.9) range. However, the scores on the public/private LB were at least 20 points lower. Obviously, we can (partly) explain this discrepancy by our ability to predict the 'Other' class. However, it also raises the question of how well our models generalize to new data (<em>i.e.</em> data points from new centers/hospitals). Even when using <a href=\"https://huggingface.co/owkin/phikon\" target=\"_blank\">Phikon</a>, it is likely that differences in tissue preparation, tissue staining or differences in scanner type and magnification across centers still hinder generalization.</p></li>\n</ul>\n<h2>2. Matter detection and tiling</h2>\n<p>Whole Slide Images (WSI) in digital pathology are often too large and cannot be directly analyzed using convolutional neural networks. The WSI in this competition were no exception. A well-established workaround consists in splitting regions containing tissue into smaller patches (e.g. 224 x 224 px or 512 x 512 px). As a result, a WSI can be seen as a collection of hundreds, thousands of patches. Although Tissue Microarrays (TMA) were much smaller, these images were also split into patches.</p>\n<h3>2.1. Matter detection</h3>\n<p>In order to detect the regions of the WSI (or TMA) which contain tissue, we employed Otsu thresholding. This thresholding was applied to the thumbnail image, in the HSV color space. Although this method is not perfect, we found that it worked quite well on the images from this competition.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F705762%2Fd3dbaa1da15d2cbbfa5b7ce32381c3ad%2F1020_matter_detection.png?generation=1704733858157340&amp;alt=media\" alt=\"\"></p>\n<h3>2.2. Tiling</h3>\n<h4>Patch size</h4>\n<p>We set the patch size to 224 x 224 px for WSI. According to the <a href=\"https://www.kaggle.com/competitions/UBC-OCEAN/data\" target=\"_blank\">data description</a>, the train set is composed of a majority of WSI at magnification 20x and few (25) TMA at magnification 40x. We hypothesized that given the low number of TMA in the train set, learning would be more efficient if we standardized all images (WSI or TMA) to a 20x resolution. Therefore, we set the patch size to 448 x 448 px for TMA; These patches were then resized to 224 x 224 px. As illustrated below [left: a 448 x 448 px patch from 91.png (TMA) resized to 224 x 224 px; Right: a 224 x 224 px patch from 4211.png (WSI)], cells have roughly the same size in resized tiles from TMA and tiles from WSI. At test time, we used a logistic regression (LR) to detect if an image is a TMA. The LR was trained on features extracted from the thumbnail of train images using a pretrained (ImageNet) ResNet18.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F705762%2F88d1e1c133cfaac9dee1cce00fb977a4%2Fsample_tiles.png?generation=1704733902739438&amp;alt=media\" alt=\"\"></p>\n<h4>Runtime 🕕</h4>\n<ul>\n<li><p>For efficient tiling, we developed custom C code which splits a large PNG image into patches and saves them to disk (also as PNG images). This C code leverages the <code>libpng</code> library. In particular, it uses the <code>png_read_row</code> function to limit the amount of data read and loaded into RAM. This C code was easily compiled in a Kaggle notebook. The performance of the compiled code is likely to be similar to pyvips.</p></li>\n<li><p>With a patch size of 224 x 224 px, WSI have - on average - more than 10 000 patches. Processing that many patches was not feasible given the 12 hours runtime limit. With a limit of <strong>200 patches</strong> per image, our first submissions successfully ran in approximately 10 hours. Later, we used <a href=\"https://www.ray.io/\" target=\"_blank\">ray</a> to process images in pairs: matter detection + tiling + feature extraction for two images at a time on a single P100 GPU. Ray would spawn two processes, each using 0.5 GPU and 2 CPU cores. As a result, our submissions successfully ran in less than 7 hours. The limiting resource was the RAM; With more RAM, we could have processed 4 images at a time (on a single P100). Note that the 200 patches limit only applies to WSI. Given that TMA are small images, the number of 448 x 448 px patches hardly ever exceeded 50. Most TMA had less than 30 patches.</p></li>\n</ul>\n<h2>3. Feature extraction</h2>\n<p><a href=\"https://huggingface.co/owkin/phikon\" target=\"_blank\">Phikon</a> is a ViT-Base pre-trained with <a href=\"https://github.com/bytedance/ibot\" target=\"_blank\">iBOT</a> on 40M tiles from the TCGA dataset (📝 see our <a href=\"https://www.medrxiv.org/content/10.1101/2023.07.21.23292757v1\" target=\"_blank\">paper</a> for more detailed info). We benchmarked Phikon against multiple backbones including <a href=\"https://github.com/Xiyue-Wang/TransPath\" target=\"_blank\">CTransPath</a>, <a href=\"https://github.com/lunit-io/benchmark-ssl-pathology#pre-trained-weights\" target=\"_blank\">LUNIT</a> and <a href=\"https://github.com/facebookresearch/dinov2\" target=\"_blank\">DinoV2</a>. Phikon outperformed these models in our local cross-validation tests and showed no improvement when combined in an ensemble. For each input patch, a <code>(3, 224, 224)</code> tensor, Phikon outputs a 768-dimensional embedding vector. Therefore, a WSI or a TMA is represented as a 2D tensor with shape <code>(n_patches, 768)</code>.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F705762%2F13ea1a1f74c768e3f1c5075e25ea282b%2Fbackbones_scores.png?generation=1704734075166073&amp;alt=media\" alt=\"\"></p>\n<h2>4. Subtypes classification</h2>\n<p>In order to predict the cancer subtypes from the extracted features, we considered several Multiple Instance Learning (MIL) models from our public GitHub repository <a href=\"https://github.com/owkin/HistoSSLscaling\" target=\"_blank\">HistoSSLScaling</a>: <a href=\"https://arxiv.org/pdf/1802.02212.pdf\" target=\"_blank\">Chowder</a>, <a href=\"https://arxiv.org/abs/1802.04712\" target=\"_blank\">DeepMIL</a>, <a href=\"https://arxiv.org/abs/2011.08939\" target=\"_blank\">DSMIL</a>, <a href=\"https://arxiv.org/abs/2106.00908\" target=\"_blank\">TransMIL</a> and MeanPool. Chowder outperformed others including DeepMIL, MeanPool, and DSMIL.</p>\n<p>Before the competition deadline, we selected two submissions. One submission scored <strong>0.64</strong>/0.64 and consisted of an ensemble of 5 Chowder, 5 DeepMIL, 5 DSMIL and 5 MeanPool trained upon Phikon, CTranspath and LUNIT. We will detail the <em>winning submission</em> which only used Chowder and Phikon.</p>\n<h3>4.1. Chowder's architecture</h3>\n<p>The input dimension of Chowder was set to 768 (<em>i.e.</em> the dimension of Phikon embeddings) and its output dimension to 5. The first layer of Chowder is a <code>TilesMLP</code> layer with a hidden dimension of 192. The <code>n_top</code> and <code>n_bottom</code> values of its <code>ExtremeLayer</code> layer were both set to 10. The last layer of Chowder is a <code>MLP</code> with hidden dimension 96, a dropout rate of 30% and Sigmoid activation. We refer the reader to <a href=\"https://github.com/owkin/HistoSSLscaling/blob/main/rl_benchmarks/models/slide_models/chowder.py\" target=\"_blank\">this Python file</a> for an implementation of Chowder. These hyperparameters were manually selected (no hyperparameter tuning).</p>\n<h3>4.2. Cross-validation</h3>\n<p>Locally, we used stratified 5-fold cross-validation (CV) to estimate the predictive performance of our model. For each fold, four-fifths of the data were used to create a train-validation split (75%-25%) and the remaining fifth of the data was used as a test set. This stratified 5-fold CV was repeated 3 times (with different seeds changing how data was shuffled). We designed the stratified 5-fold CV to ensure that each test set would contain at least 60% of TMA. Hence, most of the TMA were used for evaluation and not for training.</p>\n<h3>4.3. Training</h3>\n<p>Our models were trained for a maximum of 30 epochs. The validation set was used for early stopping (using the validation balanced accuracy as stopping criterion) with a patience parameter of 4 epochs. In addition to this, we used the <a href=\"https://pytorch.org/docs/stable/generated/torch.optim.AdamW.html\" target=\"_blank\">AdamW</a> optimizer with a constant learning rate of 0.0001 and weight decay of 0.001. The loss was the <a href=\"https://pytorch.org/docs/stable/generated/torch.nn.CrossEntropyLoss.html\" target=\"_blank\">Cross-Entropy (CE) loss</a>.</p>\n<p>Two strategies were implemented to mitigate class imbalance:</p>\n<ol>\n<li><a href=\"https://pytorch.org/docs/stable/data.html#torch.utils.data.WeightedRandomSampler\" target=\"_blank\">Weighted sampling</a> to create balanced batches for training,</li>\n<li>Using class weights in <a href=\"https://pytorch.org/docs/stable/generated/torch.nn.CrossEntropyLoss.html#torch.nn.CrossEntropyLoss\" target=\"_blank\">CE loss</a>.</li>\n</ol>\n<h3>4.4. The ensembling trick</h3>\n<p>Chowder can be quite sensitive to weight initialization. Instead of training a single Chowder model, we decided to train an ensemble of N=50 Chowder models. The Chowder models in the ensemble only differ by their initialization. We found the ensemble to be more stable during training - and more efficient - than a single Chowder model. The Python code below is copied from our winning submission:</p>\n<pre><code> (nn.ModuleList):\n\n     () -&gt; :\n        ().__init__(modules=models)\n\n     () -&gt; torch.Tensor:\n        \n        predictions, scores = [], []\n         model  self:\n            logits_, scores_ = model(x, mask)\n            predictions.append(logits_.unsqueeze(-))\n            scores.append(torch.mean(scores_, dim=, keepdim=).unsqueeze(-))\n        predictions = torch.cat(predictions, dim=)\n        scores = torch.cat(scores, dim=)\n         predictions, scores\n\n\nchowder_models = [Chowder(**chowder_kwargs)  _  ()]\n\nmodel = ModelEnsemble(chowder_models)\n</code></pre>\n<h3>4.5. Submissions</h3>\n<p>Three repetitions of stratified 5-fold CV with an ensemble of 50 Chowder lead to a great number of Chowder models! Through several submissions, we noticed that it was more efficient to select specific repetitions and folds rather than ensembling all the 3 x 5 x 50 models. Our winning submission is the <strong>average prediction of 65 Chowder models</strong> trained on different data splits. We calibrated these models using a logistic regression on their internal validation set (CV), which appeared to yield a slight improvement on the public LB. After calibration, we added a “model filtering” step: we selected only a subset of the 50 Chowder models in an ensemble based on the performance of the calibrated models on their internal test set (CV).</p>\n<h2>5. Outlier detection</h2>\n<p>With balanced accuracy as the metric, correctly predicting the 'Other' category, worth 16.6 points (100/6), was key! Our strategy to identify outliers raised our public leaderboard score from 0.59 to 0.64, highlighting its importance as the most challenging class.</p>\n<p>We found that using a threshold on the entropy of predictions, calculated as H = -sum(p*log(p)), was most effective for us, with high entropy indicating uncertainty in predictions. As no outliers were provided, we calibrated this threshold based on the public leaderboard.</p>\n<h2>6. Some (un)successful ideas</h2>\n<ul>\n<li><p>To account for inter-center variability, we tried several color normalization schemes (Vahadane, Reinhard) but did not observe any improvement doing so. This finding is aligned with <a href=\"https://www.nature.com/articles/s41598-023-46619-6\" target=\"_blank\">recent publications</a> suggesting that staining normalization does not improve the performance of models for histopathological classification tasks.</p></li>\n<li><p>Increasing the number of patches for TMA using a sliding window (with 30% to 80% overlap between two consecutive patches). This method significantly increased the runtime of our submissions (most failing with <code>Notebook Timeout</code>) and did not provide any performance improvement.</p></li>\n<li><p>We explored the idea of identifying normal cases using a tumor detection model trained on the provided annotations. We binarized the annotations (tumor=1, stroma/necrosis=0) and trained a logistic regression on patch features to identify patches containing tumor. The percentage of tumor patches in a WSI/TMA would be used to identify normal cases. This idea did not provide any performance improvement on the public LB.</p></li>\n<li><p>We used the <a href=\"https://docs.ray.io/en/latest/tune/index.html\" target=\"_blank\">Ray Tune</a> library to do hyperparameter tuning with <a href=\"https://arxiv.org/pdf/1802.02212.pdf\" target=\"_blank\">Chowder</a> and <a href=\"https://arxiv.org/abs/1802.04712\" target=\"_blank\">DeepMIL</a>. The hyperparameters we optimized for were: batch size, number of training epochs, learning rate, dimensions of hidden layers in MLP, activation functions. The hyperparameter tuning resulted in an increase of our local CV scores, but in a decrease of our submissions scores on the public LB. We hypothesized that hyperparameter tuning led to overfitting the train set and dropped the idea.</p></li>\n<li><p>[Successful idea 💡] Fine-tuning Phikon. Here, we’re not talking about fine-tuning in a conventional way; Instead, we mean pretraining a ViT-Base, initialized with Phikon’s weights, using <a href=\"https://github.com/bytedance/ibot\" target=\"_blank\">iBOT</a> on patches from the images in the train set. To do so, we extracted a total of 6.5M patches (224 x 224 px RGB) from the images in the train set. Following the recent paper from <a href=\"https://arxiv.org/abs/2309.16588\" target=\"_blank\">Darcet et al. 2023</a> and the work of <a href=\"https://github.com/facebookresearch/dinov2\" target=\"_blank\">Dino V2</a>, we added 4 register tokens to the ViT-Base. This ViT was trained for a single epoch with an initial learning rate of 0.0005 and batch size (per device) of 32. A single epoch took 2.5 days on 2 NVIDIA P100 GPUs. In a submission, we combined Chowder models trained on features extracted with the <em>original Phikon</em> and Chowder models trained on features extracted with the <em>fine-tuned Phikon</em> (with register tokens). This submission scored 0.62/<strong>0.67</strong> (not selected for the final evaluation).</p></li>\n<li><p>[Successful idea 💡] Using the variance of predictions (across models in an ensemble of Chowder) to identify outliers. A submission implementing this idea (along with the entropy of predictions) scored 0.59/<strong>0.67</strong> (not selected for the final evaluation).</p></li>\n</ul>\n<h2>7. Notes</h2>\n<p>The magnification of the images proved to be a key information. We found that resizing the TMA patches to match those from WSI allowed us to have consistent performance over all images in the test set. Keeping the original pyramidal image would have saved the participants from cumbersome image processing and could have led to overall better performances. </p>\n<p>For each image (WSI or TMA), we converted the extracted patches to grayscale and applied contrast equalization; Phikon features were then averaged across all available patches. We applied UMAP dimension reduction to these averaged representations and noticed that 38 WSI could be told apart from the remaining 500 images. . Red cells can be used as a scale to compare the resolution levels. Comparing red-cells’ size in slide 431 (among the main group of slides) and slide 4 (among the 38 odd slides) shows there is a substantial difference in resolution between the two groups, with an estimated x12 magnification for the odd slides, instead of the normal x20 magnification. This could be the result of a failed conversion from the original .svs (or .tiff) format of the slide to .png, resulting from the selection of the wrong zoom level. </p>\n<p>Such resolution variances are significant: 8% of all slides exhibited incorrect zoom levels impacting model performance. Notably, 20% (9 out of 47) of all LGSC slides had an incorrect zoom level. Finally, this UMAP also shows that after downsampling the TMA, their features are mixed with those of the WSI, this confirms our ability to use a single approach for these two data types.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F705762%2Fac94b6191f21b8c8184f40f03800a74f%2Fumap.png?generation=1704734453560448&amp;alt=media\" alt=\"\"></p>\n<p>The ids of the odd looking slides are the following:    [4,970,1080,2097,3222,3511,3881,9509,12159,12244,13364,13387,15583,15871,25604,26124,29888,31300,31793,32035,32192,33839,34688,34720,40079,40639,41099,44432,44700,49995,51215,52308,52784,53402,61100,63298,63836,64629]</p>\n<p>jbschiratti, on behalf of Owkin's team</p>",
  "messages": [
    {
      "id": 2592618,
      "postDate": "2024-01-08T17:22:46.267Z",
      "content": "<h1>1st Place Solution 🥇 [Owkin] -- Phikon &amp; Chowder</h1>\n<h2>Introduction</h2>\n<p>First of all, we would like to thank the University of British Columbia (UBC) for this exceptional multi-centric cohort and the Kaggle staff for organizing this competition. We got into this competition to showcase the efficiency and robustness of Phikon (<a href=\"https://huggingface.co/owkin/phikon\" target=\"_blank\">model card</a>, <a href=\"https://www.medrxiv.org/content/10.1101/2023.07.21.23292757v1\" target=\"_blank\">paper</a>, <a href=\"https://huggingface.co/blog/EazyAl/phikon\" target=\"_blank\">blog post</a>), the foundation model (FM) for digital pathology made available to the community by <a href=\"https://www.owkin.com/\" target=\"_blank\">Owkin</a> last November. We are very pleased with the outcome and really enjoyed participating in this competition.</p>\n<p>Our solution is straightforward: we trained an ensemble of <a href=\"https://arxiv.org/pdf/1802.02212.pdf\" target=\"_blank\">Chowder</a> models on top of Phikon tile embeddings. We used high entropy predictions to detect outliers. We did not use extra training data nor annotations (other than the ones provided by the organizers). Our <em>winning submission</em> submission scored 0.64/<strong>0.66</strong> (public/private) and our <em>top submission</em> scored 0.62/<strong>0.68</strong>. These submissions run in approximately 6 hours 🚀</p>\n<p>Our code is available <a href=\"https://www.kaggle.com/code/jbschiratti/winning-submission\" target=\"_blank\">here</a>. We cleaned our code (removed comments, unused code, added sections…) and created the <code>winning_submission</code> notebook. After a late submission, this notebook scores 0.63/<strong>0.66</strong>. The slight difference on the public LB is likely due to differences in the sampling of patches.</p>\n<p>Our solution write-up is structured as follows:</p>\n<ol>\n<li><a href=\"#1-our-main-takeaways\">Main takeaways</a></li>\n<li><a href=\"#2-matter-detection-and-tiling\">Matter detection and tiling</a></li>\n<li><a href=\"#3-feature-extraction\">Feature extraction</a></li>\n<li><a href=\"#4-subtypes-classification\">Subtypes classification</a></li>\n<li><a href=\"#5-outlier-detection\">Outlier detection</a></li>\n<li><a href=\"#6-some-(un)successful-ideas\">Some (un)successful ideas</a></li>\n<li><a href=\"#7-notes\">Notes</a></li>\n</ol>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F705762%2Fe5d2f490c74b954bf94ce12871659276%2Fovr_pipeline.png?generation=1704733650042787&amp;alt=media\" alt=\"Our pipeline\"></p>\n<h2>1. Our main takeaways</h2>\n<ul>\n<li><p>Foundation models (or domain-specific large vision models) are the way of the future. Our results further validate the effectiveness of <a href=\"https://huggingface.co/owkin/phikon\" target=\"_blank\">Phikon</a>, Owkin's foundation model for digital pathology. The next frontier is multimodality: combining spatial-omics, imaging, clinical data - from genotype to phenotype - and blending it with medical knowledge and reasoning powered by Large Language Models (LLMs).</p></li>\n<li><p>Occam's razor: our simple and efficient pipeline outperformed more complex approaches. In particular, <a href=\"https://arxiv.org/pdf/1802.02212.pdf\" target=\"_blank\">Chowder</a> - a Multiple Instance Learning (MIL) model - is still on par with more recent MIL models (e.g. <a href=\"https://arxiv.org/abs/2106.00908\" target=\"_blank\">TransMIL</a>, <a href=\"https://arxiv.org/abs/2203.12081\" target=\"_blank\">DTFD-MIL</a>), especially when combined with ensembling techniques. As opposed to more elaborate MIL models, Chowder is also interpretable.</p></li>\n<li><p>This competition was not easy! The PNG image format posed a major challenge: loading large PNG images with standard libraries (<a href=\"https://pillow.readthedocs.io/en/stable/\" target=\"_blank\">Pillow</a>, <a href=\"https://opencv.org/\" target=\"_blank\">OpenCV</a>) was time consuming and used a lot of RAM. Standard formats in digital pathology (SVS, TIFF, NDPI) store data pyramidally to prevent the need for loading the entire images into RAM. Although Kaggle staff members <a href=\"https://www.kaggle.com/competitions/UBC-OCEAN/discussion/446688\" target=\"_blank\">acknowledged that PNG format was a bad choice</a>, pyramidal images were not made available to the participants. Furthermore, either by design or as a result of converting pyramidal images to PNG, useful metadata such as mpp (image resolution in microns per pixels) and ICC profile (if available) were stripped from the images. In addition to this, working with images at different resolutions and dealing with outliers (rare variants and normal cases) in the test set made this competition challenging (and quite interesting!).</p></li>\n<li><p>How well do our models generalize? Locally, in cross-validation (CV), the balanced accuracy scores were in the (0.8, 0.9) range. However, the scores on the public/private LB were at least 20 points lower. Obviously, we can (partly) explain this discrepancy by our ability to predict the 'Other' class. However, it also raises the question of how well our models generalize to new data (<em>i.e.</em> data points from new centers/hospitals). Even when using <a href=\"https://huggingface.co/owkin/phikon\" target=\"_blank\">Phikon</a>, it is likely that differences in tissue preparation, tissue staining or differences in scanner type and magnification across centers still hinder generalization.</p></li>\n</ul>\n<h2>2. Matter detection and tiling</h2>\n<p>Whole Slide Images (WSI) in digital pathology are often too large and cannot be directly analyzed using convolutional neural networks. The WSI in this competition were no exception. A well-established workaround consists in splitting regions containing tissue into smaller patches (e.g. 224 x 224 px or 512 x 512 px). As a result, a WSI can be seen as a collection of hundreds, thousands of patches. Although Tissue Microarrays (TMA) were much smaller, these images were also split into patches.</p>\n<h3>2.1. Matter detection</h3>\n<p>In order to detect the regions of the WSI (or TMA) which contain tissue, we employed Otsu thresholding. This thresholding was applied to the thumbnail image, in the HSV color space. Although this method is not perfect, we found that it worked quite well on the images from this competition.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F705762%2Fd3dbaa1da15d2cbbfa5b7ce32381c3ad%2F1020_matter_detection.png?generation=1704733858157340&amp;alt=media\" alt=\"\"></p>\n<h3>2.2. Tiling</h3>\n<h4>Patch size</h4>\n<p>We set the patch size to 224 x 224 px for WSI. According to the <a href=\"https://www.kaggle.com/competitions/UBC-OCEAN/data\" target=\"_blank\">data description</a>, the train set is composed of a majority of WSI at magnification 20x and few (25) TMA at magnification 40x. We hypothesized that given the low number of TMA in the train set, learning would be more efficient if we standardized all images (WSI or TMA) to a 20x resolution. Therefore, we set the patch size to 448 x 448 px for TMA; These patches were then resized to 224 x 224 px. As illustrated below [left: a 448 x 448 px patch from 91.png (TMA) resized to 224 x 224 px; Right: a 224 x 224 px patch from 4211.png (WSI)], cells have roughly the same size in resized tiles from TMA and tiles from WSI. At test time, we used a logistic regression (LR) to detect if an image is a TMA. The LR was trained on features extracted from the thumbnail of train images using a pretrained (ImageNet) ResNet18.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F705762%2F88d1e1c133cfaac9dee1cce00fb977a4%2Fsample_tiles.png?generation=1704733902739438&amp;alt=media\" alt=\"\"></p>\n<h4>Runtime 🕕</h4>\n<ul>\n<li><p>For efficient tiling, we developed custom C code which splits a large PNG image into patches and saves them to disk (also as PNG images). This C code leverages the <code>libpng</code> library. In particular, it uses the <code>png_read_row</code> function to limit the amount of data read and loaded into RAM. This C code was easily compiled in a Kaggle notebook. The performance of the compiled code is likely to be similar to pyvips.</p></li>\n<li><p>With a patch size of 224 x 224 px, WSI have - on average - more than 10 000 patches. Processing that many patches was not feasible given the 12 hours runtime limit. With a limit of <strong>200 patches</strong> per image, our first submissions successfully ran in approximately 10 hours. Later, we used <a href=\"https://www.ray.io/\" target=\"_blank\">ray</a> to process images in pairs: matter detection + tiling + feature extraction for two images at a time on a single P100 GPU. Ray would spawn two processes, each using 0.5 GPU and 2 CPU cores. As a result, our submissions successfully ran in less than 7 hours. The limiting resource was the RAM; With more RAM, we could have processed 4 images at a time (on a single P100). Note that the 200 patches limit only applies to WSI. Given that TMA are small images, the number of 448 x 448 px patches hardly ever exceeded 50. Most TMA had less than 30 patches.</p></li>\n</ul>\n<h2>3. Feature extraction</h2>\n<p><a href=\"https://huggingface.co/owkin/phikon\" target=\"_blank\">Phikon</a> is a ViT-Base pre-trained with <a href=\"https://github.com/bytedance/ibot\" target=\"_blank\">iBOT</a> on 40M tiles from the TCGA dataset (📝 see our <a href=\"https://www.medrxiv.org/content/10.1101/2023.07.21.23292757v1\" target=\"_blank\">paper</a> for more detailed info). We benchmarked Phikon against multiple backbones including <a href=\"https://github.com/Xiyue-Wang/TransPath\" target=\"_blank\">CTransPath</a>, <a href=\"https://github.com/lunit-io/benchmark-ssl-pathology#pre-trained-weights\" target=\"_blank\">LUNIT</a> and <a href=\"https://github.com/facebookresearch/dinov2\" target=\"_blank\">DinoV2</a>. Phikon outperformed these models in our local cross-validation tests and showed no improvement when combined in an ensemble. For each input patch, a <code>(3, 224, 224)</code> tensor, Phikon outputs a 768-dimensional embedding vector. Therefore, a WSI or a TMA is represented as a 2D tensor with shape <code>(n_patches, 768)</code>.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F705762%2F13ea1a1f74c768e3f1c5075e25ea282b%2Fbackbones_scores.png?generation=1704734075166073&amp;alt=media\" alt=\"\"></p>\n<h2>4. Subtypes classification</h2>\n<p>In order to predict the cancer subtypes from the extracted features, we considered several Multiple Instance Learning (MIL) models from our public GitHub repository <a href=\"https://github.com/owkin/HistoSSLscaling\" target=\"_blank\">HistoSSLScaling</a>: <a href=\"https://arxiv.org/pdf/1802.02212.pdf\" target=\"_blank\">Chowder</a>, <a href=\"https://arxiv.org/abs/1802.04712\" target=\"_blank\">DeepMIL</a>, <a href=\"https://arxiv.org/abs/2011.08939\" target=\"_blank\">DSMIL</a>, <a href=\"https://arxiv.org/abs/2106.00908\" target=\"_blank\">TransMIL</a> and MeanPool. Chowder outperformed others including DeepMIL, MeanPool, and DSMIL.</p>\n<p>Before the competition deadline, we selected two submissions. One submission scored <strong>0.64</strong>/0.64 and consisted of an ensemble of 5 Chowder, 5 DeepMIL, 5 DSMIL and 5 MeanPool trained upon Phikon, CTranspath and LUNIT. We will detail the <em>winning submission</em> which only used Chowder and Phikon.</p>\n<h3>4.1. Chowder's architecture</h3>\n<p>The input dimension of Chowder was set to 768 (<em>i.e.</em> the dimension of Phikon embeddings) and its output dimension to 5. The first layer of Chowder is a <code>TilesMLP</code> layer with a hidden dimension of 192. The <code>n_top</code> and <code>n_bottom</code> values of its <code>ExtremeLayer</code> layer were both set to 10. The last layer of Chowder is a <code>MLP</code> with hidden dimension 96, a dropout rate of 30% and Sigmoid activation. We refer the reader to <a href=\"https://github.com/owkin/HistoSSLscaling/blob/main/rl_benchmarks/models/slide_models/chowder.py\" target=\"_blank\">this Python file</a> for an implementation of Chowder. These hyperparameters were manually selected (no hyperparameter tuning).</p>\n<h3>4.2. Cross-validation</h3>\n<p>Locally, we used stratified 5-fold cross-validation (CV) to estimate the predictive performance of our model. For each fold, four-fifths of the data were used to create a train-validation split (75%-25%) and the remaining fifth of the data was used as a test set. This stratified 5-fold CV was repeated 3 times (with different seeds changing how data was shuffled). We designed the stratified 5-fold CV to ensure that each test set would contain at least 60% of TMA. Hence, most of the TMA were used for evaluation and not for training.</p>\n<h3>4.3. Training</h3>\n<p>Our models were trained for a maximum of 30 epochs. The validation set was used for early stopping (using the validation balanced accuracy as stopping criterion) with a patience parameter of 4 epochs. In addition to this, we used the <a href=\"https://pytorch.org/docs/stable/generated/torch.optim.AdamW.html\" target=\"_blank\">AdamW</a> optimizer with a constant learning rate of 0.0001 and weight decay of 0.001. The loss was the <a href=\"https://pytorch.org/docs/stable/generated/torch.nn.CrossEntropyLoss.html\" target=\"_blank\">Cross-Entropy (CE) loss</a>.</p>\n<p>Two strategies were implemented to mitigate class imbalance:</p>\n<ol>\n<li><a href=\"https://pytorch.org/docs/stable/data.html#torch.utils.data.WeightedRandomSampler\" target=\"_blank\">Weighted sampling</a> to create balanced batches for training,</li>\n<li>Using class weights in <a href=\"https://pytorch.org/docs/stable/generated/torch.nn.CrossEntropyLoss.html#torch.nn.CrossEntropyLoss\" target=\"_blank\">CE loss</a>.</li>\n</ol>\n<h3>4.4. The ensembling trick</h3>\n<p>Chowder can be quite sensitive to weight initialization. Instead of training a single Chowder model, we decided to train an ensemble of N=50 Chowder models. The Chowder models in the ensemble only differ by their initialization. We found the ensemble to be more stable during training - and more efficient - than a single Chowder model. The Python code below is copied from our winning submission:</p>\n<pre><code> (nn.ModuleList):\n\n     () -&gt; :\n        ().__init__(modules=models)\n\n     () -&gt; torch.Tensor:\n        \n        predictions, scores = [], []\n         model  self:\n            logits_, scores_ = model(x, mask)\n            predictions.append(logits_.unsqueeze(-))\n            scores.append(torch.mean(scores_, dim=, keepdim=).unsqueeze(-))\n        predictions = torch.cat(predictions, dim=)\n        scores = torch.cat(scores, dim=)\n         predictions, scores\n\n\nchowder_models = [Chowder(**chowder_kwargs)  _  ()]\n\nmodel = ModelEnsemble(chowder_models)\n</code></pre>\n<h3>4.5. Submissions</h3>\n<p>Three repetitions of stratified 5-fold CV with an ensemble of 50 Chowder lead to a great number of Chowder models! Through several submissions, we noticed that it was more efficient to select specific repetitions and folds rather than ensembling all the 3 x 5 x 50 models. Our winning submission is the <strong>average prediction of 65 Chowder models</strong> trained on different data splits. We calibrated these models using a logistic regression on their internal validation set (CV), which appeared to yield a slight improvement on the public LB. After calibration, we added a “model filtering” step: we selected only a subset of the 50 Chowder models in an ensemble based on the performance of the calibrated models on their internal test set (CV).</p>\n<h2>5. Outlier detection</h2>\n<p>With balanced accuracy as the metric, correctly predicting the 'Other' category, worth 16.6 points (100/6), was key! Our strategy to identify outliers raised our public leaderboard score from 0.59 to 0.64, highlighting its importance as the most challenging class.</p>\n<p>We found that using a threshold on the entropy of predictions, calculated as H = -sum(p*log(p)), was most effective for us, with high entropy indicating uncertainty in predictions. As no outliers were provided, we calibrated this threshold based on the public leaderboard.</p>\n<h2>6. Some (un)successful ideas</h2>\n<ul>\n<li><p>To account for inter-center variability, we tried several color normalization schemes (Vahadane, Reinhard) but did not observe any improvement doing so. This finding is aligned with <a href=\"https://www.nature.com/articles/s41598-023-46619-6\" target=\"_blank\">recent publications</a> suggesting that staining normalization does not improve the performance of models for histopathological classification tasks.</p></li>\n<li><p>Increasing the number of patches for TMA using a sliding window (with 30% to 80% overlap between two consecutive patches). This method significantly increased the runtime of our submissions (most failing with <code>Notebook Timeout</code>) and did not provide any performance improvement.</p></li>\n<li><p>We explored the idea of identifying normal cases using a tumor detection model trained on the provided annotations. We binarized the annotations (tumor=1, stroma/necrosis=0) and trained a logistic regression on patch features to identify patches containing tumor. The percentage of tumor patches in a WSI/TMA would be used to identify normal cases. This idea did not provide any performance improvement on the public LB.</p></li>\n<li><p>We used the <a href=\"https://docs.ray.io/en/latest/tune/index.html\" target=\"_blank\">Ray Tune</a> library to do hyperparameter tuning with <a href=\"https://arxiv.org/pdf/1802.02212.pdf\" target=\"_blank\">Chowder</a> and <a href=\"https://arxiv.org/abs/1802.04712\" target=\"_blank\">DeepMIL</a>. The hyperparameters we optimized for were: batch size, number of training epochs, learning rate, dimensions of hidden layers in MLP, activation functions. The hyperparameter tuning resulted in an increase of our local CV scores, but in a decrease of our submissions scores on the public LB. We hypothesized that hyperparameter tuning led to overfitting the train set and dropped the idea.</p></li>\n<li><p>[Successful idea 💡] Fine-tuning Phikon. Here, we’re not talking about fine-tuning in a conventional way; Instead, we mean pretraining a ViT-Base, initialized with Phikon’s weights, using <a href=\"https://github.com/bytedance/ibot\" target=\"_blank\">iBOT</a> on patches from the images in the train set. To do so, we extracted a total of 6.5M patches (224 x 224 px RGB) from the images in the train set. Following the recent paper from <a href=\"https://arxiv.org/abs/2309.16588\" target=\"_blank\">Darcet et al. 2023</a> and the work of <a href=\"https://github.com/facebookresearch/dinov2\" target=\"_blank\">Dino V2</a>, we added 4 register tokens to the ViT-Base. This ViT was trained for a single epoch with an initial learning rate of 0.0005 and batch size (per device) of 32. A single epoch took 2.5 days on 2 NVIDIA P100 GPUs. In a submission, we combined Chowder models trained on features extracted with the <em>original Phikon</em> and Chowder models trained on features extracted with the <em>fine-tuned Phikon</em> (with register tokens). This submission scored 0.62/<strong>0.67</strong> (not selected for the final evaluation).</p></li>\n<li><p>[Successful idea 💡] Using the variance of predictions (across models in an ensemble of Chowder) to identify outliers. A submission implementing this idea (along with the entropy of predictions) scored 0.59/<strong>0.67</strong> (not selected for the final evaluation).</p></li>\n</ul>\n<h2>7. Notes</h2>\n<p>The magnification of the images proved to be a key information. We found that resizing the TMA patches to match those from WSI allowed us to have consistent performance over all images in the test set. Keeping the original pyramidal image would have saved the participants from cumbersome image processing and could have led to overall better performances. </p>\n<p>For each image (WSI or TMA), we converted the extracted patches to grayscale and applied contrast equalization; Phikon features were then averaged across all available patches. We applied UMAP dimension reduction to these averaged representations and noticed that 38 WSI could be told apart from the remaining 500 images. . Red cells can be used as a scale to compare the resolution levels. Comparing red-cells’ size in slide 431 (among the main group of slides) and slide 4 (among the 38 odd slides) shows there is a substantial difference in resolution between the two groups, with an estimated x12 magnification for the odd slides, instead of the normal x20 magnification. This could be the result of a failed conversion from the original .svs (or .tiff) format of the slide to .png, resulting from the selection of the wrong zoom level. </p>\n<p>Such resolution variances are significant: 8% of all slides exhibited incorrect zoom levels impacting model performance. Notably, 20% (9 out of 47) of all LGSC slides had an incorrect zoom level. Finally, this UMAP also shows that after downsampling the TMA, their features are mixed with those of the WSI, this confirms our ability to use a single approach for these two data types.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F705762%2Fac94b6191f21b8c8184f40f03800a74f%2Fumap.png?generation=1704734453560448&amp;alt=media\" alt=\"\"></p>\n<p>The ids of the odd looking slides are the following:    [4,970,1080,2097,3222,3511,3881,9509,12159,12244,13364,13387,15583,15871,25604,26124,29888,31300,31793,32035,32192,33839,34688,34720,40079,40639,41099,44432,44700,49995,51215,52308,52784,53402,61100,63298,63836,64629]</p>\n<p>jbschiratti, on behalf of Owkin's team</p>",
      "rawMarkdown": "# 1st Place Solution 🥇 [Owkin] -- Phikon & Chowder\n\n## Introduction\n\nFirst of all, we would like to thank the University of British Columbia (UBC) for this exceptional multi-centric cohort and the Kaggle staff for organizing this competition. We got into this competition to showcase the efficiency and robustness of Phikon ([model card](https://huggingface.co/owkin/phikon), [paper](https://www.medrxiv.org/content/10.1101/2023.07.21.23292757v1), [blog post](https://huggingface.co/blog/EazyAl/phikon)), the foundation model (FM) for digital pathology made available to the community by [Owkin](https://www.owkin.com/) last November. We are very pleased with the outcome and really enjoyed participating in this competition.\n\nOur solution is straightforward: we trained an ensemble of [Chowder](https://arxiv.org/pdf/1802.02212.pdf) models on top of Phikon tile embeddings. We used high entropy predictions to detect outliers. We did not use extra training data nor annotations (other than the ones provided by the organizers). Our *winning submission* submission scored 0.64/**0.66** (public/private) and our *top submission* scored 0.62/**0.68**. These submissions run in approximately 6 hours 🚀\n\nOur code is available [here](https://www.kaggle.com/code/jbschiratti/winning-submission). We cleaned our code (removed comments, unused code, added sections…) and created the `winning_submission` notebook. After a late submission, this notebook scores 0.63/**0.66**. The slight difference on the public LB is likely due to differences in the sampling of patches.\n\nOur solution write-up is structured as follows:\n\n1. [Main takeaways](#1-our-main-takeaways)\n2. [Matter detection and tiling](#2-matter-detection-and-tiling)\n3. [Feature extraction](#3-feature-extraction)\n4. [Subtypes classification](#4-subtypes-classification)\n5. [Outlier detection](#5-outlier-detection)\n6. [Some (un)successful ideas](#6-some-(un)successful-ideas)\n7. [Notes](#7-notes)\n\n![Our pipeline](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F705762%2Fe5d2f490c74b954bf94ce12871659276%2Fovr_pipeline.png?generation=1704733650042787&alt=media)\n\n## 1. Our main takeaways\n\n* Foundation models (or domain-specific large vision models) are the way of the future. Our results further validate the effectiveness of [Phikon](https://huggingface.co/owkin/phikon), Owkin's foundation model for digital pathology. The next frontier is multimodality: combining spatial-omics, imaging, clinical data - from genotype to phenotype - and blending it with medical knowledge and reasoning powered by Large Language Models (LLMs).\n\n* Occam's razor: our simple and efficient pipeline outperformed more complex approaches. In particular, [Chowder](https://arxiv.org/pdf/1802.02212.pdf) - a Multiple Instance Learning (MIL) model - is still on par with more recent MIL models (e.g. [TransMIL](https://arxiv.org/abs/2106.00908), [DTFD-MIL](https://arxiv.org/abs/2203.12081)), especially when combined with ensembling techniques. As opposed to more elaborate MIL models, Chowder is also interpretable.\n\n* This competition was not easy! The PNG image format posed a major challenge: loading large PNG images with standard libraries ([Pillow](https://pillow.readthedocs.io/en/stable/), [OpenCV](https://opencv.org/)) was time consuming and used a lot of RAM. Standard formats in digital pathology (SVS, TIFF, NDPI) store data pyramidally to prevent the need for loading the entire images into RAM. Although Kaggle staff members [acknowledged that PNG format was a bad choice](https://www.kaggle.com/competitions/UBC-OCEAN/discussion/446688), pyramidal images were not made available to the participants. Furthermore, either by design or as a result of converting pyramidal images to PNG, useful metadata such as mpp (image resolution in microns per pixels) and ICC profile (if available) were stripped from the images. In addition to this, working with images at different resolutions and dealing with outliers (rare variants and normal cases) in the test set made this competition challenging (and quite interesting!).\n\n* How well do our models generalize? Locally, in cross-validation (CV), the balanced accuracy scores were in the (0.8, 0.9) range. However, the scores on the public/private LB were at least 20 points lower. Obviously, we can (partly) explain this discrepancy by our ability to predict the 'Other' class. However, it also raises the question of how well our models generalize to new data (_i.e._ data points from new centers/hospitals). Even when using [Phikon](https://huggingface.co/owkin/phikon), it is likely that differences in tissue preparation, tissue staining or differences in scanner type and magnification across centers still hinder generalization.\n\n## 2. Matter detection and tiling\n\nWhole Slide Images (WSI) in digital pathology are often too large and cannot be directly analyzed using convolutional neural networks. The WSI in this competition were no exception. A well-established workaround consists in splitting regions containing tissue into smaller patches (e.g. 224 x 224 px or 512 x 512 px). As a result, a WSI can be seen as a collection of hundreds, thousands of patches. Although Tissue Microarrays (TMA) were much smaller, these images were also split into patches.\n\n### 2.1. Matter detection\n\nIn order to detect the regions of the WSI (or TMA) which contain tissue, we employed Otsu thresholding. This thresholding was applied to the thumbnail image, in the HSV color space. Although this method is not perfect, we found that it worked quite well on the images from this competition.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F705762%2Fd3dbaa1da15d2cbbfa5b7ce32381c3ad%2F1020_matter_detection.png?generation=1704733858157340&alt=media)\n\n### 2.2. Tiling\n\n#### Patch size\n\nWe set the patch size to 224 x 224 px for WSI. According to the [data description](https://www.kaggle.com/competitions/UBC-OCEAN/data), the train set is composed of a majority of WSI at magnification 20x and few (25) TMA at magnification 40x. We hypothesized that given the low number of TMA in the train set, learning would be more efficient if we standardized all images (WSI or TMA) to a 20x resolution. Therefore, we set the patch size to 448 x 448 px for TMA; These patches were then resized to 224 x 224 px. As illustrated below [left: a 448 x 448 px patch from 91.png (TMA) resized to 224 x 224 px; Right: a 224 x 224 px patch from 4211.png (WSI)], cells have roughly the same size in resized tiles from TMA and tiles from WSI. At test time, we used a logistic regression (LR) to detect if an image is a TMA. The LR was trained on features extracted from the thumbnail of train images using a pretrained (ImageNet) ResNet18.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F705762%2F88d1e1c133cfaac9dee1cce00fb977a4%2Fsample_tiles.png?generation=1704733902739438&alt=media)\n\n#### Runtime 🕕\n\n* For efficient tiling, we developed custom C code which splits a large PNG image into patches and saves them to disk (also as PNG images). This C code leverages the `libpng` library. In particular, it uses the `png_read_row` function to limit the amount of data read and loaded into RAM. This C code was easily compiled in a Kaggle notebook. The performance of the compiled code is likely to be similar to pyvips.\n\n* With a patch size of 224 x 224 px, WSI have - on average - more than 10 000 patches. Processing that many patches was not feasible given the 12 hours runtime limit. With a limit of **200 patches** per image, our first submissions successfully ran in approximately 10 hours. Later, we used [ray](https://www.ray.io/) to process images in pairs: matter detection + tiling + feature extraction for two images at a time on a single P100 GPU. Ray would spawn two processes, each using 0.5 GPU and 2 CPU cores. As a result, our submissions successfully ran in less than 7 hours. The limiting resource was the RAM; With more RAM, we could have processed 4 images at a time (on a single P100). Note that the 200 patches limit only applies to WSI. Given that TMA are small images, the number of 448 x 448 px patches hardly ever exceeded 50. Most TMA had less than 30 patches.\n\n## 3. Feature extraction\n\n[Phikon](https://huggingface.co/owkin/phikon) is a ViT-Base pre-trained with [iBOT](https://github.com/bytedance/ibot) on 40M tiles from the TCGA dataset (📝 see our [paper](https://www.medrxiv.org/content/10.1101/2023.07.21.23292757v1) for more detailed info). We benchmarked Phikon against multiple backbones including [CTransPath](https://github.com/Xiyue-Wang/TransPath), [LUNIT](https://github.com/lunit-io/benchmark-ssl-pathology#pre-trained-weights) and [DinoV2](https://github.com/facebookresearch/dinov2). Phikon outperformed these models in our local cross-validation tests and showed no improvement when combined in an ensemble. For each input patch, a `(3, 224, 224)` tensor, Phikon outputs a 768-dimensional embedding vector. Therefore, a WSI or a TMA is represented as a 2D tensor with shape `(n_patches, 768)`.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F705762%2F13ea1a1f74c768e3f1c5075e25ea282b%2Fbackbones_scores.png?generation=1704734075166073&alt=media)\n\n## 4. Subtypes classification\n\nIn order to predict the cancer subtypes from the extracted features, we considered several Multiple Instance Learning (MIL) models from our public GitHub repository [HistoSSLScaling](https://github.com/owkin/HistoSSLscaling): [Chowder](https://arxiv.org/pdf/1802.02212.pdf), [DeepMIL](https://arxiv.org/abs/1802.04712), [DSMIL](https://arxiv.org/abs/2011.08939), [TransMIL](https://arxiv.org/abs/2106.00908) and MeanPool. Chowder outperformed others including DeepMIL, MeanPool, and DSMIL.\n\nBefore the competition deadline, we selected two submissions. One submission scored **0.64**/0.64 and consisted of an ensemble of 5 Chowder, 5 DeepMIL, 5 DSMIL and 5 MeanPool trained upon Phikon, CTranspath and LUNIT. We will detail the _winning submission_ which only used Chowder and Phikon.\n\n### 4.1. Chowder's architecture\n\nThe input dimension of Chowder was set to 768 (_i.e._ the dimension of Phikon embeddings) and its output dimension to 5. The first layer of Chowder is a `TilesMLP` layer with a hidden dimension of 192. The `n_top` and `n_bottom` values of its `ExtremeLayer` layer were both set to 10. The last layer of Chowder is a `MLP` with hidden dimension 96, a dropout rate of 30% and Sigmoid activation. We refer the reader to [this Python file](https://github.com/owkin/HistoSSLscaling/blob/main/rl_benchmarks/models/slide_models/chowder.py) for an implementation of Chowder. These hyperparameters were manually selected (no hyperparameter tuning).\n\n### 4.2. Cross-validation\n\nLocally, we used stratified 5-fold cross-validation (CV) to estimate the predictive performance of our model. For each fold, four-fifths of the data were used to create a train-validation split (75%-25%) and the remaining fifth of the data was used as a test set. This stratified 5-fold CV was repeated 3 times (with different seeds changing how data was shuffled). We designed the stratified 5-fold CV to ensure that each test set would contain at least 60% of TMA. Hence, most of the TMA were used for evaluation and not for training.\n\n### 4.3. Training\n\nOur models were trained for a maximum of 30 epochs. The validation set was used for early stopping (using the validation balanced accuracy as stopping criterion) with a patience parameter of 4 epochs. In addition to this, we used the [AdamW](https://pytorch.org/docs/stable/generated/torch.optim.AdamW.html) optimizer with a constant learning rate of 0.0001 and weight decay of 0.001. The loss was the [Cross-Entropy (CE) loss](https://pytorch.org/docs/stable/generated/torch.nn.CrossEntropyLoss.html).\n\nTwo strategies were implemented to mitigate class imbalance:\n\n1. [Weighted sampling](https://pytorch.org/docs/stable/data.html#torch.utils.data.WeightedRandomSampler) to create balanced batches for training,\n2. Using class weights in [CE loss](https://pytorch.org/docs/stable/generated/torch.nn.CrossEntropyLoss.html#torch.nn.CrossEntropyLoss).\n\n### 4.4. The ensembling trick\n\nChowder can be quite sensitive to weight initialization. Instead of training a single Chowder model, we decided to train an ensemble of N=50 Chowder models. The Chowder models in the ensemble only differ by their initialization. We found the ensemble to be more stable during training - and more efficient - than a single Chowder model. The Python code below is copied from our winning submission:\n\n```python\nclass ModelEnsemble(nn.ModuleList):\n\n    def __init__(self, models: List[nn.Module]) -> None:\n        super().__init__(modules=models)\n\n    def forward(self,\n                x: torch.Tensor,\n                mask: Optional[torch.BoolTensor] = None) -> torch.Tensor:\n        \"\"\"Forward pass.\"\"\"\n        predictions, scores = [], []\n        for model in self:\n            logits_, scores_ = model(x, mask)\n            predictions.append(logits_.unsqueeze(-1))\n            scores.append(torch.mean(scores_, dim=1, keepdim=True).unsqueeze(-1))\n        predictions = torch.cat(predictions, dim=2)\n        scores = torch.cat(scores, dim=2)\n        return predictions, scores\n\n\nchowder_models = [Chowder(**chowder_kwargs) for _ in range(50)]\n\nmodel = ModelEnsemble(chowder_models)\n```\n\n### 4.5. Submissions\n\nThree repetitions of stratified 5-fold CV with an ensemble of 50 Chowder lead to a great number of Chowder models! Through several submissions, we noticed that it was more efficient to select specific repetitions and folds rather than ensembling all the 3 x 5 x 50 models. Our winning submission is the **average prediction of 65 Chowder models** trained on different data splits. We calibrated these models using a logistic regression on their internal validation set (CV), which appeared to yield a slight improvement on the public LB. After calibration, we added a “model filtering” step: we selected only a subset of the 50 Chowder models in an ensemble based on the performance of the calibrated models on their internal test set (CV).\n\n## 5. Outlier detection\n\nWith balanced accuracy as the metric, correctly predicting the 'Other' category, worth 16.6 points (100/6), was key! Our strategy to identify outliers raised our public leaderboard score from 0.59 to 0.64, highlighting its importance as the most challenging class.\n\nWe found that using a threshold on the entropy of predictions, calculated as H = -sum(p*log(p)), was most effective for us, with high entropy indicating uncertainty in predictions. As no outliers were provided, we calibrated this threshold based on the public leaderboard.\n\n## 6. Some (un)successful ideas\n\n* To account for inter-center variability, we tried several color normalization schemes (Vahadane, Reinhard) but did not observe any improvement doing so. This finding is aligned with [recent publications](https://www.nature.com/articles/s41598-023-46619-6) suggesting that staining normalization does not improve the performance of models for histopathological classification tasks.\n\n* Increasing the number of patches for TMA using a sliding window (with 30% to 80% overlap between two consecutive patches). This method significantly increased the runtime of our submissions (most failing with `Notebook Timeout`) and did not provide any performance improvement.\n\n* We explored the idea of identifying normal cases using a tumor detection model trained on the provided annotations. We binarized the annotations (tumor=1, stroma/necrosis=0) and trained a logistic regression on patch features to identify patches containing tumor. The percentage of tumor patches in a WSI/TMA would be used to identify normal cases. This idea did not provide any performance improvement on the public LB.\n\n* We used the [Ray Tune](https://docs.ray.io/en/latest/tune/index.html) library to do hyperparameter tuning with [Chowder](https://arxiv.org/pdf/1802.02212.pdf) and [DeepMIL](https://arxiv.org/abs/1802.04712). The hyperparameters we optimized for were: batch size, number of training epochs, learning rate, dimensions of hidden layers in MLP, activation functions. The hyperparameter tuning resulted in an increase of our local CV scores, but in a decrease of our submissions scores on the public LB. We hypothesized that hyperparameter tuning led to overfitting the train set and dropped the idea.\n\n* [Successful idea 💡] Fine-tuning Phikon. Here, we’re not talking about fine-tuning in a conventional way; Instead, we mean pretraining a ViT-Base, initialized with Phikon’s weights, using [iBOT](https://github.com/bytedance/ibot) on patches from the images in the train set. To do so, we extracted a total of 6.5M patches (224 x 224 px RGB) from the images in the train set. Following the recent paper from [Darcet et al. 2023](https://arxiv.org/abs/2309.16588) and the work of [Dino V2](https://github.com/facebookresearch/dinov2), we added 4 register tokens to the ViT-Base. This ViT was trained for a single epoch with an initial learning rate of 0.0005 and batch size (per device) of 32. A single epoch took 2.5 days on 2 NVIDIA P100 GPUs. In a submission, we combined Chowder models trained on features extracted with the _original Phikon_ and Chowder models trained on features extracted with the _fine-tuned Phikon_ (with register tokens). This submission scored 0.62/**0.67** (not selected for the final evaluation).\n\n* [Successful idea 💡] Using the variance of predictions (across models in an ensemble of Chowder) to identify outliers. A submission implementing this idea (along with the entropy of predictions) scored 0.59/**0.67** (not selected for the final evaluation).\n\n## 7. Notes\n\nThe magnification of the images proved to be a key information. We found that resizing the TMA patches to match those from WSI allowed us to have consistent performance over all images in the test set. Keeping the original pyramidal image would have saved the participants from cumbersome image processing and could have led to overall better performances. \n\nFor each image (WSI or TMA), we converted the extracted patches to grayscale and applied contrast equalization; Phikon features were then averaged across all available patches. We applied UMAP dimension reduction to these averaged representations and noticed that 38 WSI could be told apart from the remaining 500 images. . Red cells can be used as a scale to compare the resolution levels. Comparing red-cells’ size in slide 431 (among the main group of slides) and slide 4 (among the 38 odd slides) shows there is a substantial difference in resolution between the two groups, with an estimated x12 magnification for the odd slides, instead of the normal x20 magnification. This could be the result of a failed conversion from the original .svs (or .tiff) format of the slide to .png, resulting from the selection of the wrong zoom level. \n\nSuch resolution variances are significant: 8% of all slides exhibited incorrect zoom levels impacting model performance. Notably, 20% (9 out of 47) of all LGSC slides had an incorrect zoom level. Finally, this UMAP also shows that after downsampling the TMA, their features are mixed with those of the WSI, this confirms our ability to use a single approach for these two data types.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F705762%2Fac94b6191f21b8c8184f40f03800a74f%2Fumap.png?generation=1704734453560448&alt=media)\n\nThe ids of the odd looking slides are the following:\t[4,970,1080,2097,3222,3511,3881,9509,12159,12244,13364,13387,15583,15871,25604,26124,29888,31300,31793,32035,32192,33839,34688,34720,40079,40639,41099,44432,44700,49995,51215,52308,52784,53402,61100,63298,63836,64629]\n\njbschiratti, on behalf of Owkin's team\n\n",
      "votes": 65
    },
    {
      "id": 2608219,
      "postDate": "2024-01-18T16:15:01.017Z",
      "content": "<p>Congratulations on this large win! Very impressive piece of work!</p>\n<p>I did not participate in this competition but I would like to understand why the <a href=\"https://github.com/owkin/HistoSSLscaling/blob/main/LICENSE.txt\" target=\"_blank\">Owkin non-commercial license</a> of Phikon is valid for this competition?<br>\nThe winner license rule seems quite clear \"License that in no event limits commercial use of such code or model containing or depending on such code.\" </p>\n<p>Has this been discussed with organizers somewhere ?</p>",
      "rawMarkdown": "Congratulations on this large win! Very impressive piece of work!\n\nI did not participate in this competition but I would like to understand why the [Owkin non-commercial license](https://github.com/owkin/HistoSSLscaling/blob/main/LICENSE.txt) of Phikon is valid for this competition?\nThe winner license rule seems quite clear \"License that in no event limits commercial use of such code or model containing or depending on such code.\" \n\nHas this been discussed with organizers somewhere ?",
      "votes": 3,
      "replies": [
        {
          "id": 2608287,
          "postDate": "2024-01-18T17:04:36.953Z",
          "content": "<p>Thank you for your interest in our work. <br>\nWe are currently in the process of discussing this with the Kaggle staff and the organizing team (UBC).</p>",
          "rawMarkdown": "Thank you for your interest in our work. \nWe are currently in the process of discussing this with the Kaggle staff and the organizing team (UBC).",
          "votes": 2,
          "replies": [
            {
              "id": 2617488,
              "postDate": "2024-01-24T09:16:55.597Z",
              "content": "<p>Looks like the competition is finalized, what was the logic to accept this license?</p>",
              "rawMarkdown": "Looks like the competition is finalized, what was the logic to accept this license?",
              "votes": 1
            },
            {
              "id": 2630279,
              "postDate": "2024-02-01T08:34:54.453Z",
              "content": "<p>Discussions are still ongoing.</p>",
              "rawMarkdown": "Discussions are still ongoing.",
              "votes": 1
            },
            {
              "id": 3207071,
              "postDate": "2025-05-22T07:22:44.880Z",
              "content": "<p>Hi. Would you please update on the last question. Your response might be helpful for future</p>",
              "rawMarkdown": "Hi. Would you please update on the last question. Your response might be helpful for future"
            }
          ]
        }
      ]
    },
    {
      "id": 2593017,
      "postDate": "2024-01-09T02:06:54Z",
      "content": "<p>This is the best advertisement of Phikon. <a href=\"https://www.kaggle.com/jbschiratti\" target=\"_blank\">@jbschiratti</a> Congratulations on the championship and thank you for the contribution to the community. </p>",
      "rawMarkdown": "This is the best advertisement of Phikon. @jbschiratti Congratulations on the championship and thank you for the contribution to the community. ",
      "votes": 4,
      "replies": [
        {
          "id": 2593096,
          "postDate": "2024-01-09T03:59:54.113Z",
          "content": "<p>And the part of UMAP is very insightful. Thank you! </p>",
          "rawMarkdown": "And the part of UMAP is very insightful. Thank you! "
        },
        {
          "id": 2593359,
          "postDate": "2024-01-09T07:29:12.743Z",
          "content": "<p><a href=\"https://www.kaggle.com/forcewithme\" target=\"_blank\">@forcewithme</a>, we entered this competition to highlight Phikon. Among the top 10 teams, only we and <a href=\"https://www.kaggle.com/xdu4yangzhou\" target=\"_blank\">@xdu4yangzhou</a> (Team XDU YYDS!) utilized it, demonstrating the necessity for some advertising 😊</p>",
          "rawMarkdown": "@forcewithme, we entered this competition to highlight Phikon. Among the top 10 teams, only we and @xdu4yangzhou (Team XDU YYDS!) utilized it, demonstrating the necessity for some advertising 😊",
          "votes": 6,
          "replies": [
            {
              "id": 2596275,
              "postDate": "2024-01-11T02:25:16.903Z",
              "content": "<p>Haha, Phikon is all you need! Thanks a lot for making this excellent work open-sourced!</p>",
              "rawMarkdown": "Haha, Phikon is all you need! Thanks a lot for making this excellent work open-sourced!",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2640344,
      "postDate": "2024-02-06T20:50:50.970Z",
      "content": "<p>What would you say about <a href=\"https://arxiv.org/abs/2312.03558\" target=\"_blank\">LongVIT</a>? Have you tried it for the task or for your startup purposes? </p>",
      "rawMarkdown": "What would you say about [LongVIT](https://arxiv.org/abs/2312.03558)? Have you tried it for the task or for your startup purposes? ",
      "votes": 1,
      "replies": [
        {
          "id": 2648390,
          "postDate": "2024-02-12T08:42:45.217Z",
          "content": "<p>We did not look into it. Looks interesting, though! Did you make a (late) submission using this LongViT?</p>",
          "rawMarkdown": "We did not look into it. Looks interesting, though! Did you make a (late) submission using this LongViT?",
          "replies": [
            {
              "id": 2648682,
              "postDate": "2024-02-12T11:03:53.723Z",
              "content": "<p>I have not managed to run their code for this task as far as it was challenging to inference in 10 hours. But who knows, it could work</p>",
              "rawMarkdown": "I have not managed to run their code for this task as far as it was challenging to inference in 10 hours. But who knows, it could work"
            }
          ]
        }
      ]
    },
    {
      "id": 2605255,
      "postDate": "2024-01-17T01:31:04.513Z",
      "content": "<p>Congratulations to your team for winning first place!<br>\nOur team also tried using <code>Phikon</code>, but the performance was not very satisfactory. I think there may be two reasons:</p>\n<ol>\n<li>Our team mainly conducted experiments under 10x, while you and other teams using <code>Phikon</code> mainly conducted experiments under 20x.</li>\n<li>We haven’t tried <code>Chowder</code>, and you and other teams using <code>Phikon</code> mainly did experiments under 20x. Maybe <code>Chowder</code> is better combined with <code>Phikon</code>?</li>\n</ol>",
      "rawMarkdown": "Congratulations to your team for winning first place!\nOur team also tried using `Phikon`, but the performance was not very satisfactory. I think there may be two reasons:\n1. Our team mainly conducted experiments under 10x, while you and other teams using `Phikon` mainly conducted experiments under 20x.\n2. We haven’t tried `Chowder`, and you and other teams using `Phikon` mainly did experiments under 20x. Maybe `Chowder` is better combined with `Phikon`?",
      "votes": 1,
      "replies": [
        {
          "id": 2608299,
          "postDate": "2024-01-18T17:12:41.150Z",
          "content": "<p>Thank you for your comment.</p>\n<p>Indeed, it seems that working at 20x resolution is preferable with Phikon. As explained in the associated <a href=\"https://www.medrxiv.org/content/10.1101/2023.07.21.23292757v2.full.pdf\" target=\"_blank\">publication</a>, Phikon was pretrained on tiles at mpp 0.5 (equiv. 20x). </p>\n<p>For <em>this</em> specific competition/task, Chowder worked well with Phikon embeddings. </p>",
          "rawMarkdown": "Thank you for your comment.\n\nIndeed, it seems that working at 20x resolution is preferable with Phikon. As explained in the associated [publication](https://www.medrxiv.org/content/10.1101/2023.07.21.23292757v2.full.pdf), Phikon was pretrained on tiles at mpp 0.5 (equiv. 20x). \n\nFor *this* specific competition/task, Chowder worked well with Phikon embeddings. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 2602961,
      "postDate": "2024-01-15T13:22:18.380Z",
      "content": "<p>Congratulations! And thank you for that amazing explanation.<br>\nI have a doubt regarding how did you choose the 200 patches to be analyzed for each image. It was a random selection over all of them or did you choose them with some criteria? After reduce the procesing time did you try to increase the number of patches in order to try to improve your results?</p>",
      "rawMarkdown": "Congratulations! And thank you for that amazing explanation.\nI have a doubt regarding how did you choose the 200 patches to be analyzed for each image. It was a random selection over all of them or did you choose them with some criteria? After reduce the procesing time did you try to increase the number of patches in order to try to improve your results?",
      "votes": 1,
      "replies": [
        {
          "id": 2602981,
          "postDate": "2024-01-15T13:47:44.840Z",
          "content": "<p>Thanks ! It was indeed a random selection over patches that are within the matter mask (computed using otsu thresholding). Once we got the inference time down to 6 hours we did try to add more patches - 400 instead of 200, but we did not observe any substantial improvement of the score. It is worth noting that increasing the maximum amount of tiles per slides only affects the prediction of WSIs because TMAs only include less than 50 patches in any case. <br>\nWe also tried to create a segmentation model using the annotations to produce a more restrained matter mask and directly select information-rich tiles but this did not lead to an increased score either…</p>",
          "rawMarkdown": "Thanks ! It was indeed a random selection over patches that are within the matter mask (computed using otsu thresholding). Once we got the inference time down to 6 hours we did try to add more patches - 400 instead of 200, but we did not observe any substantial improvement of the score. It is worth noting that increasing the maximum amount of tiles per slides only affects the prediction of WSIs because TMAs only include less than 50 patches in any case. \nWe also tried to create a segmentation model using the annotations to produce a more restrained matter mask and directly select information-rich tiles but this did not lead to an increased score either...",
          "votes": 2,
          "replies": [
            {
              "id": 2611670,
              "postDate": "2024-01-20T21:46:29.647Z",
              "content": "<p>Thank you for your answer! Totally clear now. I have another question, did you use other alternatives for outlier detection instead of entropy? </p>",
              "rawMarkdown": "Thank you for your answer! Totally clear now. I have another question, did you use other alternatives for outlier detection instead of entropy? "
            },
            {
              "id": 2611742,
              "postDate": "2024-01-21T00:00:41.250Z",
              "content": "<p>We also tried max(logits), L3(logits) and the variance over the models of the ensemble (average for each class). We finally tried to detect \"normal\" cases using a tumor detection model in addition with entropy (0.59 public / 0.67 private). For all these methods we had to do multiple submissions to get the optimal one on the public LB</p>",
              "rawMarkdown": "We also tried max(logits), L3(logits) and the variance over the models of the ensemble (average for each class). We finally tried to detect \"normal\" cases using a tumor detection model in addition with entropy (0.59 public / 0.67 private). For all these methods we had to do multiple submissions to get the optimal one on the public LB",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2602151,
      "postDate": "2024-01-14T23:33:49.100Z",
      "content": "<p>Congratulations and amazing work, I appreciate the thorough explanation of your solution and the effectiveness of domain-specific large vision models. As a student, it is very insightful and helpful to learn from.</p>",
      "rawMarkdown": "Congratulations and amazing work, I appreciate the thorough explanation of your solution and the effectiveness of domain-specific large vision models. As a student, it is very insightful and helpful to learn from.",
      "votes": 2
    },
    {
      "id": 2602080,
      "postDate": "2024-01-14T21:37:57.447Z",
      "content": "<p>Huge congratulations on your win! I wish i would have known about Phikon before so thanks a lot for promoting it in such a great manner. Medical foundation models are definitely the future!<br>\nOur team focused a lot on creating a segmentation model to extract valid tumorous patches from the WSIs. Did you consider the same thing and do you think it could have give you even bette results? </p>",
      "rawMarkdown": "Huge congratulations on your win! I wish i would have known about Phikon before so thanks a lot for promoting it in such a great manner. Medical foundation models are definitely the future!\nOur team focused a lot on creating a segmentation model to extract valid tumorous patches from the WSIs. Did you consider the same thing and do you think it could have give you even bette results? ",
      "votes": 2,
      "replies": [
        {
          "id": 2602509,
          "postDate": "2024-01-15T06:52:42.743Z",
          "content": "<p>Yes, at some point we developed our own segmentation model (trained on the annotated data) to identify tumor regions on new images and sample patches in them. It turns out that it didn't improve our submissions scores on the public LB. Therefore,we decided not to use this segmentation model to guide the sampling of patches. A posteriori, I'm thinking that non-tumor patches may also matter for subtypes classification (TBD with a pathologist).</p>",
          "rawMarkdown": "Yes, at some point we developed our own segmentation model (trained on the annotated data) to identify tumor regions on new images and sample patches in them. It turns out that it didn't improve our submissions scores on the public LB. Therefore,we decided not to use this segmentation model to guide the sampling of patches. A posteriori, I'm thinking that non-tumor patches may also matter for subtypes classification (TBD with a pathologist).",
          "votes": 1
        }
      ]
    },
    {
      "id": 2598345,
      "postDate": "2024-01-12T09:45:16.363Z",
      "content": "<p>Firstly, I must commend you and the entire Owkin team for the comprehensive and insightful breakdown of your winning solution. The strategic approach, from leveraging Phikon's foundation model to the effective use of Chowder for classification, is truly impressive. Your work not only showcases the power of domain-specific models in digital pathology but also sets a benchmark for future competitions and research in the field.</p>\n<p>One question that stands out to me, which if elaborated upon could greatly benefit the community, is regarding the fine-tuning of Phikon with iBOT on patches from the train set. Could you provide more details on the impact this had on the model's performance, particularly in terms of generalization to new data, and whether you believe this approach could be a standard practice for adapting foundation models to specific datasets in digital pathology?</p>",
      "rawMarkdown": "Firstly, I must commend you and the entire Owkin team for the comprehensive and insightful breakdown of your winning solution. The strategic approach, from leveraging Phikon's foundation model to the effective use of Chowder for classification, is truly impressive. Your work not only showcases the power of domain-specific models in digital pathology but also sets a benchmark for future competitions and research in the field.\n\nOne question that stands out to me, which if elaborated upon could greatly benefit the community, is regarding the fine-tuning of Phikon with iBOT on patches from the train set. Could you provide more details on the impact this had on the model's performance, particularly in terms of generalization to new data, and whether you believe this approach could be a standard practice for adapting foundation models to specific datasets in digital pathology?",
      "votes": 2,
      "replies": [
        {
          "id": 2598688,
          "postDate": "2024-01-12T13:16:24.413Z",
          "content": "<p>\"Fine-tuning Phikon\" should be understood as: pre-train with iBOT a ViT-Base, initialized with Phikon's weights, on a large dataset of patches (extracted from the images in the train set). To do so, we just used the code from the <a href=\"https://github.com/bytedance/ibot\" target=\"_blank\">iBOT Github repository</a>. Locally (in CV), the results were slightly worse with those \"fine-tuned\" versions of Phikon (see the fourth figure of our write-up). As a result, the \"fine-tuned\" version of Phikon only did not improve our scores (public/private). However, in a submission (not selected for final evaluation), we combined Chowder models trained on top of the <em>original</em> Phikon + Chowder models trained on top of the <em>fine-tuned</em> Phikon. This submission scored 0.67 on the private LB. Overall, this idea of \"fine-tuning Phikon\" is promising but further work/research is needed.</p>",
          "rawMarkdown": "\"Fine-tuning Phikon\" should be understood as: pre-train with iBOT a ViT-Base, initialized with Phikon's weights, on a large dataset of patches (extracted from the images in the train set). To do so, we just used the code from the [iBOT Github repository](https://github.com/bytedance/ibot). Locally (in CV), the results were slightly worse with those \"fine-tuned\" versions of Phikon (see the fourth figure of our write-up). As a result, the \"fine-tuned\" version of Phikon only did not improve our scores (public/private). However, in a submission (not selected for final evaluation), we combined Chowder models trained on top of the _original_ Phikon + Chowder models trained on top of the _fine-tuned_ Phikon. This submission scored 0.67 on the private LB. Overall, this idea of \"fine-tuning Phikon\" is promising but further work/research is needed.",
          "votes": 2,
          "replies": [
            {
              "id": 2598718,
              "postDate": "2024-01-12T13:26:45.810Z",
              "content": "<p>Awesome. Thanks for the reply!</p>",
              "rawMarkdown": "Awesome. Thanks for the reply!"
            }
          ]
        }
      ]
    },
    {
      "id": 2598317,
      "postDate": "2024-01-12T09:31:01.257Z",
      "content": "<p>Great work, congratulation!<br>\nWhen you say \" in cross-validation (CV), the balanced accuracy scores were in the (0.8, 0.9)\", when you do train/valid/test separation, have you mixed all patches and randomly separated them, or have you also separated them by the original image ids? I.e. if it is possible that the train/valid/test sets could contain patches from the same original WSI/TMA image?</p>",
      "rawMarkdown": "Great work, congratulation!\nWhen you say \" in cross-validation (CV), the balanced accuracy scores were in the (0.8, 0.9)\", when you do train/valid/test separation, have you mixed all patches and randomly separated them, or have you also separated them by the original image ids? I.e. if it is possible that the train/valid/test sets could contain patches from the same original WSI/TMA image?\n",
      "votes": 2,
      "replies": [
        {
          "id": 2598331,
          "postDate": "2024-01-12T09:38:57.793Z",
          "content": "<p>The split was performed at the image level not at the patch level so the shift in performance is more likely related to new centers or scanners / outliers</p>",
          "rawMarkdown": "The split was performed at the image level not at the patch level so the shift in performance is more likely related to new centers or scanners / outliers",
          "votes": 2,
          "replies": [
            {
              "id": 2598519,
              "postDate": "2024-01-12T11:31:35.983Z",
              "content": "<p>Thanks for the answer, I will try it to see if I can also get such high CV scores with image level separation. <br>\nAnother question is seems you have not utilized the <a href=\"https://www.kaggle.com/datasets/sohier/ubc-ovarian-cancer-competition-supplemental-masks\" target=\"_blank\">supplemental-mask data</a> in the process? Could you share some thoughts on that?</p>",
              "rawMarkdown": "Thanks for the answer, I will try it to see if I can also get such high CV scores with image level separation. \nAnother question is seems you have not utilized the [supplemental-mask data](https://www.kaggle.com/datasets/sohier/ubc-ovarian-cancer-competition-supplemental-masks) in the process? Could you share some thoughts on that?"
            },
            {
              "id": 2598574,
              "postDate": "2024-01-12T12:16:53.227Z",
              "content": "<p>We did some experiments with the annotated data (masks) but it did not increase our score on the public LB. </p>\n<p>For instance, we used the annotated data to train a tumor prediction model at the patch level (i.e. a model which predicts if a given patch contains tumor or not). To do so, we binarized the annotations (tumor=1, stroma/necrosis=0). Once trained, we used model (at test time) to quantify the percentage of tumor patches in an image. If this percentage was lower than a given threshold (e.g. 1%), we would predict \"Other\". Although it seemed like a good idea (especially to identify the \"normal cases\"), it did not improve (nor decrease) our scores.</p>",
              "rawMarkdown": "We did some experiments with the annotated data (masks) but it did not increase our score on the public LB. \n\nFor instance, we used the annotated data to train a tumor prediction model at the patch level (i.e. a model which predicts if a given patch contains tumor or not). To do so, we binarized the annotations (tumor=1, stroma/necrosis=0). Once trained, we used model (at test time) to quantify the percentage of tumor patches in an image. If this percentage was lower than a given threshold (e.g. 1%), we would predict \"Other\". Although it seemed like a good idea (especially to identify the \"normal cases\"), it did not improve (nor decrease) our scores.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2593839,
      "postDate": "2024-01-09T13:44:21.497Z",
      "content": "<p>Wow congratulations, I guess more focus on domain specific models really helps. </p>",
      "rawMarkdown": "Wow congratulations, I guess more focus on domain specific models really helps. ",
      "votes": 2,
      "replies": [
        {
          "id": 2593889,
          "postDate": "2024-01-09T14:26:05.770Z",
          "content": "<p>Thanks! Indeed, we are convinced that <em>domain-specific</em> large vision models are the way to go. Andrew Ng shared <a href=\"https://www.linkedin.com/posts/andrewyng_the-lvm-large-vision-model-revolution-is-activity-7137483177714995200-nxlM?utm_source=share&amp;utm_medium=member_desktop\" target=\"_blank\">a post on LinkedIn</a> about a month ago with a similar point of view.</p>",
          "rawMarkdown": "Thanks! Indeed, we are convinced that _domain-specific_ large vision models are the way to go. Andrew Ng shared [a post on LinkedIn](https://www.linkedin.com/posts/andrewyng_the-lvm-large-vision-model-revolution-is-activity-7137483177714995200-nxlM?utm_source=share&utm_medium=member_desktop) about a month ago with a similar point of view.",
          "votes": 2
        }
      ]
    },
    {
      "id": 3207073,
      "postDate": "2025-05-22T07:30:09.797Z",
      "content": "<p>First of all, congratulations on your victory! Hats off to your spirit and hard work. Could you please share your research paper? I am curious if you have prepared one based on your proposed solution.</p>",
      "rawMarkdown": "First of all, congratulations on your victory! Hats off to your spirit and hard work. Could you please share your research paper? I am curious if you have prepared one based on your proposed solution."
    },
    {
      "id": 3199707,
      "postDate": "2025-05-11T11:18:21.420Z",
      "content": "<p>Can we reverse engineer the final solution notebook on our laptops with 24GB ram ? or do we need to use google collab with GPUs and alot of ram ? </p>",
      "rawMarkdown": "Can we reverse engineer the final solution notebook on our laptops with 24GB ram ? or do we need to use google collab with GPUs and alot of ram ? "
    },
    {
      "id": 3029691,
      "postDate": "2024-10-27T15:59:46.770Z",
      "content": "<p>There is an IEEE paper that is academic fraud. The title is \"OCEAN - Ovarian Cancer subtypE clAssification and outlier detectionioN using DenseNet121\". It claims that only convolutional networks were used without mentioning multi-instance learning. Then, 2,000 WSI images from the OECEN-UBC competition that our contestants could not get were used for training to obtain a classification accuracy of 99.7%. However, our first place winner only had a test accuracy of 0.6 and a 5-fold cross-validation accuracy of less than 90%<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16846649%2Fff519f283467f5bcf2e5416645a08cfa%2F_20241027234938.png?generation=1730044736614051&amp;alt=media\" alt=\"\"><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16846649%2F2692818a6d51b91285ace54d23d8b383%2Ffake_acc.png?generation=1730044746592544&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "There is an IEEE paper that is academic fraud. The title is \"OCEAN - Ovarian Cancer subtypE clAssification and outlier detectionioN using DenseNet121\". It claims that only convolutional networks were used without mentioning multi-instance learning. Then, 2,000 WSI images from the OECEN-UBC competition that our contestants could not get were used for training to obtain a classification accuracy of 99.7%. However, our first place winner only had a test accuracy of 0.6 and a 5-fold cross-validation accuracy of less than 90%![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16846649%2Fff519f283467f5bcf2e5416645a08cfa%2F_20241027234938.png?generation=1730044736614051&alt=media)![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16846649%2F2692818a6d51b91285ace54d23d8b383%2Ffake_acc.png?generation=1730044746592544&alt=media)"
    },
    {
      "id": 2644472,
      "postDate": "2024-02-09T14:04:57.583Z",
      "content": "<p>Congratulations on the amazing work. I am new to Kaggle competitions, and I wanted to understand your strategy for working with huge datasets. Did you download it to work locally? </p>",
      "rawMarkdown": "Congratulations on the amazing work. I am new to Kaggle competitions, and I wanted to understand your strategy for working with huge datasets. Did you download it to work locally? ",
      "replies": [
        {
          "id": 2648387,
          "postDate": "2024-02-12T08:41:34.293Z",
          "content": "<p>Indeed, we downloaded the training data locally. The models were trained locally.</p>",
          "rawMarkdown": "Indeed, we downloaded the training data locally. The models were trained locally.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2628318,
      "postDate": "2024-01-31T08:16:29.937Z",
      "rawMarkdown": "",
      "votes": 3,
      "isDeleted": true
    },
    {
      "id": 2966914,
      "postDate": "2024-08-22T11:06:22.823Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!"
    }
  ],
  "comments": [
    {
      "id": 2608219,
      "author_name": "Optimo",
      "author_url": "",
      "post_date": "2024-01-18T16:15:01.017000",
      "content": "<p>Congratulations on this large win! Very impressive piece of work!</p>\n<p>I did not participate in this competition but I would like to understand why the <a href=\"https://github.com/owkin/HistoSSLscaling/blob/main/LICENSE.txt\" target=\"_blank\">Owkin non-commercial license</a> of Phikon is valid for this competition?<br>\nThe winner license rule seems quite clear \"License that in no event limits commercial use of such code or model containing or depending on such code.\" </p>\n<p>Has this been discussed with organizers somewhere ?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2608287,
          "author_name": "jibounet",
          "author_url": "",
          "post_date": "2024-01-18T17:04:36.953000",
          "content": "<p>Thank you for your interest in our work. <br>\nWe are currently in the process of discussing this with the Kaggle staff and the organizing team (UBC).</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2617488,
              "author_name": "Optimo",
              "author_url": "",
              "post_date": "2024-01-24T09:16:55.597000",
              "content": "<p>Looks like the competition is finalized, what was the logic to accept this license?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2630279,
              "author_name": "jibounet",
              "author_url": "",
              "post_date": "2024-02-01T08:34:54.453000",
              "content": "<p>Discussions are still ongoing.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3207071,
              "author_name": "Awais Ahmed",
              "author_url": "",
              "post_date": "2025-05-22T07:22:44.880000",
              "content": "<p>Hi. Would you please update on the last question. Your response might be helpful for future</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2593017,
      "author_name": "ForcewithMe",
      "author_url": "",
      "post_date": "2024-01-09T02:06:54",
      "content": "<p>This is the best advertisement of Phikon. <a href=\"https://www.kaggle.com/jbschiratti\" target=\"_blank\">@jbschiratti</a> Congratulations on the championship and thank you for the contribution to the community. </p>",
      "votes": 4,
      "replies": [
        {
          "id": 2593096,
          "author_name": "ForcewithMe",
          "author_url": "",
          "post_date": "2024-01-09T03:59:54.113000",
          "content": "<p>And the part of UMAP is very insightful. Thank you! </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2593359,
          "author_name": "simjeg",
          "author_url": "",
          "post_date": "2024-01-09T07:29:12.743000",
          "content": "<p><a href=\"https://www.kaggle.com/forcewithme\" target=\"_blank\">@forcewithme</a>, we entered this competition to highlight Phikon. Among the top 10 teams, only we and <a href=\"https://www.kaggle.com/xdu4yangzhou\" target=\"_blank\">@xdu4yangzhou</a> (Team XDU YYDS!) utilized it, demonstrating the necessity for some advertising 😊</p>",
          "votes": 6,
          "replies": [
            {
              "id": 2596275,
              "author_name": "yang_zhou",
              "author_url": "",
              "post_date": "2024-01-11T02:25:16.903000",
              "content": "<p>Haha, Phikon is all you need! Thanks a lot for making this excellent work open-sourced!</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2640344,
      "author_name": "Pizzaboi",
      "author_url": "",
      "post_date": "2024-02-06T20:50:50.970000",
      "content": "<p>What would you say about <a href=\"https://arxiv.org/abs/2312.03558\" target=\"_blank\">LongVIT</a>? Have you tried it for the task or for your startup purposes? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2648390,
          "author_name": "jibounet",
          "author_url": "",
          "post_date": "2024-02-12T08:42:45.217000",
          "content": "<p>We did not look into it. Looks interesting, though! Did you make a (late) submission using this LongViT?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2648682,
              "author_name": "Pizzaboi",
              "author_url": "",
              "post_date": "2024-02-12T11:03:53.723000",
              "content": "<p>I have not managed to run their code for this task as far as it was challenging to inference in 10 hours. But who knows, it could work</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2605255,
      "author_name": "m1dsolo",
      "author_url": "",
      "post_date": "2024-01-17T01:31:04.513000",
      "content": "<p>Congratulations to your team for winning first place!<br>\nOur team also tried using <code>Phikon</code>, but the performance was not very satisfactory. I think there may be two reasons:</p>\n<ol>\n<li>Our team mainly conducted experiments under 10x, while you and other teams using <code>Phikon</code> mainly conducted experiments under 20x.</li>\n<li>We haven’t tried <code>Chowder</code>, and you and other teams using <code>Phikon</code> mainly did experiments under 20x. Maybe <code>Chowder</code> is better combined with <code>Phikon</code>?</li>\n</ol>",
      "votes": 1,
      "replies": [
        {
          "id": 2608299,
          "author_name": "jibounet",
          "author_url": "",
          "post_date": "2024-01-18T17:12:41.150000",
          "content": "<p>Thank you for your comment.</p>\n<p>Indeed, it seems that working at 20x resolution is preferable with Phikon. As explained in the associated <a href=\"https://www.medrxiv.org/content/10.1101/2023.07.21.23292757v2.full.pdf\" target=\"_blank\">publication</a>, Phikon was pretrained on tiles at mpp 0.5 (equiv. 20x). </p>\n<p>For <em>this</em> specific competition/task, Chowder worked well with Phikon embeddings. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2602961,
      "author_name": "Cristian Cuadrado",
      "author_url": "",
      "post_date": "2024-01-15T13:22:18.380000",
      "content": "<p>Congratulations! And thank you for that amazing explanation.<br>\nI have a doubt regarding how did you choose the 200 patches to be analyzed for each image. It was a random selection over all of them or did you choose them with some criteria? After reduce the procesing time did you try to increase the number of patches in order to try to improve your results?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2602981,
          "author_name": "Etienne Andrier",
          "author_url": "",
          "post_date": "2024-01-15T13:47:44.840000",
          "content": "<p>Thanks ! It was indeed a random selection over patches that are within the matter mask (computed using otsu thresholding). Once we got the inference time down to 6 hours we did try to add more patches - 400 instead of 200, but we did not observe any substantial improvement of the score. It is worth noting that increasing the maximum amount of tiles per slides only affects the prediction of WSIs because TMAs only include less than 50 patches in any case. <br>\nWe also tried to create a segmentation model using the annotations to produce a more restrained matter mask and directly select information-rich tiles but this did not lead to an increased score either…</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2611670,
              "author_name": "Cristian Cuadrado",
              "author_url": "",
              "post_date": "2024-01-20T21:46:29.647000",
              "content": "<p>Thank you for your answer! Totally clear now. I have another question, did you use other alternatives for outlier detection instead of entropy? </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2611742,
              "author_name": "simjeg",
              "author_url": "",
              "post_date": "2024-01-21T00:00:41.250000",
              "content": "<p>We also tried max(logits), L3(logits) and the variance over the models of the ensemble (average for each class). We finally tried to detect \"normal\" cases using a tumor detection model in addition with entropy (0.59 public / 0.67 private). For all these methods we had to do multiple submissions to get the optimal one on the public LB</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2602151,
      "author_name": "Leo",
      "author_url": "",
      "post_date": "2024-01-14T23:33:49.100000",
      "content": "<p>Congratulations and amazing work, I appreciate the thorough explanation of your solution and the effectiveness of domain-specific large vision models. As a student, it is very insightful and helpful to learn from.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2602080,
      "author_name": "Jonas Gjerris",
      "author_url": "",
      "post_date": "2024-01-14T21:37:57.447000",
      "content": "<p>Huge congratulations on your win! I wish i would have known about Phikon before so thanks a lot for promoting it in such a great manner. Medical foundation models are definitely the future!<br>\nOur team focused a lot on creating a segmentation model to extract valid tumorous patches from the WSIs. Did you consider the same thing and do you think it could have give you even bette results? </p>",
      "votes": 2,
      "replies": [
        {
          "id": 2602509,
          "author_name": "jibounet",
          "author_url": "",
          "post_date": "2024-01-15T06:52:42.743000",
          "content": "<p>Yes, at some point we developed our own segmentation model (trained on the annotated data) to identify tumor regions on new images and sample patches in them. It turns out that it didn't improve our submissions scores on the public LB. Therefore,we decided not to use this segmentation model to guide the sampling of patches. A posteriori, I'm thinking that non-tumor patches may also matter for subtypes classification (TBD with a pathologist).</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2598345,
      "author_name": "Nishant Singhal",
      "author_url": "",
      "post_date": "2024-01-12T09:45:16.363000",
      "content": "<p>Firstly, I must commend you and the entire Owkin team for the comprehensive and insightful breakdown of your winning solution. The strategic approach, from leveraging Phikon's foundation model to the effective use of Chowder for classification, is truly impressive. Your work not only showcases the power of domain-specific models in digital pathology but also sets a benchmark for future competitions and research in the field.</p>\n<p>One question that stands out to me, which if elaborated upon could greatly benefit the community, is regarding the fine-tuning of Phikon with iBOT on patches from the train set. Could you provide more details on the impact this had on the model's performance, particularly in terms of generalization to new data, and whether you believe this approach could be a standard practice for adapting foundation models to specific datasets in digital pathology?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2598688,
          "author_name": "jibounet",
          "author_url": "",
          "post_date": "2024-01-12T13:16:24.413000",
          "content": "<p>\"Fine-tuning Phikon\" should be understood as: pre-train with iBOT a ViT-Base, initialized with Phikon's weights, on a large dataset of patches (extracted from the images in the train set). To do so, we just used the code from the <a href=\"https://github.com/bytedance/ibot\" target=\"_blank\">iBOT Github repository</a>. Locally (in CV), the results were slightly worse with those \"fine-tuned\" versions of Phikon (see the fourth figure of our write-up). As a result, the \"fine-tuned\" version of Phikon only did not improve our scores (public/private). However, in a submission (not selected for final evaluation), we combined Chowder models trained on top of the <em>original</em> Phikon + Chowder models trained on top of the <em>fine-tuned</em> Phikon. This submission scored 0.67 on the private LB. Overall, this idea of \"fine-tuning Phikon\" is promising but further work/research is needed.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2598718,
              "author_name": "Nishant Singhal",
              "author_url": "",
              "post_date": "2024-01-12T13:26:45.810000",
              "content": "<p>Awesome. Thanks for the reply!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2598317,
      "author_name": "Dong - CUB",
      "author_url": "",
      "post_date": "2024-01-12T09:31:01.257000",
      "content": "<p>Great work, congratulation!<br>\nWhen you say \" in cross-validation (CV), the balanced accuracy scores were in the (0.8, 0.9)\", when you do train/valid/test separation, have you mixed all patches and randomly separated them, or have you also separated them by the original image ids? I.e. if it is possible that the train/valid/test sets could contain patches from the same original WSI/TMA image?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2598331,
          "author_name": "simjeg",
          "author_url": "",
          "post_date": "2024-01-12T09:38:57.793000",
          "content": "<p>The split was performed at the image level not at the patch level so the shift in performance is more likely related to new centers or scanners / outliers</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2598519,
              "author_name": "Dong - CUB",
              "author_url": "",
              "post_date": "2024-01-12T11:31:35.983000",
              "content": "<p>Thanks for the answer, I will try it to see if I can also get such high CV scores with image level separation. <br>\nAnother question is seems you have not utilized the <a href=\"https://www.kaggle.com/datasets/sohier/ubc-ovarian-cancer-competition-supplemental-masks\" target=\"_blank\">supplemental-mask data</a> in the process? Could you share some thoughts on that?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2598574,
              "author_name": "jibounet",
              "author_url": "",
              "post_date": "2024-01-12T12:16:53.227000",
              "content": "<p>We did some experiments with the annotated data (masks) but it did not increase our score on the public LB. </p>\n<p>For instance, we used the annotated data to train a tumor prediction model at the patch level (i.e. a model which predicts if a given patch contains tumor or not). To do so, we binarized the annotations (tumor=1, stroma/necrosis=0). Once trained, we used model (at test time) to quantify the percentage of tumor patches in an image. If this percentage was lower than a given threshold (e.g. 1%), we would predict \"Other\". Although it seemed like a good idea (especially to identify the \"normal cases\"), it did not improve (nor decrease) our scores.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2593839,
      "author_name": "samu2505",
      "author_url": "",
      "post_date": "2024-01-09T13:44:21.497000",
      "content": "<p>Wow congratulations, I guess more focus on domain specific models really helps. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 2593889,
          "author_name": "jibounet",
          "author_url": "",
          "post_date": "2024-01-09T14:26:05.770000",
          "content": "<p>Thanks! Indeed, we are convinced that <em>domain-specific</em> large vision models are the way to go. Andrew Ng shared <a href=\"https://www.linkedin.com/posts/andrewyng_the-lvm-large-vision-model-revolution-is-activity-7137483177714995200-nxlM?utm_source=share&amp;utm_medium=member_desktop\" target=\"_blank\">a post on LinkedIn</a> about a month ago with a similar point of view.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 3207073,
      "author_name": "Awais Ahmed",
      "author_url": "",
      "post_date": "2025-05-22T07:30:09.797000",
      "content": "<p>First of all, congratulations on your victory! Hats off to your spirit and hard work. Could you please share your research paper? I am curious if you have prepared one based on your proposed solution.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3199707,
      "author_name": "Shahzad Hassan",
      "author_url": "",
      "post_date": "2025-05-11T11:18:21.420000",
      "content": "<p>Can we reverse engineer the final solution notebook on our laptops with 24GB ram ? or do we need to use google collab with GPUs and alot of ram ? </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3029691,
      "author_name": "Metavers",
      "author_url": "",
      "post_date": "2024-10-27T15:59:46.770000",
      "content": "<p>There is an IEEE paper that is academic fraud. The title is \"OCEAN - Ovarian Cancer subtypE clAssification and outlier detectionioN using DenseNet121\". It claims that only convolutional networks were used without mentioning multi-instance learning. Then, 2,000 WSI images from the OECEN-UBC competition that our contestants could not get were used for training to obtain a classification accuracy of 99.7%. However, our first place winner only had a test accuracy of 0.6 and a 5-fold cross-validation accuracy of less than 90%<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16846649%2Fff519f283467f5bcf2e5416645a08cfa%2F_20241027234938.png?generation=1730044736614051&amp;alt=media\" alt=\"\"><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16846649%2F2692818a6d51b91285ace54d23d8b383%2Ffake_acc.png?generation=1730044746592544&amp;alt=media\" alt=\"\"></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2644472,
      "author_name": "Kellen Guimaraes",
      "author_url": "",
      "post_date": "2024-02-09T14:04:57.583000",
      "content": "<p>Congratulations on the amazing work. I am new to Kaggle competitions, and I wanted to understand your strategy for working with huge datasets. Did you download it to work locally? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2648387,
          "author_name": "jibounet",
          "author_url": "",
          "post_date": "2024-02-12T08:41:34.293000",
          "content": "<p>Indeed, we downloaded the training data locally. The models were trained locally.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2628318,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-01-31T08:16:29.937000",
      "content": "",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2966914,
      "author_name": "Leo0081",
      "author_url": "",
      "post_date": "2024-08-22T11:06:22.823000",
      "content": "<p>Thanks for sharing!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2592618": "# 1st Place Solution 🥇 [Owkin] -- Phikon & Chowder\n\n## Introduction\n\nFirst of all, we would like to thank the University of British Columbia (UBC) for this exceptional multi-centric cohort and the Kaggle staff for organizing this competition. We got into this competition to showcase the efficiency and robustness of Phikon ([model card](https://huggingface.co/owkin/phikon), [paper](https://www.medrxiv.org/content/10.1101/2023.07.21.23292757v1), [blog post](https://huggingface.co/blog/EazyAl/phikon)), the foundation model (FM) for digital pathology made available to the community by [Owkin](https://www.owkin.com/) last November. We are very pleased with the outcome and really enjoyed participating in this competition.\n\nOur solution is straightforward: we trained an ensemble of [Chowder](https://arxiv.org/pdf/1802.02212.pdf) models on top of Phikon tile embeddings. We used high entropy predictions to detect outliers. We did not use extra training data nor annotations (other than the ones provided by the organizers). Our *winning submission* submission scored 0.64/**0.66** (public/private) and our *top submission* scored 0.62/**0.68**. These submissions run in approximately 6 hours 🚀\n\nOur code is available [here](https://www.kaggle.com/code/jbschiratti/winning-submission). We cleaned our code (removed comments, unused code, added sections…) and created the `winning_submission` notebook. After a late submission, this notebook scores 0.63/**0.66**. The slight difference on the public LB is likely due to differences in the sampling of patches.\n\nOur solution write-up is structured as follows:\n\n1. [Main takeaways](#1-our-main-takeaways)\n2. [Matter detection and tiling](#2-matter-detection-and-tiling)\n3. [Feature extraction](#3-feature-extraction)\n4. [Subtypes classification](#4-subtypes-classification)\n5. [Outlier detection](#5-outlier-detection)\n6. [Some (un)successful ideas](#6-some-(un)successful-ideas)\n7. [Notes](#7-notes)\n\n![Our pipeline](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F705762%2Fe5d2f490c74b954bf94ce12871659276%2Fovr_pipeline.png?generation=1704733650042787&alt=media)\n\n## 1. Our main takeaways\n\n* Foundation models (or domain-specific large vision models) are the way of the future. Our results further validate the effectiveness of [Phikon](https://huggingface.co/owkin/phikon), Owkin's foundation model for digital pathology. The next frontier is multimodality: combining spatial-omics, imaging, clinical data - from genotype to phenotype - and blending it with medical knowledge and reasoning powered by Large Language Models (LLMs).\n\n* Occam's razor: our simple and efficient pipeline outperformed more complex approaches. In particular, [Chowder](https://arxiv.org/pdf/1802.02212.pdf) - a Multiple Instance Learning (MIL) model - is still on par with more recent MIL models (e.g. [TransMIL](https://arxiv.org/abs/2106.00908), [DTFD-MIL](https://arxiv.org/abs/2203.12081)), especially when combined with ensembling techniques. As opposed to more elaborate MIL models, Chowder is also interpretable.\n\n* This competition was not easy! The PNG image format posed a major challenge: loading large PNG images with standard libraries ([Pillow](https://pillow.readthedocs.io/en/stable/), [OpenCV](https://opencv.org/)) was time consuming and used a lot of RAM. Standard formats in digital pathology (SVS, TIFF, NDPI) store data pyramidally to prevent the need for loading the entire images into RAM. Although Kaggle staff members [acknowledged that PNG format was a bad choice](https://www.kaggle.com/competitions/UBC-OCEAN/discussion/446688), pyramidal images were not made available to the participants. Furthermore, either by design or as a result of converting pyramidal images to PNG, useful metadata such as mpp (image resolution in microns per pixels) and ICC profile (if available) were stripped from the images. In addition to this, working with images at different resolutions and dealing with outliers (rare variants and normal cases) in the test set made this competition challenging (and quite interesting!).\n\n* How well do our models generalize? Locally, in cross-validation (CV), the balanced accuracy scores were in the (0.8, 0.9) range. However, the scores on the public/private LB were at least 20 points lower. Obviously, we can (partly) explain this discrepancy by our ability to predict the 'Other' class. However, it also raises the question of how well our models generalize to new data (_i.e._ data points from new centers/hospitals). Even when using [Phikon](https://huggingface.co/owkin/phikon), it is likely that differences in tissue preparation, tissue staining or differences in scanner type and magnification across centers still hinder generalization.\n\n## 2. Matter detection and tiling\n\nWhole Slide Images (WSI) in digital pathology are often too large and cannot be directly analyzed using convolutional neural networks. The WSI in this competition were no exception. A well-established workaround consists in splitting regions containing tissue into smaller patches (e.g. 224 x 224 px or 512 x 512 px). As a result, a WSI can be seen as a collection of hundreds, thousands of patches. Although Tissue Microarrays (TMA) were much smaller, these images were also split into patches.\n\n### 2.1. Matter detection\n\nIn order to detect the regions of the WSI (or TMA) which contain tissue, we employed Otsu thresholding. This thresholding was applied to the thumbnail image, in the HSV color space. Although this method is not perfect, we found that it worked quite well on the images from this competition.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F705762%2Fd3dbaa1da15d2cbbfa5b7ce32381c3ad%2F1020_matter_detection.png?generation=1704733858157340&alt=media)\n\n### 2.2. Tiling\n\n#### Patch size\n\nWe set the patch size to 224 x 224 px for WSI. According to the [data description](https://www.kaggle.com/competitions/UBC-OCEAN/data), the train set is composed of a majority of WSI at magnification 20x and few (25) TMA at magnification 40x. We hypothesized that given the low number of TMA in the train set, learning would be more efficient if we standardized all images (WSI or TMA) to a 20x resolution. Therefore, we set the patch size to 448 x 448 px for TMA; These patches were then resized to 224 x 224 px. As illustrated below [left: a 448 x 448 px patch from 91.png (TMA) resized to 224 x 224 px; Right: a 224 x 224 px patch from 4211.png (WSI)], cells have roughly the same size in resized tiles from TMA and tiles from WSI. At test time, we used a logistic regression (LR) to detect if an image is a TMA. The LR was trained on features extracted from the thumbnail of train images using a pretrained (ImageNet) ResNet18.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F705762%2F88d1e1c133cfaac9dee1cce00fb977a4%2Fsample_tiles.png?generation=1704733902739438&alt=media)\n\n#### Runtime 🕕\n\n* For efficient tiling, we developed custom C code which splits a large PNG image into patches and saves them to disk (also as PNG images). This C code leverages the `libpng` library. In particular, it uses the `png_read_row` function to limit the amount of data read and loaded into RAM. This C code was easily compiled in a Kaggle notebook. The performance of the compiled code is likely to be similar to pyvips.\n\n* With a patch size of 224 x 224 px, WSI have - on average - more than 10 000 patches. Processing that many patches was not feasible given the 12 hours runtime limit. With a limit of **200 patches** per image, our first submissions successfully ran in approximately 10 hours. Later, we used [ray](https://www.ray.io/) to process images in pairs: matter detection + tiling + feature extraction for two images at a time on a single P100 GPU. Ray would spawn two processes, each using 0.5 GPU and 2 CPU cores. As a result, our submissions successfully ran in less than 7 hours. The limiting resource was the RAM; With more RAM, we could have processed 4 images at a time (on a single P100). Note that the 200 patches limit only applies to WSI. Given that TMA are small images, the number of 448 x 448 px patches hardly ever exceeded 50. Most TMA had less than 30 patches.\n\n## 3. Feature extraction\n\n[Phikon](https://huggingface.co/owkin/phikon) is a ViT-Base pre-trained with [iBOT](https://github.com/bytedance/ibot) on 40M tiles from the TCGA dataset (📝 see our [paper](https://www.medrxiv.org/content/10.1101/2023.07.21.23292757v1) for more detailed info). We benchmarked Phikon against multiple backbones including [CTransPath](https://github.com/Xiyue-Wang/TransPath), [LUNIT](https://github.com/lunit-io/benchmark-ssl-pathology#pre-trained-weights) and [DinoV2](https://github.com/facebookresearch/dinov2). Phikon outperformed these models in our local cross-validation tests and showed no improvement when combined in an ensemble. For each input patch, a `(3, 224, 224)` tensor, Phikon outputs a 768-dimensional embedding vector. Therefore, a WSI or a TMA is represented as a 2D tensor with shape `(n_patches, 768)`.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F705762%2F13ea1a1f74c768e3f1c5075e25ea282b%2Fbackbones_scores.png?generation=1704734075166073&alt=media)\n\n## 4. Subtypes classification\n\nIn order to predict the cancer subtypes from the extracted features, we considered several Multiple Instance Learning (MIL) models from our public GitHub repository [HistoSSLScaling](https://github.com/owkin/HistoSSLscaling): [Chowder](https://arxiv.org/pdf/1802.02212.pdf), [DeepMIL](https://arxiv.org/abs/1802.04712), [DSMIL](https://arxiv.org/abs/2011.08939), [TransMIL](https://arxiv.org/abs/2106.00908) and MeanPool. Chowder outperformed others including DeepMIL, MeanPool, and DSMIL.\n\nBefore the competition deadline, we selected two submissions. One submission scored **0.64**/0.64 and consisted of an ensemble of 5 Chowder, 5 DeepMIL, 5 DSMIL and 5 MeanPool trained upon Phikon, CTranspath and LUNIT. We will detail the _winning submission_ which only used Chowder and Phikon.\n\n### 4.1. Chowder's architecture\n\nThe input dimension of Chowder was set to 768 (_i.e._ the dimension of Phikon embeddings) and its output dimension to 5. The first layer of Chowder is a `TilesMLP` layer with a hidden dimension of 192. The `n_top` and `n_bottom` values of its `ExtremeLayer` layer were both set to 10. The last layer of Chowder is a `MLP` with hidden dimension 96, a dropout rate of 30% and Sigmoid activation. We refer the reader to [this Python file](https://github.com/owkin/HistoSSLscaling/blob/main/rl_benchmarks/models/slide_models/chowder.py) for an implementation of Chowder. These hyperparameters were manually selected (no hyperparameter tuning).\n\n### 4.2. Cross-validation\n\nLocally, we used stratified 5-fold cross-validation (CV) to estimate the predictive performance of our model. For each fold, four-fifths of the data were used to create a train-validation split (75%-25%) and the remaining fifth of the data was used as a test set. This stratified 5-fold CV was repeated 3 times (with different seeds changing how data was shuffled). We designed the stratified 5-fold CV to ensure that each test set would contain at least 60% of TMA. Hence, most of the TMA were used for evaluation and not for training.\n\n### 4.3. Training\n\nOur models were trained for a maximum of 30 epochs. The validation set was used for early stopping (using the validation balanced accuracy as stopping criterion) with a patience parameter of 4 epochs. In addition to this, we used the [AdamW](https://pytorch.org/docs/stable/generated/torch.optim.AdamW.html) optimizer with a constant learning rate of 0.0001 and weight decay of 0.001. The loss was the [Cross-Entropy (CE) loss](https://pytorch.org/docs/stable/generated/torch.nn.CrossEntropyLoss.html).\n\nTwo strategies were implemented to mitigate class imbalance:\n\n1. [Weighted sampling](https://pytorch.org/docs/stable/data.html#torch.utils.data.WeightedRandomSampler) to create balanced batches for training,\n2. Using class weights in [CE loss](https://pytorch.org/docs/stable/generated/torch.nn.CrossEntropyLoss.html#torch.nn.CrossEntropyLoss).\n\n### 4.4. The ensembling trick\n\nChowder can be quite sensitive to weight initialization. Instead of training a single Chowder model, we decided to train an ensemble of N=50 Chowder models. The Chowder models in the ensemble only differ by their initialization. We found the ensemble to be more stable during training - and more efficient - than a single Chowder model. The Python code below is copied from our winning submission:\n\n```python\nclass ModelEnsemble(nn.ModuleList):\n\n    def __init__(self, models: List[nn.Module]) -> None:\n        super().__init__(modules=models)\n\n    def forward(self,\n                x: torch.Tensor,\n                mask: Optional[torch.BoolTensor] = None) -> torch.Tensor:\n        \"\"\"Forward pass.\"\"\"\n        predictions, scores = [], []\n        for model in self:\n            logits_, scores_ = model(x, mask)\n            predictions.append(logits_.unsqueeze(-1))\n            scores.append(torch.mean(scores_, dim=1, keepdim=True).unsqueeze(-1))\n        predictions = torch.cat(predictions, dim=2)\n        scores = torch.cat(scores, dim=2)\n        return predictions, scores\n\n\nchowder_models = [Chowder(**chowder_kwargs) for _ in range(50)]\n\nmodel = ModelEnsemble(chowder_models)\n```\n\n### 4.5. Submissions\n\nThree repetitions of stratified 5-fold CV with an ensemble of 50 Chowder lead to a great number of Chowder models! Through several submissions, we noticed that it was more efficient to select specific repetitions and folds rather than ensembling all the 3 x 5 x 50 models. Our winning submission is the **average prediction of 65 Chowder models** trained on different data splits. We calibrated these models using a logistic regression on their internal validation set (CV), which appeared to yield a slight improvement on the public LB. After calibration, we added a “model filtering” step: we selected only a subset of the 50 Chowder models in an ensemble based on the performance of the calibrated models on their internal test set (CV).\n\n## 5. Outlier detection\n\nWith balanced accuracy as the metric, correctly predicting the 'Other' category, worth 16.6 points (100/6), was key! Our strategy to identify outliers raised our public leaderboard score from 0.59 to 0.64, highlighting its importance as the most challenging class.\n\nWe found that using a threshold on the entropy of predictions, calculated as H = -sum(p*log(p)), was most effective for us, with high entropy indicating uncertainty in predictions. As no outliers were provided, we calibrated this threshold based on the public leaderboard.\n\n## 6. Some (un)successful ideas\n\n* To account for inter-center variability, we tried several color normalization schemes (Vahadane, Reinhard) but did not observe any improvement doing so. This finding is aligned with [recent publications](https://www.nature.com/articles/s41598-023-46619-6) suggesting that staining normalization does not improve the performance of models for histopathological classification tasks.\n\n* Increasing the number of patches for TMA using a sliding window (with 30% to 80% overlap between two consecutive patches). This method significantly increased the runtime of our submissions (most failing with `Notebook Timeout`) and did not provide any performance improvement.\n\n* We explored the idea of identifying normal cases using a tumor detection model trained on the provided annotations. We binarized the annotations (tumor=1, stroma/necrosis=0) and trained a logistic regression on patch features to identify patches containing tumor. The percentage of tumor patches in a WSI/TMA would be used to identify normal cases. This idea did not provide any performance improvement on the public LB.\n\n* We used the [Ray Tune](https://docs.ray.io/en/latest/tune/index.html) library to do hyperparameter tuning with [Chowder](https://arxiv.org/pdf/1802.02212.pdf) and [DeepMIL](https://arxiv.org/abs/1802.04712). The hyperparameters we optimized for were: batch size, number of training epochs, learning rate, dimensions of hidden layers in MLP, activation functions. The hyperparameter tuning resulted in an increase of our local CV scores, but in a decrease of our submissions scores on the public LB. We hypothesized that hyperparameter tuning led to overfitting the train set and dropped the idea.\n\n* [Successful idea 💡] Fine-tuning Phikon. Here, we’re not talking about fine-tuning in a conventional way; Instead, we mean pretraining a ViT-Base, initialized with Phikon’s weights, using [iBOT](https://github.com/bytedance/ibot) on patches from the images in the train set. To do so, we extracted a total of 6.5M patches (224 x 224 px RGB) from the images in the train set. Following the recent paper from [Darcet et al. 2023](https://arxiv.org/abs/2309.16588) and the work of [Dino V2](https://github.com/facebookresearch/dinov2), we added 4 register tokens to the ViT-Base. This ViT was trained for a single epoch with an initial learning rate of 0.0005 and batch size (per device) of 32. A single epoch took 2.5 days on 2 NVIDIA P100 GPUs. In a submission, we combined Chowder models trained on features extracted with the _original Phikon_ and Chowder models trained on features extracted with the _fine-tuned Phikon_ (with register tokens). This submission scored 0.62/**0.67** (not selected for the final evaluation).\n\n* [Successful idea 💡] Using the variance of predictions (across models in an ensemble of Chowder) to identify outliers. A submission implementing this idea (along with the entropy of predictions) scored 0.59/**0.67** (not selected for the final evaluation).\n\n## 7. Notes\n\nThe magnification of the images proved to be a key information. We found that resizing the TMA patches to match those from WSI allowed us to have consistent performance over all images in the test set. Keeping the original pyramidal image would have saved the participants from cumbersome image processing and could have led to overall better performances. \n\nFor each image (WSI or TMA), we converted the extracted patches to grayscale and applied contrast equalization; Phikon features were then averaged across all available patches. We applied UMAP dimension reduction to these averaged representations and noticed that 38 WSI could be told apart from the remaining 500 images. . Red cells can be used as a scale to compare the resolution levels. Comparing red-cells’ size in slide 431 (among the main group of slides) and slide 4 (among the 38 odd slides) shows there is a substantial difference in resolution between the two groups, with an estimated x12 magnification for the odd slides, instead of the normal x20 magnification. This could be the result of a failed conversion from the original .svs (or .tiff) format of the slide to .png, resulting from the selection of the wrong zoom level. \n\nSuch resolution variances are significant: 8% of all slides exhibited incorrect zoom levels impacting model performance. Notably, 20% (9 out of 47) of all LGSC slides had an incorrect zoom level. Finally, this UMAP also shows that after downsampling the TMA, their features are mixed with those of the WSI, this confirms our ability to use a single approach for these two data types.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F705762%2Fac94b6191f21b8c8184f40f03800a74f%2Fumap.png?generation=1704734453560448&alt=media)\n\nThe ids of the odd looking slides are the following:\t[4,970,1080,2097,3222,3511,3881,9509,12159,12244,13364,13387,15583,15871,25604,26124,29888,31300,31793,32035,32192,33839,34688,34720,40079,40639,41099,44432,44700,49995,51215,52308,52784,53402,61100,63298,63836,64629]\n\njbschiratti, on behalf of Owkin's team\n\n",
    "2608219": "Congratulations on this large win! Very impressive piece of work!\n\nI did not participate in this competition but I would like to understand why the [Owkin non-commercial license](https://github.com/owkin/HistoSSLscaling/blob/main/LICENSE.txt) of Phikon is valid for this competition?\nThe winner license rule seems quite clear \"License that in no event limits commercial use of such code or model containing or depending on such code.\" \n\nHas this been discussed with organizers somewhere ?",
    "2593017": "This is the best advertisement of Phikon. @jbschiratti Congratulations on the championship and thank you for the contribution to the community. ",
    "2640344": "What would you say about [LongVIT](https://arxiv.org/abs/2312.03558)? Have you tried it for the task or for your startup purposes? ",
    "2605255": "Congratulations to your team for winning first place!\nOur team also tried using `Phikon`, but the performance was not very satisfactory. I think there may be two reasons:\n1. Our team mainly conducted experiments under 10x, while you and other teams using `Phikon` mainly conducted experiments under 20x.\n2. We haven’t tried `Chowder`, and you and other teams using `Phikon` mainly did experiments under 20x. Maybe `Chowder` is better combined with `Phikon`?",
    "2602961": "Congratulations! And thank you for that amazing explanation.\nI have a doubt regarding how did you choose the 200 patches to be analyzed for each image. It was a random selection over all of them or did you choose them with some criteria? After reduce the procesing time did you try to increase the number of patches in order to try to improve your results?",
    "2602151": "Congratulations and amazing work, I appreciate the thorough explanation of your solution and the effectiveness of domain-specific large vision models. As a student, it is very insightful and helpful to learn from.",
    "2602080": "Huge congratulations on your win! I wish i would have known about Phikon before so thanks a lot for promoting it in such a great manner. Medical foundation models are definitely the future!\nOur team focused a lot on creating a segmentation model to extract valid tumorous patches from the WSIs. Did you consider the same thing and do you think it could have give you even bette results? ",
    "2598345": "Firstly, I must commend you and the entire Owkin team for the comprehensive and insightful breakdown of your winning solution. The strategic approach, from leveraging Phikon's foundation model to the effective use of Chowder for classification, is truly impressive. Your work not only showcases the power of domain-specific models in digital pathology but also sets a benchmark for future competitions and research in the field.\n\nOne question that stands out to me, which if elaborated upon could greatly benefit the community, is regarding the fine-tuning of Phikon with iBOT on patches from the train set. Could you provide more details on the impact this had on the model's performance, particularly in terms of generalization to new data, and whether you believe this approach could be a standard practice for adapting foundation models to specific datasets in digital pathology?",
    "2598317": "Great work, congratulation!\nWhen you say \" in cross-validation (CV), the balanced accuracy scores were in the (0.8, 0.9)\", when you do train/valid/test separation, have you mixed all patches and randomly separated them, or have you also separated them by the original image ids? I.e. if it is possible that the train/valid/test sets could contain patches from the same original WSI/TMA image?\n",
    "2593839": "Wow congratulations, I guess more focus on domain specific models really helps. ",
    "3207073": "First of all, congratulations on your victory! Hats off to your spirit and hard work. Could you please share your research paper? I am curious if you have prepared one based on your proposed solution.",
    "3199707": "Can we reverse engineer the final solution notebook on our laptops with 24GB ram ? or do we need to use google collab with GPUs and alot of ram ? ",
    "3029691": "There is an IEEE paper that is academic fraud. The title is \"OCEAN - Ovarian Cancer subtypE clAssification and outlier detectionioN using DenseNet121\". It claims that only convolutional networks were used without mentioning multi-instance learning. Then, 2,000 WSI images from the OECEN-UBC competition that our contestants could not get were used for training to obtain a classification accuracy of 99.7%. However, our first place winner only had a test accuracy of 0.6 and a 5-fold cross-validation accuracy of less than 90%![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16846649%2Fff519f283467f5bcf2e5416645a08cfa%2F_20241027234938.png?generation=1730044736614051&alt=media)![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16846649%2F2692818a6d51b91285ace54d23d8b383%2Ffake_acc.png?generation=1730044746592544&alt=media)",
    "2644472": "Congratulations on the amazing work. I am new to Kaggle competitions, and I wanted to understand your strategy for working with huge datasets. Did you download it to work locally? ",
    "2628318": "",
    "2966914": "Thanks for sharing!"
  }
}