{
  "id": 465476,
  "title": "[14th Place Notes]  Image Agumentation + Domain Adaptation + ABMIL",
  "url": "/competitions/UBC-OCEAN/discussion/465476",
  "author_name": "Yannan Chen",
  "post_date": "2024-01-04T12:23:21.242000",
  "votes": 11,
  "comment_count": 0,
  "views": 0,
  "content": "<p>This was my first Kaggle competition, an enjoyable and educational journey. <br>\nHere are some key takeaways from my experience:</p>\n<h1>==================================================</h1>\n<h1>Basic structure of my model</h1>\n<p>1 backbone <strong>(densenet201)</strong> for instance-level feature extraction &gt;&gt; 2 <strong>ABMIL</strong> models (TMA and WSI separated) for bag-level classification</p>\n<h1>==================================================</h1>\n<h1>Training procedure</h1>\n<ul>\n<li>Use the <strong>152 WSI masks</strong> to extract tiles whose types I'm certain of for <strong>backbone training.</strong></li>\n<li><strong>Lock the backbone</strong> for feature extraction, applying <strong>MIL training</strong> on <strong>513 WSI dataset</strong> with <strong>a light attantion model that has its own classifier</strong>. The bag classifier's params are inherited from the backbone and is applied with a low learning rate during MIL training.</li>\n<li>Use TMA samples  for <strong>unsupervised domain adaptation</strong> training and monitoring the model's perfomance on TMA during the whole train process.</li>\n</ul>\n<h1>==================================================</h1>\n<h1>Inference procedure</h1>\n<ul>\n<li>All images are patched into 224×224 tiles: <br>\nWSI is scaled down by 0.5, TMA is scaled down by 0.25;<br>\nWSI is patched in grid of 224, TMA is patched in grid of 120;<br>\nthe maximum number of tiles for one bag is set at 512 (this is for large WSIs);</li>\n<li>A single backbone is shared for feature extraction of both WSIs and TMAs. It also performs instance-level prediction, and only tiles that are classified as cancer types will be sent to attention models. (heathy or dead tiles are eliminated)</li>\n<li>Two attention models dedicated to WSI and TMA separately transform features of tiles into one bag label, a softmax confidence threshold of 0.4 is used to re-label low-confidence predicitons as \"Other\". </li>\n</ul>\n<h1>==================================================</h1>\n<h1>Breakthroughs during exploration</h1>\n<h2>Image Augmentation</h2>\n<p>Learning that both TMA and WSI can vary significantly in staining, color, and clarity, I implemented extensive augmentation techniques. This included custom tools like using a circular mask to make WSI tiles resemble TMA more closely. Based on my LB perfomrance, I think Image Augmentation is a critical step for improving the models' genralization ability on TMA. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9629127%2F8da84bf07bd9986b2f1594b7aa2e81f9%2F2024-01-04%20211824.png?generation=1704374341582625&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9629127%2F3a7515caf75739270a9242fdbbea5937%2F2024-01-04%20212152.png?generation=1704374567405661&amp;alt=media\" alt=\"\"></p>\n<h2>Domain Adaption</h2>\n<p>Still, the model trained on WSI tiles performs worse than my expection on TMA. So I used <strong>domain adversial training techniques</strong>, from classic <strong>DANN</strong>, <strong>heurstic domain adaptaion</strong>, to <strong>toAlign</strong>. The key idea is to <strong>use the limited TMA images to help the model extract more task-related and less domain-related features without revealing their labels</strong>. This is the second and most critical step for my score boost on LB.</p>\n<p>During training, I use TMA acc to actively monitoring the model's transfering ability on TMA:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9629127%2Fcd8350f60477009890400cb196944b47%2F2024-01-04%20200547.png?generation=1704369980249966&amp;alt=media\" alt=\"\"></p>\n<p><strong><em>related resources:</em></strong><br>\nDomain-Adversarial Training of Neural Networks: <a href=\"https://arxiv.org/abs/1505.07818\" target=\"_blank\">https://arxiv.org/abs/1505.07818</a><br>\nHeuristic Domain Adaptation: <a href=\"https://arxiv.org/abs/2011.14540\" target=\"_blank\">https://arxiv.org/abs/2011.14540</a><br>\nToAlign: Task-oriented Alignment for Unsupervised Domain Adaptation: <a href=\"https://arxiv.org/abs/2106.10812\" target=\"_blank\">https://arxiv.org/abs/2106.10812</a></p>\n<h2>AB-MIL (Attention-based Deep Multiple Instance Learning)</h2>\n<p>I used <strong>the most basic attention-based MIL model</strong> with <strong>a self-attention kernel</strong> whose impact I'm unclear of. Based on my observation, if the feature extractor is trained well, the model performed adequately with basic MIL (max/mean scoring) on public LB. However, AB-MIL offered much better accuracy on my local valid dataset, therefore theoretically more stable and superior performance.</p>\n<p><strong><em>related resources:</em></strong><br>\nAttention-based Deep Multiple Instance Learning: <a href=\"https://arxiv.org/abs/1802.04712\" target=\"_blank\">https://arxiv.org/abs/1802.04712</a><br>\nKernel Self-Attention in Deep Multiple Instance Learning: <a href=\"https://arxiv.org/abs/2005.12991\" target=\"_blank\">https://arxiv.org/abs/2005.12991</a></p>\n<h1>==================================================</h1>\n<h1>Approaches that I found not quite useful</h1>\n<ul>\n<li><strong>Feature augmentation</strong><br>\nI aimed to further narrow the gap between WSI and TMA after applying various image augmentation methods. Attempting to add noise directly to the extracted features, however, proved ineffective.<br>\n<strong><em>related resources:</em></strong><br>\nA Simple Feature Augmentation for Domain Generalization: <a href=\"https://openaccess.thecvf.com/content/ICCV2021/papers/Li_A_Simple_Feature_Augmentation_for_Domain_Generalization_ICCV_2021_paper.pdf\" target=\"_blank\">https://openaccess.thecvf.com/content/ICCV2021/papers/Li_A_Simple_Feature_Augmentation_for_Domain_Generalization_ICCV_2021_paper.pdf</a></li>\n<li><strong>Switch backbones</strong><br>\nI tried different types of backbones, from classic reset to popular efficientnet, None of them perfomed better than densenet with my pipeline. </li>\n<li><strong>Traditional anomaly detection techniques for outlier detection</strong><br>\nI tried Isolation Forest, DBSCAN on the features I extracted and found that these methods couldn't even tell the existing cancer types apart. I soon releazed that there was no way these methods could surpass my specially trained classifiers. To me this is an absoulte wrong path.</li>\n</ul>\n<h1>==================================================</h1>\n<h1>Potential improvements in the future</h1>\n<ul>\n<li>Fundimental training techniques:<br>\nLabel Smooth<br>\nMixup and CutMix<br>\nCV for the best backbone+ABMIL combination<br>\nMedian averaging for basic MIL approach instead of max/mean<br>\nUse BCEWithLogitsLoss. Unlike CrossEntropy, it computes independently for each label.<br>\nTrain in 16-bit float to increase speed, usually doesn't hurt performance</li>\n<li>Smarter ways to disthiguish outlier:<br>\nUse Entropy to thresholding<br>\nPredict bottom 5 or 10 percentile scores as Others<br>\nTrain directly with external other cancer types<br>\nsynthesize images with exsiting types as \"Other\" data for training (I doubt its validity, but seems to work as well)</li>\n<li>I didn't use Model Ensemble at all. There are multiple ways to use ensemble:<br>\nEnsemble of backbones of the same structure but trained on different scales<br>\nEnsemble of backbones of different strutrues<br>\nEnsemble of backbones (like ConvNext, HoRNet, EfficientNetV1, and EfficientNetV2) predicts different sets of labels for the same input, using sigmoid activations to combine / compare  independent probabilities across models<br>\nEnsemble of attention models of different structures</li>\n<li>Despite my unsuccessful tests with different backbones, many teams with top LB scores credited models specifically trained on Pathology dataset, which I think should be of vital help:<br>\nCTransPath ( <a href=\"https://github.com/Xiyue-Wang/TransPath\" target=\"_blank\">https://github.com/Xiyue-Wang/TransPath</a> )<br>\nLunitDINO ( <a href=\"https://github.com/lunit-io/benchmark-ssl-pathology\" target=\"_blank\">https://github.com/lunit-io/benchmark-ssl-pathology</a> )<br>\niBOT-ViT ( <a href=\"https://github.com/owkin/HistoSSLscaling\" target=\"_blank\">https://github.com/owkin/HistoSSLscaling</a> )</li>\n<li>Try more sophisticated MIL attention models:<br>\nDTFD-MIL ( <a href=\"https://arxiv.org/abs/2203.12081\" target=\"_blank\">https://arxiv.org/abs/2203.12081</a> )<br>\nTransMIL ( <a href=\"https://arxiv.org/abs/2106.00908\" target=\"_blank\">https://arxiv.org/abs/2106.00908</a> )<br>\nCLAM ( <a href=\"https://github.com/mahmoodlab/CLAM\" target=\"_blank\">https://github.com/mahmoodlab/CLAM</a> )<br>\nDSMIL ( <a href=\"https://github.com/binli123/dsmil-wsi\" target=\"_blank\">https://github.com/binli123/dsmil-wsi</a> )<br>\nPerceiver ( <a href=\"https://github.com/cgtuebingen/DualQueryMIL\" target=\"_blank\">https://github.com/cgtuebingen/DualQueryMIL</a> )</li>\n</ul>\n<h1>==================================================</h1>\n<p>My notebook link: <br>\n<a href=\"https://www.kaggle.com/code/yannan90/ubc-submit-att\" target=\"_blank\">https://www.kaggle.com/code/yannan90/ubc-submit-att</a></p>",
  "messages": [
    {
      "id": 2586817,
      "postDate": "2024-01-04T12:23:21.243Z",
      "content": "<p>This was my first Kaggle competition, an enjoyable and educational journey. <br>\nHere are some key takeaways from my experience:</p>\n<h1>==================================================</h1>\n<h1>Basic structure of my model</h1>\n<p>1 backbone <strong>(densenet201)</strong> for instance-level feature extraction &gt;&gt; 2 <strong>ABMIL</strong> models (TMA and WSI separated) for bag-level classification</p>\n<h1>==================================================</h1>\n<h1>Training procedure</h1>\n<ul>\n<li>Use the <strong>152 WSI masks</strong> to extract tiles whose types I'm certain of for <strong>backbone training.</strong></li>\n<li><strong>Lock the backbone</strong> for feature extraction, applying <strong>MIL training</strong> on <strong>513 WSI dataset</strong> with <strong>a light attantion model that has its own classifier</strong>. The bag classifier's params are inherited from the backbone and is applied with a low learning rate during MIL training.</li>\n<li>Use TMA samples  for <strong>unsupervised domain adaptation</strong> training and monitoring the model's perfomance on TMA during the whole train process.</li>\n</ul>\n<h1>==================================================</h1>\n<h1>Inference procedure</h1>\n<ul>\n<li>All images are patched into 224×224 tiles: <br>\nWSI is scaled down by 0.5, TMA is scaled down by 0.25;<br>\nWSI is patched in grid of 224, TMA is patched in grid of 120;<br>\nthe maximum number of tiles for one bag is set at 512 (this is for large WSIs);</li>\n<li>A single backbone is shared for feature extraction of both WSIs and TMAs. It also performs instance-level prediction, and only tiles that are classified as cancer types will be sent to attention models. (heathy or dead tiles are eliminated)</li>\n<li>Two attention models dedicated to WSI and TMA separately transform features of tiles into one bag label, a softmax confidence threshold of 0.4 is used to re-label low-confidence predicitons as \"Other\". </li>\n</ul>\n<h1>==================================================</h1>\n<h1>Breakthroughs during exploration</h1>\n<h2>Image Augmentation</h2>\n<p>Learning that both TMA and WSI can vary significantly in staining, color, and clarity, I implemented extensive augmentation techniques. This included custom tools like using a circular mask to make WSI tiles resemble TMA more closely. Based on my LB perfomrance, I think Image Augmentation is a critical step for improving the models' genralization ability on TMA. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9629127%2F8da84bf07bd9986b2f1594b7aa2e81f9%2F2024-01-04%20211824.png?generation=1704374341582625&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9629127%2F3a7515caf75739270a9242fdbbea5937%2F2024-01-04%20212152.png?generation=1704374567405661&amp;alt=media\" alt=\"\"></p>\n<h2>Domain Adaption</h2>\n<p>Still, the model trained on WSI tiles performs worse than my expection on TMA. So I used <strong>domain adversial training techniques</strong>, from classic <strong>DANN</strong>, <strong>heurstic domain adaptaion</strong>, to <strong>toAlign</strong>. The key idea is to <strong>use the limited TMA images to help the model extract more task-related and less domain-related features without revealing their labels</strong>. This is the second and most critical step for my score boost on LB.</p>\n<p>During training, I use TMA acc to actively monitoring the model's transfering ability on TMA:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9629127%2Fcd8350f60477009890400cb196944b47%2F2024-01-04%20200547.png?generation=1704369980249966&amp;alt=media\" alt=\"\"></p>\n<p><strong><em>related resources:</em></strong><br>\nDomain-Adversarial Training of Neural Networks: <a href=\"https://arxiv.org/abs/1505.07818\" target=\"_blank\">https://arxiv.org/abs/1505.07818</a><br>\nHeuristic Domain Adaptation: <a href=\"https://arxiv.org/abs/2011.14540\" target=\"_blank\">https://arxiv.org/abs/2011.14540</a><br>\nToAlign: Task-oriented Alignment for Unsupervised Domain Adaptation: <a href=\"https://arxiv.org/abs/2106.10812\" target=\"_blank\">https://arxiv.org/abs/2106.10812</a></p>\n<h2>AB-MIL (Attention-based Deep Multiple Instance Learning)</h2>\n<p>I used <strong>the most basic attention-based MIL model</strong> with <strong>a self-attention kernel</strong> whose impact I'm unclear of. Based on my observation, if the feature extractor is trained well, the model performed adequately with basic MIL (max/mean scoring) on public LB. However, AB-MIL offered much better accuracy on my local valid dataset, therefore theoretically more stable and superior performance.</p>\n<p><strong><em>related resources:</em></strong><br>\nAttention-based Deep Multiple Instance Learning: <a href=\"https://arxiv.org/abs/1802.04712\" target=\"_blank\">https://arxiv.org/abs/1802.04712</a><br>\nKernel Self-Attention in Deep Multiple Instance Learning: <a href=\"https://arxiv.org/abs/2005.12991\" target=\"_blank\">https://arxiv.org/abs/2005.12991</a></p>\n<h1>==================================================</h1>\n<h1>Approaches that I found not quite useful</h1>\n<ul>\n<li><strong>Feature augmentation</strong><br>\nI aimed to further narrow the gap between WSI and TMA after applying various image augmentation methods. Attempting to add noise directly to the extracted features, however, proved ineffective.<br>\n<strong><em>related resources:</em></strong><br>\nA Simple Feature Augmentation for Domain Generalization: <a href=\"https://openaccess.thecvf.com/content/ICCV2021/papers/Li_A_Simple_Feature_Augmentation_for_Domain_Generalization_ICCV_2021_paper.pdf\" target=\"_blank\">https://openaccess.thecvf.com/content/ICCV2021/papers/Li_A_Simple_Feature_Augmentation_for_Domain_Generalization_ICCV_2021_paper.pdf</a></li>\n<li><strong>Switch backbones</strong><br>\nI tried different types of backbones, from classic reset to popular efficientnet, None of them perfomed better than densenet with my pipeline. </li>\n<li><strong>Traditional anomaly detection techniques for outlier detection</strong><br>\nI tried Isolation Forest, DBSCAN on the features I extracted and found that these methods couldn't even tell the existing cancer types apart. I soon releazed that there was no way these methods could surpass my specially trained classifiers. To me this is an absoulte wrong path.</li>\n</ul>\n<h1>==================================================</h1>\n<h1>Potential improvements in the future</h1>\n<ul>\n<li>Fundimental training techniques:<br>\nLabel Smooth<br>\nMixup and CutMix<br>\nCV for the best backbone+ABMIL combination<br>\nMedian averaging for basic MIL approach instead of max/mean<br>\nUse BCEWithLogitsLoss. Unlike CrossEntropy, it computes independently for each label.<br>\nTrain in 16-bit float to increase speed, usually doesn't hurt performance</li>\n<li>Smarter ways to disthiguish outlier:<br>\nUse Entropy to thresholding<br>\nPredict bottom 5 or 10 percentile scores as Others<br>\nTrain directly with external other cancer types<br>\nsynthesize images with exsiting types as \"Other\" data for training (I doubt its validity, but seems to work as well)</li>\n<li>I didn't use Model Ensemble at all. There are multiple ways to use ensemble:<br>\nEnsemble of backbones of the same structure but trained on different scales<br>\nEnsemble of backbones of different strutrues<br>\nEnsemble of backbones (like ConvNext, HoRNet, EfficientNetV1, and EfficientNetV2) predicts different sets of labels for the same input, using sigmoid activations to combine / compare  independent probabilities across models<br>\nEnsemble of attention models of different structures</li>\n<li>Despite my unsuccessful tests with different backbones, many teams with top LB scores credited models specifically trained on Pathology dataset, which I think should be of vital help:<br>\nCTransPath ( <a href=\"https://github.com/Xiyue-Wang/TransPath\" target=\"_blank\">https://github.com/Xiyue-Wang/TransPath</a> )<br>\nLunitDINO ( <a href=\"https://github.com/lunit-io/benchmark-ssl-pathology\" target=\"_blank\">https://github.com/lunit-io/benchmark-ssl-pathology</a> )<br>\niBOT-ViT ( <a href=\"https://github.com/owkin/HistoSSLscaling\" target=\"_blank\">https://github.com/owkin/HistoSSLscaling</a> )</li>\n<li>Try more sophisticated MIL attention models:<br>\nDTFD-MIL ( <a href=\"https://arxiv.org/abs/2203.12081\" target=\"_blank\">https://arxiv.org/abs/2203.12081</a> )<br>\nTransMIL ( <a href=\"https://arxiv.org/abs/2106.00908\" target=\"_blank\">https://arxiv.org/abs/2106.00908</a> )<br>\nCLAM ( <a href=\"https://github.com/mahmoodlab/CLAM\" target=\"_blank\">https://github.com/mahmoodlab/CLAM</a> )<br>\nDSMIL ( <a href=\"https://github.com/binli123/dsmil-wsi\" target=\"_blank\">https://github.com/binli123/dsmil-wsi</a> )<br>\nPerceiver ( <a href=\"https://github.com/cgtuebingen/DualQueryMIL\" target=\"_blank\">https://github.com/cgtuebingen/DualQueryMIL</a> )</li>\n</ul>\n<h1>==================================================</h1>\n<p>My notebook link: <br>\n<a href=\"https://www.kaggle.com/code/yannan90/ubc-submit-att\" target=\"_blank\">https://www.kaggle.com/code/yannan90/ubc-submit-att</a></p>",
      "rawMarkdown": "This was my first Kaggle competition, an enjoyable and educational journey. \nHere are some key takeaways from my experience:\n\n# ==================================================\n\n#  Basic structure of my model\n\n1 backbone **(densenet201)** for instance-level feature extraction >> 2 **ABMIL** models (TMA and WSI separated) for bag-level classification\n\n# ==================================================\n\n# Training procedure\n\n- Use the **152 WSI masks** to extract tiles whose types I'm certain of for **backbone training.**\n- **Lock the backbone** for feature extraction, applying **MIL training** on **513 WSI dataset** with **a light attantion model that has its own classifier**. The bag classifier's params are inherited from the backbone and is applied with a low learning rate during MIL training.\n- Use TMA samples  for **unsupervised domain adaptation** training and monitoring the model's perfomance on TMA during the whole train process.\n\n# ==================================================\n\n# Inference procedure\n\n- All images are patched into 224×224 tiles: \nWSI is scaled down by 0.5, TMA is scaled down by 0.25;\nWSI is patched in grid of 224, TMA is patched in grid of 120;\nthe maximum number of tiles for one bag is set at 512 (this is for large WSIs);\n- A single backbone is shared for feature extraction of both WSIs and TMAs. It also performs instance-level prediction, and only tiles that are classified as cancer types will be sent to attention models. (heathy or dead tiles are eliminated)\n- Two attention models dedicated to WSI and TMA separately transform features of tiles into one bag label, a softmax confidence threshold of 0.4 is used to re-label low-confidence predicitons as \"Other\". \n\n# ==================================================\n\n# Breakthroughs during exploration\n\n## Image Augmentation\n\nLearning that both TMA and WSI can vary significantly in staining, color, and clarity, I implemented extensive augmentation techniques. This included custom tools like using a circular mask to make WSI tiles resemble TMA more closely. Based on my LB perfomrance, I think Image Augmentation is a critical step for improving the models' genralization ability on TMA. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9629127%2F8da84bf07bd9986b2f1594b7aa2e81f9%2F2024-01-04%20211824.png?generation=1704374341582625&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9629127%2F3a7515caf75739270a9242fdbbea5937%2F2024-01-04%20212152.png?generation=1704374567405661&alt=media)\n\n##  Domain Adaption\n\nStill, the model trained on WSI tiles performs worse than my expection on TMA. So I used **domain adversial training techniques**, from classic **DANN**, **heurstic domain adaptaion**, to **toAlign**. The key idea is to **use the limited TMA images to help the model extract more task-related and less domain-related features without revealing their labels**. This is the second and most critical step for my score boost on LB.\n\nDuring training, I use TMA acc to actively monitoring the model's transfering ability on TMA:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9629127%2Fcd8350f60477009890400cb196944b47%2F2024-01-04%20200547.png?generation=1704369980249966&alt=media)\n\n***related resources:***\nDomain-Adversarial Training of Neural Networks: https://arxiv.org/abs/1505.07818\nHeuristic Domain Adaptation: https://arxiv.org/abs/2011.14540\nToAlign: Task-oriented Alignment for Unsupervised Domain Adaptation: https://arxiv.org/abs/2106.10812\n\n## AB-MIL (Attention-based Deep Multiple Instance Learning)\n\nI used **the most basic attention-based MIL model** with **a self-attention kernel** whose impact I'm unclear of. Based on my observation, if the feature extractor is trained well, the model performed adequately with basic MIL (max/mean scoring) on public LB. However, AB-MIL offered much better accuracy on my local valid dataset, therefore theoretically more stable and superior performance.\n\n***related resources:***\nAttention-based Deep Multiple Instance Learning: https://arxiv.org/abs/1802.04712\nKernel Self-Attention in Deep Multiple Instance Learning: https://arxiv.org/abs/2005.12991\n\n# ==================================================\n\n# Approaches that I found not quite useful\n\n- **Feature augmentation**\nI aimed to further narrow the gap between WSI and TMA after applying various image augmentation methods. Attempting to add noise directly to the extracted features, however, proved ineffective.\n***related resources:***\nA Simple Feature Augmentation for Domain Generalization: https://openaccess.thecvf.com/content/ICCV2021/papers/Li_A_Simple_Feature_Augmentation_for_Domain_Generalization_ICCV_2021_paper.pdf\n- **Switch backbones**\nI tried different types of backbones, from classic reset to popular efficientnet, None of them perfomed better than densenet with my pipeline. \n- **Traditional anomaly detection techniques for outlier detection**\nI tried Isolation Forest, DBSCAN on the features I extracted and found that these methods couldn't even tell the existing cancer types apart. I soon releazed that there was no way these methods could surpass my specially trained classifiers. To me this is an absoulte wrong path.\n# ==================================================\n\n# Potential improvements in the future\n\n- Fundimental training techniques:\nLabel Smooth\nMixup and CutMix\nCV for the best backbone+ABMIL combination\nMedian averaging for basic MIL approach instead of max/mean\nUse BCEWithLogitsLoss. Unlike CrossEntropy, it computes independently for each label.\nTrain in 16-bit float to increase speed, usually doesn't hurt performance\n- Smarter ways to disthiguish outlier:\nUse Entropy to thresholding\nPredict bottom 5 or 10 percentile scores as Others\nTrain directly with external other cancer types\nsynthesize images with exsiting types as \"Other\" data for training (I doubt its validity, but seems to work as well)\n- I didn't use Model Ensemble at all. There are multiple ways to use ensemble:\nEnsemble of backbones of the same structure but trained on different scales\nEnsemble of backbones of different strutrues\nEnsemble of backbones (like ConvNext, HoRNet, EfficientNetV1, and EfficientNetV2) predicts different sets of labels for the same input, using sigmoid activations to combine / compare  independent probabilities across models\nEnsemble of attention models of different structures\n- Despite my unsuccessful tests with different backbones, many teams with top LB scores credited models specifically trained on Pathology dataset, which I think should be of vital help:\nCTransPath ( https://github.com/Xiyue-Wang/TransPath )\nLunitDINO ( https://github.com/lunit-io/benchmark-ssl-pathology )\niBOT-ViT ( https://github.com/owkin/HistoSSLscaling )\n- Try more sophisticated MIL attention models:\nDTFD-MIL ( https://arxiv.org/abs/2203.12081 )\nTransMIL ( https://arxiv.org/abs/2106.00908 )\nCLAM ( https://github.com/mahmoodlab/CLAM )\nDSMIL ( https://github.com/binli123/dsmil-wsi )\nPerceiver ( https://github.com/cgtuebingen/DualQueryMIL )\n\n# ==================================================\n\nMy notebook link: \nhttps://www.kaggle.com/code/yannan90/ubc-submit-att\n",
      "votes": 11
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2586817": "This was my first Kaggle competition, an enjoyable and educational journey. \nHere are some key takeaways from my experience:\n\n# ==================================================\n\n#  Basic structure of my model\n\n1 backbone **(densenet201)** for instance-level feature extraction >> 2 **ABMIL** models (TMA and WSI separated) for bag-level classification\n\n# ==================================================\n\n# Training procedure\n\n- Use the **152 WSI masks** to extract tiles whose types I'm certain of for **backbone training.**\n- **Lock the backbone** for feature extraction, applying **MIL training** on **513 WSI dataset** with **a light attantion model that has its own classifier**. The bag classifier's params are inherited from the backbone and is applied with a low learning rate during MIL training.\n- Use TMA samples  for **unsupervised domain adaptation** training and monitoring the model's perfomance on TMA during the whole train process.\n\n# ==================================================\n\n# Inference procedure\n\n- All images are patched into 224×224 tiles: \nWSI is scaled down by 0.5, TMA is scaled down by 0.25;\nWSI is patched in grid of 224, TMA is patched in grid of 120;\nthe maximum number of tiles for one bag is set at 512 (this is for large WSIs);\n- A single backbone is shared for feature extraction of both WSIs and TMAs. It also performs instance-level prediction, and only tiles that are classified as cancer types will be sent to attention models. (heathy or dead tiles are eliminated)\n- Two attention models dedicated to WSI and TMA separately transform features of tiles into one bag label, a softmax confidence threshold of 0.4 is used to re-label low-confidence predicitons as \"Other\". \n\n# ==================================================\n\n# Breakthroughs during exploration\n\n## Image Augmentation\n\nLearning that both TMA and WSI can vary significantly in staining, color, and clarity, I implemented extensive augmentation techniques. This included custom tools like using a circular mask to make WSI tiles resemble TMA more closely. Based on my LB perfomrance, I think Image Augmentation is a critical step for improving the models' genralization ability on TMA. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9629127%2F8da84bf07bd9986b2f1594b7aa2e81f9%2F2024-01-04%20211824.png?generation=1704374341582625&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9629127%2F3a7515caf75739270a9242fdbbea5937%2F2024-01-04%20212152.png?generation=1704374567405661&alt=media)\n\n##  Domain Adaption\n\nStill, the model trained on WSI tiles performs worse than my expection on TMA. So I used **domain adversial training techniques**, from classic **DANN**, **heurstic domain adaptaion**, to **toAlign**. The key idea is to **use the limited TMA images to help the model extract more task-related and less domain-related features without revealing their labels**. This is the second and most critical step for my score boost on LB.\n\nDuring training, I use TMA acc to actively monitoring the model's transfering ability on TMA:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9629127%2Fcd8350f60477009890400cb196944b47%2F2024-01-04%20200547.png?generation=1704369980249966&alt=media)\n\n***related resources:***\nDomain-Adversarial Training of Neural Networks: https://arxiv.org/abs/1505.07818\nHeuristic Domain Adaptation: https://arxiv.org/abs/2011.14540\nToAlign: Task-oriented Alignment for Unsupervised Domain Adaptation: https://arxiv.org/abs/2106.10812\n\n## AB-MIL (Attention-based Deep Multiple Instance Learning)\n\nI used **the most basic attention-based MIL model** with **a self-attention kernel** whose impact I'm unclear of. Based on my observation, if the feature extractor is trained well, the model performed adequately with basic MIL (max/mean scoring) on public LB. However, AB-MIL offered much better accuracy on my local valid dataset, therefore theoretically more stable and superior performance.\n\n***related resources:***\nAttention-based Deep Multiple Instance Learning: https://arxiv.org/abs/1802.04712\nKernel Self-Attention in Deep Multiple Instance Learning: https://arxiv.org/abs/2005.12991\n\n# ==================================================\n\n# Approaches that I found not quite useful\n\n- **Feature augmentation**\nI aimed to further narrow the gap between WSI and TMA after applying various image augmentation methods. Attempting to add noise directly to the extracted features, however, proved ineffective.\n***related resources:***\nA Simple Feature Augmentation for Domain Generalization: https://openaccess.thecvf.com/content/ICCV2021/papers/Li_A_Simple_Feature_Augmentation_for_Domain_Generalization_ICCV_2021_paper.pdf\n- **Switch backbones**\nI tried different types of backbones, from classic reset to popular efficientnet, None of them perfomed better than densenet with my pipeline. \n- **Traditional anomaly detection techniques for outlier detection**\nI tried Isolation Forest, DBSCAN on the features I extracted and found that these methods couldn't even tell the existing cancer types apart. I soon releazed that there was no way these methods could surpass my specially trained classifiers. To me this is an absoulte wrong path.\n# ==================================================\n\n# Potential improvements in the future\n\n- Fundimental training techniques:\nLabel Smooth\nMixup and CutMix\nCV for the best backbone+ABMIL combination\nMedian averaging for basic MIL approach instead of max/mean\nUse BCEWithLogitsLoss. Unlike CrossEntropy, it computes independently for each label.\nTrain in 16-bit float to increase speed, usually doesn't hurt performance\n- Smarter ways to disthiguish outlier:\nUse Entropy to thresholding\nPredict bottom 5 or 10 percentile scores as Others\nTrain directly with external other cancer types\nsynthesize images with exsiting types as \"Other\" data for training (I doubt its validity, but seems to work as well)\n- I didn't use Model Ensemble at all. There are multiple ways to use ensemble:\nEnsemble of backbones of the same structure but trained on different scales\nEnsemble of backbones of different strutrues\nEnsemble of backbones (like ConvNext, HoRNet, EfficientNetV1, and EfficientNetV2) predicts different sets of labels for the same input, using sigmoid activations to combine / compare  independent probabilities across models\nEnsemble of attention models of different structures\n- Despite my unsuccessful tests with different backbones, many teams with top LB scores credited models specifically trained on Pathology dataset, which I think should be of vital help:\nCTransPath ( https://github.com/Xiyue-Wang/TransPath )\nLunitDINO ( https://github.com/lunit-io/benchmark-ssl-pathology )\niBOT-ViT ( https://github.com/owkin/HistoSSLscaling )\n- Try more sophisticated MIL attention models:\nDTFD-MIL ( https://arxiv.org/abs/2203.12081 )\nTransMIL ( https://arxiv.org/abs/2106.00908 )\nCLAM ( https://github.com/mahmoodlab/CLAM )\nDSMIL ( https://github.com/binli123/dsmil-wsi )\nPerceiver ( https://github.com/cgtuebingen/DualQueryMIL )\n\n# ==================================================\n\nMy notebook link: \nhttps://www.kaggle.com/code/yannan90/ubc-submit-att\n"
  }
}