{
  "id": 357898,
  "title": "25th Solution - Custom MIL with timm backbones",
  "url": "/competitions/mayo-clinic-strip-ai/discussion/357898",
  "author_name": "Gunes Evitan",
  "post_date": "2022-10-06T04:29:40.094000",
  "votes": 25,
  "comment_count": 4,
  "views": 0,
  "content": "<p>When I first started this competition 2 weeks ago, I got RSNA-MICCAI Brain Tumor Radiogenomic Classification flashbacks. These two competitions are quite similar in terms of low signal but this one was slightly better and less random. I had some decent experiences in that competition and I applied those things here.</p>\n<p>Inference Notebook: <a href=\"https://www.kaggle.com/code/gunesevitan/mayo-clinic-strip-ai-inference\" target=\"_blank\">https://www.kaggle.com/code/gunesevitan/mayo-clinic-strip-ai-inference</a><br>\nGitHub Repository: <a href=\"https://github.com/gunesevitan/mayo-clinic-strip-ai\" target=\"_blank\">https://github.com/gunesevitan/mayo-clinic-strip-ai</a></p>\n<h2>Dataset</h2>\n<p>I created a pre-computed dataset of instances for training more efficiently. Pre-computed dataset consist of 16 instances of size 1024x1024 with 3 channels (RGB) per image. Dataset creation process can be simplified to</p>\n<ul>\n<li>Compress image with JPEG compression (100% JPEG quality)</li>\n<li>Resize longest edge to 20.000 pixels</li>\n<li>Extract non-overlapping instances and pad them with white background</li>\n<li>Sort instances by their sums in descending order and take top 16 instances</li>\n</ul>\n<h2>Models</h2>\n<p>I used Multiple Instance Learning (MIL) models for processing instances. My MIL model is slightly different than the one shared in this competition. I think the model in <a href=\"https://www.kaggle.com/code/analokamus/a-sample-of-multi-instance-learning-model\" target=\"_blank\">this</a> notebook concatenates feature maps on height dimension which felt weird to me, so I tried different pooling and concatenation methods. </p>\n<pre><code>class ConvolutionalMultiInstanceLearningModel(nn.Module):\n\n    def __init__(self, n_instances, model_name, pretrained, freeze_parameters, aggregation, head_class, head_args):\n\n        super(ConvolutionalMultiInstanceLearningModel, self).__init__()\n\n        self.backbone = timm.create_model(\n            model_name=model_name,\n            pretrained=pretrained,\n            num_classes=head_args['n_classes']\n        )\n\n        if freeze_parameters is not None:\n            # Freeze all parameters in backbone\n            if freeze_parameters == 'all':\n                for parameter in self.backbone.parameters():\n                    parameter.requires_grad = False\n            else:\n                # Freeze specified parameters in backbone\n                for group in freeze_parameters:\n                    if isinstance(self.backbone, timm.models.DenseNet):\n                        for parameter in self.backbone.features[group].parameters():\n                            parameter.requires_grad = False\n                    elif isinstance(self.backbone, timm.models.EfficientNet):\n                        for parameter in self.backbone.blocks[group].parameters():\n                            parameter.requires_grad = False\n\n        self.aggregation = aggregation\n        n_classifier_features = self.backbone.get_classifier().in_features\n        input_features = (n_classifier_features * n_instances) if self.aggregation == 'concat' else n_classifier_features\n        self.classification_head = eval(head_class)(input_features=input_features, **head_args)\n\n    def forward(self, x):\n\n        # Stack instances on batch dimension before passing input to feature extractor\n        input_batch_size, input_instance, input_channel, input_height, input_width = x.shape\n        x = x.view(input_batch_size * input_instance, input_channel, input_height, input_width)\n        x = self.backbone.forward_features(x)\n        feature_batch_size, feature_channel, feature_height, feature_width = x.shape\n\n        if self.aggregation == 'avg':\n            # Average feature maps of multiple instances\n            x = x.contiguous().view(input_batch_size, input_instance, feature_channel, feature_height, feature_width)\n            x = torch.mean(x, dim=1)\n        elif self.aggregation == 'max':\n            # Max feature maps of multiple instances\n            x = x.contiguous().view(input_batch_size, input_instance, feature_channel, feature_height, feature_width)\n            x = torch.max(x, dim=1)[0]\n        elif self.aggregation == 'logsumexp':\n            # LogSumExp feature maps of multiple instances\n            x = x.contiguous().view(input_batch_size, input_instance, feature_channel, feature_height, feature_width)\n            x = torch.logsumexp(x, dim=1)\n        elif self.aggregation == 'concat':\n            # Stack feature maps on channel dimension\n            x = x.contiguous().view(input_batch_size, input_instance * feature_channel, feature_height, feature_width)\n\n        output = self.classification_head(x)\n        return output\n\n    def __init__(self, n_instances, model_name, pretrained, freeze_parameters, aggregation, head_class, head_args):\n\n        super(ConvolutionalMultiInstanceLearningModel, self).__init__()\n\n        self.backbone = timm.create_model(\n            model_name=model_name,\n            pretrained=pretrained,\n            num_classes=head_args['n_classes']\n        )\n\n        if freeze_parameters is not None:\n            # Freeze all parameters in backbone\n            if freeze_parameters == 'all':\n                for parameter in self.backbone.parameters():\n                    parameter.requires_grad = False\n            else:\n                # Freeze specified parameters in backbone\n                for group in freeze_parameters:\n                    if isinstance(self.backbone, timm.models.DenseNet):\n                        for parameter in self.backbone.features[group].parameters():\n                            parameter.requires_grad = False\n                    elif isinstance(self.backbone, timm.models.EfficientNet):\n                        for parameter in self.backbone.blocks[group].parameters():\n                            parameter.requires_grad = False\n\n        self.aggregation = aggregation\n        n_classifier_features = self.backbone.get_classifier().in_features\n        input_features = (n_classifier_features * n_instances) if self.aggregation == 'concat' else n_classifier_features\n        self.classification_head = eval(head_class)(input_features=input_features, **head_args)\n\n    def forward(self, x):\n\n        # Stack instances on batch dimension before passing input to feature extractor\n        input_batch_size, input_instance, input_channel, input_height, input_width = x.shape\n        x = x.view(input_batch_size * input_instance, input_channel, input_height, input_width)\n        x = self.backbone.forward_features(x)\n        feature_batch_size, feature_channel, feature_height, feature_width = x.shape\n\n        if self.aggregation == 'avg':\n            # Average feature maps of multiple instances\n            x = x.contiguous().view(input_batch_size, input_instance, feature_channel, feature_height, feature_width)\n            x = torch.mean(x, dim=1)\n        elif self.aggregation == 'max':\n            # Max feature maps of multiple instances\n            x = x.contiguous().view(input_batch_size, input_instance, feature_channel, feature_height, feature_width)\n            x = torch.max(x, dim=1)[0]\n        elif self.aggregation == 'logsumexp':\n            # LogSumExp feature maps of multiple instances\n            x = x.contiguous().view(input_batch_size, input_instance, feature_channel, feature_height, feature_width)\n            x = torch.logsumexp(x, dim=1)\n        elif self.aggregation == 'concat':\n            # Stack feature maps on channel dimension\n            x = x.contiguous().view(input_batch_size, input_instance * feature_channel, feature_height, feature_width)\n\n        output = self.classification_head(x)\n        return output\n</code></pre>\n<p>Aggregation parameter tunes the MIL pooling or aggregating part of the model.</p>\n<ul>\n<li><code>avg</code>: average all feature maps on instance dimension</li>\n<li><code>max</code>: max all feature maps on instance dimension (this was mentioned on many papers)</li>\n<li><code>logsumexp</code>: logsumexmp all feature maps on instance (I saw this on one paper so I added it as well)</li>\n<li><code>concat</code>: concatenate all feature maps on channel dimension</li>\n</ul>\n<p><code>concat</code> worked best for me even though the feature map counts multiplied with number of instances. I was working on an attention mechanism for instance feature maps but I didn't have enough time and I eventually ditched the idea.</p>\n<h2>Validation</h2>\n<p>5 folds of cross-validation is used as the validation scheme. Splits are stratified on <code>binary_encoded_label</code>, <code>patient_id</code> doesn't overlap between folds, and dataset is shuffled before splitting. Best random seed is selected by minimizing the standard deviation of <code>binary_encoded_label</code> mean between folds which ensures target means of folds are close to each other.</p>\n<h2>Training</h2>\n<p>Loss Function: Negative log likelihood loss is used as the loss function. Binary labels are converted to multi-class labels by simply doing</p>\n<pre><code>inputs = torch.sigmoid(inputs)\ninputs = torch.cat(((1 - inputs), inputs), dim=-1)\n</code></pre>\n<p>Optimizer: AdamW with 1e-5 initial learning rate and learning rate is cycled to 1e-6 back-and-forth after every 300 steps.</p>\n<p>Batch Size: Training: 4 - Validation: 8</p>\n<p>Early Stopping: Early Stopping is triggered after 5 epochs with no validation loss improvement</p>\n<h2>Training Augmentations</h2>\n<p>Training augmentations are applied to all instances randomly.</p>\n<ul>\n<li>Resize 256x256 or 224x224</li>\n<li>Luminosity standardization by 10% chance</li>\n<li>Horizontal flip by 50% chance</li>\n<li>Vertical flip by 50% chance</li>\n<li>Random 90-degree rotation by 25% chance</li>\n<li>Random hue-saturation-value by 25% chance</li>\n<li>Coarse dropout or pixel dropout by 10% chance</li>\n<li>Normalize by instance dataset statistics</li>\n</ul>\n<h2>Ensemble</h2>\n<p>8 models are used in the ensemble. Those models are selected based on their ROC AUC scores. Models with OOF ROC AUC score greater than 0.65 are qualified for the ensemble. My best single model was DenseNetBlur121d which is quite interesting and worth more exploring.</p>\n<ul>\n<li>DenseNet121 (ImageNet pre-trained weights) - 256x256 16 instances</li>\n<li>DenseNetBlur121d (ImageNet pre-trained weights) - 256x256 16 instances</li>\n<li>DenseNet169 (ImageNet pre-trained weights) - 256x256 16 instances</li>\n<li>EfficientNetB2 (ImageNet pre-trained weights) - 256x256 16 instances</li>\n<li>EfficientNetV2 Tiny (ImageNet pre-trained weights) - 256x256 16 instances</li>\n<li>CoaT Lite Mini (ImageNet pre-trained weights) - 224x224 16 instances</li>\n<li>PoolFormer 24 (ImageNet pre-trained weights) - 224x224 16 instances</li>\n<li>Swin Transformer Tiny Patch 4 Window 7 (ImageNet pre-trained weights) - 224x224 16 instances</li>\n</ul>\n<p>Blending is utilized on selected models. Blend weights are found by trial and error.</p>\n<pre><code>blend_weights = {\n    'mil_densenet121_16_256': 0.10,\n    'mil_densenet169_16_256': 0.10,\n    'mil_densenetblur121d_16_256': 0.25,\n    'mil_efficientnetb2_16_256': 0.125,\n    'mil_efficientnetv2rwt_16_256': 0.125,\n    'mil_coatlitemini_16_224': 0.10,\n    'mil_poolformer24_16_224': 0.10,\n    'mil_swintinypatch4window7_16_224': 0.10\n}\n</code></pre>\n<p>Linear stacking is also utilized on selected models by fitting a linear regression model on training predictions and targets.</p>\n<h2>Post-processing</h2>\n<p>Predictions are averaged among patient_id groups and 0.21 is added to aggregated predictions for making the prediction 0.5 centered. That constant is selected by maximizing the OOF score.</p>\n<h2>Results</h2>\n<p>Quantitative results are provided as cross-validation, public and private leaderboard scores.<br>\nBest single model and ensemble cross-validation scores are highlighted.</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>CV ROC AUC</th>\n<th>CV Weighted Log Loss</th>\n<th>Public Leaderboard</th>\n<th>Private Leaderboard</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>DenseNet121</td>\n<td>0.6501</td>\n<td></td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>DenseNetBlur121d</td>\n<td><strong>0.6628</strong></td>\n<td></td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>DenseNet169</td>\n<td>0.6551</td>\n<td></td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>EfficientNetB2</td>\n<td>0.6501</td>\n<td></td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>EfficientNetV2 Tiny</td>\n<td>0.6500</td>\n<td></td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>CoaT Lite Mini</td>\n<td>0.6503</td>\n<td></td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>PoolFormer 24</td>\n<td>0.6536</td>\n<td></td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>Swin Transformer Tiny Patch 4 Window 7</td>\n<td>0.6497</td>\n<td></td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>Blend + patient_id Aggregation + Adjustment</td>\n<td>0.6822</td>\n<td>0.6471</td>\n<td>0.6</td>\n<td><strong>0.67897</strong></td>\n</tr>\n<tr>\n<td>Stack + patient_id Aggregation + Adjustment</td>\n<td><strong>0.6842</strong></td>\n<td><strong>0.6387</strong></td>\n<td><strong>0.6</strong></td>\n<td>0.68666</td>\n</tr>\n</tbody>\n</table>\n<p>I started overfitting when I added more and more models to ensemble. My stack was also overfitting a bit but my best submission is my blend.</p>\n<h2>What didn't work</h2>\n<ul>\n<li>Decaying learning rate</li>\n<li>Freezing backbone parameters</li>\n<li>Average and max MIL pooling</li>\n<li>Lots of other timm models</li>\n<li>Binary cross-entropy loss, focal loss, weighted log loss</li>\n<li>Label smoothing</li>\n<li>Different number of instances with larger sizes</li>\n</ul>",
  "messages": [
    {
      "id": 1974097,
      "postDate": "2022-10-06T04:29:40.093Z",
      "content": "<p>When I first started this competition 2 weeks ago, I got RSNA-MICCAI Brain Tumor Radiogenomic Classification flashbacks. These two competitions are quite similar in terms of low signal but this one was slightly better and less random. I had some decent experiences in that competition and I applied those things here.</p>\n<p>Inference Notebook: <a href=\"https://www.kaggle.com/code/gunesevitan/mayo-clinic-strip-ai-inference\" target=\"_blank\">https://www.kaggle.com/code/gunesevitan/mayo-clinic-strip-ai-inference</a><br>\nGitHub Repository: <a href=\"https://github.com/gunesevitan/mayo-clinic-strip-ai\" target=\"_blank\">https://github.com/gunesevitan/mayo-clinic-strip-ai</a></p>\n<h2>Dataset</h2>\n<p>I created a pre-computed dataset of instances for training more efficiently. Pre-computed dataset consist of 16 instances of size 1024x1024 with 3 channels (RGB) per image. Dataset creation process can be simplified to</p>\n<ul>\n<li>Compress image with JPEG compression (100% JPEG quality)</li>\n<li>Resize longest edge to 20.000 pixels</li>\n<li>Extract non-overlapping instances and pad them with white background</li>\n<li>Sort instances by their sums in descending order and take top 16 instances</li>\n</ul>\n<h2>Models</h2>\n<p>I used Multiple Instance Learning (MIL) models for processing instances. My MIL model is slightly different than the one shared in this competition. I think the model in <a href=\"https://www.kaggle.com/code/analokamus/a-sample-of-multi-instance-learning-model\" target=\"_blank\">this</a> notebook concatenates feature maps on height dimension which felt weird to me, so I tried different pooling and concatenation methods. </p>\n<pre><code>class ConvolutionalMultiInstanceLearningModel(nn.Module):\n\n    def __init__(self, n_instances, model_name, pretrained, freeze_parameters, aggregation, head_class, head_args):\n\n        super(ConvolutionalMultiInstanceLearningModel, self).__init__()\n\n        self.backbone = timm.create_model(\n            model_name=model_name,\n            pretrained=pretrained,\n            num_classes=head_args['n_classes']\n        )\n\n        if freeze_parameters is not None:\n            # Freeze all parameters in backbone\n            if freeze_parameters == 'all':\n                for parameter in self.backbone.parameters():\n                    parameter.requires_grad = False\n            else:\n                # Freeze specified parameters in backbone\n                for group in freeze_parameters:\n                    if isinstance(self.backbone, timm.models.DenseNet):\n                        for parameter in self.backbone.features[group].parameters():\n                            parameter.requires_grad = False\n                    elif isinstance(self.backbone, timm.models.EfficientNet):\n                        for parameter in self.backbone.blocks[group].parameters():\n                            parameter.requires_grad = False\n\n        self.aggregation = aggregation\n        n_classifier_features = self.backbone.get_classifier().in_features\n        input_features = (n_classifier_features * n_instances) if self.aggregation == 'concat' else n_classifier_features\n        self.classification_head = eval(head_class)(input_features=input_features, **head_args)\n\n    def forward(self, x):\n\n        # Stack instances on batch dimension before passing input to feature extractor\n        input_batch_size, input_instance, input_channel, input_height, input_width = x.shape\n        x = x.view(input_batch_size * input_instance, input_channel, input_height, input_width)\n        x = self.backbone.forward_features(x)\n        feature_batch_size, feature_channel, feature_height, feature_width = x.shape\n\n        if self.aggregation == 'avg':\n            # Average feature maps of multiple instances\n            x = x.contiguous().view(input_batch_size, input_instance, feature_channel, feature_height, feature_width)\n            x = torch.mean(x, dim=1)\n        elif self.aggregation == 'max':\n            # Max feature maps of multiple instances\n            x = x.contiguous().view(input_batch_size, input_instance, feature_channel, feature_height, feature_width)\n            x = torch.max(x, dim=1)[0]\n        elif self.aggregation == 'logsumexp':\n            # LogSumExp feature maps of multiple instances\n            x = x.contiguous().view(input_batch_size, input_instance, feature_channel, feature_height, feature_width)\n            x = torch.logsumexp(x, dim=1)\n        elif self.aggregation == 'concat':\n            # Stack feature maps on channel dimension\n            x = x.contiguous().view(input_batch_size, input_instance * feature_channel, feature_height, feature_width)\n\n        output = self.classification_head(x)\n        return output\n\n    def __init__(self, n_instances, model_name, pretrained, freeze_parameters, aggregation, head_class, head_args):\n\n        super(ConvolutionalMultiInstanceLearningModel, self).__init__()\n\n        self.backbone = timm.create_model(\n            model_name=model_name,\n            pretrained=pretrained,\n            num_classes=head_args['n_classes']\n        )\n\n        if freeze_parameters is not None:\n            # Freeze all parameters in backbone\n            if freeze_parameters == 'all':\n                for parameter in self.backbone.parameters():\n                    parameter.requires_grad = False\n            else:\n                # Freeze specified parameters in backbone\n                for group in freeze_parameters:\n                    if isinstance(self.backbone, timm.models.DenseNet):\n                        for parameter in self.backbone.features[group].parameters():\n                            parameter.requires_grad = False\n                    elif isinstance(self.backbone, timm.models.EfficientNet):\n                        for parameter in self.backbone.blocks[group].parameters():\n                            parameter.requires_grad = False\n\n        self.aggregation = aggregation\n        n_classifier_features = self.backbone.get_classifier().in_features\n        input_features = (n_classifier_features * n_instances) if self.aggregation == 'concat' else n_classifier_features\n        self.classification_head = eval(head_class)(input_features=input_features, **head_args)\n\n    def forward(self, x):\n\n        # Stack instances on batch dimension before passing input to feature extractor\n        input_batch_size, input_instance, input_channel, input_height, input_width = x.shape\n        x = x.view(input_batch_size * input_instance, input_channel, input_height, input_width)\n        x = self.backbone.forward_features(x)\n        feature_batch_size, feature_channel, feature_height, feature_width = x.shape\n\n        if self.aggregation == 'avg':\n            # Average feature maps of multiple instances\n            x = x.contiguous().view(input_batch_size, input_instance, feature_channel, feature_height, feature_width)\n            x = torch.mean(x, dim=1)\n        elif self.aggregation == 'max':\n            # Max feature maps of multiple instances\n            x = x.contiguous().view(input_batch_size, input_instance, feature_channel, feature_height, feature_width)\n            x = torch.max(x, dim=1)[0]\n        elif self.aggregation == 'logsumexp':\n            # LogSumExp feature maps of multiple instances\n            x = x.contiguous().view(input_batch_size, input_instance, feature_channel, feature_height, feature_width)\n            x = torch.logsumexp(x, dim=1)\n        elif self.aggregation == 'concat':\n            # Stack feature maps on channel dimension\n            x = x.contiguous().view(input_batch_size, input_instance * feature_channel, feature_height, feature_width)\n\n        output = self.classification_head(x)\n        return output\n</code></pre>\n<p>Aggregation parameter tunes the MIL pooling or aggregating part of the model.</p>\n<ul>\n<li><code>avg</code>: average all feature maps on instance dimension</li>\n<li><code>max</code>: max all feature maps on instance dimension (this was mentioned on many papers)</li>\n<li><code>logsumexp</code>: logsumexmp all feature maps on instance (I saw this on one paper so I added it as well)</li>\n<li><code>concat</code>: concatenate all feature maps on channel dimension</li>\n</ul>\n<p><code>concat</code> worked best for me even though the feature map counts multiplied with number of instances. I was working on an attention mechanism for instance feature maps but I didn't have enough time and I eventually ditched the idea.</p>\n<h2>Validation</h2>\n<p>5 folds of cross-validation is used as the validation scheme. Splits are stratified on <code>binary_encoded_label</code>, <code>patient_id</code> doesn't overlap between folds, and dataset is shuffled before splitting. Best random seed is selected by minimizing the standard deviation of <code>binary_encoded_label</code> mean between folds which ensures target means of folds are close to each other.</p>\n<h2>Training</h2>\n<p>Loss Function: Negative log likelihood loss is used as the loss function. Binary labels are converted to multi-class labels by simply doing</p>\n<pre><code>inputs = torch.sigmoid(inputs)\ninputs = torch.cat(((1 - inputs), inputs), dim=-1)\n</code></pre>\n<p>Optimizer: AdamW with 1e-5 initial learning rate and learning rate is cycled to 1e-6 back-and-forth after every 300 steps.</p>\n<p>Batch Size: Training: 4 - Validation: 8</p>\n<p>Early Stopping: Early Stopping is triggered after 5 epochs with no validation loss improvement</p>\n<h2>Training Augmentations</h2>\n<p>Training augmentations are applied to all instances randomly.</p>\n<ul>\n<li>Resize 256x256 or 224x224</li>\n<li>Luminosity standardization by 10% chance</li>\n<li>Horizontal flip by 50% chance</li>\n<li>Vertical flip by 50% chance</li>\n<li>Random 90-degree rotation by 25% chance</li>\n<li>Random hue-saturation-value by 25% chance</li>\n<li>Coarse dropout or pixel dropout by 10% chance</li>\n<li>Normalize by instance dataset statistics</li>\n</ul>\n<h2>Ensemble</h2>\n<p>8 models are used in the ensemble. Those models are selected based on their ROC AUC scores. Models with OOF ROC AUC score greater than 0.65 are qualified for the ensemble. My best single model was DenseNetBlur121d which is quite interesting and worth more exploring.</p>\n<ul>\n<li>DenseNet121 (ImageNet pre-trained weights) - 256x256 16 instances</li>\n<li>DenseNetBlur121d (ImageNet pre-trained weights) - 256x256 16 instances</li>\n<li>DenseNet169 (ImageNet pre-trained weights) - 256x256 16 instances</li>\n<li>EfficientNetB2 (ImageNet pre-trained weights) - 256x256 16 instances</li>\n<li>EfficientNetV2 Tiny (ImageNet pre-trained weights) - 256x256 16 instances</li>\n<li>CoaT Lite Mini (ImageNet pre-trained weights) - 224x224 16 instances</li>\n<li>PoolFormer 24 (ImageNet pre-trained weights) - 224x224 16 instances</li>\n<li>Swin Transformer Tiny Patch 4 Window 7 (ImageNet pre-trained weights) - 224x224 16 instances</li>\n</ul>\n<p>Blending is utilized on selected models. Blend weights are found by trial and error.</p>\n<pre><code>blend_weights = {\n    'mil_densenet121_16_256': 0.10,\n    'mil_densenet169_16_256': 0.10,\n    'mil_densenetblur121d_16_256': 0.25,\n    'mil_efficientnetb2_16_256': 0.125,\n    'mil_efficientnetv2rwt_16_256': 0.125,\n    'mil_coatlitemini_16_224': 0.10,\n    'mil_poolformer24_16_224': 0.10,\n    'mil_swintinypatch4window7_16_224': 0.10\n}\n</code></pre>\n<p>Linear stacking is also utilized on selected models by fitting a linear regression model on training predictions and targets.</p>\n<h2>Post-processing</h2>\n<p>Predictions are averaged among patient_id groups and 0.21 is added to aggregated predictions for making the prediction 0.5 centered. That constant is selected by maximizing the OOF score.</p>\n<h2>Results</h2>\n<p>Quantitative results are provided as cross-validation, public and private leaderboard scores.<br>\nBest single model and ensemble cross-validation scores are highlighted.</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>CV ROC AUC</th>\n<th>CV Weighted Log Loss</th>\n<th>Public Leaderboard</th>\n<th>Private Leaderboard</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>DenseNet121</td>\n<td>0.6501</td>\n<td></td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>DenseNetBlur121d</td>\n<td><strong>0.6628</strong></td>\n<td></td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>DenseNet169</td>\n<td>0.6551</td>\n<td></td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>EfficientNetB2</td>\n<td>0.6501</td>\n<td></td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>EfficientNetV2 Tiny</td>\n<td>0.6500</td>\n<td></td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>CoaT Lite Mini</td>\n<td>0.6503</td>\n<td></td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>PoolFormer 24</td>\n<td>0.6536</td>\n<td></td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>Swin Transformer Tiny Patch 4 Window 7</td>\n<td>0.6497</td>\n<td></td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>Blend + patient_id Aggregation + Adjustment</td>\n<td>0.6822</td>\n<td>0.6471</td>\n<td>0.6</td>\n<td><strong>0.67897</strong></td>\n</tr>\n<tr>\n<td>Stack + patient_id Aggregation + Adjustment</td>\n<td><strong>0.6842</strong></td>\n<td><strong>0.6387</strong></td>\n<td><strong>0.6</strong></td>\n<td>0.68666</td>\n</tr>\n</tbody>\n</table>\n<p>I started overfitting when I added more and more models to ensemble. My stack was also overfitting a bit but my best submission is my blend.</p>\n<h2>What didn't work</h2>\n<ul>\n<li>Decaying learning rate</li>\n<li>Freezing backbone parameters</li>\n<li>Average and max MIL pooling</li>\n<li>Lots of other timm models</li>\n<li>Binary cross-entropy loss, focal loss, weighted log loss</li>\n<li>Label smoothing</li>\n<li>Different number of instances with larger sizes</li>\n</ul>",
      "rawMarkdown": "When I first started this competition 2 weeks ago, I got RSNA-MICCAI Brain Tumor Radiogenomic Classification flashbacks. These two competitions are quite similar in terms of low signal but this one was slightly better and less random. I had some decent experiences in that competition and I applied those things here.\n\nInference Notebook: https://www.kaggle.com/code/gunesevitan/mayo-clinic-strip-ai-inference\nGitHub Repository: https://github.com/gunesevitan/mayo-clinic-strip-ai\n\n## Dataset\n\nI created a pre-computed dataset of instances for training more efficiently. Pre-computed dataset consist of 16 instances of size 1024x1024 with 3 channels (RGB) per image. Dataset creation process can be simplified to\n\n* Compress image with JPEG compression (100% JPEG quality)\n* Resize longest edge to 20.000 pixels\n* Extract non-overlapping instances and pad them with white background\n* Sort instances by their sums in descending order and take top 16 instances\n\n## Models\n\nI used Multiple Instance Learning (MIL) models for processing instances. My MIL model is slightly different than the one shared in this competition. I think the model in [this](https://www.kaggle.com/code/analokamus/a-sample-of-multi-instance-learning-model) notebook concatenates feature maps on height dimension which felt weird to me, so I tried different pooling and concatenation methods. \n\n```\nclass ConvolutionalMultiInstanceLearningModel(nn.Module):\n\n    def __init__(self, n_instances, model_name, pretrained, freeze_parameters, aggregation, head_class, head_args):\n\n        super(ConvolutionalMultiInstanceLearningModel, self).__init__()\n\n        self.backbone = timm.create_model(\n            model_name=model_name,\n            pretrained=pretrained,\n            num_classes=head_args['n_classes']\n        )\n\n        if freeze_parameters is not None:\n            # Freeze all parameters in backbone\n            if freeze_parameters == 'all':\n                for parameter in self.backbone.parameters():\n                    parameter.requires_grad = False\n            else:\n                # Freeze specified parameters in backbone\n                for group in freeze_parameters:\n                    if isinstance(self.backbone, timm.models.DenseNet):\n                        for parameter in self.backbone.features[group].parameters():\n                            parameter.requires_grad = False\n                    elif isinstance(self.backbone, timm.models.EfficientNet):\n                        for parameter in self.backbone.blocks[group].parameters():\n                            parameter.requires_grad = False\n\n        self.aggregation = aggregation\n        n_classifier_features = self.backbone.get_classifier().in_features\n        input_features = (n_classifier_features * n_instances) if self.aggregation == 'concat' else n_classifier_features\n        self.classification_head = eval(head_class)(input_features=input_features, **head_args)\n\n    def forward(self, x):\n\n        # Stack instances on batch dimension before passing input to feature extractor\n        input_batch_size, input_instance, input_channel, input_height, input_width = x.shape\n        x = x.view(input_batch_size * input_instance, input_channel, input_height, input_width)\n        x = self.backbone.forward_features(x)\n        feature_batch_size, feature_channel, feature_height, feature_width = x.shape\n\n        if self.aggregation == 'avg':\n            # Average feature maps of multiple instances\n            x = x.contiguous().view(input_batch_size, input_instance, feature_channel, feature_height, feature_width)\n            x = torch.mean(x, dim=1)\n        elif self.aggregation == 'max':\n            # Max feature maps of multiple instances\n            x = x.contiguous().view(input_batch_size, input_instance, feature_channel, feature_height, feature_width)\n            x = torch.max(x, dim=1)[0]\n        elif self.aggregation == 'logsumexp':\n            # LogSumExp feature maps of multiple instances\n            x = x.contiguous().view(input_batch_size, input_instance, feature_channel, feature_height, feature_width)\n            x = torch.logsumexp(x, dim=1)\n        elif self.aggregation == 'concat':\n            # Stack feature maps on channel dimension\n            x = x.contiguous().view(input_batch_size, input_instance * feature_channel, feature_height, feature_width)\n\n        output = self.classification_head(x)\n        return output\n\n    def __init__(self, n_instances, model_name, pretrained, freeze_parameters, aggregation, head_class, head_args):\n\n        super(ConvolutionalMultiInstanceLearningModel, self).__init__()\n\n        self.backbone = timm.create_model(\n            model_name=model_name,\n            pretrained=pretrained,\n            num_classes=head_args['n_classes']\n        )\n\n        if freeze_parameters is not None:\n            # Freeze all parameters in backbone\n            if freeze_parameters == 'all':\n                for parameter in self.backbone.parameters():\n                    parameter.requires_grad = False\n            else:\n                # Freeze specified parameters in backbone\n                for group in freeze_parameters:\n                    if isinstance(self.backbone, timm.models.DenseNet):\n                        for parameter in self.backbone.features[group].parameters():\n                            parameter.requires_grad = False\n                    elif isinstance(self.backbone, timm.models.EfficientNet):\n                        for parameter in self.backbone.blocks[group].parameters():\n                            parameter.requires_grad = False\n\n        self.aggregation = aggregation\n        n_classifier_features = self.backbone.get_classifier().in_features\n        input_features = (n_classifier_features * n_instances) if self.aggregation == 'concat' else n_classifier_features\n        self.classification_head = eval(head_class)(input_features=input_features, **head_args)\n\n    def forward(self, x):\n\n        # Stack instances on batch dimension before passing input to feature extractor\n        input_batch_size, input_instance, input_channel, input_height, input_width = x.shape\n        x = x.view(input_batch_size * input_instance, input_channel, input_height, input_width)\n        x = self.backbone.forward_features(x)\n        feature_batch_size, feature_channel, feature_height, feature_width = x.shape\n\n        if self.aggregation == 'avg':\n            # Average feature maps of multiple instances\n            x = x.contiguous().view(input_batch_size, input_instance, feature_channel, feature_height, feature_width)\n            x = torch.mean(x, dim=1)\n        elif self.aggregation == 'max':\n            # Max feature maps of multiple instances\n            x = x.contiguous().view(input_batch_size, input_instance, feature_channel, feature_height, feature_width)\n            x = torch.max(x, dim=1)[0]\n        elif self.aggregation == 'logsumexp':\n            # LogSumExp feature maps of multiple instances\n            x = x.contiguous().view(input_batch_size, input_instance, feature_channel, feature_height, feature_width)\n            x = torch.logsumexp(x, dim=1)\n        elif self.aggregation == 'concat':\n            # Stack feature maps on channel dimension\n            x = x.contiguous().view(input_batch_size, input_instance * feature_channel, feature_height, feature_width)\n\n        output = self.classification_head(x)\n        return output\n```\n\nAggregation parameter tunes the MIL pooling or aggregating part of the model.\n* `avg`: average all feature maps on instance dimension\n* `max`: max all feature maps on instance dimension (this was mentioned on many papers)\n* `logsumexp`: logsumexmp all feature maps on instance (I saw this on one paper so I added it as well)\n* `concat`: concatenate all feature maps on channel dimension\n\n`concat` worked best for me even though the feature map counts multiplied with number of instances. I was working on an attention mechanism for instance feature maps but I didn't have enough time and I eventually ditched the idea.\n\n## Validation\n\n5 folds of cross-validation is used as the validation scheme. Splits are stratified on `binary_encoded_label`, `patient_id` doesn't overlap between folds, and dataset is shuffled before splitting. Best random seed is selected by minimizing the standard deviation of `binary_encoded_label` mean between folds which ensures target means of folds are close to each other.\n\n## Training\n\nLoss Function: Negative log likelihood loss is used as the loss function. Binary labels are converted to multi-class labels by simply doing\n```\n\ninputs = torch.sigmoid(inputs)\ninputs = torch.cat(((1 - inputs), inputs), dim=-1)\n```\n\nOptimizer: AdamW with 1e-5 initial learning rate and learning rate is cycled to 1e-6 back-and-forth after every 300 steps.\n\nBatch Size: Training: 4 - Validation: 8\n\nEarly Stopping: Early Stopping is triggered after 5 epochs with no validation loss improvement\n\n## Training Augmentations\n\nTraining augmentations are applied to all instances randomly.\n\n* Resize 256x256 or 224x224\n* Luminosity standardization by 10% chance\n* Horizontal flip by 50% chance\n* Vertical flip by 50% chance\n* Random 90-degree rotation by 25% chance\n* Random hue-saturation-value by 25% chance\n* Coarse dropout or pixel dropout by 10% chance\n* Normalize by instance dataset statistics\n\n## Ensemble\n\n8 models are used in the ensemble. Those models are selected based on their ROC AUC scores. Models with OOF ROC AUC score greater than 0.65 are qualified for the ensemble. My best single model was DenseNetBlur121d which is quite interesting and worth more exploring.\n\n* DenseNet121 (ImageNet pre-trained weights) - 256x256 16 instances\n* DenseNetBlur121d (ImageNet pre-trained weights) - 256x256 16 instances\n* DenseNet169 (ImageNet pre-trained weights) - 256x256 16 instances\n* EfficientNetB2 (ImageNet pre-trained weights) - 256x256 16 instances\n* EfficientNetV2 Tiny (ImageNet pre-trained weights) - 256x256 16 instances\n* CoaT Lite Mini (ImageNet pre-trained weights) - 224x224 16 instances\n* PoolFormer 24 (ImageNet pre-trained weights) - 224x224 16 instances\n* Swin Transformer Tiny Patch 4 Window 7 (ImageNet pre-trained weights) - 224x224 16 instances\n\nBlending is utilized on selected models. Blend weights are found by trial and error.\n\n```\nblend_weights = {\n    'mil_densenet121_16_256': 0.10,\n    'mil_densenet169_16_256': 0.10,\n    'mil_densenetblur121d_16_256': 0.25,\n    'mil_efficientnetb2_16_256': 0.125,\n    'mil_efficientnetv2rwt_16_256': 0.125,\n    'mil_coatlitemini_16_224': 0.10,\n    'mil_poolformer24_16_224': 0.10,\n    'mil_swintinypatch4window7_16_224': 0.10\n}\n\n```\nLinear stacking is also utilized on selected models by fitting a linear regression model on training predictions and targets.\n\n## Post-processing\n\nPredictions are averaged among patient_id groups and 0.21 is added to aggregated predictions for making the prediction 0.5 centered. That constant is selected by maximizing the OOF score.\n\n## Results\n\nQuantitative results are provided as cross-validation, public and private leaderboard scores.\nBest single model and ensemble cross-validation scores are highlighted.\n\n|                                             | CV ROC AUC | CV Weighted Log Loss | Public Leaderboard | Private Leaderboard |\n|---------------------------------------------|------------|----------------------|--------------------|---------------------|\n| DenseNet121                                 | 0.6501     |                      |                    |                     |\n| DenseNetBlur121d                            | **0.6628** |                      |                    |                     |\n| DenseNet169                                 | 0.6551     |                      |                    |                     |\n| EfficientNetB2                              | 0.6501     |                      |                    |                     |\n| EfficientNetV2 Tiny                         | 0.6500     |                      |                    |                     |\n| CoaT Lite Mini                              | 0.6503     |                      |                    |                     |\n| PoolFormer 24                               | 0.6536     |                      |                    |                     |\n| Swin Transformer Tiny Patch 4 Window 7      | 0.6497     |                      |                    |                     |\n| Blend + patient_id Aggregation + Adjustment | 0.6822     | 0.6471               | 0.6                | **0.67897**                    |\n| Stack + patient_id Aggregation + Adjustment | **0.6842** | **0.6387**           | **0.6**            | 0.68666                    |\n\nI started overfitting when I added more and more models to ensemble. My stack was also overfitting a bit but my best submission is my blend.\n\n## What didn't work\n\n* Decaying learning rate\n* Freezing backbone parameters\n* Average and max MIL pooling\n* Lots of other timm models\n* Binary cross-entropy loss, focal loss, weighted log loss\n* Label smoothing\n* Different number of instances with larger sizes\n",
      "votes": 25
    },
    {
      "id": 1974662,
      "postDate": "2022-10-06T11:35:33.363Z",
      "content": "<p>Congratz on surviving the shake-up !</p>\n<blockquote>\n  <p>I got RSNA-MICCAI Brain Tumor Radiogenomic Classification flashbacks. These two competitions are quite similar in terms of low signal but this one was slightly better and less random.</p>\n</blockquote>\n<p>I am completely with you on this one ^</p>",
      "rawMarkdown": "Congratz on surviving the shake-up !\n\n> I got RSNA-MICCAI Brain Tumor Radiogenomic Classification flashbacks. These two competitions are quite similar in terms of low signal but this one was slightly better and less random.\n\nI am completely with you on this one ^",
      "votes": 1
    },
    {
      "id": 1974267,
      "postDate": "2022-10-06T06:43:57.523Z",
      "content": "<p>Nice work!   But I think when you train with freeze weight mode,  the BN layer also will update the running mean/var value, maybe it will affect your result.</p>",
      "rawMarkdown": "Nice work!   But I think when you train with freeze weight mode,  the BN layer also will update the running mean/var value, maybe it will affect your result.",
      "votes": 1,
      "replies": [
        {
          "id": 1974724,
          "postDate": "2022-10-06T12:09:54.870Z",
          "content": "<p>You are right and I think that's the desired way when fine-tuning from ImageNet weights.</p>",
          "rawMarkdown": "You are right and I think that's the desired way when fine-tuning from ImageNet weights."
        }
      ]
    },
    {
      "id": 1974254,
      "postDate": "2022-10-06T06:35:06.477Z",
      "content": "<blockquote>\n  <p>0.21 is added to aggregated predictions for making the prediction 0.5 centered. </p>\n</blockquote>\n<p>I also done this, however I added +10 then rescaled (guess depends how far your predictions are from mean), I tested this on a few submissions and it greatly decreased the log loss sometimes by up to 0.2. </p>",
      "rawMarkdown": "> 0.21 is added to aggregated predictions for making the prediction 0.5 centered. \n\nI also done this, however I added +10 then rescaled (guess depends how far your predictions are from mean), I tested this on a few submissions and it greatly decreased the log loss sometimes by up to 0.2. ",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 1974662,
      "author_name": "Theo Viel",
      "author_url": "",
      "post_date": "2022-10-06T11:35:33.363000",
      "content": "<p>Congratz on surviving the shake-up !</p>\n<blockquote>\n  <p>I got RSNA-MICCAI Brain Tumor Radiogenomic Classification flashbacks. These two competitions are quite similar in terms of low signal but this one was slightly better and less random.</p>\n</blockquote>\n<p>I am completely with you on this one ^</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1974267,
      "author_name": "KKY",
      "author_url": "",
      "post_date": "2022-10-06T06:43:57.523000",
      "content": "<p>Nice work!   But I think when you train with freeze weight mode,  the BN layer also will update the running mean/var value, maybe it will affect your result.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1974724,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2022-10-06T12:09:54.870000",
          "content": "<p>You are right and I think that's the desired way when fine-tuning from ImageNet weights.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1974254,
      "author_name": "JM",
      "author_url": "",
      "post_date": "2022-10-06T06:35:06.477000",
      "content": "<blockquote>\n  <p>0.21 is added to aggregated predictions for making the prediction 0.5 centered. </p>\n</blockquote>\n<p>I also done this, however I added +10 then rescaled (guess depends how far your predictions are from mean), I tested this on a few submissions and it greatly decreased the log loss sometimes by up to 0.2. </p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1974097": "When I first started this competition 2 weeks ago, I got RSNA-MICCAI Brain Tumor Radiogenomic Classification flashbacks. These two competitions are quite similar in terms of low signal but this one was slightly better and less random. I had some decent experiences in that competition and I applied those things here.\n\nInference Notebook: https://www.kaggle.com/code/gunesevitan/mayo-clinic-strip-ai-inference\nGitHub Repository: https://github.com/gunesevitan/mayo-clinic-strip-ai\n\n## Dataset\n\nI created a pre-computed dataset of instances for training more efficiently. Pre-computed dataset consist of 16 instances of size 1024x1024 with 3 channels (RGB) per image. Dataset creation process can be simplified to\n\n* Compress image with JPEG compression (100% JPEG quality)\n* Resize longest edge to 20.000 pixels\n* Extract non-overlapping instances and pad them with white background\n* Sort instances by their sums in descending order and take top 16 instances\n\n## Models\n\nI used Multiple Instance Learning (MIL) models for processing instances. My MIL model is slightly different than the one shared in this competition. I think the model in [this](https://www.kaggle.com/code/analokamus/a-sample-of-multi-instance-learning-model) notebook concatenates feature maps on height dimension which felt weird to me, so I tried different pooling and concatenation methods. \n\n```\nclass ConvolutionalMultiInstanceLearningModel(nn.Module):\n\n    def __init__(self, n_instances, model_name, pretrained, freeze_parameters, aggregation, head_class, head_args):\n\n        super(ConvolutionalMultiInstanceLearningModel, self).__init__()\n\n        self.backbone = timm.create_model(\n            model_name=model_name,\n            pretrained=pretrained,\n            num_classes=head_args['n_classes']\n        )\n\n        if freeze_parameters is not None:\n            # Freeze all parameters in backbone\n            if freeze_parameters == 'all':\n                for parameter in self.backbone.parameters():\n                    parameter.requires_grad = False\n            else:\n                # Freeze specified parameters in backbone\n                for group in freeze_parameters:\n                    if isinstance(self.backbone, timm.models.DenseNet):\n                        for parameter in self.backbone.features[group].parameters():\n                            parameter.requires_grad = False\n                    elif isinstance(self.backbone, timm.models.EfficientNet):\n                        for parameter in self.backbone.blocks[group].parameters():\n                            parameter.requires_grad = False\n\n        self.aggregation = aggregation\n        n_classifier_features = self.backbone.get_classifier().in_features\n        input_features = (n_classifier_features * n_instances) if self.aggregation == 'concat' else n_classifier_features\n        self.classification_head = eval(head_class)(input_features=input_features, **head_args)\n\n    def forward(self, x):\n\n        # Stack instances on batch dimension before passing input to feature extractor\n        input_batch_size, input_instance, input_channel, input_height, input_width = x.shape\n        x = x.view(input_batch_size * input_instance, input_channel, input_height, input_width)\n        x = self.backbone.forward_features(x)\n        feature_batch_size, feature_channel, feature_height, feature_width = x.shape\n\n        if self.aggregation == 'avg':\n            # Average feature maps of multiple instances\n            x = x.contiguous().view(input_batch_size, input_instance, feature_channel, feature_height, feature_width)\n            x = torch.mean(x, dim=1)\n        elif self.aggregation == 'max':\n            # Max feature maps of multiple instances\n            x = x.contiguous().view(input_batch_size, input_instance, feature_channel, feature_height, feature_width)\n            x = torch.max(x, dim=1)[0]\n        elif self.aggregation == 'logsumexp':\n            # LogSumExp feature maps of multiple instances\n            x = x.contiguous().view(input_batch_size, input_instance, feature_channel, feature_height, feature_width)\n            x = torch.logsumexp(x, dim=1)\n        elif self.aggregation == 'concat':\n            # Stack feature maps on channel dimension\n            x = x.contiguous().view(input_batch_size, input_instance * feature_channel, feature_height, feature_width)\n\n        output = self.classification_head(x)\n        return output\n\n    def __init__(self, n_instances, model_name, pretrained, freeze_parameters, aggregation, head_class, head_args):\n\n        super(ConvolutionalMultiInstanceLearningModel, self).__init__()\n\n        self.backbone = timm.create_model(\n            model_name=model_name,\n            pretrained=pretrained,\n            num_classes=head_args['n_classes']\n        )\n\n        if freeze_parameters is not None:\n            # Freeze all parameters in backbone\n            if freeze_parameters == 'all':\n                for parameter in self.backbone.parameters():\n                    parameter.requires_grad = False\n            else:\n                # Freeze specified parameters in backbone\n                for group in freeze_parameters:\n                    if isinstance(self.backbone, timm.models.DenseNet):\n                        for parameter in self.backbone.features[group].parameters():\n                            parameter.requires_grad = False\n                    elif isinstance(self.backbone, timm.models.EfficientNet):\n                        for parameter in self.backbone.blocks[group].parameters():\n                            parameter.requires_grad = False\n\n        self.aggregation = aggregation\n        n_classifier_features = self.backbone.get_classifier().in_features\n        input_features = (n_classifier_features * n_instances) if self.aggregation == 'concat' else n_classifier_features\n        self.classification_head = eval(head_class)(input_features=input_features, **head_args)\n\n    def forward(self, x):\n\n        # Stack instances on batch dimension before passing input to feature extractor\n        input_batch_size, input_instance, input_channel, input_height, input_width = x.shape\n        x = x.view(input_batch_size * input_instance, input_channel, input_height, input_width)\n        x = self.backbone.forward_features(x)\n        feature_batch_size, feature_channel, feature_height, feature_width = x.shape\n\n        if self.aggregation == 'avg':\n            # Average feature maps of multiple instances\n            x = x.contiguous().view(input_batch_size, input_instance, feature_channel, feature_height, feature_width)\n            x = torch.mean(x, dim=1)\n        elif self.aggregation == 'max':\n            # Max feature maps of multiple instances\n            x = x.contiguous().view(input_batch_size, input_instance, feature_channel, feature_height, feature_width)\n            x = torch.max(x, dim=1)[0]\n        elif self.aggregation == 'logsumexp':\n            # LogSumExp feature maps of multiple instances\n            x = x.contiguous().view(input_batch_size, input_instance, feature_channel, feature_height, feature_width)\n            x = torch.logsumexp(x, dim=1)\n        elif self.aggregation == 'concat':\n            # Stack feature maps on channel dimension\n            x = x.contiguous().view(input_batch_size, input_instance * feature_channel, feature_height, feature_width)\n\n        output = self.classification_head(x)\n        return output\n```\n\nAggregation parameter tunes the MIL pooling or aggregating part of the model.\n* `avg`: average all feature maps on instance dimension\n* `max`: max all feature maps on instance dimension (this was mentioned on many papers)\n* `logsumexp`: logsumexmp all feature maps on instance (I saw this on one paper so I added it as well)\n* `concat`: concatenate all feature maps on channel dimension\n\n`concat` worked best for me even though the feature map counts multiplied with number of instances. I was working on an attention mechanism for instance feature maps but I didn't have enough time and I eventually ditched the idea.\n\n## Validation\n\n5 folds of cross-validation is used as the validation scheme. Splits are stratified on `binary_encoded_label`, `patient_id` doesn't overlap between folds, and dataset is shuffled before splitting. Best random seed is selected by minimizing the standard deviation of `binary_encoded_label` mean between folds which ensures target means of folds are close to each other.\n\n## Training\n\nLoss Function: Negative log likelihood loss is used as the loss function. Binary labels are converted to multi-class labels by simply doing\n```\n\ninputs = torch.sigmoid(inputs)\ninputs = torch.cat(((1 - inputs), inputs), dim=-1)\n```\n\nOptimizer: AdamW with 1e-5 initial learning rate and learning rate is cycled to 1e-6 back-and-forth after every 300 steps.\n\nBatch Size: Training: 4 - Validation: 8\n\nEarly Stopping: Early Stopping is triggered after 5 epochs with no validation loss improvement\n\n## Training Augmentations\n\nTraining augmentations are applied to all instances randomly.\n\n* Resize 256x256 or 224x224\n* Luminosity standardization by 10% chance\n* Horizontal flip by 50% chance\n* Vertical flip by 50% chance\n* Random 90-degree rotation by 25% chance\n* Random hue-saturation-value by 25% chance\n* Coarse dropout or pixel dropout by 10% chance\n* Normalize by instance dataset statistics\n\n## Ensemble\n\n8 models are used in the ensemble. Those models are selected based on their ROC AUC scores. Models with OOF ROC AUC score greater than 0.65 are qualified for the ensemble. My best single model was DenseNetBlur121d which is quite interesting and worth more exploring.\n\n* DenseNet121 (ImageNet pre-trained weights) - 256x256 16 instances\n* DenseNetBlur121d (ImageNet pre-trained weights) - 256x256 16 instances\n* DenseNet169 (ImageNet pre-trained weights) - 256x256 16 instances\n* EfficientNetB2 (ImageNet pre-trained weights) - 256x256 16 instances\n* EfficientNetV2 Tiny (ImageNet pre-trained weights) - 256x256 16 instances\n* CoaT Lite Mini (ImageNet pre-trained weights) - 224x224 16 instances\n* PoolFormer 24 (ImageNet pre-trained weights) - 224x224 16 instances\n* Swin Transformer Tiny Patch 4 Window 7 (ImageNet pre-trained weights) - 224x224 16 instances\n\nBlending is utilized on selected models. Blend weights are found by trial and error.\n\n```\nblend_weights = {\n    'mil_densenet121_16_256': 0.10,\n    'mil_densenet169_16_256': 0.10,\n    'mil_densenetblur121d_16_256': 0.25,\n    'mil_efficientnetb2_16_256': 0.125,\n    'mil_efficientnetv2rwt_16_256': 0.125,\n    'mil_coatlitemini_16_224': 0.10,\n    'mil_poolformer24_16_224': 0.10,\n    'mil_swintinypatch4window7_16_224': 0.10\n}\n\n```\nLinear stacking is also utilized on selected models by fitting a linear regression model on training predictions and targets.\n\n## Post-processing\n\nPredictions are averaged among patient_id groups and 0.21 is added to aggregated predictions for making the prediction 0.5 centered. That constant is selected by maximizing the OOF score.\n\n## Results\n\nQuantitative results are provided as cross-validation, public and private leaderboard scores.\nBest single model and ensemble cross-validation scores are highlighted.\n\n|                                             | CV ROC AUC | CV Weighted Log Loss | Public Leaderboard | Private Leaderboard |\n|---------------------------------------------|------------|----------------------|--------------------|---------------------|\n| DenseNet121                                 | 0.6501     |                      |                    |                     |\n| DenseNetBlur121d                            | **0.6628** |                      |                    |                     |\n| DenseNet169                                 | 0.6551     |                      |                    |                     |\n| EfficientNetB2                              | 0.6501     |                      |                    |                     |\n| EfficientNetV2 Tiny                         | 0.6500     |                      |                    |                     |\n| CoaT Lite Mini                              | 0.6503     |                      |                    |                     |\n| PoolFormer 24                               | 0.6536     |                      |                    |                     |\n| Swin Transformer Tiny Patch 4 Window 7      | 0.6497     |                      |                    |                     |\n| Blend + patient_id Aggregation + Adjustment | 0.6822     | 0.6471               | 0.6                | **0.67897**                    |\n| Stack + patient_id Aggregation + Adjustment | **0.6842** | **0.6387**           | **0.6**            | 0.68666                    |\n\nI started overfitting when I added more and more models to ensemble. My stack was also overfitting a bit but my best submission is my blend.\n\n## What didn't work\n\n* Decaying learning rate\n* Freezing backbone parameters\n* Average and max MIL pooling\n* Lots of other timm models\n* Binary cross-entropy loss, focal loss, weighted log loss\n* Label smoothing\n* Different number of instances with larger sizes\n",
    "1974662": "Congratz on surviving the shake-up !\n\n> I got RSNA-MICCAI Brain Tumor Radiogenomic Classification flashbacks. These two competitions are quite similar in terms of low signal but this one was slightly better and less random.\n\nI am completely with you on this one ^",
    "1974267": "Nice work!   But I think when you train with freeze weight mode,  the BN layer also will update the running mean/var value, maybe it will affect your result.",
    "1974254": "> 0.21 is added to aggregated predictions for making the prediction 0.5 centered. \n\nI also done this, however I added +10 then rescaled (guess depends how far your predictions are from mean), I tested this on a few submissions and it greatly decreased the log loss sometimes by up to 0.2. "
  }
}