{
  "id": 432282,
  "title": "30th Place Solution for the Google Research - Identify Contrails to Reduce Global Warming Competition",
  "url": "/competitions/google-research-identify-contrails-reduce-global-warming/discussion/432282",
  "author_name": "delai50",
  "post_date": "2023-08-16T18:55:49.505000",
  "votes": 8,
  "comment_count": 0,
  "views": 0,
  "content": "<p>First of all, we would like to thank Kaggle and the host for such a funny (and <em>nothing-works-well-for-us</em>) competition. This is not a fancy solution but we feel proud of taking the right steps to generalize well to the private test set (it was indeed, 42 positions shake up 😃).</p>\n<h1>1. <strong>Context</strong></h1>\n<ul>\n<li><em>Business context</em>: <a href=\"https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/overview\" target=\"_blank\">https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/overview</a></li>\n<li><em>Data context</em>: <a href=\"https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/data\" target=\"_blank\">https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/data</a></li>\n</ul>\n<h1>2. <strong>Overview of the Approach</strong></h1>\n<p>Our final solution was an ensemble of 8 models which combines different encoders and image sizes. Each model configuration was trained in turn 3 times with different random seeds (24 models in total) and always with all available data (train+validation).</p>\n<h3><strong>Data preprocessing</strong></h3>\n<p>ASH false color images, only 2D models with labeled frame and hard labels.</p>\n<h3><strong>Validation schema</strong></h3>\n<p>Provided train/validation split.</p>\n<h1>3. <strong>Overview of the Approach</strong></h1>\n<h3>3.1. <strong>Individual models and final ensemble</strong></h3>\n<p>All models are based on UNet architecture from SMP. We combine them using a weighted average of the probabilities dropped by single models, where weights were determined optimizing the global dice score on the validation set using Optuna. The final CV score was <strong>0.6825</strong> (which may be a little biased because as we mentioned we optimized the weights directly using the validation set).</p>\n<table>\n<thead>\n<tr>\n<th>Weight</th>\n<th>Encoder</th>\n<th>Image size</th>\n<th>LR</th>\n<th>Epochs</th>\n<th>BS</th>\n<th>CV score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.02</td>\n<td>timm-resnest101e</td>\n<td>384</td>\n<td>5e-4</td>\n<td>20</td>\n<td>48</td>\n<td>0.64397</td>\n</tr>\n<tr>\n<td>0.07</td>\n<td>maxvit_base_tf_384.in21k_ft_in1k</td>\n<td>384</td>\n<td>1e-4</td>\n<td>20</td>\n<td>16</td>\n<td>0.65699</td>\n</tr>\n<tr>\n<td>0.18</td>\n<td>maxvit_base_tf_512</td>\n<td>512</td>\n<td>1e-4</td>\n<td>40</td>\n<td>8*</td>\n<td>0.66414</td>\n</tr>\n<tr>\n<td>0.02</td>\n<td>tf_efficientnetv2_b3.in1k</td>\n<td>768</td>\n<td>5e-4</td>\n<td>35</td>\n<td>24*</td>\n<td>0.65178</td>\n</tr>\n<tr>\n<td>0.23</td>\n<td>tf_efficientnetv2_m.in21k_ft_in1k</td>\n<td>512</td>\n<td>5e-4</td>\n<td>35</td>\n<td>24*</td>\n<td>0.6525</td>\n</tr>\n<tr>\n<td>0.14</td>\n<td>tf_efficientnetv2_l.in21k_ft_in1k</td>\n<td>384</td>\n<td>5e-4</td>\n<td>35</td>\n<td>24*</td>\n<td>0.64538</td>\n</tr>\n<tr>\n<td>0.11</td>\n<td>maxvit_large_tf_224.in1k</td>\n<td>448</td>\n<td>1e-4</td>\n<td>40</td>\n<td>8*</td>\n<td>0.66407</td>\n</tr>\n<tr>\n<td>0.23</td>\n<td>maxvit_large_tf_384.in1k</td>\n<td>384</td>\n<td>1e-4</td>\n<td>40</td>\n<td>8*</td>\n<td>0.66782</td>\n</tr>\n</tbody>\n</table>\n<p>*gradient accumulation of 2.</p>\n<h3>3.2. <strong>Threshold selection</strong></h3>\n<p>The optimal threshold for the ensemble was 0.43, which was determined running a for loop varying the threshold from 0 to 1 with 0.1 increments and calculating the corresponding global dice. We observed a relatively wide maximum peaking at 0.43 which give us enough confidence to choose that threshold.</p>\n<h3>3.3. <strong>Training procedure</strong></h3>\n<p>We used CosineAnnealingLR and AdamW. Maybe it is important to note that we always tried to choose the number of epochs such that the last epoch is the best one. This generalizes better than choosing the best checkpoint during training, especially in our case, where we observed an irregular behavior of the training curves. Finally, once a model was validated, we took the same configuration and train it again using all the available data.</p>\n<h3>3.4. <strong>Augmentations</strong></h3>\n<p>We didn't notice the mask shift, so flipping and rotation didn't work for us. Other augmentations that worked were:</p>\n<pre><code>A.Compose([  \n     A.RandomResizedCrop(height=, width=, p=),  \n     A.RandomBrightnessContrast(p=)  \n])  \nCutMix(p=)  \nMixUp(p=)\n</code></pre>\n<p>These are our implementations of CutMix and MixUp:</p>\n<pre><code> ():\n    W = size[]\n    H = size[]\n    cut_rat = np.sqrt( - lam)\n    cut_w = (W * cut_rat)\n    cut_h = (H * cut_rat)\n\n    \n    cx = np.random.randint(W)\n    cy = np.random.randint(H)\n\n    bbx1 = np.clip(cx - cut_w // , , W)\n    bby1 = np.clip(cy - cut_h // , , H)\n    bbx2 = np.clip(cx + cut_w // , , W)\n    bby2 = np.clip(cy + cut_h // , , H)\n     bbx1, bby1, bbx2, bby2\n\n ():\n    indices = torch.randperm(data.size())\n    \n    shuffled_target = target[indices]\n\n    lam = np.clip(np.random.beta(alpha, alpha),,)\n    bbx1, bby1, bbx2, bby2 = rand_bbox(data.size(), lam)\n    new_data = data.clone()\n    new_data[:, :, bby1:bby2, bbx1:bbx2] = data[indices, :, bby1:bby2, bbx1:bbx2]\n    \n    lam =  - ((bbx2 - bbx1) * (bby2 - bby1) / (data.size()[-] * data.size()[-]))\n    \n     new_data, target, shuffled_target, lam\n\n ():\n     alpha &gt; , \n     x.size() &gt; , \n\n    lam = np.random.beta(alpha, alpha)\n    rand_index = torch.randperm(x.size()[])\n    mixed_x = lam * x + ( - lam) * x[rand_index, :]\n    target_a, target_b = y, y[rand_index]\n     mixed_x, target_a, target_b, lam\n</code></pre>\n<p>And losses are calculated as:</p>\n<pre><code> self.cfg.mixup  torch.rand()[] &lt; self.cfg.mixup_p:\n    mix_images, target_a, target_b, lam = mixup(images, masks, alpha=self.cfg.mixup_alpha)\n    logits = self(mix_images)\n    loss = self.loss_fn(logits, target_a) * lam + ( - lam) * self.loss_fn(logits, target_b)\n\n self.cfg.cutmix  torch.rand()[] &lt; self.cfg.cutmix_p:\n    mix_images, target_a, target_b, lam = cutmix(images, masks, alpha=self.cfg.cutmix_alpha)\n    logits = self(mix_images)\n    loss = self.loss_fn(logits, target_a) * lam + ( - lam) * self.loss_fn(logits, target_b)\n</code></pre>\n<h3>3.5. <strong>What didn't work</strong></h3>\n<p>A ton of things:</p>\n<ul>\n<li>Using other bands/other data schema.</li>\n<li>TTA.</li>\n<li>We tried to combine the former and the latter frames with the labeled one at each stage of the UNet without success (only tried concatenation).</li>\n<li>At the end of the competition we discovered that pseudolabels give a big boost for some models (not the strongest ones) but we didn't have enough time to properly experiment with it.</li>\n<li>…</li>\n</ul>\n<h1><strong>References</strong></h1>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/inversion/visualizing-contrails\" target=\"_blank\">ASH color scheme</a></li>\n<li><a href=\"https://arxiv.org/abs/2304.02122\" target=\"_blank\">OpenContrails: Benchmarking Contrail Detection on GOES-16 ABI</a></li>\n<li><a href=\"https://www.kaggle.com/egortrushin\" target=\"_blank\">@egortrushin</a> <a href=\"https://www.kaggle.com/code/egortrushin/gr-icrgw-training-with-4-folds\" target=\"_blank\">public notebook</a> was our starting point.</li>\n</ul>\n<h1><strong>Acknowledgments</strong></h1>\n<p>To my wonderful teammates <a href=\"https://www.kaggle.com/maxdiazbattan\" target=\"_blank\">@maxdiazbattan</a> <a href=\"https://www.kaggle.com/edomingo\" target=\"_blank\">@edomingo</a> <a href=\"https://www.kaggle.com/juansensio\" target=\"_blank\">@juansensio</a> (if I forgot something please feel free to write it in the comments).</p>",
  "messages": [
    {
      "id": 2394213,
      "postDate": "2023-08-16T18:55:49.507Z",
      "content": "<p>First of all, we would like to thank Kaggle and the host for such a funny (and <em>nothing-works-well-for-us</em>) competition. This is not a fancy solution but we feel proud of taking the right steps to generalize well to the private test set (it was indeed, 42 positions shake up 😃).</p>\n<h1>1. <strong>Context</strong></h1>\n<ul>\n<li><em>Business context</em>: <a href=\"https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/overview\" target=\"_blank\">https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/overview</a></li>\n<li><em>Data context</em>: <a href=\"https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/data\" target=\"_blank\">https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/data</a></li>\n</ul>\n<h1>2. <strong>Overview of the Approach</strong></h1>\n<p>Our final solution was an ensemble of 8 models which combines different encoders and image sizes. Each model configuration was trained in turn 3 times with different random seeds (24 models in total) and always with all available data (train+validation).</p>\n<h3><strong>Data preprocessing</strong></h3>\n<p>ASH false color images, only 2D models with labeled frame and hard labels.</p>\n<h3><strong>Validation schema</strong></h3>\n<p>Provided train/validation split.</p>\n<h1>3. <strong>Overview of the Approach</strong></h1>\n<h3>3.1. <strong>Individual models and final ensemble</strong></h3>\n<p>All models are based on UNet architecture from SMP. We combine them using a weighted average of the probabilities dropped by single models, where weights were determined optimizing the global dice score on the validation set using Optuna. The final CV score was <strong>0.6825</strong> (which may be a little biased because as we mentioned we optimized the weights directly using the validation set).</p>\n<table>\n<thead>\n<tr>\n<th>Weight</th>\n<th>Encoder</th>\n<th>Image size</th>\n<th>LR</th>\n<th>Epochs</th>\n<th>BS</th>\n<th>CV score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.02</td>\n<td>timm-resnest101e</td>\n<td>384</td>\n<td>5e-4</td>\n<td>20</td>\n<td>48</td>\n<td>0.64397</td>\n</tr>\n<tr>\n<td>0.07</td>\n<td>maxvit_base_tf_384.in21k_ft_in1k</td>\n<td>384</td>\n<td>1e-4</td>\n<td>20</td>\n<td>16</td>\n<td>0.65699</td>\n</tr>\n<tr>\n<td>0.18</td>\n<td>maxvit_base_tf_512</td>\n<td>512</td>\n<td>1e-4</td>\n<td>40</td>\n<td>8*</td>\n<td>0.66414</td>\n</tr>\n<tr>\n<td>0.02</td>\n<td>tf_efficientnetv2_b3.in1k</td>\n<td>768</td>\n<td>5e-4</td>\n<td>35</td>\n<td>24*</td>\n<td>0.65178</td>\n</tr>\n<tr>\n<td>0.23</td>\n<td>tf_efficientnetv2_m.in21k_ft_in1k</td>\n<td>512</td>\n<td>5e-4</td>\n<td>35</td>\n<td>24*</td>\n<td>0.6525</td>\n</tr>\n<tr>\n<td>0.14</td>\n<td>tf_efficientnetv2_l.in21k_ft_in1k</td>\n<td>384</td>\n<td>5e-4</td>\n<td>35</td>\n<td>24*</td>\n<td>0.64538</td>\n</tr>\n<tr>\n<td>0.11</td>\n<td>maxvit_large_tf_224.in1k</td>\n<td>448</td>\n<td>1e-4</td>\n<td>40</td>\n<td>8*</td>\n<td>0.66407</td>\n</tr>\n<tr>\n<td>0.23</td>\n<td>maxvit_large_tf_384.in1k</td>\n<td>384</td>\n<td>1e-4</td>\n<td>40</td>\n<td>8*</td>\n<td>0.66782</td>\n</tr>\n</tbody>\n</table>\n<p>*gradient accumulation of 2.</p>\n<h3>3.2. <strong>Threshold selection</strong></h3>\n<p>The optimal threshold for the ensemble was 0.43, which was determined running a for loop varying the threshold from 0 to 1 with 0.1 increments and calculating the corresponding global dice. We observed a relatively wide maximum peaking at 0.43 which give us enough confidence to choose that threshold.</p>\n<h3>3.3. <strong>Training procedure</strong></h3>\n<p>We used CosineAnnealingLR and AdamW. Maybe it is important to note that we always tried to choose the number of epochs such that the last epoch is the best one. This generalizes better than choosing the best checkpoint during training, especially in our case, where we observed an irregular behavior of the training curves. Finally, once a model was validated, we took the same configuration and train it again using all the available data.</p>\n<h3>3.4. <strong>Augmentations</strong></h3>\n<p>We didn't notice the mask shift, so flipping and rotation didn't work for us. Other augmentations that worked were:</p>\n<pre><code>A.Compose([  \n     A.RandomResizedCrop(height=, width=, p=),  \n     A.RandomBrightnessContrast(p=)  \n])  \nCutMix(p=)  \nMixUp(p=)\n</code></pre>\n<p>These are our implementations of CutMix and MixUp:</p>\n<pre><code> ():\n    W = size[]\n    H = size[]\n    cut_rat = np.sqrt( - lam)\n    cut_w = (W * cut_rat)\n    cut_h = (H * cut_rat)\n\n    \n    cx = np.random.randint(W)\n    cy = np.random.randint(H)\n\n    bbx1 = np.clip(cx - cut_w // , , W)\n    bby1 = np.clip(cy - cut_h // , , H)\n    bbx2 = np.clip(cx + cut_w // , , W)\n    bby2 = np.clip(cy + cut_h // , , H)\n     bbx1, bby1, bbx2, bby2\n\n ():\n    indices = torch.randperm(data.size())\n    \n    shuffled_target = target[indices]\n\n    lam = np.clip(np.random.beta(alpha, alpha),,)\n    bbx1, bby1, bbx2, bby2 = rand_bbox(data.size(), lam)\n    new_data = data.clone()\n    new_data[:, :, bby1:bby2, bbx1:bbx2] = data[indices, :, bby1:bby2, bbx1:bbx2]\n    \n    lam =  - ((bbx2 - bbx1) * (bby2 - bby1) / (data.size()[-] * data.size()[-]))\n    \n     new_data, target, shuffled_target, lam\n\n ():\n     alpha &gt; , \n     x.size() &gt; , \n\n    lam = np.random.beta(alpha, alpha)\n    rand_index = torch.randperm(x.size()[])\n    mixed_x = lam * x + ( - lam) * x[rand_index, :]\n    target_a, target_b = y, y[rand_index]\n     mixed_x, target_a, target_b, lam\n</code></pre>\n<p>And losses are calculated as:</p>\n<pre><code> self.cfg.mixup  torch.rand()[] &lt; self.cfg.mixup_p:\n    mix_images, target_a, target_b, lam = mixup(images, masks, alpha=self.cfg.mixup_alpha)\n    logits = self(mix_images)\n    loss = self.loss_fn(logits, target_a) * lam + ( - lam) * self.loss_fn(logits, target_b)\n\n self.cfg.cutmix  torch.rand()[] &lt; self.cfg.cutmix_p:\n    mix_images, target_a, target_b, lam = cutmix(images, masks, alpha=self.cfg.cutmix_alpha)\n    logits = self(mix_images)\n    loss = self.loss_fn(logits, target_a) * lam + ( - lam) * self.loss_fn(logits, target_b)\n</code></pre>\n<h3>3.5. <strong>What didn't work</strong></h3>\n<p>A ton of things:</p>\n<ul>\n<li>Using other bands/other data schema.</li>\n<li>TTA.</li>\n<li>We tried to combine the former and the latter frames with the labeled one at each stage of the UNet without success (only tried concatenation).</li>\n<li>At the end of the competition we discovered that pseudolabels give a big boost for some models (not the strongest ones) but we didn't have enough time to properly experiment with it.</li>\n<li>…</li>\n</ul>\n<h1><strong>References</strong></h1>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/inversion/visualizing-contrails\" target=\"_blank\">ASH color scheme</a></li>\n<li><a href=\"https://arxiv.org/abs/2304.02122\" target=\"_blank\">OpenContrails: Benchmarking Contrail Detection on GOES-16 ABI</a></li>\n<li><a href=\"https://www.kaggle.com/egortrushin\" target=\"_blank\">@egortrushin</a> <a href=\"https://www.kaggle.com/code/egortrushin/gr-icrgw-training-with-4-folds\" target=\"_blank\">public notebook</a> was our starting point.</li>\n</ul>\n<h1><strong>Acknowledgments</strong></h1>\n<p>To my wonderful teammates <a href=\"https://www.kaggle.com/maxdiazbattan\" target=\"_blank\">@maxdiazbattan</a> <a href=\"https://www.kaggle.com/edomingo\" target=\"_blank\">@edomingo</a> <a href=\"https://www.kaggle.com/juansensio\" target=\"_blank\">@juansensio</a> (if I forgot something please feel free to write it in the comments).</p>",
      "rawMarkdown": "First of all, we would like to thank Kaggle and the host for such a funny (and *nothing-works-well-for-us*) competition. This is not a fancy solution but we feel proud of taking the right steps to generalize well to the private test set (it was indeed, 42 positions shake up 😃).\n\n# 1. **Context**\n- *Business context*: https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/overview\n- *Data context*: https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/data\n\n# 2. **Overview of the Approach**\nOur final solution was an ensemble of 8 models which combines different encoders and image sizes. Each model configuration was trained in turn 3 times with different random seeds (24 models in total) and always with all available data (train+validation).\n\n### **Data preprocessing**\nASH false color images, only 2D models with labeled frame and hard labels.\n\n### **Validation schema**\nProvided train/validation split.\n\n# 3. **Overview of the Approach**\n\n### 3.1. **Individual models and final ensemble**\nAll models are based on UNet architecture from SMP. We combine them using a weighted average of the probabilities dropped by single models, where weights were determined optimizing the global dice score on the validation set using Optuna. The final CV score was **0.6825** (which may be a little biased because as we mentioned we optimized the weights directly using the validation set).\n\n| Weight | Encoder| Image size | LR | Epochs | BS | CV score |\n| --- | --- | --- | --- | --- | --- | --- |\n| 0.02 | timm-resnest101e | 384 | 5e-4 | 20 | 48 | 0.64397\n| 0.07 | maxvit_base_tf_384.in21k_ft_in1k | 384 | 1e-4 | 20 | 16 | 0.65699 \n| 0.18 | maxvit_base_tf_512 | 512 | 1e-4 | 40 | 8* | 0.66414\n| 0.02 | tf_efficientnetv2_b3.in1k | 768 | 5e-4 | 35 | 24* | 0.65178\n| 0.23 | tf_efficientnetv2_m.in21k_ft_in1k | 512 | 5e-4 | 35 | 24* | 0.6525\n| 0.14 | tf_efficientnetv2_l.in21k_ft_in1k | 384 | 5e-4 | 35 | 24* | 0.64538\n| 0.11 | maxvit_large_tf_224.in1k | 448 | 1e-4 | 40 | 8* | 0.66407\n| 0.23 | maxvit_large_tf_384.in1k | 384 | 1e-4 | 40 | 8* | 0.66782\n\n*gradient accumulation of 2.\n \n### 3.2. **Threshold selection**\nThe optimal threshold for the ensemble was 0.43, which was determined running a for loop varying the threshold from 0 to 1 with 0.1 increments and calculating the corresponding global dice. We observed a relatively wide maximum peaking at 0.43 which give us enough confidence to choose that threshold.\n\n### 3.3. **Training procedure**\nWe used CosineAnnealingLR and AdamW. Maybe it is important to note that we always tried to choose the number of epochs such that the last epoch is the best one. This generalizes better than choosing the best checkpoint during training, especially in our case, where we observed an irregular behavior of the training curves. Finally, once a model was validated, we took the same configuration and train it again using all the available data.\n\n### 3.4. **Augmentations**\nWe didn't notice the mask shift, so flipping and rotation didn't work for us. Other augmentations that worked were:\n\n```python\nA.Compose([  \n     A.RandomResizedCrop(height=256, width=256, p=0.5),  \n     A.RandomBrightnessContrast(p=0.5)  \n])  \nCutMix(p=0.5)  \nMixUp(p=0.5)\n```\n\nThese are our implementations of CutMix and MixUp:\n```python\ndef rand_bbox(size, lam):\n    W = size[2]\n    H = size[3]\n    cut_rat = np.sqrt(1. - lam)\n    cut_w = int(W * cut_rat)\n    cut_h = int(H * cut_rat)\n\n    # uniform\n    cx = np.random.randint(W)\n    cy = np.random.randint(H)\n\n    bbx1 = np.clip(cx - cut_w // 2, 0, W)\n    bby1 = np.clip(cy - cut_h // 2, 0, H)\n    bbx2 = np.clip(cx + cut_w // 2, 0, W)\n    bby2 = np.clip(cy + cut_h // 2, 0, H)\n    return bbx1, bby1, bbx2, bby2\n\ndef cutmix(data, target, alpha):\n    indices = torch.randperm(data.size(0))\n    # shuffled_data = data[indices]\n    shuffled_target = target[indices]\n\n    lam = np.clip(np.random.beta(alpha, alpha),0.3,0.4)\n    bbx1, bby1, bbx2, bby2 = rand_bbox(data.size(), lam)\n    new_data = data.clone()\n    new_data[:, :, bby1:bby2, bbx1:bbx2] = data[indices, :, bby1:bby2, bbx1:bbx2]\n    # adjust lambda to exactly match pixel ratio\n    lam = 1 - ((bbx2 - bbx1) * (bby2 - bby1) / (data.size()[-1] * data.size()[-2]))\n    # targets = (target, shuffled_target, lam)\n    return new_data, target, shuffled_target, lam\n\ndef mixup(x: torch.Tensor, y: torch.Tensor, alpha: float = 1.0):\n    assert alpha > 0, \"alpha should be larger than 0\"\n    assert x.size(0) > 1, \"Mixup cannot be applied to a single instance.\"\n\n    lam = np.random.beta(alpha, alpha)\n    rand_index = torch.randperm(x.size()[0])\n    mixed_x = lam * x + (1 - lam) * x[rand_index, :]\n    target_a, target_b = y, y[rand_index]\n    return mixed_x, target_a, target_b, lam\n```\nAnd losses are calculated as:\n```\nif self.cfg.mixup and torch.rand(1)[0] < self.cfg.mixup_p:\n    mix_images, target_a, target_b, lam = mixup(images, masks, alpha=self.cfg.mixup_alpha)\n    logits = self(mix_images)\n    loss = self.loss_fn(logits, target_a) * lam + (1 - lam) * self.loss_fn(logits, target_b)\n                \nelif self.cfg.cutmix and torch.rand(1)[0] < self.cfg.cutmix_p:\n    mix_images, target_a, target_b, lam = cutmix(images, masks, alpha=self.cfg.cutmix_alpha)\n    logits = self(mix_images)\n    loss = self.loss_fn(logits, target_a) * lam + (1 - lam) * self.loss_fn(logits, target_b)\n```\n\n### 3.5. **What didn't work**\nA ton of things:\n- Using other bands/other data schema.\n- TTA.\n- We tried to combine the former and the latter frames with the labeled one at each stage of the UNet without success (only tried concatenation).\n- At the end of the competition we discovered that pseudolabels give a big boost for some models (not the strongest ones) but we didn't have enough time to properly experiment with it.\n- ...\n\n# **References**\n- [ASH color scheme](https://www.kaggle.com/code/inversion/visualizing-contrails)\n- [OpenContrails: Benchmarking Contrail Detection on GOES-16 ABI](https://arxiv.org/abs/2304.02122)\n- @egortrushin [public notebook](https://www.kaggle.com/code/egortrushin/gr-icrgw-training-with-4-folds) was our starting point.\n\n# **Acknowledgments**\nTo my wonderful teammates @maxdiazbattan @edomingo @juansensio (if I forgot something please feel free to write it in the comments).",
      "votes": 8
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2394213": "First of all, we would like to thank Kaggle and the host for such a funny (and *nothing-works-well-for-us*) competition. This is not a fancy solution but we feel proud of taking the right steps to generalize well to the private test set (it was indeed, 42 positions shake up 😃).\n\n# 1. **Context**\n- *Business context*: https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/overview\n- *Data context*: https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/data\n\n# 2. **Overview of the Approach**\nOur final solution was an ensemble of 8 models which combines different encoders and image sizes. Each model configuration was trained in turn 3 times with different random seeds (24 models in total) and always with all available data (train+validation).\n\n### **Data preprocessing**\nASH false color images, only 2D models with labeled frame and hard labels.\n\n### **Validation schema**\nProvided train/validation split.\n\n# 3. **Overview of the Approach**\n\n### 3.1. **Individual models and final ensemble**\nAll models are based on UNet architecture from SMP. We combine them using a weighted average of the probabilities dropped by single models, where weights were determined optimizing the global dice score on the validation set using Optuna. The final CV score was **0.6825** (which may be a little biased because as we mentioned we optimized the weights directly using the validation set).\n\n| Weight | Encoder| Image size | LR | Epochs | BS | CV score |\n| --- | --- | --- | --- | --- | --- | --- |\n| 0.02 | timm-resnest101e | 384 | 5e-4 | 20 | 48 | 0.64397\n| 0.07 | maxvit_base_tf_384.in21k_ft_in1k | 384 | 1e-4 | 20 | 16 | 0.65699 \n| 0.18 | maxvit_base_tf_512 | 512 | 1e-4 | 40 | 8* | 0.66414\n| 0.02 | tf_efficientnetv2_b3.in1k | 768 | 5e-4 | 35 | 24* | 0.65178\n| 0.23 | tf_efficientnetv2_m.in21k_ft_in1k | 512 | 5e-4 | 35 | 24* | 0.6525\n| 0.14 | tf_efficientnetv2_l.in21k_ft_in1k | 384 | 5e-4 | 35 | 24* | 0.64538\n| 0.11 | maxvit_large_tf_224.in1k | 448 | 1e-4 | 40 | 8* | 0.66407\n| 0.23 | maxvit_large_tf_384.in1k | 384 | 1e-4 | 40 | 8* | 0.66782\n\n*gradient accumulation of 2.\n \n### 3.2. **Threshold selection**\nThe optimal threshold for the ensemble was 0.43, which was determined running a for loop varying the threshold from 0 to 1 with 0.1 increments and calculating the corresponding global dice. We observed a relatively wide maximum peaking at 0.43 which give us enough confidence to choose that threshold.\n\n### 3.3. **Training procedure**\nWe used CosineAnnealingLR and AdamW. Maybe it is important to note that we always tried to choose the number of epochs such that the last epoch is the best one. This generalizes better than choosing the best checkpoint during training, especially in our case, where we observed an irregular behavior of the training curves. Finally, once a model was validated, we took the same configuration and train it again using all the available data.\n\n### 3.4. **Augmentations**\nWe didn't notice the mask shift, so flipping and rotation didn't work for us. Other augmentations that worked were:\n\n```python\nA.Compose([  \n     A.RandomResizedCrop(height=256, width=256, p=0.5),  \n     A.RandomBrightnessContrast(p=0.5)  \n])  \nCutMix(p=0.5)  \nMixUp(p=0.5)\n```\n\nThese are our implementations of CutMix and MixUp:\n```python\ndef rand_bbox(size, lam):\n    W = size[2]\n    H = size[3]\n    cut_rat = np.sqrt(1. - lam)\n    cut_w = int(W * cut_rat)\n    cut_h = int(H * cut_rat)\n\n    # uniform\n    cx = np.random.randint(W)\n    cy = np.random.randint(H)\n\n    bbx1 = np.clip(cx - cut_w // 2, 0, W)\n    bby1 = np.clip(cy - cut_h // 2, 0, H)\n    bbx2 = np.clip(cx + cut_w // 2, 0, W)\n    bby2 = np.clip(cy + cut_h // 2, 0, H)\n    return bbx1, bby1, bbx2, bby2\n\ndef cutmix(data, target, alpha):\n    indices = torch.randperm(data.size(0))\n    # shuffled_data = data[indices]\n    shuffled_target = target[indices]\n\n    lam = np.clip(np.random.beta(alpha, alpha),0.3,0.4)\n    bbx1, bby1, bbx2, bby2 = rand_bbox(data.size(), lam)\n    new_data = data.clone()\n    new_data[:, :, bby1:bby2, bbx1:bbx2] = data[indices, :, bby1:bby2, bbx1:bbx2]\n    # adjust lambda to exactly match pixel ratio\n    lam = 1 - ((bbx2 - bbx1) * (bby2 - bby1) / (data.size()[-1] * data.size()[-2]))\n    # targets = (target, shuffled_target, lam)\n    return new_data, target, shuffled_target, lam\n\ndef mixup(x: torch.Tensor, y: torch.Tensor, alpha: float = 1.0):\n    assert alpha > 0, \"alpha should be larger than 0\"\n    assert x.size(0) > 1, \"Mixup cannot be applied to a single instance.\"\n\n    lam = np.random.beta(alpha, alpha)\n    rand_index = torch.randperm(x.size()[0])\n    mixed_x = lam * x + (1 - lam) * x[rand_index, :]\n    target_a, target_b = y, y[rand_index]\n    return mixed_x, target_a, target_b, lam\n```\nAnd losses are calculated as:\n```\nif self.cfg.mixup and torch.rand(1)[0] < self.cfg.mixup_p:\n    mix_images, target_a, target_b, lam = mixup(images, masks, alpha=self.cfg.mixup_alpha)\n    logits = self(mix_images)\n    loss = self.loss_fn(logits, target_a) * lam + (1 - lam) * self.loss_fn(logits, target_b)\n                \nelif self.cfg.cutmix and torch.rand(1)[0] < self.cfg.cutmix_p:\n    mix_images, target_a, target_b, lam = cutmix(images, masks, alpha=self.cfg.cutmix_alpha)\n    logits = self(mix_images)\n    loss = self.loss_fn(logits, target_a) * lam + (1 - lam) * self.loss_fn(logits, target_b)\n```\n\n### 3.5. **What didn't work**\nA ton of things:\n- Using other bands/other data schema.\n- TTA.\n- We tried to combine the former and the latter frames with the labeled one at each stage of the UNet without success (only tried concatenation).\n- At the end of the competition we discovered that pseudolabels give a big boost for some models (not the strongest ones) but we didn't have enough time to properly experiment with it.\n- ...\n\n# **References**\n- [ASH color scheme](https://www.kaggle.com/code/inversion/visualizing-contrails)\n- [OpenContrails: Benchmarking Contrail Detection on GOES-16 ABI](https://arxiv.org/abs/2304.02122)\n- @egortrushin [public notebook](https://www.kaggle.com/code/egortrushin/gr-icrgw-training-with-4-folds) was our starting point.\n\n# **Acknowledgments**\nTo my wonderful teammates @maxdiazbattan @edomingo @juansensio (if I forgot something please feel free to write it in the comments)."
  }
}