{
  "id": 432690,
  "title": "11th Place Solution Write-up",
  "url": "/competitions/google-research-identify-contrails-reduce-global-warming/discussion/432690",
  "author_name": "Ioannis M",
  "post_date": "2023-08-18T12:51:40.303000",
  "votes": 16,
  "comment_count": 2,
  "views": 0,
  "content": "<h2><strong>Summary [TLDR]</strong></h2>\n<ul>\n<li>Weighted ensemble (optimal weights by Optuna)</li>\n<li>Multi-stage long training runs (70-100 epochs)</li>\n<li>Pseudo-labeling (time slices = 1-4, 6-8)</li>\n<li>SWA [3]</li>\n<li>TTA d4 / FlipTransforms [4]</li>\n<li>Image size x2 </li>\n</ul>\n<h2><strong>Modelling Approach</strong></h2>\n<p>We used <code>human_pixel_masks</code> as ground truth. We tried to incorporate <code>human_individual_labels</code> in different ways, but couldn’t improve using them. We didn’t try to average them and train on the soft labels. Our ensemble includes a model that calculates the losses on all individual masks and includes the one with the largest loss as an aux loss.</p>\n<p>We tried to improve a single model using temporal information and also couldn’t improve over 2d setup, so we moved to create a diverse ensemble using diff encoders and decoders. We used different CNNs and transformers as backbones and we also had models with a transformer decoder (UNetFormer).</p>\n<p>Many experiments conducted to utilise the additional data in the form of additional time periods so we incorporate that information in at least some way if not directly modelling for it. What ended up working for the pseudo setup in terms of both improving the single model and adding to the ensemble was the following:</p>\n<ul>\n<li>Balance the batch to continue ½ pseudo ½ real examples.</li>\n<li>Apply heavy augmentation on the pseudo data and light to normal augmentation on real data.</li>\n<li>Mixup the pseudo examples with the real examples.</li>\n</ul>\n<h2><strong>Training setup</strong></h2>\n<ul>\n<li>Optimizer: AdamW </li>\n<li>LR: 1.0e-3 or 1.0e-4</li>\n<li>Scheduler: Cosine Decay with 2 epochs warmup.</li>\n<li>Augs: Flips / RandomResizedCrop / RandAugment / Mixup / Mosaic</li>\n<li>Mixed loss: BCE / Dice / Focal </li>\n</ul>\n<h2><strong>CV strategy</strong></h2>\n<p>For most of our experiments we used the split provided by the hosts, i.e. train on full train data and evaluate on <code>validation.csv</code>. We did earlier some experiments with KFold but OOF CV was overestimated and the CV scores on the holdout (validation.csv) weren’t good. </p>\n<h4>Evaluation</h4>\n<p>We used 2 variants for evaluation against ground truth labels (256x256): <br>\n i) Model outputs x2 size --&gt; <code>interpolate(256, mode='nearest')</code> --&gt; calculate Dice score <br>\n ii) Model outputs directly 256 size </p>\n<p>PS: Most of our experiments used <code>(i)</code>, although we believe 2nd setup without interpolation is more reliable  </p>\n<h2><strong>Architectures / Backbones</strong></h2>\n<p>Final ensemble </p>\n<h4><code>CV: 0.7026</code> /  <code>Public LB:  0.71868</code> / <code>Private LB: 0.71059</code></h4>\n<table>\n<thead>\n<tr>\n<th>Backbone</th>\n<th>Architecture</th>\n<th>CV</th>\n<th>TTA</th>\n<th>Pseudo</th>\n<th>Mixup</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>efficientnet-b7</td>\n<td>Unet (SMP)</td>\n<td>0.6428</td>\n<td>Yes</td>\n<td>No</td>\n<td>Yes</td>\n</tr>\n<tr>\n<td>tf_efficientnet_b8</td>\n<td>Unet++ (Timm)</td>\n<td>0.6704</td>\n<td>No</td>\n<td>No</td>\n<td>Yes</td>\n</tr>\n<tr>\n<td>tf_efficientnet_b8</td>\n<td>Unet++ (Timm)</td>\n<td>0.6752</td>\n<td>Yes</td>\n<td>No</td>\n<td>Yes</td>\n</tr>\n<tr>\n<td>maxvit_small_tf_512</td>\n<td>Unet (Timm)</td>\n<td>0.6677</td>\n<td>No</td>\n<td>No</td>\n<td>Yes</td>\n</tr>\n<tr>\n<td>convnext_base</td>\n<td>Unet++ (Timm)</td>\n<td>0.6618</td>\n<td>No</td>\n<td>Yes</td>\n<td>Yes</td>\n</tr>\n<tr>\n<td>convnext_large</td>\n<td>Unet (Timm)</td>\n<td>0.6670</td>\n<td>No</td>\n<td>Yes</td>\n<td>Yes</td>\n</tr>\n<tr>\n<td>timm-resnest200e</td>\n<td>Unet</td>\n<td>0.6869</td>\n<td>Yes</td>\n<td>No</td>\n<td>No</td>\n</tr>\n<tr>\n<td>timm-resnest200e</td>\n<td>Unet</td>\n<td>0.6853</td>\n<td>Yes</td>\n<td>No</td>\n<td>Yes</td>\n</tr>\n<tr>\n<td>timm-resnest200e</td>\n<td>Unet</td>\n<td>0.6885</td>\n<td>Yes</td>\n<td>Yes</td>\n<td>Yes</td>\n</tr>\n<tr>\n<td>convnext_large_384_in22ft1k</td>\n<td>UnetFormer</td>\n<td>0.6788</td>\n<td>Yes</td>\n<td>No</td>\n<td>Yes</td>\n</tr>\n<tr>\n<td>tf_efficientnet_b6</td>\n<td>UnetFormer</td>\n<td>0.6532</td>\n<td>No</td>\n<td>No</td>\n<td>No</td>\n</tr>\n<tr>\n<td>eva02_large_patch14_448</td>\n<td>Unet</td>\n<td>0.6542</td>\n<td>Yes</td>\n<td>No</td>\n<td>No</td>\n</tr>\n<tr>\n<td>convnext_large_384_in22ft1k</td>\n<td>UnetFormer</td>\n<td>0.6841</td>\n<td>Yes</td>\n<td>No</td>\n<td>Yes</td>\n</tr>\n</tbody>\n</table>\n<h4>The selected ensemble was our <strong>best CV</strong> / <strong>best Public LB (6th place)</strong> and the <strong>best PVT LB (11th place)</strong></h4>\n<h2>CV-LB Plot</h2>\n<p><img src=\"https://raw.githubusercontent.com/i-mein/Kaggle_images/main/2023-08/CV-PVT.png\" alt=\"\"></p>\n<p><img src=\"https://raw.githubusercontent.com/i-mein/Kaggle_images/main/2023-08/CV-Public.png\" alt=\"\"></p>\n<h2><strong>Things didn't work</strong></h2>\n<ul>\n<li>Use of ConvLSTM, UTAE, 3D-Unets to leverage information from other time slices </li>\n<li>Post-processing with t-2, t-1, t slices </li>\n<li>diff image resolutions (384, 768, 1024) </li>\n</ul>\n<h2><strong>Acknowledgements</strong></h2>\n<p>We would like to thank hosts and Kaggle team for organizing such an interesting research competition. We are also grateful to all kagglers that share their ideas in discussions and notebooks. </p>\n<h2>Libraries</h2>\n<p>[1] <a href=\"https://github.com/huggingface/pytorch-image-models\" target=\"_blank\">Timm</a> <br>\n[2] <a href=\"https://github.com/qubvel/segmentation_models.pytorch\" target=\"_blank\">segmentation_models.pytorch (SMP)</a> <br>\n[3] <a href=\"https://github.com/izmailovpavel/contrib_swa_examples\" target=\"_blank\">SWA implementation</a><br>\n[4] <a href=\"https://github.com/qubvel/ttach\" target=\"_blank\">ttach</a></p>\n<h2><strong>Team GIRYN</strong></h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2596066%2Fc292aef766c287db0c594c6ed07ababb%2FScreenshot%202023-08-18%20at%2015.39.37.png?generation=1692362680181936&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li><a href=\"https://www.kaggle.com/rohitsingh9990\" target=\"_blank\">Rohit Singh (Dracarys)</a></li>\n<li><a href=\"https://www.kaggle.com/phoenix9032\" target=\"_blank\">Nihjar (doomsday)</a> </li>\n<li><a href=\"https://www.kaggle.com/imeintanis\" target=\"_blank\">Ioannis Meintanis</a></li>\n<li><a href=\"https://www.kaggle.com/yousof9\" target=\"_blank\">Yousef Rabi</a></li>\n<li><a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">Giba</a></li>\n</ul>",
  "messages": [
    {
      "id": 2396738,
      "postDate": "2023-08-18T12:51:40.303Z",
      "content": "<h2><strong>Summary [TLDR]</strong></h2>\n<ul>\n<li>Weighted ensemble (optimal weights by Optuna)</li>\n<li>Multi-stage long training runs (70-100 epochs)</li>\n<li>Pseudo-labeling (time slices = 1-4, 6-8)</li>\n<li>SWA [3]</li>\n<li>TTA d4 / FlipTransforms [4]</li>\n<li>Image size x2 </li>\n</ul>\n<h2><strong>Modelling Approach</strong></h2>\n<p>We used <code>human_pixel_masks</code> as ground truth. We tried to incorporate <code>human_individual_labels</code> in different ways, but couldn’t improve using them. We didn’t try to average them and train on the soft labels. Our ensemble includes a model that calculates the losses on all individual masks and includes the one with the largest loss as an aux loss.</p>\n<p>We tried to improve a single model using temporal information and also couldn’t improve over 2d setup, so we moved to create a diverse ensemble using diff encoders and decoders. We used different CNNs and transformers as backbones and we also had models with a transformer decoder (UNetFormer).</p>\n<p>Many experiments conducted to utilise the additional data in the form of additional time periods so we incorporate that information in at least some way if not directly modelling for it. What ended up working for the pseudo setup in terms of both improving the single model and adding to the ensemble was the following:</p>\n<ul>\n<li>Balance the batch to continue ½ pseudo ½ real examples.</li>\n<li>Apply heavy augmentation on the pseudo data and light to normal augmentation on real data.</li>\n<li>Mixup the pseudo examples with the real examples.</li>\n</ul>\n<h2><strong>Training setup</strong></h2>\n<ul>\n<li>Optimizer: AdamW </li>\n<li>LR: 1.0e-3 or 1.0e-4</li>\n<li>Scheduler: Cosine Decay with 2 epochs warmup.</li>\n<li>Augs: Flips / RandomResizedCrop / RandAugment / Mixup / Mosaic</li>\n<li>Mixed loss: BCE / Dice / Focal </li>\n</ul>\n<h2><strong>CV strategy</strong></h2>\n<p>For most of our experiments we used the split provided by the hosts, i.e. train on full train data and evaluate on <code>validation.csv</code>. We did earlier some experiments with KFold but OOF CV was overestimated and the CV scores on the holdout (validation.csv) weren’t good. </p>\n<h4>Evaluation</h4>\n<p>We used 2 variants for evaluation against ground truth labels (256x256): <br>\n i) Model outputs x2 size --&gt; <code>interpolate(256, mode='nearest')</code> --&gt; calculate Dice score <br>\n ii) Model outputs directly 256 size </p>\n<p>PS: Most of our experiments used <code>(i)</code>, although we believe 2nd setup without interpolation is more reliable  </p>\n<h2><strong>Architectures / Backbones</strong></h2>\n<p>Final ensemble </p>\n<h4><code>CV: 0.7026</code> /  <code>Public LB:  0.71868</code> / <code>Private LB: 0.71059</code></h4>\n<table>\n<thead>\n<tr>\n<th>Backbone</th>\n<th>Architecture</th>\n<th>CV</th>\n<th>TTA</th>\n<th>Pseudo</th>\n<th>Mixup</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>efficientnet-b7</td>\n<td>Unet (SMP)</td>\n<td>0.6428</td>\n<td>Yes</td>\n<td>No</td>\n<td>Yes</td>\n</tr>\n<tr>\n<td>tf_efficientnet_b8</td>\n<td>Unet++ (Timm)</td>\n<td>0.6704</td>\n<td>No</td>\n<td>No</td>\n<td>Yes</td>\n</tr>\n<tr>\n<td>tf_efficientnet_b8</td>\n<td>Unet++ (Timm)</td>\n<td>0.6752</td>\n<td>Yes</td>\n<td>No</td>\n<td>Yes</td>\n</tr>\n<tr>\n<td>maxvit_small_tf_512</td>\n<td>Unet (Timm)</td>\n<td>0.6677</td>\n<td>No</td>\n<td>No</td>\n<td>Yes</td>\n</tr>\n<tr>\n<td>convnext_base</td>\n<td>Unet++ (Timm)</td>\n<td>0.6618</td>\n<td>No</td>\n<td>Yes</td>\n<td>Yes</td>\n</tr>\n<tr>\n<td>convnext_large</td>\n<td>Unet (Timm)</td>\n<td>0.6670</td>\n<td>No</td>\n<td>Yes</td>\n<td>Yes</td>\n</tr>\n<tr>\n<td>timm-resnest200e</td>\n<td>Unet</td>\n<td>0.6869</td>\n<td>Yes</td>\n<td>No</td>\n<td>No</td>\n</tr>\n<tr>\n<td>timm-resnest200e</td>\n<td>Unet</td>\n<td>0.6853</td>\n<td>Yes</td>\n<td>No</td>\n<td>Yes</td>\n</tr>\n<tr>\n<td>timm-resnest200e</td>\n<td>Unet</td>\n<td>0.6885</td>\n<td>Yes</td>\n<td>Yes</td>\n<td>Yes</td>\n</tr>\n<tr>\n<td>convnext_large_384_in22ft1k</td>\n<td>UnetFormer</td>\n<td>0.6788</td>\n<td>Yes</td>\n<td>No</td>\n<td>Yes</td>\n</tr>\n<tr>\n<td>tf_efficientnet_b6</td>\n<td>UnetFormer</td>\n<td>0.6532</td>\n<td>No</td>\n<td>No</td>\n<td>No</td>\n</tr>\n<tr>\n<td>eva02_large_patch14_448</td>\n<td>Unet</td>\n<td>0.6542</td>\n<td>Yes</td>\n<td>No</td>\n<td>No</td>\n</tr>\n<tr>\n<td>convnext_large_384_in22ft1k</td>\n<td>UnetFormer</td>\n<td>0.6841</td>\n<td>Yes</td>\n<td>No</td>\n<td>Yes</td>\n</tr>\n</tbody>\n</table>\n<h4>The selected ensemble was our <strong>best CV</strong> / <strong>best Public LB (6th place)</strong> and the <strong>best PVT LB (11th place)</strong></h4>\n<h2>CV-LB Plot</h2>\n<p><img src=\"https://raw.githubusercontent.com/i-mein/Kaggle_images/main/2023-08/CV-PVT.png\" alt=\"\"></p>\n<p><img src=\"https://raw.githubusercontent.com/i-mein/Kaggle_images/main/2023-08/CV-Public.png\" alt=\"\"></p>\n<h2><strong>Things didn't work</strong></h2>\n<ul>\n<li>Use of ConvLSTM, UTAE, 3D-Unets to leverage information from other time slices </li>\n<li>Post-processing with t-2, t-1, t slices </li>\n<li>diff image resolutions (384, 768, 1024) </li>\n</ul>\n<h2><strong>Acknowledgements</strong></h2>\n<p>We would like to thank hosts and Kaggle team for organizing such an interesting research competition. We are also grateful to all kagglers that share their ideas in discussions and notebooks. </p>\n<h2>Libraries</h2>\n<p>[1] <a href=\"https://github.com/huggingface/pytorch-image-models\" target=\"_blank\">Timm</a> <br>\n[2] <a href=\"https://github.com/qubvel/segmentation_models.pytorch\" target=\"_blank\">segmentation_models.pytorch (SMP)</a> <br>\n[3] <a href=\"https://github.com/izmailovpavel/contrib_swa_examples\" target=\"_blank\">SWA implementation</a><br>\n[4] <a href=\"https://github.com/qubvel/ttach\" target=\"_blank\">ttach</a></p>\n<h2><strong>Team GIRYN</strong></h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2596066%2Fc292aef766c287db0c594c6ed07ababb%2FScreenshot%202023-08-18%20at%2015.39.37.png?generation=1692362680181936&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li><a href=\"https://www.kaggle.com/rohitsingh9990\" target=\"_blank\">Rohit Singh (Dracarys)</a></li>\n<li><a href=\"https://www.kaggle.com/phoenix9032\" target=\"_blank\">Nihjar (doomsday)</a> </li>\n<li><a href=\"https://www.kaggle.com/imeintanis\" target=\"_blank\">Ioannis Meintanis</a></li>\n<li><a href=\"https://www.kaggle.com/yousof9\" target=\"_blank\">Yousef Rabi</a></li>\n<li><a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">Giba</a></li>\n</ul>",
      "rawMarkdown": "## **Summary [TLDR]**\n\n- Weighted ensemble (optimal weights by Optuna)\n- Multi-stage long training runs (70-100 epochs)\n- Pseudo-labeling (time slices = 1-4, 6-8)\n- SWA [3]\n- TTA d4 / FlipTransforms [4]\n- Image size x2 \n\n\n## **Modelling Approach** \n\nWe used `human_pixel_masks` as ground truth. We tried to incorporate `human_individual_labels` in different ways, but couldn’t improve using them. We didn’t try to average them and train on the soft labels. Our ensemble includes a model that calculates the losses on all individual masks and includes the one with the largest loss as an aux loss.\n\nWe tried to improve a single model using temporal information and also couldn’t improve over 2d setup, so we moved to create a diverse ensemble using diff encoders and decoders. We used different CNNs and transformers as backbones and we also had models with a transformer decoder (UNetFormer).\n\nMany experiments conducted to utilise the additional data in the form of additional time periods so we incorporate that information in at least some way if not directly modelling for it. What ended up working for the pseudo setup in terms of both improving the single model and adding to the ensemble was the following:\n\n- Balance the batch to continue ½ pseudo ½ real examples.\n- Apply heavy augmentation on the pseudo data and light to normal augmentation on real data.\n- Mixup the pseudo examples with the real examples.\n\n\n## **Training setup**\n\n- Optimizer: AdamW \n- LR: 1.0e-3 or 1.0e-4\n- Scheduler: Cosine Decay with 2 epochs warmup.\n- Augs: Flips / RandomResizedCrop / RandAugment / Mixup / Mosaic\n- Mixed loss: BCE / Dice / Focal \n\n\n## **CV strategy**\n\nFor most of our experiments we used the split provided by the hosts, i.e. train on full train data and evaluate on `validation.csv`. We did earlier some experiments with KFold but OOF CV was overestimated and the CV scores on the holdout (validation.csv) weren’t good. \n\n#### Evaluation\nWe used 2 variants for evaluation against ground truth labels (256x256): \n i) Model outputs x2 size --> `interpolate(256, mode='nearest')` --> calculate Dice score \n ii) Model outputs directly 256 size \n\nPS: Most of our experiments used `(i)`, although we believe 2nd setup without interpolation is more reliable  \n\n\n## **Architectures / Backbones**\n\nFinal ensemble \n#### `CV: 0.7026` /  `Public LB:  0.71868` / `Private LB: 0.71059`\n\n\n| Backbone | Architecture | CV | TTA | Pseudo | Mixup |\n| --- | --- | --- | --- | --- | --- |\n| efficientnet-b7     | Unet (SMP)       | 0.6428 | Yes | No | Yes |\n| tf_efficientnet_b8 | Unet++ (Timm) | 0.6704 | No  | No | Yes |\n| tf_efficientnet_b8 | Unet++ (Timm) | 0.6752 | Yes | No | Yes |\n| maxvit_small_tf_512 | Unet (Timm)  | 0.6677 | No  | No | Yes |\n| convnext_base    | Unet++ (Timm)  | 0.6618 | No  | Yes | Yes |\n| convnext_large    | Unet (Timm)  | 0.6670 | No  | Yes | Yes |\n| timm-resnest200e | Unet   | 0.6869 | Yes  | No | No |\n| timm-resnest200e | Unet   | 0.6853 | Yes  | No | Yes |\n| timm-resnest200e | Unet   | 0.6885 | Yes  | Yes | Yes |\n| convnext_large_384_in22ft1k | UnetFormer | 0.6788 | Yes  | No | Yes |\n| tf_efficientnet_b6 | UnetFormer | 0.6532 | No  | No | No |\n| eva02_large_patch14_448 | Unet | 0.6542 | Yes  | No | No |\n| convnext_large_384_in22ft1k | UnetFormer | 0.6841 | Yes  | No | Yes |\n\n\n#### The selected ensemble was our **best CV** / **best Public LB (6th place)** and the **best PVT LB (11th place)**\n\n\n## CV-LB Plot \n\n![](https://raw.githubusercontent.com/i-mein/Kaggle_images/main/2023-08/CV-PVT.png)\n\n![](https://raw.githubusercontent.com/i-mein/Kaggle_images/main/2023-08/CV-Public.png)\n\n\n## **Things didn't work**\n\n- Use of ConvLSTM, UTAE, 3D-Unets to leverage information from other time slices \n- Post-processing with t-2, t-1, t slices \n- diff image resolutions (384, 768, 1024) \n\n## **Acknowledgements**\n\nWe would like to thank hosts and Kaggle team for organizing such an interesting research competition. We are also grateful to all kagglers that share their ideas in discussions and notebooks. \n\n## Libraries\n[1] [Timm](https://github.com/huggingface/pytorch-image-models) \n[2] [segmentation_models.pytorch (SMP)](https://github.com/qubvel/segmentation_models.pytorch) \n[3] [SWA implementation](https://github.com/izmailovpavel/contrib_swa_examples)\n[4] [ttach](https://github.com/qubvel/ttach)\n\n\n## **Team GIRYN**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2596066%2Fc292aef766c287db0c594c6ed07ababb%2FScreenshot%202023-08-18%20at%2015.39.37.png?generation=1692362680181936&alt=media)\n\n- [Rohit Singh (Dracarys)](https://www.kaggle.com/rohitsingh9990)\n- [Nihjar (doomsday)](https://www.kaggle.com/phoenix9032) \n- [Ioannis Meintanis](https://www.kaggle.com/imeintanis)\n- [Yousef Rabi](https://www.kaggle.com/yousof9)\n- [Giba](https://www.kaggle.com/titericz)",
      "votes": 16
    },
    {
      "id": 2396988,
      "postDate": "2023-08-18T16:30:59.630Z",
      "content": "<p>Congratulations guys and thanks for sharing! I wanted to ask you a couple of questions: </p>\n<p>What was the secret behind the 0.68+ timm-resnest200e models? It is impressive how high you scored with this backbone. Was the long training run?</p>\n<blockquote>\n  <p>Multi-stage long training runs (70-100 epochs)</p>\n</blockquote>\n<p>What do you mean with multi-stage in this case?</p>\n<blockquote>\n  <p>SWA</p>\n</blockquote>\n<p>How much SWA improved your scores? I have tried it a lot of times in the past but I couldn't get a boost. Could you share how to implement It/use It properly? 😅</p>\n<blockquote>\n  <p>TTA d4</p>\n</blockquote>\n<p>What is d4? 😅</p>\n<p>Thanks again</p>",
      "rawMarkdown": "Congratulations guys and thanks for sharing! I wanted to ask you a couple of questions: \n\nWhat was the secret behind the 0.68+ timm-resnest200e models? It is impressive how high you scored with this backbone. Was the long training run?\n\n>Multi-stage long training runs (70-100 epochs)\n\nWhat do you mean with multi-stage in this case?\n\n>SWA\n\nHow much SWA improved your scores? I have tried it a lot of times in the past but I couldn't get a boost. Could you share how to implement It/use It properly? 😅\n\n> TTA d4\n\nWhat is d4? 😅\n\nThanks again",
      "votes": 3,
      "replies": [
        {
          "id": 2397134,
          "postDate": "2023-08-18T18:41:01.440Z",
          "content": "<p>Hi thanks for your questions. </p>\n<blockquote>\n  <p>What was the secret behind the 0.68+ timm-resnest200e models? It is impressive how high you scored with this backbone. Was the long training run?</p>\n</blockquote>\n<p>I'd say was due to the multi-stage training and the training setup (balanced dataloader, aggresive augs on pseudo to avoid overfit) </p>\n<p>In general by multi-stage here we mean</p>\n<ul>\n<li>i) train with mixed loss and t=5 image only for 50-80 epochs</li>\n<li>ii) finetune with pseudo images (t=1-4, 6-8) and dice loss for 20-50 epochs </li>\n<li>iii) continue finetuning for 10-20 epochs more with SWA (disable mixup/cutmix)</li>\n</ul>\n<p>to have more diversity each of us used the above as guideline with some variants on number of epochs, number of stages, pseudo slices to include etc - particularly in my runs I pretrained first with pseudo and finetuned with t=5 only, dice loss and SWA. Also, in some exp (ii) and (iii) combined to a single one.</p>\n<p>ps: the model you are referring has been trained by <a href=\"https://www.kaggle.com/yousof9\" target=\"_blank\">Yousef Rabi</a>, he may add some details if I missed anything</p>\n<blockquote>\n  <p>SWA</p>\n</blockquote>\n<p>well SWA helped a lot not only to improve scores (up to ~ 0.01 in some cases) but also to stabilise/smooth loss curves.<br>\nFor the implementation we adopt it from here <br>\n<a href=\"https://github.com/izmailovpavel/contrib_swa_examples\" target=\"_blank\">repo</a><br>\n<a href=\"https://pytorch.org/blog/stochastic-weight-averaging-in-pytorch\" target=\"_blank\">blogpost</a></p>\n<blockquote>\n  <p>TTA d4</p>\n</blockquote>\n<p><code>d4_transform --&gt; (flips + rotation 0, 90, 180, 270)</code><br>\nsee here <a href=\"https://github.com/qubvel/ttach\" target=\"_blank\">qubvel/ttach</a></p>",
          "rawMarkdown": "Hi thanks for your questions. \n\n> What was the secret behind the 0.68+ timm-resnest200e models? It is impressive how high you scored with this backbone. Was the long training run?\n\nI'd say was due to the multi-stage training and the training setup (balanced dataloader, aggresive augs on pseudo to avoid overfit) \n\nIn general by multi-stage here we mean\n- i) train with mixed loss and t=5 image only for 50-80 epochs\n- ii) finetune with pseudo images (t=1-4, 6-8) and dice loss for 20-50 epochs \n- iii) continue finetuning for 10-20 epochs more with SWA (disable mixup/cutmix)\n\nto have more diversity each of us used the above as guideline with some variants on number of epochs, number of stages, pseudo slices to include etc - particularly in my runs I pretrained first with pseudo and finetuned with t=5 only, dice loss and SWA. Also, in some exp (ii) and (iii) combined to a single one.\n\nps: the model you are referring has been trained by [Yousef Rabi](https://www.kaggle.com/yousof9), he may add some details if I missed anything\n\n> SWA\n\nwell SWA helped a lot not only to improve scores (up to ~ 0.01 in some cases) but also to stabilise/smooth loss curves.\nFor the implementation we adopt it from here \n[repo](https://github.com/izmailovpavel/contrib_swa_examples)\n[blogpost](https://pytorch.org/blog/stochastic-weight-averaging-in-pytorch)\n\n> TTA d4\n\n`d4_transform --> (flips + rotation 0, 90, 180, 270)`\nsee here [qubvel/ttach](https://github.com/qubvel/ttach)\n",
          "votes": 4
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2396988,
      "author_name": "delai50",
      "author_url": "",
      "post_date": "2023-08-18T16:30:59.630000",
      "content": "<p>Congratulations guys and thanks for sharing! I wanted to ask you a couple of questions: </p>\n<p>What was the secret behind the 0.68+ timm-resnest200e models? It is impressive how high you scored with this backbone. Was the long training run?</p>\n<blockquote>\n  <p>Multi-stage long training runs (70-100 epochs)</p>\n</blockquote>\n<p>What do you mean with multi-stage in this case?</p>\n<blockquote>\n  <p>SWA</p>\n</blockquote>\n<p>How much SWA improved your scores? I have tried it a lot of times in the past but I couldn't get a boost. Could you share how to implement It/use It properly? 😅</p>\n<blockquote>\n  <p>TTA d4</p>\n</blockquote>\n<p>What is d4? 😅</p>\n<p>Thanks again</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2397134,
          "author_name": "Ioannis M",
          "author_url": "",
          "post_date": "2023-08-18T18:41:01.440000",
          "content": "<p>Hi thanks for your questions. </p>\n<blockquote>\n  <p>What was the secret behind the 0.68+ timm-resnest200e models? It is impressive how high you scored with this backbone. Was the long training run?</p>\n</blockquote>\n<p>I'd say was due to the multi-stage training and the training setup (balanced dataloader, aggresive augs on pseudo to avoid overfit) </p>\n<p>In general by multi-stage here we mean</p>\n<ul>\n<li>i) train with mixed loss and t=5 image only for 50-80 epochs</li>\n<li>ii) finetune with pseudo images (t=1-4, 6-8) and dice loss for 20-50 epochs </li>\n<li>iii) continue finetuning for 10-20 epochs more with SWA (disable mixup/cutmix)</li>\n</ul>\n<p>to have more diversity each of us used the above as guideline with some variants on number of epochs, number of stages, pseudo slices to include etc - particularly in my runs I pretrained first with pseudo and finetuned with t=5 only, dice loss and SWA. Also, in some exp (ii) and (iii) combined to a single one.</p>\n<p>ps: the model you are referring has been trained by <a href=\"https://www.kaggle.com/yousof9\" target=\"_blank\">Yousef Rabi</a>, he may add some details if I missed anything</p>\n<blockquote>\n  <p>SWA</p>\n</blockquote>\n<p>well SWA helped a lot not only to improve scores (up to ~ 0.01 in some cases) but also to stabilise/smooth loss curves.<br>\nFor the implementation we adopt it from here <br>\n<a href=\"https://github.com/izmailovpavel/contrib_swa_examples\" target=\"_blank\">repo</a><br>\n<a href=\"https://pytorch.org/blog/stochastic-weight-averaging-in-pytorch\" target=\"_blank\">blogpost</a></p>\n<blockquote>\n  <p>TTA d4</p>\n</blockquote>\n<p><code>d4_transform --&gt; (flips + rotation 0, 90, 180, 270)</code><br>\nsee here <a href=\"https://github.com/qubvel/ttach\" target=\"_blank\">qubvel/ttach</a></p>",
          "votes": 4,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2396738": "## **Summary [TLDR]**\n\n- Weighted ensemble (optimal weights by Optuna)\n- Multi-stage long training runs (70-100 epochs)\n- Pseudo-labeling (time slices = 1-4, 6-8)\n- SWA [3]\n- TTA d4 / FlipTransforms [4]\n- Image size x2 \n\n\n## **Modelling Approach** \n\nWe used `human_pixel_masks` as ground truth. We tried to incorporate `human_individual_labels` in different ways, but couldn’t improve using them. We didn’t try to average them and train on the soft labels. Our ensemble includes a model that calculates the losses on all individual masks and includes the one with the largest loss as an aux loss.\n\nWe tried to improve a single model using temporal information and also couldn’t improve over 2d setup, so we moved to create a diverse ensemble using diff encoders and decoders. We used different CNNs and transformers as backbones and we also had models with a transformer decoder (UNetFormer).\n\nMany experiments conducted to utilise the additional data in the form of additional time periods so we incorporate that information in at least some way if not directly modelling for it. What ended up working for the pseudo setup in terms of both improving the single model and adding to the ensemble was the following:\n\n- Balance the batch to continue ½ pseudo ½ real examples.\n- Apply heavy augmentation on the pseudo data and light to normal augmentation on real data.\n- Mixup the pseudo examples with the real examples.\n\n\n## **Training setup**\n\n- Optimizer: AdamW \n- LR: 1.0e-3 or 1.0e-4\n- Scheduler: Cosine Decay with 2 epochs warmup.\n- Augs: Flips / RandomResizedCrop / RandAugment / Mixup / Mosaic\n- Mixed loss: BCE / Dice / Focal \n\n\n## **CV strategy**\n\nFor most of our experiments we used the split provided by the hosts, i.e. train on full train data and evaluate on `validation.csv`. We did earlier some experiments with KFold but OOF CV was overestimated and the CV scores on the holdout (validation.csv) weren’t good. \n\n#### Evaluation\nWe used 2 variants for evaluation against ground truth labels (256x256): \n i) Model outputs x2 size --> `interpolate(256, mode='nearest')` --> calculate Dice score \n ii) Model outputs directly 256 size \n\nPS: Most of our experiments used `(i)`, although we believe 2nd setup without interpolation is more reliable  \n\n\n## **Architectures / Backbones**\n\nFinal ensemble \n#### `CV: 0.7026` /  `Public LB:  0.71868` / `Private LB: 0.71059`\n\n\n| Backbone | Architecture | CV | TTA | Pseudo | Mixup |\n| --- | --- | --- | --- | --- | --- |\n| efficientnet-b7     | Unet (SMP)       | 0.6428 | Yes | No | Yes |\n| tf_efficientnet_b8 | Unet++ (Timm) | 0.6704 | No  | No | Yes |\n| tf_efficientnet_b8 | Unet++ (Timm) | 0.6752 | Yes | No | Yes |\n| maxvit_small_tf_512 | Unet (Timm)  | 0.6677 | No  | No | Yes |\n| convnext_base    | Unet++ (Timm)  | 0.6618 | No  | Yes | Yes |\n| convnext_large    | Unet (Timm)  | 0.6670 | No  | Yes | Yes |\n| timm-resnest200e | Unet   | 0.6869 | Yes  | No | No |\n| timm-resnest200e | Unet   | 0.6853 | Yes  | No | Yes |\n| timm-resnest200e | Unet   | 0.6885 | Yes  | Yes | Yes |\n| convnext_large_384_in22ft1k | UnetFormer | 0.6788 | Yes  | No | Yes |\n| tf_efficientnet_b6 | UnetFormer | 0.6532 | No  | No | No |\n| eva02_large_patch14_448 | Unet | 0.6542 | Yes  | No | No |\n| convnext_large_384_in22ft1k | UnetFormer | 0.6841 | Yes  | No | Yes |\n\n\n#### The selected ensemble was our **best CV** / **best Public LB (6th place)** and the **best PVT LB (11th place)**\n\n\n## CV-LB Plot \n\n![](https://raw.githubusercontent.com/i-mein/Kaggle_images/main/2023-08/CV-PVT.png)\n\n![](https://raw.githubusercontent.com/i-mein/Kaggle_images/main/2023-08/CV-Public.png)\n\n\n## **Things didn't work**\n\n- Use of ConvLSTM, UTAE, 3D-Unets to leverage information from other time slices \n- Post-processing with t-2, t-1, t slices \n- diff image resolutions (384, 768, 1024) \n\n## **Acknowledgements**\n\nWe would like to thank hosts and Kaggle team for organizing such an interesting research competition. We are also grateful to all kagglers that share their ideas in discussions and notebooks. \n\n## Libraries\n[1] [Timm](https://github.com/huggingface/pytorch-image-models) \n[2] [segmentation_models.pytorch (SMP)](https://github.com/qubvel/segmentation_models.pytorch) \n[3] [SWA implementation](https://github.com/izmailovpavel/contrib_swa_examples)\n[4] [ttach](https://github.com/qubvel/ttach)\n\n\n## **Team GIRYN**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2596066%2Fc292aef766c287db0c594c6ed07ababb%2FScreenshot%202023-08-18%20at%2015.39.37.png?generation=1692362680181936&alt=media)\n\n- [Rohit Singh (Dracarys)](https://www.kaggle.com/rohitsingh9990)\n- [Nihjar (doomsday)](https://www.kaggle.com/phoenix9032) \n- [Ioannis Meintanis](https://www.kaggle.com/imeintanis)\n- [Yousef Rabi](https://www.kaggle.com/yousof9)\n- [Giba](https://www.kaggle.com/titericz)",
    "2396988": "Congratulations guys and thanks for sharing! I wanted to ask you a couple of questions: \n\nWhat was the secret behind the 0.68+ timm-resnest200e models? It is impressive how high you scored with this backbone. Was the long training run?\n\n>Multi-stage long training runs (70-100 epochs)\n\nWhat do you mean with multi-stage in this case?\n\n>SWA\n\nHow much SWA improved your scores? I have tried it a lot of times in the past but I couldn't get a boost. Could you share how to implement It/use It properly? 😅\n\n> TTA d4\n\nWhat is d4? 😅\n\nThanks again"
  }
}