{
  "id": 430549,
  "title": "5th place solution (best single model, private LB 0.71443)",
  "url": "/competitions/google-research-identify-contrails-reduce-global-warming/discussion/430549",
  "author_name": "Johnny Lee",
  "post_date": "2023-08-10T08:41:25.714000",
  "votes": 41,
  "comment_count": 15,
  "views": 0,
  "content": "<p>Thanks to the organizers for the great competition. Also thanks to my teammates: <a href=\"https://www.kaggle.com/imakarov\" target=\"_blank\">Ilya Makarov</a>. I had a lot of fun and learned a lot.<br>\nI'd like to share our solution and what I learned from this competition. I hope it will be helpful for you.</p>\n<h2>Solution overview of my part</h2>\n<h4>Models</h4>\n<p>I use 5 models as follows:</p>\n<ul>\n<li>Single frame model 1: EfficientNetV2L + UNet, 512x512, crop to 480x480 for training.</li>\n<li>Single frame model 2: EfficientNetV2L + UNet, 768x768, crop to 512x512 for training.</li>\n<li>Multi frames model 1: EfficientNetV2L + customized 3D UNet 1, 256x256, use all 8 frames.</li>\n<li>Multi frames model 2: EfficientNetV2L + customized 3D UNet 2, 512x512, crop to 480x480 for training, use 5 frames only.</li>\n<li>Multi frames model 3: EfficientNetV2L + customized 3D UNet 3, 512x512, crop to 480x480 for training, use 5 frames only.<br>\n<strong>(<a href=\"https://www.kaggle.com/wuliaokaola/icrgw-submission-single-model-lb-0-714-5th-place\" target=\"_blank\">Best single model, late submission, private LB: 0.71443, public LB 0.70803</a>)</strong></li>\n</ul>\n<p>About the customized 3D UNet<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2006644%2F2130a20eb0d042ac9d48d1ec2d7d88f9%2Funet3d_.png?generation=1691656552129914&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>Apply the same encoder(backbone) to each frame</li>\n<li>Combine the encoder outputs by Conv3D at each level. ConvLSTM2D also works, but it's slower.</li>\n<li>Use UNet as decoder to get the final output.</li>\n</ul>\n<h4>Augmentation</h4>\n<ul>\n<li>Filp LR, Filp UD, Rotate 90 (for TTA8)</li>\n<li>Noise by channel, noise by pixel</li>\n<li>Random dropout frame (multi frames model only)</li>\n<li>Random crop</li>\n</ul>\n<p>At the early stage of the competition, I noticed that if I don't use flip and rot90, the model will be overfitting quickly. But if I use the flip and rot90, the result became worse. I was confused for a long time. Finally, I found the reason. As mentioned in <a href=\"https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/discussion/430479#2382723\" target=\"_blank\">other discussions</a>, The image and the mask are not aligned. The mask is a littel bit shifted to the bottom right. So if we use flip or rot90, we need to align them. The magic number for me is: x = 0.408, y = 0.453. </p>\n<p>Tried but not worked:</p>\n<ul>\n<li>Random rotate</li>\n<li>CoarseDropout</li>\n<li>Dropout by channel</li>\n</ul>\n<h4>Target and loss function</h4>\n<p>I used 2 targets:</p>\n<ul>\n<li>Target 1: Sigmoid to 1 channel. weight: 75% (human_pixel_masks.npy)</li>\n<li>Target 2: Softmax to 5 channel. weight: 25% (human_individual_masks.npy. 0 for not contrails, 1 for 25% of labelers, 2 for 50% of labelers, 3 for 75% of labelers, 4 for 100% of labelers) </li>\n</ul>\n<p>The loss function is weighted binary cross entropy + dice loss. <br>\nAnd use Lion for optimizer.　(Lion: <a href=\"https://github.com/keras-team/keras/blob/v2.13.1/keras/optimizers/lion.py\" target=\"_blank\">https://github.com/keras-team/keras/blob/v2.13.1/keras/optimizers/lion.py</a>)</p>\n<h4>Prediction</h4>\n<ul>\n<li>Ensemble by 5 models with TTA8 (flip LR, flip UD, rotate 90)<br>\nThe TTA8 is very important. It can improve the score by 0.005~0.01.</li>\n</ul>\n<p>The best result is public LB: 0.71243, private LB: 0.71756. But unfortunately, we didn't choice it as our final submission. </p>\n<h4>Other experiments</h4>\n<ul>\n<li>I tried to use all bands, which can give a improvement of 0.000x. But it make the submission too complex, and take too much time. So I didn't use it in the final submission.</li>\n<li>UNet + Tranformer (256, 128: UNet, 64, 32: Transformer). Using the trasformer to get the information of other frames. It works but not better than UNet. Maybe because the transformer is not pretrained.</li>\n</ul>\n<h4>Update</h4>\n<ul>\n<li>20230811 Added a late submission score of the best single model.</li>\n</ul>",
  "messages": [
    {
      "id": 2383212,
      "postDate": "2023-08-10T08:41:25.713Z",
      "content": "<p>Thanks to the organizers for the great competition. Also thanks to my teammates: <a href=\"https://www.kaggle.com/imakarov\" target=\"_blank\">Ilya Makarov</a>. I had a lot of fun and learned a lot.<br>\nI'd like to share our solution and what I learned from this competition. I hope it will be helpful for you.</p>\n<h2>Solution overview of my part</h2>\n<h4>Models</h4>\n<p>I use 5 models as follows:</p>\n<ul>\n<li>Single frame model 1: EfficientNetV2L + UNet, 512x512, crop to 480x480 for training.</li>\n<li>Single frame model 2: EfficientNetV2L + UNet, 768x768, crop to 512x512 for training.</li>\n<li>Multi frames model 1: EfficientNetV2L + customized 3D UNet 1, 256x256, use all 8 frames.</li>\n<li>Multi frames model 2: EfficientNetV2L + customized 3D UNet 2, 512x512, crop to 480x480 for training, use 5 frames only.</li>\n<li>Multi frames model 3: EfficientNetV2L + customized 3D UNet 3, 512x512, crop to 480x480 for training, use 5 frames only.<br>\n<strong>(<a href=\"https://www.kaggle.com/wuliaokaola/icrgw-submission-single-model-lb-0-714-5th-place\" target=\"_blank\">Best single model, late submission, private LB: 0.71443, public LB 0.70803</a>)</strong></li>\n</ul>\n<p>About the customized 3D UNet<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2006644%2F2130a20eb0d042ac9d48d1ec2d7d88f9%2Funet3d_.png?generation=1691656552129914&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>Apply the same encoder(backbone) to each frame</li>\n<li>Combine the encoder outputs by Conv3D at each level. ConvLSTM2D also works, but it's slower.</li>\n<li>Use UNet as decoder to get the final output.</li>\n</ul>\n<h4>Augmentation</h4>\n<ul>\n<li>Filp LR, Filp UD, Rotate 90 (for TTA8)</li>\n<li>Noise by channel, noise by pixel</li>\n<li>Random dropout frame (multi frames model only)</li>\n<li>Random crop</li>\n</ul>\n<p>At the early stage of the competition, I noticed that if I don't use flip and rot90, the model will be overfitting quickly. But if I use the flip and rot90, the result became worse. I was confused for a long time. Finally, I found the reason. As mentioned in <a href=\"https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/discussion/430479#2382723\" target=\"_blank\">other discussions</a>, The image and the mask are not aligned. The mask is a littel bit shifted to the bottom right. So if we use flip or rot90, we need to align them. The magic number for me is: x = 0.408, y = 0.453. </p>\n<p>Tried but not worked:</p>\n<ul>\n<li>Random rotate</li>\n<li>CoarseDropout</li>\n<li>Dropout by channel</li>\n</ul>\n<h4>Target and loss function</h4>\n<p>I used 2 targets:</p>\n<ul>\n<li>Target 1: Sigmoid to 1 channel. weight: 75% (human_pixel_masks.npy)</li>\n<li>Target 2: Softmax to 5 channel. weight: 25% (human_individual_masks.npy. 0 for not contrails, 1 for 25% of labelers, 2 for 50% of labelers, 3 for 75% of labelers, 4 for 100% of labelers) </li>\n</ul>\n<p>The loss function is weighted binary cross entropy + dice loss. <br>\nAnd use Lion for optimizer.　(Lion: <a href=\"https://github.com/keras-team/keras/blob/v2.13.1/keras/optimizers/lion.py\" target=\"_blank\">https://github.com/keras-team/keras/blob/v2.13.1/keras/optimizers/lion.py</a>)</p>\n<h4>Prediction</h4>\n<ul>\n<li>Ensemble by 5 models with TTA8 (flip LR, flip UD, rotate 90)<br>\nThe TTA8 is very important. It can improve the score by 0.005~0.01.</li>\n</ul>\n<p>The best result is public LB: 0.71243, private LB: 0.71756. But unfortunately, we didn't choice it as our final submission. </p>\n<h4>Other experiments</h4>\n<ul>\n<li>I tried to use all bands, which can give a improvement of 0.000x. But it make the submission too complex, and take too much time. So I didn't use it in the final submission.</li>\n<li>UNet + Tranformer (256, 128: UNet, 64, 32: Transformer). Using the trasformer to get the information of other frames. It works but not better than UNet. Maybe because the transformer is not pretrained.</li>\n</ul>\n<h4>Update</h4>\n<ul>\n<li>20230811 Added a late submission score of the best single model.</li>\n</ul>",
      "rawMarkdown": "Thanks to the organizers for the great competition. Also thanks to my teammates: [Ilya Makarov](https://www.kaggle.com/imakarov). I had a lot of fun and learned a lot.\nI'd like to share our solution and what I learned from this competition. I hope it will be helpful for you.\n\n## Solution overview of my part\n\n#### Models\nI use 5 models as follows:\n- Single frame model 1: EfficientNetV2L + UNet, 512x512, crop to 480x480 for training.\n- Single frame model 2: EfficientNetV2L + UNet, 768x768, crop to 512x512 for training.\n- Multi frames model 1: EfficientNetV2L + customized 3D UNet 1, 256x256, use all 8 frames.\n- Multi frames model 2: EfficientNetV2L + customized 3D UNet 2, 512x512, crop to 480x480 for training, use 5 frames only.\n- Multi frames model 3: EfficientNetV2L + customized 3D UNet 3, 512x512, crop to 480x480 for training, use 5 frames only.\n**([Best single model, late submission, private LB: 0.71443, public LB 0.70803](https://www.kaggle.com/wuliaokaola/icrgw-submission-single-model-lb-0-714-5th-place))**\n\nAbout the customized 3D UNet\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2006644%2F2130a20eb0d042ac9d48d1ec2d7d88f9%2Funet3d_.png?generation=1691656552129914&alt=media)\n- Apply the same encoder(backbone) to each frame\n- Combine the encoder outputs by Conv3D at each level. ConvLSTM2D also works, but it's slower.\n- Use UNet as decoder to get the final output.\n\n#### Augmentation\n- Filp LR, Filp UD, Rotate 90 (for TTA8)\n- Noise by channel, noise by pixel\n- Random dropout frame (multi frames model only)\n- Random crop\n\nAt the early stage of the competition, I noticed that if I don't use flip and rot90, the model will be overfitting quickly. But if I use the flip and rot90, the result became worse. I was confused for a long time. Finally, I found the reason. As mentioned in [other discussions](https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/discussion/430479#2382723), The image and the mask are not aligned. The mask is a littel bit shifted to the bottom right. So if we use flip or rot90, we need to align them. The magic number for me is: x = 0.408, y = 0.453. \n\nTried but not worked:\n- Random rotate\n- CoarseDropout\n- Dropout by channel\n\n\n#### Target and loss function\nI used 2 targets:\n- Target 1: Sigmoid to 1 channel. weight: 75% (human_pixel_masks.npy)\n- Target 2: Softmax to 5 channel. weight: 25% (human_individual_masks.npy. 0 for not contrails, 1 for 25% of labelers, 2 for 50% of labelers, 3 for 75% of labelers, 4 for 100% of labelers) \n\nThe loss function is weighted binary cross entropy + dice loss. \nAnd use Lion for optimizer.　(Lion: https://github.com/keras-team/keras/blob/v2.13.1/keras/optimizers/lion.py)\n\n#### Prediction\n- Ensemble by 5 models with TTA8 (flip LR, flip UD, rotate 90)\nThe TTA8 is very important. It can improve the score by 0.005~0.01.\n\nThe best result is public LB: 0.71243, private LB: 0.71756. But unfortunately, we didn't choice it as our final submission. \n\n#### Other experiments\n- I tried to use all bands, which can give a improvement of 0.000x. But it make the submission too complex, and take too much time. So I didn't use it in the final submission.\n- UNet + Tranformer (256, 128: UNet, 64, 32: Transformer). Using the trasformer to get the information of other frames. It works but not better than UNet. Maybe because the transformer is not pretrained.\n\n#### Update\n- 20230811 Added a late submission score of the best single model.",
      "votes": 41
    },
    {
      "id": 2383878,
      "postDate": "2023-08-10T16:01:17.137Z",
      "content": "<p>I would also like to thank the organizers and Kaggle for making this interesting competition! A particular thank you goes to my great teammate <a href=\"https://www.kaggle.com/wuliaokaola\" target=\"_blank\">@wuliaokaola</a> who have discovered many ingredients of the secret sauce which turned a good solution into the top-scoring one.</p>\n<p>My part of the solution consisted of 3 Unet++ models (with <code>eca_nfnet_l1</code>, <code>dm_nfnet_f1</code>, <code>efficientnetv2-l</code> encoders) trained on images with aligned masks and 2 Unet++ models (with <code>eca_nfnet_l1</code>, <code>efficientnetv2-l</code> encoders) trained on non-aligned images. For training I used 352x352 and 448x448 ash images of timestamp 4 with <code>human_pixel_masks</code> as well as ash images of one of the timestamps 0 to 3 pseudo-labelled with an ensemble of timestamp-4 models. The pseudo-labelling turned out to be useful and improved both the local CV and the LB by ~0.01 but adding more pseudo-labelled timestamps did not improve the score any further. A combination of the focal loss with a pretty aggressive LR scheduler (constant lr of <code>4e-3</code> for 75% of the training followed by the cosine annealing) allowed to train a 4-fold (15 epochs each) model with all images from the train and validation folders with 2 timestamps in a reasonable time of 16 to 20 h on one 3090. Additionally, tuning the threshold when converting to a binary mask has improved the scores to some extent. Overall, such an ensemble scored 0.706 on the public LB and 0.696 on the private one landing it in the silver top 40.</p>",
      "rawMarkdown": "I would also like to thank the organizers and Kaggle for making this interesting competition! A particular thank you goes to my great teammate @wuliaokaola who have discovered many ingredients of the secret sauce which turned a good solution into the top-scoring one.\n\nMy part of the solution consisted of 3 Unet++ models (with `eca_nfnet_l1`, `dm_nfnet_f1`, `efficientnetv2-l` encoders) trained on images with aligned masks and 2 Unet++ models (with `eca_nfnet_l1`, `efficientnetv2-l` encoders) trained on non-aligned images. For training I used 352x352 and 448x448 ash images of timestamp 4 with `human_pixel_masks` as well as ash images of one of the timestamps 0 to 3 pseudo-labelled with an ensemble of timestamp-4 models. The pseudo-labelling turned out to be useful and improved both the local CV and the LB by ~0.01 but adding more pseudo-labelled timestamps did not improve the score any further. A combination of the focal loss with a pretty aggressive LR scheduler (constant lr of `4e-3` for 75% of the training followed by the cosine annealing) allowed to train a 4-fold (15 epochs each) model with all images from the train and validation folders with 2 timestamps in a reasonable time of 16 to 20 h on one 3090. Additionally, tuning the threshold when converting to a binary mask has improved the scores to some extent. Overall, such an ensemble scored 0.706 on the public LB and 0.696 on the private one landing it in the silver top 40.",
      "votes": 5
    },
    {
      "id": 2385632,
      "postDate": "2023-08-11T13:22:44.540Z",
      "content": "<p><a href=\"https://www.kaggle.com/wuliaokaola\" target=\"_blank\">@wuliaokaola</a> congratulations on your results!</p>\n<p>May I ask how you ended up with such precise number for the mask shift ? x = 0.408, y = 0.453 ?</p>",
      "rawMarkdown": "@wuliaokaola congratulations on your results!\n\nMay I ask how you ended up with such precise number for the mask shift ? x = 0.408, y = 0.453 ?\n",
      "votes": 1,
      "replies": [
        {
          "id": 2385847,
          "postDate": "2023-08-11T15:31:13.933Z",
          "content": "<p>Thank you. It came from a grid search.</p>",
          "rawMarkdown": "Thank you. It came from a grid search.",
          "replies": [
            {
              "id": 2385877,
              "postDate": "2023-08-11T15:49:22.633Z",
              "content": "<p>You mean you treated like a parameter and see which one works best ?</p>",
              "rawMarkdown": "You mean you treated like a parameter and see which one works best ?"
            },
            {
              "id": 2386624,
              "postDate": "2023-08-12T04:41:39.513Z",
              "content": "<ul>\n<li>Train a small and simple model without any augmentation.</li>\n<li>Inference the flipped (LR) images with the model.</li>\n<li>Shift the prediction by 2x to get the maximum score. (grid search)</li>\n</ul>\n<p>By using this method I got x = 0.408 and y = 0.453</p>",
              "rawMarkdown": "- Train a small and simple model without any augmentation.\n- Inference the flipped (LR) images with the model.\n- Shift the prediction by 2x to get the maximum score. (grid search)\n\nBy using this method I got x = 0.408 and y = 0.453",
              "votes": 1
            },
            {
              "id": 2413595,
              "postDate": "2023-08-29T03:27:16.517Z",
              "content": "<p>I`m sorry to ask may be a novice question, what  tool  did you used to implement grid search?</p>",
              "rawMarkdown": "I`m sorry to ask may be a novice question, what  tool  did you used to implement grid search?"
            },
            {
              "id": 2414953,
              "postDate": "2023-08-30T03:48:31.313Z",
              "content": "<p>It's nothing but just a for loop. </p>",
              "rawMarkdown": "It's nothing but just a for loop. "
            }
          ]
        }
      ]
    },
    {
      "id": 2383980,
      "postDate": "2023-08-10T17:31:35.873Z",
      "content": "<p>congrats guys on the gold!! <br>\ndid you train on full train set and eval on the validation.csv split from the hosts? <br>\nMay I ask what is the CV of your best single model and your ensemble please? </p>\n<blockquote>\n  <p>The best result is public LB: 0.71243, private LB: 0.71756.</p>\n</blockquote>\n<p>That is quite huge improvement from public to private </p>",
      "rawMarkdown": "congrats guys on the gold!! \ndid you train on full train set and eval on the validation.csv split from the hosts? \nMay I ask what is the CV of your best single model and your ensemble please? \n\n> The best result is public LB: 0.71243, private LB: 0.71756.\n\nThat is quite huge improvement from public to private ",
      "votes": 1,
      "replies": [
        {
          "id": 2383998,
          "postDate": "2023-08-10T17:55:30.280Z",
          "content": "<p>Thanks :-)</p>\n<p>Some models were trained on the train set only and some on both sets combined (in this case CV was OOF Dice).</p>\n<p>Regarding the CV for the best models, <a href=\"https://www.kaggle.com/wuliaokaola\" target=\"_blank\">@wuliaokaola</a> might answer this question soon.</p>\n<p>The best private LB was 26th best public LB among our submissions and we unfortunately could not recognize it as a potentially best-scoring submission.</p>",
          "rawMarkdown": "Thanks :-)\n\nSome models were trained on the train set only and some on both sets combined (in this case CV was OOF Dice).\n\nRegarding the CV for the best models, @wuliaokaola might answer this question soon.\n\nThe best private LB was 26th best public LB among our submissions and we unfortunately could not recognize it as a potentially best-scoring submission.",
          "votes": 2
        },
        {
          "id": 2384448,
          "postDate": "2023-08-11T01:23:37.627Z",
          "content": "<blockquote>\n  <p>did you train on full train set and eval on the validation.csv split from the hosts?</p>\n</blockquote>\n<p>Yes, I used validation folder as OOF.</p>\n<blockquote>\n  <p>May I ask what is the CV of your best single model and your ensemble please?</p>\n</blockquote>\n<p>The best single model score of OOF is 0.6959, and the ensemble score of OOF is 0.6993.</p>",
          "rawMarkdown": ">did you train on full train set and eval on the validation.csv split from the hosts?\n\nYes, I used validation folder as OOF.\n\n>May I ask what is the CV of your best single model and your ensemble please?\n\nThe best single model score of OOF is 0.6959, and the ensemble score of OOF is 0.6993.",
          "votes": 1,
          "replies": [
            {
              "id": 2384907,
              "postDate": "2023-08-11T05:51:18.350Z",
              "content": "<p>The LB score (late submission) of the best single model is private LB: 0.71443, public LB 0.70803.</p>",
              "rawMarkdown": "The LB score (late submission) of the best single model is private LB: 0.71443, public LB 0.70803."
            },
            {
              "id": 2385841,
              "postDate": "2023-08-11T15:24:05.977Z",
              "content": "<p>Which one configuration from your set?</p>",
              "rawMarkdown": "Which one configuration from your set?"
            },
            {
              "id": 2385845,
              "postDate": "2023-08-11T15:29:00.563Z",
              "content": "<p>Multi frames model 3: EfficientNetV2L + customized 3D UNet 3, 512x512, crop to 480x480 for training, use 5 frames only.<br>\n(Best single model, late submission, private LB: 0.71443, public LB 0.70803)</p>",
              "rawMarkdown": "Multi frames model 3: EfficientNetV2L + customized 3D UNet 3, 512x512, crop to 480x480 for training, use 5 frames only.\n(Best single model, late submission, private LB: 0.71443, public LB 0.70803)"
            },
            {
              "id": 2390302,
              "postDate": "2023-08-14T13:38:01.090Z",
              "content": "<p>So you crop the images to 480x480 during training but made inference with 512x512, right? Is this a trick to improve accuracy or just to reduce training time?</p>\n<p>Congratulations and thanks for sharing!</p>",
              "rawMarkdown": "So you crop the images to 480x480 during training but made inference with 512x512, right? Is this a trick to improve accuracy or just to reduce training time?\n\nCongratulations and thanks for sharing!"
            },
            {
              "id": 2391605,
              "postDate": "2023-08-15T08:44:07.177Z",
              "content": "<p>Yes, cropping can reduce training time. I trained the base model with 512x512, and it takes 702S per epoch. When fine-tuning the model, I cropped the images to 480x480, and it takes 597S per epoch. The accuracy was better than the base model. But I'm not sure if it's because of the cropping, because I also changed the augmentation parameters and the learning rate.</p>",
              "rawMarkdown": "Yes, cropping can reduce training time. I trained the base model with 512x512, and it takes 702S per epoch. When fine-tuning the model, I cropped the images to 480x480, and it takes 597S per epoch. The accuracy was better than the base model. But I'm not sure if it's because of the cropping, because I also changed the augmentation parameters and the learning rate.",
              "votes": 1
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2383878,
      "author_name": "Ilya Makarov",
      "author_url": "",
      "post_date": "2023-08-10T16:01:17.137000",
      "content": "<p>I would also like to thank the organizers and Kaggle for making this interesting competition! A particular thank you goes to my great teammate <a href=\"https://www.kaggle.com/wuliaokaola\" target=\"_blank\">@wuliaokaola</a> who have discovered many ingredients of the secret sauce which turned a good solution into the top-scoring one.</p>\n<p>My part of the solution consisted of 3 Unet++ models (with <code>eca_nfnet_l1</code>, <code>dm_nfnet_f1</code>, <code>efficientnetv2-l</code> encoders) trained on images with aligned masks and 2 Unet++ models (with <code>eca_nfnet_l1</code>, <code>efficientnetv2-l</code> encoders) trained on non-aligned images. For training I used 352x352 and 448x448 ash images of timestamp 4 with <code>human_pixel_masks</code> as well as ash images of one of the timestamps 0 to 3 pseudo-labelled with an ensemble of timestamp-4 models. The pseudo-labelling turned out to be useful and improved both the local CV and the LB by ~0.01 but adding more pseudo-labelled timestamps did not improve the score any further. A combination of the focal loss with a pretty aggressive LR scheduler (constant lr of <code>4e-3</code> for 75% of the training followed by the cosine annealing) allowed to train a 4-fold (15 epochs each) model with all images from the train and validation folders with 2 timestamps in a reasonable time of 16 to 20 h on one 3090. Additionally, tuning the threshold when converting to a binary mask has improved the scores to some extent. Overall, such an ensemble scored 0.706 on the public LB and 0.696 on the private one landing it in the silver top 40.</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 2385632,
      "author_name": "Optimo",
      "author_url": "",
      "post_date": "2023-08-11T13:22:44.540000",
      "content": "<p><a href=\"https://www.kaggle.com/wuliaokaola\" target=\"_blank\">@wuliaokaola</a> congratulations on your results!</p>\n<p>May I ask how you ended up with such precise number for the mask shift ? x = 0.408, y = 0.453 ?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2385847,
          "author_name": "Johnny Lee",
          "author_url": "",
          "post_date": "2023-08-11T15:31:13.933000",
          "content": "<p>Thank you. It came from a grid search.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2385877,
              "author_name": "Optimo",
              "author_url": "",
              "post_date": "2023-08-11T15:49:22.633000",
              "content": "<p>You mean you treated like a parameter and see which one works best ?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2386624,
              "author_name": "Johnny Lee",
              "author_url": "",
              "post_date": "2023-08-12T04:41:39.513000",
              "content": "<ul>\n<li>Train a small and simple model without any augmentation.</li>\n<li>Inference the flipped (LR) images with the model.</li>\n<li>Shift the prediction by 2x to get the maximum score. (grid search)</li>\n</ul>\n<p>By using this method I got x = 0.408 and y = 0.453</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2413595,
              "author_name": "kongweihao",
              "author_url": "",
              "post_date": "2023-08-29T03:27:16.517000",
              "content": "<p>I`m sorry to ask may be a novice question, what  tool  did you used to implement grid search?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2414953,
              "author_name": "Johnny Lee",
              "author_url": "",
              "post_date": "2023-08-30T03:48:31.313000",
              "content": "<p>It's nothing but just a for loop. </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2383980,
      "author_name": "Ioannis M",
      "author_url": "",
      "post_date": "2023-08-10T17:31:35.873000",
      "content": "<p>congrats guys on the gold!! <br>\ndid you train on full train set and eval on the validation.csv split from the hosts? <br>\nMay I ask what is the CV of your best single model and your ensemble please? </p>\n<blockquote>\n  <p>The best result is public LB: 0.71243, private LB: 0.71756.</p>\n</blockquote>\n<p>That is quite huge improvement from public to private </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2383998,
          "author_name": "Ilya Makarov",
          "author_url": "",
          "post_date": "2023-08-10T17:55:30.280000",
          "content": "<p>Thanks :-)</p>\n<p>Some models were trained on the train set only and some on both sets combined (in this case CV was OOF Dice).</p>\n<p>Regarding the CV for the best models, <a href=\"https://www.kaggle.com/wuliaokaola\" target=\"_blank\">@wuliaokaola</a> might answer this question soon.</p>\n<p>The best private LB was 26th best public LB among our submissions and we unfortunately could not recognize it as a potentially best-scoring submission.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2384448,
          "author_name": "Johnny Lee",
          "author_url": "",
          "post_date": "2023-08-11T01:23:37.627000",
          "content": "<blockquote>\n  <p>did you train on full train set and eval on the validation.csv split from the hosts?</p>\n</blockquote>\n<p>Yes, I used validation folder as OOF.</p>\n<blockquote>\n  <p>May I ask what is the CV of your best single model and your ensemble please?</p>\n</blockquote>\n<p>The best single model score of OOF is 0.6959, and the ensemble score of OOF is 0.6993.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2384907,
              "author_name": "Johnny Lee",
              "author_url": "",
              "post_date": "2023-08-11T05:51:18.350000",
              "content": "<p>The LB score (late submission) of the best single model is private LB: 0.71443, public LB 0.70803.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2385841,
              "author_name": "Remek Kinas",
              "author_url": "",
              "post_date": "2023-08-11T15:24:05.977000",
              "content": "<p>Which one configuration from your set?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2385845,
              "author_name": "Johnny Lee",
              "author_url": "",
              "post_date": "2023-08-11T15:29:00.563000",
              "content": "<p>Multi frames model 3: EfficientNetV2L + customized 3D UNet 3, 512x512, crop to 480x480 for training, use 5 frames only.<br>\n(Best single model, late submission, private LB: 0.71443, public LB 0.70803)</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2390302,
              "author_name": "delai50",
              "author_url": "",
              "post_date": "2023-08-14T13:38:01.090000",
              "content": "<p>So you crop the images to 480x480 during training but made inference with 512x512, right? Is this a trick to improve accuracy or just to reduce training time?</p>\n<p>Congratulations and thanks for sharing!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2391605,
              "author_name": "Johnny Lee",
              "author_url": "",
              "post_date": "2023-08-15T08:44:07.177000",
              "content": "<p>Yes, cropping can reduce training time. I trained the base model with 512x512, and it takes 702S per epoch. When fine-tuning the model, I cropped the images to 480x480, and it takes 597S per epoch. The accuracy was better than the base model. But I'm not sure if it's because of the cropping, because I also changed the augmentation parameters and the learning rate.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2383212": "Thanks to the organizers for the great competition. Also thanks to my teammates: [Ilya Makarov](https://www.kaggle.com/imakarov). I had a lot of fun and learned a lot.\nI'd like to share our solution and what I learned from this competition. I hope it will be helpful for you.\n\n## Solution overview of my part\n\n#### Models\nI use 5 models as follows:\n- Single frame model 1: EfficientNetV2L + UNet, 512x512, crop to 480x480 for training.\n- Single frame model 2: EfficientNetV2L + UNet, 768x768, crop to 512x512 for training.\n- Multi frames model 1: EfficientNetV2L + customized 3D UNet 1, 256x256, use all 8 frames.\n- Multi frames model 2: EfficientNetV2L + customized 3D UNet 2, 512x512, crop to 480x480 for training, use 5 frames only.\n- Multi frames model 3: EfficientNetV2L + customized 3D UNet 3, 512x512, crop to 480x480 for training, use 5 frames only.\n**([Best single model, late submission, private LB: 0.71443, public LB 0.70803](https://www.kaggle.com/wuliaokaola/icrgw-submission-single-model-lb-0-714-5th-place))**\n\nAbout the customized 3D UNet\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2006644%2F2130a20eb0d042ac9d48d1ec2d7d88f9%2Funet3d_.png?generation=1691656552129914&alt=media)\n- Apply the same encoder(backbone) to each frame\n- Combine the encoder outputs by Conv3D at each level. ConvLSTM2D also works, but it's slower.\n- Use UNet as decoder to get the final output.\n\n#### Augmentation\n- Filp LR, Filp UD, Rotate 90 (for TTA8)\n- Noise by channel, noise by pixel\n- Random dropout frame (multi frames model only)\n- Random crop\n\nAt the early stage of the competition, I noticed that if I don't use flip and rot90, the model will be overfitting quickly. But if I use the flip and rot90, the result became worse. I was confused for a long time. Finally, I found the reason. As mentioned in [other discussions](https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/discussion/430479#2382723), The image and the mask are not aligned. The mask is a littel bit shifted to the bottom right. So if we use flip or rot90, we need to align them. The magic number for me is: x = 0.408, y = 0.453. \n\nTried but not worked:\n- Random rotate\n- CoarseDropout\n- Dropout by channel\n\n\n#### Target and loss function\nI used 2 targets:\n- Target 1: Sigmoid to 1 channel. weight: 75% (human_pixel_masks.npy)\n- Target 2: Softmax to 5 channel. weight: 25% (human_individual_masks.npy. 0 for not contrails, 1 for 25% of labelers, 2 for 50% of labelers, 3 for 75% of labelers, 4 for 100% of labelers) \n\nThe loss function is weighted binary cross entropy + dice loss. \nAnd use Lion for optimizer.　(Lion: https://github.com/keras-team/keras/blob/v2.13.1/keras/optimizers/lion.py)\n\n#### Prediction\n- Ensemble by 5 models with TTA8 (flip LR, flip UD, rotate 90)\nThe TTA8 is very important. It can improve the score by 0.005~0.01.\n\nThe best result is public LB: 0.71243, private LB: 0.71756. But unfortunately, we didn't choice it as our final submission. \n\n#### Other experiments\n- I tried to use all bands, which can give a improvement of 0.000x. But it make the submission too complex, and take too much time. So I didn't use it in the final submission.\n- UNet + Tranformer (256, 128: UNet, 64, 32: Transformer). Using the trasformer to get the information of other frames. It works but not better than UNet. Maybe because the transformer is not pretrained.\n\n#### Update\n- 20230811 Added a late submission score of the best single model.",
    "2383878": "I would also like to thank the organizers and Kaggle for making this interesting competition! A particular thank you goes to my great teammate @wuliaokaola who have discovered many ingredients of the secret sauce which turned a good solution into the top-scoring one.\n\nMy part of the solution consisted of 3 Unet++ models (with `eca_nfnet_l1`, `dm_nfnet_f1`, `efficientnetv2-l` encoders) trained on images with aligned masks and 2 Unet++ models (with `eca_nfnet_l1`, `efficientnetv2-l` encoders) trained on non-aligned images. For training I used 352x352 and 448x448 ash images of timestamp 4 with `human_pixel_masks` as well as ash images of one of the timestamps 0 to 3 pseudo-labelled with an ensemble of timestamp-4 models. The pseudo-labelling turned out to be useful and improved both the local CV and the LB by ~0.01 but adding more pseudo-labelled timestamps did not improve the score any further. A combination of the focal loss with a pretty aggressive LR scheduler (constant lr of `4e-3` for 75% of the training followed by the cosine annealing) allowed to train a 4-fold (15 epochs each) model with all images from the train and validation folders with 2 timestamps in a reasonable time of 16 to 20 h on one 3090. Additionally, tuning the threshold when converting to a binary mask has improved the scores to some extent. Overall, such an ensemble scored 0.706 on the public LB and 0.696 on the private one landing it in the silver top 40.",
    "2385632": "@wuliaokaola congratulations on your results!\n\nMay I ask how you ended up with such precise number for the mask shift ? x = 0.408, y = 0.453 ?\n",
    "2383980": "congrats guys on the gold!! \ndid you train on full train set and eval on the validation.csv split from the hosts? \nMay I ask what is the CV of your best single model and your ensemble please? \n\n> The best result is public LB: 0.71243, private LB: 0.71756.\n\nThat is quite huge improvement from public to private "
  }
}