{
  "id": 430540,
  "title": "18th place solution",
  "url": "/competitions/google-research-identify-contrails-reduce-global-warming/discussion/430540",
  "author_name": "Jin Niu",
  "post_date": "2023-08-10T08:02:23.195000",
  "votes": 25,
  "comment_count": 8,
  "views": 0,
  "content": "<h1>18th place solution</h1>\n<p>Thanks to Google Research and Kaggle for organizing the interesting competition. Here are our summary of this competition.</p>\n<h2>Summary</h2>\n<ul>\n<li>data preprocessing: ash color images</li>\n<li>train/val split: only use the provided train/val split</li>\n<li>architecture: Unet with squeeze and excitation (SE) blocks</li>\n<li>encoders (including image size):<ul>\n<li>coat_lite_medium (256×256, 384×384, 512×512, 768×768, 1024×1024)</li>\n<li>convnext_base (512×512)</li>\n<li>convnextv2_base (384×384, 512×512)</li>\n<li>convnextv2_tiny (512×512)</li>\n<li>resnest101e (512×512)</li>\n<li>tf_efficientnet_b6 (512×512)</li>\n<li>tf_efficientnetv2_s (512×512)</li></ul></li>\n<li>loss: dice loss</li>\n<li>no tta</li>\n</ul>\n<h2>Model</h2>\n<p>To reduce the computational cost, at the beginning of the competition, my teammate and I mainly conducted our own experiments on Unet + efficientnet-b0 independently, and the best model achieved a validation score of 0.6540 and a public leaderboard score of 0.666. Later, we used the well-performed method on more powerful encoders such as coat and convnext. Our best single model is Unet + coat_lite_medium, and the final ensemble model is Unet + coat_lite_medium, convnext_base, convnextv2_base, convnextv2_tiny, resnest101e, tf_efficientnet_b6 and tf_efficientnetv2_s in some different image size.</p>\n<h2>Training strategy</h2>\n<p>Considering that training with large images has better performance but takes longer time, we adopted the method of progressive learning, starting with small images and gradually increasing the size. Taking the coat_lite_medium (512×512) as an example, the model is trained in 5 stages (256-&gt;320-&gt;384-&gt;448-&gt;512). Flip, rotate90 and transpose augmentations cause worse training performance, and relevant tta also brought worse lb score, which confused us for a long time. So we only used affine transform with minor rotation to train the model and get validation score 0.6756 at an earlier time.</p>\n<p>I thought these flipped and tranposed images are somewhat different from the original images, although I don't quite understand the reason. (Tascj and hengck23 have pointed out the issue and I will have a try to address the problem later.)</p>\n<p>These flipped and transposed images can be regarded as samples with flaw. Pretraining on flawed data and finetune on good data is often an effective approach, therefore, we didn't give up flip and transpose augmentations, and one more stage was added to the training process: in the first 5 stages, flip, transpose, affine and perspective augmentations were applied; in the last stage, only affine (or perspective) augmentations were utilized. It improved our validation score from 0.6756 to 0.6838.</p>\n<p>Main settings of the training stages are as follows:</p>\n<table>\n<thead>\n<tr>\n<th>stage</th>\n<th>image size</th>\n<th>epochs</th>\n<th>batch size</th>\n<th>initial lr</th>\n<th>augmentations</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>256×256</td>\n<td>30</td>\n<td>64</td>\n<td>6e-4</td>\n<td>h flip, v flip, transpose, affine (or perspective)</td>\n</tr>\n<tr>\n<td>2</td>\n<td>320×320</td>\n<td>6</td>\n<td>40</td>\n<td>4.5e-4</td>\n<td>h flip, v flip, transpose, affine (or perspective)</td>\n</tr>\n<tr>\n<td>3</td>\n<td>384×384</td>\n<td>4</td>\n<td>28</td>\n<td>3e-4</td>\n<td>h flip, v flip, transpose, affine (or perspective)</td>\n</tr>\n<tr>\n<td>4</td>\n<td>448×448</td>\n<td>4</td>\n<td>20</td>\n<td>2e-4</td>\n<td>h flip, v flip, transpose, affine (or perspective)</td>\n</tr>\n<tr>\n<td>5</td>\n<td>512×512</td>\n<td>8</td>\n<td>16</td>\n<td>1.5e-4</td>\n<td>h flip, v flip, transpose, affine (or perspective)</td>\n</tr>\n<tr>\n<td>6</td>\n<td>512×512</td>\n<td>8</td>\n<td>16</td>\n<td>7.5e-5</td>\n<td>affine (or perspective)</td>\n</tr>\n</tbody>\n</table>\n<h2>Loss function</h2>\n<p>smp.losses.DiceLoss or a very simple batch dice loss</p>\n<pre><code> ():\n    y_pred = y_pred.view(-)\n    y_true = y_true.view(-)\n     - ( * (y_true * y_pred).()) / ((y_true.() + y_pred.()) + smooth)\n</code></pre>\n<h2>Things didn't work (maybe we didn't do it right)</h2>\n<ul>\n<li>color augmentation (RandomBrightnessContrast, HueSaturationValue, RandomGamma, etc.)</li>\n<li>mixup augmentation (mixup of two images and their masks)</li>\n<li>larger image size (768×768, 1024×1024)</li>\n<li>Pseudo labels are used for the images before and after the labeled images. We tried to train on pseudo labels and finetune on labeled images, but it didn't have noticeable effect.</li>\n<li>We tried to integrate the data at multiple time points to the input, but it didn't work.</li>\n</ul>\n<p>Finaly, thanks to my teammate for bringing me a pleasant competition experience.</p>",
  "messages": [
    {
      "id": 2383170,
      "postDate": "2023-08-10T08:02:23.197Z",
      "content": "<h1>18th place solution</h1>\n<p>Thanks to Google Research and Kaggle for organizing the interesting competition. Here are our summary of this competition.</p>\n<h2>Summary</h2>\n<ul>\n<li>data preprocessing: ash color images</li>\n<li>train/val split: only use the provided train/val split</li>\n<li>architecture: Unet with squeeze and excitation (SE) blocks</li>\n<li>encoders (including image size):<ul>\n<li>coat_lite_medium (256×256, 384×384, 512×512, 768×768, 1024×1024)</li>\n<li>convnext_base (512×512)</li>\n<li>convnextv2_base (384×384, 512×512)</li>\n<li>convnextv2_tiny (512×512)</li>\n<li>resnest101e (512×512)</li>\n<li>tf_efficientnet_b6 (512×512)</li>\n<li>tf_efficientnetv2_s (512×512)</li></ul></li>\n<li>loss: dice loss</li>\n<li>no tta</li>\n</ul>\n<h2>Model</h2>\n<p>To reduce the computational cost, at the beginning of the competition, my teammate and I mainly conducted our own experiments on Unet + efficientnet-b0 independently, and the best model achieved a validation score of 0.6540 and a public leaderboard score of 0.666. Later, we used the well-performed method on more powerful encoders such as coat and convnext. Our best single model is Unet + coat_lite_medium, and the final ensemble model is Unet + coat_lite_medium, convnext_base, convnextv2_base, convnextv2_tiny, resnest101e, tf_efficientnet_b6 and tf_efficientnetv2_s in some different image size.</p>\n<h2>Training strategy</h2>\n<p>Considering that training with large images has better performance but takes longer time, we adopted the method of progressive learning, starting with small images and gradually increasing the size. Taking the coat_lite_medium (512×512) as an example, the model is trained in 5 stages (256-&gt;320-&gt;384-&gt;448-&gt;512). Flip, rotate90 and transpose augmentations cause worse training performance, and relevant tta also brought worse lb score, which confused us for a long time. So we only used affine transform with minor rotation to train the model and get validation score 0.6756 at an earlier time.</p>\n<p>I thought these flipped and tranposed images are somewhat different from the original images, although I don't quite understand the reason. (Tascj and hengck23 have pointed out the issue and I will have a try to address the problem later.)</p>\n<p>These flipped and transposed images can be regarded as samples with flaw. Pretraining on flawed data and finetune on good data is often an effective approach, therefore, we didn't give up flip and transpose augmentations, and one more stage was added to the training process: in the first 5 stages, flip, transpose, affine and perspective augmentations were applied; in the last stage, only affine (or perspective) augmentations were utilized. It improved our validation score from 0.6756 to 0.6838.</p>\n<p>Main settings of the training stages are as follows:</p>\n<table>\n<thead>\n<tr>\n<th>stage</th>\n<th>image size</th>\n<th>epochs</th>\n<th>batch size</th>\n<th>initial lr</th>\n<th>augmentations</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>256×256</td>\n<td>30</td>\n<td>64</td>\n<td>6e-4</td>\n<td>h flip, v flip, transpose, affine (or perspective)</td>\n</tr>\n<tr>\n<td>2</td>\n<td>320×320</td>\n<td>6</td>\n<td>40</td>\n<td>4.5e-4</td>\n<td>h flip, v flip, transpose, affine (or perspective)</td>\n</tr>\n<tr>\n<td>3</td>\n<td>384×384</td>\n<td>4</td>\n<td>28</td>\n<td>3e-4</td>\n<td>h flip, v flip, transpose, affine (or perspective)</td>\n</tr>\n<tr>\n<td>4</td>\n<td>448×448</td>\n<td>4</td>\n<td>20</td>\n<td>2e-4</td>\n<td>h flip, v flip, transpose, affine (or perspective)</td>\n</tr>\n<tr>\n<td>5</td>\n<td>512×512</td>\n<td>8</td>\n<td>16</td>\n<td>1.5e-4</td>\n<td>h flip, v flip, transpose, affine (or perspective)</td>\n</tr>\n<tr>\n<td>6</td>\n<td>512×512</td>\n<td>8</td>\n<td>16</td>\n<td>7.5e-5</td>\n<td>affine (or perspective)</td>\n</tr>\n</tbody>\n</table>\n<h2>Loss function</h2>\n<p>smp.losses.DiceLoss or a very simple batch dice loss</p>\n<pre><code> ():\n    y_pred = y_pred.view(-)\n    y_true = y_true.view(-)\n     - ( * (y_true * y_pred).()) / ((y_true.() + y_pred.()) + smooth)\n</code></pre>\n<h2>Things didn't work (maybe we didn't do it right)</h2>\n<ul>\n<li>color augmentation (RandomBrightnessContrast, HueSaturationValue, RandomGamma, etc.)</li>\n<li>mixup augmentation (mixup of two images and their masks)</li>\n<li>larger image size (768×768, 1024×1024)</li>\n<li>Pseudo labels are used for the images before and after the labeled images. We tried to train on pseudo labels and finetune on labeled images, but it didn't have noticeable effect.</li>\n<li>We tried to integrate the data at multiple time points to the input, but it didn't work.</li>\n</ul>\n<p>Finaly, thanks to my teammate for bringing me a pleasant competition experience.</p>",
      "rawMarkdown": "# 18th place solution\n\nThanks to Google Research and Kaggle for organizing the interesting competition. Here are our summary of this competition.\n\n## Summary\n\n- data preprocessing: ash color images\n- train/val split: only use the provided train/val split\n- architecture: Unet with squeeze and excitation (SE) blocks\n- encoders (including image size):\n  - coat_lite_medium (256×256, 384×384, 512×512, 768×768, 1024×1024)\n  - convnext_base (512×512)\n  - convnextv2_base (384×384, 512×512)\n  - convnextv2_tiny (512×512)\n  - resnest101e (512×512)\n  - tf_efficientnet_b6 (512×512)\n  - tf_efficientnetv2_s (512×512)\n- loss: dice loss\n- no tta\n\n## Model\n\nTo reduce the computational cost, at the beginning of the competition, my teammate and I mainly conducted our own experiments on Unet + efficientnet-b0 independently, and the best model achieved a validation score of 0.6540 and a public leaderboard score of 0.666. Later, we used the well-performed method on more powerful encoders such as coat and convnext. Our best single model is Unet + coat_lite_medium, and the final ensemble model is Unet + coat_lite_medium, convnext_base, convnextv2_base, convnextv2_tiny, resnest101e, tf_efficientnet_b6 and tf_efficientnetv2_s in some different image size.\n\n## Training strategy\n\nConsidering that training with large images has better performance but takes longer time, we adopted the method of progressive learning, starting with small images and gradually increasing the size. Taking the coat_lite_medium (512×512) as an example, the model is trained in 5 stages (256->320->384->448->512). Flip, rotate90 and transpose augmentations cause worse training performance, and relevant tta also brought worse lb score, which confused us for a long time. So we only used affine transform with minor rotation to train the model and get validation score 0.6756 at an earlier time.\n\nI thought these flipped and tranposed images are somewhat different from the original images, although I don't quite understand the reason. (Tascj and hengck23 have pointed out the issue and I will have a try to address the problem later.)\n\nThese flipped and transposed images can be regarded as samples with flaw. Pretraining on flawed data and finetune on good data is often an effective approach, therefore, we didn't give up flip and transpose augmentations, and one more stage was added to the training process: in the first 5 stages, flip, transpose, affine and perspective augmentations were applied; in the last stage, only affine (or perspective) augmentations were utilized. It improved our validation score from 0.6756 to 0.6838.\n\nMain settings of the training stages are as follows:\n\n| stage | image size | epochs | batch size | initial lr | augmentations |\n| :----: | :----: | :----: | :----: | :----: | :----: |\n| 1 | 256×256 | 30 | 64 | 6e-4 | h flip, v flip, transpose, affine (or perspective) |\n| 2 | 320×320 | 6 | 40 | 4.5e-4 | h flip, v flip, transpose, affine (or perspective) |\n| 3 | 384×384 | 4 | 28 | 3e-4 | h flip, v flip, transpose, affine (or perspective) |\n| 4 | 448×448 | 4 | 20 | 2e-4 | h flip, v flip, transpose, affine (or perspective) |\n| 5 | 512×512 | 8 | 16 | 1.5e-4 | h flip, v flip, transpose, affine (or perspective) |\n| 6 | 512×512 | 8 | 16 | 7.5e-5 | affine (or perspective) |\n\n## Loss function\n\nsmp.losses.DiceLoss or a very simple batch dice loss\n\n```python\ndef batch_dice_loss(y_pred, y_true, smooth=0.01):\n    y_pred = y_pred.view(-1)\n    y_true = y_true.view(-1)\n    return - (2. * (y_true * y_pred).sum()) / ((y_true.sum() + y_pred.sum()) + smooth)\n```\n\n## Things didn't work (maybe we didn't do it right)\n\n- color augmentation (RandomBrightnessContrast, HueSaturationValue, RandomGamma, etc.)\n- mixup augmentation (mixup of two images and their masks)\n- larger image size (768×768, 1024×1024)\n- Pseudo labels are used for the images before and after the labeled images. We tried to train on pseudo labels and finetune on labeled images, but it didn't have noticeable effect.\n- We tried to integrate the data at multiple time points to the input, but it didn't work.\n\nFinaly, thanks to my teammate for bringing me a pleasant competition experience.",
      "votes": 24
    },
    {
      "id": 2385073,
      "postDate": "2023-08-11T06:59:32.487Z",
      "content": "<p>Thank you for a nice write-up.<br>\nI have a question. You said, \"I mainly conducted our own experiments on Unet + efficientnet-b0\", but what kind of experiment is done? I also start with small-scale experiments, but ideas that worked fine with small models frequently fail when I change into larger models because of un-scalability. Could you kindly give me any advice?</p>",
      "rawMarkdown": "Thank you for a nice write-up.\nI have a question. You said, \"I mainly conducted our own experiments on Unet + efficientnet-b0\", but what kind of experiment is done? I also start with small-scale experiments, but ideas that worked fine with small models frequently fail when I change into larger models because of un-scalability. Could you kindly give me any advice?",
      "votes": 3,
      "replies": [
        {
          "id": 2386581,
          "postDate": "2023-08-12T04:00:39.633Z",
          "content": "<p>It's true that some methods that have shown improvement on small models may not necessarily yield positive results on larger models. Due to the differences in model parameter size and complexity, appropriate adjustments in learning rate and weight decay may be required. However, suitable image augmentation techniques are generally effective for all models, and most of our experiments involve adjusting augmentations at different training stages.</p>",
          "rawMarkdown": "It's true that some methods that have shown improvement on small models may not necessarily yield positive results on larger models. Due to the differences in model parameter size and complexity, appropriate adjustments in learning rate and weight decay may be required. However, suitable image augmentation techniques are generally effective for all models, and most of our experiments involve adjusting augmentations at different training stages.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2383328,
      "postDate": "2023-08-10T10:01:49.607Z",
      "content": "<p>Congratulations! Really nice score! </p>\n<p>One question regarding Unet + Convnext - what repo did you use to combine Unet/Convnext? (MMSegmentation, SMP or your custom implementation)? Thank you!</p>",
      "rawMarkdown": "Congratulations! Really nice score! \n\nOne question regarding Unet + Convnext - what repo did you use to combine Unet/Convnext? (MMSegmentation, SMP or your custom implementation)? Thank you!",
      "votes": 4,
      "replies": [
        {
          "id": 2383423,
          "postDate": "2023-08-10T10:51:05.453Z",
          "content": "<p>We implemented our own version of Unet, similar to the top solution in \"Hacking the Human Body\" competition. (<a href=\"https://www.kaggle.com/code/victorsd/2nd-place-inference\" target=\"_blank\">https://www.kaggle.com/code/victorsd/2nd-place-inference</a>)</p>\n<p>When creating models based on timm, use <code>features_only=True</code> to extract feature maps at each stride to build the Unet. </p>\n<pre><code>encoder = timm.create_model(encoder_name, features_only=, pretrained=)\nfm1, fm2, fm3, fm4, fm5 = encoder(x)\n</code></pre>",
          "rawMarkdown": "We implemented our own version of Unet, similar to the top solution in \"Hacking the Human Body\" competition. (https://www.kaggle.com/code/victorsd/2nd-place-inference)\n\nWhen creating models based on timm, use `features_only=True` to extract feature maps at each stride to build the Unet. \n\n```python\nencoder = timm.create_model(encoder_name, features_only=True, pretrained=True)\nfm1, fm2, fm3, fm4, fm5 = encoder(x)\n```",
          "votes": 9,
          "replies": [
            {
              "id": 2383472,
              "postDate": "2023-08-10T11:33:02.390Z",
              "content": "<p>Thank you a lot! <br>\nOnce again - congratulations!</p>",
              "rawMarkdown": "Thank you a lot! \nOnce again - congratulations!",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2385313,
      "postDate": "2023-08-11T09:06:51.380Z",
      "content": "<p>Congratulations and thanks for sharing! </p>\n<p>(Without taking into account the flipping/transpose augmentations) Based on your general experience, what is the difference in accuracy between doing progressive learning and training directly on the highest resolution?</p>\n<p>Thanks!</p>",
      "rawMarkdown": "Congratulations and thanks for sharing! \n\n(Without taking into account the flipping/transpose augmentations) Based on your general experience, what is the difference in accuracy between doing progressive learning and training directly on the highest resolution?\n\nThanks!",
      "votes": 1,
      "replies": [
        {
          "id": 2386593,
          "postDate": "2023-08-12T04:16:54.630Z",
          "content": "<p>In theory, progressive learning is expected to better adapt to multi-scale features. However, in our limited comparative experiments, the results on the validation set for both methods were similar overall. The main advantage of the progressive learning approach is that it helps us save a significant amount of time, approximately half the time required for 512×512 images in one single model.</p>",
          "rawMarkdown": "In theory, progressive learning is expected to better adapt to multi-scale features. However, in our limited comparative experiments, the results on the validation set for both methods were similar overall. The main advantage of the progressive learning approach is that it helps us save a significant amount of time, approximately half the time required for 512×512 images in one single model.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2386599,
      "postDate": "2023-08-12T04:20:37.457Z",
      "content": "<p>Thanks for Sharing!</p>\n<p>It is very helpful.</p>",
      "rawMarkdown": "Thanks for Sharing!\n\nIt is very helpful."
    }
  ],
  "comments": [
    {
      "id": 2385073,
      "author_name": "luddite^",
      "author_url": "",
      "post_date": "2023-08-11T06:59:32.487000",
      "content": "<p>Thank you for a nice write-up.<br>\nI have a question. You said, \"I mainly conducted our own experiments on Unet + efficientnet-b0\", but what kind of experiment is done? I also start with small-scale experiments, but ideas that worked fine with small models frequently fail when I change into larger models because of un-scalability. Could you kindly give me any advice?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2386581,
          "author_name": "Jin Niu",
          "author_url": "",
          "post_date": "2023-08-12T04:00:39.633000",
          "content": "<p>It's true that some methods that have shown improvement on small models may not necessarily yield positive results on larger models. Due to the differences in model parameter size and complexity, appropriate adjustments in learning rate and weight decay may be required. However, suitable image augmentation techniques are generally effective for all models, and most of our experiments involve adjusting augmentations at different training stages.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2383328,
      "author_name": "Remek Kinas",
      "author_url": "",
      "post_date": "2023-08-10T10:01:49.607000",
      "content": "<p>Congratulations! Really nice score! </p>\n<p>One question regarding Unet + Convnext - what repo did you use to combine Unet/Convnext? (MMSegmentation, SMP or your custom implementation)? Thank you!</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2383423,
          "author_name": "Jin Niu",
          "author_url": "",
          "post_date": "2023-08-10T10:51:05.453000",
          "content": "<p>We implemented our own version of Unet, similar to the top solution in \"Hacking the Human Body\" competition. (<a href=\"https://www.kaggle.com/code/victorsd/2nd-place-inference\" target=\"_blank\">https://www.kaggle.com/code/victorsd/2nd-place-inference</a>)</p>\n<p>When creating models based on timm, use <code>features_only=True</code> to extract feature maps at each stride to build the Unet. </p>\n<pre><code>encoder = timm.create_model(encoder_name, features_only=, pretrained=)\nfm1, fm2, fm3, fm4, fm5 = encoder(x)\n</code></pre>",
          "votes": 9,
          "replies": [
            {
              "id": 2383472,
              "author_name": "Remek Kinas",
              "author_url": "",
              "post_date": "2023-08-10T11:33:02.390000",
              "content": "<p>Thank you a lot! <br>\nOnce again - congratulations!</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2385313,
      "author_name": "delai50",
      "author_url": "",
      "post_date": "2023-08-11T09:06:51.380000",
      "content": "<p>Congratulations and thanks for sharing! </p>\n<p>(Without taking into account the flipping/transpose augmentations) Based on your general experience, what is the difference in accuracy between doing progressive learning and training directly on the highest resolution?</p>\n<p>Thanks!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2386593,
          "author_name": "Jin Niu",
          "author_url": "",
          "post_date": "2023-08-12T04:16:54.630000",
          "content": "<p>In theory, progressive learning is expected to better adapt to multi-scale features. However, in our limited comparative experiments, the results on the validation set for both methods were similar overall. The main advantage of the progressive learning approach is that it helps us save a significant amount of time, approximately half the time required for 512×512 images in one single model.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2386599,
      "author_name": "Kanika Thakur",
      "author_url": "",
      "post_date": "2023-08-12T04:20:37.457000",
      "content": "<p>Thanks for Sharing!</p>\n<p>It is very helpful.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2383170": "# 18th place solution\n\nThanks to Google Research and Kaggle for organizing the interesting competition. Here are our summary of this competition.\n\n## Summary\n\n- data preprocessing: ash color images\n- train/val split: only use the provided train/val split\n- architecture: Unet with squeeze and excitation (SE) blocks\n- encoders (including image size):\n  - coat_lite_medium (256×256, 384×384, 512×512, 768×768, 1024×1024)\n  - convnext_base (512×512)\n  - convnextv2_base (384×384, 512×512)\n  - convnextv2_tiny (512×512)\n  - resnest101e (512×512)\n  - tf_efficientnet_b6 (512×512)\n  - tf_efficientnetv2_s (512×512)\n- loss: dice loss\n- no tta\n\n## Model\n\nTo reduce the computational cost, at the beginning of the competition, my teammate and I mainly conducted our own experiments on Unet + efficientnet-b0 independently, and the best model achieved a validation score of 0.6540 and a public leaderboard score of 0.666. Later, we used the well-performed method on more powerful encoders such as coat and convnext. Our best single model is Unet + coat_lite_medium, and the final ensemble model is Unet + coat_lite_medium, convnext_base, convnextv2_base, convnextv2_tiny, resnest101e, tf_efficientnet_b6 and tf_efficientnetv2_s in some different image size.\n\n## Training strategy\n\nConsidering that training with large images has better performance but takes longer time, we adopted the method of progressive learning, starting with small images and gradually increasing the size. Taking the coat_lite_medium (512×512) as an example, the model is trained in 5 stages (256->320->384->448->512). Flip, rotate90 and transpose augmentations cause worse training performance, and relevant tta also brought worse lb score, which confused us for a long time. So we only used affine transform with minor rotation to train the model and get validation score 0.6756 at an earlier time.\n\nI thought these flipped and tranposed images are somewhat different from the original images, although I don't quite understand the reason. (Tascj and hengck23 have pointed out the issue and I will have a try to address the problem later.)\n\nThese flipped and transposed images can be regarded as samples with flaw. Pretraining on flawed data and finetune on good data is often an effective approach, therefore, we didn't give up flip and transpose augmentations, and one more stage was added to the training process: in the first 5 stages, flip, transpose, affine and perspective augmentations were applied; in the last stage, only affine (or perspective) augmentations were utilized. It improved our validation score from 0.6756 to 0.6838.\n\nMain settings of the training stages are as follows:\n\n| stage | image size | epochs | batch size | initial lr | augmentations |\n| :----: | :----: | :----: | :----: | :----: | :----: |\n| 1 | 256×256 | 30 | 64 | 6e-4 | h flip, v flip, transpose, affine (or perspective) |\n| 2 | 320×320 | 6 | 40 | 4.5e-4 | h flip, v flip, transpose, affine (or perspective) |\n| 3 | 384×384 | 4 | 28 | 3e-4 | h flip, v flip, transpose, affine (or perspective) |\n| 4 | 448×448 | 4 | 20 | 2e-4 | h flip, v flip, transpose, affine (or perspective) |\n| 5 | 512×512 | 8 | 16 | 1.5e-4 | h flip, v flip, transpose, affine (or perspective) |\n| 6 | 512×512 | 8 | 16 | 7.5e-5 | affine (or perspective) |\n\n## Loss function\n\nsmp.losses.DiceLoss or a very simple batch dice loss\n\n```python\ndef batch_dice_loss(y_pred, y_true, smooth=0.01):\n    y_pred = y_pred.view(-1)\n    y_true = y_true.view(-1)\n    return - (2. * (y_true * y_pred).sum()) / ((y_true.sum() + y_pred.sum()) + smooth)\n```\n\n## Things didn't work (maybe we didn't do it right)\n\n- color augmentation (RandomBrightnessContrast, HueSaturationValue, RandomGamma, etc.)\n- mixup augmentation (mixup of two images and their masks)\n- larger image size (768×768, 1024×1024)\n- Pseudo labels are used for the images before and after the labeled images. We tried to train on pseudo labels and finetune on labeled images, but it didn't have noticeable effect.\n- We tried to integrate the data at multiple time points to the input, but it didn't work.\n\nFinaly, thanks to my teammate for bringing me a pleasant competition experience.",
    "2385073": "Thank you for a nice write-up.\nI have a question. You said, \"I mainly conducted our own experiments on Unet + efficientnet-b0\", but what kind of experiment is done? I also start with small-scale experiments, but ideas that worked fine with small models frequently fail when I change into larger models because of un-scalability. Could you kindly give me any advice?",
    "2383328": "Congratulations! Really nice score! \n\nOne question regarding Unet + Convnext - what repo did you use to combine Unet/Convnext? (MMSegmentation, SMP or your custom implementation)? Thank you!",
    "2385313": "Congratulations and thanks for sharing! \n\n(Without taking into account the flipping/transpose augmentations) Based on your general experience, what is the difference in accuracy between doing progressive learning and training directly on the highest resolution?\n\nThanks!",
    "2386599": "Thanks for Sharing!\n\nIt is very helpful."
  }
}