{
  "id": 430483,
  "title": "15th place solution",
  "url": "/competitions/google-research-identify-contrails-reduce-global-warming/discussion/430483",
  "author_name": "stakahashi",
  "post_date": "2023-08-10T02:33:40.569000",
  "votes": 22,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Thank you for organizers hosting the competition. I also appreciate my team mates cooperation.</p>\n<h1>Summary</h1>\n<p>Our solution has nothing special, it's just ensemble of 4 EfficientNetV2 with different image sizes.</p>\n<ul>\n<li>2x l with 768x768 image size (different seeds)</li>\n<li>1x l with 1024x1024 image size</li>\n<li>1x s with 1536x1536 image size</li>\n</ul>\n<p>First 3 models give similar score around 0.690 ~ 0.692 for <code>validation</code> folder images and the last gives 0.685. Ensemble score is 0.6936.</p>\n<h1>What worked</h1>\n<ul>\n<li>Augmentations<ul>\n<li>HV Flip</li>\n<li>Rotate</li>\n<li>Random scale then crop</li></ul></li>\n<li>TTA<ul>\n<li>HV Flip</li>\n<li>Rotate</li></ul></li>\n<li>set <code>\"drop_path_rate\": 0.2</code> and <code>\"drop_rate\": 0.2</code></li>\n<li>Adding time steps 3 and 5 as pseudo labels</li>\n<li>Focal loss</li>\n</ul>\n<p>But I guess improvements from above things are pretty small compared with addressing label misalignment and using individual human annotated mask as new labels as mentioned other discussions (I'm plan to do experiment later).</p>\n<h1>What not worked</h1>\n<ul>\n<li>Other augmentations<ul>\n<li>ColorJitter</li>\n<li>HueSaturationValue</li>\n<li>RandomBrightnessContrast</li>\n<li>RandomGamma</li>\n<li>RandomFog</li>\n<li>RandomShadow</li>\n<li>CoarseDropout</li>\n<li>MixUp</li></ul></li>\n<li>Adding more time steps as pseudo labels</li>\n<li>Adding pseudo label rounds</li>\n<li>I'm really focused on training 2.5D and 3D models, but nothing worked at all.<ul>\n<li>Just 3D model (e.g. Resnet3D)</li>\n<li>3D encoder + 2.5D decoder</li>\n<li>Input other time frames predictions as additional channels</li>\n<li>Stack 2 Unets, first predicts segmentation masks for multiple time frames, second takes those masks as input by stacking channel dimensions</li></ul></li>\n</ul>",
  "messages": [
    {
      "id": 2382754,
      "postDate": "2023-08-10T02:33:40.570Z",
      "content": "<p>Thank you for organizers hosting the competition. I also appreciate my team mates cooperation.</p>\n<h1>Summary</h1>\n<p>Our solution has nothing special, it's just ensemble of 4 EfficientNetV2 with different image sizes.</p>\n<ul>\n<li>2x l with 768x768 image size (different seeds)</li>\n<li>1x l with 1024x1024 image size</li>\n<li>1x s with 1536x1536 image size</li>\n</ul>\n<p>First 3 models give similar score around 0.690 ~ 0.692 for <code>validation</code> folder images and the last gives 0.685. Ensemble score is 0.6936.</p>\n<h1>What worked</h1>\n<ul>\n<li>Augmentations<ul>\n<li>HV Flip</li>\n<li>Rotate</li>\n<li>Random scale then crop</li></ul></li>\n<li>TTA<ul>\n<li>HV Flip</li>\n<li>Rotate</li></ul></li>\n<li>set <code>\"drop_path_rate\": 0.2</code> and <code>\"drop_rate\": 0.2</code></li>\n<li>Adding time steps 3 and 5 as pseudo labels</li>\n<li>Focal loss</li>\n</ul>\n<p>But I guess improvements from above things are pretty small compared with addressing label misalignment and using individual human annotated mask as new labels as mentioned other discussions (I'm plan to do experiment later).</p>\n<h1>What not worked</h1>\n<ul>\n<li>Other augmentations<ul>\n<li>ColorJitter</li>\n<li>HueSaturationValue</li>\n<li>RandomBrightnessContrast</li>\n<li>RandomGamma</li>\n<li>RandomFog</li>\n<li>RandomShadow</li>\n<li>CoarseDropout</li>\n<li>MixUp</li></ul></li>\n<li>Adding more time steps as pseudo labels</li>\n<li>Adding pseudo label rounds</li>\n<li>I'm really focused on training 2.5D and 3D models, but nothing worked at all.<ul>\n<li>Just 3D model (e.g. Resnet3D)</li>\n<li>3D encoder + 2.5D decoder</li>\n<li>Input other time frames predictions as additional channels</li>\n<li>Stack 2 Unets, first predicts segmentation masks for multiple time frames, second takes those masks as input by stacking channel dimensions</li></ul></li>\n</ul>",
      "rawMarkdown": "Thank you for organizers hosting the competition. I also appreciate my team mates cooperation.\n\n# Summary\n\nOur solution has nothing special, it's just ensemble of 4 EfficientNetV2 with different image sizes.\n\n- 2x l with 768x768 image size (different seeds)\n- 1x l with 1024x1024 image size\n- 1x s with 1536x1536 image size\n\nFirst 3 models give similar score around 0.690 ~ 0.692 for `validation` folder images and the last gives 0.685. Ensemble score is 0.6936.\n\n# What worked\n\n- Augmentations\n    - HV Flip\n    - Rotate\n    - Random scale then crop\n- TTA\n    - HV Flip\n    - Rotate\n- set `\"drop_path_rate\": 0.2` and `\"drop_rate\": 0.2`\n- Adding time steps 3 and 5 as pseudo labels\n- Focal loss\n\nBut I guess improvements from above things are pretty small compared with addressing label misalignment and using individual human annotated mask as new labels as mentioned other discussions (I'm plan to do experiment later).\n\n# What not worked\n\n- Other augmentations\n    - ColorJitter\n    - HueSaturationValue\n    - RandomBrightnessContrast\n    - RandomGamma\n    - RandomFog\n    - RandomShadow\n    - CoarseDropout\n    - MixUp\n- Adding more time steps as pseudo labels\n- Adding pseudo label rounds\n- I'm really focused on training 2.5D and 3D models, but nothing worked at all.\n    - Just 3D model (e.g. Resnet3D)\n    - 3D encoder + 2.5D decoder\n    - Input other time frames predictions as additional channels\n    - Stack 2 Unets, first predicts segmentation masks for multiple time frames, second takes those masks as input by stacking channel dimensions",
      "votes": 22
    },
    {
      "id": 2383301,
      "postDate": "2023-08-10T09:43:50.747Z",
      "content": "<p>When you trained on individual annotations, did you just add them to the whole dataset or did you pass as the input to the model several images at once? </p>\n<p>Also, how did you manage flips and rotations augmentations to work?</p>",
      "rawMarkdown": "When you trained on individual annotations, did you just add them to the whole dataset or did you pass as the input to the model several images at once? \n\nAlso, how did you manage flips and rotations augmentations to work?",
      "votes": 1,
      "replies": [
        {
          "id": 2383571,
          "postDate": "2023-08-10T12:58:34.113Z",
          "content": "<p>For the first question, I haven't trained on individual annotations. Just ground truth label is used.</p>\n<p>For the second, I'm not sure why they are worked but I guess</p>\n<ul>\n<li>Pseudo label mitigates label issue because pseudo label may not be misaligned.</li>\n<li>Also large image size may alleviates label issue because the effect of label misalignment is maybe smaller for large masks compared with small one.</li>\n</ul>",
          "rawMarkdown": "For the first question, I haven't trained on individual annotations. Just ground truth label is used.\n\nFor the second, I'm not sure why they are worked but I guess\n\n- Pseudo label mitigates label issue because pseudo label may not be misaligned.\n- Also large image size may alleviates label issue because the effect of label misalignment is maybe smaller for large masks compared with small one.",
          "votes": 2,
          "replies": [
            {
              "id": 2383592,
              "postDate": "2023-08-10T13:08:31.703Z",
              "content": "<p>Oh, I have meant that you used pseudo labels for other frames. Did you just add them to the dataset?</p>",
              "rawMarkdown": "Oh, I have meant that you used pseudo labels for other frames. Did you just add them to the dataset?"
            },
            {
              "id": 2390120,
              "postDate": "2023-08-14T11:44:08.657Z",
              "content": "<p>Sorry for my late reply.<br>\nYes, pseudo labels are just added to the dataset.</p>",
              "rawMarkdown": "Sorry for my late reply.\nYes, pseudo labels are just added to the dataset."
            }
          ]
        }
      ]
    },
    {
      "id": 2382828,
      "postDate": "2023-08-10T04:12:33.703Z",
      "content": "<p>Congratulations for achieving top position in the competition.<br>\nThanks for sharing valuable information. </p>",
      "rawMarkdown": "Congratulations for achieving top position in the competition.\nThanks for sharing valuable information. ",
      "votes": 1
    },
    {
      "id": 2384727,
      "postDate": "2023-08-11T04:01:48.760Z",
      "content": "<p>Thank you for sharing your excellent write-up.<br>\nAbout image sizes, I couldn't observe performance improvements with 512*512 or larger image sizes. I used albumentations.Resize(interpolation=cv2.INTER_LINEAR). <br>\nDid you do something special to resize the image? How did the score improve with larger images? </p>",
      "rawMarkdown": "Thank you for sharing your excellent write-up.\nAbout image sizes, I couldn't observe performance improvements with 512*512 or larger image sizes. I used albumentations.Resize(interpolation=cv2.INTER_LINEAR). \nDid you do something special to resize the image? How did the score improve with larger images? ",
      "votes": 2,
      "replies": [
        {
          "id": 2390132,
          "postDate": "2023-08-14T11:53:23.080Z",
          "content": "<p>Sorry for my late reply.<br>\nI didn't do anything special to resize image, just used <code>torch.nn.functional.interpolate</code> with <code>mode=\"bilinear\"</code>.</p>\n<p>The score of each image sizes are as follows.</p>\n<ul>\n<li>512x512: 0.6885</li>\n<li>768x768: 0.6906 and 0.6921 (two different seeds)</li>\n<li>1024x1024: 0.6901</li>\n</ul>",
          "rawMarkdown": "Sorry for my late reply.\nI didn't do anything special to resize image, just used `torch.nn.functional.interpolate` with `mode=\"bilinear\"`.\n\nThe score of each image sizes are as follows.\n- 512x512: 0.6885\n- 768x768: 0.6906 and 0.6921 (two different seeds)\n- 1024x1024: 0.6901",
          "votes": 3,
          "replies": [
            {
              "id": 2390250,
              "postDate": "2023-08-14T13:13:20.277Z",
              "content": "<p>I see.. <br>\nAnyway, thank you for the reply.</p>",
              "rawMarkdown": "I see.. \nAnyway, thank you for the reply."
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2383301,
      "author_name": "Man of the year",
      "author_url": "",
      "post_date": "2023-08-10T09:43:50.747000",
      "content": "<p>When you trained on individual annotations, did you just add them to the whole dataset or did you pass as the input to the model several images at once? </p>\n<p>Also, how did you manage flips and rotations augmentations to work?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2383571,
          "author_name": "stakahashi",
          "author_url": "",
          "post_date": "2023-08-10T12:58:34.113000",
          "content": "<p>For the first question, I haven't trained on individual annotations. Just ground truth label is used.</p>\n<p>For the second, I'm not sure why they are worked but I guess</p>\n<ul>\n<li>Pseudo label mitigates label issue because pseudo label may not be misaligned.</li>\n<li>Also large image size may alleviates label issue because the effect of label misalignment is maybe smaller for large masks compared with small one.</li>\n</ul>",
          "votes": 2,
          "replies": [
            {
              "id": 2383592,
              "author_name": "Man of the year",
              "author_url": "",
              "post_date": "2023-08-10T13:08:31.703000",
              "content": "<p>Oh, I have meant that you used pseudo labels for other frames. Did you just add them to the dataset?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2390120,
              "author_name": "stakahashi",
              "author_url": "",
              "post_date": "2023-08-14T11:44:08.657000",
              "content": "<p>Sorry for my late reply.<br>\nYes, pseudo labels are just added to the dataset.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2382828,
      "author_name": "C R Suthikshn Kumar",
      "author_url": "",
      "post_date": "2023-08-10T04:12:33.703000",
      "content": "<p>Congratulations for achieving top position in the competition.<br>\nThanks for sharing valuable information. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2384727,
      "author_name": "luddite^",
      "author_url": "",
      "post_date": "2023-08-11T04:01:48.760000",
      "content": "<p>Thank you for sharing your excellent write-up.<br>\nAbout image sizes, I couldn't observe performance improvements with 512*512 or larger image sizes. I used albumentations.Resize(interpolation=cv2.INTER_LINEAR). <br>\nDid you do something special to resize the image? How did the score improve with larger images? </p>",
      "votes": 2,
      "replies": [
        {
          "id": 2390132,
          "author_name": "stakahashi",
          "author_url": "",
          "post_date": "2023-08-14T11:53:23.080000",
          "content": "<p>Sorry for my late reply.<br>\nI didn't do anything special to resize image, just used <code>torch.nn.functional.interpolate</code> with <code>mode=\"bilinear\"</code>.</p>\n<p>The score of each image sizes are as follows.</p>\n<ul>\n<li>512x512: 0.6885</li>\n<li>768x768: 0.6906 and 0.6921 (two different seeds)</li>\n<li>1024x1024: 0.6901</li>\n</ul>",
          "votes": 3,
          "replies": [
            {
              "id": 2390250,
              "author_name": "luddite^",
              "author_url": "",
              "post_date": "2023-08-14T13:13:20.277000",
              "content": "<p>I see.. <br>\nAnyway, thank you for the reply.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2382754": "Thank you for organizers hosting the competition. I also appreciate my team mates cooperation.\n\n# Summary\n\nOur solution has nothing special, it's just ensemble of 4 EfficientNetV2 with different image sizes.\n\n- 2x l with 768x768 image size (different seeds)\n- 1x l with 1024x1024 image size\n- 1x s with 1536x1536 image size\n\nFirst 3 models give similar score around 0.690 ~ 0.692 for `validation` folder images and the last gives 0.685. Ensemble score is 0.6936.\n\n# What worked\n\n- Augmentations\n    - HV Flip\n    - Rotate\n    - Random scale then crop\n- TTA\n    - HV Flip\n    - Rotate\n- set `\"drop_path_rate\": 0.2` and `\"drop_rate\": 0.2`\n- Adding time steps 3 and 5 as pseudo labels\n- Focal loss\n\nBut I guess improvements from above things are pretty small compared with addressing label misalignment and using individual human annotated mask as new labels as mentioned other discussions (I'm plan to do experiment later).\n\n# What not worked\n\n- Other augmentations\n    - ColorJitter\n    - HueSaturationValue\n    - RandomBrightnessContrast\n    - RandomGamma\n    - RandomFog\n    - RandomShadow\n    - CoarseDropout\n    - MixUp\n- Adding more time steps as pseudo labels\n- Adding pseudo label rounds\n- I'm really focused on training 2.5D and 3D models, but nothing worked at all.\n    - Just 3D model (e.g. Resnet3D)\n    - 3D encoder + 2.5D decoder\n    - Input other time frames predictions as additional channels\n    - Stack 2 Unets, first predicts segmentation masks for multiple time frames, second takes those masks as input by stacking channel dimensions",
    "2383301": "When you trained on individual annotations, did you just add them to the whole dataset or did you pass as the input to the model several images at once? \n\nAlso, how did you manage flips and rotations augmentations to work?",
    "2382828": "Congratulations for achieving top position in the competition.\nThanks for sharing valuable information. ",
    "2384727": "Thank you for sharing your excellent write-up.\nAbout image sizes, I couldn't observe performance improvements with 512*512 or larger image sizes. I used albumentations.Resize(interpolation=cv2.INTER_LINEAR). \nDid you do something special to resize the image? How did the score improve with larger images? "
  }
}