{
  "id": 413068,
  "title": "Main Findings after the first two Competition Weeks",
  "url": "/competitions/google-research-identify-contrails-reduce-global-warming/discussion/413068",
  "author_name": "Jan H",
  "post_date": "2023-05-26T16:30:28.547000",
  "votes": 55,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Hey everyone,</p>\n<p>after 2 weeks of the competition, I want to share my main findings with you and hope that some of you also share some of their experiences.</p>\n<p>I first started by implementing a simple U-Net model on the ASH-RGB images. After training this model with several parameter tunes (learning rate, batch size, augmentations) my best score was around 0.41, which did not seem too promising.. Then I decided to read the preprint and found out that they used the DeepLabV3+ model, which I then decided to switch to. I found a nice implementation of DeepLabV3+ for tensorflow (<a href=\"https://keras.io/examples/vision/deeplabv3_plus/\" target=\"_blank\">https://keras.io/examples/vision/deeplabv3_plus/</a>) and used this code to build the model. This model is also currently the base for my best model (Score 0.61) in the leaderboars, which I trained on the ASH-RGB images. I also tried to train a model on the 6 channels proposed in the preprint, which resulted in poor performance (only around 0.4 in the LB score). I tried a lot of different parameter settings and experimented with resizing the images and different augmentations. For my model, it worked best to use a 512x512 input image size with only a crop augmentation (zooming in and using this image for training). Considering the backbone, I use a resnet50, since using the resnet101 as proposed in the preprint did not result in a higher score but increased training time. I also tried horizontally and vertically flipping the images, which didnt result in a performance boost. Currently, I am considering using other augmentations, such as changing contrast, hue etc., but I am unsure whether these will bring an increase in performance. Since the training and val dataset have a lot of images which do not contain any contrails, I tried to train the model with less of the images that do not contain any contrail, which increased the dice score for my \"private\" test dataset but resulted in less performance on the leaderboard one. I also tried different loss functions, such as WCE, dice loss and jacard loss. The WCE brought the worst results, whereas dice and jacard loss performed similarly. I also tried different train/test/val splits (60/20/20, 80/10/10) and the differences between these two were not that big. Further, I implemented a linear learning rate warmup, as proposed in the preprint, which did not seem to help the training, but instead increased the training time coupled with a higher final validation error…</p>\n<p>So summing up, my personal experience is that a more basic model with less changes and modifications (such as callbacks, augmentations etc.) seems to perform better than a highly modified one.</p>\n<p>Currently, I am unsure of how to go on. I tried to use resizing the images to 784x784, but even with a batch size of 4 I got an out of memory error during training. Right now, I am not sure whether I should try to further improve the DeepLabV3+ model (experiment with other backbones, further modify the hyperparameters etc.) or whether it might be more promising to try other semantic segmentation models. As far as I know, DeepLabV3+ is one of the SOTA-models for semantic segmentation, such that I dont know if changing the model might help…</p>\n<p>This are my experiences so far. I hope that it might help some of the other participants to make a decision and maybe safe some time for the own model design. If you have any suggestions or want to share your own experiences, feel free! </p>\n<p>Good luck to everyone for finding a nice solution!</p>",
  "messages": [
    {
      "id": 2276272,
      "postDate": "2023-05-26T16:30:28.547Z",
      "content": "<p>Hey everyone,</p>\n<p>after 2 weeks of the competition, I want to share my main findings with you and hope that some of you also share some of their experiences.</p>\n<p>I first started by implementing a simple U-Net model on the ASH-RGB images. After training this model with several parameter tunes (learning rate, batch size, augmentations) my best score was around 0.41, which did not seem too promising.. Then I decided to read the preprint and found out that they used the DeepLabV3+ model, which I then decided to switch to. I found a nice implementation of DeepLabV3+ for tensorflow (<a href=\"https://keras.io/examples/vision/deeplabv3_plus/\" target=\"_blank\">https://keras.io/examples/vision/deeplabv3_plus/</a>) and used this code to build the model. This model is also currently the base for my best model (Score 0.61) in the leaderboars, which I trained on the ASH-RGB images. I also tried to train a model on the 6 channels proposed in the preprint, which resulted in poor performance (only around 0.4 in the LB score). I tried a lot of different parameter settings and experimented with resizing the images and different augmentations. For my model, it worked best to use a 512x512 input image size with only a crop augmentation (zooming in and using this image for training). Considering the backbone, I use a resnet50, since using the resnet101 as proposed in the preprint did not result in a higher score but increased training time. I also tried horizontally and vertically flipping the images, which didnt result in a performance boost. Currently, I am considering using other augmentations, such as changing contrast, hue etc., but I am unsure whether these will bring an increase in performance. Since the training and val dataset have a lot of images which do not contain any contrails, I tried to train the model with less of the images that do not contain any contrail, which increased the dice score for my \"private\" test dataset but resulted in less performance on the leaderboard one. I also tried different loss functions, such as WCE, dice loss and jacard loss. The WCE brought the worst results, whereas dice and jacard loss performed similarly. I also tried different train/test/val splits (60/20/20, 80/10/10) and the differences between these two were not that big. Further, I implemented a linear learning rate warmup, as proposed in the preprint, which did not seem to help the training, but instead increased the training time coupled with a higher final validation error…</p>\n<p>So summing up, my personal experience is that a more basic model with less changes and modifications (such as callbacks, augmentations etc.) seems to perform better than a highly modified one.</p>\n<p>Currently, I am unsure of how to go on. I tried to use resizing the images to 784x784, but even with a batch size of 4 I got an out of memory error during training. Right now, I am not sure whether I should try to further improve the DeepLabV3+ model (experiment with other backbones, further modify the hyperparameters etc.) or whether it might be more promising to try other semantic segmentation models. As far as I know, DeepLabV3+ is one of the SOTA-models for semantic segmentation, such that I dont know if changing the model might help…</p>\n<p>This are my experiences so far. I hope that it might help some of the other participants to make a decision and maybe safe some time for the own model design. If you have any suggestions or want to share your own experiences, feel free! </p>\n<p>Good luck to everyone for finding a nice solution!</p>",
      "rawMarkdown": "Hey everyone,\n\nafter 2 weeks of the competition, I want to share my main findings with you and hope that some of you also share some of their experiences.\n\nI first started by implementing a simple U-Net model on the ASH-RGB images. After training this model with several parameter tunes (learning rate, batch size, augmentations) my best score was around 0.41, which did not seem too promising.. Then I decided to read the preprint and found out that they used the DeepLabV3+ model, which I then decided to switch to. I found a nice implementation of DeepLabV3+ for tensorflow (https://keras.io/examples/vision/deeplabv3_plus/) and used this code to build the model. This model is also currently the base for my best model (Score 0.61) in the leaderboars, which I trained on the ASH-RGB images. I also tried to train a model on the 6 channels proposed in the preprint, which resulted in poor performance (only around 0.4 in the LB score). I tried a lot of different parameter settings and experimented with resizing the images and different augmentations. For my model, it worked best to use a 512x512 input image size with only a crop augmentation (zooming in and using this image for training). Considering the backbone, I use a resnet50, since using the resnet101 as proposed in the preprint did not result in a higher score but increased training time. I also tried horizontally and vertically flipping the images, which didnt result in a performance boost. Currently, I am considering using other augmentations, such as changing contrast, hue etc., but I am unsure whether these will bring an increase in performance. Since the training and val dataset have a lot of images which do not contain any contrails, I tried to train the model with less of the images that do not contain any contrail, which increased the dice score for my \"private\" test dataset but resulted in less performance on the leaderboard one. I also tried different loss functions, such as WCE, dice loss and jacard loss. The WCE brought the worst results, whereas dice and jacard loss performed similarly. I also tried different train/test/val splits (60/20/20, 80/10/10) and the differences between these two were not that big. Further, I implemented a linear learning rate warmup, as proposed in the preprint, which did not seem to help the training, but instead increased the training time coupled with a higher final validation error...\n\nSo summing up, my personal experience is that a more basic model with less changes and modifications (such as callbacks, augmentations etc.) seems to perform better than a highly modified one.\n\nCurrently, I am unsure of how to go on. I tried to use resizing the images to 784x784, but even with a batch size of 4 I got an out of memory error during training. Right now, I am not sure whether I should try to further improve the DeepLabV3+ model (experiment with other backbones, further modify the hyperparameters etc.) or whether it might be more promising to try other semantic segmentation models. As far as I know, DeepLabV3+ is one of the SOTA-models for semantic segmentation, such that I dont know if changing the model might help...\n\nThis are my experiences so far. I hope that it might help some of the other participants to make a decision and maybe safe some time for the own model design. If you have any suggestions or want to share your own experiences, feel free! \n\nGood luck to everyone for finding a nice solution!",
      "votes": 53
    },
    {
      "id": 2276374,
      "postDate": "2023-05-26T18:37:05.527Z",
      "content": "<p>I'm guessing there will be significant shakeups because,</p>\n<blockquote>\n  <p>The hidden test set is approximately the same size (± 5%) as the validation set</p>\n</blockquote>\n<p>The validation set consists of 1856 samples, this means that the hidden test set is approximately 1763-1949 samples.<br>\nAs the public LB is calculated using just 15% of the test data, the public LB is based on just 264-292 samples<br>\nConclusion: The public LB is not very informative so we should put more emphasis on the CV score</p>\n<p>Good luck!</p>",
      "rawMarkdown": "I'm guessing there will be significant shakeups because,\n>The hidden test set is approximately the same size (± 5%) as the validation set\n\nThe validation set consists of 1856 samples, this means that the hidden test set is approximately 1763-1949 samples.\nAs the public LB is calculated using just 15% of the test data, the public LB is based on just 264-292 samples\nConclusion: The public LB is not very informative so we should put more emphasis on the CV score\n\nGood luck!",
      "votes": 4,
      "replies": [
        {
          "id": 2276414,
          "postDate": "2023-05-26T19:21:34.470Z",
          "content": "<p>So you think minimizing your own validation / test dataset loss is more promising than focussing on the public LB score? I did not yet think about the size of the dataset used for scoring, but your comment makes sense!</p>",
          "rawMarkdown": "So you think minimizing your own validation / test dataset loss is more promising than focussing on the public LB score? I did not yet think about the size of the dataset used for scoring, but your comment makes sense!"
        }
      ]
    },
    {
      "id": 2276413,
      "postDate": "2023-05-26T19:19:31.670Z",
      "content": "<p>How did you train the model with 6 channels ? Did you simply take 0th timestamp and stacked 6 color bands between 8th to 15th bands ? Or Did you take 3 bands (ash-rgb ) and stacked together 2 timestamps to make it 6 channel ?</p>",
      "rawMarkdown": "How did you train the model with 6 channels ? Did you simply take 0th timestamp and stacked 6 color bands between 8th to 15th bands ? Or Did you take 3 bands (ash-rgb ) and stacked together 2 timestamps to make it 6 channel ?",
      "votes": 1,
      "replies": [
        {
          "id": 2276416,
          "postDate": "2023-05-26T19:22:42.750Z",
          "content": "<p>I took the 6 channels proposed in the preprint paper (mostly the same as the ones used for Ash Rgb). One downside is that you cannot use a pretrained model, since these are most often trained on 3 channel RGB images.</p>\n<p>May i ask what kind of model you are using? Also a 3 channel ASH RGB DeepLab model?</p>",
          "rawMarkdown": "I took the 6 channels proposed in the preprint paper (mostly the same as the ones used for Ash Rgb). One downside is that you cannot use a pretrained model, since these are most often trained on 3 channel RGB images.\n\nMay i ask what kind of model you are using? Also a 3 channel ASH RGB DeepLab model?",
          "replies": [
            {
              "id": 2276564,
              "postDate": "2023-05-27T02:06:08.710Z",
              "content": "<p>If the other layers are the same, you should still be able to use a pretrained model. For the input convolutional layer, you can then randomly initialize, or replicate the weights from the pretrained model.</p>",
              "rawMarkdown": "If the other layers are the same, you should still be able to use a pretrained model. For the input convolutional layer, you can then randomly initialize, or replicate the weights from the pretrained model."
            },
            {
              "id": 2276569,
              "postDate": "2023-05-27T02:28:39.130Z",
              "content": "<p>The problem is about the number of channels. Since the pretrained model used 3 channels, it does not work for me</p>",
              "rawMarkdown": "The problem is about the number of channels. Since the pretrained model used 3 channels, it does not work for me"
            }
          ]
        }
      ]
    },
    {
      "id": 2282663,
      "postDate": "2023-05-31T18:13:06.273Z",
      "content": "<p>also I tried different upscaling ratios with my baseline model resnet34-unet - and x2 upscaling gives better results (lb 517-&gt; 598), x3 and x4 gives the same results but training time increases</p>",
      "rawMarkdown": "also I tried different upscaling ratios with my baseline model resnet34-unet - and x2 upscaling gives better results (lb 517-> 598), x3 and x4 gives the same results but training time increases",
      "votes": 2,
      "replies": [
        {
          "id": 2282671,
          "postDate": "2023-05-31T18:18:21.740Z",
          "content": "<p>With upscaling you mean the size of the images? So that upscaling x2 = 512x512 px?</p>",
          "rawMarkdown": "With upscaling you mean the size of the images? So that upscaling x2 = 512x512 px?"
        }
      ]
    },
    {
      "id": 2291794,
      "postDate": "2023-06-07T20:20:59.490Z",
      "content": "<p>Have you tried adding temporal context like in the preprint? It seemed to give them a decent boost. There also may be better ways of adding temporal context than in the preprint and dabbling with 2.5D techniques might be interesting.</p>",
      "rawMarkdown": "Have you tried adding temporal context like in the preprint? It seemed to give them a decent boost. There also may be better ways of adding temporal context than in the preprint and dabbling with 2.5D techniques might be interesting."
    },
    {
      "id": 2288442,
      "postDate": "2023-06-05T11:42:20.077Z",
      "content": "<p>How big is your model?<br>\nI got 0.503 on the leaderboard after training a U-Net model with 1.11M parameters (occupying 13.4 MB) using dice loss on 256x256 ASH-RGB images.<br>\nI only used samples from the train folder: 16k for training and 4k for validation. </p>",
      "rawMarkdown": "How big is your model?\nI got 0.503 on the leaderboard after training a U-Net model with 1.11M parameters (occupying 13.4 MB) using dice loss on 256x256 ASH-RGB images.\nI only used samples from the train folder: 16k for training and 4k for validation. "
    },
    {
      "id": 2276537,
      "postDate": "2023-05-27T00:28:07.417Z",
      "content": "<p>Thank you for sharing! <br>\nMay I ask some details about training with dice loss? For me, it's pretty hard to train.</p>",
      "rawMarkdown": "Thank you for sharing! \nMay I ask some details about training with dice loss? For me, it's pretty hard to train."
    },
    {
      "id": 2276410,
      "postDate": "2023-05-26T19:15:29.783Z",
      "content": "<p>Thanks ..great insights ! </p>",
      "rawMarkdown": "Thanks ..great insights ! ",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 2276374,
      "author_name": "MD Mushfirat Mohaimin",
      "author_url": "",
      "post_date": "2023-05-26T18:37:05.527000",
      "content": "<p>I'm guessing there will be significant shakeups because,</p>\n<blockquote>\n  <p>The hidden test set is approximately the same size (± 5%) as the validation set</p>\n</blockquote>\n<p>The validation set consists of 1856 samples, this means that the hidden test set is approximately 1763-1949 samples.<br>\nAs the public LB is calculated using just 15% of the test data, the public LB is based on just 264-292 samples<br>\nConclusion: The public LB is not very informative so we should put more emphasis on the CV score</p>\n<p>Good luck!</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2276414,
          "author_name": "Jan H",
          "author_url": "",
          "post_date": "2023-05-26T19:21:34.470000",
          "content": "<p>So you think minimizing your own validation / test dataset loss is more promising than focussing on the public LB score? I did not yet think about the size of the dataset used for scoring, but your comment makes sense!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2276413,
      "author_name": "Nirjhar Roy",
      "author_url": "",
      "post_date": "2023-05-26T19:19:31.670000",
      "content": "<p>How did you train the model with 6 channels ? Did you simply take 0th timestamp and stacked 6 color bands between 8th to 15th bands ? Or Did you take 3 bands (ash-rgb ) and stacked together 2 timestamps to make it 6 channel ?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2276416,
          "author_name": "Jan H",
          "author_url": "",
          "post_date": "2023-05-26T19:22:42.750000",
          "content": "<p>I took the 6 channels proposed in the preprint paper (mostly the same as the ones used for Ash Rgb). One downside is that you cannot use a pretrained model, since these are most often trained on 3 channel RGB images.</p>\n<p>May i ask what kind of model you are using? Also a 3 channel ASH RGB DeepLab model?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2276564,
              "author_name": "Jonathan",
              "author_url": "",
              "post_date": "2023-05-27T02:06:08.710000",
              "content": "<p>If the other layers are the same, you should still be able to use a pretrained model. For the input convolutional layer, you can then randomly initialize, or replicate the weights from the pretrained model.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2276569,
              "author_name": "Jan H",
              "author_url": "",
              "post_date": "2023-05-27T02:28:39.130000",
              "content": "<p>The problem is about the number of channels. Since the pretrained model used 3 channels, it does not work for me</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2282663,
      "author_name": "Kostiantyn Maksymov",
      "author_url": "",
      "post_date": "2023-05-31T18:13:06.273000",
      "content": "<p>also I tried different upscaling ratios with my baseline model resnet34-unet - and x2 upscaling gives better results (lb 517-&gt; 598), x3 and x4 gives the same results but training time increases</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2282671,
          "author_name": "Jan H",
          "author_url": "",
          "post_date": "2023-05-31T18:18:21.740000",
          "content": "<p>With upscaling you mean the size of the images? So that upscaling x2 = 512x512 px?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2291794,
      "author_name": "Ari",
      "author_url": "",
      "post_date": "2023-06-07T20:20:59.490000",
      "content": "<p>Have you tried adding temporal context like in the preprint? It seemed to give them a decent boost. There also may be better ways of adding temporal context than in the preprint and dabbling with 2.5D techniques might be interesting.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2288442,
      "author_name": "MD Mushfirat Mohaimin",
      "author_url": "",
      "post_date": "2023-06-05T11:42:20.077000",
      "content": "<p>How big is your model?<br>\nI got 0.503 on the leaderboard after training a U-Net model with 1.11M parameters (occupying 13.4 MB) using dice loss on 256x256 ASH-RGB images.<br>\nI only used samples from the train folder: 16k for training and 4k for validation. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2276537,
      "author_name": "LUPIN11",
      "author_url": "",
      "post_date": "2023-05-27T00:28:07.417000",
      "content": "<p>Thank you for sharing! <br>\nMay I ask some details about training with dice loss? For me, it's pretty hard to train.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2276410,
      "author_name": "Nirjhar Roy",
      "author_url": "",
      "post_date": "2023-05-26T19:15:29.783000",
      "content": "<p>Thanks ..great insights ! </p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2276272": "Hey everyone,\n\nafter 2 weeks of the competition, I want to share my main findings with you and hope that some of you also share some of their experiences.\n\nI first started by implementing a simple U-Net model on the ASH-RGB images. After training this model with several parameter tunes (learning rate, batch size, augmentations) my best score was around 0.41, which did not seem too promising.. Then I decided to read the preprint and found out that they used the DeepLabV3+ model, which I then decided to switch to. I found a nice implementation of DeepLabV3+ for tensorflow (https://keras.io/examples/vision/deeplabv3_plus/) and used this code to build the model. This model is also currently the base for my best model (Score 0.61) in the leaderboars, which I trained on the ASH-RGB images. I also tried to train a model on the 6 channels proposed in the preprint, which resulted in poor performance (only around 0.4 in the LB score). I tried a lot of different parameter settings and experimented with resizing the images and different augmentations. For my model, it worked best to use a 512x512 input image size with only a crop augmentation (zooming in and using this image for training). Considering the backbone, I use a resnet50, since using the resnet101 as proposed in the preprint did not result in a higher score but increased training time. I also tried horizontally and vertically flipping the images, which didnt result in a performance boost. Currently, I am considering using other augmentations, such as changing contrast, hue etc., but I am unsure whether these will bring an increase in performance. Since the training and val dataset have a lot of images which do not contain any contrails, I tried to train the model with less of the images that do not contain any contrail, which increased the dice score for my \"private\" test dataset but resulted in less performance on the leaderboard one. I also tried different loss functions, such as WCE, dice loss and jacard loss. The WCE brought the worst results, whereas dice and jacard loss performed similarly. I also tried different train/test/val splits (60/20/20, 80/10/10) and the differences between these two were not that big. Further, I implemented a linear learning rate warmup, as proposed in the preprint, which did not seem to help the training, but instead increased the training time coupled with a higher final validation error...\n\nSo summing up, my personal experience is that a more basic model with less changes and modifications (such as callbacks, augmentations etc.) seems to perform better than a highly modified one.\n\nCurrently, I am unsure of how to go on. I tried to use resizing the images to 784x784, but even with a batch size of 4 I got an out of memory error during training. Right now, I am not sure whether I should try to further improve the DeepLabV3+ model (experiment with other backbones, further modify the hyperparameters etc.) or whether it might be more promising to try other semantic segmentation models. As far as I know, DeepLabV3+ is one of the SOTA-models for semantic segmentation, such that I dont know if changing the model might help...\n\nThis are my experiences so far. I hope that it might help some of the other participants to make a decision and maybe safe some time for the own model design. If you have any suggestions or want to share your own experiences, feel free! \n\nGood luck to everyone for finding a nice solution!",
    "2276374": "I'm guessing there will be significant shakeups because,\n>The hidden test set is approximately the same size (± 5%) as the validation set\n\nThe validation set consists of 1856 samples, this means that the hidden test set is approximately 1763-1949 samples.\nAs the public LB is calculated using just 15% of the test data, the public LB is based on just 264-292 samples\nConclusion: The public LB is not very informative so we should put more emphasis on the CV score\n\nGood luck!",
    "2276413": "How did you train the model with 6 channels ? Did you simply take 0th timestamp and stacked 6 color bands between 8th to 15th bands ? Or Did you take 3 bands (ash-rgb ) and stacked together 2 timestamps to make it 6 channel ?",
    "2282663": "also I tried different upscaling ratios with my baseline model resnet34-unet - and x2 upscaling gives better results (lb 517-> 598), x3 and x4 gives the same results but training time increases",
    "2291794": "Have you tried adding temporal context like in the preprint? It seemed to give them a decent boost. There also may be better ways of adding temporal context than in the preprint and dabbling with 2.5D techniques might be interesting.",
    "2288442": "How big is your model?\nI got 0.503 on the leaderboard after training a U-Net model with 1.11M parameters (occupying 13.4 MB) using dice loss on 256x256 ASH-RGB images.\nI only used samples from the train folder: 16k for training and 4k for validation. ",
    "2276537": "Thank you for sharing! \nMay I ask some details about training with dice loss? For me, it's pretty hard to train.",
    "2276410": "Thanks ..great insights ! "
  }
}