{
  "id": 432502,
  "title": "92nd Place Solution for the Google Research - Identify Contrails to Reduce Global Warming Competition",
  "url": "/competitions/google-research-identify-contrails-reduce-global-warming/discussion/432502",
  "author_name": "Man of the year",
  "post_date": "2023-08-17T17:53:22.825000",
  "votes": 19,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Welcome! Сongratulations to everyone on finishing this great competition. It was a honor for us to compete with you all. Makes me happy to see, that AI solutions can be beneficial for the environment and nature. </p>\n<p>Our team ( <a href=\"https://www.kaggle.com/slavabarkov\" target=\"_blank\">@slavabarkov</a> and me) would like to present our approach. We didn't get the SOTA result, but I believe, that we got some insights. </p>\n<h1>Context section</h1>\n<ul>\n<li>Business context: <a href=\"http://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming\" target=\"_blank\">www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming</a></li>\n<li>Data context: <a href=\"https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/data\" target=\"_blank\">https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/data</a></li>\n</ul>\n<h1>Overview of the Approach</h1>\n<p>Our best solution was an ensemble of several models with equal weights: efficientnet-b5, b6, b7, b8 and eca_nfnet_l2 with the Unet decoder. All models were trained on 512x512 images (Ash Color images, shout-out to <a href=\"https://www.kaggle.com/shashwatraman\" target=\"_blank\">@shashwatraman</a>) and DICE loss. We used threshold=0.3 for the pixels final prediction. It got 0.685 on the public lb and 0.681 on the private. </p>\n<p>Our validation score correlated rather nice with the public lb. We used the StratifiedKFold validation (4 folds), stratified by the masks size. I'll describe our approach in the next paragraph more precisely.</p>\n<h1>Details of the submission</h1>\n<h3>What was special about the submission</h3>\n<p>We had a straightforward robust approach, that has a nice performance, without using the SOTA models. We developed a thoughtful validation strategy, that can be used in the future competitions. </p>\n<p>We also had ideas that we did not manage to work right (we joined the competition last week). However, according to the winning solutions, we moved in the right direction.</p>\n<h3>Validation strategy</h3>\n<p>I didn't like the idea of random folds, since folds might contain the images of different complexity level. Some folds might contain too many empty masks or too big masks. I thought, that models might struggle with prediction of empty masks, for example. Therefore, it was a nice approach to divide images equally on folds, based on their masks size. One needs just to sum the total amount of masks pixels. Then you should create bins to label ranges of masks sizes. For empty masks we created a separate bin. Then you just stratify data, based on these labels. </p>\n<p>Code example:</p>\n<pre><code>\nskf = StratifiedKFold(\n    n_splits=config[][],\n    shuffle=,\n    random_state=config[][],\n)\n\n\ndf[] = pd.qcut(\n    df[df[] &gt; ][],\n    q=config[][] - ,\n    labels=,\n)\n\ndf.loc[df[] == , ] = -\n\n\n\n fold_number, (train_index, val_index)  (\n    skf.split(df, df[])\n):\n    df.loc[val_index, ] = (fold_number)\n\n\ndf.drop(columns=[], inplace=)\n</code></pre>\n<h3>Other details about the models</h3>\n<p>For segmentation models we used the smp torch library. We have tried different backbones from the timm library. However, we found out that an efficientnet family was the best, resnet was not successful in our case. We also liked the nfnet models, they were faster and also could be trained on bigger image sizes. However, 512x512 worked the best. We also found out, that models hyperparameters was not that important: lr, scheduler - does not mean too much. </p>\n<p>Unet decoder also was the best one, others were worse or showed almost the same performance. </p>\n<h3>Ideas that didn't work in our case (but they can)</h3>\n<h3>Pretraining on the synthetic dataset</h3>\n<p>That was a really cool idea to generate the similar background images by generative networks and draw masks on them to pretrain the models. I think that it is a very perspective idea in future works. In our case it didn't work good, trained models showed nice performance on these images, but this didn't improve the quality on the competition dataset. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2590337%2Fc04e6ac0c033d17bac52b6405ea828ee%2Fphoto_2023-08-05_00-49-53.jpg?generation=1692294981919634&amp;alt=media\" alt=\"\"></p>\n<h3>Training on the soft labels</h3>\n<p>Instead of using the ground truth masks, we can average the predictions by annotators. We preprocessed the data and then just didn't have the time to test it. But this can work nice, since there are controversial images in the dataset.</p>\n<h3>Pseudo labeling other frames</h3>\n<p>That's a great approach to make the data bigger. We did this for the 2,3,5,6 frames and trained models with the DICE loss. It was a total mistake. We trained on binary masks, but we should use BCE loss and soft labels in this case. Another mistake was that we used the whole ensemble to label this data, so we could not really validate the models on previous folds. Our validation score improved, but the lb did not. I believe it is due to the hard labels and knowledge (data) leakage from the models trained on other folds. </p>\n<h3>Training the additional classifier</h3>\n<p>I wanted to train a classifier model just to predict, whether the image has the empty mask or not. I was really concerned about the empty masks, since the wrong predictions on these images will get the high error. I tried different backbones, got about 0.9 accuracy and in the pipeline it didn't work well then. Didn't have time to improve this step.</p>\n<h3>Training code</h3>\n<p>We have created a python project with loggers (wandb), which runs from the command line with YAML config. You can find it on the github: <a href=\"https://github.com/25icecreamflavors/contrails\" target=\"_blank\">https://github.com/25icecreamflavors/contrails</a></p>\n<h1>Sources</h1>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/shashwatraman/contrails-dataset-ash-color\" target=\"_blank\">https://www.kaggle.com/code/shashwatraman/contrails-dataset-ash-color</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/understanding_cloud_organization/discussion/118080\" target=\"_blank\">https://www.kaggle.com/competitions/understanding_cloud_organization/discussion/118080</a></li>\n<li><a href=\"https://github.com/25icecreamflavors/contrails\" target=\"_blank\">https://github.com/25icecreamflavors/contrails</a></li>\n</ul>",
  "messages": [
    {
      "id": 2395659,
      "postDate": "2023-08-17T17:53:22.827Z",
      "content": "<p>Welcome! Сongratulations to everyone on finishing this great competition. It was a honor for us to compete with you all. Makes me happy to see, that AI solutions can be beneficial for the environment and nature. </p>\n<p>Our team ( <a href=\"https://www.kaggle.com/slavabarkov\" target=\"_blank\">@slavabarkov</a> and me) would like to present our approach. We didn't get the SOTA result, but I believe, that we got some insights. </p>\n<h1>Context section</h1>\n<ul>\n<li>Business context: <a href=\"http://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming\" target=\"_blank\">www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming</a></li>\n<li>Data context: <a href=\"https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/data\" target=\"_blank\">https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/data</a></li>\n</ul>\n<h1>Overview of the Approach</h1>\n<p>Our best solution was an ensemble of several models with equal weights: efficientnet-b5, b6, b7, b8 and eca_nfnet_l2 with the Unet decoder. All models were trained on 512x512 images (Ash Color images, shout-out to <a href=\"https://www.kaggle.com/shashwatraman\" target=\"_blank\">@shashwatraman</a>) and DICE loss. We used threshold=0.3 for the pixels final prediction. It got 0.685 on the public lb and 0.681 on the private. </p>\n<p>Our validation score correlated rather nice with the public lb. We used the StratifiedKFold validation (4 folds), stratified by the masks size. I'll describe our approach in the next paragraph more precisely.</p>\n<h1>Details of the submission</h1>\n<h3>What was special about the submission</h3>\n<p>We had a straightforward robust approach, that has a nice performance, without using the SOTA models. We developed a thoughtful validation strategy, that can be used in the future competitions. </p>\n<p>We also had ideas that we did not manage to work right (we joined the competition last week). However, according to the winning solutions, we moved in the right direction.</p>\n<h3>Validation strategy</h3>\n<p>I didn't like the idea of random folds, since folds might contain the images of different complexity level. Some folds might contain too many empty masks or too big masks. I thought, that models might struggle with prediction of empty masks, for example. Therefore, it was a nice approach to divide images equally on folds, based on their masks size. One needs just to sum the total amount of masks pixels. Then you should create bins to label ranges of masks sizes. For empty masks we created a separate bin. Then you just stratify data, based on these labels. </p>\n<p>Code example:</p>\n<pre><code>\nskf = StratifiedKFold(\n    n_splits=config[][],\n    shuffle=,\n    random_state=config[][],\n)\n\n\ndf[] = pd.qcut(\n    df[df[] &gt; ][],\n    q=config[][] - ,\n    labels=,\n)\n\ndf.loc[df[] == , ] = -\n\n\n\n fold_number, (train_index, val_index)  (\n    skf.split(df, df[])\n):\n    df.loc[val_index, ] = (fold_number)\n\n\ndf.drop(columns=[], inplace=)\n</code></pre>\n<h3>Other details about the models</h3>\n<p>For segmentation models we used the smp torch library. We have tried different backbones from the timm library. However, we found out that an efficientnet family was the best, resnet was not successful in our case. We also liked the nfnet models, they were faster and also could be trained on bigger image sizes. However, 512x512 worked the best. We also found out, that models hyperparameters was not that important: lr, scheduler - does not mean too much. </p>\n<p>Unet decoder also was the best one, others were worse or showed almost the same performance. </p>\n<h3>Ideas that didn't work in our case (but they can)</h3>\n<h3>Pretraining on the synthetic dataset</h3>\n<p>That was a really cool idea to generate the similar background images by generative networks and draw masks on them to pretrain the models. I think that it is a very perspective idea in future works. In our case it didn't work good, trained models showed nice performance on these images, but this didn't improve the quality on the competition dataset. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2590337%2Fc04e6ac0c033d17bac52b6405ea828ee%2Fphoto_2023-08-05_00-49-53.jpg?generation=1692294981919634&amp;alt=media\" alt=\"\"></p>\n<h3>Training on the soft labels</h3>\n<p>Instead of using the ground truth masks, we can average the predictions by annotators. We preprocessed the data and then just didn't have the time to test it. But this can work nice, since there are controversial images in the dataset.</p>\n<h3>Pseudo labeling other frames</h3>\n<p>That's a great approach to make the data bigger. We did this for the 2,3,5,6 frames and trained models with the DICE loss. It was a total mistake. We trained on binary masks, but we should use BCE loss and soft labels in this case. Another mistake was that we used the whole ensemble to label this data, so we could not really validate the models on previous folds. Our validation score improved, but the lb did not. I believe it is due to the hard labels and knowledge (data) leakage from the models trained on other folds. </p>\n<h3>Training the additional classifier</h3>\n<p>I wanted to train a classifier model just to predict, whether the image has the empty mask or not. I was really concerned about the empty masks, since the wrong predictions on these images will get the high error. I tried different backbones, got about 0.9 accuracy and in the pipeline it didn't work well then. Didn't have time to improve this step.</p>\n<h3>Training code</h3>\n<p>We have created a python project with loggers (wandb), which runs from the command line with YAML config. You can find it on the github: <a href=\"https://github.com/25icecreamflavors/contrails\" target=\"_blank\">https://github.com/25icecreamflavors/contrails</a></p>\n<h1>Sources</h1>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/shashwatraman/contrails-dataset-ash-color\" target=\"_blank\">https://www.kaggle.com/code/shashwatraman/contrails-dataset-ash-color</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/understanding_cloud_organization/discussion/118080\" target=\"_blank\">https://www.kaggle.com/competitions/understanding_cloud_organization/discussion/118080</a></li>\n<li><a href=\"https://github.com/25icecreamflavors/contrails\" target=\"_blank\">https://github.com/25icecreamflavors/contrails</a></li>\n</ul>",
      "rawMarkdown": "Welcome! Сongratulations to everyone on finishing this great competition. It was a honor for us to compete with you all. Makes me happy to see, that AI solutions can be beneficial for the environment and nature. \n\nOur team ( @slavabarkov and me) would like to present our approach. We didn't get the SOTA result, but I believe, that we got some insights. \n\n# Context section\n\n- Business context: www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming\n- Data context: https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/data\n\n# Overview of the Approach\nOur best solution was an ensemble of several models with equal weights: efficientnet-b5, b6, b7, b8 and eca_nfnet_l2 with the Unet decoder. All models were trained on 512x512 images (Ash Color images, shout-out to @shashwatraman) and DICE loss. We used threshold=0.3 for the pixels final prediction. It got 0.685 on the public lb and 0.681 on the private. \n\nOur validation score correlated rather nice with the public lb. We used the StratifiedKFold validation (4 folds), stratified by the masks size. I'll describe our approach in the next paragraph more precisely.\n\n# Details of the submission\n### What was special about the submission\n\nWe had a straightforward robust approach, that has a nice performance, without using the SOTA models. We developed a thoughtful validation strategy, that can be used in the future competitions. \n\nWe also had ideas that we did not manage to work right (we joined the competition last week). However, according to the winning solutions, we moved in the right direction.\n\n### Validation strategy\n\nI didn't like the idea of random folds, since folds might contain the images of different complexity level. Some folds might contain too many empty masks or too big masks. I thought, that models might struggle with prediction of empty masks, for example. Therefore, it was a nice approach to divide images equally on folds, based on their masks size. One needs just to sum the total amount of masks pixels. Then you should create bins to label ranges of masks sizes. For empty masks we created a separate bin. Then you just stratify data, based on these labels. \n\nCode example:\n```python\n# Initialize StratifiedKFold\nskf = StratifiedKFold(\n    n_splits=config[\"folds\"][\"n_splits\"],\n    shuffle=True,\n    random_state=config[\"folds\"][\"random_state\"],\n)\n# Create a temporary \"mask_size_bin\" column to handle bins\n# for non-zero sizes\ndf[\"mask_size_bin\"] = pd.qcut(\n    df[df[\"mask_size\"] > 0][\"mask_size\"],\n    q=config[\"folds\"][\"bins\"] - 1,\n    labels=False,\n)\n# Assign a unique bin for zero-size masks\ndf.loc[df[\"mask_size\"] == 0, \"mask_size_bin\"] = -1\n\n# Stratify based on the \"mask_size_bin\" column and\n# assign fold indices\nfor fold_number, (train_index, val_index) in enumerate(\n    skf.split(df, df[\"mask_size_bin\"])\n):\n    df.loc[val_index, \"kfold\"] = int(fold_number)\n\n# Drop the temporary \"mask_size_bin\" column\ndf.drop(columns=[\"mask_size_bin\"], inplace=True)\n```\n\n### Other details about the models\n\nFor segmentation models we used the smp torch library. We have tried different backbones from the timm library. However, we found out that an efficientnet family was the best, resnet was not successful in our case. We also liked the nfnet models, they were faster and also could be trained on bigger image sizes. However, 512x512 worked the best. We also found out, that models hyperparameters was not that important: lr, scheduler - does not mean too much. \n\nUnet decoder also was the best one, others were worse or showed almost the same performance. \n\n### Ideas that didn't work in our case (but they can)\n\n### Pretraining on the synthetic dataset\n\nThat was a really cool idea to generate the similar background images by generative networks and draw masks on them to pretrain the models. I think that it is a very perspective idea in future works. In our case it didn't work good, trained models showed nice performance on these images, but this didn't improve the quality on the competition dataset. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2590337%2Fc04e6ac0c033d17bac52b6405ea828ee%2Fphoto_2023-08-05_00-49-53.jpg?generation=1692294981919634&alt=media)\n\n### Training on the soft labels\n\nInstead of using the ground truth masks, we can average the predictions by annotators. We preprocessed the data and then just didn't have the time to test it. But this can work nice, since there are controversial images in the dataset.\n\n### Pseudo labeling other frames\n\nThat's a great approach to make the data bigger. We did this for the 2,3,5,6 frames and trained models with the DICE loss. It was a total mistake. We trained on binary masks, but we should use BCE loss and soft labels in this case. Another mistake was that we used the whole ensemble to label this data, so we could not really validate the models on previous folds. Our validation score improved, but the lb did not. I believe it is due to the hard labels and knowledge (data) leakage from the models trained on other folds. \n\n### Training the additional classifier\n\nI wanted to train a classifier model just to predict, whether the image has the empty mask or not. I was really concerned about the empty masks, since the wrong predictions on these images will get the high error. I tried different backbones, got about 0.9 accuracy and in the pipeline it didn't work well then. Didn't have time to improve this step.\n\n### Training code\n\nWe have created a python project with loggers (wandb), which runs from the command line with YAML config. You can find it on the github: https://github.com/25icecreamflavors/contrails\n# Sources\n- https://www.kaggle.com/code/shashwatraman/contrails-dataset-ash-color\n- https://www.kaggle.com/competitions/understanding_cloud_organization/discussion/118080\n- https://github.com/25icecreamflavors/contrails",
      "votes": 17
    },
    {
      "id": 2395834,
      "postDate": "2023-08-17T21:20:34.390Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2395834,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-08-17T21:20:34.390000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2395659": "Welcome! Сongratulations to everyone on finishing this great competition. It was a honor for us to compete with you all. Makes me happy to see, that AI solutions can be beneficial for the environment and nature. \n\nOur team ( @slavabarkov and me) would like to present our approach. We didn't get the SOTA result, but I believe, that we got some insights. \n\n# Context section\n\n- Business context: www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming\n- Data context: https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/data\n\n# Overview of the Approach\nOur best solution was an ensemble of several models with equal weights: efficientnet-b5, b6, b7, b8 and eca_nfnet_l2 with the Unet decoder. All models were trained on 512x512 images (Ash Color images, shout-out to @shashwatraman) and DICE loss. We used threshold=0.3 for the pixels final prediction. It got 0.685 on the public lb and 0.681 on the private. \n\nOur validation score correlated rather nice with the public lb. We used the StratifiedKFold validation (4 folds), stratified by the masks size. I'll describe our approach in the next paragraph more precisely.\n\n# Details of the submission\n### What was special about the submission\n\nWe had a straightforward robust approach, that has a nice performance, without using the SOTA models. We developed a thoughtful validation strategy, that can be used in the future competitions. \n\nWe also had ideas that we did not manage to work right (we joined the competition last week). However, according to the winning solutions, we moved in the right direction.\n\n### Validation strategy\n\nI didn't like the idea of random folds, since folds might contain the images of different complexity level. Some folds might contain too many empty masks or too big masks. I thought, that models might struggle with prediction of empty masks, for example. Therefore, it was a nice approach to divide images equally on folds, based on their masks size. One needs just to sum the total amount of masks pixels. Then you should create bins to label ranges of masks sizes. For empty masks we created a separate bin. Then you just stratify data, based on these labels. \n\nCode example:\n```python\n# Initialize StratifiedKFold\nskf = StratifiedKFold(\n    n_splits=config[\"folds\"][\"n_splits\"],\n    shuffle=True,\n    random_state=config[\"folds\"][\"random_state\"],\n)\n# Create a temporary \"mask_size_bin\" column to handle bins\n# for non-zero sizes\ndf[\"mask_size_bin\"] = pd.qcut(\n    df[df[\"mask_size\"] > 0][\"mask_size\"],\n    q=config[\"folds\"][\"bins\"] - 1,\n    labels=False,\n)\n# Assign a unique bin for zero-size masks\ndf.loc[df[\"mask_size\"] == 0, \"mask_size_bin\"] = -1\n\n# Stratify based on the \"mask_size_bin\" column and\n# assign fold indices\nfor fold_number, (train_index, val_index) in enumerate(\n    skf.split(df, df[\"mask_size_bin\"])\n):\n    df.loc[val_index, \"kfold\"] = int(fold_number)\n\n# Drop the temporary \"mask_size_bin\" column\ndf.drop(columns=[\"mask_size_bin\"], inplace=True)\n```\n\n### Other details about the models\n\nFor segmentation models we used the smp torch library. We have tried different backbones from the timm library. However, we found out that an efficientnet family was the best, resnet was not successful in our case. We also liked the nfnet models, they were faster and also could be trained on bigger image sizes. However, 512x512 worked the best. We also found out, that models hyperparameters was not that important: lr, scheduler - does not mean too much. \n\nUnet decoder also was the best one, others were worse or showed almost the same performance. \n\n### Ideas that didn't work in our case (but they can)\n\n### Pretraining on the synthetic dataset\n\nThat was a really cool idea to generate the similar background images by generative networks and draw masks on them to pretrain the models. I think that it is a very perspective idea in future works. In our case it didn't work good, trained models showed nice performance on these images, but this didn't improve the quality on the competition dataset. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2590337%2Fc04e6ac0c033d17bac52b6405ea828ee%2Fphoto_2023-08-05_00-49-53.jpg?generation=1692294981919634&alt=media)\n\n### Training on the soft labels\n\nInstead of using the ground truth masks, we can average the predictions by annotators. We preprocessed the data and then just didn't have the time to test it. But this can work nice, since there are controversial images in the dataset.\n\n### Pseudo labeling other frames\n\nThat's a great approach to make the data bigger. We did this for the 2,3,5,6 frames and trained models with the DICE loss. It was a total mistake. We trained on binary masks, but we should use BCE loss and soft labels in this case. Another mistake was that we used the whole ensemble to label this data, so we could not really validate the models on previous folds. Our validation score improved, but the lb did not. I believe it is due to the hard labels and knowledge (data) leakage from the models trained on other folds. \n\n### Training the additional classifier\n\nI wanted to train a classifier model just to predict, whether the image has the empty mask or not. I was really concerned about the empty masks, since the wrong predictions on these images will get the high error. I tried different backbones, got about 0.9 accuracy and in the pipeline it didn't work well then. Didn't have time to improve this step.\n\n### Training code\n\nWe have created a python project with loggers (wandb), which runs from the command line with YAML config. You can find it on the github: https://github.com/25icecreamflavors/contrails\n# Sources\n- https://www.kaggle.com/code/shashwatraman/contrails-dataset-ash-color\n- https://www.kaggle.com/competitions/understanding_cloud_organization/discussion/118080\n- https://github.com/25icecreamflavors/contrails",
    "2395834": ""
  }
}