{
  "id": 432998,
  "title": "4th place solution",
  "url": "/competitions/google-research-identify-contrails-reduce-global-warming/discussion/432998",
  "author_name": "Ivan Panshin",
  "post_date": "2023-08-19T21:46:31.342000",
  "votes": 18,
  "comment_count": 11,
  "views": 0,
  "content": "<p>This was a fantastic competition - thanks to organizers and my team. Past 3 months have not been easy, but I’m glad that in the end it was worth it.</p>\n<h3>Overview</h3>\n<ul>\n<li>U-Net models with a diverse set of encoders (both convolution-based, and transformer-based).</li>\n<li>Creating pseudo-labels for the unlabelled part of the data (7/8 of all the data) using folds and ensembles.</li>\n<li>Composite loss (CE, Dice, Focal).</li>\n<li>EMA + SWA.</li>\n<li>High resolution (512 + 768).</li>\n<li>4TTA (hflip, rot90, rot270).</li>\n</ul>\n<h3>Base models</h3>\n<p>Quite early we realized that heavy models and long training work well here, so we mostly ran experiments with backbones like effnet_v2_large and heavier. The final solution includes:</p>\n<ul>\n<li>effnet_v2_l.</li>\n<li>effnet_v2_xl.</li>\n<li>effnet_l2.</li>\n<li>maxvit.</li>\n</ul>\n<p>It takes about a week to train with 4xA6000 (most of the training time goes into creation of good pseudo-labels).</p>\n<h3>Validation</h3>\n<p>After we figured out that some geometric augmentations don’t really work (for example, vertical flips), we removed them from the augmentation pipeline. Additionally, the 4TTA validation and regular one didn’t have a perfect correlation, so the validation during training was performed with 4TTA, which gave a significant boost early in the competition. </p>\n<p>In terms of split - the splitting proposed by organizers was used. </p>\n<h3>Pseudo labels</h3>\n<p>The original dataset (of 20519 records, each of 8 images) is already quite big for U-Net like models, however we wanted to expand it further. That’s why we created pseudo-labels for the whole dataset by training 3 models with 4 folds for 2 rounds. The following scheme illustrates the idea:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2755695%2F0bf39e12f3004d2d843bea9aa2944dee%2F2023-08-20%2000.42.48.jpg?generation=1692481385008583&amp;alt=media\" alt=\"\"></p>\n<p>Then we train final models with a sampler: 50% original data, 50% soft pseudo labels.</p>\n<h3>Additional tricks</h3>\n<p>Not a lot of things worked well here, so we added some simple tricks to improve the pipeline a bit. Those things include:</p>\n<ul>\n<li>Weighted loss between CE, Dice, and Focal. </li>\n<li>EMA during training.</li>\n<li>SWA on checkpoints after training.</li>\n<li>Remove BN from the decoder in U-Net. </li>\n<li>Add different resolutions in final ensemble (512 + 768).</li>\n</ul>\n<h3>Things that didn’t work</h3>\n<ul>\n<li>Twersky loss (with focus on either FN or FP).</li>\n<li>Post-processing to remove FP. </li>\n<li>3D models (including conv-lstm).</li>\n<li>Training on individual annotations (shame on us for not trying to train with mean of annotations).</li>\n<li>Figuring-out why geometric augmentations don’t work (our guess was annotators’ bias. Turns out, it was conversion bias that other top teams found).</li>\n<li>Heavy augmentations.</li>\n<li>Adding classifier.</li>\n<li>Predicting additional frames or using additional channels during training.</li>\n<li>Validation based on geography. </li>\n<li>Training only with positive data.</li>\n</ul>\n<h3>Links</h3>\n<ul>\n<li>Inference kernel: <a href=\"https://www.kaggle.com/code/selimsef/kdl-unet-768-inference-contrails\" target=\"_blank\">link</a>.</li>\n<li>GitHub repo with training code: <a href=\"https://github.com/selimsef/kaggle-identify-contrails-4th/\" target=\"_blank\">link</a>.</li>\n</ul>",
  "messages": [
    {
      "id": 2398737,
      "postDate": "2023-08-19T21:46:31.343Z",
      "content": "<p>This was a fantastic competition - thanks to organizers and my team. Past 3 months have not been easy, but I’m glad that in the end it was worth it.</p>\n<h3>Overview</h3>\n<ul>\n<li>U-Net models with a diverse set of encoders (both convolution-based, and transformer-based).</li>\n<li>Creating pseudo-labels for the unlabelled part of the data (7/8 of all the data) using folds and ensembles.</li>\n<li>Composite loss (CE, Dice, Focal).</li>\n<li>EMA + SWA.</li>\n<li>High resolution (512 + 768).</li>\n<li>4TTA (hflip, rot90, rot270).</li>\n</ul>\n<h3>Base models</h3>\n<p>Quite early we realized that heavy models and long training work well here, so we mostly ran experiments with backbones like effnet_v2_large and heavier. The final solution includes:</p>\n<ul>\n<li>effnet_v2_l.</li>\n<li>effnet_v2_xl.</li>\n<li>effnet_l2.</li>\n<li>maxvit.</li>\n</ul>\n<p>It takes about a week to train with 4xA6000 (most of the training time goes into creation of good pseudo-labels).</p>\n<h3>Validation</h3>\n<p>After we figured out that some geometric augmentations don’t really work (for example, vertical flips), we removed them from the augmentation pipeline. Additionally, the 4TTA validation and regular one didn’t have a perfect correlation, so the validation during training was performed with 4TTA, which gave a significant boost early in the competition. </p>\n<p>In terms of split - the splitting proposed by organizers was used. </p>\n<h3>Pseudo labels</h3>\n<p>The original dataset (of 20519 records, each of 8 images) is already quite big for U-Net like models, however we wanted to expand it further. That’s why we created pseudo-labels for the whole dataset by training 3 models with 4 folds for 2 rounds. The following scheme illustrates the idea:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2755695%2F0bf39e12f3004d2d843bea9aa2944dee%2F2023-08-20%2000.42.48.jpg?generation=1692481385008583&amp;alt=media\" alt=\"\"></p>\n<p>Then we train final models with a sampler: 50% original data, 50% soft pseudo labels.</p>\n<h3>Additional tricks</h3>\n<p>Not a lot of things worked well here, so we added some simple tricks to improve the pipeline a bit. Those things include:</p>\n<ul>\n<li>Weighted loss between CE, Dice, and Focal. </li>\n<li>EMA during training.</li>\n<li>SWA on checkpoints after training.</li>\n<li>Remove BN from the decoder in U-Net. </li>\n<li>Add different resolutions in final ensemble (512 + 768).</li>\n</ul>\n<h3>Things that didn’t work</h3>\n<ul>\n<li>Twersky loss (with focus on either FN or FP).</li>\n<li>Post-processing to remove FP. </li>\n<li>3D models (including conv-lstm).</li>\n<li>Training on individual annotations (shame on us for not trying to train with mean of annotations).</li>\n<li>Figuring-out why geometric augmentations don’t work (our guess was annotators’ bias. Turns out, it was conversion bias that other top teams found).</li>\n<li>Heavy augmentations.</li>\n<li>Adding classifier.</li>\n<li>Predicting additional frames or using additional channels during training.</li>\n<li>Validation based on geography. </li>\n<li>Training only with positive data.</li>\n</ul>\n<h3>Links</h3>\n<ul>\n<li>Inference kernel: <a href=\"https://www.kaggle.com/code/selimsef/kdl-unet-768-inference-contrails\" target=\"_blank\">link</a>.</li>\n<li>GitHub repo with training code: <a href=\"https://github.com/selimsef/kaggle-identify-contrails-4th/\" target=\"_blank\">link</a>.</li>\n</ul>",
      "rawMarkdown": "This was a fantastic competition - thanks to organizers and my team. Past 3 months have not been easy, but I’m glad that in the end it was worth it.\n\n### Overview\n\n- U-Net models with a diverse set of encoders (both convolution-based, and transformer-based).\n- Creating pseudo-labels for the unlabelled part of the data (7/8 of all the data) using folds and ensembles.\n- Composite loss (CE, Dice, Focal).\n- EMA + SWA.\n- High resolution (512 + 768).\n- 4TTA (hflip, rot90, rot270).\n\n### Base models\n\nQuite early we realized that heavy models and long training work well here, so we mostly ran experiments with backbones like effnet_v2_large and heavier. The final solution includes:\n\n- effnet_v2_l.\n- effnet_v2_xl.\n- effnet_l2.\n- maxvit.\n\nIt takes about a week to train with 4xA6000 (most of the training time goes into creation of good pseudo-labels).\n\n### Validation\n\nAfter we figured out that some geometric augmentations don’t really work (for example, vertical flips), we removed them from the augmentation pipeline. Additionally, the 4TTA validation and regular one didn’t have a perfect correlation, so the validation during training was performed with 4TTA, which gave a significant boost early in the competition. \n\nIn terms of split - the splitting proposed by organizers was used. \n\n### Pseudo labels\n\nThe original dataset (of 20519 records, each of 8 images) is already quite big for U-Net like models, however we wanted to expand it further. That’s why we created pseudo-labels for the whole dataset by training 3 models with 4 folds for 2 rounds. The following scheme illustrates the idea:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2755695%2F0bf39e12f3004d2d843bea9aa2944dee%2F2023-08-20%2000.42.48.jpg?generation=1692481385008583&alt=media)\n\nThen we train final models with a sampler: 50% original data, 50% soft pseudo labels.\n\n### Additional tricks\nNot a lot of things worked well here, so we added some simple tricks to improve the pipeline a bit. Those things include:\n\n- Weighted loss between CE, Dice, and Focal. \n- EMA during training.\n- SWA on checkpoints after training.\n- Remove BN from the decoder in U-Net. \n- Add different resolutions in final ensemble (512 + 768).\n\n### Things that didn’t work\n\n- Twersky loss (with focus on either FN or FP).\n- Post-processing to remove FP. \n- 3D models (including conv-lstm).\n- Training on individual annotations (shame on us for not trying to train with mean of annotations).\n- Figuring-out why geometric augmentations don’t work (our guess was annotators’ bias. Turns out, it was conversion bias that other top teams found).\n- Heavy augmentations.\n- Adding classifier.\n- Predicting additional frames or using additional channels during training.\n- Validation based on geography. \n- Training only with positive data.\n\n### Links \n\t\n- Inference kernel: [link](https://www.kaggle.com/code/selimsef/kdl-unet-768-inference-contrails).\n- GitHub repo with training code: [link](https://github.com/selimsef/kaggle-identify-contrails-4th/).\n",
      "votes": 18
    },
    {
      "id": 2401561,
      "postDate": "2023-08-21T16:53:22.817Z",
      "content": "<p>Congratulations and thanks for sharing, it is a very impressive result, and even more w/o correcting the labels and using the mean of annotations.</p>\n<p>In this case, what would be the advantage of doing pseudolabeling using folds (as you did) vs doing pseudolabeling using all training data at once (i.e. train the 3 models with all available 4th slice training data to predict the 5-8 slices)? </p>",
      "rawMarkdown": "Congratulations and thanks for sharing, it is a very impressive result, and even more w/o correcting the labels and using the mean of annotations.\n\nIn this case, what would be the advantage of doing pseudolabeling using folds (as you did) vs doing pseudolabeling using all training data at once (i.e. train the 3 models with all available 4th slice training data to predict the 5-8 slices)? ",
      "votes": 1,
      "replies": [
        {
          "id": 2401580,
          "postDate": "2023-08-21T17:02:24.637Z",
          "content": "<p>With OOF one can run a few rounds of pseudolabeling. <br>\nOtherwise confirmation bias will lead to zero improvements after the 1st round.</p>",
          "rawMarkdown": "With OOF one can run a few rounds of pseudolabeling. \nOtherwise confirmation bias will lead to zero improvements after the 1st round.",
          "votes": 4,
          "replies": [
            {
              "id": 2403559,
              "postDate": "2023-08-22T18:08:10.020Z",
              "content": "<p><a href=\"https://www.kaggle.com/selimsef\" target=\"_blank\">@selimsef</a> could you clarify the thing about the OOF please? Am I correct, that you firstly train the model on 3 folds, leaving the 4th fold for the validation? Then you do pseudo. Thereafter you train models again on the previous 3 folds+pseudo. And then you are doing pseudo again (the 2nd round)? So basically, do you leave the same fold for the validation each time? </p>",
              "rawMarkdown": "@selimsef could you clarify the thing about the OOF please? Am I correct, that you firstly train the model on 3 folds, leaving the 4th fold for the validation? Then you do pseudo. Thereafter you train models again on the previous 3 folds+pseudo. And then you are doing pseudo again (the 2nd round)? So basically, do you leave the same fold for the validation each time? "
            },
            {
              "id": 2403591,
              "postDate": "2023-08-22T18:22:27.100Z",
              "content": "<p>Generally, it's crucial to make out-of-fold (OOF) predictions at each stage (standard k-fold approach). That's the only requirement in this process. So it will work even if you train with 3-4-5-6-7 folds in 5 stage pseudo it will work as well.</p>",
              "rawMarkdown": "Generally, it's crucial to make out-of-fold (OOF) predictions at each stage (standard k-fold approach). That's the only requirement in this process. So it will work even if you train with 3-4-5-6-7 folds in 5 stage pseudo it will work as well.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2398747,
      "postDate": "2023-08-19T22:26:09.033Z",
      "content": "<p>Wow! what a laborious project! Congratulations <a href=\"https://www.kaggle.com/ivanpan\" target=\"_blank\">@ivanpan</a> <br>\nTransformers seem to be very useful in various AI problems. Can you please suggest me a good source to learn them?</p>",
      "rawMarkdown": "Wow! what a laborious project! Congratulations @ivanpan \nTransformers seem to be very useful in various AI problems. Can you please suggest me a good source to learn them?",
      "votes": 2,
      "replies": [
        {
          "id": 2402745,
          "postDate": "2023-08-22T10:07:32.503Z",
          "content": "<p>Sure. The main idea that you should understand before diving into any transformers (whether tailored for Computer Vision or NLP or anything else) is attention mechanism that was proposed (or at least popularized) in <a href=\"https://arxiv.org/abs/1706.03762\" target=\"_blank\">this paper</a>.</p>\n<p>I heavily recommend <a href=\"http://jalammar.github.io/illustrated-transformer/\" target=\"_blank\">this article</a> that explains it the best. </p>\n<p>After you're done with the fundamentals, just google \"TRANSFORMER_NAME explained medium\" to understand any particular architecture. For implementations - look either at <a href=\"https://github.com/huggingface/pytorch-image-models\" target=\"_blank\">timm</a> or <a href=\"http://huggingface.co/\" target=\"_blank\">HuggingFace</a>.  </p>\n<p>To give you pointers, for NLP check out BERT, T5 and GPT models. For CV - ViT and Swin. </p>",
          "rawMarkdown": "Sure. The main idea that you should understand before diving into any transformers (whether tailored for Computer Vision or NLP or anything else) is attention mechanism that was proposed (or at least popularized) in [this paper](https://arxiv.org/abs/1706.03762).\n\nI heavily recommend [this article](http://jalammar.github.io/illustrated-transformer/) that explains it the best. \n\nAfter you're done with the fundamentals, just google \"TRANSFORMER_NAME explained medium\" to understand any particular architecture. For implementations - look either at [timm](https://github.com/huggingface/pytorch-image-models) or [HuggingFace](http://huggingface.co/).  \n\nTo give you pointers, for NLP check out BERT, T5 and GPT models. For CV - ViT and Swin. ",
          "votes": 2,
          "replies": [
            {
              "id": 2402888,
              "postDate": "2023-08-22T11:38:37.103Z",
              "content": "<p>Thanks a lot. you really helped me🙏</p>",
              "rawMarkdown": "Thanks a lot. you really helped me🙏",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2413580,
      "postDate": "2023-08-29T03:04:52.660Z",
      "content": "<p>I'm sorry to ask three very novice questions:</p>\n<ol>\n<li>Is TTA only used during the testing phase? If so, what needs to be done during the validation phase?</li>\n<li>Where does the data from the pseudo label stage come from? What is the confidence level of the pseudo label to be used? And what does OOF in the picture mean?</li>\n<li>How did you find hyperparameters such as Composite loss and backbone? Did you use some automl tools? If so, which tool are you using?</li>\n</ol>\n<p>Thank you very much, if you can answer</p>",
      "rawMarkdown": "I'm sorry to ask three very novice questions:\n\n1. Is TTA only used during the testing phase? If so, what needs to be done during the validation phase?\n2. Where does the data from the pseudo label stage come from? What is the confidence level of the pseudo label to be used? And what does OOF in the picture mean?\n3. How did you find hyperparameters such as Composite loss and backbone? Did you use some automl tools? If so, which tool are you using?\n\nThank you very much, if you can answer"
    },
    {
      "id": 2398881,
      "postDate": "2023-08-20T03:29:36.413Z",
      "content": "<p>Wow very complex and great work <a href=\"https://www.kaggle.com/ivanpan\" target=\"_blank\">@ivanpan</a> !</p>",
      "rawMarkdown": "Wow very complex and great work @ivanpan !"
    },
    {
      "id": 2413575,
      "postDate": "2023-08-29T02:51:33.860Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2413572,
      "postDate": "2023-08-29T02:47:23.953Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2401561,
      "author_name": "delai50",
      "author_url": "",
      "post_date": "2023-08-21T16:53:22.817000",
      "content": "<p>Congratulations and thanks for sharing, it is a very impressive result, and even more w/o correcting the labels and using the mean of annotations.</p>\n<p>In this case, what would be the advantage of doing pseudolabeling using folds (as you did) vs doing pseudolabeling using all training data at once (i.e. train the 3 models with all available 4th slice training data to predict the 5-8 slices)? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2401580,
          "author_name": "Selim Seferbekov",
          "author_url": "",
          "post_date": "2023-08-21T17:02:24.637000",
          "content": "<p>With OOF one can run a few rounds of pseudolabeling. <br>\nOtherwise confirmation bias will lead to zero improvements after the 1st round.</p>",
          "votes": 4,
          "replies": [
            {
              "id": 2403559,
              "author_name": "Man of the year",
              "author_url": "",
              "post_date": "2023-08-22T18:08:10.020000",
              "content": "<p><a href=\"https://www.kaggle.com/selimsef\" target=\"_blank\">@selimsef</a> could you clarify the thing about the OOF please? Am I correct, that you firstly train the model on 3 folds, leaving the 4th fold for the validation? Then you do pseudo. Thereafter you train models again on the previous 3 folds+pseudo. And then you are doing pseudo again (the 2nd round)? So basically, do you leave the same fold for the validation each time? </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2403591,
              "author_name": "Selim Seferbekov",
              "author_url": "",
              "post_date": "2023-08-22T18:22:27.100000",
              "content": "<p>Generally, it's crucial to make out-of-fold (OOF) predictions at each stage (standard k-fold approach). That's the only requirement in this process. So it will work even if you train with 3-4-5-6-7 folds in 5 stage pseudo it will work as well.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2398747,
      "author_name": "NewToAI_i",
      "author_url": "",
      "post_date": "2023-08-19T22:26:09.033000",
      "content": "<p>Wow! what a laborious project! Congratulations <a href=\"https://www.kaggle.com/ivanpan\" target=\"_blank\">@ivanpan</a> <br>\nTransformers seem to be very useful in various AI problems. Can you please suggest me a good source to learn them?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2402745,
          "author_name": "Ivan Panshin",
          "author_url": "",
          "post_date": "2023-08-22T10:07:32.503000",
          "content": "<p>Sure. The main idea that you should understand before diving into any transformers (whether tailored for Computer Vision or NLP or anything else) is attention mechanism that was proposed (or at least popularized) in <a href=\"https://arxiv.org/abs/1706.03762\" target=\"_blank\">this paper</a>.</p>\n<p>I heavily recommend <a href=\"http://jalammar.github.io/illustrated-transformer/\" target=\"_blank\">this article</a> that explains it the best. </p>\n<p>After you're done with the fundamentals, just google \"TRANSFORMER_NAME explained medium\" to understand any particular architecture. For implementations - look either at <a href=\"https://github.com/huggingface/pytorch-image-models\" target=\"_blank\">timm</a> or <a href=\"http://huggingface.co/\" target=\"_blank\">HuggingFace</a>.  </p>\n<p>To give you pointers, for NLP check out BERT, T5 and GPT models. For CV - ViT and Swin. </p>",
          "votes": 2,
          "replies": [
            {
              "id": 2402888,
              "author_name": "NewToAI_i",
              "author_url": "",
              "post_date": "2023-08-22T11:38:37.103000",
              "content": "<p>Thanks a lot. you really helped me🙏</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2413580,
      "author_name": "kongweihao",
      "author_url": "",
      "post_date": "2023-08-29T03:04:52.660000",
      "content": "<p>I'm sorry to ask three very novice questions:</p>\n<ol>\n<li>Is TTA only used during the testing phase? If so, what needs to be done during the validation phase?</li>\n<li>Where does the data from the pseudo label stage come from? What is the confidence level of the pseudo label to be used? And what does OOF in the picture mean?</li>\n<li>How did you find hyperparameters such as Composite loss and backbone? Did you use some automl tools? If so, which tool are you using?</li>\n</ol>\n<p>Thank you very much, if you can answer</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2398881,
      "author_name": "Arvind (Yetirajan) Narayanan Iyengar",
      "author_url": "",
      "post_date": "2023-08-20T03:29:36.413000",
      "content": "<p>Wow very complex and great work <a href=\"https://www.kaggle.com/ivanpan\" target=\"_blank\">@ivanpan</a> !</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2413575,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-08-29T02:51:33.860000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2413572,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-08-29T02:47:23.953000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2398737": "This was a fantastic competition - thanks to organizers and my team. Past 3 months have not been easy, but I’m glad that in the end it was worth it.\n\n### Overview\n\n- U-Net models with a diverse set of encoders (both convolution-based, and transformer-based).\n- Creating pseudo-labels for the unlabelled part of the data (7/8 of all the data) using folds and ensembles.\n- Composite loss (CE, Dice, Focal).\n- EMA + SWA.\n- High resolution (512 + 768).\n- 4TTA (hflip, rot90, rot270).\n\n### Base models\n\nQuite early we realized that heavy models and long training work well here, so we mostly ran experiments with backbones like effnet_v2_large and heavier. The final solution includes:\n\n- effnet_v2_l.\n- effnet_v2_xl.\n- effnet_l2.\n- maxvit.\n\nIt takes about a week to train with 4xA6000 (most of the training time goes into creation of good pseudo-labels).\n\n### Validation\n\nAfter we figured out that some geometric augmentations don’t really work (for example, vertical flips), we removed them from the augmentation pipeline. Additionally, the 4TTA validation and regular one didn’t have a perfect correlation, so the validation during training was performed with 4TTA, which gave a significant boost early in the competition. \n\nIn terms of split - the splitting proposed by organizers was used. \n\n### Pseudo labels\n\nThe original dataset (of 20519 records, each of 8 images) is already quite big for U-Net like models, however we wanted to expand it further. That’s why we created pseudo-labels for the whole dataset by training 3 models with 4 folds for 2 rounds. The following scheme illustrates the idea:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2755695%2F0bf39e12f3004d2d843bea9aa2944dee%2F2023-08-20%2000.42.48.jpg?generation=1692481385008583&alt=media)\n\nThen we train final models with a sampler: 50% original data, 50% soft pseudo labels.\n\n### Additional tricks\nNot a lot of things worked well here, so we added some simple tricks to improve the pipeline a bit. Those things include:\n\n- Weighted loss between CE, Dice, and Focal. \n- EMA during training.\n- SWA on checkpoints after training.\n- Remove BN from the decoder in U-Net. \n- Add different resolutions in final ensemble (512 + 768).\n\n### Things that didn’t work\n\n- Twersky loss (with focus on either FN or FP).\n- Post-processing to remove FP. \n- 3D models (including conv-lstm).\n- Training on individual annotations (shame on us for not trying to train with mean of annotations).\n- Figuring-out why geometric augmentations don’t work (our guess was annotators’ bias. Turns out, it was conversion bias that other top teams found).\n- Heavy augmentations.\n- Adding classifier.\n- Predicting additional frames or using additional channels during training.\n- Validation based on geography. \n- Training only with positive data.\n\n### Links \n\t\n- Inference kernel: [link](https://www.kaggle.com/code/selimsef/kdl-unet-768-inference-contrails).\n- GitHub repo with training code: [link](https://github.com/selimsef/kaggle-identify-contrails-4th/).\n",
    "2401561": "Congratulations and thanks for sharing, it is a very impressive result, and even more w/o correcting the labels and using the mean of annotations.\n\nIn this case, what would be the advantage of doing pseudolabeling using folds (as you did) vs doing pseudolabeling using all training data at once (i.e. train the 3 models with all available 4th slice training data to predict the 5-8 slices)? ",
    "2398747": "Wow! what a laborious project! Congratulations @ivanpan \nTransformers seem to be very useful in various AI problems. Can you please suggest me a good source to learn them?",
    "2413580": "I'm sorry to ask three very novice questions:\n\n1. Is TTA only used during the testing phase? If so, what needs to be done during the validation phase?\n2. Where does the data from the pseudo label stage come from? What is the confidence level of the pseudo label to be used? And what does OOF in the picture mean?\n3. How did you find hyperparameters such as Composite loss and backbone? Did you use some automl tools? If so, which tool are you using?\n\nThank you very much, if you can answer",
    "2398881": "Wow very complex and great work @ivanpan !",
    "2413575": "",
    "2413572": ""
  }
}