{
  "id": 430581,
  "title": "6th place solution (CV: 0.7069 / Public LB: 0.706 / Private LB: 0.713) ",
  "url": "/competitions/google-research-identify-contrails-reduce-global-warming/discussion/430581",
  "author_name": "tattaka",
  "post_date": "2023-08-10T10:50:44.238000",
  "votes": 39,
  "comment_count": 13,
  "views": 0,
  "content": "<p>First of all, we would like to thank the organizers of the competition and the Kaggle team for the competition hosting and all the participants who have shared their knowledge so generously. </p>\n<h1>Summary</h1>\n<ul>\n<li>2 stage pipeline: classification and segmentation</li>\n<li>ensemble of 2.5D model using 5 or 7 frame and 11ch input 2D model</li>\n<li>soft label using individual label</li>\n<li>percentile threshold</li>\n</ul>\n<h1>Pipeline</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1745801%2Fdeb33f0f5a0349334ed541ec39c2e353%2F2023-08-10%2016.57.28.png?generation=1691664174212060&amp;alt=media\" alt=\"\"></p>\n<p>This pipeline has been used frequently in past semantic segmentation competitions.    <br>\nThe advantage of this pipeline is that it creates two models, one trained on pos_only and the other on all data, thereby increasing the ensemble effect.    <br>\nIt also saves inference time by allowing more models to be assigned to images for which screening has determined that a mask is present.    </p>\n<h1>Models Detail</h1>\n<h2>Common Settings</h2>\n<p>Both 2.5D and 2D models were trained for both hard and soft labels.  <br>\nThe soft label is created as follows</p>\n<pre><code>label = self.np_load(\n    os.path.join()\n).astype(np.float32)\nlabel = np.clip((label * ).(-) / label.shape[-], , )\n</code></pre>\n<p>All models were trained by all the data in train folder and validated by the one in validation folder. </p>\n<h2>2.5D Model</h2>\n<p>We used a modified <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/417255\" target=\"_blank\">2.5D model used in Vesuvius Challenge</a>.<br>\nThe main changes are as follows</p>\n<ul>\n<li>Apply 3D CNN to feature maps of all resolutions</li>\n<li>3D For the output of the 3D CNN, feature maps corresponding to the target frame are extracted and given to UNet  </li>\n<li>In the binary classification model, the 3DCNN is applied to the smallest resolution feature map and input to the FC layer in the same way.    </li>\n</ul>\n<p>Backbones used were resnetrs101, resnest101e, swin_base_patch4_window12, swinv2_base_window16, convnext_base, convnext_large, etc.  <br>\nThe input frame is set to 5 or 7 frames, with the target frame in the middle.    <br>\nInput resolution of 384 or 512 are used for the model.  <br>\nIt was trained using <code>BCE+Dice*0.25</code> for the pos only model and <code>BCE</code> for the all data model.  <br>\nAlso, different augmentation is used for each of the classification model, all data seg model, and pos only seg model. (possibly not optimized and therefore not appropriate).    </p>\n<pre><code>\nalbu.Flip(p=),\nalbu.RandomRotate90(p=),\nalbu.ShiftScaleRotate(p=, scale_limit=, shift_limit=),\nalbu.GridDistortion(num_steps=, distort_limit=, p=),\nalbu.CoarseDropout(max_height=, max_width=, fill_value=, mask_fill_value=, p=,)\n\n\nalbu.ShiftScaleRotate(p=, scale_limit=, shift_limit=),\n\n\nalbu.Flip(p=),\nalbu.RandomRotate90(p=),\nalbu.ShiftScaleRotate(p=, scale_limit=, shift_limit=, rotate_limit=),\nalbu.Rotate(limit=, p=),\n</code></pre>\n<ul>\n<li>EMA(decay=0.998)</li>\n</ul>\n<h2>2D Model</h2>\n<p>For 2D models, we simply used Unet implemented in Segmentation Models Pytorch. The backbones used were res2net50d, regnetz_d8, regnetz_d32, and regnetz_e8.</p>\n<p>Characteristically, the 2D models were trained on 11 channels, i.e. band15 - band14, band14 - band11 and all 9 bands. Learning with this input does not perform well in short epochs (e.g. 20~30 epochs), but in long epochs (specifically 100 epochs), it obtained better validation scores than with false color input.</p>\n<p>The 11ch models also had one curious point. They had poor public score but good private score. I believe this is what pushed us to the prize zone.</p>\n<p>Other training settings are as follows(shared by all data seg model and pos only seg one):</p>\n<ul>\n<li>Input<ul>\n<li>size: 384x384</li>\n<li>each channels are standardized by subtracting the global mean and dividing by the global variance of the channel</li></ul></li>\n<li>Loss: <code>0.9*BCE + 0.1*Dice</code></li>\n<li>Data Augmentation</li>\n</ul>\n<pre><code>albu.HorizontalFlip(p=),\nalbu.VerticalFlip(p=),\nalbu.ShiftScaleRotate(p=, rotate_limit=)\nalbu.RandomResizedCrop(p: , scale=[, ], height=, width=)\n</code></pre>\n<ul>\n<li>EMA(decay=0.9999)</li>\n</ul>\n<h1>Postprocessing</h1>\n<p>From the time taken to inference, the percentage of empty masks in the entire test was found to be about the same as in validation set. (about 30%).  <br>\nTherefore, we used validation set to optimize the percentile threshold for both the classification part and segmentation part.  <br>\nAs a result, the percentile threshold is better than the optimized fixed threshold. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1745801%2F48a3c5b9e9ed108faaed1514dac63485%2F2023-08-10%2017.20.42.png?generation=1691664304846451&amp;alt=media\" alt=\"\"></p>\n<h1>Not Works</h1>\n<ul>\n<li>pseudo label・mean teacher</li>\n<li>2D pretrain -&gt; 2.5D finetuning<ul>\n<li>unstable</li></ul></li>\n<li>more augmentations</li>\n<li>mixup・cutmix・label smoothing</li>\n<li>TTA</li>\n<li>efficientnet</li>\n</ul>\n<h1>Training Code and Inference Notebook</h1>\n<p>We have released the training code and the inference notebook for the best submission.</p>\n<ul>\n<li>training code (2.5D model part): <a href=\"https://github.com/tattaka/google-research-identify-contrails-reduce-global-warming\" target=\"_blank\">https://github.com/tattaka/google-research-identify-contrails-reduce-global-warming</a></li>\n<li>training code (2D model part): <a href=\"https://github.com/tawatawara/kaggle-google-research-identify-contrails-reduce-global-warming\" target=\"_blank\">https://github.com/tawatawara/kaggle-google-research-identify-contrails-reduce-global-warming</a></li>\n<li>inference code: <a href=\"https://www.kaggle.com/code/tattaka/contrail-submission-ensemble?scriptVersionId=139455030\" target=\"_blank\">https://www.kaggle.com/code/tattaka/contrail-submission-ensemble?scriptVersionId=139455030</a></li>\n</ul>",
  "messages": [
    {
      "id": 2383422,
      "postDate": "2023-08-10T10:50:44.237Z",
      "content": "<p>First of all, we would like to thank the organizers of the competition and the Kaggle team for the competition hosting and all the participants who have shared their knowledge so generously. </p>\n<h1>Summary</h1>\n<ul>\n<li>2 stage pipeline: classification and segmentation</li>\n<li>ensemble of 2.5D model using 5 or 7 frame and 11ch input 2D model</li>\n<li>soft label using individual label</li>\n<li>percentile threshold</li>\n</ul>\n<h1>Pipeline</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1745801%2Fdeb33f0f5a0349334ed541ec39c2e353%2F2023-08-10%2016.57.28.png?generation=1691664174212060&amp;alt=media\" alt=\"\"></p>\n<p>This pipeline has been used frequently in past semantic segmentation competitions.    <br>\nThe advantage of this pipeline is that it creates two models, one trained on pos_only and the other on all data, thereby increasing the ensemble effect.    <br>\nIt also saves inference time by allowing more models to be assigned to images for which screening has determined that a mask is present.    </p>\n<h1>Models Detail</h1>\n<h2>Common Settings</h2>\n<p>Both 2.5D and 2D models were trained for both hard and soft labels.  <br>\nThe soft label is created as follows</p>\n<pre><code>label = self.np_load(\n    os.path.join()\n).astype(np.float32)\nlabel = np.clip((label * ).(-) / label.shape[-], , )\n</code></pre>\n<p>All models were trained by all the data in train folder and validated by the one in validation folder. </p>\n<h2>2.5D Model</h2>\n<p>We used a modified <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/417255\" target=\"_blank\">2.5D model used in Vesuvius Challenge</a>.<br>\nThe main changes are as follows</p>\n<ul>\n<li>Apply 3D CNN to feature maps of all resolutions</li>\n<li>3D For the output of the 3D CNN, feature maps corresponding to the target frame are extracted and given to UNet  </li>\n<li>In the binary classification model, the 3DCNN is applied to the smallest resolution feature map and input to the FC layer in the same way.    </li>\n</ul>\n<p>Backbones used were resnetrs101, resnest101e, swin_base_patch4_window12, swinv2_base_window16, convnext_base, convnext_large, etc.  <br>\nThe input frame is set to 5 or 7 frames, with the target frame in the middle.    <br>\nInput resolution of 384 or 512 are used for the model.  <br>\nIt was trained using <code>BCE+Dice*0.25</code> for the pos only model and <code>BCE</code> for the all data model.  <br>\nAlso, different augmentation is used for each of the classification model, all data seg model, and pos only seg model. (possibly not optimized and therefore not appropriate).    </p>\n<pre><code>\nalbu.Flip(p=),\nalbu.RandomRotate90(p=),\nalbu.ShiftScaleRotate(p=, scale_limit=, shift_limit=),\nalbu.GridDistortion(num_steps=, distort_limit=, p=),\nalbu.CoarseDropout(max_height=, max_width=, fill_value=, mask_fill_value=, p=,)\n\n\nalbu.ShiftScaleRotate(p=, scale_limit=, shift_limit=),\n\n\nalbu.Flip(p=),\nalbu.RandomRotate90(p=),\nalbu.ShiftScaleRotate(p=, scale_limit=, shift_limit=, rotate_limit=),\nalbu.Rotate(limit=, p=),\n</code></pre>\n<ul>\n<li>EMA(decay=0.998)</li>\n</ul>\n<h2>2D Model</h2>\n<p>For 2D models, we simply used Unet implemented in Segmentation Models Pytorch. The backbones used were res2net50d, regnetz_d8, regnetz_d32, and regnetz_e8.</p>\n<p>Characteristically, the 2D models were trained on 11 channels, i.e. band15 - band14, band14 - band11 and all 9 bands. Learning with this input does not perform well in short epochs (e.g. 20~30 epochs), but in long epochs (specifically 100 epochs), it obtained better validation scores than with false color input.</p>\n<p>The 11ch models also had one curious point. They had poor public score but good private score. I believe this is what pushed us to the prize zone.</p>\n<p>Other training settings are as follows(shared by all data seg model and pos only seg one):</p>\n<ul>\n<li>Input<ul>\n<li>size: 384x384</li>\n<li>each channels are standardized by subtracting the global mean and dividing by the global variance of the channel</li></ul></li>\n<li>Loss: <code>0.9*BCE + 0.1*Dice</code></li>\n<li>Data Augmentation</li>\n</ul>\n<pre><code>albu.HorizontalFlip(p=),\nalbu.VerticalFlip(p=),\nalbu.ShiftScaleRotate(p=, rotate_limit=)\nalbu.RandomResizedCrop(p: , scale=[, ], height=, width=)\n</code></pre>\n<ul>\n<li>EMA(decay=0.9999)</li>\n</ul>\n<h1>Postprocessing</h1>\n<p>From the time taken to inference, the percentage of empty masks in the entire test was found to be about the same as in validation set. (about 30%).  <br>\nTherefore, we used validation set to optimize the percentile threshold for both the classification part and segmentation part.  <br>\nAs a result, the percentile threshold is better than the optimized fixed threshold. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1745801%2F48a3c5b9e9ed108faaed1514dac63485%2F2023-08-10%2017.20.42.png?generation=1691664304846451&amp;alt=media\" alt=\"\"></p>\n<h1>Not Works</h1>\n<ul>\n<li>pseudo label・mean teacher</li>\n<li>2D pretrain -&gt; 2.5D finetuning<ul>\n<li>unstable</li></ul></li>\n<li>more augmentations</li>\n<li>mixup・cutmix・label smoothing</li>\n<li>TTA</li>\n<li>efficientnet</li>\n</ul>\n<h1>Training Code and Inference Notebook</h1>\n<p>We have released the training code and the inference notebook for the best submission.</p>\n<ul>\n<li>training code (2.5D model part): <a href=\"https://github.com/tattaka/google-research-identify-contrails-reduce-global-warming\" target=\"_blank\">https://github.com/tattaka/google-research-identify-contrails-reduce-global-warming</a></li>\n<li>training code (2D model part): <a href=\"https://github.com/tawatawara/kaggle-google-research-identify-contrails-reduce-global-warming\" target=\"_blank\">https://github.com/tawatawara/kaggle-google-research-identify-contrails-reduce-global-warming</a></li>\n<li>inference code: <a href=\"https://www.kaggle.com/code/tattaka/contrail-submission-ensemble?scriptVersionId=139455030\" target=\"_blank\">https://www.kaggle.com/code/tattaka/contrail-submission-ensemble?scriptVersionId=139455030</a></li>\n</ul>",
      "rawMarkdown": "First of all, we would like to thank the organizers of the competition and the Kaggle team for the competition hosting and all the participants who have shared their knowledge so generously. \n\n# Summary\n* 2 stage pipeline: classification and segmentation\n* ensemble of 2.5D model using 5 or 7 frame and 11ch input 2D model\n* soft label using individual label\n* percentile threshold\n\n# Pipeline\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1745801%2Fdeb33f0f5a0349334ed541ec39c2e353%2F2023-08-10%2016.57.28.png?generation=1691664174212060&alt=media)\n\nThis pipeline has been used frequently in past semantic segmentation competitions.    \nThe advantage of this pipeline is that it creates two models, one trained on pos_only and the other on all data, thereby increasing the ensemble effect.    \nIt also saves inference time by allowing more models to be assigned to images for which screening has determined that a mask is present.    \n\n# Models Detail\n## Common Settings\nBoth 2.5D and 2D models were trained for both hard and soft labels.  \nThe soft label is created as follows\n``` python\nlabel = self.np_load(\n    os.path.join(\"mask_individual.npy\")\n).astype(np.float32)\nlabel = np.clip((label * 2).sum(-1) / label.shape[-1], 0, 1)\n```\n\nAll models were trained by all the data in train folder and validated by the one in validation folder. \n\n## 2.5D Model\nWe used a modified [2.5D model used in Vesuvius Challenge](https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/417255).\nThe main changes are as follows\n* Apply 3D CNN to feature maps of all resolutions\n* 3D For the output of the 3D CNN, feature maps corresponding to the target frame are extracted and given to UNet  \n* In the binary classification model, the 3DCNN is applied to the smallest resolution feature map and input to the FC layer in the same way.    \n\nBackbones used were resnetrs101, resnest101e, swin_base_patch4_window12, swinv2_base_window16, convnext_base, convnext_large, etc.  \nThe input frame is set to 5 or 7 frames, with the target frame in the middle.    \nInput resolution of 384 or 512 are used for the model.  \nIt was trained using `BCE+Dice*0.25` for the pos only model and `BCE` for the all data model.  \nAlso, different augmentation is used for each of the classification model, all data seg model, and pos only seg model. (possibly not optimized and therefore not appropriate).    \n\n``` python\n# classification\nalbu.Flip(p=0.5),\nalbu.RandomRotate90(p=0.5),\nalbu.ShiftScaleRotate(p=0.5, scale_limit=0.3, shift_limit=0.1),\nalbu.GridDistortion(num_steps=5, distort_limit=0.3, p=0.5),\nalbu.CoarseDropout(max_height=16, max_width=16, fill_value=255, mask_fill_value=0, p=0.5,)\n\n# all data segmentation\nalbu.ShiftScaleRotate(p=0.5, scale_limit=0.3, shift_limit=0.1),\n\n# pos only segmentation\nalbu.Flip(p=0.25),\nalbu.RandomRotate90(p=0.25),\nalbu.ShiftScaleRotate(p=0.5, scale_limit=0.3, shift_limit=0.1, rotate_limit=0),\nalbu.Rotate(limit=45, p=0.25),\n```\n* EMA(decay=0.998)\n\n## 2D Model\nFor 2D models, we simply used Unet implemented in Segmentation Models Pytorch. The backbones used were res2net50d, regnetz_d8, regnetz_d32, and regnetz_e8.\n\nCharacteristically, the 2D models were trained on 11 channels, i.e. band15 - band14, band14 - band11 and all 9 bands. Learning with this input does not perform well in short epochs (e.g. 20~30 epochs), but in long epochs (specifically 100 epochs), it obtained better validation scores than with false color input.\n\nThe 11ch models also had one curious point. They had poor public score but good private score. I believe this is what pushed us to the prize zone.\n\nOther training settings are as follows(shared by all data seg model and pos only seg one):\n* Input\n    * size: 384x384\n    * each channels are standardized by subtracting the global mean and dividing by the global variance of the channel\n* Loss: `0.9*BCE + 0.1*Dice`\n* Data Augmentation\n``` python\nalbu.HorizontalFlip(p=0.5),\nalbu.VerticalFlip(p=0.5),\nalbu.ShiftScaleRotate(p=0.5, rotate_limit=90)\nalbu.RandomResizedCrop(p: 1.0, scale=[0.875, 1.0], height=256, width=256)\n```\n* EMA(decay=0.9999)\n\n# Postprocessing\nFrom the time taken to inference, the percentage of empty masks in the entire test was found to be about the same as in validation set. (about 30%).  \nTherefore, we used validation set to optimize the percentile threshold for both the classification part and segmentation part.  \nAs a result, the percentile threshold is better than the optimized fixed threshold. \n \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1745801%2F48a3c5b9e9ed108faaed1514dac63485%2F2023-08-10%2017.20.42.png?generation=1691664304846451&alt=media)\n\n# Not Works\n* pseudo label・mean teacher\n* 2D pretrain -> 2.5D finetuning\n  * unstable\n* more augmentations\n* mixup・cutmix・label smoothing\n* TTA\n* efficientnet\n\n# Training Code and Inference Notebook\nWe have released the training code and the inference notebook for the best submission.\n\n* training code (2.5D model part): https://github.com/tattaka/google-research-identify-contrails-reduce-global-warming\n* training code (2D model part): https://github.com/tawatawara/kaggle-google-research-identify-contrails-reduce-global-warming\n* inference code: https://www.kaggle.com/code/tattaka/contrail-submission-ensemble?scriptVersionId=139455030",
      "votes": 38
    },
    {
      "id": 2386414,
      "postDate": "2023-08-11T23:28:54.980Z",
      "content": "<p><strong>Update</strong><br>\nWe have released the training code(2.5D model part) and the inference notebook for the best submission.</p>\n<ul>\n<li>training code (2.5D model part): <a href=\"https://github.com/tattaka/google-research-identify-contrails-reduce-global-warming\" target=\"_blank\">https://github.com/tattaka/google-research-identify-contrails-reduce-global-warming</a></li>\n<li>inference code: <a href=\"https://www.kaggle.com/code/tattaka/contrail-submission-ensemble?scriptVersionId=139455030\" target=\"_blank\">https://www.kaggle.com/code/tattaka/contrail-submission-ensemble?scriptVersionId=139455030</a></li>\n</ul>",
      "rawMarkdown": "**Update**\nWe have released the training code(2.5D model part) and the inference notebook for the best submission.\n\n* training code (2.5D model part): https://github.com/tattaka/google-research-identify-contrails-reduce-global-warming\n* inference code: https://www.kaggle.com/code/tattaka/contrail-submission-ensemble?scriptVersionId=139455030",
      "votes": 5,
      "replies": [
        {
          "id": 2394546,
          "postDate": "2023-08-17T02:19:20.093Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 2386427,
      "postDate": "2023-08-11T23:56:03.057Z",
      "content": "<p>Sorry for late.</p>\n<p>We have released  the training code for 2D model part:  <br>\n<a href=\"https://github.com/tawatawara/kaggle-google-research-identify-contrails-reduce-global-warming\" target=\"_blank\">https://github.com/tawatawara/kaggle-google-research-identify-contrails-reduce-global-warming</a></p>",
      "rawMarkdown": "Sorry for late.\n\nWe have released  the training code for 2D model part:  \nhttps://github.com/tawatawara/kaggle-google-research-identify-contrails-reduce-global-warming",
      "votes": 1
    },
    {
      "id": 2393427,
      "postDate": "2023-08-16T10:03:16.487Z",
      "content": "<p>Congratulations and thanks a lot for sharing!</p>\n<p>Let me ask you a couple of questions 😃:</p>\n<ul>\n<li>If I am correct, the purpose of classification models is to determine whether the predicted mask is empty or not, and they are only applied to segmentation models trained on all data. However, what is the role of the \"top 100 logits mean per image\"?</li>\n</ul>\n<blockquote>\n  <p>From the time taken to inference, the percentage of empty masks in the entire test was found to be about the same as in validation set. (about 30%)</p>\n</blockquote>\n<ul>\n<li>Could you elaborate a bit more about the steps to find that?</li>\n</ul>",
      "rawMarkdown": "Congratulations and thanks a lot for sharing!\n\nLet me ask you a couple of questions 😃:\n\n- If I am correct, the purpose of classification models is to determine whether the predicted mask is empty or not, and they are only applied to segmentation models trained on all data. However, what is the role of the \"top 100 logits mean per image\"?\n\n>From the time taken to inference, the percentage of empty masks in the entire test was found to be about the same as in validation set. (about 30%)\n\n- Could you elaborate a bit more about the steps to find that?\n",
      "replies": [
        {
          "id": 2395102,
          "postDate": "2023-08-17T09:34:31.320Z",
          "content": "<blockquote>\n  <p>If I am correct, the purpose of classification models is to determine whether the predicted mask is empty or not, and they are only applied to segmentation models trained on all data. However, what is the role of the \"top 100 logits mean per image\"?</p>\n</blockquote>\n<p>\"top 100 logits mean per image\" is a variation of the classification model using the segmentation model.<br>\nA similar approach was used in a previous competition, and again the top 100 logits mean was more accurate for binary classification than the classification model.</p>\n<blockquote>\n  <p>Could you elaborate a bit more about the steps to find that?</p>\n</blockquote>\n<p>In binary classification, we have both submissions with a percentile threshold and sub with the percentile converted to a fixed value and used as the threshold. Comparing the execution times of the two submissions, we can see that the percentage of empty mask images differs when there is a change. (since the size of the test set and the validation set are the same)</p>",
          "rawMarkdown": ">If I am correct, the purpose of classification models is to determine whether the predicted mask is empty or not, and they are only applied to segmentation models trained on all data. However, what is the role of the \"top 100 logits mean per image\"?\n\n\n\"top 100 logits mean per image\" is a variation of the classification model using the segmentation model.\nA similar approach was used in a previous competition, and again the top 100 logits mean was more accurate for binary classification than the classification model.\n\n>Could you elaborate a bit more about the steps to find that?\n\nIn binary classification, we have both submissions with a percentile threshold and sub with the percentile converted to a fixed value and used as the threshold. Comparing the execution times of the two submissions, we can see that the percentage of empty mask images differs when there is a change. (since the size of the test set and the validation set are the same)",
          "votes": 2
        }
      ]
    },
    {
      "id": 2386573,
      "postDate": "2023-08-12T03:51:09.497Z",
      "content": "<p>That is very interesting. Thank you for sharing your approach!</p>",
      "rawMarkdown": "That is very interesting. Thank you for sharing your approach!\n"
    },
    {
      "id": 2383895,
      "postDate": "2023-08-10T16:11:51.447Z",
      "content": "<p>Congratulations!<br>\nYour pipeline is very attractive, do you plan to publish your code in the future?</p>",
      "rawMarkdown": "Congratulations!\nYour pipeline is very attractive, do you plan to publish your code in the future?",
      "replies": [
        {
          "id": 2384203,
          "postDate": "2023-08-10T22:29:12.110Z",
          "content": "<p>Yes, we plan to release the code for both learning and inference.<br>\nWe are in the process of organizing it, so please be patient.</p>",
          "rawMarkdown": "Yes, we plan to release the code for both learning and inference.\nWe are in the process of organizing it, so please be patient.",
          "votes": 2,
          "replies": [
            {
              "id": 2384232,
              "postDate": "2023-08-10T23:22:42.747Z",
              "content": "<p>I read your Vesuvius Challenge solution and tried to do something similar but what I got was something like score 0.5 ¯_(ツ)_/¯. Looking forward to seeing your code.<br>\nCongratulations to the gold shake up!</p>",
              "rawMarkdown": "I read your Vesuvius Challenge solution and tried to do something similar but what I got was something like score 0.5 ¯\\_(ツ)_/¯. Looking forward to seeing your code.\nCongratulations to the gold shake up!"
            },
            {
              "id": 2384241,
              "postDate": "2023-08-10T23:34:42.447Z",
              "content": "<p>Thank you and congratulations to you too on your solo win. I will read your solution carefully.</p>\n<p>This may be due to the limited number of suitable backbones for training in my experiment. Our 2.5D UNet with resnetrs101 resulted in a gain of 0.01 compared to the 2D UNet.</p>",
              "rawMarkdown": "Thank you and congratulations to you too on your solo win. I will read your solution carefully.\n\nThis may be due to the limited number of suitable backbones for training in my experiment. Our 2.5D UNet with resnetrs101 resulted in a gain of 0.01 compared to the 2D UNet.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2383581,
      "postDate": "2023-08-10T13:03:12.347Z",
      "content": "<p>Did you use 2d models for classification as well? What was the performance (accuracy) of pure classifications models?</p>",
      "rawMarkdown": "Did you use 2d models for classification as well? What was the performance (accuracy) of pure classifications models?",
      "replies": [
        {
          "id": 2383595,
          "postDate": "2023-08-10T13:09:26.180Z",
          "content": "<p>We did not add a 2D model to the binary classification model; use it as a variation of the segmentation model. The classification accuracy is about 0.95.</p>",
          "rawMarkdown": "We did not add a 2D model to the binary classification model; use it as a variation of the segmentation model. The classification accuracy is about 0.95.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2383469,
      "postDate": "2023-08-10T11:29:20.767Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/tattaka\" target=\"_blank\">@tattaka</a> </p>",
      "rawMarkdown": "Congrats @tattaka "
    }
  ],
  "comments": [
    {
      "id": 2386414,
      "author_name": "tattaka",
      "author_url": "",
      "post_date": "2023-08-11T23:28:54.980000",
      "content": "<p><strong>Update</strong><br>\nWe have released the training code(2.5D model part) and the inference notebook for the best submission.</p>\n<ul>\n<li>training code (2.5D model part): <a href=\"https://github.com/tattaka/google-research-identify-contrails-reduce-global-warming\" target=\"_blank\">https://github.com/tattaka/google-research-identify-contrails-reduce-global-warming</a></li>\n<li>inference code: <a href=\"https://www.kaggle.com/code/tattaka/contrail-submission-ensemble?scriptVersionId=139455030\" target=\"_blank\">https://www.kaggle.com/code/tattaka/contrail-submission-ensemble?scriptVersionId=139455030</a></li>\n</ul>",
      "votes": 5,
      "replies": [
        {
          "id": 2394546,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-08-17T02:19:20.093000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2386427,
      "author_name": "Tawara",
      "author_url": "",
      "post_date": "2023-08-11T23:56:03.057000",
      "content": "<p>Sorry for late.</p>\n<p>We have released  the training code for 2D model part:  <br>\n<a href=\"https://github.com/tawatawara/kaggle-google-research-identify-contrails-reduce-global-warming\" target=\"_blank\">https://github.com/tawatawara/kaggle-google-research-identify-contrails-reduce-global-warming</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2393427,
      "author_name": "delai50",
      "author_url": "",
      "post_date": "2023-08-16T10:03:16.487000",
      "content": "<p>Congratulations and thanks a lot for sharing!</p>\n<p>Let me ask you a couple of questions 😃:</p>\n<ul>\n<li>If I am correct, the purpose of classification models is to determine whether the predicted mask is empty or not, and they are only applied to segmentation models trained on all data. However, what is the role of the \"top 100 logits mean per image\"?</li>\n</ul>\n<blockquote>\n  <p>From the time taken to inference, the percentage of empty masks in the entire test was found to be about the same as in validation set. (about 30%)</p>\n</blockquote>\n<ul>\n<li>Could you elaborate a bit more about the steps to find that?</li>\n</ul>",
      "votes": 0,
      "replies": [
        {
          "id": 2395102,
          "author_name": "tattaka",
          "author_url": "",
          "post_date": "2023-08-17T09:34:31.320000",
          "content": "<blockquote>\n  <p>If I am correct, the purpose of classification models is to determine whether the predicted mask is empty or not, and they are only applied to segmentation models trained on all data. However, what is the role of the \"top 100 logits mean per image\"?</p>\n</blockquote>\n<p>\"top 100 logits mean per image\" is a variation of the classification model using the segmentation model.<br>\nA similar approach was used in a previous competition, and again the top 100 logits mean was more accurate for binary classification than the classification model.</p>\n<blockquote>\n  <p>Could you elaborate a bit more about the steps to find that?</p>\n</blockquote>\n<p>In binary classification, we have both submissions with a percentile threshold and sub with the percentile converted to a fixed value and used as the threshold. Comparing the execution times of the two submissions, we can see that the percentage of empty mask images differs when there is a change. (since the size of the test set and the validation set are the same)</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2386573,
      "author_name": "NA",
      "author_url": "",
      "post_date": "2023-08-12T03:51:09.497000",
      "content": "<p>That is very interesting. Thank you for sharing your approach!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2383895,
      "author_name": "Never$",
      "author_url": "",
      "post_date": "2023-08-10T16:11:51.447000",
      "content": "<p>Congratulations!<br>\nYour pipeline is very attractive, do you plan to publish your code in the future?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2384203,
          "author_name": "tattaka",
          "author_url": "",
          "post_date": "2023-08-10T22:29:12.110000",
          "content": "<p>Yes, we plan to release the code for both learning and inference.<br>\nWe are in the process of organizing it, so please be patient.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2384232,
              "author_name": "🐢 Jun Koda",
              "author_url": "",
              "post_date": "2023-08-10T23:22:42.747000",
              "content": "<p>I read your Vesuvius Challenge solution and tried to do something similar but what I got was something like score 0.5 ¯_(ツ)_/¯. Looking forward to seeing your code.<br>\nCongratulations to the gold shake up!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2384241,
              "author_name": "tattaka",
              "author_url": "",
              "post_date": "2023-08-10T23:34:42.447000",
              "content": "<p>Thank you and congratulations to you too on your solo win. I will read your solution carefully.</p>\n<p>This may be due to the limited number of suitable backbones for training in my experiment. Our 2.5D UNet with resnetrs101 resulted in a gain of 0.01 compared to the 2D UNet.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2383581,
      "author_name": "Man of the year",
      "author_url": "",
      "post_date": "2023-08-10T13:03:12.347000",
      "content": "<p>Did you use 2d models for classification as well? What was the performance (accuracy) of pure classifications models?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2383595,
          "author_name": "tattaka",
          "author_url": "",
          "post_date": "2023-08-10T13:09:26.180000",
          "content": "<p>We did not add a 2D model to the binary classification model; use it as a variation of the segmentation model. The classification accuracy is about 0.95.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2383469,
      "author_name": "Vijay Joshi",
      "author_url": "",
      "post_date": "2023-08-10T11:29:20.767000",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/tattaka\" target=\"_blank\">@tattaka</a> </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2383422": "First of all, we would like to thank the organizers of the competition and the Kaggle team for the competition hosting and all the participants who have shared their knowledge so generously. \n\n# Summary\n* 2 stage pipeline: classification and segmentation\n* ensemble of 2.5D model using 5 or 7 frame and 11ch input 2D model\n* soft label using individual label\n* percentile threshold\n\n# Pipeline\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1745801%2Fdeb33f0f5a0349334ed541ec39c2e353%2F2023-08-10%2016.57.28.png?generation=1691664174212060&alt=media)\n\nThis pipeline has been used frequently in past semantic segmentation competitions.    \nThe advantage of this pipeline is that it creates two models, one trained on pos_only and the other on all data, thereby increasing the ensemble effect.    \nIt also saves inference time by allowing more models to be assigned to images for which screening has determined that a mask is present.    \n\n# Models Detail\n## Common Settings\nBoth 2.5D and 2D models were trained for both hard and soft labels.  \nThe soft label is created as follows\n``` python\nlabel = self.np_load(\n    os.path.join(\"mask_individual.npy\")\n).astype(np.float32)\nlabel = np.clip((label * 2).sum(-1) / label.shape[-1], 0, 1)\n```\n\nAll models were trained by all the data in train folder and validated by the one in validation folder. \n\n## 2.5D Model\nWe used a modified [2.5D model used in Vesuvius Challenge](https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/417255).\nThe main changes are as follows\n* Apply 3D CNN to feature maps of all resolutions\n* 3D For the output of the 3D CNN, feature maps corresponding to the target frame are extracted and given to UNet  \n* In the binary classification model, the 3DCNN is applied to the smallest resolution feature map and input to the FC layer in the same way.    \n\nBackbones used were resnetrs101, resnest101e, swin_base_patch4_window12, swinv2_base_window16, convnext_base, convnext_large, etc.  \nThe input frame is set to 5 or 7 frames, with the target frame in the middle.    \nInput resolution of 384 or 512 are used for the model.  \nIt was trained using `BCE+Dice*0.25` for the pos only model and `BCE` for the all data model.  \nAlso, different augmentation is used for each of the classification model, all data seg model, and pos only seg model. (possibly not optimized and therefore not appropriate).    \n\n``` python\n# classification\nalbu.Flip(p=0.5),\nalbu.RandomRotate90(p=0.5),\nalbu.ShiftScaleRotate(p=0.5, scale_limit=0.3, shift_limit=0.1),\nalbu.GridDistortion(num_steps=5, distort_limit=0.3, p=0.5),\nalbu.CoarseDropout(max_height=16, max_width=16, fill_value=255, mask_fill_value=0, p=0.5,)\n\n# all data segmentation\nalbu.ShiftScaleRotate(p=0.5, scale_limit=0.3, shift_limit=0.1),\n\n# pos only segmentation\nalbu.Flip(p=0.25),\nalbu.RandomRotate90(p=0.25),\nalbu.ShiftScaleRotate(p=0.5, scale_limit=0.3, shift_limit=0.1, rotate_limit=0),\nalbu.Rotate(limit=45, p=0.25),\n```\n* EMA(decay=0.998)\n\n## 2D Model\nFor 2D models, we simply used Unet implemented in Segmentation Models Pytorch. The backbones used were res2net50d, regnetz_d8, regnetz_d32, and regnetz_e8.\n\nCharacteristically, the 2D models were trained on 11 channels, i.e. band15 - band14, band14 - band11 and all 9 bands. Learning with this input does not perform well in short epochs (e.g. 20~30 epochs), but in long epochs (specifically 100 epochs), it obtained better validation scores than with false color input.\n\nThe 11ch models also had one curious point. They had poor public score but good private score. I believe this is what pushed us to the prize zone.\n\nOther training settings are as follows(shared by all data seg model and pos only seg one):\n* Input\n    * size: 384x384\n    * each channels are standardized by subtracting the global mean and dividing by the global variance of the channel\n* Loss: `0.9*BCE + 0.1*Dice`\n* Data Augmentation\n``` python\nalbu.HorizontalFlip(p=0.5),\nalbu.VerticalFlip(p=0.5),\nalbu.ShiftScaleRotate(p=0.5, rotate_limit=90)\nalbu.RandomResizedCrop(p: 1.0, scale=[0.875, 1.0], height=256, width=256)\n```\n* EMA(decay=0.9999)\n\n# Postprocessing\nFrom the time taken to inference, the percentage of empty masks in the entire test was found to be about the same as in validation set. (about 30%).  \nTherefore, we used validation set to optimize the percentile threshold for both the classification part and segmentation part.  \nAs a result, the percentile threshold is better than the optimized fixed threshold. \n \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1745801%2F48a3c5b9e9ed108faaed1514dac63485%2F2023-08-10%2017.20.42.png?generation=1691664304846451&alt=media)\n\n# Not Works\n* pseudo label・mean teacher\n* 2D pretrain -> 2.5D finetuning\n  * unstable\n* more augmentations\n* mixup・cutmix・label smoothing\n* TTA\n* efficientnet\n\n# Training Code and Inference Notebook\nWe have released the training code and the inference notebook for the best submission.\n\n* training code (2.5D model part): https://github.com/tattaka/google-research-identify-contrails-reduce-global-warming\n* training code (2D model part): https://github.com/tawatawara/kaggle-google-research-identify-contrails-reduce-global-warming\n* inference code: https://www.kaggle.com/code/tattaka/contrail-submission-ensemble?scriptVersionId=139455030",
    "2386414": "**Update**\nWe have released the training code(2.5D model part) and the inference notebook for the best submission.\n\n* training code (2.5D model part): https://github.com/tattaka/google-research-identify-contrails-reduce-global-warming\n* inference code: https://www.kaggle.com/code/tattaka/contrail-submission-ensemble?scriptVersionId=139455030",
    "2386427": "Sorry for late.\n\nWe have released  the training code for 2D model part:  \nhttps://github.com/tawatawara/kaggle-google-research-identify-contrails-reduce-global-warming",
    "2393427": "Congratulations and thanks a lot for sharing!\n\nLet me ask you a couple of questions 😃:\n\n- If I am correct, the purpose of classification models is to determine whether the predicted mask is empty or not, and they are only applied to segmentation models trained on all data. However, what is the role of the \"top 100 logits mean per image\"?\n\n>From the time taken to inference, the percentage of empty masks in the entire test was found to be about the same as in validation set. (about 30%)\n\n- Could you elaborate a bit more about the steps to find that?\n",
    "2386573": "That is very interesting. Thank you for sharing your approach!\n",
    "2383895": "Congratulations!\nYour pipeline is very attractive, do you plan to publish your code in the future?",
    "2383581": "Did you use 2d models for classification as well? What was the performance (accuracy) of pure classifications models?",
    "2383469": "Congrats @tattaka "
  }
}