{
  "id": 430794,
  "title": "25th Place Solution for the Google Research Identify Contrails to Reduce Global Warming Competition",
  "url": "/competitions/google-research-identify-contrails-reduce-global-warming/discussion/430794",
  "author_name": "Bilzard",
  "post_date": "2023-08-11T03:08:11.738000",
  "votes": 29,
  "comment_count": 10,
  "views": 0,
  "content": "<h2>Context</h2>\n<p>Business context: <a href=\"https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/overview\" target=\"_blank\">https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/overview</a><br>\nData context: <a href=\"https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/data\" target=\"_blank\">https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/data</a></p>\n<h2>Overview of the Approach</h2>\n<p>My approach is simplified version of the 3D model found in the author's paper [1].</p>\n<ul>\n<li>extracting feature maps of each time frames in 2D backbone</li>\n<li>modulating feature map from neighboring time frame's feature map, and obtains modulated 2D feature maps (<strong>Temporal Feature Modulator</strong>)</li>\n<li>passed 2D feature maps to 2D decoder (U-Net)</li>\n</ul>\n<h2>Details of the submission</h2>\n<h3>Temporal Feature Modulator</h3>\n<p>First, I hypothesized that features on earlier layers are not contributed much to the final prediction, since contrails are drifted from frames to frames. So I come up with the idea of only applying temporal modulation in the later layers of feature maps.</p>\n<p>I tried some experiments with applying temporal modulation to earlier layers, and found only applying last layer gives the best performance. Note that the rest of the feature maps are just sliced by the current time frame (T=4).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F99cce57c10157369fcc5c263ebfff4be%2Fcontrail.drawio.png?generation=1691723146332719&amp;alt=media\" alt=\"\"></p>\n<h3>Data processing &amp; Augmentations</h3>\n<p>In order for faster loading and saving disk space, I quantized false-color image to <code>uint8</code>.</p>\n<p>I did it based on the assumption that 'if humans label at <code>uint8</code> resolution, then training an AI model at that resolution would be sufficient'.</p>\n<p>As other participants are already discussed in publicly, any kind of geometric augmentations fails to improve models' performance (after competition, it turns ouf to be because of label misalignment though).</p>\n<p>So I just applied light augmentations of <code>RandomResizedCrop</code> and <code>HorizontalFlip</code>.</p>\n<pre><code>cfg.geometric_transform = A.Compose(\n    [\n        A.RandomResizedCrop(height=, width=, scale=(, ), p=),\n        A.HorizontalFlip(p=),\n    ]\n)\n</code></pre>\n<h3>CV strategy</h3>\n<p>My CV strategy is very naive. I only trained by images on <code>training</code> frames and evaluated on <code>validation</code> images.</p>\n<p>In the final submission, it might have been advisable to use all the images for training or to take the average across folds, but considering the remaining time until the competition's end, it was not realistic, so I gave up.</p>\n<h3>Ensemble &amp; TTA</h3>\n<p>The ensemble policy is:</p>\n<ul>\n<li>TTA: <code>hflip</code> (x2)</li>\n<li>seed ensemble: x4</li>\n<li>different backbones and input resolutions: x3</li>\n</ul>\n<p>I tested several backbones including recently published ones, and find EfficientNet and PVTv2 as ensemble seeds have good training time &amp; performance tradeoff.</p>\n<p>The final submission was composed of 24 models in total.</p>\n<table>\n<thead>\n<tr>\n<th>backbone</th>\n<th>used time frames</th>\n<th>input resolution</th>\n<th>TTA</th>\n<th>random seed</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>efficientnet_b3</td>\n<td>[1, 2, 3, 4]</td>\n<td>768</td>\n<td>hflip (x2)</td>\n<td>x4</td>\n</tr>\n<tr>\n<td>pvt_v2_b3</td>\n<td>[1, 2, 3, 4]</td>\n<td>640</td>\n<td>hflip (x2)</td>\n<td>x4</td>\n</tr>\n<tr>\n<td>pvt_v2_b5</td>\n<td>[1, 2, 3, 4]</td>\n<td>512</td>\n<td>hflip (x2)</td>\n<td>x4</td>\n</tr>\n</tbody>\n</table>\n<p>The CV score on validation set is <strong>0.692</strong>, and private LB score is <strong>0.700</strong>.</p>\n<h3>Other findings to be noted</h3>\n<h4>Technical tips of training with larger resolutions</h4>\n<p>In my moderate machine environment (RTX 3090 Ti x1), Training on higher resolution tend to get CUDA memory overflow error, or result in too small batch sizes. So I used gradient check-pointing and FP16 training to overcome this issue.</p>\n<h4>Geometric Distribution</h4>\n<p>Using metadata, I found the geometric distribution of contrails are very different between training and validation sets.</p>\n<p>As the authors state in paper [1], I believe there are geographical observation points that is only appeared in the training data, and not included in the validation. I also implemented a fold split based on geographical distribution, but I was unable to effectively utilize this information.</p>\n<blockquote>\n  <p>To further boost the number of positives in the dataset, we also included some GOES-16 ABI imagery at locations in the US where Google Street View images of the sky contained contrails.</p>\n</blockquote>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fca513333ca52eca20a62c9ed0a5cc082%2Fgeometric_distributions.png?generation=1691723168388936&amp;alt=media\" alt=\"\"></p>\n<h3>Things that did not worked</h3>\n<ul>\n<li>PL</li>\n<li>randomly chose individual labels</li>\n<li>DeepLabv3 as a decoder</li>\n</ul>\n<h2>Acknowledgements</h2>\n<p>Thank you for holding this competition. In addition to tackling an interesting challenge, I was able to review and recap all the methods I used in previous CV competitions.</p>\n<h2>Sources</h2>\n<ul>\n<li>[1] <a href=\"https://arxiv.org/abs/2304.02122\" target=\"_blank\">Joe et.al., OpenContrails: Benchmarking Contrail Detection on GOES-16 ABI, 2023</a></li>\n<li>[2] <a href=\"https://github.com/opencv/opencv/issues/11784\" target=\"_blank\">warpAffine: correct coordinate system, documentation and incorrect usage</a></li>\n</ul>\n<h2>Appendix</h2>\n<h3>A. Effect of misalignment correction and soft labels</h3>\n<p>In late submission, I tested how misalignment of labels affect the score.<br>\nI also tested the effect of soft labels as some participants said effective.</p>\n<h4>Baseline model</h4>\n<p>To reduce training cost, I simply tested on the basic 2D U-Net architecture where only current frame is used, and no temporal feature modulator.<br>\nOther setups are like these:</p>\n<table>\n<thead>\n<tr>\n<th>Parameter</th>\n<th>Value</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>backend</td>\n<td>efficientnet_b3</td>\n</tr>\n<tr>\n<td>crop_size</td>\n<td>512x512</td>\n</tr>\n<tr>\n<td>loss</td>\n<td>BCE(pos_weight=8) + Dice</td>\n</tr>\n<tr>\n<td>augmentation</td>\n<td>ShiftScaleRotate(p=0.6), hflip</td>\n</tr>\n<tr>\n<td>ensemble</td>\n<td>4x seed ensemble</td>\n</tr>\n<tr>\n<td>TTA</td>\n<td>hflip</td>\n</tr>\n</tbody>\n</table>\n<h4>Misalignment correction</h4>\n<p>Since label was misaligned by (0.5, 0.5) pixels, I upscale the input image by 2x, and shifted (1, 1) pixels. It is simpler than correcting labels on training time and inverse transform them on inference time.<br>\nCorrection was implemented by the below algorithm.<br>\nNote that additional shift coefficient (0.5, 0.5) is added to transformation matrix as a workaround of the open issue of OpenCV's <code>warpAffine</code> function[2].</p>\n<pre><code>M = np.array([[, ,  + ], [, ,  + ]], dtype=np.float32)\nimg = cv2.warpAffine(\n    img, M, ( * W,  * H), flags=cv2.INTER_LINEAR, borderMode=cv2.BORDER_CONSTANT\n)\n</code></pre>\n<h4>Soft Label</h4>\n<p>Soft labels are generated from <code>human_individual_masks.npy</code>.<br>\nSince the data tab on the competition page saids:</p>\n<blockquote>\n  <p>Pixels were considered a contrail when &gt;50% of the labelers annotated it as such. Individual annotations (<code>human_individual_masks.npy</code>) as well as the aggregated ground truth annotations (<code>human_pixel_masks.npy</code>) are included in the training data.</p>\n</blockquote>\n<p>I created the soft labels in the below processing:</p>\n<ul>\n<li>take the mean of individual annotations</li>\n<li>multiplied by 2</li>\n<li>clip by (0, 1)</li>\n</ul>\n<pre><code>human_pixel_mask = np.clip(individual_pixel_mask.mean(-) * , , )\n</code></pre>\n<h4>Result</h4>\n<p>I trained three types of models and tested on validation data (CV) and leader board(Public and Private):</p>\n<ol>\n<li>baseline</li>\n<li>baseline with misalignment correction (MC)</li>\n<li>baseline with MC and soft labels (SL)</li>\n</ol>\n<p>The result shows both MC and SL has independent positive gains by 0.59% and 0.70% respectively on the private LBs.</p>\n<table>\n<thead>\n<tr>\n<th>description</th>\n<th>ensemble seeds</th>\n<th>crop_size</th>\n<th>CV</th>\n<th>Public</th>\n<th>Private</th>\n<th>gain</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>baseline</td>\n<td>4x</td>\n<td>512</td>\n<td>0.6672</td>\n<td>0.67818</td>\n<td>0.67669</td>\n<td>-</td>\n</tr>\n<tr>\n<td>+MC</td>\n<td>4x</td>\n<td>512</td>\n<td>0.6708</td>\n<td>0.68131</td>\n<td>0.68259</td>\n<td>0.59%</td>\n</tr>\n<tr>\n<td>+MC +SL</td>\n<td>4x</td>\n<td>512</td>\n<td>0.6766</td>\n<td>0.69623</td>\n<td>0.68956</td>\n<td>1.29%</td>\n</tr>\n</tbody>\n</table>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F85d08acac7f886c3270dd1da47992811%2FEffect%20of%20misalignment%20correction%20%20soft%20labels.png?generation=1691897900692345&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 2384662,
      "postDate": "2023-08-11T03:08:11.737Z",
      "content": "<h2>Context</h2>\n<p>Business context: <a href=\"https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/overview\" target=\"_blank\">https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/overview</a><br>\nData context: <a href=\"https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/data\" target=\"_blank\">https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/data</a></p>\n<h2>Overview of the Approach</h2>\n<p>My approach is simplified version of the 3D model found in the author's paper [1].</p>\n<ul>\n<li>extracting feature maps of each time frames in 2D backbone</li>\n<li>modulating feature map from neighboring time frame's feature map, and obtains modulated 2D feature maps (<strong>Temporal Feature Modulator</strong>)</li>\n<li>passed 2D feature maps to 2D decoder (U-Net)</li>\n</ul>\n<h2>Details of the submission</h2>\n<h3>Temporal Feature Modulator</h3>\n<p>First, I hypothesized that features on earlier layers are not contributed much to the final prediction, since contrails are drifted from frames to frames. So I come up with the idea of only applying temporal modulation in the later layers of feature maps.</p>\n<p>I tried some experiments with applying temporal modulation to earlier layers, and found only applying last layer gives the best performance. Note that the rest of the feature maps are just sliced by the current time frame (T=4).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F99cce57c10157369fcc5c263ebfff4be%2Fcontrail.drawio.png?generation=1691723146332719&amp;alt=media\" alt=\"\"></p>\n<h3>Data processing &amp; Augmentations</h3>\n<p>In order for faster loading and saving disk space, I quantized false-color image to <code>uint8</code>.</p>\n<p>I did it based on the assumption that 'if humans label at <code>uint8</code> resolution, then training an AI model at that resolution would be sufficient'.</p>\n<p>As other participants are already discussed in publicly, any kind of geometric augmentations fails to improve models' performance (after competition, it turns ouf to be because of label misalignment though).</p>\n<p>So I just applied light augmentations of <code>RandomResizedCrop</code> and <code>HorizontalFlip</code>.</p>\n<pre><code>cfg.geometric_transform = A.Compose(\n    [\n        A.RandomResizedCrop(height=, width=, scale=(, ), p=),\n        A.HorizontalFlip(p=),\n    ]\n)\n</code></pre>\n<h3>CV strategy</h3>\n<p>My CV strategy is very naive. I only trained by images on <code>training</code> frames and evaluated on <code>validation</code> images.</p>\n<p>In the final submission, it might have been advisable to use all the images for training or to take the average across folds, but considering the remaining time until the competition's end, it was not realistic, so I gave up.</p>\n<h3>Ensemble &amp; TTA</h3>\n<p>The ensemble policy is:</p>\n<ul>\n<li>TTA: <code>hflip</code> (x2)</li>\n<li>seed ensemble: x4</li>\n<li>different backbones and input resolutions: x3</li>\n</ul>\n<p>I tested several backbones including recently published ones, and find EfficientNet and PVTv2 as ensemble seeds have good training time &amp; performance tradeoff.</p>\n<p>The final submission was composed of 24 models in total.</p>\n<table>\n<thead>\n<tr>\n<th>backbone</th>\n<th>used time frames</th>\n<th>input resolution</th>\n<th>TTA</th>\n<th>random seed</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>efficientnet_b3</td>\n<td>[1, 2, 3, 4]</td>\n<td>768</td>\n<td>hflip (x2)</td>\n<td>x4</td>\n</tr>\n<tr>\n<td>pvt_v2_b3</td>\n<td>[1, 2, 3, 4]</td>\n<td>640</td>\n<td>hflip (x2)</td>\n<td>x4</td>\n</tr>\n<tr>\n<td>pvt_v2_b5</td>\n<td>[1, 2, 3, 4]</td>\n<td>512</td>\n<td>hflip (x2)</td>\n<td>x4</td>\n</tr>\n</tbody>\n</table>\n<p>The CV score on validation set is <strong>0.692</strong>, and private LB score is <strong>0.700</strong>.</p>\n<h3>Other findings to be noted</h3>\n<h4>Technical tips of training with larger resolutions</h4>\n<p>In my moderate machine environment (RTX 3090 Ti x1), Training on higher resolution tend to get CUDA memory overflow error, or result in too small batch sizes. So I used gradient check-pointing and FP16 training to overcome this issue.</p>\n<h4>Geometric Distribution</h4>\n<p>Using metadata, I found the geometric distribution of contrails are very different between training and validation sets.</p>\n<p>As the authors state in paper [1], I believe there are geographical observation points that is only appeared in the training data, and not included in the validation. I also implemented a fold split based on geographical distribution, but I was unable to effectively utilize this information.</p>\n<blockquote>\n  <p>To further boost the number of positives in the dataset, we also included some GOES-16 ABI imagery at locations in the US where Google Street View images of the sky contained contrails.</p>\n</blockquote>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fca513333ca52eca20a62c9ed0a5cc082%2Fgeometric_distributions.png?generation=1691723168388936&amp;alt=media\" alt=\"\"></p>\n<h3>Things that did not worked</h3>\n<ul>\n<li>PL</li>\n<li>randomly chose individual labels</li>\n<li>DeepLabv3 as a decoder</li>\n</ul>\n<h2>Acknowledgements</h2>\n<p>Thank you for holding this competition. In addition to tackling an interesting challenge, I was able to review and recap all the methods I used in previous CV competitions.</p>\n<h2>Sources</h2>\n<ul>\n<li>[1] <a href=\"https://arxiv.org/abs/2304.02122\" target=\"_blank\">Joe et.al., OpenContrails: Benchmarking Contrail Detection on GOES-16 ABI, 2023</a></li>\n<li>[2] <a href=\"https://github.com/opencv/opencv/issues/11784\" target=\"_blank\">warpAffine: correct coordinate system, documentation and incorrect usage</a></li>\n</ul>\n<h2>Appendix</h2>\n<h3>A. Effect of misalignment correction and soft labels</h3>\n<p>In late submission, I tested how misalignment of labels affect the score.<br>\nI also tested the effect of soft labels as some participants said effective.</p>\n<h4>Baseline model</h4>\n<p>To reduce training cost, I simply tested on the basic 2D U-Net architecture where only current frame is used, and no temporal feature modulator.<br>\nOther setups are like these:</p>\n<table>\n<thead>\n<tr>\n<th>Parameter</th>\n<th>Value</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>backend</td>\n<td>efficientnet_b3</td>\n</tr>\n<tr>\n<td>crop_size</td>\n<td>512x512</td>\n</tr>\n<tr>\n<td>loss</td>\n<td>BCE(pos_weight=8) + Dice</td>\n</tr>\n<tr>\n<td>augmentation</td>\n<td>ShiftScaleRotate(p=0.6), hflip</td>\n</tr>\n<tr>\n<td>ensemble</td>\n<td>4x seed ensemble</td>\n</tr>\n<tr>\n<td>TTA</td>\n<td>hflip</td>\n</tr>\n</tbody>\n</table>\n<h4>Misalignment correction</h4>\n<p>Since label was misaligned by (0.5, 0.5) pixels, I upscale the input image by 2x, and shifted (1, 1) pixels. It is simpler than correcting labels on training time and inverse transform them on inference time.<br>\nCorrection was implemented by the below algorithm.<br>\nNote that additional shift coefficient (0.5, 0.5) is added to transformation matrix as a workaround of the open issue of OpenCV's <code>warpAffine</code> function[2].</p>\n<pre><code>M = np.array([[, ,  + ], [, ,  + ]], dtype=np.float32)\nimg = cv2.warpAffine(\n    img, M, ( * W,  * H), flags=cv2.INTER_LINEAR, borderMode=cv2.BORDER_CONSTANT\n)\n</code></pre>\n<h4>Soft Label</h4>\n<p>Soft labels are generated from <code>human_individual_masks.npy</code>.<br>\nSince the data tab on the competition page saids:</p>\n<blockquote>\n  <p>Pixels were considered a contrail when &gt;50% of the labelers annotated it as such. Individual annotations (<code>human_individual_masks.npy</code>) as well as the aggregated ground truth annotations (<code>human_pixel_masks.npy</code>) are included in the training data.</p>\n</blockquote>\n<p>I created the soft labels in the below processing:</p>\n<ul>\n<li>take the mean of individual annotations</li>\n<li>multiplied by 2</li>\n<li>clip by (0, 1)</li>\n</ul>\n<pre><code>human_pixel_mask = np.clip(individual_pixel_mask.mean(-) * , , )\n</code></pre>\n<h4>Result</h4>\n<p>I trained three types of models and tested on validation data (CV) and leader board(Public and Private):</p>\n<ol>\n<li>baseline</li>\n<li>baseline with misalignment correction (MC)</li>\n<li>baseline with MC and soft labels (SL)</li>\n</ol>\n<p>The result shows both MC and SL has independent positive gains by 0.59% and 0.70% respectively on the private LBs.</p>\n<table>\n<thead>\n<tr>\n<th>description</th>\n<th>ensemble seeds</th>\n<th>crop_size</th>\n<th>CV</th>\n<th>Public</th>\n<th>Private</th>\n<th>gain</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>baseline</td>\n<td>4x</td>\n<td>512</td>\n<td>0.6672</td>\n<td>0.67818</td>\n<td>0.67669</td>\n<td>-</td>\n</tr>\n<tr>\n<td>+MC</td>\n<td>4x</td>\n<td>512</td>\n<td>0.6708</td>\n<td>0.68131</td>\n<td>0.68259</td>\n<td>0.59%</td>\n</tr>\n<tr>\n<td>+MC +SL</td>\n<td>4x</td>\n<td>512</td>\n<td>0.6766</td>\n<td>0.69623</td>\n<td>0.68956</td>\n<td>1.29%</td>\n</tr>\n</tbody>\n</table>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F85d08acac7f886c3270dd1da47992811%2FEffect%20of%20misalignment%20correction%20%20soft%20labels.png?generation=1691897900692345&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "## Context\n\nBusiness context: https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/overview\nData context: https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/data\n\n## Overview of the Approach\n\nMy approach is simplified version of the 3D model found in the author's paper [1].\n\n* extracting feature maps of each time frames in 2D backbone\n* modulating feature map from neighboring time frame's feature map, and obtains modulated 2D feature maps (**Temporal Feature Modulator**)\n* passed 2D feature maps to 2D decoder (U-Net)\n\n## Details of the submission\n\n### Temporal Feature Modulator\n\nFirst, I hypothesized that features on earlier layers are not contributed much to the final prediction, since contrails are drifted from frames to frames. So I come up with the idea of only applying temporal modulation in the later layers of feature maps.\n\nI tried some experiments with applying temporal modulation to earlier layers, and found only applying last layer gives the best performance. Note that the rest of the feature maps are just sliced by the current time frame (T=4).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F99cce57c10157369fcc5c263ebfff4be%2Fcontrail.drawio.png?generation=1691723146332719&alt=media)\n\n### Data processing & Augmentations\n\nIn order for faster loading and saving disk space, I quantized false-color image to `uint8`.\n\nI did it based on the assumption that 'if humans label at `uint8` resolution, then training an AI model at that resolution would be sufficient'.\n\nAs other participants are already discussed in publicly, any kind of geometric augmentations fails to improve models' performance (after competition, it turns ouf to be because of label misalignment though).\n\nSo I just applied light augmentations of `RandomResizedCrop` and `HorizontalFlip`.\n\n```python\ncfg.geometric_transform = A.Compose(\n    [\n        A.RandomResizedCrop(height=256, width=256, scale=(0.75, 1.0), p=0.6),\n        A.HorizontalFlip(p=0.5),\n    ]\n)\n```\n\n### CV strategy\n\nMy CV strategy is very naive. I only trained by images on `training` frames and evaluated on `validation` images.\n\nIn the final submission, it might have been advisable to use all the images for training or to take the average across folds, but considering the remaining time until the competition's end, it was not realistic, so I gave up.\n\n\n### Ensemble & TTA\n\nThe ensemble policy is:\n\n* TTA: `hflip` (x2)\n* seed ensemble: x4\n* different backbones and input resolutions: x3\n\nI tested several backbones including recently published ones, and find EfficientNet and PVTv2 as ensemble seeds have good training time & performance tradeoff.\n\nThe final submission was composed of 24 models in total.\n\n| backbone       | used time frames | input resolution | TTA   | random seed |\n|----------------|------------------|------------------|-------|-------------|\n| efficientnet_b3| [1, 2, 3, 4]     | 768              | hflip (x2) | x4          |\n| pvt_v2_b3      | [1, 2, 3, 4]     | 640              | hflip (x2) | x4          |\n| pvt_v2_b5      | [1, 2, 3, 4]     | 512              | hflip (x2) | x4          |\n\nThe CV score on validation set is **0.692**, and private LB score is **0.700**.\n\n### Other findings to be noted\n\n#### Technical tips of training with larger resolutions\n\nIn my moderate machine environment (RTX 3090 Ti x1), Training on higher resolution tend to get CUDA memory overflow error, or result in too small batch sizes. So I used gradient check-pointing and FP16 training to overcome this issue.\n\n#### Geometric Distribution\n\nUsing metadata, I found the geometric distribution of contrails are very different between training and validation sets.\n\nAs the authors state in paper [1], I believe there are geographical observation points that is only appeared in the training data, and not included in the validation. I also implemented a fold split based on geographical distribution, but I was unable to effectively utilize this information.\n\n> To further boost the number of positives in the dataset, we also included some GOES-16 ABI imagery at locations in the US where Google Street View images of the sky contained contrails.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fca513333ca52eca20a62c9ed0a5cc082%2Fgeometric_distributions.png?generation=1691723168388936&alt=media)\n\n### Things that did not worked\n\n* PL\n* randomly chose individual labels\n* DeepLabv3 as a decoder\n\n\n\n## Acknowledgements\n\nThank you for holding this competition. In addition to tackling an interesting challenge, I was able to review and recap all the methods I used in previous CV competitions.\n\n## Sources\n\n- [1] [Joe et.al., OpenContrails: Benchmarking Contrail Detection on GOES-16 ABI, 2023](https://arxiv.org/abs/2304.02122)\n- [2] [warpAffine: correct coordinate system, documentation and incorrect usage](https://github.com/opencv/opencv/issues/11784)\n\n## Appendix\n\n### A. Effect of misalignment correction and soft labels\n\nIn late submission, I tested how misalignment of labels affect the score.\nI also tested the effect of soft labels as some participants said effective.\n\n#### Baseline model\n\nTo reduce training cost, I simply tested on the basic 2D U-Net architecture where only current frame is used, and no temporal feature modulator.\nOther setups are like these:\n\n| Parameter     | Value                             |\n|---------------|-----------------------------------|\n| backend       | efficientnet_b3                   |\n| crop_size     | 512x512                           |\n| loss          | BCE(pos_weight=8) + Dice          |\n| augmentation  | ShiftScaleRotate(p=0.6), hflip    |\n| ensemble      | 4x seed ensemble                  |\n| TTA           | hflip                             |\n\n#### Misalignment correction\n\nSince label was misaligned by (0.5, 0.5) pixels, I upscale the input image by 2x, and shifted (1, 1) pixels. It is simpler than correcting labels on training time and inverse transform them on inference time.\nCorrection was implemented by the below algorithm.\nNote that additional shift coefficient (0.5, 0.5) is added to transformation matrix as a workaround of the open issue of OpenCV's `warpAffine` function[2].\n\n```python\nM = np.array([[2, 0, 1 + 0.5], [0, 2, 1 + 0.5]], dtype=np.float32)\nimg = cv2.warpAffine(\n    img, M, (2 * W, 2 * H), flags=cv2.INTER_LINEAR, borderMode=cv2.BORDER_CONSTANT\n)\n```\n\n#### Soft Label\n\nSoft labels are generated from `human_individual_masks.npy`.\nSince the data tab on the competition page saids:\n\n> Pixels were considered a contrail when >50% of the labelers annotated it as such. Individual annotations (`human_individual_masks.npy`) as well as the aggregated ground truth annotations (`human_pixel_masks.npy`) are included in the training data.\n\nI created the soft labels in the below processing:\n- take the mean of individual annotations\n- multiplied by 2\n- clip by (0, 1)\n\n```python\nhuman_pixel_mask = np.clip(individual_pixel_mask.mean(-1) * 2, 0, 1)\n```\n\n#### Result\n\nI trained three types of models and tested on validation data (CV) and leader board(Public and Private):\n\n1. baseline\n2. baseline with misalignment correction (MC)\n3. baseline with MC and soft labels (SL)\n\nThe result shows both MC and SL has independent positive gains by 0.59% and 0.70% respectively on the private LBs.\n\n| description                 | ensemble seeds | crop_size | CV     | Public  | Private | gain  |\n|-----------------------------|----------|-----------|--------|---------|---------|-------|\n| baseline                    | 4x        | 512       | 0.6672 | 0.67818 | 0.67669 | -     |\n| +MC          | 4x        | 512       | 0.6708 | 0.68131 | 0.68259 | 0.59% |\n| +MC +SL | 4x      | 512       | 0.6766 | 0.69623 | 0.68956 | 1.29% |\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F85d08acac7f886c3270dd1da47992811%2FEffect%20of%20misalignment%20correction%20%20soft%20labels.png?generation=1691897900692345&alt=media)",
      "votes": 29
    },
    {
      "id": 2385453,
      "postDate": "2023-08-11T10:51:01.403Z",
      "content": "<p>Interesting!</p>\n<p>Have you considered frame t=4 as skip connection to the decoder? </p>",
      "rawMarkdown": "Interesting!\n\nHave you considered frame t=4 as skip connection to the decoder? ",
      "votes": 1,
      "replies": [
        {
          "id": 2385527,
          "postDate": "2023-08-11T11:55:49.463Z",
          "content": "<p><a href=\"https://www.kaggle.com/mayurimk\" target=\"_blank\">@mayurimk</a></p>\n<blockquote>\n  <p>Have you considered frame t=4 as skip connection to the decoder?</p>\n</blockquote>\n<p>Well, it was once in my mind, but never tested. I believe it would stabilize training especially on deeper modulator (in my case, N=3). I will put this on stock of my arsenal. Thank you for pointing this out.</p>",
          "rawMarkdown": "@mayurimk\n\n> Have you considered frame t=4 as skip connection to the decoder?\n\nWell, it was once in my mind, but never tested. I believe it would stabilize training especially on deeper modulator (in my case, N=3). I will put this on stock of my arsenal. Thank you for pointing this out."
        }
      ]
    },
    {
      "id": 2385324,
      "postDate": "2023-08-11T09:13:44.223Z",
      "content": "<p>Many congratulations on the solo sliver and thank you for the great write up!</p>\n<p>Could you please clarify this part:</p>\n<blockquote>\n  <p>Temporal Feature Modulator</p>\n  <p>First, I hypothesized that features on earlier layers are not contributed much to the final prediction, since contrails are drifted from frames to frames. So I come up with the idea of only applying temporal modulation in the later layers of feature maps.</p>\n  <p>I tried some experiments with applying temporal modulation to earlier layers, and found only applying last layer gives the best performance. Note that the rest of the feature maps are just sliced by the current time frame (T=4).</p>\n</blockquote>",
      "rawMarkdown": "Many congratulations on the solo sliver and thank you for the great write up!\n\nCould you please clarify this part:\n>Temporal Feature Modulator\n\n>First, I hypothesized that features on earlier layers are not contributed much to the final prediction, since contrails are drifted from frames to frames. So I come up with the idea of only applying temporal modulation in the later layers of feature maps.\n>\nI tried some experiments with applying temporal modulation to earlier layers, and found only applying last layer gives the best performance. Note that the rest of the feature maps are just sliced by the current time frame (T=4).",
      "votes": 1,
      "replies": [
        {
          "id": 2385347,
          "postDate": "2023-08-11T09:25:20.130Z",
          "content": "<p><a href=\"https://www.kaggle.com/tonymarkchris\" target=\"_blank\">@tonymarkchris</a> <br>\nDoes this helps?</p>\n<pre><code>     ():\n        cfg = self.cfg\n\n        crop_size = cfg.crop_size  self.training  cfg.infer_crop_size\n        x = batch[]\n        B, C, T, H, W = x.shape\n\n        x = rearrange(x, )\n         (H != crop_size)  (W != crop_size):\n            x = F.interpolate(x, size=(crop_size, crop_size), mode=, align_corners=)\n\n        features = self.encoder(x)\n         i  ((features[:-])):\n            x = rearrange(features[i], , t=T)\n            x = x[:, :, self.current_frame_relative, :, :]\n            features[i] = x\n\n        x = features[-]\n        x = rearrange(x, , t=T)\n        x = self.temporal_encoder(x)\n        x = x[:, :, self.current_frame_relative - cfg.temporal_offset, :, :]\n        features[-] = x\n\n        decoder_output = self.decoder(*features)\n        logit = self.segmentation_head(decoder_output)\n\n        _, _, h, w = logit.shape\n         (h != H)  (w != W):\n            logit = F.interpolate(logit, size=(H, W), mode=, align_corners=)\n\n        output = (logit=logit)\n\n         output\n</code></pre>",
          "rawMarkdown": "@tonymarkchris \nDoes this helps?\n\n```python\n    def forward(self, batch):\n        cfg = self.cfg\n\n        crop_size = cfg.crop_size if self.training else cfg.infer_crop_size\n        x = batch[\"image\"]\n        B, C, T, H, W = x.shape\n\n        x = rearrange(x, \"b c t h w -> (b t) c h w\")\n        if (H != crop_size) or (W != crop_size):\n            x = F.interpolate(x, size=(crop_size, crop_size), mode=\"bilinear\", align_corners=False)\n\n        features = self.encoder(x)\n        for i in range(len(features[:-1])):\n            x = rearrange(features[i], \"(b t) c h w -> b c t h w\", t=T)\n            x = x[:, :, self.current_frame_relative, :, :]\n            features[i] = x\n\n        x = features[-1]\n        x = rearrange(x, \"(b t) c h w -> b c t h w\", t=T)\n        x = self.temporal_encoder(x)\n        x = x[:, :, self.current_frame_relative - cfg.temporal_offset, :, :]\n        features[-1] = x\n\n        decoder_output = self.decoder(*features)\n        logit = self.segmentation_head(decoder_output)\n\n        _, _, h, w = logit.shape\n        if (h != H) or (w != W):\n            logit = F.interpolate(logit, size=(H, W), mode=\"bilinear\", align_corners=False)\n\n        output = dict(logit=logit)\n\n        return output\n```",
          "votes": 3,
          "replies": [
            {
              "id": 2385665,
              "postDate": "2023-08-11T13:40:54.817Z",
              "content": "<p>thanks for sharing, can u also share the encoder code snippet?</p>",
              "rawMarkdown": "thanks for sharing, can u also share the encoder code snippet?"
            },
            {
              "id": 2402194,
              "postDate": "2023-08-22T03:49:30.203Z",
              "content": "<p>Well, it's nothing special. It is just as I wrote on the block diagram (3 layers of Conv3D + BN + ReLU). I think you can implement it by yourself.</p>",
              "rawMarkdown": "Well, it's nothing special. It is just as I wrote on the block diagram (3 layers of Conv3D + BN + ReLU). I think you can implement it by yourself."
            }
          ]
        }
      ]
    },
    {
      "id": 2398367,
      "postDate": "2023-08-19T16:17:41.670Z",
      "content": "<p>Congratulations and thanks for sharing such a detailed writeup! <br>\nDo you know why HFlip worked for you even though you didn't correct label misalignment?</p>",
      "rawMarkdown": "Congratulations and thanks for sharing such a detailed writeup! \nDo you know why HFlip worked for you even though you didn't correct label misalignment?",
      "replies": [
        {
          "id": 2398713,
          "postDate": "2023-08-19T21:23:33.553Z",
          "content": "<p>Thanks. Well, actually, I didn’t conducted detailed comparison (which I should have done). However, I believe augmentation in this competition (without misalignment correction) has both positive and negative effect.</p>\n<p>In the negative effect, v/h flips and rotation surely introduces label noise, but after enough steps of training, the model would learn “expected contrail label” which is close to mean of flipped/rotated contrails.</p>\n<p>On the positive effect, the augmentation would inject model with certain amount of generalization as in the general context of machine learning. Additionally, the larger model with time frame information might even have chance to learn flipped/rotated direction, where the model could learn shifted labels correctly.</p>\n<p>What discussed above is just hypothesis, but considering some top solution says they conducted intense augmentation with larger models, I believe my hypothesis is somehow true.</p>",
          "rawMarkdown": "Thanks. Well, actually, I didn’t conducted detailed comparison (which I should have done). However, I believe augmentation in this competition (without misalignment correction) has both positive and negative effect.\n\nIn the negative effect, v/h flips and rotation surely introduces label noise, but after enough steps of training, the model would learn “expected contrail label” which is close to mean of flipped/rotated contrails.\n\nOn the positive effect, the augmentation would inject model with certain amount of generalization as in the general context of machine learning. Additionally, the larger model with time frame information might even have chance to learn flipped/rotated direction, where the model could learn shifted labels correctly.\n\nWhat discussed above is just hypothesis, but considering some top solution says they conducted intense augmentation with larger models, I believe my hypothesis is somehow true."
        }
      ]
    },
    {
      "id": 2388440,
      "postDate": "2023-08-13T11:43:26.357Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> Thank you for sharing your solution and after deadline experiments <br>\nI'm wondering why multiply *2  here<br>\n<code>human_pixel_mask = np.clip(individual_pixel_mask.mean(-1) * 2, 0, 1)</code></p>",
      "rawMarkdown": "Hi @tatamikenn Thank you for sharing your solution and after deadline experiments \nI'm wondering why multiply *2  here\n`human_pixel_mask = np.clip(individual_pixel_mask.mean(-1) * 2, 0, 1)`\n",
      "replies": [
        {
          "id": 2388446,
          "postDate": "2023-08-13T11:51:01.217Z",
          "content": "<p><a href=\"https://www.kaggle.com/RB\" target=\"_blank\">@RB</a><br>\nBecause I want to make the pixel value greater than 0.5 makes 1 (not 0.5). However, I believe it's not big difference to simply taking average, since setting proper threshold should gave us equivalent result.</p>",
          "rawMarkdown": "@RB\nBecause I want to make the pixel value greater than 0.5 makes 1 (not 0.5). However, I believe it's not big difference to simply taking average, since setting proper threshold should gave us equivalent result.",
          "votes": 2
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2385453,
      "author_name": "mayurimk",
      "author_url": "",
      "post_date": "2023-08-11T10:51:01.403000",
      "content": "<p>Interesting!</p>\n<p>Have you considered frame t=4 as skip connection to the decoder? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2385527,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2023-08-11T11:55:49.463000",
          "content": "<p><a href=\"https://www.kaggle.com/mayurimk\" target=\"_blank\">@mayurimk</a></p>\n<blockquote>\n  <p>Have you considered frame t=4 as skip connection to the decoder?</p>\n</blockquote>\n<p>Well, it was once in my mind, but never tested. I believe it would stabilize training especially on deeper modulator (in my case, N=3). I will put this on stock of my arsenal. Thank you for pointing this out.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2385324,
      "author_name": "DJ_Xia",
      "author_url": "",
      "post_date": "2023-08-11T09:13:44.223000",
      "content": "<p>Many congratulations on the solo sliver and thank you for the great write up!</p>\n<p>Could you please clarify this part:</p>\n<blockquote>\n  <p>Temporal Feature Modulator</p>\n  <p>First, I hypothesized that features on earlier layers are not contributed much to the final prediction, since contrails are drifted from frames to frames. So I come up with the idea of only applying temporal modulation in the later layers of feature maps.</p>\n  <p>I tried some experiments with applying temporal modulation to earlier layers, and found only applying last layer gives the best performance. Note that the rest of the feature maps are just sliced by the current time frame (T=4).</p>\n</blockquote>",
      "votes": 1,
      "replies": [
        {
          "id": 2385347,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2023-08-11T09:25:20.130000",
          "content": "<p><a href=\"https://www.kaggle.com/tonymarkchris\" target=\"_blank\">@tonymarkchris</a> <br>\nDoes this helps?</p>\n<pre><code>     ():\n        cfg = self.cfg\n\n        crop_size = cfg.crop_size  self.training  cfg.infer_crop_size\n        x = batch[]\n        B, C, T, H, W = x.shape\n\n        x = rearrange(x, )\n         (H != crop_size)  (W != crop_size):\n            x = F.interpolate(x, size=(crop_size, crop_size), mode=, align_corners=)\n\n        features = self.encoder(x)\n         i  ((features[:-])):\n            x = rearrange(features[i], , t=T)\n            x = x[:, :, self.current_frame_relative, :, :]\n            features[i] = x\n\n        x = features[-]\n        x = rearrange(x, , t=T)\n        x = self.temporal_encoder(x)\n        x = x[:, :, self.current_frame_relative - cfg.temporal_offset, :, :]\n        features[-] = x\n\n        decoder_output = self.decoder(*features)\n        logit = self.segmentation_head(decoder_output)\n\n        _, _, h, w = logit.shape\n         (h != H)  (w != W):\n            logit = F.interpolate(logit, size=(H, W), mode=, align_corners=)\n\n        output = (logit=logit)\n\n         output\n</code></pre>",
          "votes": 3,
          "replies": [
            {
              "id": 2385665,
              "author_name": "DJ_Xia",
              "author_url": "",
              "post_date": "2023-08-11T13:40:54.817000",
              "content": "<p>thanks for sharing, can u also share the encoder code snippet?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2402194,
              "author_name": "Bilzard",
              "author_url": "",
              "post_date": "2023-08-22T03:49:30.203000",
              "content": "<p>Well, it's nothing special. It is just as I wrote on the block diagram (3 layers of Conv3D + BN + ReLU). I think you can implement it by yourself.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2398367,
      "author_name": "delai50",
      "author_url": "",
      "post_date": "2023-08-19T16:17:41.670000",
      "content": "<p>Congratulations and thanks for sharing such a detailed writeup! <br>\nDo you know why HFlip worked for you even though you didn't correct label misalignment?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2398713,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2023-08-19T21:23:33.553000",
          "content": "<p>Thanks. Well, actually, I didn’t conducted detailed comparison (which I should have done). However, I believe augmentation in this competition (without misalignment correction) has both positive and negative effect.</p>\n<p>In the negative effect, v/h flips and rotation surely introduces label noise, but after enough steps of training, the model would learn “expected contrail label” which is close to mean of flipped/rotated contrails.</p>\n<p>On the positive effect, the augmentation would inject model with certain amount of generalization as in the general context of machine learning. Additionally, the larger model with time frame information might even have chance to learn flipped/rotated direction, where the model could learn shifted labels correctly.</p>\n<p>What discussed above is just hypothesis, but considering some top solution says they conducted intense augmentation with larger models, I believe my hypothesis is somehow true.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2388440,
      "author_name": "RB",
      "author_url": "",
      "post_date": "2023-08-13T11:43:26.357000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> Thank you for sharing your solution and after deadline experiments <br>\nI'm wondering why multiply *2  here<br>\n<code>human_pixel_mask = np.clip(individual_pixel_mask.mean(-1) * 2, 0, 1)</code></p>",
      "votes": 0,
      "replies": [
        {
          "id": 2388446,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2023-08-13T11:51:01.217000",
          "content": "<p><a href=\"https://www.kaggle.com/RB\" target=\"_blank\">@RB</a><br>\nBecause I want to make the pixel value greater than 0.5 makes 1 (not 0.5). However, I believe it's not big difference to simply taking average, since setting proper threshold should gave us equivalent result.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2384662": "## Context\n\nBusiness context: https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/overview\nData context: https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/data\n\n## Overview of the Approach\n\nMy approach is simplified version of the 3D model found in the author's paper [1].\n\n* extracting feature maps of each time frames in 2D backbone\n* modulating feature map from neighboring time frame's feature map, and obtains modulated 2D feature maps (**Temporal Feature Modulator**)\n* passed 2D feature maps to 2D decoder (U-Net)\n\n## Details of the submission\n\n### Temporal Feature Modulator\n\nFirst, I hypothesized that features on earlier layers are not contributed much to the final prediction, since contrails are drifted from frames to frames. So I come up with the idea of only applying temporal modulation in the later layers of feature maps.\n\nI tried some experiments with applying temporal modulation to earlier layers, and found only applying last layer gives the best performance. Note that the rest of the feature maps are just sliced by the current time frame (T=4).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F99cce57c10157369fcc5c263ebfff4be%2Fcontrail.drawio.png?generation=1691723146332719&alt=media)\n\n### Data processing & Augmentations\n\nIn order for faster loading and saving disk space, I quantized false-color image to `uint8`.\n\nI did it based on the assumption that 'if humans label at `uint8` resolution, then training an AI model at that resolution would be sufficient'.\n\nAs other participants are already discussed in publicly, any kind of geometric augmentations fails to improve models' performance (after competition, it turns ouf to be because of label misalignment though).\n\nSo I just applied light augmentations of `RandomResizedCrop` and `HorizontalFlip`.\n\n```python\ncfg.geometric_transform = A.Compose(\n    [\n        A.RandomResizedCrop(height=256, width=256, scale=(0.75, 1.0), p=0.6),\n        A.HorizontalFlip(p=0.5),\n    ]\n)\n```\n\n### CV strategy\n\nMy CV strategy is very naive. I only trained by images on `training` frames and evaluated on `validation` images.\n\nIn the final submission, it might have been advisable to use all the images for training or to take the average across folds, but considering the remaining time until the competition's end, it was not realistic, so I gave up.\n\n\n### Ensemble & TTA\n\nThe ensemble policy is:\n\n* TTA: `hflip` (x2)\n* seed ensemble: x4\n* different backbones and input resolutions: x3\n\nI tested several backbones including recently published ones, and find EfficientNet and PVTv2 as ensemble seeds have good training time & performance tradeoff.\n\nThe final submission was composed of 24 models in total.\n\n| backbone       | used time frames | input resolution | TTA   | random seed |\n|----------------|------------------|------------------|-------|-------------|\n| efficientnet_b3| [1, 2, 3, 4]     | 768              | hflip (x2) | x4          |\n| pvt_v2_b3      | [1, 2, 3, 4]     | 640              | hflip (x2) | x4          |\n| pvt_v2_b5      | [1, 2, 3, 4]     | 512              | hflip (x2) | x4          |\n\nThe CV score on validation set is **0.692**, and private LB score is **0.700**.\n\n### Other findings to be noted\n\n#### Technical tips of training with larger resolutions\n\nIn my moderate machine environment (RTX 3090 Ti x1), Training on higher resolution tend to get CUDA memory overflow error, or result in too small batch sizes. So I used gradient check-pointing and FP16 training to overcome this issue.\n\n#### Geometric Distribution\n\nUsing metadata, I found the geometric distribution of contrails are very different between training and validation sets.\n\nAs the authors state in paper [1], I believe there are geographical observation points that is only appeared in the training data, and not included in the validation. I also implemented a fold split based on geographical distribution, but I was unable to effectively utilize this information.\n\n> To further boost the number of positives in the dataset, we also included some GOES-16 ABI imagery at locations in the US where Google Street View images of the sky contained contrails.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fca513333ca52eca20a62c9ed0a5cc082%2Fgeometric_distributions.png?generation=1691723168388936&alt=media)\n\n### Things that did not worked\n\n* PL\n* randomly chose individual labels\n* DeepLabv3 as a decoder\n\n\n\n## Acknowledgements\n\nThank you for holding this competition. In addition to tackling an interesting challenge, I was able to review and recap all the methods I used in previous CV competitions.\n\n## Sources\n\n- [1] [Joe et.al., OpenContrails: Benchmarking Contrail Detection on GOES-16 ABI, 2023](https://arxiv.org/abs/2304.02122)\n- [2] [warpAffine: correct coordinate system, documentation and incorrect usage](https://github.com/opencv/opencv/issues/11784)\n\n## Appendix\n\n### A. Effect of misalignment correction and soft labels\n\nIn late submission, I tested how misalignment of labels affect the score.\nI also tested the effect of soft labels as some participants said effective.\n\n#### Baseline model\n\nTo reduce training cost, I simply tested on the basic 2D U-Net architecture where only current frame is used, and no temporal feature modulator.\nOther setups are like these:\n\n| Parameter     | Value                             |\n|---------------|-----------------------------------|\n| backend       | efficientnet_b3                   |\n| crop_size     | 512x512                           |\n| loss          | BCE(pos_weight=8) + Dice          |\n| augmentation  | ShiftScaleRotate(p=0.6), hflip    |\n| ensemble      | 4x seed ensemble                  |\n| TTA           | hflip                             |\n\n#### Misalignment correction\n\nSince label was misaligned by (0.5, 0.5) pixels, I upscale the input image by 2x, and shifted (1, 1) pixels. It is simpler than correcting labels on training time and inverse transform them on inference time.\nCorrection was implemented by the below algorithm.\nNote that additional shift coefficient (0.5, 0.5) is added to transformation matrix as a workaround of the open issue of OpenCV's `warpAffine` function[2].\n\n```python\nM = np.array([[2, 0, 1 + 0.5], [0, 2, 1 + 0.5]], dtype=np.float32)\nimg = cv2.warpAffine(\n    img, M, (2 * W, 2 * H), flags=cv2.INTER_LINEAR, borderMode=cv2.BORDER_CONSTANT\n)\n```\n\n#### Soft Label\n\nSoft labels are generated from `human_individual_masks.npy`.\nSince the data tab on the competition page saids:\n\n> Pixels were considered a contrail when >50% of the labelers annotated it as such. Individual annotations (`human_individual_masks.npy`) as well as the aggregated ground truth annotations (`human_pixel_masks.npy`) are included in the training data.\n\nI created the soft labels in the below processing:\n- take the mean of individual annotations\n- multiplied by 2\n- clip by (0, 1)\n\n```python\nhuman_pixel_mask = np.clip(individual_pixel_mask.mean(-1) * 2, 0, 1)\n```\n\n#### Result\n\nI trained three types of models and tested on validation data (CV) and leader board(Public and Private):\n\n1. baseline\n2. baseline with misalignment correction (MC)\n3. baseline with MC and soft labels (SL)\n\nThe result shows both MC and SL has independent positive gains by 0.59% and 0.70% respectively on the private LBs.\n\n| description                 | ensemble seeds | crop_size | CV     | Public  | Private | gain  |\n|-----------------------------|----------|-----------|--------|---------|---------|-------|\n| baseline                    | 4x        | 512       | 0.6672 | 0.67818 | 0.67669 | -     |\n| +MC          | 4x        | 512       | 0.6708 | 0.68131 | 0.68259 | 0.59% |\n| +MC +SL | 4x      | 512       | 0.6766 | 0.69623 | 0.68956 | 1.29% |\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F85d08acac7f886c3270dd1da47992811%2FEffect%20of%20misalignment%20correction%20%20soft%20labels.png?generation=1691897900692345&alt=media)",
    "2385453": "Interesting!\n\nHave you considered frame t=4 as skip connection to the decoder? ",
    "2385324": "Many congratulations on the solo sliver and thank you for the great write up!\n\nCould you please clarify this part:\n>Temporal Feature Modulator\n\n>First, I hypothesized that features on earlier layers are not contributed much to the final prediction, since contrails are drifted from frames to frames. So I come up with the idea of only applying temporal modulation in the later layers of feature maps.\n>\nI tried some experiments with applying temporal modulation to earlier layers, and found only applying last layer gives the best performance. Note that the rest of the feature maps are just sliced by the current time frame (T=4).",
    "2398367": "Congratulations and thanks for sharing such a detailed writeup! \nDo you know why HFlip worked for you even though you didn't correct label misalignment?",
    "2388440": "Hi @tatamikenn Thank you for sharing your solution and after deadline experiments \nI'm wondering why multiply *2  here\n`human_pixel_mask = np.clip(individual_pixel_mask.mean(-1) * 2, 0, 1)`\n"
  }
}