{
  "id": 430491,
  "title": "2nd place solution",
  "url": "/competitions/google-research-identify-contrails-reduce-global-warming/discussion/430491",
  "author_name": "Iafoss",
  "post_date": "2023-08-10T03:22:46.512000",
  "votes": 111,
  "comment_count": 73,
  "views": 0,
  "content": "<h1>Summary</h1>\n<ul>\n<li>Customized U-shape network with CoaT, NeXtViT, SAM-B, and tf_efficientnetv2_s backbones</li>\n<li>×2 (with pixel shuffle up-scaling of predictions to 256×256) and ×4 input upscaling: pixel-level accuracy of predictions is critical</li>\n<li>Sequence of images: LSTM, Transformer, Convolutional temporal mixing</li>\n<li>No flip and 90-degree rotation augmentation (nor TTA) because masks are shifted</li>\n<li>Train on soft labels: label = average of all annotators</li>\n<li>BCE + dice Lovasz loss</li>\n</ul>\n<h1>Code</h1>\n<p>My and <a href=\"https://www.kaggle.com/drhabib\" target=\"_blank\">@drhabib</a> 's <a href=\"https://github.com/DrHB/2nd-place-contrails\" target=\"_blank\">part</a> (includes the best performing CoaT_ULSTM model)<br>\n<a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> 's <a href=\"https://github.com/TheoViel/kaggle_contrails\" target=\"_blank\">part</a><br>\n<a href=\"https://www.kaggle.com/code/theoviel/contrails-inference-comb\" target=\"_blank\">Inference</a></p>\n<h1>Introduction</h1>\n<p>Our team would like to thank the organizers and Kaggle for making this competition possible. Also, I want to express my gratitude to my outstanding teammates <a href=\"https://www.kaggle.com/drhabib\" target=\"_blank\">@drhabib</a> and <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> for their incredible contribution toward our final result. </p>\n<h1>Details</h1>\n<p>This competition has two main challenges: (1) noisy labels and (2) pixel-level accuracy requirements. (1) If one checks the annotation from all annotators, a major disagreement is quite apparent. This label noise challenge may be partially addressed by using soft labels (the average of all annotator labels) during training. However, since the evaluation requires hard labels, the major model failure is the prediction of contrails near the decision boundary, which cannot be avoided. (2) The thing that could really be addressed by the models is achieving pixel-level accuracy of predictions. Pixel-level accuracy is quite important since contrails are only several pixels thick, and even a mistake by one pixel in the mask boundary in the lateral direction may result in a significant decrease in the dice score, used as a metric. For such a task, a typical solution is up-sampling the input of the model or replacing the final linear up-sampling in segmentation models with transposed convolution or pixel shuffle up-sampling. This modification gave us a substantial boost in initial experiments over the direct use of models on the original resolution.</p>\n<h2>Data</h2>\n<p>Following the <a href=\"https://arxiv.org/pdf/2304.02122.pdf\" target=\"_blank\">organizer's publication</a>, we used “ash” false color images considering the 12 μm band, the difference between 12 and 11 μm bands, and the difference between 11 and 8 μm bands, respectively. We also tried to consider all bands or expand the “ash” color images with 8, 10, and 12 μm, but it resulted in lower performance (the pre-trained weights of the first convolution were replicated). We hypothesize that the best performance of ash color images may be a consequence of the use of the ash color images by annotators, and biased label noise. In our experiments, we up-sample the input with bi-cubic interpolation by ×2 for CoaT and NeXtViT models, and by ×4 for SAM model. EfficientNet model uses the original input but the stride in the first convolutional block is set to 1, which is equivalent to ×4 input up-sampling.</p>\n<h2>Model</h2>\n<p>The interesting thing about this competition is that the model is required to do both (1) tracking the global pixel dependencies because contrails are quite elongated, and (2) capable of generating predictions with pixel-level accuracy because contrails are only several pixels thick. (1) can be addressed with transformers, while (2) is addressed with convolutional networks or local attention. In the competition, we tried to address these points independently with SAM-B ViT backbone and EfficientNet v2 purely convolutional network. In addition, there is a class of networks, transformers with a hierarchical structure, which addresses both aspects of the considered problem, (1) + (2), simultaneously. This class includes CoaT, which resulted in the best single model performance in our experiments, NeXtViT, MaxViT (which we missed), and MiT transformer from Segformer (unfortunately, it cannot be used in the competition because of the license).<br>\nFor the decoder we either considered a vanilla U-Net decoder (EfficientNet models) or a more customized set of U-blocks using pixel shuffle up-sampling and factorized FPN (CoaT, NeXtViT, SAM). In these blocks, we also replaced the BatchNorm with LayerNorm2d to have proper model convergence at the low batch size and used GELU instead of ReLU activation. <br>\nThe models are schematically illustrated in the figure below.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2F16e90aba2cb8ee7ac94837e09455688f%2F4.png?generation=1692419847964897&amp;alt=media\" alt=\"\"></p>\n<h2>Temporal mixing</h2>\n<p>In addition to our initial runs with single-frame models, we performed experiments with image sequences. In our best setup image, multi-frame models get ~0.01 improvement.<br>\nSince the images are taken with a considerable temporal delay, and the displacement of clouds between frames is huge (~10-30 pixels) and is deteriorated by input up-sampling, video models based on 3D convolutions or window attention (VideoSwin) are not effective in the context of the considered problem. For the same reason, early mixing or mixing the predictions of sequential models is also not expected to work. Therefore, mixing at the intermediate feature map scales, such as res/32 and res/16, is most promising: even if clouds/contrails are misplaced, the feature maps at low resolution are reasonably aligned. In our experiments with the CoaT model, the best performance is achieved if two low-resolution feature maps (res/32 and res/16) are considered. The output of the temporal mixing modules as well as the output of the remaining feature maps are pooled at frame 4 and the decoder sees only the input corresponding to the considered frame, as illustrated in Figure 1.<br>\nAs temporal mixing modules, we considered LSTMs (which achieved the best performance), 1D temporal convolutions, and Transformers. LSTM (1 layer) and 1D temporal convolutions perform mixing along the temporal dimension only, while the spatial dependences between features are considered in the backbone and the decoder. The transformer-based mixing was proposed to perform an implicit feature map registration for a large displacement of contrails/clouds because it works with temporal and spatial mixing of features simultaneously and can match them even if they are not spatially aligned. <br>\nThis mixing is performed in the following way. To each feature map, corresponding to a given frame, we add encoding to distinguish it from others, then we flatten tokens from all frames into a single sequence and process it with 2 transformer decoder blocks. Key and value input to the transformer is represented by the concatenated sequence of all frames, while the query is a sequence of tokens from the 4-th frame only. Unfortunately, transformer-based mixing ended up with slightly lower performance, and the training with it is less stable. However, we kept both approaches in the final ensemble, benefitting from the improved diversity.<br>\nIn our experiments with CoaT/NeXtViT we followed the <a href=\"https://arxiv.org/pdf/2304.02122.pdf\" target=\"_blank\">organizer's publication</a> and selected the first 5 frames as the input. Both LSTM and Transformer mixing were considered. In our experiments with EfficientNet we considered 4 frame input (2nd – 5th, 2 frames before and 1 frame after the annotated frame) with bi-directional LSTM. This frame selection is dictated by the instruction to annotators to have a contrail at least on two subsequent frames. With the SAM model, we considered 1D-convolutional mixing with 3 frames (1 frame before and 1 frame after the annotation) because of the heavy VRAM requirement for this model, which was also an issue in inference. <br>\nOne idea we had is explicit registration of input frames to have contrails and clouds aligned between frames (temporal mixing is the most effective), for example considering optical flow. Unfortunately, the images contain not only clouds but also the Earth's surface, which does not move, or several layers of clouds moving in different directions. Therefore, we could not come up with any methods to proceed with the frame registration task. We were also considering performing registration based on the predicted contrails, but since contrails appear and disappear from frame to frame, this approach is also not feasible.</p>\n<h2>Pseudo-labels (PL)</h2>\n<p>In addition to the competition data, we have collected a dataset “Contrails GOES16 Images May” (single frame setup) [9]. We split the images into 256×256 tiles with partial overlap and generated the masks based on the ensemble of best single-frame models. This data is shared as a Kaggle dataset [10]. This data was used for pretraining the models in a single-frame setup in some experiments, followed by finetuning the models either in single-frame or multi-frame setups. The use of PL has drastically improved the performance of individual models (folds), but the diversity of the models is reduced even if each fold of the finetuned model uses an independently trained PL model. Therefore, the performance on the average prediction over all folds is comparable to one for the setup trained without PL.<br>\nWe also have experimented with PL training on both external data + unannotated competition frames, but it did not give any benefit in finetuning a single frame model. Since such pretraining could affect multi-frame models, in the production experiments we used only external data at the PL step.</p>\n<h2>Training</h2>\n<p>CoaT and NeXtViT models: Training is performed for 24 epochs with Over9000 (Radam+LAMB+LookAhead) optimizer, learning rate of 3.5e-4, weight decay of 0.01. It appeared that this optimizer has superior compatibility with CoaT. In the case of PL pretraining, it is performed for 12 epochs on the external data, and then the models are finetuned for 12 epochs for CoaT and 18 epochs for NeXtViT models.<br>\nEfficientNet models: Pretraining is performed in a single frame setup for 100 epochs with additional CutMix augmentation, learning rate is 1e-3, AdamW optimizer, weight decay 0.2. During pretraining the fraction of PL data is linearly decreased to zero. Then the model is fine-tuned in the image sequence setup with 3e-5 for the encoder and 1e-4 for the decoder. The addition of external data for pretraining resulted in much stronger 2D models, hence multi-frame EfficientNet models were not considered in the final ensemble. For diversity, we also pre-trained models for 200 epochs, and in a full-fit setup (training on all data, including validation set).<br>\nSAM models: Our final submission includes only single-frame SAM-B ViT-based models. During training, we, first, train the model with a frozen encoder for 10 epochs, and then we unfreeze the model and continue for 20 more epochs. AdamW optimizer is used with the learning rate of 4e-6 for the encoder and 4e-5 for the decoder.<br>\nDuring training, we used ShiftRotateScale, RandomGamma, RandomBrightnessContrast, MotionBlur, and GaussianBlur augmentations. Flip and 90-rotate augmentation were not used because they resulted in worse performance. The training of all production models is performed for 5-6 folds our folds are trained on the entire train dataset with different random seeds, while the evaluation is performed on the provided validation set. This strategy is used because evaluation in a standard K-fold split of the training data overestimates CV (cross-validation), which may be caused by spatial overlaps of tiles in the training dataset.</p>\n<h2>Loss</h2>\n<p>As a loss function, we used BCE + dice Lovasz loss. The second term is a modified version of a Lovasz loss [12] where (1) ReLU is replaced with 1 + ELU, (2) a symmetric version is used, and (3) instead of IoU we use a surrogate function maximizing dice. The second term provides an insignificant boost to the individual model performance. However, this term gives a wider maximum of the dice with respect to the threshold selection, i.e. weaker dependence of the performance on the threshold selection. This property is vital to avoid shakeup at the private leaderboard because of the improper threshold and is preferable for more effective ensembling.</p>\n<h2>Results</h2>\n<p>The performance of our final models is summarized in the table below. Ex refers to models trained in PL setup (EfficientNet is trained with PL only). The postfix U refers to the single frame model, UT – transformer-based temporal mixing, and ULSTM – LSTM-based temporal mixing. The threshold is selected based on the search maximizing CV value. For most models, it is 0.46-0.50.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2F5b5010ef692bccaca91f04bdd8ae6be3%2F5.png?generation=1692423621715751&amp;alt=media\" alt=\"\"></p>\n<h2>Best Single model</h2>\n<p>Our best single mode is CoaT, which for 5 folds gets 0.7039 CV, 0.71790 private, and 0.71243 public LB. Single fold CV is 0.6960+-0.0003 evaluated on val set (our folds are trained on the train set with different seeds because eval in a standard Kfold split of train data overestimates CV due to possible spatial overlaps). <strong>We could take top-4 in the competition with this single model</strong>. </p>\n<h2>Ensemble</h2>\n<p>Our best ensemble consists of 8 models with weights selected to make approximately equal contributions from all components: CV is 0.7140 (on eval set), 0.72574 public and 0.72304 private LB. The models and their weight are summarized in the table below.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2F2ddf4778766d0b5aa44a24a81409d460%2F6.png?generation=1692421380873655&amp;alt=media\" alt=\"\"><br>\nEfficientNet includes 3 model setups (100 epochs, 200 epochs, and full-fit) with the total efficient weight equal to 1.</p>\n<h2>Model Execution Time</h2>\n<p>CoaT and NeXtViT models: Training time is 2 hours for single-frame and 8 hours for multi-frame models per fold on 2×RTX4090 GPUs. The inference time at Kaggle (P100) for multi-frame models is 35 minutes for 5 folds.<br>\nEfficientNet models: Training time is 3 hours 45 minutes for 100 epochs on 8×V100 GPUs per fold. Inference at Kaggle (P100) takes 20 minutes for 6 folds including data loading.<br>\nSAM models: Training time is 16 hours per fold on 2×A6000 GPUs. Inference at Kaggle (P100) takes 48 minutes for 5 folds including data loading.</p>\n<h2>Things we missed</h2>\n<p>(1) During the competition one surprising finding for us was that flips and 90-degree rotation augmentation in model training resulted in worse performance. It was quite surprising, but, unfortunately, we did not think sufficiently about the origin of this behavior and attributed it to prevailing winds and a particular shape of clouds and contrails. After the competition, it appeared that the masks were shifted by 0.5 pixels, which made any flips inapplicable, as we discovered. If the mask issue was fixed in our setup, flip/90-rotation augmentation could lead to an 8-fold increase of the effective training dataset size and enables further performance boost at the inference stage due to test time augmentation. So, potentially the performance of our models can be noticeably improved. <br>\n(2) PL single model pertaining on the external data gives moderate CV improvement to the single-frame model (interesting that at private LB this improvement is quite large). Meanwhile, this single-frame pretraining didn't give much to sequence models. So, external sequence data collection + proper reprojection could potentially server as a better pertaining strategy for sequence models, and we may expect a boost comparable to single-frame models.</p>",
  "messages": [
    {
      "id": 2382786,
      "postDate": "2023-08-10T03:22:46.513Z",
      "content": "<h1>Summary</h1>\n<ul>\n<li>Customized U-shape network with CoaT, NeXtViT, SAM-B, and tf_efficientnetv2_s backbones</li>\n<li>×2 (with pixel shuffle up-scaling of predictions to 256×256) and ×4 input upscaling: pixel-level accuracy of predictions is critical</li>\n<li>Sequence of images: LSTM, Transformer, Convolutional temporal mixing</li>\n<li>No flip and 90-degree rotation augmentation (nor TTA) because masks are shifted</li>\n<li>Train on soft labels: label = average of all annotators</li>\n<li>BCE + dice Lovasz loss</li>\n</ul>\n<h1>Code</h1>\n<p>My and <a href=\"https://www.kaggle.com/drhabib\" target=\"_blank\">@drhabib</a> 's <a href=\"https://github.com/DrHB/2nd-place-contrails\" target=\"_blank\">part</a> (includes the best performing CoaT_ULSTM model)<br>\n<a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> 's <a href=\"https://github.com/TheoViel/kaggle_contrails\" target=\"_blank\">part</a><br>\n<a href=\"https://www.kaggle.com/code/theoviel/contrails-inference-comb\" target=\"_blank\">Inference</a></p>\n<h1>Introduction</h1>\n<p>Our team would like to thank the organizers and Kaggle for making this competition possible. Also, I want to express my gratitude to my outstanding teammates <a href=\"https://www.kaggle.com/drhabib\" target=\"_blank\">@drhabib</a> and <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> for their incredible contribution toward our final result. </p>\n<h1>Details</h1>\n<p>This competition has two main challenges: (1) noisy labels and (2) pixel-level accuracy requirements. (1) If one checks the annotation from all annotators, a major disagreement is quite apparent. This label noise challenge may be partially addressed by using soft labels (the average of all annotator labels) during training. However, since the evaluation requires hard labels, the major model failure is the prediction of contrails near the decision boundary, which cannot be avoided. (2) The thing that could really be addressed by the models is achieving pixel-level accuracy of predictions. Pixel-level accuracy is quite important since contrails are only several pixels thick, and even a mistake by one pixel in the mask boundary in the lateral direction may result in a significant decrease in the dice score, used as a metric. For such a task, a typical solution is up-sampling the input of the model or replacing the final linear up-sampling in segmentation models with transposed convolution or pixel shuffle up-sampling. This modification gave us a substantial boost in initial experiments over the direct use of models on the original resolution.</p>\n<h2>Data</h2>\n<p>Following the <a href=\"https://arxiv.org/pdf/2304.02122.pdf\" target=\"_blank\">organizer's publication</a>, we used “ash” false color images considering the 12 μm band, the difference between 12 and 11 μm bands, and the difference between 11 and 8 μm bands, respectively. We also tried to consider all bands or expand the “ash” color images with 8, 10, and 12 μm, but it resulted in lower performance (the pre-trained weights of the first convolution were replicated). We hypothesize that the best performance of ash color images may be a consequence of the use of the ash color images by annotators, and biased label noise. In our experiments, we up-sample the input with bi-cubic interpolation by ×2 for CoaT and NeXtViT models, and by ×4 for SAM model. EfficientNet model uses the original input but the stride in the first convolutional block is set to 1, which is equivalent to ×4 input up-sampling.</p>\n<h2>Model</h2>\n<p>The interesting thing about this competition is that the model is required to do both (1) tracking the global pixel dependencies because contrails are quite elongated, and (2) capable of generating predictions with pixel-level accuracy because contrails are only several pixels thick. (1) can be addressed with transformers, while (2) is addressed with convolutional networks or local attention. In the competition, we tried to address these points independently with SAM-B ViT backbone and EfficientNet v2 purely convolutional network. In addition, there is a class of networks, transformers with a hierarchical structure, which addresses both aspects of the considered problem, (1) + (2), simultaneously. This class includes CoaT, which resulted in the best single model performance in our experiments, NeXtViT, MaxViT (which we missed), and MiT transformer from Segformer (unfortunately, it cannot be used in the competition because of the license).<br>\nFor the decoder we either considered a vanilla U-Net decoder (EfficientNet models) or a more customized set of U-blocks using pixel shuffle up-sampling and factorized FPN (CoaT, NeXtViT, SAM). In these blocks, we also replaced the BatchNorm with LayerNorm2d to have proper model convergence at the low batch size and used GELU instead of ReLU activation. <br>\nThe models are schematically illustrated in the figure below.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2F16e90aba2cb8ee7ac94837e09455688f%2F4.png?generation=1692419847964897&amp;alt=media\" alt=\"\"></p>\n<h2>Temporal mixing</h2>\n<p>In addition to our initial runs with single-frame models, we performed experiments with image sequences. In our best setup image, multi-frame models get ~0.01 improvement.<br>\nSince the images are taken with a considerable temporal delay, and the displacement of clouds between frames is huge (~10-30 pixels) and is deteriorated by input up-sampling, video models based on 3D convolutions or window attention (VideoSwin) are not effective in the context of the considered problem. For the same reason, early mixing or mixing the predictions of sequential models is also not expected to work. Therefore, mixing at the intermediate feature map scales, such as res/32 and res/16, is most promising: even if clouds/contrails are misplaced, the feature maps at low resolution are reasonably aligned. In our experiments with the CoaT model, the best performance is achieved if two low-resolution feature maps (res/32 and res/16) are considered. The output of the temporal mixing modules as well as the output of the remaining feature maps are pooled at frame 4 and the decoder sees only the input corresponding to the considered frame, as illustrated in Figure 1.<br>\nAs temporal mixing modules, we considered LSTMs (which achieved the best performance), 1D temporal convolutions, and Transformers. LSTM (1 layer) and 1D temporal convolutions perform mixing along the temporal dimension only, while the spatial dependences between features are considered in the backbone and the decoder. The transformer-based mixing was proposed to perform an implicit feature map registration for a large displacement of contrails/clouds because it works with temporal and spatial mixing of features simultaneously and can match them even if they are not spatially aligned. <br>\nThis mixing is performed in the following way. To each feature map, corresponding to a given frame, we add encoding to distinguish it from others, then we flatten tokens from all frames into a single sequence and process it with 2 transformer decoder blocks. Key and value input to the transformer is represented by the concatenated sequence of all frames, while the query is a sequence of tokens from the 4-th frame only. Unfortunately, transformer-based mixing ended up with slightly lower performance, and the training with it is less stable. However, we kept both approaches in the final ensemble, benefitting from the improved diversity.<br>\nIn our experiments with CoaT/NeXtViT we followed the <a href=\"https://arxiv.org/pdf/2304.02122.pdf\" target=\"_blank\">organizer's publication</a> and selected the first 5 frames as the input. Both LSTM and Transformer mixing were considered. In our experiments with EfficientNet we considered 4 frame input (2nd – 5th, 2 frames before and 1 frame after the annotated frame) with bi-directional LSTM. This frame selection is dictated by the instruction to annotators to have a contrail at least on two subsequent frames. With the SAM model, we considered 1D-convolutional mixing with 3 frames (1 frame before and 1 frame after the annotation) because of the heavy VRAM requirement for this model, which was also an issue in inference. <br>\nOne idea we had is explicit registration of input frames to have contrails and clouds aligned between frames (temporal mixing is the most effective), for example considering optical flow. Unfortunately, the images contain not only clouds but also the Earth's surface, which does not move, or several layers of clouds moving in different directions. Therefore, we could not come up with any methods to proceed with the frame registration task. We were also considering performing registration based on the predicted contrails, but since contrails appear and disappear from frame to frame, this approach is also not feasible.</p>\n<h2>Pseudo-labels (PL)</h2>\n<p>In addition to the competition data, we have collected a dataset “Contrails GOES16 Images May” (single frame setup) [9]. We split the images into 256×256 tiles with partial overlap and generated the masks based on the ensemble of best single-frame models. This data is shared as a Kaggle dataset [10]. This data was used for pretraining the models in a single-frame setup in some experiments, followed by finetuning the models either in single-frame or multi-frame setups. The use of PL has drastically improved the performance of individual models (folds), but the diversity of the models is reduced even if each fold of the finetuned model uses an independently trained PL model. Therefore, the performance on the average prediction over all folds is comparable to one for the setup trained without PL.<br>\nWe also have experimented with PL training on both external data + unannotated competition frames, but it did not give any benefit in finetuning a single frame model. Since such pretraining could affect multi-frame models, in the production experiments we used only external data at the PL step.</p>\n<h2>Training</h2>\n<p>CoaT and NeXtViT models: Training is performed for 24 epochs with Over9000 (Radam+LAMB+LookAhead) optimizer, learning rate of 3.5e-4, weight decay of 0.01. It appeared that this optimizer has superior compatibility with CoaT. In the case of PL pretraining, it is performed for 12 epochs on the external data, and then the models are finetuned for 12 epochs for CoaT and 18 epochs for NeXtViT models.<br>\nEfficientNet models: Pretraining is performed in a single frame setup for 100 epochs with additional CutMix augmentation, learning rate is 1e-3, AdamW optimizer, weight decay 0.2. During pretraining the fraction of PL data is linearly decreased to zero. Then the model is fine-tuned in the image sequence setup with 3e-5 for the encoder and 1e-4 for the decoder. The addition of external data for pretraining resulted in much stronger 2D models, hence multi-frame EfficientNet models were not considered in the final ensemble. For diversity, we also pre-trained models for 200 epochs, and in a full-fit setup (training on all data, including validation set).<br>\nSAM models: Our final submission includes only single-frame SAM-B ViT-based models. During training, we, first, train the model with a frozen encoder for 10 epochs, and then we unfreeze the model and continue for 20 more epochs. AdamW optimizer is used with the learning rate of 4e-6 for the encoder and 4e-5 for the decoder.<br>\nDuring training, we used ShiftRotateScale, RandomGamma, RandomBrightnessContrast, MotionBlur, and GaussianBlur augmentations. Flip and 90-rotate augmentation were not used because they resulted in worse performance. The training of all production models is performed for 5-6 folds our folds are trained on the entire train dataset with different random seeds, while the evaluation is performed on the provided validation set. This strategy is used because evaluation in a standard K-fold split of the training data overestimates CV (cross-validation), which may be caused by spatial overlaps of tiles in the training dataset.</p>\n<h2>Loss</h2>\n<p>As a loss function, we used BCE + dice Lovasz loss. The second term is a modified version of a Lovasz loss [12] where (1) ReLU is replaced with 1 + ELU, (2) a symmetric version is used, and (3) instead of IoU we use a surrogate function maximizing dice. The second term provides an insignificant boost to the individual model performance. However, this term gives a wider maximum of the dice with respect to the threshold selection, i.e. weaker dependence of the performance on the threshold selection. This property is vital to avoid shakeup at the private leaderboard because of the improper threshold and is preferable for more effective ensembling.</p>\n<h2>Results</h2>\n<p>The performance of our final models is summarized in the table below. Ex refers to models trained in PL setup (EfficientNet is trained with PL only). The postfix U refers to the single frame model, UT – transformer-based temporal mixing, and ULSTM – LSTM-based temporal mixing. The threshold is selected based on the search maximizing CV value. For most models, it is 0.46-0.50.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2F5b5010ef692bccaca91f04bdd8ae6be3%2F5.png?generation=1692423621715751&amp;alt=media\" alt=\"\"></p>\n<h2>Best Single model</h2>\n<p>Our best single mode is CoaT, which for 5 folds gets 0.7039 CV, 0.71790 private, and 0.71243 public LB. Single fold CV is 0.6960+-0.0003 evaluated on val set (our folds are trained on the train set with different seeds because eval in a standard Kfold split of train data overestimates CV due to possible spatial overlaps). <strong>We could take top-4 in the competition with this single model</strong>. </p>\n<h2>Ensemble</h2>\n<p>Our best ensemble consists of 8 models with weights selected to make approximately equal contributions from all components: CV is 0.7140 (on eval set), 0.72574 public and 0.72304 private LB. The models and their weight are summarized in the table below.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2F2ddf4778766d0b5aa44a24a81409d460%2F6.png?generation=1692421380873655&amp;alt=media\" alt=\"\"><br>\nEfficientNet includes 3 model setups (100 epochs, 200 epochs, and full-fit) with the total efficient weight equal to 1.</p>\n<h2>Model Execution Time</h2>\n<p>CoaT and NeXtViT models: Training time is 2 hours for single-frame and 8 hours for multi-frame models per fold on 2×RTX4090 GPUs. The inference time at Kaggle (P100) for multi-frame models is 35 minutes for 5 folds.<br>\nEfficientNet models: Training time is 3 hours 45 minutes for 100 epochs on 8×V100 GPUs per fold. Inference at Kaggle (P100) takes 20 minutes for 6 folds including data loading.<br>\nSAM models: Training time is 16 hours per fold on 2×A6000 GPUs. Inference at Kaggle (P100) takes 48 minutes for 5 folds including data loading.</p>\n<h2>Things we missed</h2>\n<p>(1) During the competition one surprising finding for us was that flips and 90-degree rotation augmentation in model training resulted in worse performance. It was quite surprising, but, unfortunately, we did not think sufficiently about the origin of this behavior and attributed it to prevailing winds and a particular shape of clouds and contrails. After the competition, it appeared that the masks were shifted by 0.5 pixels, which made any flips inapplicable, as we discovered. If the mask issue was fixed in our setup, flip/90-rotation augmentation could lead to an 8-fold increase of the effective training dataset size and enables further performance boost at the inference stage due to test time augmentation. So, potentially the performance of our models can be noticeably improved. <br>\n(2) PL single model pertaining on the external data gives moderate CV improvement to the single-frame model (interesting that at private LB this improvement is quite large). Meanwhile, this single-frame pretraining didn't give much to sequence models. So, external sequence data collection + proper reprojection could potentially server as a better pertaining strategy for sequence models, and we may expect a boost comparable to single-frame models.</p>",
      "rawMarkdown": "# Summary\n- Customized U-shape network with CoaT, NeXtViT, SAM-B, and tf_efficientnetv2_s backbones\n- ×2 (with pixel shuffle up-scaling of predictions to 256×256) and ×4 input upscaling: pixel-level accuracy of predictions is critical\n- Sequence of images: LSTM, Transformer, Convolutional temporal mixing\n- No flip and 90-degree rotation augmentation (nor TTA) because masks are shifted\n- Train on soft labels: label = average of all annotators\n- BCE + dice Lovasz loss\n\n# Code\nMy and @drhabib 's [part](https://github.com/DrHB/2nd-place-contrails) (includes the best performing CoaT_ULSTM model)\n@theoviel 's [part](https://github.com/TheoViel/kaggle_contrails)\n[Inference](https://www.kaggle.com/code/theoviel/contrails-inference-comb)\n\n# Introduction\nOur team would like to thank the organizers and Kaggle for making this competition possible. Also, I want to express my gratitude to my outstanding teammates @drhabib and @theoviel for their incredible contribution toward our final result. \n\n# Details\nThis competition has two main challenges: (1) noisy labels and (2) pixel-level accuracy requirements. (1) If one checks the annotation from all annotators, a major disagreement is quite apparent. This label noise challenge may be partially addressed by using soft labels (the average of all annotator labels) during training. However, since the evaluation requires hard labels, the major model failure is the prediction of contrails near the decision boundary, which cannot be avoided. (2) The thing that could really be addressed by the models is achieving pixel-level accuracy of predictions. Pixel-level accuracy is quite important since contrails are only several pixels thick, and even a mistake by one pixel in the mask boundary in the lateral direction may result in a significant decrease in the dice score, used as a metric. For such a task, a typical solution is up-sampling the input of the model or replacing the final linear up-sampling in segmentation models with transposed convolution or pixel shuffle up-sampling. This modification gave us a substantial boost in initial experiments over the direct use of models on the original resolution.\n\n## Data\nFollowing the [organizer's publication](https://arxiv.org/pdf/2304.02122.pdf), we used “ash” false color images considering the 12 μm band, the difference between 12 and 11 μm bands, and the difference between 11 and 8 μm bands, respectively. We also tried to consider all bands or expand the “ash” color images with 8, 10, and 12 μm, but it resulted in lower performance (the pre-trained weights of the first convolution were replicated). We hypothesize that the best performance of ash color images may be a consequence of the use of the ash color images by annotators, and biased label noise. In our experiments, we up-sample the input with bi-cubic interpolation by ×2 for CoaT and NeXtViT models, and by ×4 for SAM model. EfficientNet model uses the original input but the stride in the first convolutional block is set to 1, which is equivalent to ×4 input up-sampling.\n\n## Model\nThe interesting thing about this competition is that the model is required to do both (1) tracking the global pixel dependencies because contrails are quite elongated, and (2) capable of generating predictions with pixel-level accuracy because contrails are only several pixels thick. (1) can be addressed with transformers, while (2) is addressed with convolutional networks or local attention. In the competition, we tried to address these points independently with SAM-B ViT backbone and EfficientNet v2 purely convolutional network. In addition, there is a class of networks, transformers with a hierarchical structure, which addresses both aspects of the considered problem, (1) + (2), simultaneously. This class includes CoaT, which resulted in the best single model performance in our experiments, NeXtViT, MaxViT (which we missed), and MiT transformer from Segformer (unfortunately, it cannot be used in the competition because of the license).\nFor the decoder we either considered a vanilla U-Net decoder (EfficientNet models) or a more customized set of U-blocks using pixel shuffle up-sampling and factorized FPN (CoaT, NeXtViT, SAM). In these blocks, we also replaced the BatchNorm with LayerNorm2d to have proper model convergence at the low batch size and used GELU instead of ReLU activation. \nThe models are schematically illustrated in the figure below.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2F16e90aba2cb8ee7ac94837e09455688f%2F4.png?generation=1692419847964897&alt=media)\n\n## Temporal mixing\nIn addition to our initial runs with single-frame models, we performed experiments with image sequences. In our best setup image, multi-frame models get ~0.01 improvement.\nSince the images are taken with a considerable temporal delay, and the displacement of clouds between frames is huge (~10-30 pixels) and is deteriorated by input up-sampling, video models based on 3D convolutions or window attention (VideoSwin) are not effective in the context of the considered problem. For the same reason, early mixing or mixing the predictions of sequential models is also not expected to work. Therefore, mixing at the intermediate feature map scales, such as res/32 and res/16, is most promising: even if clouds/contrails are misplaced, the feature maps at low resolution are reasonably aligned. In our experiments with the CoaT model, the best performance is achieved if two low-resolution feature maps (res/32 and res/16) are considered. The output of the temporal mixing modules as well as the output of the remaining feature maps are pooled at frame 4 and the decoder sees only the input corresponding to the considered frame, as illustrated in Figure 1.\nAs temporal mixing modules, we considered LSTMs (which achieved the best performance), 1D temporal convolutions, and Transformers. LSTM (1 layer) and 1D temporal convolutions perform mixing along the temporal dimension only, while the spatial dependences between features are considered in the backbone and the decoder. The transformer-based mixing was proposed to perform an implicit feature map registration for a large displacement of contrails/clouds because it works with temporal and spatial mixing of features simultaneously and can match them even if they are not spatially aligned. \nThis mixing is performed in the following way. To each feature map, corresponding to a given frame, we add encoding to distinguish it from others, then we flatten tokens from all frames into a single sequence and process it with 2 transformer decoder blocks. Key and value input to the transformer is represented by the concatenated sequence of all frames, while the query is a sequence of tokens from the 4-th frame only. Unfortunately, transformer-based mixing ended up with slightly lower performance, and the training with it is less stable. However, we kept both approaches in the final ensemble, benefitting from the improved diversity.\nIn our experiments with CoaT/NeXtViT we followed the [organizer's publication](https://arxiv.org/pdf/2304.02122.pdf) and selected the first 5 frames as the input. Both LSTM and Transformer mixing were considered. In our experiments with EfficientNet we considered 4 frame input (2nd – 5th, 2 frames before and 1 frame after the annotated frame) with bi-directional LSTM. This frame selection is dictated by the instruction to annotators to have a contrail at least on two subsequent frames. With the SAM model, we considered 1D-convolutional mixing with 3 frames (1 frame before and 1 frame after the annotation) because of the heavy VRAM requirement for this model, which was also an issue in inference. \nOne idea we had is explicit registration of input frames to have contrails and clouds aligned between frames (temporal mixing is the most effective), for example considering optical flow. Unfortunately, the images contain not only clouds but also the Earth's surface, which does not move, or several layers of clouds moving in different directions. Therefore, we could not come up with any methods to proceed with the frame registration task. We were also considering performing registration based on the predicted contrails, but since contrails appear and disappear from frame to frame, this approach is also not feasible.\n\n##Pseudo-labels (PL)\nIn addition to the competition data, we have collected a dataset “Contrails GOES16 Images May” (single frame setup) [9]. We split the images into 256×256 tiles with partial overlap and generated the masks based on the ensemble of best single-frame models. This data is shared as a Kaggle dataset [10]. This data was used for pretraining the models in a single-frame setup in some experiments, followed by finetuning the models either in single-frame or multi-frame setups. The use of PL has drastically improved the performance of individual models (folds), but the diversity of the models is reduced even if each fold of the finetuned model uses an independently trained PL model. Therefore, the performance on the average prediction over all folds is comparable to one for the setup trained without PL.\nWe also have experimented with PL training on both external data + unannotated competition frames, but it did not give any benefit in finetuning a single frame model. Since such pretraining could affect multi-frame models, in the production experiments we used only external data at the PL step.\n\n## Training\nCoaT and NeXtViT models: Training is performed for 24 epochs with Over9000 (Radam+LAMB+LookAhead) optimizer, learning rate of 3.5e-4, weight decay of 0.01. It appeared that this optimizer has superior compatibility with CoaT. In the case of PL pretraining, it is performed for 12 epochs on the external data, and then the models are finetuned for 12 epochs for CoaT and 18 epochs for NeXtViT models.\nEfficientNet models: Pretraining is performed in a single frame setup for 100 epochs with additional CutMix augmentation, learning rate is 1e-3, AdamW optimizer, weight decay 0.2. During pretraining the fraction of PL data is linearly decreased to zero. Then the model is fine-tuned in the image sequence setup with 3e-5 for the encoder and 1e-4 for the decoder. The addition of external data for pretraining resulted in much stronger 2D models, hence multi-frame EfficientNet models were not considered in the final ensemble. For diversity, we also pre-trained models for 200 epochs, and in a full-fit setup (training on all data, including validation set).\nSAM models: Our final submission includes only single-frame SAM-B ViT-based models. During training, we, first, train the model with a frozen encoder for 10 epochs, and then we unfreeze the model and continue for 20 more epochs. AdamW optimizer is used with the learning rate of 4e-6 for the encoder and 4e-5 for the decoder.\nDuring training, we used ShiftRotateScale, RandomGamma, RandomBrightnessContrast, MotionBlur, and GaussianBlur augmentations. Flip and 90-rotate augmentation were not used because they resulted in worse performance. The training of all production models is performed for 5-6 folds our folds are trained on the entire train dataset with different random seeds, while the evaluation is performed on the provided validation set. This strategy is used because evaluation in a standard K-fold split of the training data overestimates CV (cross-validation), which may be caused by spatial overlaps of tiles in the training dataset.\n\n## Loss\nAs a loss function, we used BCE + dice Lovasz loss. The second term is a modified version of a Lovasz loss [12] where (1) ReLU is replaced with 1 + ELU, (2) a symmetric version is used, and (3) instead of IoU we use a surrogate function maximizing dice. The second term provides an insignificant boost to the individual model performance. However, this term gives a wider maximum of the dice with respect to the threshold selection, i.e. weaker dependence of the performance on the threshold selection. This property is vital to avoid shakeup at the private leaderboard because of the improper threshold and is preferable for more effective ensembling.\n\n## Results\nThe performance of our final models is summarized in the table below. Ex refers to models trained in PL setup (EfficientNet is trained with PL only). The postfix U refers to the single frame model, UT – transformer-based temporal mixing, and ULSTM – LSTM-based temporal mixing. The threshold is selected based on the search maximizing CV value. For most models, it is 0.46-0.50.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2F5b5010ef692bccaca91f04bdd8ae6be3%2F5.png?generation=1692423621715751&alt=media)\n\n## Best Single model\nOur best single mode is CoaT, which for 5 folds gets 0.7039 CV, 0.71790 private, and 0.71243 public LB. Single fold CV is 0.6960+-0.0003 evaluated on val set (our folds are trained on the train set with different seeds because eval in a standard Kfold split of train data overestimates CV due to possible spatial overlaps). **We could take top-4 in the competition with this single model**. \n\n## Ensemble\nOur best ensemble consists of 8 models with weights selected to make approximately equal contributions from all components: CV is 0.7140 (on eval set), 0.72574 public and 0.72304 private LB. The models and their weight are summarized in the table below.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2F2ddf4778766d0b5aa44a24a81409d460%2F6.png?generation=1692421380873655&alt=media)\nEfficientNet includes 3 model setups (100 epochs, 200 epochs, and full-fit) with the total efficient weight equal to 1.\n\n##Model Execution Time\nCoaT and NeXtViT models: Training time is 2 hours for single-frame and 8 hours for multi-frame models per fold on 2×RTX4090 GPUs. The inference time at Kaggle (P100) for multi-frame models is 35 minutes for 5 folds.\nEfficientNet models: Training time is 3 hours 45 minutes for 100 epochs on 8×V100 GPUs per fold. Inference at Kaggle (P100) takes 20 minutes for 6 folds including data loading.\nSAM models: Training time is 16 hours per fold on 2×A6000 GPUs. Inference at Kaggle (P100) takes 48 minutes for 5 folds including data loading.\n\n## Things we missed\n(1) During the competition one surprising finding for us was that flips and 90-degree rotation augmentation in model training resulted in worse performance. It was quite surprising, but, unfortunately, we did not think sufficiently about the origin of this behavior and attributed it to prevailing winds and a particular shape of clouds and contrails. After the competition, it appeared that the masks were shifted by 0.5 pixels, which made any flips inapplicable, as we discovered. If the mask issue was fixed in our setup, flip/90-rotation augmentation could lead to an 8-fold increase of the effective training dataset size and enables further performance boost at the inference stage due to test time augmentation. So, potentially the performance of our models can be noticeably improved. \n(2) PL single model pertaining on the external data gives moderate CV improvement to the single-frame model (interesting that at private LB this improvement is quite large). Meanwhile, this single-frame pretraining didn't give much to sequence models. So, external sequence data collection + proper reprojection could potentially server as a better pertaining strategy for sequence models, and we may expect a boost comparable to single-frame models.",
      "votes": 111
    },
    {
      "id": 2382925,
      "postDate": "2023-08-10T05:28:09.093Z",
      "content": "<p>Congratulations! Beautiful solution, thank you for sharing it with us. Looking forward to studying your code.</p>\n<p>We knew you had a unique approach that gained your team a nice separation throughout. Even during the competition I was excited to learn about it and it didn't disappoint.</p>\n<p>We had a feeling the temporal information was important and tried a bunch of stuff like 3d models and sequence models to no avail, so we dropped the idea as computation was expensive.</p>\n<p>We also had a feeling that there was important information in the individual masks and tried a bunch of stuff like using random individual masks as auxiliary signal or different classes, but the simplest approach of averaging the individual masks didn't occur to us, nice one.</p>",
      "rawMarkdown": "Congratulations! Beautiful solution, thank you for sharing it with us. Looking forward to studying your code.\n\nWe knew you had a unique approach that gained your team a nice separation throughout. Even during the competition I was excited to learn about it and it didn't disappoint.\n\nWe had a feeling the temporal information was important and tried a bunch of stuff like 3d models and sequence models to no avail, so we dropped the idea as computation was expensive.\n\nWe also had a feeling that there was important information in the individual masks and tried a bunch of stuff like using random individual masks as auxiliary signal or different classes, but the simplest approach of averaging the individual masks didn't occur to us, nice one.",
      "votes": 8
    },
    {
      "id": 2382849,
      "postDate": "2023-08-10T04:27:04.067Z",
      "content": "<p>thanks for the writeup.</p>\n<p>i am interested in usage of tempral information. how much did it increase your CV and LB.???</p>\n<hr>\n<p>I spent the last week in fusion of multiple frams, early fusion, late fusion, etc …<br>\nI also tried optcal flow (RAFT) for registration and  augmentation (find the flow and then wrap image and mask)<br>\nnone of them really work well.<br>\ni only have 0.004~0.005 improvement when using 3 consective frames (over 2 different models i tried)<br>\nThis is worse than expected  (the paper can get 0.01 to 0.015 improvement). I wondered why????</p>\n<p>i don't think large movement  is the cause of poor performance.<br>\n(on a side note, actually you can use temporal superresolution, i.e. predict intermediate frames with frame interpolation. <a href=\"https://www.youtube.com/watch?v=MjViy6kyiqs)\" target=\"_blank\">https://www.youtube.com/watch?v=MjViy6kyiqs)</a>. Frame interpolation results may be wrong, but as long as it is \"consistent\", it can \"provide temporal features\"</p>",
      "rawMarkdown": "thanks for the writeup.\n\ni am interested in usage of tempral information. how much did it increase your CV and LB.???\n\n---\n\nI spent the last week in fusion of multiple frams, early fusion, late fusion, etc ...\nI also tried optcal flow (RAFT) for registration and  augmentation (find the flow and then wrap image and mask)\nnone of them really work well.\ni only have 0.004~0.005 improvement when using 3 consective frames (over 2 different models i tried)\nThis is worse than expected  (the paper can get 0.01 to 0.015 improvement). I wondered why????\n\ni don't think large movement  is the cause of poor performance.\n(on a side note, actually you can use temporal superresolution, i.e. predict intermediate frames with frame interpolation. https://www.youtube.com/watch?v=MjViy6kyiqs). Frame interpolation results may be wrong, but as long as it is \"consistent\", it can \"provide temporal features\"",
      "votes": 5,
      "replies": [
        {
          "id": 2382872,
          "postDate": "2023-08-10T04:43:18.023Z",
          "content": "<p>Here are some of the results from my notes (I list public CV):<br>\nCoaT_U (single image): 0.6887+-0.0014, avr pred 0.6983 (th 0.47), LB 0.706<br>\nExCoaT_U (single image, PL): 0.6952+-0.0008, avr pred 0.7002 (th 0.49)</p>\n<p>NeXtViT_U (single image): 0.6848 +- 0.0013, avr pred 0.6917 (th 0.47), LB 0.710<br>\nExNeXtViT_U (single image, PL): 0.6904 +- 0.0006, avr pred 0.6941 (th 0.48)</p>\n<p>CoaT_ULSTM (5 image seq): 0.6960+-0.0003, avr pred 0.7039 (th 0.48), LB 0.712<br>\nEX Coat_ULSTM (5 image seq, PL): 0.7027+-0.0002, avr pred 0.7064 (th 0.49), LB 721</p>\n<p>CoaT_UT (5 image seq, transformer mixing): 0.6950+-0.0013, avr pred 0.7055 (th 0.47), LB 0.722<br>\nEX Coat_UT (5 image seq, transformer mixing, PL): 0.6997+-0.0008, avr pred 0.7040 (th 0.49), LB 0.716</p>\n<p>NeXtViT_ULSTM (5 image seq): 0.6951+-0.0004, avr pred 0.7014 (th 0.50), LB 0.713<br>\nEx NeXtViT_ULSTM (5 image seq, PL): 0.6988 +- 0.0007, avr pred 0.7025 (th 0.50)</p>\n<p>Overall <strong>sequence model gives 0.005-0.01 boost</strong>. In my case the key thing was mixing at res/32 and res/16 feature maps followed with up-sampling U blocks, like in the single frame model. I used 0-5 frames, Theo was using 3-6  frames. 8 frames didn't perform well, but I didn't run full 5 fold for this setup.</p>",
          "rawMarkdown": "Here are some of the results from my notes (I list public CV):\nCoaT_U (single image): 0.6887+-0.0014, avr pred 0.6983 (th 0.47), LB 0.706\nExCoaT_U (single image, PL): 0.6952+-0.0008, avr pred 0.7002 (th 0.49)\n\nNeXtViT_U (single image): 0.6848 +- 0.0013, avr pred 0.6917 (th 0.47), LB 0.710\nExNeXtViT_U (single image, PL): 0.6904 +- 0.0006, avr pred 0.6941 (th 0.48)\n\nCoaT_ULSTM (5 image seq): 0.6960+-0.0003, avr pred 0.7039 (th 0.48), LB 0.712\nEX Coat_ULSTM (5 image seq, PL): 0.7027+-0.0002, avr pred 0.7064 (th 0.49), LB 721\n\nCoaT_UT (5 image seq, transformer mixing): 0.6950+-0.0013, avr pred 0.7055 (th 0.47), LB 0.722\nEX Coat_UT (5 image seq, transformer mixing, PL): 0.6997+-0.0008, avr pred 0.7040 (th 0.49), LB 0.716\n\nNeXtViT_ULSTM (5 image seq): 0.6951+-0.0004, avr pred 0.7014 (th 0.50), LB 0.713\nEx NeXtViT_ULSTM (5 image seq, PL): 0.6988 +- 0.0007, avr pred 0.7025 (th 0.50)\n\nOverall **sequence model gives 0.005-0.01 boost**. In my case the key thing was mixing at res/32 and res/16 feature maps followed with up-sampling U blocks, like in the single frame model. I used 0-5 frames, Theo was using 3-6  frames. 8 frames didn't perform well, but I didn't run full 5 fold for this setup.",
          "votes": 6,
          "replies": [
            {
              "id": 2382890,
              "postDate": "2023-08-10T04:54:55.670Z",
              "content": "<p>thanks for the details.</p>\n<p>i will try again later. i actually use two stage approach: first model is just trained on single frame. second models uses features from first one. both models are u-net. there is only improvement if i train e2e. it doesn't work if i freeze the first model. my mixing is done for all 4 scales. maybe that is the reason.</p>\n<p>lastly, congrats to you good ranking and score! 👍 👍 👍</p>\n<hr>\n<p>\"i don't think large movement is the cause of poor performance.\"</p>\n<p>this is because cloud/hurricane motion forecast or precipitation/rainfall prediction are also using large interval satellite images</p>",
              "rawMarkdown": "thanks for the details.\n\ni will try again later. i actually use two stage approach: first model is just trained on single frame. second models uses features from first one. both models are u-net. there is only improvement if i train e2e. it doesn't work if i freeze the first model. my mixing is done for all 4 scales. maybe that is the reason.\n\nlastly, congrats to you good ranking and score! 👍 👍 👍\n\n---\n\"i don't think large movement is the cause of poor performance.\"\n\nthis is because cloud/hurricane motion forecast or precipitation/rainfall prediction are also using large interval satellite images",
              "votes": 2
            },
            {
              "id": 2382897,
              "postDate": "2023-08-10T05:07:46.827Z",
              "content": "<p>Thanks so much. </p>\n<p>In my experiments use of 4 scales was worse than use of 2 scales, but I didn't run this experiment in 5 folds.</p>\n<p>In PL training I used a kind of 2 stage approach: train single frame model on PL and then add temporal mixing a and trained on the provided data (no freezing). Even If I trained 5 single frame distinct models models, the overall diversity was quite bad, and this model was not very helpful for the ensemble.</p>\n<p>Regarding large movent between frames, I expect it is a bad thing because LSTM mixing does not consider spatial misalignment. If the contrail is shifted from one pixel at res/32 feature map to another, LSTM cannot track it properly because it is applied only in the temporal dimension( So displacement of clouds  by the receptive field of the model may kill the advantage of the seq model.</p>",
              "rawMarkdown": "Thanks so much. \n\nIn my experiments use of 4 scales was worse than use of 2 scales, but I didn't run this experiment in 5 folds.\n\nIn PL training I used a kind of 2 stage approach: train single frame model on PL and then add temporal mixing a and trained on the provided data (no freezing). Even If I trained 5 single frame distinct models models, the overall diversity was quite bad, and this model was not very helpful for the ensemble.\n\nRegarding large movent between frames, I expect it is a bad thing because LSTM mixing does not consider spatial misalignment. If the contrail is shifted from one pixel at res/32 feature map to another, LSTM cannot track it properly because it is applied only in the temporal dimension( So displacement of clouds  by the receptive field of the model may kill the advantage of the seq model.",
              "votes": 1
            },
            {
              "id": 2414200,
              "postDate": "2023-08-29T13:14:19.903Z",
              "content": "<p>Can you tell me how to reproduce these coat models? It seems that even hugging faces and <a href=\"https://www.semanticscholar.org/\" target=\"_blank\">https://www.semanticscholar.org/</a> cannot find them</p>",
              "rawMarkdown": "Can you tell me how to reproduce these coat models? It seems that even hugging faces and https://www.semanticscholar.org/ cannot find them"
            }
          ]
        }
      ]
    },
    {
      "id": 2388583,
      "postDate": "2023-08-13T14:02:45.770Z",
      "content": "<p>Congratulations on your 2nd place and thanks for the great writeup!<br>\nI'm interested in temporal mixing by LSTM. Do you stack each pixel to a batch dimension as input to LSTM, i.e., (n_batch * H * W, n_frame, C)?</p>",
      "rawMarkdown": "Congratulations on your 2nd place and thanks for the great writeup!\nI'm interested in temporal mixing by LSTM. Do you stack each pixel to a batch dimension as input to LSTM, i.e., (n_batch * H * W, n_frame, C)?",
      "votes": 3,
      "replies": [
        {
          "id": 2389379,
          "postDate": "2023-08-14T02:42:42.380Z",
          "content": "<p>Thanks. Yes, exactly. I did it only at the last two, i.e. res/16 and res/32, feature maps because of too large change between input frames.</p>",
          "rawMarkdown": "Thanks. Yes, exactly. I did it only at the last two, i.e. res/16 and res/32, feature maps because of too large change between input frames.",
          "votes": 3,
          "replies": [
            {
              "id": 2389535,
              "postDate": "2023-08-14T05:48:39.370Z",
              "content": "<p>I see. That makes sense, and that's why I used conv3d temporal mixing. Did LSTM and attention perform much better than conv-based temporal mixing?</p>",
              "rawMarkdown": "I see. That makes sense, and that's why I used conv3d temporal mixing. Did LSTM and attention perform much better than conv-based temporal mixing?",
              "votes": 3
            },
            {
              "id": 2390434,
              "postDate": "2023-08-14T14:54:35.713Z",
              "content": "<p>I expect LSTM should perform pretty similarly to conv temporal mixing because it shares the same characteristics: mixing assuming that features are spatially aligned. 3d conv may have slightly better robustness to frame misalignment because you are looking spatially during mixing. But it may be addressed by the previous 2d convs, i.e. 2d+1d ~ 3d. LSTM may have slightly better behavior for longer seq. Unfortunately, I didn't experiment with conv mixing, and I do not remember if my teammates did anything on that.</p>\n<p>Regarding transformer mixing, its idea is to do intrinsic image registration for shifts larger than ~32 pixels. Conceptually it is different from LSTM/conv, and that's why it was option 2 I investigated. Unfortunately, it performed worse, which may be a results of (1) the difficulty of training such mixing model, or (2) too large misalignments are rare enough, and they are well compensated by 2d convs before and after the mixing.</p>",
              "rawMarkdown": "I expect LSTM should perform pretty similarly to conv temporal mixing because it shares the same characteristics: mixing assuming that features are spatially aligned. 3d conv may have slightly better robustness to frame misalignment because you are looking spatially during mixing. But it may be addressed by the previous 2d convs, i.e. 2d+1d ~ 3d. LSTM may have slightly better behavior for longer seq. Unfortunately, I didn't experiment with conv mixing, and I do not remember if my teammates did anything on that.\n\nRegarding transformer mixing, its idea is to do intrinsic image registration for shifts larger than ~32 pixels. Conceptually it is different from LSTM/conv, and that's why it was option 2 I investigated. Unfortunately, it performed worse, which may be a results of (1) the difficulty of training such mixing model, or (2) too large misalignments are rare enough, and they are well compensated by 2d convs before and after the mixing.",
              "votes": 2
            },
            {
              "id": 2390653,
              "postDate": "2023-08-14T17:07:37.523Z",
              "content": "<p>In which sense features are not spatially aligned in this case? When you feed the LSTM with a sequence of images/frames, each pixel represent the same location on all images/frames of the sequence, right?</p>",
              "rawMarkdown": "In which sense features are not spatially aligned in this case? When you feed the LSTM with a sequence of images/frames, each pixel represent the same location on all images/frames of the sequence, right?"
            },
            {
              "id": 2390731,
              "postDate": "2023-08-14T17:41:03.153Z",
              "content": "<p>The simplistic spatial location here is not meaningful: you try to track contrails and clouds with the seq model. Cloud movement between frames is huge, and images are very far away from being properly registered to track clouds( </p>",
              "rawMarkdown": "The simplistic spatial location here is not meaningful: you try to track contrails and clouds with the seq model. Cloud movement between frames is huge, and images are very far away from being properly registered to track clouds( ",
              "votes": 2
            },
            {
              "id": 2392302,
              "postDate": "2023-08-15T15:31:24.950Z",
              "content": "<p>Thanks for the detailed response! I learned a lot.</p>",
              "rawMarkdown": "Thanks for the detailed response! I learned a lot."
            }
          ]
        }
      ]
    },
    {
      "id": 2387504,
      "postDate": "2023-08-12T17:51:03.897Z",
      "content": "<p>Congratulations! A very unique solution.</p>\n<p>Quick question. Sorry, if I ask, but I don't understand how a transformer is used as a mixer here. Does it have anything to do with the MLP mixer idea or are you just passing multiple frames to a transformer? </p>\n<p>Also, waiting for you to release the code. Thanks for the solution.</p>",
      "rawMarkdown": "Congratulations! A very unique solution.\n\nQuick question. Sorry, if I ask, but I don't understand how a transformer is used as a mixer here. Does it have anything to do with the MLP mixer idea or are you just passing multiple frames to a transformer? \n\nAlso, waiting for you to release the code. Thanks for the solution.",
      "votes": 3,
      "replies": [
        {
          "id": 2387863,
          "postDate": "2023-08-13T03:15:44.540Z",
          "content": "<p>It is just (1) flattening the feature map into a list of tokens for res/16 and res/12 maps; then (2) adding encoding to represent the time for each considered feature map in the sequence; (3) rearranging all sequences for a given image series into one long sequence; applying transformer with q=considered frame only seq and k,v=concatenated seq; (4) rearranging the thing back and doing standard U block upsampling.<br>\nNow only inference code is released, but you can check the definition of CoaT_UT model <a href=\"https://www.kaggle.com/datasets/iafoss/contrails-model-def1\" target=\"_blank\">here</a>.</p>",
          "rawMarkdown": "It is just (1) flattening the feature map into a list of tokens for res/16 and res/12 maps; then (2) adding encoding to represent the time for each considered feature map in the sequence; (3) rearranging all sequences for a given image series into one long sequence; applying transformer with q=considered frame only seq and k,v=concatenated seq; (4) rearranging the thing back and doing standard U block upsampling.\nNow only inference code is released, but you can check the definition of CoaT_UT model [here](https://www.kaggle.com/datasets/iafoss/contrails-model-def1).",
          "votes": 2,
          "replies": [
            {
              "id": 2388020,
              "postDate": "2023-08-13T05:47:08.270Z",
              "content": "<p>Wonderful. Thank you for clarifying my doubt and sharing the definition of the model.</p>",
              "rawMarkdown": "Wonderful. Thank you for clarifying my doubt and sharing the definition of the model."
            }
          ]
        }
      ]
    },
    {
      "id": 2385782,
      "postDate": "2023-08-11T14:46:02.233Z",
      "content": "<p>Congratulations to you three and thanks a lot for such a detailed summary of your solution. It is incredible the level of the top spot solutions, it makes me feel dumb and happy at the same time 😄:</p>\n<blockquote>\n  <p>The best mixing strategy based on our experiments was using LSTM. Unfortunately, LSTM assumes spatial alignment of features, and we tried to use a transformer applied to flattened token sequence […].</p>\n</blockquote>\n<ul>\n<li><p>Maybe a dumb questions but, if LSTM was better based on your CV results, why don't continue with it despite of the spatial misalignment of features?</p></li>\n<li><p>Do you have any method to speed up experimentation in competitions like this one where training takes a lot of time? For example, when trying new things when you have to train your models for 100-200 epochs.</p></li>\n</ul>\n<blockquote>\n  <p>One of the key things for this particular data (pixel accuracy of masks is needed) was to avoid using torch.resize at the end. Typical segmentation models do x4 linear resizing of masks to get the same mask res as the input. So, I gave x2 larger input image and added a pixel unshuffle x2 upscaling block to the model instead of typical final convolution+resizing.</p>\n</blockquote>\n<ul>\n<li>I am trying to fully get what you mentioned above, let me write what I understood:</li>\n</ul>\n<ol>\n<li>You want use an input image size higher than original (let's imagine 512x512) but resizing the mask to 512x512 via linear interpolation is not desirable because introduces inaccuracies.</li>\n<li>Then you keep the mask size as 256x256 but you have to transform the output of the decoder (512x512) to 256x256 but using torch.resize is not desirable neither.</li>\n<li>The solution is doing pixel unshuffle at the end, which transforms (BS, C, 512, 512) to (BS, C*2, 256, 256). Is that correct?</li>\n</ol>\n<p>Thanks!</p>",
      "rawMarkdown": "Congratulations to you three and thanks a lot for such a detailed summary of your solution. It is incredible the level of the top spot solutions, it makes me feel dumb and happy at the same time 😄:\n\n>The best mixing strategy based on our experiments was using LSTM. Unfortunately, LSTM assumes spatial alignment of features, and we tried to use a transformer applied to flattened token sequence [...].\n\n- Maybe a dumb questions but, if LSTM was better based on your CV results, why don't continue with it despite of the spatial misalignment of features?\n\n- Do you have any method to speed up experimentation in competitions like this one where training takes a lot of time? For example, when trying new things when you have to train your models for 100-200 epochs.\n\n>One of the key things for this particular data (pixel accuracy of masks is needed) was to avoid using torch.resize at the end. Typical segmentation models do x4 linear resizing of masks to get the same mask res as the input. So, I gave x2 larger input image and added a pixel unshuffle x2 upscaling block to the model instead of typical final convolution+resizing.\n\n- I am trying to fully get what you mentioned above, let me write what I understood:\n1. You want use an input image size higher than original (let's imagine 512x512) but resizing the mask to 512x512 via linear interpolation is not desirable because introduces inaccuracies.\n2. Then you keep the mask size as 256x256 but you have to transform the output of the decoder (512x512) to 256x256 but using torch.resize is not desirable neither.\n3. The solution is doing pixel unshuffle at the end, which transforms (BS, C, 512, 512) to (BS, C*2, 256, 256). Is that correct?\n\nThanks!\n",
      "votes": 3,
      "replies": [
        {
          "id": 2385934,
          "postDate": "2023-08-11T16:30:38.503Z",
          "content": "<p>LSTM worked the best for me, but we kept other things for diversity. My training was 24 epochs long, so it took just ~2 hours on 512 input, but you could accelerate the things at initial experiments using smaller res. Seq models took probably 8 hours. The issue, though, was the results may be noisy sometimes, and 5 folds are needed to make a relatable concussion on the experiment. If you have to train 100-200 epochs, you can do initial dev on shorter training and then upscale the things for the best setup.</p>\n<p>The typical output of segmentation models is at x4 lower res, which then is linearly upsampled. So if u give x4 input you do not need to do upsampling of masks at the end. For x2 input, the output wd be 128x128, and additional x2 upscaling with pixel shuffle block (instead of linear upsampling) is needed.</p>",
          "rawMarkdown": "LSTM worked the best for me, but we kept other things for diversity. My training was 24 epochs long, so it took just ~2 hours on 512 input, but you could accelerate the things at initial experiments using smaller res. Seq models took probably 8 hours. The issue, though, was the results may be noisy sometimes, and 5 folds are needed to make a relatable concussion on the experiment. If you have to train 100-200 epochs, you can do initial dev on shorter training and then upscale the things for the best setup.\n\nThe typical output of segmentation models is at x4 lower res, which then is linearly upsampled. So if u give x4 input you do not need to do upsampling of masks at the end. For x2 input, the output wd be 128x128, and additional x2 upscaling with pixel shuffle block (instead of linear upsampling) is needed.",
          "votes": 3
        }
      ]
    },
    {
      "id": 2382865,
      "postDate": "2023-08-10T04:37:15.270Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a>! Did you train on upsampled masks (i.e. 512x512 or 1024x1024), or did you downsample the predictions back to 256x256 after your conv/pixelshuffle upsampling? Regardless, thank you for sharing your solution.</p>",
      "rawMarkdown": "Congrats @iafoss! Did you train on upsampled masks (i.e. 512x512 or 1024x1024), or did you downsample the predictions back to 256x256 after your conv/pixelshuffle upsampling? Regardless, thank you for sharing your solution.",
      "votes": 3,
      "replies": [
        {
          "id": 2382883,
          "postDate": "2023-08-10T04:52:25.517Z",
          "content": "<p>I just didn't up-sample the masks in the model keeping 256 size as the output.</p>",
          "rawMarkdown": "I just didn't up-sample the masks in the model keeping 256 size as the output.",
          "votes": 6,
          "replies": [
            {
              "id": 2383585,
              "postDate": "2023-08-10T13:05:30.923Z",
              "content": "<p>How did you do that? simple torch.resize at the end ?</p>\n<p>I tried to add an extra conv layer in the final head so that the model outputs 256 even with images 512 as inputs, also tried to remove the last upsample block but without much success.</p>",
              "rawMarkdown": "How did you do that? simple torch.resize at the end ?\n\nI tried to add an extra conv layer in the final head so that the model outputs 256 even with images 512 as inputs, also tried to remove the last upsample block but without much success.",
              "votes": 1
            },
            {
              "id": 2383739,
              "postDate": "2023-08-10T14:36:32.400Z",
              "content": "<p>One of the key things for this particular data (pixel accuracy of masks is needed) was to avoid using torch.resize at the end. Typical segmentation models do x4 linear resizing of masks to get the same mask res as the input. So, I gave x2 larger input image and added a pixel unshuffle x2 upscaling block to the model instead of typical final convolution+resizing.</p>",
              "rawMarkdown": "One of the key things for this particular data (pixel accuracy of masks is needed) was to avoid using torch.resize at the end. Typical segmentation models do x4 linear resizing of masks to get the same mask res as the input. So, I gave x2 larger input image and added a pixel unshuffle x2 upscaling block to the model instead of typical final convolution+resizing.",
              "votes": 4
            },
            {
              "id": 2383767,
              "postDate": "2023-08-10T14:52:41.417Z",
              "content": "<p>Ok thank you! But what is a \"pixel unshuffle x2 upscaling block\" exactly? 😅</p>",
              "rawMarkdown": "Ok thank you! But what is a \"pixel unshuffle x2 upscaling block\" exactly? 😅",
              "votes": 2
            },
            {
              "id": 2383806,
              "postDate": "2023-08-10T15:07:37.300Z",
              "content": "<p>You just unshuffle the input from Bx4CxHxW to BxCx2Hx2W in the way that 4 sequential values from the channel dimension go into 2x2 pixels. And I have conv + LayerNorm2D + GELU + conv after and before this shuffle. We will share the code a bit later.</p>",
              "rawMarkdown": "You just unshuffle the input from Bx4CxHxW to BxCx2Hx2W in the way that 4 sequential values from the channel dimension go into 2x2 pixels. And I have conv + LayerNorm2D + GELU + conv after and before this shuffle. We will share the code a bit later.",
              "votes": 4
            },
            {
              "id": 2385946,
              "postDate": "2023-08-11T16:37:49.640Z",
              "content": "<p>Quite tricky ! would you have a rough estimate on the impact of this wrt. to a simpler torch.resize ? </p>",
              "rawMarkdown": "Quite tricky ! would you have a rough estimate on the impact of this wrt. to a simpler torch.resize ? ",
              "votes": 1
            },
            {
              "id": 2394440,
              "postDate": "2023-08-16T22:32:55.410Z",
              "content": "<p>I didn't check it in this competition, but the negative effect of linear upscaling should be quite strong here because you need to reproduce several pixel-wide masks with pixel accuracy(</p>",
              "rawMarkdown": "I didn't check it in this competition, but the negative effect of linear upscaling should be quite strong here because you need to reproduce several pixel-wide masks with pixel accuracy(",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2396774,
      "postDate": "2023-08-18T13:12:36.827Z",
      "content": "<p>Congrats for the 2nd place guys! <br>\nThanks for the awesome write-up, we have to study and learn from you. Especially, starting from how you describe the research challenges and explaining your way of thinking is very crucial.</p>\n<p>1) threshold = 0.5 or tuned in val set ?<br>\n2) BTW your very 1st submission that scored ~0.713 on public (If I recall correct) was the Coat_ULSTM single ? </p>",
      "rawMarkdown": "Congrats for the 2nd place guys! \nThanks for the awesome write-up, we have to study and learn from you. Especially, starting from how you describe the research challenges and explaining your way of thinking is very crucial.\n\n1) threshold = 0.5 or tuned in val set ?\n2) BTW your very 1st submission that scored ~0.713 on public (If I recall correct) was the Coat_ULSTM single ? \n",
      "votes": 1,
      "replies": [
        {
          "id": 2397530,
          "postDate": "2023-08-19T05:26:11.273Z",
          "content": "<p>Thanks, congratulations on getting gold too.</p>\n<p>We selected the threshold based on CV, but it is very close to 0.5, like 0.48. Lovasz part of the loss was helpful to reduce the dependence of the dice on the threshold selection.</p>\n<p>My early sub in this competition got 0.71614 public LB and it was a combination of Coat and NeXtViT single-frame models. It was rather a coincidence, and CV was just 0.700, and private LB is 0.70323. The next sub was Seq CoaT+NeXtViT LSTM, which got 0.71205 public LB, while CV is much higher of 0.7039, and private LB is 0.71488. All those subs are just single fold, later we switched to 5 folds. In this competition, we didn't really consider public LB, and relay only on CV on eval set. </p>",
          "rawMarkdown": "Thanks, congratulations on getting gold too.\n\nWe selected the threshold based on CV, but it is very close to 0.5, like 0.48. Lovasz part of the loss was helpful to reduce the dependence of the dice on the threshold selection.\n\nMy early sub in this competition got 0.71614 public LB and it was a combination of Coat and NeXtViT single-frame models. It was rather a coincidence, and CV was just 0.700, and private LB is 0.70323. The next sub was Seq CoaT+NeXtViT LSTM, which got 0.71205 public LB, while CV is much higher of 0.7039, and private LB is 0.71488. All those subs are just single fold, later we switched to 5 folds. In this competition, we didn't really consider public LB, and relay only on CV on eval set. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 2396381,
      "postDate": "2023-08-18T08:17:08.013Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> on winning the second place and writing this fabulous writeup! </p>\n<p>Can you also please explain the temporal mixing and what type of preprocessing did you use ?</p>",
      "rawMarkdown": "Congratulations @iafoss on winning the second place and writing this fabulous writeup! \n\nCan you also please explain the temporal mixing and what type of preprocessing did you use ?",
      "votes": 1,
      "replies": [
        {
          "id": 2397520,
          "postDate": "2023-08-19T05:13:45.747Z",
          "content": "<p>Thanks. I updated the write-up with more details. Regarding preprocessing, I do not think that we did anything spatial in addition to what we listed in the write-up.</p>",
          "rawMarkdown": "Thanks. I updated the write-up with more details. Regarding preprocessing, I do not think that we did anything spatial in addition to what we listed in the write-up."
        }
      ]
    },
    {
      "id": 2394821,
      "postDate": "2023-08-17T06:32:51.023Z",
      "content": "<p>Hi, Could you please mention what hardware you have used to train your models?</p>",
      "rawMarkdown": "Hi, Could you please mention what hardware you have used to train your models?",
      "votes": 1,
      "replies": [
        {
          "id": 2396079,
          "postDate": "2023-08-18T03:47:16.073Z",
          "content": "<p>2x4090 in my experiments</p>",
          "rawMarkdown": "2x4090 in my experiments",
          "votes": 1
        }
      ]
    },
    {
      "id": 2393084,
      "postDate": "2023-08-16T05:51:46.967Z",
      "content": "<p>Congratulations on your 2nd place and thanks for the great writeup!</p>",
      "rawMarkdown": "Congratulations on your 2nd place and thanks for the great writeup!",
      "votes": 1
    },
    {
      "id": 2383714,
      "postDate": "2023-08-10T14:24:46.777Z",
      "content": "<p>Congratulations! Nice way to do the work.</p>",
      "rawMarkdown": "Congratulations! Nice way to do the work.",
      "votes": 1
    },
    {
      "id": 2383622,
      "postDate": "2023-08-10T13:26:51.343Z",
      "content": "<p>Thanks for the write-up Iafoss and congrats on your impressive result.<br>\nA quick question. Do you train your transformer-based models from scratch or do you modify the pre-trained model so as to make it accept a different resolution? </p>",
      "rawMarkdown": "Thanks for the write-up Iafoss and congrats on your impressive result.\nA quick question. Do you train your transformer-based models from scratch or do you modify the pre-trained model so as to make it accept a different resolution? \n",
      "votes": 1,
      "replies": [
        {
          "id": 2383762,
          "postDate": "2023-08-10T14:50:47.597Z",
          "content": "<p>SAM ViT encoder works with 1024x1024 input, and we just performed x4 upscaling. For CoaT and NeXtViT the resolution is not fixed, similar to conv models. Training a transformer from scratch is quite an unfavorable thing, and if you asked how to use different res in ViT like models, the typical way is just linear resizing of the pos enc to the res u are interested in. After small fine-tuning the model works quite well.</p>",
          "rawMarkdown": "SAM ViT encoder works with 1024x1024 input, and we just performed x4 upscaling. For CoaT and NeXtViT the resolution is not fixed, similar to conv models. Training a transformer from scratch is quite an unfavorable thing, and if you asked how to use different res in ViT like models, the typical way is just linear resizing of the pos enc to the res u are interested in. After small fine-tuning the model works quite well.",
          "votes": 5
        }
      ]
    },
    {
      "id": 2383539,
      "postDate": "2023-08-10T12:31:01.473Z",
      "content": "<p>Very detailed solution, thanks for the write-up. This is also the first time I hear about the <code>Over9000</code> optimizer. Is that this repo: <a href=\"https://github.com/mgrankin/over9000\" target=\"_blank\">https://github.com/mgrankin/over9000</a> ?</p>",
      "rawMarkdown": "Very detailed solution, thanks for the write-up. This is also the first time I hear about the `Over9000` optimizer. Is that this repo: https://github.com/mgrankin/over9000 ?",
      "votes": 1,
      "replies": [
        {
          "id": 2383779,
          "postDate": "2023-08-10T14:56:13.910Z",
          "content": "<p>Yep. Previously I used it a lot because of the good compatibility of this optimizer with ssl_ResNeXt50 weights in fine-tuning, but later, working with transformers, I mostly switched to AdamW. It was interesting to see that CoaT seems to work quite well with this optimizer.</p>",
          "rawMarkdown": "Yep. Previously I used it a lot because of the good compatibility of this optimizer with ssl_ResNeXt50 weights in fine-tuning, but later, working with transformers, I mostly switched to AdamW. It was interesting to see that CoaT seems to work quite well with this optimizer.",
          "votes": 3,
          "replies": [
            {
              "id": 2383859,
              "postDate": "2023-08-10T15:44:19.370Z",
              "content": "<p>Thanks for the insights. I will add this to my tools. Personally I mostly use <strong>AdamW</strong> or <strong>Adam</strong> since they work well most of the time (despite all the recent refinements 😁).</p>",
              "rawMarkdown": "Thanks for the insights. I will add this to my tools. Personally I mostly use **AdamW** or **Adam** since they work well most of the time (despite all the recent refinements 😁)."
            }
          ]
        }
      ]
    },
    {
      "id": 2383072,
      "postDate": "2023-08-10T07:02:01.027Z",
      "content": "<p>Congratulations!<br>\nCan you please show the exact formula for the dice Lovasz loss?<br>\nHow much did you gain from using CutMix strategy?<br>\nI can't wait to see your code.</p>",
      "rawMarkdown": "Congratulations!\nCan you please show the exact formula for the dice Lovasz loss?\nHow much did you gain from using CutMix strategy?\nI can't wait to see your code.\n",
      "votes": 1,
      "replies": [
        {
          "id": 2383197,
          "postDate": "2023-08-10T08:26:08.560Z",
          "content": "<p>Cutmix gave +0.005 for v2-s models. It did not work that well for other models though<br>\nWe will release the code soon, with the implementation of the loss.</p>",
          "rawMarkdown": "Cutmix gave +0.005 for v2-s models. It did not work that well for other models though\nWe will release the code soon, with the implementation of the loss.",
          "votes": 3
        }
      ]
    },
    {
      "id": 2383032,
      "postDate": "2023-08-10T06:31:34.780Z",
      "content": "<p>thanks for sharing the solution this quickly. What was your postprocessing? single thresholding tuned on CV?</p>",
      "rawMarkdown": "thanks for sharing the solution this quickly. What was your postprocessing? single thresholding tuned on CV?",
      "votes": 1,
      "replies": [
        {
          "id": 2383039,
          "postDate": "2023-08-10T06:35:36.767Z",
          "content": "<p>yep, it worked the best. Since global dice is considered in this competition, I don't think that post-processing may really work well here.</p>",
          "rawMarkdown": "yep, it worked the best. Since global dice is considered in this competition, I don't think that post-processing may really work well here.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2383016,
      "postDate": "2023-08-10T06:22:39.653Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> on 2nd and another gold!  Thank you for the thoughtful, detailed discussion of your solution, and for including a couple things that didn't work.  It's so nice to read such a clear writeup!</p>",
      "rawMarkdown": "Congratulations @iafoss on 2nd and another gold!  Thank you for the thoughtful, detailed discussion of your solution, and for including a couple things that didn't work.  It's so nice to read such a clear writeup!",
      "votes": 1,
      "replies": [
        {
          "id": 2384464,
          "postDate": "2023-08-11T01:30:10.987Z",
          "content": "<p><a href=\"https://www.kaggle.com/socated\" target=\"_blank\">@socated</a> you are very welcome</p>",
          "rawMarkdown": "@socated you are very welcome",
          "votes": 1
        }
      ]
    },
    {
      "id": 2382851,
      "postDate": "2023-08-10T04:27:21.683Z",
      "content": "<p>Kudos <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> on this achievement. Thanks for the write up.</p>",
      "rawMarkdown": "Kudos @iafoss on this achievement. Thanks for the write up.",
      "votes": 1,
      "replies": [
        {
          "id": 2384467,
          "postDate": "2023-08-11T01:30:45.743Z",
          "content": "<p>Thank you so much</p>",
          "rawMarkdown": "Thank you so much"
        }
      ]
    },
    {
      "id": 2382827,
      "postDate": "2023-08-10T04:10:50.603Z",
      "content": "<p>Congratulations for winning the 2nd place.<br>\nWe appreciate your sharing info about the solution. </p>",
      "rawMarkdown": "Congratulations for winning the 2nd place.\nWe appreciate your sharing info about the solution. ",
      "votes": 1,
      "replies": [
        {
          "id": 2384472,
          "postDate": "2023-08-11T01:32:26.420Z",
          "content": "<p>You are welcome</p>",
          "rawMarkdown": "You are welcome"
        }
      ]
    },
    {
      "id": 2382815,
      "postDate": "2023-08-10T03:57:01.853Z",
      "content": "<p>That is very interesting. Congratulations with the second place!</p>",
      "rawMarkdown": "That is very interesting. Congratulations with the second place!",
      "votes": 1
    },
    {
      "id": 2413521,
      "postDate": "2023-08-29T01:21:19.613Z",
      "content": "<p>How can you get you backbone (CoaT, NeXtViT, SAM-B, and tf_efficientnetv2_s ) ? Did you use some automatic tools? Thank you very much. </p>",
      "rawMarkdown": "How can you get you backbone (CoaT, NeXtViT, SAM-B, and tf_efficientnetv2_s ) ? Did you use some automatic tools? Thank you very much. ",
      "votes": 2,
      "replies": [
        {
          "id": 2414796,
          "postDate": "2023-08-30T00:48:23.453Z",
          "content": "<p>For tf_efficientnetv2_s  you can use timm, which is super convenient. For others, that are not shared through timm, I just go to the corresponding github, grab files with the model definition and related things, modify it a bit in some cases to eliminate unnecessary dependences like mmseg, download weights from the repo, and just use the model.</p>",
          "rawMarkdown": "For tf_efficientnetv2_s  you can use timm, which is super convenient. For others, that are not shared through timm, I just go to the corresponding github, grab files with the model definition and related things, modify it a bit in some cases to eliminate unnecessary dependences like mmseg, download weights from the repo, and just use the model.",
          "votes": 2,
          "replies": [
            {
              "id": 2417960,
              "postDate": "2023-09-01T03:10:54.947Z",
              "content": "<p>It has inspired me a lot, thank you！</p>",
              "rawMarkdown": "It has inspired me a lot, thank you！",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2394585,
      "postDate": "2023-08-17T03:25:34.907Z",
      "content": "<p><a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> Congratulations for 2nd place. </p>\n<p>I have problem with training transformer T-Mixer because it is easy to get loss equal \"nan\".<br>\nDid you experience same problem? And could you share some tips you used on stabilizing training?</p>",
      "rawMarkdown": "@iafoss Congratulations for 2nd place. \n\nI have problem with training transformer T-Mixer because it is easy to get loss equal \"nan\".\nDid you experience same problem? And could you share some tips you used on stabilizing training?",
      "votes": 2,
      "replies": [
        {
          "id": 2394754,
          "postDate": "2023-08-17T05:44:02.350Z",
          "content": "<p>Thanks. Yep, training with it is less stable than LSTM. With CoaT model everything was ok, but the spread of the results between folds is large. For NeXtViT, meanwhile, I could not finish the training from scratch even if tried to play with lr, etc. Probably proper init may resolve the problem. Meanwhile, in the case of using a model pre-trained on single frame setup without mixing (both encoder and decoder), training of the transformer mixing is quite more stable, and I didn't have any issues during training. <br>\nI used this setup for the case of PL pertaining. However, even if I pretrained the single frame model with different seeds for each fold, the fine-tuned models had less diversity (in comparison to training from scratch), and these models were not so useful for the ensemble in the end.</p>",
          "rawMarkdown": "Thanks. Yep, training with it is less stable than LSTM. With CoaT model everything was ok, but the spread of the results between folds is large. For NeXtViT, meanwhile, I could not finish the training from scratch even if tried to play with lr, etc. Probably proper init may resolve the problem. Meanwhile, in the case of using a model pre-trained on single frame setup without mixing (both encoder and decoder), training of the transformer mixing is quite more stable, and I didn't have any issues during training. \nI used this setup for the case of PL pertaining. However, even if I pretrained the single frame model with different seeds for each fold, the fine-tuned models had less diversity (in comparison to training from scratch), and these models were not so useful for the ensemble in the end.",
          "votes": 3,
          "replies": [
            {
              "id": 2394771,
              "postDate": "2023-08-17T05:55:23.233Z",
              "content": "<blockquote>\n  <p>Meanwhile, in the case of using a model pre-trained on single frame setup without mixing (both encoder and decoder), training of the transformer mixing is quite more stable, and I didn't have any issues during training.</p>\n</blockquote>\n<p>Thanks for that tips. </p>\n<p>Unfortunately, I have trouble even with CoaT as backend. Conv 3D mixier is quite stable, but not for naive transformer T-mixier, nor T-mixier with 3d-adapted CoaT blocks.<br>\nMaybe I should tune learning rate or other parameters that rerate to regularization.</p>",
              "rawMarkdown": "> Meanwhile, in the case of using a model pre-trained on single frame setup without mixing (both encoder and decoder), training of the transformer mixing is quite more stable, and I didn't have any issues during training.\n\nThanks for that tips. \n\nUnfortunately, I have trouble even with CoaT as backend. Conv 3D mixier is quite stable, but not for naive transformer T-mixier, nor T-mixier with 3d-adapted CoaT blocks.\nMaybe I should tune learning rate or other parameters that rerate to regularization."
            },
            {
              "id": 2396078,
              "postDate": "2023-08-18T03:46:37.783Z",
              "content": "<p>The origin of this unstable behavior, to my understanding, is the label noise(<br>\nI was using 3e-4 lr, but what was more important in my case is Over9000 optimizer with epsilon=1e-4. As I remember, this optimizer not only worked better for CoaT but also gave more stable training than AdamW.</p>",
              "rawMarkdown": "The origin of this unstable behavior, to my understanding, is the label noise(\nI was using 3e-4 lr, but what was more important in my case is Over9000 optimizer with epsilon=1e-4. As I remember, this optimizer not only worked better for CoaT but also gave more stable training than AdamW.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2388436,
      "postDate": "2023-08-13T11:42:27.283Z",
      "content": "<p>Congratulation! Nice Work man.</p>",
      "rawMarkdown": "Congratulation! Nice Work man.\n",
      "votes": 2
    },
    {
      "id": 2388215,
      "postDate": "2023-08-13T08:31:39.507Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a>! You did a great job man. </p>\n<p>Your solution is amazing, very detailed.</p>",
      "rawMarkdown": "Congratulations @iafoss! You did a great job man. \n\nYour solution is amazing, very detailed.",
      "votes": 2,
      "replies": [
        {
          "id": 2388575,
          "postDate": "2023-08-13T13:54:05.457Z",
          "content": "<p>Thanks.                          </p>",
          "rawMarkdown": "Thanks.                          ",
          "votes": 2
        }
      ]
    },
    {
      "id": 2383682,
      "postDate": "2023-08-10T14:10:29.017Z",
      "content": "<p>May I ask, what configuration of machine did you use for these experiments? These experiments seem to have very high requirements for the machine</p>",
      "rawMarkdown": "May I ask, what configuration of machine did you use for these experiments? These experiments seem to have very high requirements for the machine",
      "votes": 2,
      "replies": [
        {
          "id": 2383748,
          "postDate": "2023-08-10T14:45:50.633Z",
          "content": "<p>For me it was 2x4090 machine, quite fast cards for deep learning but VRAM is an issue, like always. <a href=\"https://www.kaggle.com/drhabib\" target=\"_blank\">@drhabib</a> used 2xA6000, and <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> used his V100 cluster.</p>",
          "rawMarkdown": "For me it was 2x4090 machine, quite fast cards for deep learning but VRAM is an issue, like always. @drhabib used 2xA6000, and @theoviel used his V100 cluster.",
          "votes": 3,
          "replies": [
            {
              "id": 2384506,
              "postDate": "2023-08-11T01:46:40.203Z",
              "content": "<p>Your answer has been very enlightening to me, and I am very grateful</p>",
              "rawMarkdown": "Your answer has been very enlightening to me, and I am very grateful",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2383131,
      "postDate": "2023-08-10T07:35:36.013Z",
      "content": "<p>Can you please share how codewise you combine coat encoder and Unet decoder?<br>\nCoat (coatnet, right?) encoder is not in SMP repo, so you did it by yourself, right?</p>",
      "rawMarkdown": "Can you please share how codewise you combine coat encoder and Unet decoder?\nCoat (coatnet, right?) encoder is not in SMP repo, so you did it by yourself, right?",
      "votes": 2,
      "replies": [
        {
          "id": 2383193,
          "postDate": "2023-08-10T08:23:17.587Z",
          "content": "<p>Exactly ! We will release the code later.</p>",
          "rawMarkdown": "Exactly ! We will release the code later.",
          "votes": 6,
          "replies": [
            {
              "id": 2383222,
              "postDate": "2023-08-10T08:49:09.083Z",
              "content": "<p>Thanks for your sharing. There are a few confusion to me in your details (2).</p>",
              "rawMarkdown": "Thanks for your sharing. There are a few confusion to me in your details (2)."
            }
          ]
        }
      ]
    },
    {
      "id": 2382830,
      "postDate": "2023-08-10T04:14:23.157Z",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a>  congratulations for amazing result . Can you please give some more details on the temporal mixing ? We used only single frame model ..so curious to know about this solution.  </p>",
      "rawMarkdown": "Hello @iafoss  congratulations for amazing result . Can you please give some more details on the temporal mixing ? We used only single frame model ..so curious to know about this solution.  ",
      "votes": 2,
      "replies": [
        {
          "id": 2382888,
          "postDate": "2023-08-10T04:54:40.670Z",
          "content": "<p>We will share the code a bit later, but the thing I considered is mixing at low res feature maps because images are poorly alighted (time between images is too large). Congratulations with getting gold!</p>",
          "rawMarkdown": "We will share the code a bit later, but the thing I considered is mixing at low res feature maps because images are poorly alighted (time between images is too large). Congratulations with getting gold!",
          "votes": 5
        }
      ]
    },
    {
      "id": 2382812,
      "postDate": "2023-08-10T03:56:26.883Z",
      "content": "<p>Congratulations on the strong finish! CoaT is my favorite backbone so glad to see it performs so well here.</p>",
      "rawMarkdown": "Congratulations on the strong finish! CoaT is my favorite backbone so glad to see it performs so well here.",
      "votes": 2,
      "replies": [
        {
          "id": 2386554,
          "postDate": "2023-08-12T03:20:47.713Z",
          "content": "<p>CoaT is a really good alternative to Segformer when a restrictive license is not acceptable. A combination of global attention and consideration of details at high-resolution feature maps (in contrast to ViT architectures) is something one doesn't see often. Meanwhile, in this particular competition, it was very important. </p>",
          "rawMarkdown": "CoaT is a really good alternative to Segformer when a restrictive license is not acceptable. A combination of global attention and consideration of details at high-resolution feature maps (in contrast to ViT architectures) is something one doesn't see often. Meanwhile, in this particular competition, it was very important. ",
          "votes": 4
        }
      ]
    },
    {
      "id": 2382792,
      "postDate": "2023-08-10T03:40:47.567Z",
      "content": "<p>congrats <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a>  for 2nd place. Try to get 1st place</p>",
      "rawMarkdown": "congrats @iafoss  for 2nd place. Try to get 1st place",
      "replies": [
        {
          "id": 2382795,
          "postDate": "2023-08-10T03:44:08.957Z",
          "content": "<p>Thanks so much. It happens for us in the second time during last several months when we drop from 1st place at public to the 2nd place at private((( Getting the first place is quite a difficult thing.</p>",
          "rawMarkdown": "Thanks so much. It happens for us in the second time during last several months when we drop from 1st place at public to the 2nd place at private((( Getting the first place is quite a difficult thing.",
          "votes": 4
        }
      ]
    },
    {
      "id": 2394533,
      "postDate": "2023-08-17T02:02:25.790Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2382925,
      "author_name": "Yousef Rabi",
      "author_url": "",
      "post_date": "2023-08-10T05:28:09.093000",
      "content": "<p>Congratulations! Beautiful solution, thank you for sharing it with us. Looking forward to studying your code.</p>\n<p>We knew you had a unique approach that gained your team a nice separation throughout. Even during the competition I was excited to learn about it and it didn't disappoint.</p>\n<p>We had a feeling the temporal information was important and tried a bunch of stuff like 3d models and sequence models to no avail, so we dropped the idea as computation was expensive.</p>\n<p>We also had a feeling that there was important information in the individual masks and tried a bunch of stuff like using random individual masks as auxiliary signal or different classes, but the simplest approach of averaging the individual masks didn't occur to us, nice one.</p>",
      "votes": 8,
      "replies": []
    },
    {
      "id": 2382849,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-08-10T04:27:04.067000",
      "content": "<p>thanks for the writeup.</p>\n<p>i am interested in usage of tempral information. how much did it increase your CV and LB.???</p>\n<hr>\n<p>I spent the last week in fusion of multiple frams, early fusion, late fusion, etc …<br>\nI also tried optcal flow (RAFT) for registration and  augmentation (find the flow and then wrap image and mask)<br>\nnone of them really work well.<br>\ni only have 0.004~0.005 improvement when using 3 consective frames (over 2 different models i tried)<br>\nThis is worse than expected  (the paper can get 0.01 to 0.015 improvement). I wondered why????</p>\n<p>i don't think large movement  is the cause of poor performance.<br>\n(on a side note, actually you can use temporal superresolution, i.e. predict intermediate frames with frame interpolation. <a href=\"https://www.youtube.com/watch?v=MjViy6kyiqs)\" target=\"_blank\">https://www.youtube.com/watch?v=MjViy6kyiqs)</a>. Frame interpolation results may be wrong, but as long as it is \"consistent\", it can \"provide temporal features\"</p>",
      "votes": 5,
      "replies": [
        {
          "id": 2382872,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2023-08-10T04:43:18.023000",
          "content": "<p>Here are some of the results from my notes (I list public CV):<br>\nCoaT_U (single image): 0.6887+-0.0014, avr pred 0.6983 (th 0.47), LB 0.706<br>\nExCoaT_U (single image, PL): 0.6952+-0.0008, avr pred 0.7002 (th 0.49)</p>\n<p>NeXtViT_U (single image): 0.6848 +- 0.0013, avr pred 0.6917 (th 0.47), LB 0.710<br>\nExNeXtViT_U (single image, PL): 0.6904 +- 0.0006, avr pred 0.6941 (th 0.48)</p>\n<p>CoaT_ULSTM (5 image seq): 0.6960+-0.0003, avr pred 0.7039 (th 0.48), LB 0.712<br>\nEX Coat_ULSTM (5 image seq, PL): 0.7027+-0.0002, avr pred 0.7064 (th 0.49), LB 721</p>\n<p>CoaT_UT (5 image seq, transformer mixing): 0.6950+-0.0013, avr pred 0.7055 (th 0.47), LB 0.722<br>\nEX Coat_UT (5 image seq, transformer mixing, PL): 0.6997+-0.0008, avr pred 0.7040 (th 0.49), LB 0.716</p>\n<p>NeXtViT_ULSTM (5 image seq): 0.6951+-0.0004, avr pred 0.7014 (th 0.50), LB 0.713<br>\nEx NeXtViT_ULSTM (5 image seq, PL): 0.6988 +- 0.0007, avr pred 0.7025 (th 0.50)</p>\n<p>Overall <strong>sequence model gives 0.005-0.01 boost</strong>. In my case the key thing was mixing at res/32 and res/16 feature maps followed with up-sampling U blocks, like in the single frame model. I used 0-5 frames, Theo was using 3-6  frames. 8 frames didn't perform well, but I didn't run full 5 fold for this setup.</p>",
          "votes": 6,
          "replies": [
            {
              "id": 2382890,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-08-10T04:54:55.670000",
              "content": "<p>thanks for the details.</p>\n<p>i will try again later. i actually use two stage approach: first model is just trained on single frame. second models uses features from first one. both models are u-net. there is only improvement if i train e2e. it doesn't work if i freeze the first model. my mixing is done for all 4 scales. maybe that is the reason.</p>\n<p>lastly, congrats to you good ranking and score! 👍 👍 👍</p>\n<hr>\n<p>\"i don't think large movement is the cause of poor performance.\"</p>\n<p>this is because cloud/hurricane motion forecast or precipitation/rainfall prediction are also using large interval satellite images</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2382897,
              "author_name": "Iafoss",
              "author_url": "",
              "post_date": "2023-08-10T05:07:46.827000",
              "content": "<p>Thanks so much. </p>\n<p>In my experiments use of 4 scales was worse than use of 2 scales, but I didn't run this experiment in 5 folds.</p>\n<p>In PL training I used a kind of 2 stage approach: train single frame model on PL and then add temporal mixing a and trained on the provided data (no freezing). Even If I trained 5 single frame distinct models models, the overall diversity was quite bad, and this model was not very helpful for the ensemble.</p>\n<p>Regarding large movent between frames, I expect it is a bad thing because LSTM mixing does not consider spatial misalignment. If the contrail is shifted from one pixel at res/32 feature map to another, LSTM cannot track it properly because it is applied only in the temporal dimension( So displacement of clouds  by the receptive field of the model may kill the advantage of the seq model.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2414200,
              "author_name": "kongweihao",
              "author_url": "",
              "post_date": "2023-08-29T13:14:19.903000",
              "content": "<p>Can you tell me how to reproduce these coat models? It seems that even hugging faces and <a href=\"https://www.semanticscholar.org/\" target=\"_blank\">https://www.semanticscholar.org/</a> cannot find them</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2388583,
      "author_name": "knshnb",
      "author_url": "",
      "post_date": "2023-08-13T14:02:45.770000",
      "content": "<p>Congratulations on your 2nd place and thanks for the great writeup!<br>\nI'm interested in temporal mixing by LSTM. Do you stack each pixel to a batch dimension as input to LSTM, i.e., (n_batch * H * W, n_frame, C)?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2389379,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2023-08-14T02:42:42.380000",
          "content": "<p>Thanks. Yes, exactly. I did it only at the last two, i.e. res/16 and res/32, feature maps because of too large change between input frames.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2389535,
              "author_name": "knshnb",
              "author_url": "",
              "post_date": "2023-08-14T05:48:39.370000",
              "content": "<p>I see. That makes sense, and that's why I used conv3d temporal mixing. Did LSTM and attention perform much better than conv-based temporal mixing?</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2390434,
              "author_name": "Iafoss",
              "author_url": "",
              "post_date": "2023-08-14T14:54:35.713000",
              "content": "<p>I expect LSTM should perform pretty similarly to conv temporal mixing because it shares the same characteristics: mixing assuming that features are spatially aligned. 3d conv may have slightly better robustness to frame misalignment because you are looking spatially during mixing. But it may be addressed by the previous 2d convs, i.e. 2d+1d ~ 3d. LSTM may have slightly better behavior for longer seq. Unfortunately, I didn't experiment with conv mixing, and I do not remember if my teammates did anything on that.</p>\n<p>Regarding transformer mixing, its idea is to do intrinsic image registration for shifts larger than ~32 pixels. Conceptually it is different from LSTM/conv, and that's why it was option 2 I investigated. Unfortunately, it performed worse, which may be a results of (1) the difficulty of training such mixing model, or (2) too large misalignments are rare enough, and they are well compensated by 2d convs before and after the mixing.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2390653,
              "author_name": "delai50",
              "author_url": "",
              "post_date": "2023-08-14T17:07:37.523000",
              "content": "<p>In which sense features are not spatially aligned in this case? When you feed the LSTM with a sequence of images/frames, each pixel represent the same location on all images/frames of the sequence, right?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2390731,
              "author_name": "Iafoss",
              "author_url": "",
              "post_date": "2023-08-14T17:41:03.153000",
              "content": "<p>The simplistic spatial location here is not meaningful: you try to track contrails and clouds with the seq model. Cloud movement between frames is huge, and images are very far away from being properly registered to track clouds( </p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2392302,
              "author_name": "knshnb",
              "author_url": "",
              "post_date": "2023-08-15T15:31:24.950000",
              "content": "<p>Thanks for the detailed response! I learned a lot.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2387504,
      "author_name": "Swaroop Meher",
      "author_url": "",
      "post_date": "2023-08-12T17:51:03.897000",
      "content": "<p>Congratulations! A very unique solution.</p>\n<p>Quick question. Sorry, if I ask, but I don't understand how a transformer is used as a mixer here. Does it have anything to do with the MLP mixer idea or are you just passing multiple frames to a transformer? </p>\n<p>Also, waiting for you to release the code. Thanks for the solution.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2387863,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2023-08-13T03:15:44.540000",
          "content": "<p>It is just (1) flattening the feature map into a list of tokens for res/16 and res/12 maps; then (2) adding encoding to represent the time for each considered feature map in the sequence; (3) rearranging all sequences for a given image series into one long sequence; applying transformer with q=considered frame only seq and k,v=concatenated seq; (4) rearranging the thing back and doing standard U block upsampling.<br>\nNow only inference code is released, but you can check the definition of CoaT_UT model <a href=\"https://www.kaggle.com/datasets/iafoss/contrails-model-def1\" target=\"_blank\">here</a>.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2388020,
              "author_name": "Swaroop Meher",
              "author_url": "",
              "post_date": "2023-08-13T05:47:08.270000",
              "content": "<p>Wonderful. Thank you for clarifying my doubt and sharing the definition of the model.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2385782,
      "author_name": "delai50",
      "author_url": "",
      "post_date": "2023-08-11T14:46:02.233000",
      "content": "<p>Congratulations to you three and thanks a lot for such a detailed summary of your solution. It is incredible the level of the top spot solutions, it makes me feel dumb and happy at the same time 😄:</p>\n<blockquote>\n  <p>The best mixing strategy based on our experiments was using LSTM. Unfortunately, LSTM assumes spatial alignment of features, and we tried to use a transformer applied to flattened token sequence […].</p>\n</blockquote>\n<ul>\n<li><p>Maybe a dumb questions but, if LSTM was better based on your CV results, why don't continue with it despite of the spatial misalignment of features?</p></li>\n<li><p>Do you have any method to speed up experimentation in competitions like this one where training takes a lot of time? For example, when trying new things when you have to train your models for 100-200 epochs.</p></li>\n</ul>\n<blockquote>\n  <p>One of the key things for this particular data (pixel accuracy of masks is needed) was to avoid using torch.resize at the end. Typical segmentation models do x4 linear resizing of masks to get the same mask res as the input. So, I gave x2 larger input image and added a pixel unshuffle x2 upscaling block to the model instead of typical final convolution+resizing.</p>\n</blockquote>\n<ul>\n<li>I am trying to fully get what you mentioned above, let me write what I understood:</li>\n</ul>\n<ol>\n<li>You want use an input image size higher than original (let's imagine 512x512) but resizing the mask to 512x512 via linear interpolation is not desirable because introduces inaccuracies.</li>\n<li>Then you keep the mask size as 256x256 but you have to transform the output of the decoder (512x512) to 256x256 but using torch.resize is not desirable neither.</li>\n<li>The solution is doing pixel unshuffle at the end, which transforms (BS, C, 512, 512) to (BS, C*2, 256, 256). Is that correct?</li>\n</ol>\n<p>Thanks!</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2385934,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2023-08-11T16:30:38.503000",
          "content": "<p>LSTM worked the best for me, but we kept other things for diversity. My training was 24 epochs long, so it took just ~2 hours on 512 input, but you could accelerate the things at initial experiments using smaller res. Seq models took probably 8 hours. The issue, though, was the results may be noisy sometimes, and 5 folds are needed to make a relatable concussion on the experiment. If you have to train 100-200 epochs, you can do initial dev on shorter training and then upscale the things for the best setup.</p>\n<p>The typical output of segmentation models is at x4 lower res, which then is linearly upsampled. So if u give x4 input you do not need to do upsampling of masks at the end. For x2 input, the output wd be 128x128, and additional x2 upscaling with pixel shuffle block (instead of linear upsampling) is needed.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2382865,
      "author_name": "Romain Hardy",
      "author_url": "",
      "post_date": "2023-08-10T04:37:15.270000",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a>! Did you train on upsampled masks (i.e. 512x512 or 1024x1024), or did you downsample the predictions back to 256x256 after your conv/pixelshuffle upsampling? Regardless, thank you for sharing your solution.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2382883,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2023-08-10T04:52:25.517000",
          "content": "<p>I just didn't up-sample the masks in the model keeping 256 size as the output.</p>",
          "votes": 6,
          "replies": [
            {
              "id": 2383585,
              "author_name": "Optimo",
              "author_url": "",
              "post_date": "2023-08-10T13:05:30.923000",
              "content": "<p>How did you do that? simple torch.resize at the end ?</p>\n<p>I tried to add an extra conv layer in the final head so that the model outputs 256 even with images 512 as inputs, also tried to remove the last upsample block but without much success.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2383739,
              "author_name": "Iafoss",
              "author_url": "",
              "post_date": "2023-08-10T14:36:32.400000",
              "content": "<p>One of the key things for this particular data (pixel accuracy of masks is needed) was to avoid using torch.resize at the end. Typical segmentation models do x4 linear resizing of masks to get the same mask res as the input. So, I gave x2 larger input image and added a pixel unshuffle x2 upscaling block to the model instead of typical final convolution+resizing.</p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 2383767,
              "author_name": "Optimo",
              "author_url": "",
              "post_date": "2023-08-10T14:52:41.417000",
              "content": "<p>Ok thank you! But what is a \"pixel unshuffle x2 upscaling block\" exactly? 😅</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2383806,
              "author_name": "Iafoss",
              "author_url": "",
              "post_date": "2023-08-10T15:07:37.300000",
              "content": "<p>You just unshuffle the input from Bx4CxHxW to BxCx2Hx2W in the way that 4 sequential values from the channel dimension go into 2x2 pixels. And I have conv + LayerNorm2D + GELU + conv after and before this shuffle. We will share the code a bit later.</p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 2385946,
              "author_name": "FabienDaniel",
              "author_url": "",
              "post_date": "2023-08-11T16:37:49.640000",
              "content": "<p>Quite tricky ! would you have a rough estimate on the impact of this wrt. to a simpler torch.resize ? </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2394440,
              "author_name": "Iafoss",
              "author_url": "",
              "post_date": "2023-08-16T22:32:55.410000",
              "content": "<p>I didn't check it in this competition, but the negative effect of linear upscaling should be quite strong here because you need to reproduce several pixel-wide masks with pixel accuracy(</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2396774,
      "author_name": "Ioannis M",
      "author_url": "",
      "post_date": "2023-08-18T13:12:36.827000",
      "content": "<p>Congrats for the 2nd place guys! <br>\nThanks for the awesome write-up, we have to study and learn from you. Especially, starting from how you describe the research challenges and explaining your way of thinking is very crucial.</p>\n<p>1) threshold = 0.5 or tuned in val set ?<br>\n2) BTW your very 1st submission that scored ~0.713 on public (If I recall correct) was the Coat_ULSTM single ? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2397530,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2023-08-19T05:26:11.273000",
          "content": "<p>Thanks, congratulations on getting gold too.</p>\n<p>We selected the threshold based on CV, but it is very close to 0.5, like 0.48. Lovasz part of the loss was helpful to reduce the dependence of the dice on the threshold selection.</p>\n<p>My early sub in this competition got 0.71614 public LB and it was a combination of Coat and NeXtViT single-frame models. It was rather a coincidence, and CV was just 0.700, and private LB is 0.70323. The next sub was Seq CoaT+NeXtViT LSTM, which got 0.71205 public LB, while CV is much higher of 0.7039, and private LB is 0.71488. All those subs are just single fold, later we switched to 5 folds. In this competition, we didn't really consider public LB, and relay only on CV on eval set. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2396381,
      "author_name": "Sambit Barik",
      "author_url": "",
      "post_date": "2023-08-18T08:17:08.013000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> on winning the second place and writing this fabulous writeup! </p>\n<p>Can you also please explain the temporal mixing and what type of preprocessing did you use ?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2397520,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2023-08-19T05:13:45.747000",
          "content": "<p>Thanks. I updated the write-up with more details. Regarding preprocessing, I do not think that we did anything spatial in addition to what we listed in the write-up.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2394821,
      "author_name": "Armin Azhdehnia",
      "author_url": "",
      "post_date": "2023-08-17T06:32:51.023000",
      "content": "<p>Hi, Could you please mention what hardware you have used to train your models?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2396079,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2023-08-18T03:47:16.073000",
          "content": "<p>2x4090 in my experiments</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2393084,
      "author_name": "Bananafin",
      "author_url": "",
      "post_date": "2023-08-16T05:51:46.967000",
      "content": "<p>Congratulations on your 2nd place and thanks for the great writeup!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2383714,
      "author_name": "Shailja Kant Tiwari",
      "author_url": "",
      "post_date": "2023-08-10T14:24:46.777000",
      "content": "<p>Congratulations! Nice way to do the work.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2383622,
      "author_name": "DennisSakva",
      "author_url": "",
      "post_date": "2023-08-10T13:26:51.343000",
      "content": "<p>Thanks for the write-up Iafoss and congrats on your impressive result.<br>\nA quick question. Do you train your transformer-based models from scratch or do you modify the pre-trained model so as to make it accept a different resolution? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2383762,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2023-08-10T14:50:47.597000",
          "content": "<p>SAM ViT encoder works with 1024x1024 input, and we just performed x4 upscaling. For CoaT and NeXtViT the resolution is not fixed, similar to conv models. Training a transformer from scratch is quite an unfavorable thing, and if you asked how to use different res in ViT like models, the typical way is just linear resizing of the pos enc to the res u are interested in. After small fine-tuning the model works quite well.</p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 2383539,
      "author_name": "Yassine Alouini",
      "author_url": "",
      "post_date": "2023-08-10T12:31:01.473000",
      "content": "<p>Very detailed solution, thanks for the write-up. This is also the first time I hear about the <code>Over9000</code> optimizer. Is that this repo: <a href=\"https://github.com/mgrankin/over9000\" target=\"_blank\">https://github.com/mgrankin/over9000</a> ?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2383779,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2023-08-10T14:56:13.910000",
          "content": "<p>Yep. Previously I used it a lot because of the good compatibility of this optimizer with ssl_ResNeXt50 weights in fine-tuning, but later, working with transformers, I mostly switched to AdamW. It was interesting to see that CoaT seems to work quite well with this optimizer.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2383859,
              "author_name": "Yassine Alouini",
              "author_url": "",
              "post_date": "2023-08-10T15:44:19.370000",
              "content": "<p>Thanks for the insights. I will add this to my tools. Personally I mostly use <strong>AdamW</strong> or <strong>Adam</strong> since they work well most of the time (despite all the recent refinements 😁).</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2383072,
      "author_name": "iadduk",
      "author_url": "",
      "post_date": "2023-08-10T07:02:01.027000",
      "content": "<p>Congratulations!<br>\nCan you please show the exact formula for the dice Lovasz loss?<br>\nHow much did you gain from using CutMix strategy?<br>\nI can't wait to see your code.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2383197,
          "author_name": "Theo Viel",
          "author_url": "",
          "post_date": "2023-08-10T08:26:08.560000",
          "content": "<p>Cutmix gave +0.005 for v2-s models. It did not work that well for other models though<br>\nWe will release the code soon, with the implementation of the loss.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2383032,
      "author_name": "franchesoni",
      "author_url": "",
      "post_date": "2023-08-10T06:31:34.780000",
      "content": "<p>thanks for sharing the solution this quickly. What was your postprocessing? single thresholding tuned on CV?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2383039,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2023-08-10T06:35:36.767000",
          "content": "<p>yep, it worked the best. Since global dice is considered in this competition, I don't think that post-processing may really work well here.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2383016,
      "author_name": "Ted K",
      "author_url": "",
      "post_date": "2023-08-10T06:22:39.653000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> on 2nd and another gold!  Thank you for the thoughtful, detailed discussion of your solution, and for including a couple things that didn't work.  It's so nice to read such a clear writeup!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2384464,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2023-08-11T01:30:10.987000",
          "content": "<p><a href=\"https://www.kaggle.com/socated\" target=\"_blank\">@socated</a> you are very welcome</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2382851,
      "author_name": "Suraj",
      "author_url": "",
      "post_date": "2023-08-10T04:27:21.683000",
      "content": "<p>Kudos <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> on this achievement. Thanks for the write up.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2384467,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2023-08-11T01:30:45.743000",
          "content": "<p>Thank you so much</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2382827,
      "author_name": "C R Suthikshn Kumar",
      "author_url": "",
      "post_date": "2023-08-10T04:10:50.603000",
      "content": "<p>Congratulations for winning the 2nd place.<br>\nWe appreciate your sharing info about the solution. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2384472,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2023-08-11T01:32:26.420000",
          "content": "<p>You are welcome</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2382815,
      "author_name": "NA",
      "author_url": "",
      "post_date": "2023-08-10T03:57:01.853000",
      "content": "<p>That is very interesting. Congratulations with the second place!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2413521,
      "author_name": "kongweihao",
      "author_url": "",
      "post_date": "2023-08-29T01:21:19.613000",
      "content": "<p>How can you get you backbone (CoaT, NeXtViT, SAM-B, and tf_efficientnetv2_s ) ? Did you use some automatic tools? Thank you very much. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 2414796,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2023-08-30T00:48:23.453000",
          "content": "<p>For tf_efficientnetv2_s  you can use timm, which is super convenient. For others, that are not shared through timm, I just go to the corresponding github, grab files with the model definition and related things, modify it a bit in some cases to eliminate unnecessary dependences like mmseg, download weights from the repo, and just use the model.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2417960,
              "author_name": "kongweihao",
              "author_url": "",
              "post_date": "2023-09-01T03:10:54.947000",
              "content": "<p>It has inspired me a lot, thank you！</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2394585,
      "author_name": "Bilzard",
      "author_url": "",
      "post_date": "2023-08-17T03:25:34.907000",
      "content": "<p><a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> Congratulations for 2nd place. </p>\n<p>I have problem with training transformer T-Mixer because it is easy to get loss equal \"nan\".<br>\nDid you experience same problem? And could you share some tips you used on stabilizing training?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2394754,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2023-08-17T05:44:02.350000",
          "content": "<p>Thanks. Yep, training with it is less stable than LSTM. With CoaT model everything was ok, but the spread of the results between folds is large. For NeXtViT, meanwhile, I could not finish the training from scratch even if tried to play with lr, etc. Probably proper init may resolve the problem. Meanwhile, in the case of using a model pre-trained on single frame setup without mixing (both encoder and decoder), training of the transformer mixing is quite more stable, and I didn't have any issues during training. <br>\nI used this setup for the case of PL pertaining. However, even if I pretrained the single frame model with different seeds for each fold, the fine-tuned models had less diversity (in comparison to training from scratch), and these models were not so useful for the ensemble in the end.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2394771,
              "author_name": "Bilzard",
              "author_url": "",
              "post_date": "2023-08-17T05:55:23.233000",
              "content": "<blockquote>\n  <p>Meanwhile, in the case of using a model pre-trained on single frame setup without mixing (both encoder and decoder), training of the transformer mixing is quite more stable, and I didn't have any issues during training.</p>\n</blockquote>\n<p>Thanks for that tips. </p>\n<p>Unfortunately, I have trouble even with CoaT as backend. Conv 3D mixier is quite stable, but not for naive transformer T-mixier, nor T-mixier with 3d-adapted CoaT blocks.<br>\nMaybe I should tune learning rate or other parameters that rerate to regularization.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2396078,
              "author_name": "Iafoss",
              "author_url": "",
              "post_date": "2023-08-18T03:46:37.783000",
              "content": "<p>The origin of this unstable behavior, to my understanding, is the label noise(<br>\nI was using 3e-4 lr, but what was more important in my case is Over9000 optimizer with epsilon=1e-4. As I remember, this optimizer not only worked better for CoaT but also gave more stable training than AdamW.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2388436,
      "author_name": "Kabir_011",
      "author_url": "",
      "post_date": "2023-08-13T11:42:27.283000",
      "content": "<p>Congratulation! Nice Work man.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2388215,
      "author_name": "Muhammad Usman",
      "author_url": "",
      "post_date": "2023-08-13T08:31:39.507000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a>! You did a great job man. </p>\n<p>Your solution is amazing, very detailed.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2388575,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2023-08-13T13:54:05.457000",
          "content": "<p>Thanks.                          </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2383682,
      "author_name": "kongweihao",
      "author_url": "",
      "post_date": "2023-08-10T14:10:29.017000",
      "content": "<p>May I ask, what configuration of machine did you use for these experiments? These experiments seem to have very high requirements for the machine</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2383748,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2023-08-10T14:45:50.633000",
          "content": "<p>For me it was 2x4090 machine, quite fast cards for deep learning but VRAM is an issue, like always. <a href=\"https://www.kaggle.com/drhabib\" target=\"_blank\">@drhabib</a> used 2xA6000, and <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> used his V100 cluster.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2384506,
              "author_name": "kongweihao",
              "author_url": "",
              "post_date": "2023-08-11T01:46:40.203000",
              "content": "<p>Your answer has been very enlightening to me, and I am very grateful</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2383131,
      "author_name": "iadduk",
      "author_url": "",
      "post_date": "2023-08-10T07:35:36.013000",
      "content": "<p>Can you please share how codewise you combine coat encoder and Unet decoder?<br>\nCoat (coatnet, right?) encoder is not in SMP repo, so you did it by yourself, right?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2383193,
          "author_name": "Theo Viel",
          "author_url": "",
          "post_date": "2023-08-10T08:23:17.587000",
          "content": "<p>Exactly ! We will release the code later.</p>",
          "votes": 6,
          "replies": [
            {
              "id": 2383222,
              "author_name": "ynhuhu",
              "author_url": "",
              "post_date": "2023-08-10T08:49:09.083000",
              "content": "<p>Thanks for your sharing. There are a few confusion to me in your details (2).</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2382830,
      "author_name": "Nirjhar Roy",
      "author_url": "",
      "post_date": "2023-08-10T04:14:23.157000",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a>  congratulations for amazing result . Can you please give some more details on the temporal mixing ? We used only single frame model ..so curious to know about this solution.  </p>",
      "votes": 2,
      "replies": [
        {
          "id": 2382888,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2023-08-10T04:54:40.670000",
          "content": "<p>We will share the code a bit later, but the thing I considered is mixing at low res feature maps because images are poorly alighted (time between images is too large). Congratulations with getting gold!</p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 2382812,
      "author_name": "qdv206",
      "author_url": "",
      "post_date": "2023-08-10T03:56:26.883000",
      "content": "<p>Congratulations on the strong finish! CoaT is my favorite backbone so glad to see it performs so well here.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2386554,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2023-08-12T03:20:47.713000",
          "content": "<p>CoaT is a really good alternative to Segformer when a restrictive license is not acceptable. A combination of global attention and consideration of details at high-resolution feature maps (in contrast to ViT architectures) is something one doesn't see often. Meanwhile, in this particular competition, it was very important. </p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 2382792,
      "author_name": "Sourabh Singh",
      "author_url": "",
      "post_date": "2023-08-10T03:40:47.567000",
      "content": "<p>congrats <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a>  for 2nd place. Try to get 1st place</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2382795,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2023-08-10T03:44:08.957000",
          "content": "<p>Thanks so much. It happens for us in the second time during last several months when we drop from 1st place at public to the 2nd place at private((( Getting the first place is quite a difficult thing.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 2394533,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-08-17T02:02:25.790000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2382786": "# Summary\n- Customized U-shape network with CoaT, NeXtViT, SAM-B, and tf_efficientnetv2_s backbones\n- ×2 (with pixel shuffle up-scaling of predictions to 256×256) and ×4 input upscaling: pixel-level accuracy of predictions is critical\n- Sequence of images: LSTM, Transformer, Convolutional temporal mixing\n- No flip and 90-degree rotation augmentation (nor TTA) because masks are shifted\n- Train on soft labels: label = average of all annotators\n- BCE + dice Lovasz loss\n\n# Code\nMy and @drhabib 's [part](https://github.com/DrHB/2nd-place-contrails) (includes the best performing CoaT_ULSTM model)\n@theoviel 's [part](https://github.com/TheoViel/kaggle_contrails)\n[Inference](https://www.kaggle.com/code/theoviel/contrails-inference-comb)\n\n# Introduction\nOur team would like to thank the organizers and Kaggle for making this competition possible. Also, I want to express my gratitude to my outstanding teammates @drhabib and @theoviel for their incredible contribution toward our final result. \n\n# Details\nThis competition has two main challenges: (1) noisy labels and (2) pixel-level accuracy requirements. (1) If one checks the annotation from all annotators, a major disagreement is quite apparent. This label noise challenge may be partially addressed by using soft labels (the average of all annotator labels) during training. However, since the evaluation requires hard labels, the major model failure is the prediction of contrails near the decision boundary, which cannot be avoided. (2) The thing that could really be addressed by the models is achieving pixel-level accuracy of predictions. Pixel-level accuracy is quite important since contrails are only several pixels thick, and even a mistake by one pixel in the mask boundary in the lateral direction may result in a significant decrease in the dice score, used as a metric. For such a task, a typical solution is up-sampling the input of the model or replacing the final linear up-sampling in segmentation models with transposed convolution or pixel shuffle up-sampling. This modification gave us a substantial boost in initial experiments over the direct use of models on the original resolution.\n\n## Data\nFollowing the [organizer's publication](https://arxiv.org/pdf/2304.02122.pdf), we used “ash” false color images considering the 12 μm band, the difference between 12 and 11 μm bands, and the difference between 11 and 8 μm bands, respectively. We also tried to consider all bands or expand the “ash” color images with 8, 10, and 12 μm, but it resulted in lower performance (the pre-trained weights of the first convolution were replicated). We hypothesize that the best performance of ash color images may be a consequence of the use of the ash color images by annotators, and biased label noise. In our experiments, we up-sample the input with bi-cubic interpolation by ×2 for CoaT and NeXtViT models, and by ×4 for SAM model. EfficientNet model uses the original input but the stride in the first convolutional block is set to 1, which is equivalent to ×4 input up-sampling.\n\n## Model\nThe interesting thing about this competition is that the model is required to do both (1) tracking the global pixel dependencies because contrails are quite elongated, and (2) capable of generating predictions with pixel-level accuracy because contrails are only several pixels thick. (1) can be addressed with transformers, while (2) is addressed with convolutional networks or local attention. In the competition, we tried to address these points independently with SAM-B ViT backbone and EfficientNet v2 purely convolutional network. In addition, there is a class of networks, transformers with a hierarchical structure, which addresses both aspects of the considered problem, (1) + (2), simultaneously. This class includes CoaT, which resulted in the best single model performance in our experiments, NeXtViT, MaxViT (which we missed), and MiT transformer from Segformer (unfortunately, it cannot be used in the competition because of the license).\nFor the decoder we either considered a vanilla U-Net decoder (EfficientNet models) or a more customized set of U-blocks using pixel shuffle up-sampling and factorized FPN (CoaT, NeXtViT, SAM). In these blocks, we also replaced the BatchNorm with LayerNorm2d to have proper model convergence at the low batch size and used GELU instead of ReLU activation. \nThe models are schematically illustrated in the figure below.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2F16e90aba2cb8ee7ac94837e09455688f%2F4.png?generation=1692419847964897&alt=media)\n\n## Temporal mixing\nIn addition to our initial runs with single-frame models, we performed experiments with image sequences. In our best setup image, multi-frame models get ~0.01 improvement.\nSince the images are taken with a considerable temporal delay, and the displacement of clouds between frames is huge (~10-30 pixels) and is deteriorated by input up-sampling, video models based on 3D convolutions or window attention (VideoSwin) are not effective in the context of the considered problem. For the same reason, early mixing or mixing the predictions of sequential models is also not expected to work. Therefore, mixing at the intermediate feature map scales, such as res/32 and res/16, is most promising: even if clouds/contrails are misplaced, the feature maps at low resolution are reasonably aligned. In our experiments with the CoaT model, the best performance is achieved if two low-resolution feature maps (res/32 and res/16) are considered. The output of the temporal mixing modules as well as the output of the remaining feature maps are pooled at frame 4 and the decoder sees only the input corresponding to the considered frame, as illustrated in Figure 1.\nAs temporal mixing modules, we considered LSTMs (which achieved the best performance), 1D temporal convolutions, and Transformers. LSTM (1 layer) and 1D temporal convolutions perform mixing along the temporal dimension only, while the spatial dependences between features are considered in the backbone and the decoder. The transformer-based mixing was proposed to perform an implicit feature map registration for a large displacement of contrails/clouds because it works with temporal and spatial mixing of features simultaneously and can match them even if they are not spatially aligned. \nThis mixing is performed in the following way. To each feature map, corresponding to a given frame, we add encoding to distinguish it from others, then we flatten tokens from all frames into a single sequence and process it with 2 transformer decoder blocks. Key and value input to the transformer is represented by the concatenated sequence of all frames, while the query is a sequence of tokens from the 4-th frame only. Unfortunately, transformer-based mixing ended up with slightly lower performance, and the training with it is less stable. However, we kept both approaches in the final ensemble, benefitting from the improved diversity.\nIn our experiments with CoaT/NeXtViT we followed the [organizer's publication](https://arxiv.org/pdf/2304.02122.pdf) and selected the first 5 frames as the input. Both LSTM and Transformer mixing were considered. In our experiments with EfficientNet we considered 4 frame input (2nd – 5th, 2 frames before and 1 frame after the annotated frame) with bi-directional LSTM. This frame selection is dictated by the instruction to annotators to have a contrail at least on two subsequent frames. With the SAM model, we considered 1D-convolutional mixing with 3 frames (1 frame before and 1 frame after the annotation) because of the heavy VRAM requirement for this model, which was also an issue in inference. \nOne idea we had is explicit registration of input frames to have contrails and clouds aligned between frames (temporal mixing is the most effective), for example considering optical flow. Unfortunately, the images contain not only clouds but also the Earth's surface, which does not move, or several layers of clouds moving in different directions. Therefore, we could not come up with any methods to proceed with the frame registration task. We were also considering performing registration based on the predicted contrails, but since contrails appear and disappear from frame to frame, this approach is also not feasible.\n\n##Pseudo-labels (PL)\nIn addition to the competition data, we have collected a dataset “Contrails GOES16 Images May” (single frame setup) [9]. We split the images into 256×256 tiles with partial overlap and generated the masks based on the ensemble of best single-frame models. This data is shared as a Kaggle dataset [10]. This data was used for pretraining the models in a single-frame setup in some experiments, followed by finetuning the models either in single-frame or multi-frame setups. The use of PL has drastically improved the performance of individual models (folds), but the diversity of the models is reduced even if each fold of the finetuned model uses an independently trained PL model. Therefore, the performance on the average prediction over all folds is comparable to one for the setup trained without PL.\nWe also have experimented with PL training on both external data + unannotated competition frames, but it did not give any benefit in finetuning a single frame model. Since such pretraining could affect multi-frame models, in the production experiments we used only external data at the PL step.\n\n## Training\nCoaT and NeXtViT models: Training is performed for 24 epochs with Over9000 (Radam+LAMB+LookAhead) optimizer, learning rate of 3.5e-4, weight decay of 0.01. It appeared that this optimizer has superior compatibility with CoaT. In the case of PL pretraining, it is performed for 12 epochs on the external data, and then the models are finetuned for 12 epochs for CoaT and 18 epochs for NeXtViT models.\nEfficientNet models: Pretraining is performed in a single frame setup for 100 epochs with additional CutMix augmentation, learning rate is 1e-3, AdamW optimizer, weight decay 0.2. During pretraining the fraction of PL data is linearly decreased to zero. Then the model is fine-tuned in the image sequence setup with 3e-5 for the encoder and 1e-4 for the decoder. The addition of external data for pretraining resulted in much stronger 2D models, hence multi-frame EfficientNet models were not considered in the final ensemble. For diversity, we also pre-trained models for 200 epochs, and in a full-fit setup (training on all data, including validation set).\nSAM models: Our final submission includes only single-frame SAM-B ViT-based models. During training, we, first, train the model with a frozen encoder for 10 epochs, and then we unfreeze the model and continue for 20 more epochs. AdamW optimizer is used with the learning rate of 4e-6 for the encoder and 4e-5 for the decoder.\nDuring training, we used ShiftRotateScale, RandomGamma, RandomBrightnessContrast, MotionBlur, and GaussianBlur augmentations. Flip and 90-rotate augmentation were not used because they resulted in worse performance. The training of all production models is performed for 5-6 folds our folds are trained on the entire train dataset with different random seeds, while the evaluation is performed on the provided validation set. This strategy is used because evaluation in a standard K-fold split of the training data overestimates CV (cross-validation), which may be caused by spatial overlaps of tiles in the training dataset.\n\n## Loss\nAs a loss function, we used BCE + dice Lovasz loss. The second term is a modified version of a Lovasz loss [12] where (1) ReLU is replaced with 1 + ELU, (2) a symmetric version is used, and (3) instead of IoU we use a surrogate function maximizing dice. The second term provides an insignificant boost to the individual model performance. However, this term gives a wider maximum of the dice with respect to the threshold selection, i.e. weaker dependence of the performance on the threshold selection. This property is vital to avoid shakeup at the private leaderboard because of the improper threshold and is preferable for more effective ensembling.\n\n## Results\nThe performance of our final models is summarized in the table below. Ex refers to models trained in PL setup (EfficientNet is trained with PL only). The postfix U refers to the single frame model, UT – transformer-based temporal mixing, and ULSTM – LSTM-based temporal mixing. The threshold is selected based on the search maximizing CV value. For most models, it is 0.46-0.50.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2F5b5010ef692bccaca91f04bdd8ae6be3%2F5.png?generation=1692423621715751&alt=media)\n\n## Best Single model\nOur best single mode is CoaT, which for 5 folds gets 0.7039 CV, 0.71790 private, and 0.71243 public LB. Single fold CV is 0.6960+-0.0003 evaluated on val set (our folds are trained on the train set with different seeds because eval in a standard Kfold split of train data overestimates CV due to possible spatial overlaps). **We could take top-4 in the competition with this single model**. \n\n## Ensemble\nOur best ensemble consists of 8 models with weights selected to make approximately equal contributions from all components: CV is 0.7140 (on eval set), 0.72574 public and 0.72304 private LB. The models and their weight are summarized in the table below.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2F2ddf4778766d0b5aa44a24a81409d460%2F6.png?generation=1692421380873655&alt=media)\nEfficientNet includes 3 model setups (100 epochs, 200 epochs, and full-fit) with the total efficient weight equal to 1.\n\n##Model Execution Time\nCoaT and NeXtViT models: Training time is 2 hours for single-frame and 8 hours for multi-frame models per fold on 2×RTX4090 GPUs. The inference time at Kaggle (P100) for multi-frame models is 35 minutes for 5 folds.\nEfficientNet models: Training time is 3 hours 45 minutes for 100 epochs on 8×V100 GPUs per fold. Inference at Kaggle (P100) takes 20 minutes for 6 folds including data loading.\nSAM models: Training time is 16 hours per fold on 2×A6000 GPUs. Inference at Kaggle (P100) takes 48 minutes for 5 folds including data loading.\n\n## Things we missed\n(1) During the competition one surprising finding for us was that flips and 90-degree rotation augmentation in model training resulted in worse performance. It was quite surprising, but, unfortunately, we did not think sufficiently about the origin of this behavior and attributed it to prevailing winds and a particular shape of clouds and contrails. After the competition, it appeared that the masks were shifted by 0.5 pixels, which made any flips inapplicable, as we discovered. If the mask issue was fixed in our setup, flip/90-rotation augmentation could lead to an 8-fold increase of the effective training dataset size and enables further performance boost at the inference stage due to test time augmentation. So, potentially the performance of our models can be noticeably improved. \n(2) PL single model pertaining on the external data gives moderate CV improvement to the single-frame model (interesting that at private LB this improvement is quite large). Meanwhile, this single-frame pretraining didn't give much to sequence models. So, external sequence data collection + proper reprojection could potentially server as a better pertaining strategy for sequence models, and we may expect a boost comparable to single-frame models.",
    "2382925": "Congratulations! Beautiful solution, thank you for sharing it with us. Looking forward to studying your code.\n\nWe knew you had a unique approach that gained your team a nice separation throughout. Even during the competition I was excited to learn about it and it didn't disappoint.\n\nWe had a feeling the temporal information was important and tried a bunch of stuff like 3d models and sequence models to no avail, so we dropped the idea as computation was expensive.\n\nWe also had a feeling that there was important information in the individual masks and tried a bunch of stuff like using random individual masks as auxiliary signal or different classes, but the simplest approach of averaging the individual masks didn't occur to us, nice one.",
    "2382849": "thanks for the writeup.\n\ni am interested in usage of tempral information. how much did it increase your CV and LB.???\n\n---\n\nI spent the last week in fusion of multiple frams, early fusion, late fusion, etc ...\nI also tried optcal flow (RAFT) for registration and  augmentation (find the flow and then wrap image and mask)\nnone of them really work well.\ni only have 0.004~0.005 improvement when using 3 consective frames (over 2 different models i tried)\nThis is worse than expected  (the paper can get 0.01 to 0.015 improvement). I wondered why????\n\ni don't think large movement  is the cause of poor performance.\n(on a side note, actually you can use temporal superresolution, i.e. predict intermediate frames with frame interpolation. https://www.youtube.com/watch?v=MjViy6kyiqs). Frame interpolation results may be wrong, but as long as it is \"consistent\", it can \"provide temporal features\"",
    "2388583": "Congratulations on your 2nd place and thanks for the great writeup!\nI'm interested in temporal mixing by LSTM. Do you stack each pixel to a batch dimension as input to LSTM, i.e., (n_batch * H * W, n_frame, C)?",
    "2387504": "Congratulations! A very unique solution.\n\nQuick question. Sorry, if I ask, but I don't understand how a transformer is used as a mixer here. Does it have anything to do with the MLP mixer idea or are you just passing multiple frames to a transformer? \n\nAlso, waiting for you to release the code. Thanks for the solution.",
    "2385782": "Congratulations to you three and thanks a lot for such a detailed summary of your solution. It is incredible the level of the top spot solutions, it makes me feel dumb and happy at the same time 😄:\n\n>The best mixing strategy based on our experiments was using LSTM. Unfortunately, LSTM assumes spatial alignment of features, and we tried to use a transformer applied to flattened token sequence [...].\n\n- Maybe a dumb questions but, if LSTM was better based on your CV results, why don't continue with it despite of the spatial misalignment of features?\n\n- Do you have any method to speed up experimentation in competitions like this one where training takes a lot of time? For example, when trying new things when you have to train your models for 100-200 epochs.\n\n>One of the key things for this particular data (pixel accuracy of masks is needed) was to avoid using torch.resize at the end. Typical segmentation models do x4 linear resizing of masks to get the same mask res as the input. So, I gave x2 larger input image and added a pixel unshuffle x2 upscaling block to the model instead of typical final convolution+resizing.\n\n- I am trying to fully get what you mentioned above, let me write what I understood:\n1. You want use an input image size higher than original (let's imagine 512x512) but resizing the mask to 512x512 via linear interpolation is not desirable because introduces inaccuracies.\n2. Then you keep the mask size as 256x256 but you have to transform the output of the decoder (512x512) to 256x256 but using torch.resize is not desirable neither.\n3. The solution is doing pixel unshuffle at the end, which transforms (BS, C, 512, 512) to (BS, C*2, 256, 256). Is that correct?\n\nThanks!\n",
    "2382865": "Congrats @iafoss! Did you train on upsampled masks (i.e. 512x512 or 1024x1024), or did you downsample the predictions back to 256x256 after your conv/pixelshuffle upsampling? Regardless, thank you for sharing your solution.",
    "2396774": "Congrats for the 2nd place guys! \nThanks for the awesome write-up, we have to study and learn from you. Especially, starting from how you describe the research challenges and explaining your way of thinking is very crucial.\n\n1) threshold = 0.5 or tuned in val set ?\n2) BTW your very 1st submission that scored ~0.713 on public (If I recall correct) was the Coat_ULSTM single ? \n",
    "2396381": "Congratulations @iafoss on winning the second place and writing this fabulous writeup! \n\nCan you also please explain the temporal mixing and what type of preprocessing did you use ?",
    "2394821": "Hi, Could you please mention what hardware you have used to train your models?",
    "2393084": "Congratulations on your 2nd place and thanks for the great writeup!",
    "2383714": "Congratulations! Nice way to do the work.",
    "2383622": "Thanks for the write-up Iafoss and congrats on your impressive result.\nA quick question. Do you train your transformer-based models from scratch or do you modify the pre-trained model so as to make it accept a different resolution? \n",
    "2383539": "Very detailed solution, thanks for the write-up. This is also the first time I hear about the `Over9000` optimizer. Is that this repo: https://github.com/mgrankin/over9000 ?",
    "2383072": "Congratulations!\nCan you please show the exact formula for the dice Lovasz loss?\nHow much did you gain from using CutMix strategy?\nI can't wait to see your code.\n",
    "2383032": "thanks for sharing the solution this quickly. What was your postprocessing? single thresholding tuned on CV?",
    "2383016": "Congratulations @iafoss on 2nd and another gold!  Thank you for the thoughtful, detailed discussion of your solution, and for including a couple things that didn't work.  It's so nice to read such a clear writeup!",
    "2382851": "Kudos @iafoss on this achievement. Thanks for the write up.",
    "2382827": "Congratulations for winning the 2nd place.\nWe appreciate your sharing info about the solution. ",
    "2382815": "That is very interesting. Congratulations with the second place!",
    "2413521": "How can you get you backbone (CoaT, NeXtViT, SAM-B, and tf_efficientnetv2_s ) ? Did you use some automatic tools? Thank you very much. ",
    "2394585": "@iafoss Congratulations for 2nd place. \n\nI have problem with training transformer T-Mixer because it is easy to get loss equal \"nan\".\nDid you experience same problem? And could you share some tips you used on stabilizing training?",
    "2388436": "Congratulation! Nice Work man.\n",
    "2388215": "Congratulations @iafoss! You did a great job man. \n\nYour solution is amazing, very detailed.",
    "2383682": "May I ask, what configuration of machine did you use for these experiments? These experiments seem to have very high requirements for the machine",
    "2383131": "Can you please share how codewise you combine coat encoder and Unet decoder?\nCoat (coatnet, right?) encoder is not in SMP repo, so you did it by yourself, right?",
    "2382830": "Hello @iafoss  congratulations for amazing result . Can you please give some more details on the temporal mixing ? We used only single frame model ..so curious to know about this solution.  ",
    "2382812": "Congratulations on the strong finish! CoaT is my favorite backbone so glad to see it performs so well here.",
    "2382792": "congrats @iafoss  for 2nd place. Try to get 1st place",
    "2394533": ""
  }
}