{
  "id": 432871,
  "title": "Architectural Exploration Based on Top Solutions",
  "url": "/competitions/google-research-identify-contrails-reduce-global-warming/discussion/432871",
  "author_name": "Bilzard",
  "post_date": "2023-08-19T09:28:06.487000",
  "votes": 27,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Based on the knowledge shared on public discussion, I want to share the notable results which is obtained by experiments after the competition.<br>\nI appreciate participants who kindly shared their solutions.</p>\n<h2>Overview</h2>\n<p>First, my largest interest was</p>\n<ol>\n<li>How much gain can be obtained by correction of misalignment?</li>\n<li>How much gain can be obtained by soft labels?</li>\n</ol>\n<p>These results are already added to my solution write-up[4], so you can refer to it first if you haven't read.</p>\n<p>In addition, I tested performance gain of these architectural choices:</p>\n<ol>\n<li>backbones</li>\n<li>decoders, especially the impact of PixelShuffle</li>\n<li>T-Mixer (same as <em>Temporal Modulator</em> in my original solution)</li>\n</ol>\n<p>In this article, I discussed the experiment result of 3-5.</p>\n<h2>Experiment Detail</h2>\n<h3>Note</h3>\n<p>Given the constraints of computing resources, all experiments were conducted using a single fold and a single random state. This means the results might not be entirely robust. For more definitive conclusions, multiple trials for each architecture would be necessary.</p>\n<h3>Choice of backbones</h3>\n<p>I evaluated various backbones using a simple 2D U-Net architecture that only utilizes the current frame. The models were trained on <code>training</code> images and evaluated on <code>validation</code> images. For all experiments, I employed label correction and soft-labels. The results are presented in Table 1. Among the options, CoaT-lite medium demonstrated the best balance of performance and training time. Consequently, I adopted this backbone as the default for subsequent experiments.</p>\n<p><strong>Table 1: CV score of various backbones</strong></p>\n<table>\n<thead>\n<tr>\n<th>backbone</th>\n<th>crop_size</th>\n<th>epochs</th>\n<th>gradient checkpointing</th>\n<th>batch_size</th>\n<th>CV</th>\n<th>training time/epoch [sec]</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>efficientnet_b3</td>\n<td>512</td>\n<td>20</td>\n<td>FALSE</td>\n<td>28</td>\n<td>0.6629</td>\n<td>187</td>\n</tr>\n<tr>\n<td>maxvit_tiny_tf_512.in1k</td>\n<td>512</td>\n<td>20</td>\n<td>FALSE</td>\n<td>19</td>\n<td>0.6725</td>\n<td>396</td>\n</tr>\n<tr>\n<td>maxvit_small_tf_512.in1k</td>\n<td>512</td>\n<td>20</td>\n<td>TRUE</td>\n<td>21</td>\n<td>0.672</td>\n<td>709</td>\n</tr>\n<tr>\n<td>pvt_v2_b3</td>\n<td>512</td>\n<td>20</td>\n<td>FALSE</td>\n<td>20</td>\n<td>0.6681</td>\n<td>341</td>\n</tr>\n<tr>\n<td>coat_lite_medium</td>\n<td>512</td>\n<td>20</td>\n<td>FALSE</td>\n<td>20</td>\n<td>0.6787</td>\n<td>413</td>\n</tr>\n<tr>\n<td>coat_small</td>\n<td>512</td>\n<td>20</td>\n<td>FALSE</td>\n<td>19</td>\n<td>0.6713</td>\n<td>513</td>\n</tr>\n</tbody>\n</table>\n<h3>Changing Up-scaling to PixelShuffle</h3>\n<p>PixelShuffle[1] offers an alternative method for up-scaling resolution while reducing feature channels. Following the approach mentioned by the 2nd place team in their solution[2], I implemented the PixelShuffle block as a replacement for the traditional up-scaling block, leading to what I refer to as the PS-U-Net architecture.</p>\n<p>When comparing the U-Net with the PS-U-Net, there wasn't a significant performance gain (CV: 0.6787 to 0.6779). However, the PS-U-Net had the advantage of being more memory-efficient due to the reduced feature channels. Given this efficiency, I chose to use the PS-U-Net as the default architecture for subsequent experiments.</p>\n<h3>The Effects of Intensive Augmentation</h3>\n<p>Thanks to the correction of misalignment, we can now employ a more intensive augmentation strategy, as mentioned in the 1st place solution[3]. I adapted Koda's augmentation approach with slight modifications, as shown below:</p>\n<pre><code>cfg.geometric_transform = A.Compose(\n    [\n        A.RandomRotate90(p=),\n        A.HorizontalFlip(p=),\n        A.ShiftScaleRotate(\n            rotate_limit=, scale_limit=, border_mode=cv2.BORDER_CONSTANT, p=\n        ),\n    ]\n)\n</code></pre>\n<p>The differences are:</p>\n<ol>\n<li>slightly increase probability (0.5 -&gt; 0.6)</li>\n<li>set <code>borader_mode</code> as constant in order to avoid unintentional artifacts</li>\n</ol>\n<p>The experimental results highlighted the substantial impact of this intensive augmentation. It led to an increase in the CV score by 0.58% (from 0.6779 to 0.6837). This underscores that addressing misalignment is crucial for success in this competition.</p>\n<h3>Changing Decoder Architecture: PS-U-Net -&gt; PS-FPN</h3>\n<p>Next, I transitioned from the PS-U-Net architecture to the PS-FPN, where I replaced the up-scaling block of the FPN with PixelShuffle blocks.</p>\n<p>This change resulted in a modest improvement of 0.24% in CV scores (from 0.6837 to 0.6861). However, it's uncertain whether this gain can be solely attributed to the architectural modification. To draw a more definitive conclusion, averaging across multiple seeds would be advisable. Nonetheless, based on these preliminary results, I opted to use the PS-FPN architecture as the standard for subsequent experiments.</p>\n<h3>Comparing T-Mixer Architecture</h3>\n<p>Lastly, I tested the impact of different architectural choice of T-Mixer.</p>\n<p>Especially, I wanted to compare which mixer is superior to this task: Convolutions and Transformer. So I compared below three architecture:</p>\n<ol>\n<li>Conv3d</li>\n<li>Transformer</li>\n<li>CoaT with (3D positional encoding)</li>\n</ol>\n<p>The last pattern was just a spontaneous idea I had. However, since this backbone performs best in 2D settings, I believed it might be superior as temporal mixer. In this experiment, I evaluated on CV and LB. Additionally, I employed TTA to ensure robust evaluation.</p>\n<p>Note that training Transformer and CoaT T-Mixer were unstable, sometimes resulting in NaN predictions. Stability techniques were applied to address this (see Appendix A). In contrast, Conv3d T-Mixer training was stable without any adjustments.</p>\n<p>Table 2 displays the results. The first row represents a 2D U-Net without a T-Mixer, while the second uses a 2.5D PS-FPN with a Conv3D T-Mixer. Incorporating time frame information improved both CV and Private LB scores significantly. This indicates the value of temporal information even after alignment corrections.</p>\n<p>Table 2's 3rd and 4th rows show results for Transformer and CoaT T-Mixers. While CoaT T-Mixer had the highest CV score, its private LB score was the lowest. 3D Conv T-Mixer and Transformer T-Mixer had comparable performances.</p>\n<p>Table 3 presents ensemble model results. The best Private score came from an ensemble of three T-Mixer architectures, suggesting model diversity. Due to resource constraints, only one random seed was tested, but seed averaging or cross-fold averaging might yield better scores.</p>\n<p><strong>Table2: Comparison of Different T-Mixer Architectures</strong></p>\n<table>\n<thead>\n<tr>\n<th>id</th>\n<th>used frames</th>\n<th>T-Mixer</th>\n<th>seed</th>\n<th>TTA</th>\n<th>crop_size</th>\n<th>CV</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>[4]</td>\n<td>None</td>\n<td>x1</td>\n<td>hflip, rot(90, 180, 270)</td>\n<td>512</td>\n<td>0.691</td>\n<td>0.70244</td>\n<td>0.70462</td>\n</tr>\n<tr>\n<td>2</td>\n<td>[1, 2, 3, 4]</td>\n<td>3d Conv</td>\n<td>x1</td>\n<td>hflip, rot(90, 180, 270)</td>\n<td>512</td>\n<td>0.6978</td>\n<td>0.7147</td>\n<td>0.71681</td>\n</tr>\n<tr>\n<td>3</td>\n<td>[1, 2, 3, 4]</td>\n<td>Transformer</td>\n<td>x1</td>\n<td>hflip, rot(90, 180, 270)</td>\n<td>512</td>\n<td>0.6958</td>\n<td>0.70576</td>\n<td>0.71641</td>\n</tr>\n<tr>\n<td>4</td>\n<td>[1, 2, 3, 4]</td>\n<td>CoaT(3D pos enc)</td>\n<td>x1</td>\n<td>hflip, rot(90, 180, 270)</td>\n<td>512</td>\n<td>0.6988</td>\n<td>0.71553</td>\n<td>0.71229</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Table3: Result of Ensemble</strong></p>\n<table>\n<thead>\n<tr>\n<th>ensemble</th>\n<th>CV</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>[2, 4]</td>\n<td>0.7009</td>\n<td>0.718</td>\n<td>0.71827</td>\n</tr>\n<tr>\n<td>[2, 3, 4]</td>\n<td>0.701</td>\n<td>0.71638</td>\n<td>0.72156</td>\n</tr>\n</tbody>\n</table>\n<h2>References</h2>\n<ul>\n<li>[1] <a href=\"https://arxiv.org/abs/1609.05158v2\" target=\"_blank\">Real-Time Single Image and Video Super-Resolution Using an Efficient Sub-Pixel Convolutional Neural Network</a></li>\n<li>[2] <a href=\"https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/discussion/430491\" target=\"_blank\">https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/discussion/430491</a></li>\n<li>[3] <a href=\"https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/discussion/430618\" target=\"_blank\">https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/discussion/430618</a></li>\n<li>[4] <a href=\"https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/discussion/430794\" target=\"_blank\">https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/discussion/430794</a></li>\n<li>[5] <a href=\"https://arxiv.org/abs/2004.01461\" target=\"_blank\">Gradient Centralization: A New Optimization Technique for Deep Neural Networks</a></li>\n<li>[6] <a href=\"https://discuss.pytorch.org/t/adam-half-precision-nans/1765/5\" target=\"_blank\">Adam+Half Precision = NaNs?</a></li>\n</ul>\n<h2>Appendix</h2>\n<h3>A. Tips to stabilize Transformer/CoaT T-Mixer</h3>\n<p>During my experiments with the Transformer and CoaT T-Mixer, I noticed numerous NaN values in the model's predictions. To stabilize the training, I implemented several techniques:</p>\n<ol>\n<li>Reduced the learning rate.</li>\n<li>Applied Gradient Centralization[5].</li>\n<li>Used an epsilon value of 1e-4 wherever feasible.</li>\n</ol>\n<p>Despite trying the first two techniques, the issue persisted.</p>\n<p>It's worth noting that I consistently used AMP (Automatic Mixed Precision) throughout my experiments. Upon realizing that the issue didn't manifest in FP32 precision, I suspected that an excessively large epsilon might be causing instability, especially since similar issues have been reported with FP16 training[6].</p>\n<p>Incorporating all these techniques resolved the problem, so I employed them in CoaT and Transformer T-Mixer experiments.</p>",
  "messages": [
    {
      "id": 2397864,
      "postDate": "2023-08-19T09:28:06.487Z",
      "content": "<p>Based on the knowledge shared on public discussion, I want to share the notable results which is obtained by experiments after the competition.<br>\nI appreciate participants who kindly shared their solutions.</p>\n<h2>Overview</h2>\n<p>First, my largest interest was</p>\n<ol>\n<li>How much gain can be obtained by correction of misalignment?</li>\n<li>How much gain can be obtained by soft labels?</li>\n</ol>\n<p>These results are already added to my solution write-up[4], so you can refer to it first if you haven't read.</p>\n<p>In addition, I tested performance gain of these architectural choices:</p>\n<ol>\n<li>backbones</li>\n<li>decoders, especially the impact of PixelShuffle</li>\n<li>T-Mixer (same as <em>Temporal Modulator</em> in my original solution)</li>\n</ol>\n<p>In this article, I discussed the experiment result of 3-5.</p>\n<h2>Experiment Detail</h2>\n<h3>Note</h3>\n<p>Given the constraints of computing resources, all experiments were conducted using a single fold and a single random state. This means the results might not be entirely robust. For more definitive conclusions, multiple trials for each architecture would be necessary.</p>\n<h3>Choice of backbones</h3>\n<p>I evaluated various backbones using a simple 2D U-Net architecture that only utilizes the current frame. The models were trained on <code>training</code> images and evaluated on <code>validation</code> images. For all experiments, I employed label correction and soft-labels. The results are presented in Table 1. Among the options, CoaT-lite medium demonstrated the best balance of performance and training time. Consequently, I adopted this backbone as the default for subsequent experiments.</p>\n<p><strong>Table 1: CV score of various backbones</strong></p>\n<table>\n<thead>\n<tr>\n<th>backbone</th>\n<th>crop_size</th>\n<th>epochs</th>\n<th>gradient checkpointing</th>\n<th>batch_size</th>\n<th>CV</th>\n<th>training time/epoch [sec]</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>efficientnet_b3</td>\n<td>512</td>\n<td>20</td>\n<td>FALSE</td>\n<td>28</td>\n<td>0.6629</td>\n<td>187</td>\n</tr>\n<tr>\n<td>maxvit_tiny_tf_512.in1k</td>\n<td>512</td>\n<td>20</td>\n<td>FALSE</td>\n<td>19</td>\n<td>0.6725</td>\n<td>396</td>\n</tr>\n<tr>\n<td>maxvit_small_tf_512.in1k</td>\n<td>512</td>\n<td>20</td>\n<td>TRUE</td>\n<td>21</td>\n<td>0.672</td>\n<td>709</td>\n</tr>\n<tr>\n<td>pvt_v2_b3</td>\n<td>512</td>\n<td>20</td>\n<td>FALSE</td>\n<td>20</td>\n<td>0.6681</td>\n<td>341</td>\n</tr>\n<tr>\n<td>coat_lite_medium</td>\n<td>512</td>\n<td>20</td>\n<td>FALSE</td>\n<td>20</td>\n<td>0.6787</td>\n<td>413</td>\n</tr>\n<tr>\n<td>coat_small</td>\n<td>512</td>\n<td>20</td>\n<td>FALSE</td>\n<td>19</td>\n<td>0.6713</td>\n<td>513</td>\n</tr>\n</tbody>\n</table>\n<h3>Changing Up-scaling to PixelShuffle</h3>\n<p>PixelShuffle[1] offers an alternative method for up-scaling resolution while reducing feature channels. Following the approach mentioned by the 2nd place team in their solution[2], I implemented the PixelShuffle block as a replacement for the traditional up-scaling block, leading to what I refer to as the PS-U-Net architecture.</p>\n<p>When comparing the U-Net with the PS-U-Net, there wasn't a significant performance gain (CV: 0.6787 to 0.6779). However, the PS-U-Net had the advantage of being more memory-efficient due to the reduced feature channels. Given this efficiency, I chose to use the PS-U-Net as the default architecture for subsequent experiments.</p>\n<h3>The Effects of Intensive Augmentation</h3>\n<p>Thanks to the correction of misalignment, we can now employ a more intensive augmentation strategy, as mentioned in the 1st place solution[3]. I adapted Koda's augmentation approach with slight modifications, as shown below:</p>\n<pre><code>cfg.geometric_transform = A.Compose(\n    [\n        A.RandomRotate90(p=),\n        A.HorizontalFlip(p=),\n        A.ShiftScaleRotate(\n            rotate_limit=, scale_limit=, border_mode=cv2.BORDER_CONSTANT, p=\n        ),\n    ]\n)\n</code></pre>\n<p>The differences are:</p>\n<ol>\n<li>slightly increase probability (0.5 -&gt; 0.6)</li>\n<li>set <code>borader_mode</code> as constant in order to avoid unintentional artifacts</li>\n</ol>\n<p>The experimental results highlighted the substantial impact of this intensive augmentation. It led to an increase in the CV score by 0.58% (from 0.6779 to 0.6837). This underscores that addressing misalignment is crucial for success in this competition.</p>\n<h3>Changing Decoder Architecture: PS-U-Net -&gt; PS-FPN</h3>\n<p>Next, I transitioned from the PS-U-Net architecture to the PS-FPN, where I replaced the up-scaling block of the FPN with PixelShuffle blocks.</p>\n<p>This change resulted in a modest improvement of 0.24% in CV scores (from 0.6837 to 0.6861). However, it's uncertain whether this gain can be solely attributed to the architectural modification. To draw a more definitive conclusion, averaging across multiple seeds would be advisable. Nonetheless, based on these preliminary results, I opted to use the PS-FPN architecture as the standard for subsequent experiments.</p>\n<h3>Comparing T-Mixer Architecture</h3>\n<p>Lastly, I tested the impact of different architectural choice of T-Mixer.</p>\n<p>Especially, I wanted to compare which mixer is superior to this task: Convolutions and Transformer. So I compared below three architecture:</p>\n<ol>\n<li>Conv3d</li>\n<li>Transformer</li>\n<li>CoaT with (3D positional encoding)</li>\n</ol>\n<p>The last pattern was just a spontaneous idea I had. However, since this backbone performs best in 2D settings, I believed it might be superior as temporal mixer. In this experiment, I evaluated on CV and LB. Additionally, I employed TTA to ensure robust evaluation.</p>\n<p>Note that training Transformer and CoaT T-Mixer were unstable, sometimes resulting in NaN predictions. Stability techniques were applied to address this (see Appendix A). In contrast, Conv3d T-Mixer training was stable without any adjustments.</p>\n<p>Table 2 displays the results. The first row represents a 2D U-Net without a T-Mixer, while the second uses a 2.5D PS-FPN with a Conv3D T-Mixer. Incorporating time frame information improved both CV and Private LB scores significantly. This indicates the value of temporal information even after alignment corrections.</p>\n<p>Table 2's 3rd and 4th rows show results for Transformer and CoaT T-Mixers. While CoaT T-Mixer had the highest CV score, its private LB score was the lowest. 3D Conv T-Mixer and Transformer T-Mixer had comparable performances.</p>\n<p>Table 3 presents ensemble model results. The best Private score came from an ensemble of three T-Mixer architectures, suggesting model diversity. Due to resource constraints, only one random seed was tested, but seed averaging or cross-fold averaging might yield better scores.</p>\n<p><strong>Table2: Comparison of Different T-Mixer Architectures</strong></p>\n<table>\n<thead>\n<tr>\n<th>id</th>\n<th>used frames</th>\n<th>T-Mixer</th>\n<th>seed</th>\n<th>TTA</th>\n<th>crop_size</th>\n<th>CV</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>[4]</td>\n<td>None</td>\n<td>x1</td>\n<td>hflip, rot(90, 180, 270)</td>\n<td>512</td>\n<td>0.691</td>\n<td>0.70244</td>\n<td>0.70462</td>\n</tr>\n<tr>\n<td>2</td>\n<td>[1, 2, 3, 4]</td>\n<td>3d Conv</td>\n<td>x1</td>\n<td>hflip, rot(90, 180, 270)</td>\n<td>512</td>\n<td>0.6978</td>\n<td>0.7147</td>\n<td>0.71681</td>\n</tr>\n<tr>\n<td>3</td>\n<td>[1, 2, 3, 4]</td>\n<td>Transformer</td>\n<td>x1</td>\n<td>hflip, rot(90, 180, 270)</td>\n<td>512</td>\n<td>0.6958</td>\n<td>0.70576</td>\n<td>0.71641</td>\n</tr>\n<tr>\n<td>4</td>\n<td>[1, 2, 3, 4]</td>\n<td>CoaT(3D pos enc)</td>\n<td>x1</td>\n<td>hflip, rot(90, 180, 270)</td>\n<td>512</td>\n<td>0.6988</td>\n<td>0.71553</td>\n<td>0.71229</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Table3: Result of Ensemble</strong></p>\n<table>\n<thead>\n<tr>\n<th>ensemble</th>\n<th>CV</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>[2, 4]</td>\n<td>0.7009</td>\n<td>0.718</td>\n<td>0.71827</td>\n</tr>\n<tr>\n<td>[2, 3, 4]</td>\n<td>0.701</td>\n<td>0.71638</td>\n<td>0.72156</td>\n</tr>\n</tbody>\n</table>\n<h2>References</h2>\n<ul>\n<li>[1] <a href=\"https://arxiv.org/abs/1609.05158v2\" target=\"_blank\">Real-Time Single Image and Video Super-Resolution Using an Efficient Sub-Pixel Convolutional Neural Network</a></li>\n<li>[2] <a href=\"https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/discussion/430491\" target=\"_blank\">https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/discussion/430491</a></li>\n<li>[3] <a href=\"https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/discussion/430618\" target=\"_blank\">https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/discussion/430618</a></li>\n<li>[4] <a href=\"https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/discussion/430794\" target=\"_blank\">https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/discussion/430794</a></li>\n<li>[5] <a href=\"https://arxiv.org/abs/2004.01461\" target=\"_blank\">Gradient Centralization: A New Optimization Technique for Deep Neural Networks</a></li>\n<li>[6] <a href=\"https://discuss.pytorch.org/t/adam-half-precision-nans/1765/5\" target=\"_blank\">Adam+Half Precision = NaNs?</a></li>\n</ul>\n<h2>Appendix</h2>\n<h3>A. Tips to stabilize Transformer/CoaT T-Mixer</h3>\n<p>During my experiments with the Transformer and CoaT T-Mixer, I noticed numerous NaN values in the model's predictions. To stabilize the training, I implemented several techniques:</p>\n<ol>\n<li>Reduced the learning rate.</li>\n<li>Applied Gradient Centralization[5].</li>\n<li>Used an epsilon value of 1e-4 wherever feasible.</li>\n</ol>\n<p>Despite trying the first two techniques, the issue persisted.</p>\n<p>It's worth noting that I consistently used AMP (Automatic Mixed Precision) throughout my experiments. Upon realizing that the issue didn't manifest in FP32 precision, I suspected that an excessively large epsilon might be causing instability, especially since similar issues have been reported with FP16 training[6].</p>\n<p>Incorporating all these techniques resolved the problem, so I employed them in CoaT and Transformer T-Mixer experiments.</p>",
      "rawMarkdown": "Based on the knowledge shared on public discussion, I want to share the notable results which is obtained by experiments after the competition.\nI appreciate participants who kindly shared their solutions.\n\n## Overview\n\nFirst, my largest interest was\n\n1. How much gain can be obtained by correction of misalignment?\n2. How much gain can be obtained by soft labels?\n\nThese results are already added to my solution write-up[4], so you can refer to it first if you haven't read.\n\nIn addition, I tested performance gain of these architectural choices:\n\n3. backbones\n4. decoders, especially the impact of PixelShuffle\n5. T-Mixer (same as *Temporal Modulator* in my original solution)\n\nIn this article, I discussed the experiment result of 3-5.\n\n## Experiment Detail\n\n### Note\n\nGiven the constraints of computing resources, all experiments were conducted using a single fold and a single random state. This means the results might not be entirely robust. For more definitive conclusions, multiple trials for each architecture would be necessary.\n\n### Choice of backbones\n\nI evaluated various backbones using a simple 2D U-Net architecture that only utilizes the current frame. The models were trained on `training` images and evaluated on `validation` images. For all experiments, I employed label correction and soft-labels. The results are presented in Table 1. Among the options, CoaT-lite medium demonstrated the best balance of performance and training time. Consequently, I adopted this backbone as the default for subsequent experiments.\n\n**Table 1: CV score of various backbones**\n\n| backbone                | crop_size | epochs | gradient checkpointing | batch_size | CV     | training time/epoch [sec] |\n|-------------------------|-----------|--------|---------------|------------|--------|----------------|\n| efficientnet_b3         | 512       | 20     | FALSE         | 28         | 0.6629 | 187            |\n| maxvit_tiny_tf_512.in1k | 512       | 20     | FALSE         | 19         | 0.6725 | 396            |\n| maxvit_small_tf_512.in1k| 512       | 20     | TRUE          | 21         | 0.672  | 709            |\n| pvt_v2_b3               | 512       | 20     | FALSE         | 20         | 0.6681 | 341            |\n| coat_lite_medium        | 512       | 20     | FALSE         | 20         | 0.6787 | 413            |\n| coat_small              | 512       | 20     | FALSE         | 19         | 0.6713 | 513            |\n\n\n### Changing Up-scaling to PixelShuffle\n\nPixelShuffle[1] offers an alternative method for up-scaling resolution while reducing feature channels. Following the approach mentioned by the 2nd place team in their solution[2], I implemented the PixelShuffle block as a replacement for the traditional up-scaling block, leading to what I refer to as the PS-U-Net architecture.\n\nWhen comparing the U-Net with the PS-U-Net, there wasn't a significant performance gain (CV: 0.6787 to 0.6779). However, the PS-U-Net had the advantage of being more memory-efficient due to the reduced feature channels. Given this efficiency, I chose to use the PS-U-Net as the default architecture for subsequent experiments.\n\n### The Effects of Intensive Augmentation\n\nThanks to the correction of misalignment, we can now employ a more intensive augmentation strategy, as mentioned in the 1st place solution[3]. I adapted Koda's augmentation approach with slight modifications, as shown below:\n\n\n```python\ncfg.geometric_transform = A.Compose(\n    [\n        A.RandomRotate90(p=1),\n        A.HorizontalFlip(p=0.5),\n        A.ShiftScaleRotate(\n            rotate_limit=45, scale_limit=0.2, border_mode=cv2.BORDER_CONSTANT, p=0.6\n        ),\n    ]\n)\n```\n\nThe differences are:\n1. slightly increase probability (0.5 -> 0.6)\n2. set `borader_mode` as constant in order to avoid unintentional artifacts\n\nThe experimental results highlighted the substantial impact of this intensive augmentation. It led to an increase in the CV score by 0.58% (from 0.6779 to 0.6837). This underscores that addressing misalignment is crucial for success in this competition.\n\n### Changing Decoder Architecture: PS-U-Net -> PS-FPN\n\nNext, I transitioned from the PS-U-Net architecture to the PS-FPN, where I replaced the up-scaling block of the FPN with PixelShuffle blocks.\n\nThis change resulted in a modest improvement of 0.24% in CV scores (from 0.6837 to 0.6861). However, it's uncertain whether this gain can be solely attributed to the architectural modification. To draw a more definitive conclusion, averaging across multiple seeds would be advisable. Nonetheless, based on these preliminary results, I opted to use the PS-FPN architecture as the standard for subsequent experiments.\n\n### Comparing T-Mixer Architecture\n\nLastly, I tested the impact of different architectural choice of T-Mixer.\n\nEspecially, I wanted to compare which mixer is superior to this task: Convolutions and Transformer. So I compared below three architecture:\n\n1. Conv3d\n2. Transformer\n3. CoaT with (3D positional encoding)\n\nThe last pattern was just a spontaneous idea I had. However, since this backbone performs best in 2D settings, I believed it might be superior as temporal mixer. In this experiment, I evaluated on CV and LB. Additionally, I employed TTA to ensure robust evaluation.\n\nNote that training Transformer and CoaT T-Mixer were unstable, sometimes resulting in NaN predictions. Stability techniques were applied to address this (see Appendix A). In contrast, Conv3d T-Mixer training was stable without any adjustments.\n\nTable 2 displays the results. The first row represents a 2D U-Net without a T-Mixer, while the second uses a 2.5D PS-FPN with a Conv3D T-Mixer. Incorporating time frame information improved both CV and Private LB scores significantly. This indicates the value of temporal information even after alignment corrections.\n\nTable 2's 3rd and 4th rows show results for Transformer and CoaT T-Mixers. While CoaT T-Mixer had the highest CV score, its private LB score was the lowest. 3D Conv T-Mixer and Transformer T-Mixer had comparable performances.\n\nTable 3 presents ensemble model results. The best Private score came from an ensemble of three T-Mixer architectures, suggesting model diversity. Due to resource constraints, only one random seed was tested, but seed averaging or cross-fold averaging might yield better scores.\n\n\n**Table2: Comparison of Different T-Mixer Architectures**\n\n| id | used frames    | T-Mixer           | seed | TTA                        | crop_size | CV     | Public  | Private |\n|----|----------------|-------------------|------|----------------------------|-----------|--------|---------|---------|\n| 1  | [4]            | None              | x1   | hflip, rot(90, 180, 270)   | 512       | 0.691  | 0.70244 | 0.70462 |\n| 2  | [1, 2, 3, 4]   | 3d Conv           | x1   | hflip, rot(90, 180, 270)   | 512       | 0.6978 | 0.7147  | 0.71681 |\n| 3  | [1, 2, 3, 4]   | Transformer       | x1   | hflip, rot(90, 180, 270)   | 512       | 0.6958 | 0.70576 | 0.71641 |\n| 4  | [1, 2, 3, 4]   | CoaT(3D pos enc)  | x1   | hflip, rot(90, 180, 270)   | 512       | 0.6988 | 0.71553 | 0.71229 |\n\n**Table3: Result of Ensemble**\n\n| ensemble   | CV     | Public | Private |\n|------------|--------|--------|---------|\n| [2, 4]     | 0.7009 | 0.718  | 0.71827 |\n| [2, 3, 4]  | 0.701  | 0.71638| 0.72156 |\n\n## References\n\n* [1] [Real-Time Single Image and Video Super-Resolution Using an Efficient Sub-Pixel Convolutional Neural Network](https://arxiv.org/abs/1609.05158v2)\n* [2] https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/discussion/430491\n* [3] https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/discussion/430618\n* [4] https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/discussion/430794\n* [5] [Gradient Centralization: A New Optimization Technique for Deep Neural Networks](https://arxiv.org/abs/2004.01461)\n* [6] [Adam+Half Precision = NaNs?](https://discuss.pytorch.org/t/adam-half-precision-nans/1765/5)\n\n## Appendix\n\n### A. Tips to stabilize Transformer/CoaT T-Mixer\n\nDuring my experiments with the Transformer and CoaT T-Mixer, I noticed numerous NaN values in the model's predictions. To stabilize the training, I implemented several techniques:\n\n1. Reduced the learning rate.\n2. Applied Gradient Centralization[5].\n3. Used an epsilon value of 1e-4 wherever feasible.\n\nDespite trying the first two techniques, the issue persisted.\n\nIt's worth noting that I consistently used AMP (Automatic Mixed Precision) throughout my experiments. Upon realizing that the issue didn't manifest in FP32 precision, I suspected that an excessively large epsilon might be causing instability, especially since similar issues have been reported with FP16 training[6].\n\nIncorporating all these techniques resolved the problem, so I employed them in CoaT and Transformer T-Mixer experiments.\n",
      "votes": 27
    },
    {
      "id": 2400844,
      "postDate": "2023-08-21T09:03:26.763Z",
      "content": "<p>Awesome stuff. Reproducing what worked for others is no easy task.</p>",
      "rawMarkdown": "Awesome stuff. Reproducing what worked for others is no easy task.",
      "votes": 1,
      "replies": [
        {
          "id": 2402214,
          "postDate": "2023-08-22T04:10:06.943Z",
          "content": "<p>Thanks. I have a lot of inspiration from your team's solution. Thank you for sharing this informative write-up.</p>",
          "rawMarkdown": "Thanks. I have a lot of inspiration from your team's solution. Thank you for sharing this informative write-up.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2397968,
      "postDate": "2023-08-19T11:03:06.053Z",
      "content": "<p>What a great idea and writeup! Thanks!<br>\nCan you share the code for the 3D Conv T-Mixer part?</p>",
      "rawMarkdown": "What a great idea and writeup! Thanks!\nCan you share the code for the 3D Conv T-Mixer part?",
      "votes": 1,
      "replies": [
        {
          "id": 2399874,
          "postDate": "2023-08-20T16:40:25.397Z",
          "content": "<p>Or share the entire code! what a great ablation study! 😀</p>\n<blockquote>\n  <p>Used an epsilon value of 1e-4 wherever feasible</p>\n</blockquote>\n<p>Btw, what epsilon value are you refering to?</p>",
          "rawMarkdown": "Or share the entire code! what a great ablation study! 😀\n\n>Used an epsilon value of 1e-4 wherever feasible\n\nBtw, what epsilon value are you refering to?\n",
          "votes": 1,
          "replies": [
            {
              "id": 2400019,
              "postDate": "2023-08-20T18:21:28.793Z",
              "content": "<p>A higher value of epsilon in AdamW and the like helps with NaNs when doing mixed precision training.</p>",
              "rawMarkdown": "A higher value of epsilon in AdamW and the like helps with NaNs when doing mixed precision training.",
              "votes": 1
            },
            {
              "id": 2400024,
              "postDate": "2023-08-20T18:24:38.367Z",
              "content": "<blockquote>\n  <p>Or share the entire code! what a great ablation study! 😀</p>\n</blockquote>\n<p>Yep!<br>\nI'm asking because I've tried 3d time mixer (just a couple of 3d convs+2d reshape applied to the stride 32 encoder layer) and it was super expensive in terms of memory consumption and speed.</p>",
              "rawMarkdown": "> Or share the entire code! what a great ablation study! 😀\n\nYep!\nI'm asking because I've tried 3d time mixer (just a couple of 3d convs+2d reshape applied to the stride 32 encoder layer) and it was super expensive in terms of memory consumption and speed.\n\n",
              "votes": 1
            },
            {
              "id": 2402184,
              "postDate": "2023-08-22T03:36:58.347Z",
              "content": "<blockquote>\n  <p>Can you share the code for the 3D Conv T-Mixer part?</p>\n</blockquote>\n<p>Well, I can share this part, but it's nothing special. its just as I wrote on my original solution's block diagram[4]. I think you can easily implement it by yourself.</p>\n<blockquote>\n  <p>Or share the entire code! what a great ablation study! 😀</p>\n</blockquote>\n<p>Sorry, but I am now working on other competition, and I don't have much time for that.<br>\nYou know, my code is full of external code, and I should follow licenses for each repository.</p>\n<blockquote>\n  <p>I'm asking because I've tried 3d time mixer (just a couple of 3d convs+2d reshape applied to the stride 32 encoder layer) and it was super expensive in terms of memory consumption and speed.</p>\n</blockquote>\n<p>If you have problem with memory consumption, my advice is to use gradient checkpointing[7]. For example, some backbone in timm library already be implemented with this functionality.</p>\n<ul>\n<li>[7] <a href=\"https://pytorch.org/docs/stable/checkpoint.html\" target=\"_blank\">https://pytorch.org/docs/stable/checkpoint.html</a></li>\n</ul>",
              "rawMarkdown": "> Can you share the code for the 3D Conv T-Mixer part?\n\nWell, I can share this part, but it's nothing special. its just as I wrote on my original solution's block diagram[4]. I think you can easily implement it by yourself.\n\n> Or share the entire code! what a great ablation study! 😀\n\nSorry, but I am now working on other competition, and I don't have much time for that.\nYou know, my code is full of external code, and I should follow licenses for each repository.\n\n> I'm asking because I've tried 3d time mixer (just a couple of 3d convs+2d reshape applied to the stride 32 encoder layer) and it was super expensive in terms of memory consumption and speed.\n\nIf you have problem with memory consumption, my advice is to use gradient checkpointing[7]. For example, some backbone in timm library already be implemented with this functionality.\n\n- [7] https://pytorch.org/docs/stable/checkpoint.html"
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2400844,
      "author_name": "Theo Viel",
      "author_url": "",
      "post_date": "2023-08-21T09:03:26.763000",
      "content": "<p>Awesome stuff. Reproducing what worked for others is no easy task.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2402214,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2023-08-22T04:10:06.943000",
          "content": "<p>Thanks. I have a lot of inspiration from your team's solution. Thank you for sharing this informative write-up.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2397968,
      "author_name": "DennisSakva",
      "author_url": "",
      "post_date": "2023-08-19T11:03:06.053000",
      "content": "<p>What a great idea and writeup! Thanks!<br>\nCan you share the code for the 3D Conv T-Mixer part?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2399874,
          "author_name": "delai50",
          "author_url": "",
          "post_date": "2023-08-20T16:40:25.397000",
          "content": "<p>Or share the entire code! what a great ablation study! 😀</p>\n<blockquote>\n  <p>Used an epsilon value of 1e-4 wherever feasible</p>\n</blockquote>\n<p>Btw, what epsilon value are you refering to?</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2400019,
              "author_name": "DennisSakva",
              "author_url": "",
              "post_date": "2023-08-20T18:21:28.793000",
              "content": "<p>A higher value of epsilon in AdamW and the like helps with NaNs when doing mixed precision training.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2400024,
              "author_name": "DennisSakva",
              "author_url": "",
              "post_date": "2023-08-20T18:24:38.367000",
              "content": "<blockquote>\n  <p>Or share the entire code! what a great ablation study! 😀</p>\n</blockquote>\n<p>Yep!<br>\nI'm asking because I've tried 3d time mixer (just a couple of 3d convs+2d reshape applied to the stride 32 encoder layer) and it was super expensive in terms of memory consumption and speed.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2402184,
              "author_name": "Bilzard",
              "author_url": "",
              "post_date": "2023-08-22T03:36:58.347000",
              "content": "<blockquote>\n  <p>Can you share the code for the 3D Conv T-Mixer part?</p>\n</blockquote>\n<p>Well, I can share this part, but it's nothing special. its just as I wrote on my original solution's block diagram[4]. I think you can easily implement it by yourself.</p>\n<blockquote>\n  <p>Or share the entire code! what a great ablation study! 😀</p>\n</blockquote>\n<p>Sorry, but I am now working on other competition, and I don't have much time for that.<br>\nYou know, my code is full of external code, and I should follow licenses for each repository.</p>\n<blockquote>\n  <p>I'm asking because I've tried 3d time mixer (just a couple of 3d convs+2d reshape applied to the stride 32 encoder layer) and it was super expensive in terms of memory consumption and speed.</p>\n</blockquote>\n<p>If you have problem with memory consumption, my advice is to use gradient checkpointing[7]. For example, some backbone in timm library already be implemented with this functionality.</p>\n<ul>\n<li>[7] <a href=\"https://pytorch.org/docs/stable/checkpoint.html\" target=\"_blank\">https://pytorch.org/docs/stable/checkpoint.html</a></li>\n</ul>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2397864": "Based on the knowledge shared on public discussion, I want to share the notable results which is obtained by experiments after the competition.\nI appreciate participants who kindly shared their solutions.\n\n## Overview\n\nFirst, my largest interest was\n\n1. How much gain can be obtained by correction of misalignment?\n2. How much gain can be obtained by soft labels?\n\nThese results are already added to my solution write-up[4], so you can refer to it first if you haven't read.\n\nIn addition, I tested performance gain of these architectural choices:\n\n3. backbones\n4. decoders, especially the impact of PixelShuffle\n5. T-Mixer (same as *Temporal Modulator* in my original solution)\n\nIn this article, I discussed the experiment result of 3-5.\n\n## Experiment Detail\n\n### Note\n\nGiven the constraints of computing resources, all experiments were conducted using a single fold and a single random state. This means the results might not be entirely robust. For more definitive conclusions, multiple trials for each architecture would be necessary.\n\n### Choice of backbones\n\nI evaluated various backbones using a simple 2D U-Net architecture that only utilizes the current frame. The models were trained on `training` images and evaluated on `validation` images. For all experiments, I employed label correction and soft-labels. The results are presented in Table 1. Among the options, CoaT-lite medium demonstrated the best balance of performance and training time. Consequently, I adopted this backbone as the default for subsequent experiments.\n\n**Table 1: CV score of various backbones**\n\n| backbone                | crop_size | epochs | gradient checkpointing | batch_size | CV     | training time/epoch [sec] |\n|-------------------------|-----------|--------|---------------|------------|--------|----------------|\n| efficientnet_b3         | 512       | 20     | FALSE         | 28         | 0.6629 | 187            |\n| maxvit_tiny_tf_512.in1k | 512       | 20     | FALSE         | 19         | 0.6725 | 396            |\n| maxvit_small_tf_512.in1k| 512       | 20     | TRUE          | 21         | 0.672  | 709            |\n| pvt_v2_b3               | 512       | 20     | FALSE         | 20         | 0.6681 | 341            |\n| coat_lite_medium        | 512       | 20     | FALSE         | 20         | 0.6787 | 413            |\n| coat_small              | 512       | 20     | FALSE         | 19         | 0.6713 | 513            |\n\n\n### Changing Up-scaling to PixelShuffle\n\nPixelShuffle[1] offers an alternative method for up-scaling resolution while reducing feature channels. Following the approach mentioned by the 2nd place team in their solution[2], I implemented the PixelShuffle block as a replacement for the traditional up-scaling block, leading to what I refer to as the PS-U-Net architecture.\n\nWhen comparing the U-Net with the PS-U-Net, there wasn't a significant performance gain (CV: 0.6787 to 0.6779). However, the PS-U-Net had the advantage of being more memory-efficient due to the reduced feature channels. Given this efficiency, I chose to use the PS-U-Net as the default architecture for subsequent experiments.\n\n### The Effects of Intensive Augmentation\n\nThanks to the correction of misalignment, we can now employ a more intensive augmentation strategy, as mentioned in the 1st place solution[3]. I adapted Koda's augmentation approach with slight modifications, as shown below:\n\n\n```python\ncfg.geometric_transform = A.Compose(\n    [\n        A.RandomRotate90(p=1),\n        A.HorizontalFlip(p=0.5),\n        A.ShiftScaleRotate(\n            rotate_limit=45, scale_limit=0.2, border_mode=cv2.BORDER_CONSTANT, p=0.6\n        ),\n    ]\n)\n```\n\nThe differences are:\n1. slightly increase probability (0.5 -> 0.6)\n2. set `borader_mode` as constant in order to avoid unintentional artifacts\n\nThe experimental results highlighted the substantial impact of this intensive augmentation. It led to an increase in the CV score by 0.58% (from 0.6779 to 0.6837). This underscores that addressing misalignment is crucial for success in this competition.\n\n### Changing Decoder Architecture: PS-U-Net -> PS-FPN\n\nNext, I transitioned from the PS-U-Net architecture to the PS-FPN, where I replaced the up-scaling block of the FPN with PixelShuffle blocks.\n\nThis change resulted in a modest improvement of 0.24% in CV scores (from 0.6837 to 0.6861). However, it's uncertain whether this gain can be solely attributed to the architectural modification. To draw a more definitive conclusion, averaging across multiple seeds would be advisable. Nonetheless, based on these preliminary results, I opted to use the PS-FPN architecture as the standard for subsequent experiments.\n\n### Comparing T-Mixer Architecture\n\nLastly, I tested the impact of different architectural choice of T-Mixer.\n\nEspecially, I wanted to compare which mixer is superior to this task: Convolutions and Transformer. So I compared below three architecture:\n\n1. Conv3d\n2. Transformer\n3. CoaT with (3D positional encoding)\n\nThe last pattern was just a spontaneous idea I had. However, since this backbone performs best in 2D settings, I believed it might be superior as temporal mixer. In this experiment, I evaluated on CV and LB. Additionally, I employed TTA to ensure robust evaluation.\n\nNote that training Transformer and CoaT T-Mixer were unstable, sometimes resulting in NaN predictions. Stability techniques were applied to address this (see Appendix A). In contrast, Conv3d T-Mixer training was stable without any adjustments.\n\nTable 2 displays the results. The first row represents a 2D U-Net without a T-Mixer, while the second uses a 2.5D PS-FPN with a Conv3D T-Mixer. Incorporating time frame information improved both CV and Private LB scores significantly. This indicates the value of temporal information even after alignment corrections.\n\nTable 2's 3rd and 4th rows show results for Transformer and CoaT T-Mixers. While CoaT T-Mixer had the highest CV score, its private LB score was the lowest. 3D Conv T-Mixer and Transformer T-Mixer had comparable performances.\n\nTable 3 presents ensemble model results. The best Private score came from an ensemble of three T-Mixer architectures, suggesting model diversity. Due to resource constraints, only one random seed was tested, but seed averaging or cross-fold averaging might yield better scores.\n\n\n**Table2: Comparison of Different T-Mixer Architectures**\n\n| id | used frames    | T-Mixer           | seed | TTA                        | crop_size | CV     | Public  | Private |\n|----|----------------|-------------------|------|----------------------------|-----------|--------|---------|---------|\n| 1  | [4]            | None              | x1   | hflip, rot(90, 180, 270)   | 512       | 0.691  | 0.70244 | 0.70462 |\n| 2  | [1, 2, 3, 4]   | 3d Conv           | x1   | hflip, rot(90, 180, 270)   | 512       | 0.6978 | 0.7147  | 0.71681 |\n| 3  | [1, 2, 3, 4]   | Transformer       | x1   | hflip, rot(90, 180, 270)   | 512       | 0.6958 | 0.70576 | 0.71641 |\n| 4  | [1, 2, 3, 4]   | CoaT(3D pos enc)  | x1   | hflip, rot(90, 180, 270)   | 512       | 0.6988 | 0.71553 | 0.71229 |\n\n**Table3: Result of Ensemble**\n\n| ensemble   | CV     | Public | Private |\n|------------|--------|--------|---------|\n| [2, 4]     | 0.7009 | 0.718  | 0.71827 |\n| [2, 3, 4]  | 0.701  | 0.71638| 0.72156 |\n\n## References\n\n* [1] [Real-Time Single Image and Video Super-Resolution Using an Efficient Sub-Pixel Convolutional Neural Network](https://arxiv.org/abs/1609.05158v2)\n* [2] https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/discussion/430491\n* [3] https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/discussion/430618\n* [4] https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/discussion/430794\n* [5] [Gradient Centralization: A New Optimization Technique for Deep Neural Networks](https://arxiv.org/abs/2004.01461)\n* [6] [Adam+Half Precision = NaNs?](https://discuss.pytorch.org/t/adam-half-precision-nans/1765/5)\n\n## Appendix\n\n### A. Tips to stabilize Transformer/CoaT T-Mixer\n\nDuring my experiments with the Transformer and CoaT T-Mixer, I noticed numerous NaN values in the model's predictions. To stabilize the training, I implemented several techniques:\n\n1. Reduced the learning rate.\n2. Applied Gradient Centralization[5].\n3. Used an epsilon value of 1e-4 wherever feasible.\n\nDespite trying the first two techniques, the issue persisted.\n\nIt's worth noting that I consistently used AMP (Automatic Mixed Precision) throughout my experiments. Upon realizing that the issue didn't manifest in FP32 precision, I suspected that an excessively large epsilon might be causing instability, especially since similar issues have been reported with FP16 training[6].\n\nIncorporating all these techniques resolved the problem, so I employed them in CoaT and Transformer T-Mixer experiments.\n",
    "2400844": "Awesome stuff. Reproducing what worked for others is no easy task.",
    "2397968": "What a great idea and writeup! Thanks!\nCan you share the code for the 3D Conv T-Mixer part?"
  }
}