{
  "id": 430618,
  "title": "1st place solution",
  "url": "/competitions/google-research-identify-contrails-reduce-global-warming/discussion/430618",
  "author_name": "🐢 Jun Koda",
  "post_date": "2023-08-10T13:25:48.517000",
  "votes": 164,
  "comment_count": 53,
  "views": 0,
  "content": "<p>I thank the organizer and Kaggle for hosting this interesting competition. I also thank many Kagglers who wrote solutions in past competitions, without which I could not compete for segmentation tasks.</p>\n<h2>Overview</h2>\n<ul>\n<li>U-Net with MaxViT encoder</li>\n<li>Input is the standard ash color image</li>\n<li>Target y is soft; the average of individual masks</li>\n<li>BCE (binary cross entropy) loss</li>\n<li>U-Net trained for symmetrized label, y_sym, which is the label shifted by 0.5 pixels</li>\n<li>Shift-scale-rotate augmentation</li>\n<li>Additional tiny convolution trained to map y_sym to y</li>\n</ul>\n<h2>Rotation augmentation</h2>\n<p>Augmentation is extremely important in this competition to suppress overfit and train longer. With rotation augmentation I could train 40-50 epochs, compared to 10-20 epochs without augmentation. The test-time augmentation (TTA) is also very effective, adding ~0.006 to the score for free.</p>\n<p>This is basic but does not work as usual in this competition because the label is shifted 0.5 pixels to the right and bottom with respect to the contrails. Random rotation augmentation would shift the label to random directions and make the model impossible to learn the right-bottom aligned ground-truth labels.</p>\n<h2>Discover the 0.5-pixel shift</h2>\n<p>I was very confused when the score dropped with flip and rot90 (multiples of 90° rotation) augmentations. In physics, what symmetry the system has is the first thing to consider, and although westerly winds or Coriolis force could make the physics asymmetric under flip or rotation, I could not believe that those affect the contrail detection. I visualized a prediction with a model trained without augmentation and applied it to a 180°-rotated input image to see why the augmentation did not work.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F954117%2Ffe56c0bee942eb6591adabafbe0daf39%2Fcontrail_label_shift.png?generation=1691669316963209&amp;alt=media\" alt=\"\"> </p>\n<p>The blue-green-red stripe pattern from top to bottom shows false negative, true positive, and false positive, which means that the rotated label is shifted up compared to the predicted label; that is, the original labels are shifted down compared to the contrails. I also observed the same pattern for left and right.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F954117%2Fbb92508e8a6b1d8392819ac63cb36cd6%2Fcontrail_fish.png?generation=1691669701057157&amp;alt=media\" alt=\"\"></p>\n<p>If the labels are shifted 0.5 pixels to top left (right panel in the figure above), then the rotated label is consistent with the original label in terms of the contrail-label offset.</p>\n<ul>\n<li>I first shift the label by 0.5 pixels and create y_sym</li>\n<li>y_sym has size 512×512 and is sampled from y on a shifted regular grid with bilinear interpolation</li>\n<li>U-Net is trained with y_sym </li>\n<li>Additional tiny convolution of 5x5 with stride 2 is trained to learn the mapping from y_sym to right-bottom aligned y, using 5% of the data when the augmentations are not applied randomly (note that y cannot be augmented).</li>\n<li>At inference time, y_sym is averaged for 8 patterns of rot90 and flip TTAs. The conv5x5 is applied after the y_sym.</li>\n</ul>\n<p>The augmentation is,</p>\n<pre><code>import albumentations as A\n\nA.Compose([\n    A.RandomRotate90(=1),\n    A.HorizontalFlip(=0.5),\n    A.ShiftScaleRotate(=30, =0.2)\n])\n</code></pre>\n<p>I refer to my public notebook for the details:</p>\n<p><a href=\"https://www.kaggle.com/junkoda/base-unet-model-for-the-1st-place\" target=\"_blank\">https://www.kaggle.com/junkoda/base-unet-model-for-the-1st-place</a></p>\n<h2>Models</h2>\n<p>The final prediction is the weighted mean of two models with a threshold ~0.45. I tuned the threshold and the weights using the validation set. Both models are U-Net using maxvit_tiny_tf_512.in1k as the encoder, but one uses single time t=4, and the other use four times t = 1 - 4. The input image size is 1024×1024 for both models.</p>\n<h2>Single-time model</h2>\n<ul>\n<li>Upscale the input image to 1024×1024</li>\n<li>Drop the final decoder layer of upscale and output 512×512</li>\n</ul>\n<h2>4-panel model</h2>\n<p>In order to use the time information, I tried 3D ResNet and ConvLSTM, but I was not able to make them work at all. The only thing I could do was to pack four 512×512 image at t=1,2,3,4 into one 1024×1024 image and hope that self attention in MaxViT look at different times.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F954117%2F291ed2fb89c1e4b492f0d5890ad06868%2Fimg4.png?generation=1691671228909061&amp;alt=media\" alt=\"\"></p>\n<p>In the U-Net, I only pass the quarter of the features (H/2, W/2) to the decoder, which corresponds to the t=4 quarter. The rest is the same as the single-time model. Concatenating images in the spacial direction appear in Kaggle once in a while. I remember <a href=\"https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275433\" target=\"_blank\">CPMP's solution</a> for the G2Net blackhole merger competition, in which he stacked 3 detector data horizontally rather than stacking them as 3 channels. </p>\n<p>Training details:</p>\n<ul>\n<li>Train 5 out of k=10 folds, using the out-of-fold to monitor the validation loss</li>\n<li>0.5 epochs of linear warmup to learning rate 8e-4, followed by cosine annealing</li>\n<li>batch size 4 × gradient accumulation 2</li>\n<li>AdamW with weight_decay = 0.01</li>\n<li>35 or 40 epochs of training</li>\n</ul>\n<pre><code>Scores\n                    CV     Public  Private\n       input size\n\nSingle time   512  0.697   0.707   0.712\nSingle time  1024  0.703   0.719   0.716\n4-panel      1024  0.704   0.719   0.722\nEnsemble           0.706   0.725   0.724\n</code></pre>\n<p>All scores are for 5-fold × 8 TTA mean.</p>\n<h2>Things I was not able to do</h2>\n<ul>\n<li>I did not use pseudo labeling. I tried pseudo labeling at an early stage. It boosted the single model score a lot, but the benefit was unimpressive after 5-fold mean.</li>\n<li>I was not able to train large models. I tried MaxViT small and base, and many other models, but the improvement was unclear. I increased the input image size instead. This failure must be due to my insufficient experience. </li>\n<li>I used only positive data (dropping data with no positive labels) for model evaluation because it takes 1/2 time to train. However, I was not able to use it in the real prediction; it performs very bad for negative samples, and my classifier was not good enough to exclude such false negatives.</li>\n<li>Removing small masks is a basic technique in segmentation competition, but here, removing even 1-pixel clusters (connected components) was not a good idea. I tried to remove false positive clusters based on some cluster statistics (size, density, max density, etc), but U-Net was cleverer than my post processing.</li>\n</ul>\n<h2>Codes</h2>\n<ul>\n<li>Training: <a href=\"https://github.com/junkoda/kaggle_contrails_solution\" target=\"_blank\">https://github.com/junkoda/kaggle_contrails_solution</a></li>\n<li>Inference notebook: <a href=\"https://www.kaggle.com/code/junkoda/contrails-submit\" target=\"_blank\">https://www.kaggle.com/code/junkoda/contrails-submit</a></li>\n</ul>\n<p>Update 2023-08-20: Add links to the codes and minor english corrections.</p>",
  "messages": [
    {
      "id": 2383620,
      "postDate": "2023-08-10T13:25:48.517Z",
      "content": "<p>I thank the organizer and Kaggle for hosting this interesting competition. I also thank many Kagglers who wrote solutions in past competitions, without which I could not compete for segmentation tasks.</p>\n<h2>Overview</h2>\n<ul>\n<li>U-Net with MaxViT encoder</li>\n<li>Input is the standard ash color image</li>\n<li>Target y is soft; the average of individual masks</li>\n<li>BCE (binary cross entropy) loss</li>\n<li>U-Net trained for symmetrized label, y_sym, which is the label shifted by 0.5 pixels</li>\n<li>Shift-scale-rotate augmentation</li>\n<li>Additional tiny convolution trained to map y_sym to y</li>\n</ul>\n<h2>Rotation augmentation</h2>\n<p>Augmentation is extremely important in this competition to suppress overfit and train longer. With rotation augmentation I could train 40-50 epochs, compared to 10-20 epochs without augmentation. The test-time augmentation (TTA) is also very effective, adding ~0.006 to the score for free.</p>\n<p>This is basic but does not work as usual in this competition because the label is shifted 0.5 pixels to the right and bottom with respect to the contrails. Random rotation augmentation would shift the label to random directions and make the model impossible to learn the right-bottom aligned ground-truth labels.</p>\n<h2>Discover the 0.5-pixel shift</h2>\n<p>I was very confused when the score dropped with flip and rot90 (multiples of 90° rotation) augmentations. In physics, what symmetry the system has is the first thing to consider, and although westerly winds or Coriolis force could make the physics asymmetric under flip or rotation, I could not believe that those affect the contrail detection. I visualized a prediction with a model trained without augmentation and applied it to a 180°-rotated input image to see why the augmentation did not work.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F954117%2Ffe56c0bee942eb6591adabafbe0daf39%2Fcontrail_label_shift.png?generation=1691669316963209&amp;alt=media\" alt=\"\"> </p>\n<p>The blue-green-red stripe pattern from top to bottom shows false negative, true positive, and false positive, which means that the rotated label is shifted up compared to the predicted label; that is, the original labels are shifted down compared to the contrails. I also observed the same pattern for left and right.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F954117%2Fbb92508e8a6b1d8392819ac63cb36cd6%2Fcontrail_fish.png?generation=1691669701057157&amp;alt=media\" alt=\"\"></p>\n<p>If the labels are shifted 0.5 pixels to top left (right panel in the figure above), then the rotated label is consistent with the original label in terms of the contrail-label offset.</p>\n<ul>\n<li>I first shift the label by 0.5 pixels and create y_sym</li>\n<li>y_sym has size 512×512 and is sampled from y on a shifted regular grid with bilinear interpolation</li>\n<li>U-Net is trained with y_sym </li>\n<li>Additional tiny convolution of 5x5 with stride 2 is trained to learn the mapping from y_sym to right-bottom aligned y, using 5% of the data when the augmentations are not applied randomly (note that y cannot be augmented).</li>\n<li>At inference time, y_sym is averaged for 8 patterns of rot90 and flip TTAs. The conv5x5 is applied after the y_sym.</li>\n</ul>\n<p>The augmentation is,</p>\n<pre><code>import albumentations as A\n\nA.Compose([\n    A.RandomRotate90(=1),\n    A.HorizontalFlip(=0.5),\n    A.ShiftScaleRotate(=30, =0.2)\n])\n</code></pre>\n<p>I refer to my public notebook for the details:</p>\n<p><a href=\"https://www.kaggle.com/junkoda/base-unet-model-for-the-1st-place\" target=\"_blank\">https://www.kaggle.com/junkoda/base-unet-model-for-the-1st-place</a></p>\n<h2>Models</h2>\n<p>The final prediction is the weighted mean of two models with a threshold ~0.45. I tuned the threshold and the weights using the validation set. Both models are U-Net using maxvit_tiny_tf_512.in1k as the encoder, but one uses single time t=4, and the other use four times t = 1 - 4. The input image size is 1024×1024 for both models.</p>\n<h2>Single-time model</h2>\n<ul>\n<li>Upscale the input image to 1024×1024</li>\n<li>Drop the final decoder layer of upscale and output 512×512</li>\n</ul>\n<h2>4-panel model</h2>\n<p>In order to use the time information, I tried 3D ResNet and ConvLSTM, but I was not able to make them work at all. The only thing I could do was to pack four 512×512 image at t=1,2,3,4 into one 1024×1024 image and hope that self attention in MaxViT look at different times.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F954117%2F291ed2fb89c1e4b492f0d5890ad06868%2Fimg4.png?generation=1691671228909061&amp;alt=media\" alt=\"\"></p>\n<p>In the U-Net, I only pass the quarter of the features (H/2, W/2) to the decoder, which corresponds to the t=4 quarter. The rest is the same as the single-time model. Concatenating images in the spacial direction appear in Kaggle once in a while. I remember <a href=\"https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275433\" target=\"_blank\">CPMP's solution</a> for the G2Net blackhole merger competition, in which he stacked 3 detector data horizontally rather than stacking them as 3 channels. </p>\n<p>Training details:</p>\n<ul>\n<li>Train 5 out of k=10 folds, using the out-of-fold to monitor the validation loss</li>\n<li>0.5 epochs of linear warmup to learning rate 8e-4, followed by cosine annealing</li>\n<li>batch size 4 × gradient accumulation 2</li>\n<li>AdamW with weight_decay = 0.01</li>\n<li>35 or 40 epochs of training</li>\n</ul>\n<pre><code>Scores\n                    CV     Public  Private\n       input size\n\nSingle time   512  0.697   0.707   0.712\nSingle time  1024  0.703   0.719   0.716\n4-panel      1024  0.704   0.719   0.722\nEnsemble           0.706   0.725   0.724\n</code></pre>\n<p>All scores are for 5-fold × 8 TTA mean.</p>\n<h2>Things I was not able to do</h2>\n<ul>\n<li>I did not use pseudo labeling. I tried pseudo labeling at an early stage. It boosted the single model score a lot, but the benefit was unimpressive after 5-fold mean.</li>\n<li>I was not able to train large models. I tried MaxViT small and base, and many other models, but the improvement was unclear. I increased the input image size instead. This failure must be due to my insufficient experience. </li>\n<li>I used only positive data (dropping data with no positive labels) for model evaluation because it takes 1/2 time to train. However, I was not able to use it in the real prediction; it performs very bad for negative samples, and my classifier was not good enough to exclude such false negatives.</li>\n<li>Removing small masks is a basic technique in segmentation competition, but here, removing even 1-pixel clusters (connected components) was not a good idea. I tried to remove false positive clusters based on some cluster statistics (size, density, max density, etc), but U-Net was cleverer than my post processing.</li>\n</ul>\n<h2>Codes</h2>\n<ul>\n<li>Training: <a href=\"https://github.com/junkoda/kaggle_contrails_solution\" target=\"_blank\">https://github.com/junkoda/kaggle_contrails_solution</a></li>\n<li>Inference notebook: <a href=\"https://www.kaggle.com/code/junkoda/contrails-submit\" target=\"_blank\">https://www.kaggle.com/code/junkoda/contrails-submit</a></li>\n</ul>\n<p>Update 2023-08-20: Add links to the codes and minor english corrections.</p>",
      "rawMarkdown": "I thank the organizer and Kaggle for hosting this interesting competition. I also thank many Kagglers who wrote solutions in past competitions, without which I could not compete for segmentation tasks.\n\n## Overview\n\n* U-Net with MaxViT encoder\n* Input is the standard ash color image\n* Target y is soft; the average of individual masks\n* BCE (binary cross entropy) loss\n* U-Net trained for symmetrized label, y_sym, which is the label shifted by 0.5 pixels\n* Shift-scale-rotate augmentation\n* Additional tiny convolution trained to map y_sym to y\n\n## Rotation augmentation\n\nAugmentation is extremely important in this competition to suppress overfit and train longer. With rotation augmentation I could train 40-50 epochs, compared to 10-20 epochs without augmentation. The test-time augmentation (TTA) is also very effective, adding ~0.006 to the score for free.\n\nThis is basic but does not work as usual in this competition because the label is shifted 0.5 pixels to the right and bottom with respect to the contrails. Random rotation augmentation would shift the label to random directions and make the model impossible to learn the right-bottom aligned ground-truth labels.\n\n## Discover the 0.5-pixel shift\n\nI was very confused when the score dropped with flip and rot90 (multiples of 90° rotation) augmentations. In physics, what symmetry the system has is the first thing to consider, and although westerly winds or Coriolis force could make the physics asymmetric under flip or rotation, I could not believe that those affect the contrail detection. I visualized a prediction with a model trained without augmentation and applied it to a 180°-rotated input image to see why the augmentation did not work.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F954117%2Ffe56c0bee942eb6591adabafbe0daf39%2Fcontrail_label_shift.png?generation=1691669316963209&alt=media) \n\nThe blue-green-red stripe pattern from top to bottom shows false negative, true positive, and false positive, which means that the rotated label is shifted up compared to the predicted label; that is, the original labels are shifted down compared to the contrails. I also observed the same pattern for left and right.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F954117%2Fbb92508e8a6b1d8392819ac63cb36cd6%2Fcontrail_fish.png?generation=1691669701057157&alt=media)\n\nIf the labels are shifted 0.5 pixels to top left (right panel in the figure above), then the rotated label is consistent with the original label in terms of the contrail-label offset.\n\n* I first shift the label by 0.5 pixels and create y_sym\n* y_sym has size 512×512 and is sampled from y on a shifted regular grid with bilinear interpolation\n* U-Net is trained with y_sym \n* Additional tiny convolution of 5x5 with stride 2 is trained to learn the mapping from y_sym to right-bottom aligned y, using 5% of the data when the augmentations are not applied randomly (note that y cannot be augmented).\n* At inference time, y_sym is averaged for 8 patterns of rot90 and flip TTAs. The conv5x5 is applied after the y_sym.\n\nThe augmentation is,\n\n```\nimport albumentations as A\n\nA.Compose([\n    A.RandomRotate90(p=1),\n    A.HorizontalFlip(p=0.5),\n    A.ShiftScaleRotate(rotate_limit=30, scale_limit=0.2)\n])\n```\n\nI refer to my public notebook for the details:\n\nhttps://www.kaggle.com/junkoda/base-unet-model-for-the-1st-place\n\n\n## Models\n\nThe final prediction is the weighted mean of two models with a threshold ~0.45. I tuned the threshold and the weights using the validation set. Both models are U-Net using maxvit_tiny_tf_512.in1k as the encoder, but one uses single time t=4, and the other use four times t = 1 - 4. The input image size is 1024×1024 for both models.\n\n## Single-time model\n\n* Upscale the input image to 1024×1024\n* Drop the final decoder layer of upscale and output 512×512\n\n## 4-panel model\n\nIn order to use the time information, I tried 3D ResNet and ConvLSTM, but I was not able to make them work at all. The only thing I could do was to pack four 512×512 image at t=1,2,3,4 into one 1024×1024 image and hope that self attention in MaxViT look at different times.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F954117%2F291ed2fb89c1e4b492f0d5890ad06868%2Fimg4.png?generation=1691671228909061&alt=media)\n\nIn the U-Net, I only pass the quarter of the features (H/2, W/2) to the decoder, which corresponds to the t=4 quarter. The rest is the same as the single-time model. Concatenating images in the spacial direction appear in Kaggle once in a while. I remember [CPMP's solution](https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275433) for the G2Net blackhole merger competition, in which he stacked 3 detector data horizontally rather than stacking them as 3 channels. \n\nTraining details:\n\n* Train 5 out of k=10 folds, using the out-of-fold to monitor the validation loss\n* 0.5 epochs of linear warmup to learning rate 8e-4, followed by cosine annealing\n* batch size 4 × gradient accumulation 2\n* AdamW with weight_decay = 0.01\n* 35 or 40 epochs of training\n\n```text\nScores\n                    CV     Public  Private\n       input size\n\nSingle time   512  0.697   0.707   0.712\nSingle time  1024  0.703   0.719   0.716\n4-panel      1024  0.704   0.719   0.722\nEnsemble           0.706   0.725   0.724\n```\n\nAll scores are for 5-fold × 8 TTA mean.\n\n## Things I was not able to do\n\n* I did not use pseudo labeling. I tried pseudo labeling at an early stage. It boosted the single model score a lot, but the benefit was unimpressive after 5-fold mean.\n* I was not able to train large models. I tried MaxViT small and base, and many other models, but the improvement was unclear. I increased the input image size instead. This failure must be due to my insufficient experience. \n* I used only positive data (dropping data with no positive labels) for model evaluation because it takes 1/2 time to train. However, I was not able to use it in the real prediction; it performs very bad for negative samples, and my classifier was not good enough to exclude such false negatives.\n* Removing small masks is a basic technique in segmentation competition, but here, removing even 1-pixel clusters (connected components) was not a good idea. I tried to remove false positive clusters based on some cluster statistics (size, density, max density, etc), but U-Net was cleverer than my post processing.\n\n\n## Codes\n\n* Training: https://github.com/junkoda/kaggle_contrails_solution\n* Inference notebook: https://www.kaggle.com/code/junkoda/contrails-submit\n\nUpdate 2023-08-20: Add links to the codes and minor english corrections.\n",
      "votes": 164
    },
    {
      "id": 2384094,
      "postDate": "2023-08-10T19:39:07.697Z",
      "content": "<p>Congratz on the win ! </p>\n<p>Very elegant solution, and impressive detective work. </p>",
      "rawMarkdown": "Congratz on the win ! \n\nVery elegant solution, and impressive detective work. ",
      "votes": 7
    },
    {
      "id": 2389738,
      "postDate": "2023-08-14T08:01:39.747Z",
      "content": "<p>Would it make sense to shift an image by 0.5 pixel instead of the mask? This way you don't need to shift masks during the inference.</p>",
      "rawMarkdown": "Would it make sense to shift an image by 0.5 pixel instead of the mask? This way you don't need to shift masks during the inference.",
      "votes": 5,
      "replies": [
        {
          "id": 2389827,
          "postDate": "2023-08-14T08:52:31.073Z",
          "content": "<p>I think so, and that seems easier. Is that what TASCJ did, because he did not write about shifting back 0.5 pixel in method 2? Somehow I didn't think about that during the competition.</p>",
          "rawMarkdown": "I think so, and that seems easier. Is that what TASCJ did, because he did not write about shifting back 0.5 pixel in method 2? Somehow I didn't think about that during the competition.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2387328,
      "postDate": "2023-08-12T15:28:55.143Z",
      "content": "<p>\" 4-panel model … hope that self attention in MaxViT look at different times.\"<br>\nthis actually might work. good idea!</p>\n<p>some smart design of positional encoding (e.g. t,x,y)  or/and attention may make it work better.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fa7af14e17b7344e2ae84349c25fb20eb%2FSelection_999(2895).png?generation=1691854097296090&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "\" 4-panel model ... hope that self attention in MaxViT look at different times.\"\nthis actually might work. good idea!\n\nsome smart design of positional encoding (e.g. t,x,y)  or/and attention may make it work better.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fa7af14e17b7344e2ae84349c25fb20eb%2FSelection_999(2895).png?generation=1691854097296090&alt=media)\n",
      "votes": 3,
      "replies": [
        {
          "id": 2387588,
          "postDate": "2023-08-12T19:42:56.413Z",
          "content": "<p>Thanks! 512 x 4 panel was better than 512 single time. I wonder if was competitive compared to other teams 2.5D and 3D model.</p>",
          "rawMarkdown": "Thanks! 512 x 4 panel was better than 512 single time. I wonder if was competitive compared to other teams 2.5D and 3D model."
        }
      ]
    },
    {
      "id": 2385090,
      "postDate": "2023-08-11T07:10:04.180Z",
      "content": "<p>Congrats for the winning! Thank you for the explanation, definitely simple solutions works best most of the times!</p>",
      "rawMarkdown": "Congrats for the winning! Thank you for the explanation, definitely simple solutions works best most of the times!",
      "votes": 3
    },
    {
      "id": 2383687,
      "postDate": "2023-08-10T14:13:33.797Z",
      "content": "<p>Congrats on the solo win <a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a>, and interesting write-up.</p>\n<p>Can you explain how you shift a mask by 1/2 a pixel?</p>",
      "rawMarkdown": "Congrats on the solo win @junkoda, and interesting write-up.\n\nCan you explain how you shift a mask by 1/2 a pixel?",
      "votes": 3,
      "replies": [
        {
          "id": 2383716,
          "postDate": "2023-08-10T14:25:46.167Z",
          "content": "<p>Thanks! I apply torch.nn.functional.grid_sample function to y. It returns interpolated values on the given grid. I hope it will be clear when I finish with my public notebook.</p>",
          "rawMarkdown": "Thanks! I apply torch.nn.functional.grid_sample function to y. It returns interpolated values on the given grid. I hope it will be clear when I finish with my public notebook.",
          "votes": 4
        }
      ]
    },
    {
      "id": 2396992,
      "postDate": "2023-08-18T16:31:32.833Z",
      "content": "<p><a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a> kudos for winning the competition and writing such a nice writeup!</p>",
      "rawMarkdown": "@junkoda kudos for winning the competition and writing such a nice writeup!",
      "votes": 1
    },
    {
      "id": 2394055,
      "postDate": "2023-08-16T17:25:15.767Z",
      "content": "<p>Thanks for the great writeup. OOC, how did you find out about the 0.5 pixel trick? By plotting the data maybe?</p>",
      "rawMarkdown": "Thanks for the great writeup. OOC, how did you find out about the 0.5 pixel trick? By plotting the data maybe?",
      "votes": 1,
      "replies": [
        {
          "id": 2394556,
          "postDate": "2023-08-17T02:35:28.427Z",
          "content": "<p>Thanks! I don't remember exactly now, but I think I came up with an hypothesis first that the reason for augmentation not working is the shifted labels. Thought about 1 pixel shift and then discover the 0.5 pixel with the visualization presented above. I thought there must be a special reason that rot90 and flip augmentation do not work.</p>",
          "rawMarkdown": "Thanks! I don't remember exactly now, but I think I came up with an hypothesis first that the reason for augmentation not working is the shifted labels. Thought about 1 pixel shift and then discover the 0.5 pixel with the visualization presented above. I thought there must be a special reason that rot90 and flip augmentation do not work.",
          "votes": 3,
          "replies": [
            {
              "id": 2397034,
              "postDate": "2023-08-18T17:00:36.727Z",
              "content": "<p>Thanks for the details!</p>",
              "rawMarkdown": "Thanks for the details!"
            },
            {
              "id": 2397952,
              "postDate": "2023-08-19T10:46:14.643Z",
              "content": "<p>I rationalized it away by imaging that the model gains insights about geographical location from the direction of contrails and that information being lost when applying those augmentations. I will try to always check what happens in visualizations like you did.</p>",
              "rawMarkdown": "I rationalized it away by imaging that the model gains insights about geographical location from the direction of contrails and that information being lost when applying those augmentations. I will try to always check what happens in visualizations like you did.",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2391263,
      "postDate": "2023-08-15T04:54:26.543Z",
      "content": "<p>Great job! :)</p>",
      "rawMarkdown": "Great job! :)",
      "votes": 1
    },
    {
      "id": 2390535,
      "postDate": "2023-08-14T15:59:13.970Z",
      "content": "<p>Congratulations 🎉 Dear <a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a> for this position.</p>",
      "rawMarkdown": "Congratulations 🎉 Dear @junkoda for this position.",
      "votes": 1
    },
    {
      "id": 2387813,
      "postDate": "2023-08-13T02:24:41.400Z",
      "content": "<p>Congrats and a great work done.</p>",
      "rawMarkdown": "Congrats and a great work done.",
      "votes": 1
    },
    {
      "id": 2387182,
      "postDate": "2023-08-12T13:16:48.580Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a> and thanks for sharing your approach.</p>",
      "rawMarkdown": "Congrats @junkoda and thanks for sharing your approach.",
      "votes": 1
    },
    {
      "id": 2385094,
      "postDate": "2023-08-11T07:12:50.657Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a> </p>\n<p>Very helpful solution.</p>",
      "rawMarkdown": "Congrats @junkoda \n\nVery helpful solution.",
      "votes": 1
    },
    {
      "id": 2384802,
      "postDate": "2023-08-11T04:47:48.710Z",
      "content": "<p>awesome … congratulations !</p>",
      "rawMarkdown": "awesome ... congratulations !",
      "votes": 1
    },
    {
      "id": 2384745,
      "postDate": "2023-08-11T04:17:26.343Z",
      "content": "<p>This is wonderful. Thank you so much for sharing!</p>",
      "rawMarkdown": "This is wonderful. Thank you so much for sharing!",
      "votes": 1
    },
    {
      "id": 2384532,
      "postDate": "2023-08-11T01:59:00.340Z",
      "content": "<p>May I ask what machine you are using for the experiment? These experiments seem to require very good machines</p>",
      "rawMarkdown": "May I ask what machine you are using for the experiment? These experiments seem to require very good machines",
      "votes": 1,
      "replies": [
        {
          "id": 2384535,
          "postDate": "2023-08-11T01:59:51.040Z",
          "content": "<p>Thank you very much if you can recover</p>",
          "rawMarkdown": "Thank you very much if you can recover"
        },
        {
          "id": 2386348,
          "postDate": "2023-08-11T21:40:18.683Z",
          "content": "<p>Yes, I use multiple A100 GPUs in the last week to train the final models with 1024-size image.<br>\nUntil then, I use my local machine with one RTX3090 GPU and 1TB M2 nvme SSD (3500 MB/s bandwidth). SSD is also important to load large data quickly for many times (many epochs)</p>",
          "rawMarkdown": "Yes, I use multiple A100 GPUs in the last week to train the final models with 1024-size image.\nUntil then, I use my local machine with one RTX3090 GPU and 1TB M2 nvme SSD (3500 MB/s bandwidth). SSD is also important to load large data quickly for many times (many epochs)",
          "votes": 3,
          "replies": [
            {
              "id": 2387985,
              "postDate": "2023-08-13T05:18:43.347Z",
              "content": "<p>thank you for your answer, which has greatly inspired me</p>",
              "rawMarkdown": "thank you for your answer, which has greatly inspired me",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2384163,
      "postDate": "2023-08-10T21:15:39.237Z",
      "content": "<p>Nicely done! I will definitely make use of your technique to shift masked labels as part of the model building framework. Thanks for sharing!</p>",
      "rawMarkdown": "Nicely done! I will definitely make use of your technique to shift masked labels as part of the model building framework. Thanks for sharing!",
      "votes": 1
    },
    {
      "id": 2384000,
      "postDate": "2023-08-10T17:58:22.007Z",
      "content": "<p>Congratulations for the solo win! I would like to ask a few questions:</p>\n<ol>\n<li>You created 10 folds using train folder, train only 5 of them and validate those on given validation set, right?</li>\n<li>Why you passed only the quarter of the features to the decoder?</li>\n<li>Why you dropped the final decoder layer of upscale to output 512x512 in the case of single frame model?</li>\n</ol>\n<p>Thanks for sharing</p>",
      "rawMarkdown": "Congratulations for the solo win! I would like to ask a few questions:\n\n1. You created 10 folds using train folder, train only 5 of them and validate those on given validation set, right?\n2. Why you passed only the quarter of the features to the decoder?\n3. Why you dropped the final decoder layer of upscale to output 512x512 in the case of single frame model?\n\nThanks for sharing\n",
      "votes": 1,
      "replies": [
        {
          "id": 2384219,
          "postDate": "2023-08-10T23:09:28.153Z",
          "content": "<ol>\n<li><p>Yes. I thought 10% validation set is sufficiently large and additional 10% from 80% to 90% in the train might help compared to 5 fold. I didn't see a difference between 5 fold and 10 fold at least with small models though.</p></li>\n<li><p>The quarter is the upper-left quarter for t=4. That is what the U-Net decoder should get from skip connection to output segmentation mask at t=4.</p></li>\n<li><p>I assumed the upscale to 1024 is unnecessary and 512 is sufficient but I don't know. I didn't tried.</p></li>\n</ol>",
          "rawMarkdown": "1. Yes. I thought 10% validation set is sufficiently large and additional 10% from 80% to 90% in the train might help compared to 5 fold. I didn't see a difference between 5 fold and 10 fold at least with small models though.\n\n2. The quarter is the upper-left quarter for t=4. That is what the U-Net decoder should get from skip connection to output segmentation mask at t=4.\n\n3. I assumed the upscale to 1024 is unnecessary and 512 is sufficient but I don't know. I didn't tried.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2383999,
      "postDate": "2023-08-10T17:56:01.017Z",
      "content": "<p>Did you use different learning rates for encoder and decoder? It improved my score a lot</p>",
      "rawMarkdown": "Did you use different learning rates for encoder and decoder? It improved my score a lot",
      "votes": 1,
      "replies": [
        {
          "id": 2384222,
          "postDate": "2023-08-10T23:13:54.797Z",
          "content": "<p>I tried 10 times smaller learning rate for encoder and it improves the first 10 epochs nicely but underperformed at the end of 40 epochs. I also tried completely freezing the encoder, but that was not good at all. Training the encoder is necessary.</p>\n<p>There must be some learning rate for encoder better than the single learning rate, but I could not find it and spend time on other issues.</p>",
          "rawMarkdown": "I tried 10 times smaller learning rate for encoder and it improves the first 10 epochs nicely but underperformed at the end of 40 epochs. I also tried completely freezing the encoder, but that was not good at all. Training the encoder is necessary.\n\nThere must be some learning rate for encoder better than the single learning rate, but I could not find it and spend time on other issues.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2383644,
      "postDate": "2023-08-10T13:45:35.347Z",
      "content": "<p>Thanks for sharing the write up <a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a> ! Kudos on winning the competition.🎉</p>",
      "rawMarkdown": "Thanks for sharing the write up @junkoda ! Kudos on winning the competition.🎉",
      "votes": 1
    },
    {
      "id": 2407088,
      "postDate": "2023-08-24T19:50:26.487Z",
      "content": "<p>I was just talking with an intern about the need to get into the messiness of the data. It seems that your approach is successful because of how much you dove into the problem and asked about the structure of the data.</p>",
      "rawMarkdown": "I was just talking with an intern about the need to get into the messiness of the data. It seems that your approach is successful because of how much you dove into the problem and asked about the structure of the data.",
      "votes": 2,
      "replies": [
        {
          "id": 2407189,
          "postDate": "2023-08-24T23:32:11.233Z",
          "content": "<p>Thanks! Sounds like a very nice intern program. Understanding the data is important, but it is also difficult for deeplearning problems; often we only see how annotators are inconsistent. In this problem, it probably worked because I also had a hypothesis and visualize the data with a specific purpose. </p>",
          "rawMarkdown": "Thanks! Sounds like a very nice intern program. Understanding the data is important, but it is also difficult for deeplearning problems; often we only see how annotators are inconsistent. In this problem, it probably worked because I also had a hypothesis and visualize the data with a specific purpose. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 2389724,
      "postDate": "2023-08-14T07:53:27.320Z",
      "content": "<p>Congratz on the win!<br>\nI also just found out that the y_sym of your solution is so helpful! Thank you for sharing!</p>",
      "rawMarkdown": "Congratz on the win!\nI also just found out that the y_sym of your solution is so helpful! Thank you for sharing!",
      "votes": 2,
      "replies": [
        {
          "id": 2389829,
          "postDate": "2023-08-14T08:52:56.947Z",
          "content": "<p>You are welcome!</p>",
          "rawMarkdown": "You are welcome!"
        }
      ]
    },
    {
      "id": 2384023,
      "postDate": "2023-08-10T18:27:22.410Z",
      "content": "<p>Nice shifted mask visualization, congratz with the result! </p>",
      "rawMarkdown": "Nice shifted mask visualization, congratz with the result! ",
      "votes": 2
    },
    {
      "id": 2383736,
      "postDate": "2023-08-10T14:34:24.843Z",
      "content": "<p>Many congratulations on the solo win and thank you for the great write up!</p>\n<p>Could you please clarify this part:</p>\n<blockquote>\n  <p>Additional tiny convolution of 5x5 with stride 2 is trained to learn the mapping from y_sym to right-bottom aligned y. This is done when the augmentation is not applied for the 5% of the data (note that y cannot be augmented).</p>\n</blockquote>\n<p>Where do you put this convolution ? head of the Unet ? When do you train it? What are the 5% of the data you are referring to? Why not just training with centered masks (y_sym) and then apply the invert shift on the predictions ?</p>\n<blockquote>\n  <p>At inference time, y_sym is averaged for 8 patterns of rot90 and flip TTAs. The conv5x5 is applied to the averaged y_sym.</p>\n</blockquote>\n<p>Just to be sure I understood : basically you perform classical 8 times TTA, average the predictions and then shift the results using your trained conv5x5 ?</p>",
      "rawMarkdown": "Many congratulations on the solo win and thank you for the great write up!\n\nCould you please clarify this part:\n>Additional tiny convolution of 5x5 with stride 2 is trained to learn the mapping from y_sym to right-bottom aligned y. This is done when the augmentation is not applied for the 5% of the data (note that y cannot be augmented).\n\nWhere do you put this convolution ? head of the Unet ? When do you train it? What are the 5% of the data you are referring to? Why not just training with centered masks (y_sym) and then apply the invert shift on the predictions ?\n\n>At inference time, y_sym is averaged for 8 patterns of rot90 and flip TTAs. The conv5x5 is applied to the averaged y_sym.\n\nJust to be sure I understood : basically you perform classical 8 times TTA, average the predictions and then shift the results using your trained conv5x5 ?\n",
      "votes": 2,
      "replies": [
        {
          "id": 2383827,
          "postDate": "2023-08-10T15:19:34.347Z",
          "content": "<p>Thanks for the question, and I must agree that many reader have the same question. That's why I am going to write the notebook, and I hope it become clear with the code. I'll read my text again tomorrow and think how I can revise, too.</p>\n<ul>\n<li>where do you put the convolution?<br>\nYes it is a head,</li>\n</ul>\n<pre><code>y_sym_pred = unet(x)\ny_pred = conv5x5(y_sym_pred)\n y_sym_pred, y_pred\n</code></pre>\n<p>but both y_sym_pred and y_pred are returned and both are compared to their labels. Maybe something strange or wrong is going on here for deeplearning experts.</p>\n<ul>\n<li>when do you train?</li>\n</ul>\n<p>I always train y_sym_pred with y_sym, but I cannot train y_pred if augmentation is applied.</p>\n<p>loss = bce_loss(y_sym_pred, y_sym)</p>\n<ul>\n<li>what is 5%?</li>\n</ul>\n<p>I apply augmentation with probability 0.95, and do not with 0.05. 5% is the probability that the augmentation is not applied. When augmention is not applied, I train both,</p>\n<p>loss = bce_loss(y_sym_pred, y_sym) + bce_loss(y_pred, y)</p>\n<p>(I am not accurate about the batch, I hope that is a technical detail)</p>\n<ul>\n<li>Why not just training with centered masks (y_sym) and then apply the invert shift on the predictions ?</li>\n</ul>\n<p>Yes, you can do that with 512x512 y_sym, but I was also working with 256x256 y_sym, and I used a similar method.</p>\n<ul>\n<li>Just to be sure I understood<br>\nYes that's right. conv5x5 is outside the standard TTA.</li>\n</ul>",
          "rawMarkdown": "Thanks for the question, and I must agree that many reader have the same question. That's why I am going to write the notebook, and I hope it become clear with the code. I'll read my text again tomorrow and think how I can revise, too.\n\n- where do you put the convolution?\nYes it is a head,\n\n```\ny_sym_pred = unet(x)\ny_pred = conv5x5(y_sym_pred)\nreturn y_sym_pred, y_pred\n```\n\nbut both y_sym_pred and y_pred are returned and both are compared to their labels. Maybe something strange or wrong is going on here for deeplearning experts.\n\n- when do you train?\n\nI always train y_sym_pred with y_sym, but I cannot train y_pred if augmentation is applied.\n\nloss = bce_loss(y_sym_pred, y_sym)\n\n- what is 5%?\n\nI apply augmentation with probability 0.95, and do not with 0.05. 5% is the probability that the augmentation is not applied. When augmention is not applied, I train both,\n\nloss = bce_loss(y_sym_pred, y_sym) + bce_loss(y_pred, y)\n\n(I am not accurate about the batch, I hope that is a technical detail)\n\n- Why not just training with centered masks (y_sym) and then apply the invert shift on the predictions ?\n\nYes, you can do that with 512x512 y_sym, but I was also working with 256x256 y_sym, and I used a similar method.\n\n- Just to be sure I understood\nYes that's right. conv5x5 is outside the standard TTA.\n\n",
          "votes": 4
        }
      ]
    },
    {
      "id": 2383712,
      "postDate": "2023-08-10T14:23:44.967Z",
      "content": "<p>Super cool explanation.<br>\n\"I first shift the label by 0.5 pixel and create y_sym\"<br>\n\"Additional tiny convolution of 5x5 with stride 2 is trained to learn the mapping from y_sym to right-bottom aligned y.\"</p>\n<p>Why learning a separate convolution to shift the mask back and not just shift it 1 pixel at 512x512 and downscale to 256x256, which is equivalent to 0.5 px shift)?</p>",
      "rawMarkdown": "Super cool explanation.\n\"I first shift the label by 0.5 pixel and create y_sym\"\n\"Additional tiny convolution of 5x5 with stride 2 is trained to learn the mapping from y_sym to right-bottom aligned y.\"\n\nWhy learning a separate convolution to shift the mask back and not just shift it 1 pixel at 512x512 and downscale to 256x256, which is equivalent to 0.5 px shift)?",
      "votes": 2,
      "replies": [
        {
          "id": 2383728,
          "postDate": "2023-08-10T14:32:24.273Z",
          "content": "<p>Yes! I tried that too and both works equally well. I started from 256x256 y_sym and I used 3x3 at that time, and I continue to do that way.</p>",
          "rawMarkdown": "Yes! I tried that too and both works equally well. I started from 256x256 y_sym and I used 3x3 at that time, and I continue to do that way.",
          "votes": 3
        }
      ]
    },
    {
      "id": 2418095,
      "postDate": "2023-09-01T06:11:16.697Z",
      "content": "<p>Hello, congratulations on achieving the first-place result.</p>\n<p>I have reviewed your GitHub code and noticed that in \"unet1024/data.py,\" on line 87, the code \"x=f['x'][3,:]\" means that it is selecting data from the fourth time step, and on line 109, you are resizing the image to 1024x1024 with \"x=self.resize(x).\" This approach differs from what you mentioned in the discussion, where you described combining time steps 1-4 to create a 1024x1024 image (\"The only thing I could do was to pack four 512×512 images at t=1,2,3,4 into one 1024×1024 image\").</p>\n<p>Could you please clarify this inconsistency and explain how these two approaches might impact the final results differently?</p>",
      "rawMarkdown": "Hello, congratulations on achieving the first-place result.\n\nI have reviewed your GitHub code and noticed that in \"unet1024/data.py,\" on line 87, the code \"x=f['x'][3,:]\" means that it is selecting data from the fourth time step, and on line 109, you are resizing the image to 1024x1024 with \"x=self.resize(x).\" This approach differs from what you mentioned in the discussion, where you described combining time steps 1-4 to create a 1024x1024 image (\"The only thing I could do was to pack four 512×512 images at t=1,2,3,4 into one 1024×1024 image\").\n\nCould you please clarify this inconsistency and explain how these two approaches might impact the final results differently?",
      "replies": [
        {
          "id": 2418202,
          "postDate": "2023-09-01T07:36:02.527Z",
          "content": "<p>Hi, Thanks for checking it out. unet1024 is the single time 1024x1024 model. The 4-panel model is in vit4 directory.</p>",
          "rawMarkdown": "Hi, Thanks for checking it out. unet1024 is the single time 1024x1024 model. The 4-panel model is in vit4 directory."
        }
      ]
    },
    {
      "id": 2412261,
      "postDate": "2023-08-28T07:13:29.580Z",
      "content": "<p>May I ask how you obtained the 4-panel model? Did you use some automl tools? Thank you.</p>",
      "rawMarkdown": "May I ask how you obtained the 4-panel model? Did you use some automl tools? Thank you.",
      "replies": [
        {
          "id": 2412463,
          "postDate": "2023-08-28T09:35:54.353Z",
          "content": "<p>I modified the segmentation-models-pytorch U-Net code.<br>\nThis change is very simple but it would be interesting if ChatGPT can write models. Currently it is much easier for me to code by myself than thinking about the prompt to generate the correct code.</p>",
          "rawMarkdown": "I modified the segmentation-models-pytorch U-Net code.\nThis change is very simple but it would be interesting if ChatGPT can write models. Currently it is much easier for me to code by myself than thinking about the prompt to generate the correct code.",
          "replies": [
            {
              "id": 2413529,
              "postDate": "2023-08-29T01:31:22.783Z",
              "content": "<p>That's amazing, it must be very fulfilling. How did you know how to modify it? By feeling it? Also, I feel confused about myself. Should I use some automl tools? Do you use such tools? And what is the effect?</p>",
              "rawMarkdown": "That's amazing, it must be very fulfilling. How did you know how to modify it? By feeling it? Also, I feel confused about myself. Should I use some automl tools? Do you use such tools? And what is the effect?"
            },
            {
              "id": 2414170,
              "postDate": "2023-08-29T12:44:09.997Z",
              "content": "<p>I'm afraid I'm not the right person to ask about automl tools. It must be very efficient in real applications, but I've never used. I think the effect is, if they are competitive (and probably they are in several area) such competition would not appear in Kaggle because there is little room to make difference.</p>\n<p>To learn how to modify, I guess we have to read codes, rewrite and run them. Many simple codes are available in the internet and Kaggle is also encouraging to publish top solutions. If you like making something and see them work, you can make small models. I've created small CNN models and a mini ResNet model seeing public codes.</p>\n<p>But, writing your special model is not always important; often that's too ambitious and unsuccessful. The \"2.5D model\" adopted by many to teams is rather an exception. There are often good public Kaggle Codes/notebooks, and you can start from small improvements.</p>",
              "rawMarkdown": "I'm afraid I'm not the right person to ask about automl tools. It must be very efficient in real applications, but I've never used. I think the effect is, if they are competitive (and probably they are in several area) such competition would not appear in Kaggle because there is little room to make difference.\n\nTo learn how to modify, I guess we have to read codes, rewrite and run them. Many simple codes are available in the internet and Kaggle is also encouraging to publish top solutions. If you like making something and see them work, you can make small models. I've created small CNN models and a mini ResNet model seeing public codes.\n\nBut, writing your special model is not always important; often that's too ambitious and unsuccessful. The \"2.5D model\" adopted by many to teams is rather an exception. There are often good public Kaggle Codes/notebooks, and you can start from small improvements."
            },
            {
              "id": 2414294,
              "postDate": "2023-08-29T14:33:39.713Z",
              "content": "<p>Your model was created by yourself, which is very amazing. As a beginner, I really want to know how to achieve this.</p>",
              "rawMarkdown": "Your model was created by yourself, which is very amazing. As a beginner, I really want to know how to achieve this."
            }
          ]
        }
      ]
    },
    {
      "id": 2407168,
      "postDate": "2023-08-24T22:15:22.527Z",
      "content": "<p>congrats <a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a> and thanks for sharing your method</p>",
      "rawMarkdown": "congrats @junkoda and thanks for sharing your method"
    },
    {
      "id": 2396818,
      "postDate": "2023-08-18T14:03:58.440Z",
      "content": "<p>Congrats! Great writeup, thank you for your insights. </p>",
      "rawMarkdown": "Congrats! Great writeup, thank you for your insights. "
    },
    {
      "id": 2758454,
      "postDate": "2024-04-18T07:03:04.220Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2468601,
      "postDate": "2023-10-05T16:36:19.910Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2412257,
      "postDate": "2023-08-28T07:11:41.900Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2394532,
      "postDate": "2023-08-17T02:01:02.297Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 2393797,
      "postDate": "2023-08-16T14:29:28.453Z",
      "content": "<p>Congratulations and thanks for sharing your approach. I really didn't catch the 0.5-pixel shift. Great work👍</p>",
      "rawMarkdown": "Congratulations and thanks for sharing your approach. I really didn't catch the 0.5-pixel shift. Great work👍",
      "votes": 2,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2384094,
      "author_name": "Theo Viel",
      "author_url": "",
      "post_date": "2023-08-10T19:39:07.697000",
      "content": "<p>Congratz on the win ! </p>\n<p>Very elegant solution, and impressive detective work. </p>",
      "votes": 7,
      "replies": []
    },
    {
      "id": 2389738,
      "author_name": "DennisSakva",
      "author_url": "",
      "post_date": "2023-08-14T08:01:39.747000",
      "content": "<p>Would it make sense to shift an image by 0.5 pixel instead of the mask? This way you don't need to shift masks during the inference.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 2389827,
          "author_name": "🐢 Jun Koda",
          "author_url": "",
          "post_date": "2023-08-14T08:52:31.073000",
          "content": "<p>I think so, and that seems easier. Is that what TASCJ did, because he did not write about shifting back 0.5 pixel in method 2? Somehow I didn't think about that during the competition.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2387328,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-08-12T15:28:55.143000",
      "content": "<p>\" 4-panel model … hope that self attention in MaxViT look at different times.\"<br>\nthis actually might work. good idea!</p>\n<p>some smart design of positional encoding (e.g. t,x,y)  or/and attention may make it work better.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fa7af14e17b7344e2ae84349c25fb20eb%2FSelection_999(2895).png?generation=1691854097296090&amp;alt=media\" alt=\"\"></p>",
      "votes": 3,
      "replies": [
        {
          "id": 2387588,
          "author_name": "🐢 Jun Koda",
          "author_url": "",
          "post_date": "2023-08-12T19:42:56.413000",
          "content": "<p>Thanks! 512 x 4 panel was better than 512 single time. I wonder if was competitive compared to other teams 2.5D and 3D model.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2385090,
      "author_name": "Giovanni Cavallin",
      "author_url": "",
      "post_date": "2023-08-11T07:10:04.180000",
      "content": "<p>Congrats for the winning! Thank you for the explanation, definitely simple solutions works best most of the times!</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2383687,
      "author_name": "Bartley",
      "author_url": "",
      "post_date": "2023-08-10T14:13:33.797000",
      "content": "<p>Congrats on the solo win <a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a>, and interesting write-up.</p>\n<p>Can you explain how you shift a mask by 1/2 a pixel?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2383716,
          "author_name": "🐢 Jun Koda",
          "author_url": "",
          "post_date": "2023-08-10T14:25:46.167000",
          "content": "<p>Thanks! I apply torch.nn.functional.grid_sample function to y. It returns interpolated values on the given grid. I hope it will be clear when I finish with my public notebook.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 2396992,
      "author_name": "Sambit Barik",
      "author_url": "",
      "post_date": "2023-08-18T16:31:32.833000",
      "content": "<p><a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a> kudos for winning the competition and writing such a nice writeup!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2394055,
      "author_name": "Yassine Alouini",
      "author_url": "",
      "post_date": "2023-08-16T17:25:15.767000",
      "content": "<p>Thanks for the great writeup. OOC, how did you find out about the 0.5 pixel trick? By plotting the data maybe?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2394556,
          "author_name": "🐢 Jun Koda",
          "author_url": "",
          "post_date": "2023-08-17T02:35:28.427000",
          "content": "<p>Thanks! I don't remember exactly now, but I think I came up with an hypothesis first that the reason for augmentation not working is the shifted labels. Thought about 1 pixel shift and then discover the 0.5 pixel with the visualization presented above. I thought there must be a special reason that rot90 and flip augmentation do not work.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2397034,
              "author_name": "Yassine Alouini",
              "author_url": "",
              "post_date": "2023-08-18T17:00:36.727000",
              "content": "<p>Thanks for the details!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2397952,
              "author_name": "Raki",
              "author_url": "",
              "post_date": "2023-08-19T10:46:14.643000",
              "content": "<p>I rationalized it away by imaging that the model gains insights about geographical location from the direction of contrails and that information being lost when applying those augmentations. I will try to always check what happens in visualizations like you did.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2391263,
      "author_name": "Twit ter1",
      "author_url": "",
      "post_date": "2023-08-15T04:54:26.543000",
      "content": "<p>Great job! :)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2390535,
      "author_name": "Tariq Mahmood",
      "author_url": "",
      "post_date": "2023-08-14T15:59:13.970000",
      "content": "<p>Congratulations 🎉 Dear <a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a> for this position.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2387813,
      "author_name": "Vyshakh Krishnan T",
      "author_url": "",
      "post_date": "2023-08-13T02:24:41.400000",
      "content": "<p>Congrats and a great work done.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2387182,
      "author_name": "Shailja Kant Tiwari",
      "author_url": "",
      "post_date": "2023-08-12T13:16:48.580000",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a> and thanks for sharing your approach.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2385094,
      "author_name": "Muhammad Usman",
      "author_url": "",
      "post_date": "2023-08-11T07:12:50.657000",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a> </p>\n<p>Very helpful solution.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2384802,
      "author_name": "SAM NY",
      "author_url": "",
      "post_date": "2023-08-11T04:47:48.710000",
      "content": "<p>awesome … congratulations !</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2384745,
      "author_name": "NA",
      "author_url": "",
      "post_date": "2023-08-11T04:17:26.343000",
      "content": "<p>This is wonderful. Thank you so much for sharing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2384532,
      "author_name": "kongweihao",
      "author_url": "",
      "post_date": "2023-08-11T01:59:00.340000",
      "content": "<p>May I ask what machine you are using for the experiment? These experiments seem to require very good machines</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2384535,
          "author_name": "kongweihao",
          "author_url": "",
          "post_date": "2023-08-11T01:59:51.040000",
          "content": "<p>Thank you very much if you can recover</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2386348,
          "author_name": "🐢 Jun Koda",
          "author_url": "",
          "post_date": "2023-08-11T21:40:18.683000",
          "content": "<p>Yes, I use multiple A100 GPUs in the last week to train the final models with 1024-size image.<br>\nUntil then, I use my local machine with one RTX3090 GPU and 1TB M2 nvme SSD (3500 MB/s bandwidth). SSD is also important to load large data quickly for many times (many epochs)</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2387985,
              "author_name": "kongweihao",
              "author_url": "",
              "post_date": "2023-08-13T05:18:43.347000",
              "content": "<p>thank you for your answer, which has greatly inspired me</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2384163,
      "author_name": "Darshan Patel",
      "author_url": "",
      "post_date": "2023-08-10T21:15:39.237000",
      "content": "<p>Nicely done! I will definitely make use of your technique to shift masked labels as part of the model building framework. Thanks for sharing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2384000,
      "author_name": "delai50",
      "author_url": "",
      "post_date": "2023-08-10T17:58:22.007000",
      "content": "<p>Congratulations for the solo win! I would like to ask a few questions:</p>\n<ol>\n<li>You created 10 folds using train folder, train only 5 of them and validate those on given validation set, right?</li>\n<li>Why you passed only the quarter of the features to the decoder?</li>\n<li>Why you dropped the final decoder layer of upscale to output 512x512 in the case of single frame model?</li>\n</ol>\n<p>Thanks for sharing</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2384219,
          "author_name": "🐢 Jun Koda",
          "author_url": "",
          "post_date": "2023-08-10T23:09:28.153000",
          "content": "<ol>\n<li><p>Yes. I thought 10% validation set is sufficiently large and additional 10% from 80% to 90% in the train might help compared to 5 fold. I didn't see a difference between 5 fold and 10 fold at least with small models though.</p></li>\n<li><p>The quarter is the upper-left quarter for t=4. That is what the U-Net decoder should get from skip connection to output segmentation mask at t=4.</p></li>\n<li><p>I assumed the upscale to 1024 is unnecessary and 512 is sufficient but I don't know. I didn't tried.</p></li>\n</ol>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2383999,
      "author_name": "Patchef",
      "author_url": "",
      "post_date": "2023-08-10T17:56:01.017000",
      "content": "<p>Did you use different learning rates for encoder and decoder? It improved my score a lot</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2384222,
          "author_name": "🐢 Jun Koda",
          "author_url": "",
          "post_date": "2023-08-10T23:13:54.797000",
          "content": "<p>I tried 10 times smaller learning rate for encoder and it improves the first 10 epochs nicely but underperformed at the end of 40 epochs. I also tried completely freezing the encoder, but that was not good at all. Training the encoder is necessary.</p>\n<p>There must be some learning rate for encoder better than the single learning rate, but I could not find it and spend time on other issues.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2383644,
      "author_name": "Suraj",
      "author_url": "",
      "post_date": "2023-08-10T13:45:35.347000",
      "content": "<p>Thanks for sharing the write up <a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a> ! Kudos on winning the competition.🎉</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2407088,
      "author_name": "JeremyLoscheider",
      "author_url": "",
      "post_date": "2023-08-24T19:50:26.487000",
      "content": "<p>I was just talking with an intern about the need to get into the messiness of the data. It seems that your approach is successful because of how much you dove into the problem and asked about the structure of the data.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2407189,
          "author_name": "🐢 Jun Koda",
          "author_url": "",
          "post_date": "2023-08-24T23:32:11.233000",
          "content": "<p>Thanks! Sounds like a very nice intern program. Understanding the data is important, but it is also difficult for deeplearning problems; often we only see how annotators are inconsistent. In this problem, it probably worked because I also had a hypothesis and visualize the data with a specific purpose. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2389724,
      "author_name": "Dimas Maulana Ichsan",
      "author_url": "",
      "post_date": "2023-08-14T07:53:27.320000",
      "content": "<p>Congratz on the win!<br>\nI also just found out that the y_sym of your solution is so helpful! Thank you for sharing!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2389829,
          "author_name": "🐢 Jun Koda",
          "author_url": "",
          "post_date": "2023-08-14T08:52:56.947000",
          "content": "<p>You are welcome!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2384023,
      "author_name": "slime",
      "author_url": "",
      "post_date": "2023-08-10T18:27:22.410000",
      "content": "<p>Nice shifted mask visualization, congratz with the result! </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2383736,
      "author_name": "Optimo",
      "author_url": "",
      "post_date": "2023-08-10T14:34:24.843000",
      "content": "<p>Many congratulations on the solo win and thank you for the great write up!</p>\n<p>Could you please clarify this part:</p>\n<blockquote>\n  <p>Additional tiny convolution of 5x5 with stride 2 is trained to learn the mapping from y_sym to right-bottom aligned y. This is done when the augmentation is not applied for the 5% of the data (note that y cannot be augmented).</p>\n</blockquote>\n<p>Where do you put this convolution ? head of the Unet ? When do you train it? What are the 5% of the data you are referring to? Why not just training with centered masks (y_sym) and then apply the invert shift on the predictions ?</p>\n<blockquote>\n  <p>At inference time, y_sym is averaged for 8 patterns of rot90 and flip TTAs. The conv5x5 is applied to the averaged y_sym.</p>\n</blockquote>\n<p>Just to be sure I understood : basically you perform classical 8 times TTA, average the predictions and then shift the results using your trained conv5x5 ?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2383827,
          "author_name": "🐢 Jun Koda",
          "author_url": "",
          "post_date": "2023-08-10T15:19:34.347000",
          "content": "<p>Thanks for the question, and I must agree that many reader have the same question. That's why I am going to write the notebook, and I hope it become clear with the code. I'll read my text again tomorrow and think how I can revise, too.</p>\n<ul>\n<li>where do you put the convolution?<br>\nYes it is a head,</li>\n</ul>\n<pre><code>y_sym_pred = unet(x)\ny_pred = conv5x5(y_sym_pred)\n y_sym_pred, y_pred\n</code></pre>\n<p>but both y_sym_pred and y_pred are returned and both are compared to their labels. Maybe something strange or wrong is going on here for deeplearning experts.</p>\n<ul>\n<li>when do you train?</li>\n</ul>\n<p>I always train y_sym_pred with y_sym, but I cannot train y_pred if augmentation is applied.</p>\n<p>loss = bce_loss(y_sym_pred, y_sym)</p>\n<ul>\n<li>what is 5%?</li>\n</ul>\n<p>I apply augmentation with probability 0.95, and do not with 0.05. 5% is the probability that the augmentation is not applied. When augmention is not applied, I train both,</p>\n<p>loss = bce_loss(y_sym_pred, y_sym) + bce_loss(y_pred, y)</p>\n<p>(I am not accurate about the batch, I hope that is a technical detail)</p>\n<ul>\n<li>Why not just training with centered masks (y_sym) and then apply the invert shift on the predictions ?</li>\n</ul>\n<p>Yes, you can do that with 512x512 y_sym, but I was also working with 256x256 y_sym, and I used a similar method.</p>\n<ul>\n<li>Just to be sure I understood<br>\nYes that's right. conv5x5 is outside the standard TTA.</li>\n</ul>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 2383712,
      "author_name": "DennisSakva",
      "author_url": "",
      "post_date": "2023-08-10T14:23:44.967000",
      "content": "<p>Super cool explanation.<br>\n\"I first shift the label by 0.5 pixel and create y_sym\"<br>\n\"Additional tiny convolution of 5x5 with stride 2 is trained to learn the mapping from y_sym to right-bottom aligned y.\"</p>\n<p>Why learning a separate convolution to shift the mask back and not just shift it 1 pixel at 512x512 and downscale to 256x256, which is equivalent to 0.5 px shift)?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2383728,
          "author_name": "🐢 Jun Koda",
          "author_url": "",
          "post_date": "2023-08-10T14:32:24.273000",
          "content": "<p>Yes! I tried that too and both works equally well. I started from 256x256 y_sym and I used 3x3 at that time, and I continue to do that way.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2418095,
      "author_name": "Bananafin",
      "author_url": "",
      "post_date": "2023-09-01T06:11:16.697000",
      "content": "<p>Hello, congratulations on achieving the first-place result.</p>\n<p>I have reviewed your GitHub code and noticed that in \"unet1024/data.py,\" on line 87, the code \"x=f['x'][3,:]\" means that it is selecting data from the fourth time step, and on line 109, you are resizing the image to 1024x1024 with \"x=self.resize(x).\" This approach differs from what you mentioned in the discussion, where you described combining time steps 1-4 to create a 1024x1024 image (\"The only thing I could do was to pack four 512×512 images at t=1,2,3,4 into one 1024×1024 image\").</p>\n<p>Could you please clarify this inconsistency and explain how these two approaches might impact the final results differently?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2418202,
          "author_name": "🐢 Jun Koda",
          "author_url": "",
          "post_date": "2023-09-01T07:36:02.527000",
          "content": "<p>Hi, Thanks for checking it out. unet1024 is the single time 1024x1024 model. The 4-panel model is in vit4 directory.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2412261,
      "author_name": "kongweihao",
      "author_url": "",
      "post_date": "2023-08-28T07:13:29.580000",
      "content": "<p>May I ask how you obtained the 4-panel model? Did you use some automl tools? Thank you.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2412463,
          "author_name": "🐢 Jun Koda",
          "author_url": "",
          "post_date": "2023-08-28T09:35:54.353000",
          "content": "<p>I modified the segmentation-models-pytorch U-Net code.<br>\nThis change is very simple but it would be interesting if ChatGPT can write models. Currently it is much easier for me to code by myself than thinking about the prompt to generate the correct code.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2413529,
              "author_name": "kongweihao",
              "author_url": "",
              "post_date": "2023-08-29T01:31:22.783000",
              "content": "<p>That's amazing, it must be very fulfilling. How did you know how to modify it? By feeling it? Also, I feel confused about myself. Should I use some automl tools? Do you use such tools? And what is the effect?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2414170,
              "author_name": "🐢 Jun Koda",
              "author_url": "",
              "post_date": "2023-08-29T12:44:09.997000",
              "content": "<p>I'm afraid I'm not the right person to ask about automl tools. It must be very efficient in real applications, but I've never used. I think the effect is, if they are competitive (and probably they are in several area) such competition would not appear in Kaggle because there is little room to make difference.</p>\n<p>To learn how to modify, I guess we have to read codes, rewrite and run them. Many simple codes are available in the internet and Kaggle is also encouraging to publish top solutions. If you like making something and see them work, you can make small models. I've created small CNN models and a mini ResNet model seeing public codes.</p>\n<p>But, writing your special model is not always important; often that's too ambitious and unsuccessful. The \"2.5D model\" adopted by many to teams is rather an exception. There are often good public Kaggle Codes/notebooks, and you can start from small improvements.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2414294,
              "author_name": "kongweihao",
              "author_url": "",
              "post_date": "2023-08-29T14:33:39.713000",
              "content": "<p>Your model was created by yourself, which is very amazing. As a beginner, I really want to know how to achieve this.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2407168,
      "author_name": "NewToAI_i",
      "author_url": "",
      "post_date": "2023-08-24T22:15:22.527000",
      "content": "<p>congrats <a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a> and thanks for sharing your method</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2396818,
      "author_name": "hcdubsy",
      "author_url": "",
      "post_date": "2023-08-18T14:03:58.440000",
      "content": "<p>Congrats! Great writeup, thank you for your insights. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2758454,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-04-18T07:03:04.220000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2468601,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-10-05T16:36:19.910000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2412257,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-08-28T07:11:41.900000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2394532,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-08-17T02:01:02.297000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2393797,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-08-16T14:29:28.453000",
      "content": "",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2383620": "I thank the organizer and Kaggle for hosting this interesting competition. I also thank many Kagglers who wrote solutions in past competitions, without which I could not compete for segmentation tasks.\n\n## Overview\n\n* U-Net with MaxViT encoder\n* Input is the standard ash color image\n* Target y is soft; the average of individual masks\n* BCE (binary cross entropy) loss\n* U-Net trained for symmetrized label, y_sym, which is the label shifted by 0.5 pixels\n* Shift-scale-rotate augmentation\n* Additional tiny convolution trained to map y_sym to y\n\n## Rotation augmentation\n\nAugmentation is extremely important in this competition to suppress overfit and train longer. With rotation augmentation I could train 40-50 epochs, compared to 10-20 epochs without augmentation. The test-time augmentation (TTA) is also very effective, adding ~0.006 to the score for free.\n\nThis is basic but does not work as usual in this competition because the label is shifted 0.5 pixels to the right and bottom with respect to the contrails. Random rotation augmentation would shift the label to random directions and make the model impossible to learn the right-bottom aligned ground-truth labels.\n\n## Discover the 0.5-pixel shift\n\nI was very confused when the score dropped with flip and rot90 (multiples of 90° rotation) augmentations. In physics, what symmetry the system has is the first thing to consider, and although westerly winds or Coriolis force could make the physics asymmetric under flip or rotation, I could not believe that those affect the contrail detection. I visualized a prediction with a model trained without augmentation and applied it to a 180°-rotated input image to see why the augmentation did not work.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F954117%2Ffe56c0bee942eb6591adabafbe0daf39%2Fcontrail_label_shift.png?generation=1691669316963209&alt=media) \n\nThe blue-green-red stripe pattern from top to bottom shows false negative, true positive, and false positive, which means that the rotated label is shifted up compared to the predicted label; that is, the original labels are shifted down compared to the contrails. I also observed the same pattern for left and right.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F954117%2Fbb92508e8a6b1d8392819ac63cb36cd6%2Fcontrail_fish.png?generation=1691669701057157&alt=media)\n\nIf the labels are shifted 0.5 pixels to top left (right panel in the figure above), then the rotated label is consistent with the original label in terms of the contrail-label offset.\n\n* I first shift the label by 0.5 pixels and create y_sym\n* y_sym has size 512×512 and is sampled from y on a shifted regular grid with bilinear interpolation\n* U-Net is trained with y_sym \n* Additional tiny convolution of 5x5 with stride 2 is trained to learn the mapping from y_sym to right-bottom aligned y, using 5% of the data when the augmentations are not applied randomly (note that y cannot be augmented).\n* At inference time, y_sym is averaged for 8 patterns of rot90 and flip TTAs. The conv5x5 is applied after the y_sym.\n\nThe augmentation is,\n\n```\nimport albumentations as A\n\nA.Compose([\n    A.RandomRotate90(p=1),\n    A.HorizontalFlip(p=0.5),\n    A.ShiftScaleRotate(rotate_limit=30, scale_limit=0.2)\n])\n```\n\nI refer to my public notebook for the details:\n\nhttps://www.kaggle.com/junkoda/base-unet-model-for-the-1st-place\n\n\n## Models\n\nThe final prediction is the weighted mean of two models with a threshold ~0.45. I tuned the threshold and the weights using the validation set. Both models are U-Net using maxvit_tiny_tf_512.in1k as the encoder, but one uses single time t=4, and the other use four times t = 1 - 4. The input image size is 1024×1024 for both models.\n\n## Single-time model\n\n* Upscale the input image to 1024×1024\n* Drop the final decoder layer of upscale and output 512×512\n\n## 4-panel model\n\nIn order to use the time information, I tried 3D ResNet and ConvLSTM, but I was not able to make them work at all. The only thing I could do was to pack four 512×512 image at t=1,2,3,4 into one 1024×1024 image and hope that self attention in MaxViT look at different times.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F954117%2F291ed2fb89c1e4b492f0d5890ad06868%2Fimg4.png?generation=1691671228909061&alt=media)\n\nIn the U-Net, I only pass the quarter of the features (H/2, W/2) to the decoder, which corresponds to the t=4 quarter. The rest is the same as the single-time model. Concatenating images in the spacial direction appear in Kaggle once in a while. I remember [CPMP's solution](https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275433) for the G2Net blackhole merger competition, in which he stacked 3 detector data horizontally rather than stacking them as 3 channels. \n\nTraining details:\n\n* Train 5 out of k=10 folds, using the out-of-fold to monitor the validation loss\n* 0.5 epochs of linear warmup to learning rate 8e-4, followed by cosine annealing\n* batch size 4 × gradient accumulation 2\n* AdamW with weight_decay = 0.01\n* 35 or 40 epochs of training\n\n```text\nScores\n                    CV     Public  Private\n       input size\n\nSingle time   512  0.697   0.707   0.712\nSingle time  1024  0.703   0.719   0.716\n4-panel      1024  0.704   0.719   0.722\nEnsemble           0.706   0.725   0.724\n```\n\nAll scores are for 5-fold × 8 TTA mean.\n\n## Things I was not able to do\n\n* I did not use pseudo labeling. I tried pseudo labeling at an early stage. It boosted the single model score a lot, but the benefit was unimpressive after 5-fold mean.\n* I was not able to train large models. I tried MaxViT small and base, and many other models, but the improvement was unclear. I increased the input image size instead. This failure must be due to my insufficient experience. \n* I used only positive data (dropping data with no positive labels) for model evaluation because it takes 1/2 time to train. However, I was not able to use it in the real prediction; it performs very bad for negative samples, and my classifier was not good enough to exclude such false negatives.\n* Removing small masks is a basic technique in segmentation competition, but here, removing even 1-pixel clusters (connected components) was not a good idea. I tried to remove false positive clusters based on some cluster statistics (size, density, max density, etc), but U-Net was cleverer than my post processing.\n\n\n## Codes\n\n* Training: https://github.com/junkoda/kaggle_contrails_solution\n* Inference notebook: https://www.kaggle.com/code/junkoda/contrails-submit\n\nUpdate 2023-08-20: Add links to the codes and minor english corrections.\n",
    "2384094": "Congratz on the win ! \n\nVery elegant solution, and impressive detective work. ",
    "2389738": "Would it make sense to shift an image by 0.5 pixel instead of the mask? This way you don't need to shift masks during the inference.",
    "2387328": "\" 4-panel model ... hope that self attention in MaxViT look at different times.\"\nthis actually might work. good idea!\n\nsome smart design of positional encoding (e.g. t,x,y)  or/and attention may make it work better.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fa7af14e17b7344e2ae84349c25fb20eb%2FSelection_999(2895).png?generation=1691854097296090&alt=media)\n",
    "2385090": "Congrats for the winning! Thank you for the explanation, definitely simple solutions works best most of the times!",
    "2383687": "Congrats on the solo win @junkoda, and interesting write-up.\n\nCan you explain how you shift a mask by 1/2 a pixel?",
    "2396992": "@junkoda kudos for winning the competition and writing such a nice writeup!",
    "2394055": "Thanks for the great writeup. OOC, how did you find out about the 0.5 pixel trick? By plotting the data maybe?",
    "2391263": "Great job! :)",
    "2390535": "Congratulations 🎉 Dear @junkoda for this position.",
    "2387813": "Congrats and a great work done.",
    "2387182": "Congrats @junkoda and thanks for sharing your approach.",
    "2385094": "Congrats @junkoda \n\nVery helpful solution.",
    "2384802": "awesome ... congratulations !",
    "2384745": "This is wonderful. Thank you so much for sharing!",
    "2384532": "May I ask what machine you are using for the experiment? These experiments seem to require very good machines",
    "2384163": "Nicely done! I will definitely make use of your technique to shift masked labels as part of the model building framework. Thanks for sharing!",
    "2384000": "Congratulations for the solo win! I would like to ask a few questions:\n\n1. You created 10 folds using train folder, train only 5 of them and validate those on given validation set, right?\n2. Why you passed only the quarter of the features to the decoder?\n3. Why you dropped the final decoder layer of upscale to output 512x512 in the case of single frame model?\n\nThanks for sharing\n",
    "2383999": "Did you use different learning rates for encoder and decoder? It improved my score a lot",
    "2383644": "Thanks for sharing the write up @junkoda ! Kudos on winning the competition.🎉",
    "2407088": "I was just talking with an intern about the need to get into the messiness of the data. It seems that your approach is successful because of how much you dove into the problem and asked about the structure of the data.",
    "2389724": "Congratz on the win!\nI also just found out that the y_sym of your solution is so helpful! Thank you for sharing!",
    "2384023": "Nice shifted mask visualization, congratz with the result! ",
    "2383736": "Many congratulations on the solo win and thank you for the great write up!\n\nCould you please clarify this part:\n>Additional tiny convolution of 5x5 with stride 2 is trained to learn the mapping from y_sym to right-bottom aligned y. This is done when the augmentation is not applied for the 5% of the data (note that y cannot be augmented).\n\nWhere do you put this convolution ? head of the Unet ? When do you train it? What are the 5% of the data you are referring to? Why not just training with centered masks (y_sym) and then apply the invert shift on the predictions ?\n\n>At inference time, y_sym is averaged for 8 patterns of rot90 and flip TTAs. The conv5x5 is applied to the averaged y_sym.\n\nJust to be sure I understood : basically you perform classical 8 times TTA, average the predictions and then shift the results using your trained conv5x5 ?\n",
    "2383712": "Super cool explanation.\n\"I first shift the label by 0.5 pixel and create y_sym\"\n\"Additional tiny convolution of 5x5 with stride 2 is trained to learn the mapping from y_sym to right-bottom aligned y.\"\n\nWhy learning a separate convolution to shift the mask back and not just shift it 1 pixel at 512x512 and downscale to 256x256, which is equivalent to 0.5 px shift)?",
    "2418095": "Hello, congratulations on achieving the first-place result.\n\nI have reviewed your GitHub code and noticed that in \"unet1024/data.py,\" on line 87, the code \"x=f['x'][3,:]\" means that it is selecting data from the fourth time step, and on line 109, you are resizing the image to 1024x1024 with \"x=self.resize(x).\" This approach differs from what you mentioned in the discussion, where you described combining time steps 1-4 to create a 1024x1024 image (\"The only thing I could do was to pack four 512×512 images at t=1,2,3,4 into one 1024×1024 image\").\n\nCould you please clarify this inconsistency and explain how these two approaches might impact the final results differently?",
    "2412261": "May I ask how you obtained the 4-panel model? Did you use some automl tools? Thank you.",
    "2407168": "congrats @junkoda and thanks for sharing your method",
    "2396818": "Congrats! Great writeup, thank you for your insights. ",
    "2758454": "",
    "2468601": "",
    "2412257": "",
    "2394532": "",
    "2393797": "Congratulations and thanks for sharing your approach. I really didn't catch the 0.5-pixel shift. Great work👍"
  }
}