{
  "id": 430479,
  "title": "9th place solution",
  "url": "/competitions/google-research-identify-contrails-reduce-global-warming/discussion/430479",
  "author_name": "tascj",
  "post_date": "2023-08-10T01:10:21.009000",
  "votes": 76,
  "comment_count": 37,
  "views": 0,
  "content": "<p>Thanks to the organizers for hosting competition and congrats to all the winners.</p>\n<h1>The key finding</h1>\n<p>Many people have likely noticed that flip/rot90 augmentation doesn't work well on this dataset.</p>\n<p>My guess is that during the conversion of polygon annotations to binary masks, the leftmost point is not included while the rightmost point is, causing a misalignment between the image and the mask.</p>\n<ul>\n<li>original<ul>\n<li>img:    [0, 128,   128,    3,   4, 5, 6]</li>\n<li>mask: [0, 255a, 255b, 255pad, 0, 0, 0]</li></ul></li>\n<li>normal flip that does not work<ul>\n<li>img: [6, 5, 4, 3,      128,  128,    0]</li>\n<li>mask: [0, 0, 0, 255pad, 255b, 255a, 0]</li></ul></li>\n<li>consistent flip should be<ul>\n<li>img: [6, 5, 4, 3, 128,  128,    0]</li>\n<li>mask: [0, 0, 0, 0, 255b, 255a, 255pad]</li></ul></li>\n</ul>\n<p>I tried two solutions to address this issue.</p>\n<h3>solution 1, train with misaligned (img ,mask)</h3>\n<ul>\n<li>training<ul>\n<li>img: normal flip</li>\n<li>mask: consistent flip</li></ul></li>\n<li>inference<ul>\n<li>flip tta: normal img flip -&gt; predict -&gt; consistent mask flip</li></ul></li>\n</ul>\n<p>training/test time augmentation is a bit tricky in this case.</p>\n<h3>solution 2, train with aligned (img, mask)</h3>\n<ul>\n<li>training<ul>\n<li>shift img by +0.5 pixel</li></ul></li>\n<li>inference<ul>\n<li>shift img by +0.5 pixel</li></ul></li>\n</ul>\n<p>training/test time augmentation is as normal in this case.</p>\n<p>Both solutions worked. For simplicity, I chose solution 2. I use the following code to apply resize512&amp;shift1</p>\n<pre><code>img_affine_matrix = np.array([[2.0, 0.0, 1.5], [0.0, 2.0, 1.5]], =np.float64)\nimg = cv2.warpAffine(\n    img,\n    img_affine_matrix,\n    (512, 512),\n    =cv2.INTER_LINEAR,\n    =cv2.BORDER_CONSTANT,\n    =0,\n)\n</code></pre>\n<p>Note that you need to calibrate M for warpAffine. So it's 1.5 in the final affine matrix.</p>\n<pre><code> calibrate(M):\n    \n    [:, ] += M[:, ] * . + M[:, ] * . - .\n     M\n</code></pre>\n<h1>Models</h1>\n<p>My submission notebook is pubic <a href=\"https://www.kaggle.com/code/tascj0/contrail-submit?scriptVersionId=139432132\" target=\"_blank\">here</a>.</p>\n<p>I joined the competition quite late and only briefly explored other bands and 2.5D before giving up on them. In the end, I only used false_color images and UNet models.</p>\n<ul>\n<li>Strong backbone were very helpful.</li>\n<li>train with all individual annotations is helpful.</li>\n</ul>\n<p>I had some success with pseudo-labeling on a small model, but unfortunately, I didn't have time to apply it on larger models.</p>",
  "messages": [
    {
      "id": 2382704,
      "postDate": "2023-08-10T01:10:21.010Z",
      "content": "<p>Thanks to the organizers for hosting competition and congrats to all the winners.</p>\n<h1>The key finding</h1>\n<p>Many people have likely noticed that flip/rot90 augmentation doesn't work well on this dataset.</p>\n<p>My guess is that during the conversion of polygon annotations to binary masks, the leftmost point is not included while the rightmost point is, causing a misalignment between the image and the mask.</p>\n<ul>\n<li>original<ul>\n<li>img:    [0, 128,   128,    3,   4, 5, 6]</li>\n<li>mask: [0, 255a, 255b, 255pad, 0, 0, 0]</li></ul></li>\n<li>normal flip that does not work<ul>\n<li>img: [6, 5, 4, 3,      128,  128,    0]</li>\n<li>mask: [0, 0, 0, 255pad, 255b, 255a, 0]</li></ul></li>\n<li>consistent flip should be<ul>\n<li>img: [6, 5, 4, 3, 128,  128,    0]</li>\n<li>mask: [0, 0, 0, 0, 255b, 255a, 255pad]</li></ul></li>\n</ul>\n<p>I tried two solutions to address this issue.</p>\n<h3>solution 1, train with misaligned (img ,mask)</h3>\n<ul>\n<li>training<ul>\n<li>img: normal flip</li>\n<li>mask: consistent flip</li></ul></li>\n<li>inference<ul>\n<li>flip tta: normal img flip -&gt; predict -&gt; consistent mask flip</li></ul></li>\n</ul>\n<p>training/test time augmentation is a bit tricky in this case.</p>\n<h3>solution 2, train with aligned (img, mask)</h3>\n<ul>\n<li>training<ul>\n<li>shift img by +0.5 pixel</li></ul></li>\n<li>inference<ul>\n<li>shift img by +0.5 pixel</li></ul></li>\n</ul>\n<p>training/test time augmentation is as normal in this case.</p>\n<p>Both solutions worked. For simplicity, I chose solution 2. I use the following code to apply resize512&amp;shift1</p>\n<pre><code>img_affine_matrix = np.array([[2.0, 0.0, 1.5], [0.0, 2.0, 1.5]], =np.float64)\nimg = cv2.warpAffine(\n    img,\n    img_affine_matrix,\n    (512, 512),\n    =cv2.INTER_LINEAR,\n    =cv2.BORDER_CONSTANT,\n    =0,\n)\n</code></pre>\n<p>Note that you need to calibrate M for warpAffine. So it's 1.5 in the final affine matrix.</p>\n<pre><code> calibrate(M):\n    \n    [:, ] += M[:, ] * . + M[:, ] * . - .\n     M\n</code></pre>\n<h1>Models</h1>\n<p>My submission notebook is pubic <a href=\"https://www.kaggle.com/code/tascj0/contrail-submit?scriptVersionId=139432132\" target=\"_blank\">here</a>.</p>\n<p>I joined the competition quite late and only briefly explored other bands and 2.5D before giving up on them. In the end, I only used false_color images and UNet models.</p>\n<ul>\n<li>Strong backbone were very helpful.</li>\n<li>train with all individual annotations is helpful.</li>\n</ul>\n<p>I had some success with pseudo-labeling on a small model, but unfortunately, I didn't have time to apply it on larger models.</p>",
      "rawMarkdown": "Thanks to the organizers for hosting competition and congrats to all the winners.\n\n\n# The key finding\n\nMany people have likely noticed that flip/rot90 augmentation doesn't work well on this dataset.\n\nMy guess is that during the conversion of polygon annotations to binary masks, the leftmost point is not included while the rightmost point is, causing a misalignment between the image and the mask.\n\n* original\n   * img:    [0, 128,   128,    3,   4, 5, 6]\n   * mask: [0, 255a, 255b, 255pad, 0, 0, 0]\n* normal flip that does not work\n   * img: [6, 5, 4, 3,      128,  128,    0]\n   * mask: [0, 0, 0, 255pad, 255b, 255a, 0]\n* consistent flip should be\n    * img: [6, 5, 4, 3, 128,  128,    0]\n    * mask: [0, 0, 0, 0, 255b, 255a, 255pad]\n\nI tried two solutions to address this issue.\n\n### solution 1, train with misaligned (img ,mask)\n\n* training\n  * img: normal flip\n  * mask: consistent flip\n* inference\n  * flip tta: normal img flip -> predict -> consistent mask flip\n\ntraining/test time augmentation is a bit tricky in this case.\n\n### solution 2, train with aligned (img, mask)\n\n* training\n  * shift img by +0.5 pixel\n* inference\n  * shift img by +0.5 pixel\n\ntraining/test time augmentation is as normal in this case.\n\nBoth solutions worked. For simplicity, I chose solution 2. I use the following code to apply resize512&shift1\n\n```\nimg_affine_matrix = np.array([[2.0, 0.0, 1.5], [0.0, 2.0, 1.5]], dtype=np.float64)\nimg = cv2.warpAffine(\n    img,\n    img_affine_matrix,\n    (512, 512),\n    flags=cv2.INTER_LINEAR,\n    borderMode=cv2.BORDER_CONSTANT,\n    borderValue=0,\n)\n```\n\nNote that you need to calibrate M for warpAffine. So it's 1.5 in the final affine matrix.\n```\ndef calibrate(M):\n    # dst + 0.5 = M(src + 0.5)\n    M[:, 2] += M[:, 0] * 0.5 + M[:, 1] * 0.5 - 0.5\n    return M\n```\n\n# Models\n\nMy submission notebook is pubic [here](https://www.kaggle.com/code/tascj0/contrail-submit?scriptVersionId=139432132).\n\nI joined the competition quite late and only briefly explored other bands and 2.5D before giving up on them. In the end, I only used false_color images and UNet models.\n\n* Strong backbone were very helpful.\n* train with all individual annotations is helpful.\n\nI had some success with pseudo-labeling on a small model, but unfortunately, I didn't have time to apply it on larger models.\n",
      "votes": 76
    },
    {
      "id": 2382723,
      "postDate": "2023-08-10T01:32:59.723Z",
      "content": "<p>just to put into picture on the annotation issues of flipping</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fe61b0795055ea79c56f425587a6f5b8c%2FSelection_999(2868).png?generation=1691631177229301&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "just to put into picture on the annotation issues of flipping\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fe61b0795055ea79c56f425587a6f5b8c%2FSelection_999(2868).png?generation=1691631177229301&alt=media)",
      "votes": 25,
      "replies": [
        {
          "id": 2382724,
          "postDate": "2023-08-10T01:36:04.970Z",
          "content": "<p>note this is an issue becuase our mask is very thin and has low area.<br>\nit is not an issue if the mask area is very big.</p>\n<p>i wonder if the other public benchmark dataset has this issue? e.g. coco-instance segmentation, cityscapes, etc???</p>",
          "rawMarkdown": "note this is an issue becuase our mask is very thin and has low area.\nit is not an issue if the mask area is very big.\n\ni wonder if the other public benchmark dataset has this issue? e.g. coco-instance segmentation, cityscapes, etc???",
          "votes": 7,
          "replies": [
            {
              "id": 2382727,
              "postDate": "2023-08-10T01:51:50.947Z",
              "content": "<p>i think of a way to prove that the obersvation is correct and how to detect this in future competition:<br>\n1) just train normal image + normal truth mask<br>\n2) apply model on flipped image and get predicted mask<br>\n3) compute difference = flipped truth mask - predicted mask<br>\n<strong>4) you should see consistence difference (e.g. an offset) not random difference</strong><br>\n5) if you want, you can train a correction networkto label flipped image:<br>\n  net(predicted mask from flipped image) --&gt; flipped truth mask</p>",
              "rawMarkdown": "i think of a way to prove that the obersvation is correct and how to detect this in future competition:\n1) just train normal image + normal truth mask\n2) apply model on flipped image and get predicted mask\n3) compute difference = flipped truth mask - predicted mask\n**4) you should see consistence difference (e.g. an offset) not random difference**\n5) if you want, you can train a correction networkto label flipped image:\n  net(predicted mask from flipped image) --> flipped truth mask",
              "votes": 10
            },
            {
              "id": 2382746,
              "postDate": "2023-08-10T02:26:21.773Z",
              "content": "<p>Thanks for your explaination. I understand.</p>",
              "rawMarkdown": "Thanks for your explaination. I understand."
            },
            {
              "id": 2382874,
              "postDate": "2023-08-10T04:44:19.610Z",
              "content": "<p>Thank you on explaining how we can detect this in future competition, it really helps a lot. Is there any other way also ? (Just wanted to know)</p>",
              "rawMarkdown": "Thank you on explaining how we can detect this in future competition, it really helps a lot. Is there any other way also ? (Just wanted to know)"
            },
            {
              "id": 2382954,
              "postDate": "2023-08-10T05:52:04.357Z",
              "content": "<p>Thanks for the illustration, however, I still don't get the relation between “faulty ground truth mask\" and \"created from polygon\"</p>",
              "rawMarkdown": "Thanks for the illustration, however, I still don't get the relation between “faulty ground truth mask\" and \"created from polygon\"",
              "votes": 1
            },
            {
              "id": 2383087,
              "postDate": "2023-08-10T07:10:46.537Z",
              "content": "<p>Thanks for the images!<br>\nCan you please specify it a bit more?<br>\nDoes it mean flipping causes model confusion due to the symmetry of original image?<br>\nIt would be cool to the image, which presents the issue in from flipping along with the image showing how it should be.</p>",
              "rawMarkdown": "Thanks for the images!\nCan you please specify it a bit more?\nDoes it mean flipping causes model confusion due to the symmetry of original image?\nIt would be cool to the image, which presents the issue in from flipping along with the image showing how it should be."
            },
            {
              "id": 2383323,
              "postDate": "2023-08-10T10:00:03.407Z",
              "content": "<p>Finally I understood <a href=\"https://www.kaggle.com/tascj0\" target=\"_blank\">@tascj0</a> 's words by reading it several times word by word. It's very clear using his example:</p>\n<pre><code>original\n: \n: [, a, b, pad, , , ]\n flip that does not work\n: \n: [, , , pad, b, a, ]\nconsistent flip should be\n: \n: [, , , , b, a, pad]\n</code></pre>\n<p><strong>In short, all the ground truths are shifted to the right.</strong>  If you do <code>HFlip</code> to the mask, you'll make it shift to the left. So you need to shift the Hflipped mask to the right.</p>",
              "rawMarkdown": "Finally I understood @tascj0 's words by reading it several times word by word. It's very clear using his example:\n```\noriginal\nimg: [0, 128, 128, 3, 4, 5, 6]\nmask: [0, 255a, 255b, 255pad, 0, 0, 0]\nnormal flip that does not work\nimg: [6, 5, 4, 3, 128, 128, 0]\nmask: [0, 0, 0, 255pad, 255b, 255a, 0]\nconsistent flip should be\nimg: [6, 5, 4, 3, 128, 128, 0]\nmask: [0, 0, 0, 0, 255b, 255a, 255pad]\n```\n**In short, all the ground truths are shifted to the right.**  If you do `HFlip` to the mask, you'll make it shift to the left. So you need to shift the Hflipped mask to the right.",
              "votes": 2
            }
          ]
        },
        {
          "id": 2382767,
          "postDate": "2023-08-10T02:49:01.223Z",
          "content": "<p>Thank you for the nice picture. I think it is either a pixel missing on top&amp;left or an extra pixel on bottom&amp;right. The idea is the same anyway.</p>",
          "rawMarkdown": "Thank you for the nice picture. I think it is either a pixel missing on top&left or an extra pixel on bottom&right. The idea is the same anyway."
        },
        {
          "id": 2406274,
          "postDate": "2023-08-24T10:22:37.433Z",
          "content": "<p>The visualization is great, but I wanted to ask if the shift displayed might be reversed. <br>\nThe write-ups mention a bottom-right shift of the label in relation to the image, the displayed label appears to be top-left shifted in relation to the image. </p>",
          "rawMarkdown": "The visualization is great, but I wanted to ask if the shift displayed might be reversed. \nThe write-ups mention a bottom-right shift of the label in relation to the image, the displayed label appears to be top-left shifted in relation to the image. \n"
        }
      ]
    },
    {
      "id": 2383609,
      "postDate": "2023-08-10T13:19:53.077Z",
      "content": "<p>Congratz on the solo submission gold, very impressive.</p>\n<p>We completely missed the issues with masks, great catch. I noticed very early that flips were hurting results and assumed that was because the background was similar between images in train / val. I can't believe we still got 2nd place while overlooking that.<br>\nI guess our models learnt the pattern, but still, it most likely hurt performances.</p>",
      "rawMarkdown": "Congratz on the solo submission gold, very impressive.\n\nWe completely missed the issues with masks, great catch. I noticed very early that flips were hurting results and assumed that was because the background was similar between images in train / val. I can't believe we still got 2nd place while overlooking that.\nI guess our models learnt the pattern, but still, it most likely hurt performances.",
      "votes": 3,
      "replies": [
        {
          "id": 2384669,
          "postDate": "2023-08-11T03:12:56.290Z",
          "content": "<p>Congratulations on achieving 2nd place! Training a strong model without flip/rot90 augmentation is challenging, and it's incredible that you managed to achieve it.</p>\n<p>Yes, the model can learn such a pattern as long as you don't introduce inconsistency by adding flip/rot90. In fact, you can perform flip/rot90 augmentation while keeping the pattern (my solution 1).</p>",
          "rawMarkdown": "Congratulations on achieving 2nd place! Training a strong model without flip/rot90 augmentation is challenging, and it's incredible that you managed to achieve it.\n\nYes, the model can learn such a pattern as long as you don't introduce inconsistency by adding flip/rot90. In fact, you can perform flip/rot90 augmentation while keeping the pattern (my solution 1).",
          "votes": 1
        },
        {
          "id": 2385628,
          "postDate": "2023-08-11T13:17:38.320Z",
          "content": "<p><a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> do you think the transformer models were more likely to learn where the north was based on background image ? So that they can shift the predictions on the correct side? ^^</p>",
          "rawMarkdown": "@theoviel do you think the transformer models were more likely to learn where the north was based on background image ? So that they can shift the predictions on the correct side? ^^",
          "replies": [
            {
              "id": 2385633,
              "postDate": "2023-08-11T13:24:36.470Z",
              "content": "<p>It's possible to learn where the North (and West) is since most images cover the same area of the planet (~central America), but fairly hard. Maybe transformers can but I'm not sure whether they can shift predictions by themselves.</p>",
              "rawMarkdown": "It's possible to learn where the North (and West) is since most images cover the same area of the planet (~central America), but fairly hard. Maybe transformers can but I'm not sure whether they can shift predictions by themselves."
            }
          ]
        }
      ]
    },
    {
      "id": 2382709,
      "postDate": "2023-08-10T01:16:02.153Z",
      "content": "<p>Very impressive given the number of subs and the time you had. Congratulations</p>",
      "rawMarkdown": "Very impressive given the number of subs and the time you had. Congratulations",
      "votes": 3
    },
    {
      "id": 2383289,
      "postDate": "2023-08-10T09:35:11.883Z",
      "content": "<p>Impressive 👏👏👏👏<br>\nCongratulations!</p>",
      "rawMarkdown": "Impressive 👏👏👏👏\nCongratulations!",
      "votes": 1
    },
    {
      "id": 2383098,
      "postDate": "2023-08-10T07:16:31.137Z",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/tascj0\" target=\"_blank\">@tascj0</a> , and congratulations on your solo gold with one sub.</p>\n<blockquote>\n  <p>My guess is that during the conversion of polygon annotations to binary masks, the leftmost point is not included while the rightmost point is, causing a misalignment between the image and the mask.</p>\n</blockquote>\n<p>Why it's not \"the leftmost point is included while the rightmost point isn't\"? How did you detect this? Thanks</p>",
      "rawMarkdown": "Thanks for sharing @tascj0 , and congratulations on your solo gold with one sub.\n\n> My guess is that during the conversion of polygon annotations to binary masks, the leftmost point is not included while the rightmost point is, causing a misalignment between the image and the mask.\n\nWhy it's not \"the leftmost point is included while the rightmost point isn't\"? How did you detect this? Thanks",
      "votes": 1,
      "replies": [
        {
          "id": 2383245,
          "postDate": "2023-08-10T09:03:14.103Z",
          "content": "<blockquote>\n  <p>Why it's not \"the leftmost point is included while the rightmost point isn't\"?</p>\n</blockquote>\n<p>I could be. It could also be both included but the rightmost coordinates+1.<br>\nI tried shift img by -0.5 and 0.5 and it turned out 0.5 is correct.</p>\n<blockquote>\n  <p>How did you detect this?</p>\n</blockquote>\n<p>What we know?</p>\n<ol>\n<li>Annotators draw polygons.</li>\n<li>The competition provides binary masks.</li>\n<li>Flip augmentation doesn't work.</li>\n</ol>\n<p>It's easy to infer that there might be issues with the implementation of polygon-to-mask conversion. It's also possible that there were some problems with the annotation software in handling the coordinates.</p>",
          "rawMarkdown": "> Why it's not \"the leftmost point is included while the rightmost point isn't\"?\n\nI could be. It could also be both included but the rightmost coordinates+1.\nI tried shift img by -0.5 and 0.5 and it turned out 0.5 is correct.\n\n> How did you detect this?\n\nWhat we know?\n1. Annotators draw polygons.\n2. The competition provides binary masks.\n3. Flip augmentation doesn't work.\n\nIt's easy to infer that there might be issues with the implementation of polygon-to-mask conversion. It's also possible that there were some problems with the annotation software in handling the coordinates.",
          "votes": 4,
          "replies": [
            {
              "id": 2383316,
              "postDate": "2023-08-10T09:55:39.130Z",
              "content": "<p>Thanks for the detailed explanation. I didn't get any inspiration from these evidences😭</p>",
              "rawMarkdown": "Thanks for the detailed explanation. I didn't get any inspiration from these evidences😭"
            }
          ]
        }
      ]
    },
    {
      "id": 2383040,
      "postDate": "2023-08-10T06:37:12.993Z",
      "content": "<p><a href=\"https://www.kaggle.com/tascj0\" target=\"_blank\">@tascj0</a> Congratulations for such results with a single efficient submission!</p>",
      "rawMarkdown": "@tascj0 Congratulations for such results with a single efficient submission!",
      "votes": 1
    },
    {
      "id": 2382898,
      "postDate": "2023-08-10T05:07:57.027Z",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/tascj0\" target=\"_blank\">@tascj0</a> for sharing your work, which really helpful in our learning process.</p>",
      "rawMarkdown": "Thanks @tascj0 for sharing your work, which really helpful in our learning process.",
      "votes": 1
    },
    {
      "id": 2382832,
      "postDate": "2023-08-10T04:16:02.177Z",
      "content": "<p>Congratulations on winning top position in the Competition.<br>\nThanks for sharing your notebook and useful inputs. </p>",
      "rawMarkdown": "Congratulations on winning top position in the Competition.\nThanks for sharing your notebook and useful inputs. ",
      "votes": 1
    },
    {
      "id": 2382728,
      "postDate": "2023-08-10T01:51:58.053Z",
      "content": "<p>Can you share your validation strategy in this competition please? Also how did you do pseudo-labeling and did you used it in final ensemble for some models? As for training on the individual annotation: did you train each model on different annotation separately? I didn't understand this part actually. </p>",
      "rawMarkdown": "Can you share your validation strategy in this competition please? Also how did you do pseudo-labeling and did you used it in final ensemble for some models? As for training on the individual annotation: did you train each model on different annotation separately? I didn't understand this part actually. ",
      "votes": 1,
      "replies": [
        {
          "id": 2382771,
          "postDate": "2023-08-10T02:56:50.680Z",
          "content": "<blockquote>\n  <p>Can you share your validation strategy in this competition please? </p>\n</blockquote>\n<p>I use the train/validation folders for as train/val set.</p>\n<blockquote>\n  <p>did you used it in final ensemble for some models</p>\n</blockquote>\n<p>I tried pseudo-labeling on smaller models at 256 input. Not used in the final submission. Unannotated data is 7x of annotated data, I did not have time to do pseudo-labeling using larger models at 512 input.</p>\n<blockquote>\n  <p>training on the individual annotation</p>\n</blockquote>\n<p>Use <code>human_individual_masks</code> instead of <code>human_pixel_masks</code>.</p>",
          "rawMarkdown": "> Can you share your validation strategy in this competition please? \n\nI use the train/validation folders for as train/val set.\n\n> did you used it in final ensemble for some models\n\nI tried pseudo-labeling on smaller models at 256 input. Not used in the final submission. Unannotated data is 7x of annotated data, I did not have time to do pseudo-labeling using larger models at 512 input.\n\n> training on the individual annotation\n\nUse `human_individual_masks` instead of `human_pixel_masks`.",
          "votes": 1,
          "replies": [
            {
              "id": 2382891,
              "postDate": "2023-08-10T04:55:38.887Z",
              "content": "<p><a href=\"https://www.kaggle.com/tascj0\" target=\"_blank\">@tascj0</a> How human_individual_masks is helpfull, I mean why it performs better ?</p>",
              "rawMarkdown": "@tascj0 How human_individual_masks is helpfull, I mean why it performs better ?"
            },
            {
              "id": 2383091,
              "postDate": "2023-08-10T07:12:57.720Z",
              "content": "<p>Can you please share how did you use individual masks?<br>\nDid you use 4 individual masks for one image as more training data?</p>",
              "rawMarkdown": "Can you please share how did you use individual masks?\nDid you use 4 individual masks for one image as more training data?",
              "votes": 1
            },
            {
              "id": 2383248,
              "postDate": "2023-08-10T09:07:37.297Z",
              "content": "<ol>\n<li>when collating a batch, stack all individual masks and keep corresponding img batch indices.</li>\n</ol>\n<pre><code>def (samples):\n    out = ()\n    out[] = torch.([s[] for s in samples])\n    masks = [s[] for s in samples]\n    batch_inds = []\n    for idx, mask in (masks):\n        for _ in (mask.()):\n            batch_inds.(idx)\n    out[] = torch.(masks, dim=).()\n    out[] = torch.(batch_inds)\n    return out\n</code></pre>\n<ol>\n<li>in the training loop</li>\n</ol>\n<pre><code> \n    img = \n    mask = \n    batch_inds = \n\n    out = model(img)\n    out = out[batch_inds]\n    loss = loss_fn(out, mask)\n</code></pre>",
              "rawMarkdown": "1. when collating a batch, stack all individual masks and keep corresponding img batch indices.\n\n```\ndef collate_fn(samples):\n    out = dict()\n    out[\"img\"] = torch.stack([s[\"img\"] for s in samples])\n    masks = [s[\"mask\"] for s in samples]\n    batch_inds = []\n    for idx, mask in enumerate(masks):\n        for _ in range(mask.size(0)):\n            batch_inds.append(idx)\n    out[\"mask\"] = torch.cat(masks, dim=0).unsqueeze(1)\n    out[\"batch_inds\"] = torch.LongTensor(batch_inds)\n    return out\n```\n\n2. in the training loop\n\n```\nfor data in train_dataloader:\n    img = data[\"img\"].cuda()\n    mask = data[\"mask\"].cuda()\n    batch_inds = data[\"batch_inds\"].cuda()\n\n    out = model(img)\n    out = out[batch_inds]\n    loss = loss_fn(out, mask)\n```",
              "votes": 1
            },
            {
              "id": 2383265,
              "postDate": "2023-08-10T09:25:43.983Z",
              "content": "<p>Well done! How did you come to the conclusion, that the validation set is trustworthy in this competition? Why didn't you use other validation strategies? </p>",
              "rawMarkdown": "Well done! How did you come to the conclusion, that the validation set is trustworthy in this competition? Why didn't you use other validation strategies? "
            }
          ]
        }
      ]
    },
    {
      "id": 2382708,
      "postDate": "2023-08-10T01:15:38.183Z",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!",
      "votes": 1
    },
    {
      "id": 2413863,
      "postDate": "2023-08-29T07:46:38.083Z",
      "content": "<p>大神好，想请教3个问题:<br>\n1、如果mask向右下偏移了0.5个像素，那水平翻转和垂直翻转的时候是否要把mask分别向右和向下偏移1个像素呢？<br>\n2、在您的实验中，这个偏移0.5像素的发现对分数的改善有多大<br>\n3、您笔记里的5个backbone是如何确定的，是用了某个类似automl自动搜参工具的吗，还是凭借丰富的经验。我是新手，同时也对自动搜参很感兴趣😅您觉得自动搜参帮助大吗以及能否给菜鸟们推荐一款（如果您用过的话）<br>\n很抱歉问的几个新手问题，若能回复，将不胜感激。</p>",
      "rawMarkdown": "大神好，想请教3个问题:\n1、如果mask向右下偏移了0.5个像素，那水平翻转和垂直翻转的时候是否要把mask分别向右和向下偏移1个像素呢？\n2、在您的实验中，这个偏移0.5像素的发现对分数的改善有多大\n3、您笔记里的5个backbone是如何确定的，是用了某个类似automl自动搜参工具的吗，还是凭借丰富的经验。我是新手，同时也对自动搜参很感兴趣😅您觉得自动搜参帮助大吗以及能否给菜鸟们推荐一款（如果您用过的话）\n很抱歉问的几个新手问题，若能回复，将不胜感激。",
      "replies": [
        {
          "id": 2417947,
          "postDate": "2023-09-01T02:59:14.760Z",
          "content": "<blockquote>\n  <p>1、如果mask向右下偏移了0.5个像素，那水平翻转和垂直翻转的时候是否要把mask分别向右和向下偏移1个像素呢？</p>\n</blockquote>\n<p>I believe the misalignment is not caused by some 0.5 pixel shift, it is caused by the 1-extra-pixel on the bottom-right. A shift of 0.5 pixels can bring back symmetry.</p>\n<p>Without symmetry, a naive implementation of flip/rot90 augmentation would introduce inconsistency, which confuses the model in training.</p>\n<p>As illustrated in the main post, there are two solutions:</p>\n<ol>\n<li>Keep the misalignment (keep the asymmetry), and address the inconsistency.</li>\n<li>Fix the misalignment (bring back symmetry)</li>\n</ol>\n<blockquote>\n  <p>2、在您的实验中，这个偏移0.5像素的发现对分数的改善有多大</p>\n</blockquote>\n<p>I believe the shift itself dost not help. The gain is mainly from flip/rot90 augmentation and TTA. I suggest to try it yourself.</p>\n<blockquote>\n  <p>3、您笔记里的5个backbone是如何确定的</p>\n</blockquote>\n<p>I joined the competition late (6 days before the deadline). In the final stage I simply picked several heavy backbones that could finish training before the deadline. I did not explore much here.</p>\n<p>I rarely adjust training parameters other than the learning rate. I don't have much experience with automated parameter search. The benefits of tuning parameters might not be significant after ensembling, in my opinion, the cost-effectiveness might not be high.</p>",
          "rawMarkdown": "> 1、如果mask向右下偏移了0.5个像素，那水平翻转和垂直翻转的时候是否要把mask分别向右和向下偏移1个像素呢？\n\nI believe the misalignment is not caused by some 0.5 pixel shift, it is caused by the 1-extra-pixel on the bottom-right. A shift of 0.5 pixels can bring back symmetry.\n\nWithout symmetry, a naive implementation of flip/rot90 augmentation would introduce inconsistency, which confuses the model in training.\n\nAs illustrated in the main post, there are two solutions:\n\n1. Keep the misalignment (keep the asymmetry), and address the inconsistency.\n2. Fix the misalignment (bring back symmetry)\n\n> 2、在您的实验中，这个偏移0.5像素的发现对分数的改善有多大\n\nI believe the shift itself dost not help. The gain is mainly from flip/rot90 augmentation and TTA. I suggest to try it yourself.\n\n> 3、您笔记里的5个backbone是如何确定的\n\nI joined the competition late (6 days before the deadline). In the final stage I simply picked several heavy backbones that could finish training before the deadline. I did not explore much here.\n\nI rarely adjust training parameters other than the learning rate. I don't have much experience with automated parameter search. The benefits of tuning parameters might not be significant after ensembling, in my opinion, the cost-effectiveness might not be high."
        }
      ]
    },
    {
      "id": 2386823,
      "postDate": "2023-08-12T07:09:26.487Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> on your solo gold. It's a great achievement.</p>\n<p>Your solution is really helpful to learn from.</p>",
      "rawMarkdown": "Congrats @hengck23 on your solo gold. It's a great achievement.\n\nYour solution is really helpful to learn from."
    },
    {
      "id": 2386410,
      "postDate": "2023-08-11T23:21:11.030Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/tascj0\" target=\"_blank\">@tascj0</a> !</p>\n<p>I have a question regarding the affine matrix.</p>\n<p>If the pixels were shifted by 0.5 pixel in the original image space and you upscale the image by 2, shouldn't the shift now be 1 pixel? Where does the 1.5 pixel shift come from?</p>",
      "rawMarkdown": "Congratulations @tascj0 !\n\nI have a question regarding the affine matrix.\n\nIf the pixels were shifted by 0.5 pixel in the original image space and you upscale the image by 2, shouldn't the shift now be 1 pixel? Where does the 1.5 pixel shift come from?",
      "replies": [
        {
          "id": 2386426,
          "postDate": "2023-08-11T23:53:22.380Z",
          "content": "<p>It's about the OpenCV coordinate system.</p>\n<p>In short, the rule is <code>dst+0.5 = M(src+0.5)</code>. After calibration, the <code>1</code> becomes <code>1.5</code>.</p>",
          "rawMarkdown": "It's about the OpenCV coordinate system.\n\nIn short, the rule is ` dst+0.5 = M(src+0.5)`. After calibration, the `1` becomes `1.5`.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2383117,
      "postDate": "2023-08-10T07:28:11.793Z",
      "content": "<p>Thanks so much for the writeup.<br>\nAmazing single submission result! Takes guts to stick to your own CV.<br>\nDo you have any plans on releasing your training code?</p>",
      "rawMarkdown": "Thanks so much for the writeup.\nAmazing single submission result! Takes guts to stick to your own CV.\nDo you have any plans on releasing your training code?"
    },
    {
      "id": 2413850,
      "postDate": "2023-08-29T07:36:26.743Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2413846,
      "postDate": "2023-08-29T07:34:13.710Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2382723,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-08-10T01:32:59.723000",
      "content": "<p>just to put into picture on the annotation issues of flipping</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fe61b0795055ea79c56f425587a6f5b8c%2FSelection_999(2868).png?generation=1691631177229301&amp;alt=media\" alt=\"\"></p>",
      "votes": 25,
      "replies": [
        {
          "id": 2382724,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2023-08-10T01:36:04.970000",
          "content": "<p>note this is an issue becuase our mask is very thin and has low area.<br>\nit is not an issue if the mask area is very big.</p>\n<p>i wonder if the other public benchmark dataset has this issue? e.g. coco-instance segmentation, cityscapes, etc???</p>",
          "votes": 7,
          "replies": [
            {
              "id": 2382727,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-08-10T01:51:50.947000",
              "content": "<p>i think of a way to prove that the obersvation is correct and how to detect this in future competition:<br>\n1) just train normal image + normal truth mask<br>\n2) apply model on flipped image and get predicted mask<br>\n3) compute difference = flipped truth mask - predicted mask<br>\n<strong>4) you should see consistence difference (e.g. an offset) not random difference</strong><br>\n5) if you want, you can train a correction networkto label flipped image:<br>\n  net(predicted mask from flipped image) --&gt; flipped truth mask</p>",
              "votes": 10,
              "replies": []
            },
            {
              "id": 2382746,
              "author_name": "ynhuhu",
              "author_url": "",
              "post_date": "2023-08-10T02:26:21.773000",
              "content": "<p>Thanks for your explaination. I understand.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2382874,
              "author_name": "sahilsg",
              "author_url": "",
              "post_date": "2023-08-10T04:44:19.610000",
              "content": "<p>Thank you on explaining how we can detect this in future competition, it really helps a lot. Is there any other way also ? (Just wanted to know)</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2382954,
              "author_name": "william.wu",
              "author_url": "",
              "post_date": "2023-08-10T05:52:04.357000",
              "content": "<p>Thanks for the illustration, however, I still don't get the relation between “faulty ground truth mask\" and \"created from polygon\"</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2383087,
              "author_name": "iadduk",
              "author_url": "",
              "post_date": "2023-08-10T07:10:46.537000",
              "content": "<p>Thanks for the images!<br>\nCan you please specify it a bit more?<br>\nDoes it mean flipping causes model confusion due to the symmetry of original image?<br>\nIt would be cool to the image, which presents the issue in from flipping along with the image showing how it should be.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2383323,
              "author_name": "william.wu",
              "author_url": "",
              "post_date": "2023-08-10T10:00:03.407000",
              "content": "<p>Finally I understood <a href=\"https://www.kaggle.com/tascj0\" target=\"_blank\">@tascj0</a> 's words by reading it several times word by word. It's very clear using his example:</p>\n<pre><code>original\n: \n: [, a, b, pad, , , ]\n flip that does not work\n: \n: [, , , pad, b, a, ]\nconsistent flip should be\n: \n: [, , , , b, a, pad]\n</code></pre>\n<p><strong>In short, all the ground truths are shifted to the right.</strong>  If you do <code>HFlip</code> to the mask, you'll make it shift to the left. So you need to shift the Hflipped mask to the right.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        },
        {
          "id": 2382767,
          "author_name": "tascj",
          "author_url": "",
          "post_date": "2023-08-10T02:49:01.223000",
          "content": "<p>Thank you for the nice picture. I think it is either a pixel missing on top&amp;left or an extra pixel on bottom&amp;right. The idea is the same anyway.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2406274,
          "author_name": "Raki",
          "author_url": "",
          "post_date": "2023-08-24T10:22:37.433000",
          "content": "<p>The visualization is great, but I wanted to ask if the shift displayed might be reversed. <br>\nThe write-ups mention a bottom-right shift of the label in relation to the image, the displayed label appears to be top-left shifted in relation to the image. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2383609,
      "author_name": "Theo Viel",
      "author_url": "",
      "post_date": "2023-08-10T13:19:53.077000",
      "content": "<p>Congratz on the solo submission gold, very impressive.</p>\n<p>We completely missed the issues with masks, great catch. I noticed very early that flips were hurting results and assumed that was because the background was similar between images in train / val. I can't believe we still got 2nd place while overlooking that.<br>\nI guess our models learnt the pattern, but still, it most likely hurt performances.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2384669,
          "author_name": "tascj",
          "author_url": "",
          "post_date": "2023-08-11T03:12:56.290000",
          "content": "<p>Congratulations on achieving 2nd place! Training a strong model without flip/rot90 augmentation is challenging, and it's incredible that you managed to achieve it.</p>\n<p>Yes, the model can learn such a pattern as long as you don't introduce inconsistency by adding flip/rot90. In fact, you can perform flip/rot90 augmentation while keeping the pattern (my solution 1).</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2385628,
          "author_name": "Optimo",
          "author_url": "",
          "post_date": "2023-08-11T13:17:38.320000",
          "content": "<p><a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> do you think the transformer models were more likely to learn where the north was based on background image ? So that they can shift the predictions on the correct side? ^^</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2385633,
              "author_name": "Theo Viel",
              "author_url": "",
              "post_date": "2023-08-11T13:24:36.470000",
              "content": "<p>It's possible to learn where the North (and West) is since most images cover the same area of the planet (~central America), but fairly hard. Maybe transformers can but I'm not sure whether they can shift predictions by themselves.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2382709,
      "author_name": "JEANMPIA",
      "author_url": "",
      "post_date": "2023-08-10T01:16:02.153000",
      "content": "<p>Very impressive given the number of subs and the time you had. Congratulations</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2383289,
      "author_name": "Remek Kinas",
      "author_url": "",
      "post_date": "2023-08-10T09:35:11.883000",
      "content": "<p>Impressive 👏👏👏👏<br>\nCongratulations!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2383098,
      "author_name": "william.wu",
      "author_url": "",
      "post_date": "2023-08-10T07:16:31.137000",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/tascj0\" target=\"_blank\">@tascj0</a> , and congratulations on your solo gold with one sub.</p>\n<blockquote>\n  <p>My guess is that during the conversion of polygon annotations to binary masks, the leftmost point is not included while the rightmost point is, causing a misalignment between the image and the mask.</p>\n</blockquote>\n<p>Why it's not \"the leftmost point is included while the rightmost point isn't\"? How did you detect this? Thanks</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2383245,
          "author_name": "tascj",
          "author_url": "",
          "post_date": "2023-08-10T09:03:14.103000",
          "content": "<blockquote>\n  <p>Why it's not \"the leftmost point is included while the rightmost point isn't\"?</p>\n</blockquote>\n<p>I could be. It could also be both included but the rightmost coordinates+1.<br>\nI tried shift img by -0.5 and 0.5 and it turned out 0.5 is correct.</p>\n<blockquote>\n  <p>How did you detect this?</p>\n</blockquote>\n<p>What we know?</p>\n<ol>\n<li>Annotators draw polygons.</li>\n<li>The competition provides binary masks.</li>\n<li>Flip augmentation doesn't work.</li>\n</ol>\n<p>It's easy to infer that there might be issues with the implementation of polygon-to-mask conversion. It's also possible that there were some problems with the annotation software in handling the coordinates.</p>",
          "votes": 4,
          "replies": [
            {
              "id": 2383316,
              "author_name": "william.wu",
              "author_url": "",
              "post_date": "2023-08-10T09:55:39.130000",
              "content": "<p>Thanks for the detailed explanation. I didn't get any inspiration from these evidences😭</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2383040,
      "author_name": "MPWARE",
      "author_url": "",
      "post_date": "2023-08-10T06:37:12.993000",
      "content": "<p><a href=\"https://www.kaggle.com/tascj0\" target=\"_blank\">@tascj0</a> Congratulations for such results with a single efficient submission!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2382898,
      "author_name": "Tariq Mahmood",
      "author_url": "",
      "post_date": "2023-08-10T05:07:57.027000",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/tascj0\" target=\"_blank\">@tascj0</a> for sharing your work, which really helpful in our learning process.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2382832,
      "author_name": "C R Suthikshn Kumar",
      "author_url": "",
      "post_date": "2023-08-10T04:16:02.177000",
      "content": "<p>Congratulations on winning top position in the Competition.<br>\nThanks for sharing your notebook and useful inputs. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2382728,
      "author_name": "Man of the year",
      "author_url": "",
      "post_date": "2023-08-10T01:51:58.053000",
      "content": "<p>Can you share your validation strategy in this competition please? Also how did you do pseudo-labeling and did you used it in final ensemble for some models? As for training on the individual annotation: did you train each model on different annotation separately? I didn't understand this part actually. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2382771,
          "author_name": "tascj",
          "author_url": "",
          "post_date": "2023-08-10T02:56:50.680000",
          "content": "<blockquote>\n  <p>Can you share your validation strategy in this competition please? </p>\n</blockquote>\n<p>I use the train/validation folders for as train/val set.</p>\n<blockquote>\n  <p>did you used it in final ensemble for some models</p>\n</blockquote>\n<p>I tried pseudo-labeling on smaller models at 256 input. Not used in the final submission. Unannotated data is 7x of annotated data, I did not have time to do pseudo-labeling using larger models at 512 input.</p>\n<blockquote>\n  <p>training on the individual annotation</p>\n</blockquote>\n<p>Use <code>human_individual_masks</code> instead of <code>human_pixel_masks</code>.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2382891,
              "author_name": "sahilsg",
              "author_url": "",
              "post_date": "2023-08-10T04:55:38.887000",
              "content": "<p><a href=\"https://www.kaggle.com/tascj0\" target=\"_blank\">@tascj0</a> How human_individual_masks is helpfull, I mean why it performs better ?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2383091,
              "author_name": "iadduk",
              "author_url": "",
              "post_date": "2023-08-10T07:12:57.720000",
              "content": "<p>Can you please share how did you use individual masks?<br>\nDid you use 4 individual masks for one image as more training data?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2383248,
              "author_name": "tascj",
              "author_url": "",
              "post_date": "2023-08-10T09:07:37.297000",
              "content": "<ol>\n<li>when collating a batch, stack all individual masks and keep corresponding img batch indices.</li>\n</ol>\n<pre><code>def (samples):\n    out = ()\n    out[] = torch.([s[] for s in samples])\n    masks = [s[] for s in samples]\n    batch_inds = []\n    for idx, mask in (masks):\n        for _ in (mask.()):\n            batch_inds.(idx)\n    out[] = torch.(masks, dim=).()\n    out[] = torch.(batch_inds)\n    return out\n</code></pre>\n<ol>\n<li>in the training loop</li>\n</ol>\n<pre><code> \n    img = \n    mask = \n    batch_inds = \n\n    out = model(img)\n    out = out[batch_inds]\n    loss = loss_fn(out, mask)\n</code></pre>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2383265,
              "author_name": "Man of the year",
              "author_url": "",
              "post_date": "2023-08-10T09:25:43.983000",
              "content": "<p>Well done! How did you come to the conclusion, that the validation set is trustworthy in this competition? Why didn't you use other validation strategies? </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2382708,
      "author_name": "ynhuhu",
      "author_url": "",
      "post_date": "2023-08-10T01:15:38.183000",
      "content": "<p>Congratulations!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2413863,
      "author_name": "kongweihao",
      "author_url": "",
      "post_date": "2023-08-29T07:46:38.083000",
      "content": "<p>大神好，想请教3个问题:<br>\n1、如果mask向右下偏移了0.5个像素，那水平翻转和垂直翻转的时候是否要把mask分别向右和向下偏移1个像素呢？<br>\n2、在您的实验中，这个偏移0.5像素的发现对分数的改善有多大<br>\n3、您笔记里的5个backbone是如何确定的，是用了某个类似automl自动搜参工具的吗，还是凭借丰富的经验。我是新手，同时也对自动搜参很感兴趣😅您觉得自动搜参帮助大吗以及能否给菜鸟们推荐一款（如果您用过的话）<br>\n很抱歉问的几个新手问题，若能回复，将不胜感激。</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2417947,
          "author_name": "tascj",
          "author_url": "",
          "post_date": "2023-09-01T02:59:14.760000",
          "content": "<blockquote>\n  <p>1、如果mask向右下偏移了0.5个像素，那水平翻转和垂直翻转的时候是否要把mask分别向右和向下偏移1个像素呢？</p>\n</blockquote>\n<p>I believe the misalignment is not caused by some 0.5 pixel shift, it is caused by the 1-extra-pixel on the bottom-right. A shift of 0.5 pixels can bring back symmetry.</p>\n<p>Without symmetry, a naive implementation of flip/rot90 augmentation would introduce inconsistency, which confuses the model in training.</p>\n<p>As illustrated in the main post, there are two solutions:</p>\n<ol>\n<li>Keep the misalignment (keep the asymmetry), and address the inconsistency.</li>\n<li>Fix the misalignment (bring back symmetry)</li>\n</ol>\n<blockquote>\n  <p>2、在您的实验中，这个偏移0.5像素的发现对分数的改善有多大</p>\n</blockquote>\n<p>I believe the shift itself dost not help. The gain is mainly from flip/rot90 augmentation and TTA. I suggest to try it yourself.</p>\n<blockquote>\n  <p>3、您笔记里的5个backbone是如何确定的</p>\n</blockquote>\n<p>I joined the competition late (6 days before the deadline). In the final stage I simply picked several heavy backbones that could finish training before the deadline. I did not explore much here.</p>\n<p>I rarely adjust training parameters other than the learning rate. I don't have much experience with automated parameter search. The benefits of tuning parameters might not be significant after ensembling, in my opinion, the cost-effectiveness might not be high.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2386823,
      "author_name": "Muhammad Usman",
      "author_url": "",
      "post_date": "2023-08-12T07:09:26.487000",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> on your solo gold. It's a great achievement.</p>\n<p>Your solution is really helpful to learn from.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2386410,
      "author_name": "Yousef Rabi",
      "author_url": "",
      "post_date": "2023-08-11T23:21:11.030000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/tascj0\" target=\"_blank\">@tascj0</a> !</p>\n<p>I have a question regarding the affine matrix.</p>\n<p>If the pixels were shifted by 0.5 pixel in the original image space and you upscale the image by 2, shouldn't the shift now be 1 pixel? Where does the 1.5 pixel shift come from?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2386426,
          "author_name": "tascj",
          "author_url": "",
          "post_date": "2023-08-11T23:53:22.380000",
          "content": "<p>It's about the OpenCV coordinate system.</p>\n<p>In short, the rule is <code>dst+0.5 = M(src+0.5)</code>. After calibration, the <code>1</code> becomes <code>1.5</code>.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2383117,
      "author_name": "iadduk",
      "author_url": "",
      "post_date": "2023-08-10T07:28:11.793000",
      "content": "<p>Thanks so much for the writeup.<br>\nAmazing single submission result! Takes guts to stick to your own CV.<br>\nDo you have any plans on releasing your training code?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2413850,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-08-29T07:36:26.743000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2413846,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-08-29T07:34:13.710000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2382704": "Thanks to the organizers for hosting competition and congrats to all the winners.\n\n\n# The key finding\n\nMany people have likely noticed that flip/rot90 augmentation doesn't work well on this dataset.\n\nMy guess is that during the conversion of polygon annotations to binary masks, the leftmost point is not included while the rightmost point is, causing a misalignment between the image and the mask.\n\n* original\n   * img:    [0, 128,   128,    3,   4, 5, 6]\n   * mask: [0, 255a, 255b, 255pad, 0, 0, 0]\n* normal flip that does not work\n   * img: [6, 5, 4, 3,      128,  128,    0]\n   * mask: [0, 0, 0, 255pad, 255b, 255a, 0]\n* consistent flip should be\n    * img: [6, 5, 4, 3, 128,  128,    0]\n    * mask: [0, 0, 0, 0, 255b, 255a, 255pad]\n\nI tried two solutions to address this issue.\n\n### solution 1, train with misaligned (img ,mask)\n\n* training\n  * img: normal flip\n  * mask: consistent flip\n* inference\n  * flip tta: normal img flip -> predict -> consistent mask flip\n\ntraining/test time augmentation is a bit tricky in this case.\n\n### solution 2, train with aligned (img, mask)\n\n* training\n  * shift img by +0.5 pixel\n* inference\n  * shift img by +0.5 pixel\n\ntraining/test time augmentation is as normal in this case.\n\nBoth solutions worked. For simplicity, I chose solution 2. I use the following code to apply resize512&shift1\n\n```\nimg_affine_matrix = np.array([[2.0, 0.0, 1.5], [0.0, 2.0, 1.5]], dtype=np.float64)\nimg = cv2.warpAffine(\n    img,\n    img_affine_matrix,\n    (512, 512),\n    flags=cv2.INTER_LINEAR,\n    borderMode=cv2.BORDER_CONSTANT,\n    borderValue=0,\n)\n```\n\nNote that you need to calibrate M for warpAffine. So it's 1.5 in the final affine matrix.\n```\ndef calibrate(M):\n    # dst + 0.5 = M(src + 0.5)\n    M[:, 2] += M[:, 0] * 0.5 + M[:, 1] * 0.5 - 0.5\n    return M\n```\n\n# Models\n\nMy submission notebook is pubic [here](https://www.kaggle.com/code/tascj0/contrail-submit?scriptVersionId=139432132).\n\nI joined the competition quite late and only briefly explored other bands and 2.5D before giving up on them. In the end, I only used false_color images and UNet models.\n\n* Strong backbone were very helpful.\n* train with all individual annotations is helpful.\n\nI had some success with pseudo-labeling on a small model, but unfortunately, I didn't have time to apply it on larger models.\n",
    "2382723": "just to put into picture on the annotation issues of flipping\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fe61b0795055ea79c56f425587a6f5b8c%2FSelection_999(2868).png?generation=1691631177229301&alt=media)",
    "2383609": "Congratz on the solo submission gold, very impressive.\n\nWe completely missed the issues with masks, great catch. I noticed very early that flips were hurting results and assumed that was because the background was similar between images in train / val. I can't believe we still got 2nd place while overlooking that.\nI guess our models learnt the pattern, but still, it most likely hurt performances.",
    "2382709": "Very impressive given the number of subs and the time you had. Congratulations",
    "2383289": "Impressive 👏👏👏👏\nCongratulations!",
    "2383098": "Thanks for sharing @tascj0 , and congratulations on your solo gold with one sub.\n\n> My guess is that during the conversion of polygon annotations to binary masks, the leftmost point is not included while the rightmost point is, causing a misalignment between the image and the mask.\n\nWhy it's not \"the leftmost point is included while the rightmost point isn't\"? How did you detect this? Thanks",
    "2383040": "@tascj0 Congratulations for such results with a single efficient submission!",
    "2382898": "Thanks @tascj0 for sharing your work, which really helpful in our learning process.",
    "2382832": "Congratulations on winning top position in the Competition.\nThanks for sharing your notebook and useful inputs. ",
    "2382728": "Can you share your validation strategy in this competition please? Also how did you do pseudo-labeling and did you used it in final ensemble for some models? As for training on the individual annotation: did you train each model on different annotation separately? I didn't understand this part actually. ",
    "2382708": "Congratulations!",
    "2413863": "大神好，想请教3个问题:\n1、如果mask向右下偏移了0.5个像素，那水平翻转和垂直翻转的时候是否要把mask分别向右和向下偏移1个像素呢？\n2、在您的实验中，这个偏移0.5像素的发现对分数的改善有多大\n3、您笔记里的5个backbone是如何确定的，是用了某个类似automl自动搜参工具的吗，还是凭借丰富的经验。我是新手，同时也对自动搜参很感兴趣😅您觉得自动搜参帮助大吗以及能否给菜鸟们推荐一款（如果您用过的话）\n很抱歉问的几个新手问题，若能回复，将不胜感激。",
    "2386823": "Congrats @hengck23 on your solo gold. It's a great achievement.\n\nYour solution is really helpful to learn from.",
    "2386410": "Congratulations @tascj0 !\n\nI have a question regarding the affine matrix.\n\nIf the pixels were shifted by 0.5 pixel in the original image space and you upscale the image by 2, shouldn't the shift now be 1 pixel? Where does the 1.5 pixel shift come from?",
    "2383117": "Thanks so much for the writeup.\nAmazing single submission result! Takes guts to stick to your own CV.\nDo you have any plans on releasing your training code?",
    "2413850": "",
    "2413846": ""
  }
}