{
  "id": 430471,
  "title": "Notes about what worked and didnt",
  "url": "/competitions/google-research-identify-contrails-reduce-global-warming/discussion/430471",
  "author_name": "ryches",
  "post_date": "2023-08-09T23:36:15.196000",
  "votes": 41,
  "comment_count": 15,
  "views": 0,
  "content": "<p>I tried many different things and most of them didn't work. I wanted to summarize some of the things I figured out. What mattered and what made no difference. Some of it was kind of surprising. </p>\n<h2>Things that mattered</h2>\n<p><strong>Resolution</strong><br>\n512 was significantly better than 256. Very unintuitive given that it is operating on exactly the same data, just upscaled. The argument in the paper is that the model is using additional capacity technically making multiple predictions per pixel, but I am not completely convinced this is the mechanism of improved performance. </p>\n<p><strong>Interpolation</strong> <br>\nI think something a lot of people probably overlooked is that even just how you interpolate and upscale and downscale images makes a big difference in the context of this competition. Bilinear, bicubic, nearest etc can move your predictions significantly because the granularity on these predictions is so fine. Just the slightest bit of blurring can greatly hurt performance. I found that simply upscaling with bicubic or lanczos and downscaling with nearest method yielded significant gains. You can even try upscaling existing 256 predictions to 512 with lanczos and downscale again to 256 with nearest and you will likely see a significant boost in performance</p>\n<p>Incorrect pixels<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F87c08540388bfc5bbb1ec9c9bbd8837d%2FScreen%20Shot%202023-07-26%20at%2010.56.51%20PM.png?generation=1691622203802155&amp;alt=media\" alt=\"\"></p>\n<p><strong>Backbone</strong> <br>\nThis is a point I kind of regret. In the previous competition, I did so well with segformer that I stuck with that for a long time. It beat most of the other model configurations, unet, upernet, etc. But eventually scaling up to the larger backbones and finding the right setup could make a massive swing. From 0.62 with resnet26d vs .69 with convnext xl. </p>\n<p><strong>Using soft labels</strong><br>\nIn this competition, we were given the annotations from individuals as well as the aggregated labeled with their agreed upon labels. You can try to train against the final labels, but I found that training against the mean of the individual annotations was significantly better than training against the hard aggregations. These aggregated labels ended up with some very unnatural shapes not indicative of a contrail line like we expect. Doing the mean of individual annotations gave the model much better training signal showing which pixels were highly agreed upon and which ones were in contention without labeling them with hard 0's. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F97a8746635d566bb833df3c735a19acd%2FScreen%20Shot%202023-08-09%20at%203.50.38%20PM.png?generation=1691621461497517&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F94f1fe5fda377c546302f6f50d995f84%2FScreen%20Shot%202023-08-09%20at%203.50.46%20PM.png?generation=1691621475570972&amp;alt=media\" alt=\"\"></p>\n<p><strong>TTA</strong><br>\nThis didn't have a massive impact on performance but it was a very predictable improvement I got from doing flips and rotations at inference time to leaderboard performance. The model was trained against flips and rotations so some capacity is allocated to understanding these variations so it makes sense averaging between multiple representations would help some amount</p>\n<h2>Things that didn't work</h2>\n<p><strong>Pseudolabeling</strong><br>\nI tried various different flavors of pseudolabeling. My favorite was training a model on the 5th timestep and then making predictions against all of the other timesteps and then retraining on this dataset that had been expanded to all of the frames. I still dont know if there is some variant of this that would work but I couldnt make any gains from this. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Fd69c7d83568e864766beffd5ff9c14bb%2Fezgif.com-gif-maker%20(13).gif?generation=1691621924235769&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F6ad7122fe1602261bff07c580a688091%2Fezgif.com-gif-maker%20(14).gif?generation=1691621935864261&amp;alt=media\" alt=\"\"></p>\n<p><strong>More channels</strong><br>\nI tried a lot of experiments training against all channels and I always found it really surprising that I couldnt get any extra signal from any of them. I guess on a certain level it makes sense that the annotators only saw a certain representation so the labels would be limited by what could be seen in that representation. I showed pretty conclusively that the most predictive band on its own was band_16 and when I added it in it made the model converge more slowly which was surprising. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F68608469501cf449e44b39dedc59e248%2FScreen%20Shot%202023-08-09%20at%204.00.10%20PM.png?generation=1691622080639442&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F6c730d8dde0dbaa625c1df389833df74%2FScreen%20Shot%202023-08-09%20at%204.01.00%20PM.png?generation=1691622094488501&amp;alt=media\" alt=\"\"></p>\n<p><strong>More timesteps</strong><br>\nI tried various experiments trying to make models that had additional timesteps worth of information but it just seemed to muddle the signal. Maybe with enough training,I tried many different things and most of them didn't work. Wanted to summarize some of the things I figured out. What mattered and what made no difference. Some of it was kind of surprising. </p>\n<p><strong>Geographic sampling</strong><br>\nOne thing I realized was that we actually had a very uneven geographic distribution. Some lat longs we had hundreds of samples and others we might only have one or two. I tried to understand this relationship. I wanted to check if the model was simply learning the area and if it was likely to have a contrail or not and not generalizing as well. I tried training a model that had all positives and only one negative from each geo-location and found that it actually reached a similarly accurate solution while dropping thousands of images. I did not pursue this very far but it was interesting to see that you could ditch so much data and get a solution within .01 of training against all data. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F2903729af472eca94b02fff508646e1d%2FScreen%20Shot%202023-08-09%20at%203.47.49%20PM.png?generation=1691621289592984&amp;alt=media\" alt=\"\"></p>\n<h2>Ideas I wish I explored</h2>\n<p><strong>Some variant on tracking</strong><br>\nOne idea I really wish I had spent some more time pursuing was using something to predict per timestep and then do some form of tracking across time. Maybe you could make a model that predicts per frame and then train a secondary model that summarizes this to the frame of interest. Hard to say if this would be useful but I have to assume there is some value to the change over time.</p>\n<p><strong>Adjacent inference</strong><br>\nI plotted my aggregated error from all samples and it seemed like there was a pattern of my models being more incorrect along the borders. I did not evaluate thoroughly but I think there might be some way to put adjacent tiles and timesteps together so that you can do inference with them put together. It seems like the model struggled when you got just the edge of something coming into frame so maybe there is some way to mitigate that by giving it broader context. </p>\n<p>Error averaged across the whole validation set<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Ffe8b75ec045af4af74d15db59d2fa373%2FScreen%20Shot%202023-07-25%20at%202.28.53%20PM.png?generation=1691621231186415&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 2382655,
      "postDate": "2023-08-09T23:36:15.197Z",
      "content": "<p>I tried many different things and most of them didn't work. I wanted to summarize some of the things I figured out. What mattered and what made no difference. Some of it was kind of surprising. </p>\n<h2>Things that mattered</h2>\n<p><strong>Resolution</strong><br>\n512 was significantly better than 256. Very unintuitive given that it is operating on exactly the same data, just upscaled. The argument in the paper is that the model is using additional capacity technically making multiple predictions per pixel, but I am not completely convinced this is the mechanism of improved performance. </p>\n<p><strong>Interpolation</strong> <br>\nI think something a lot of people probably overlooked is that even just how you interpolate and upscale and downscale images makes a big difference in the context of this competition. Bilinear, bicubic, nearest etc can move your predictions significantly because the granularity on these predictions is so fine. Just the slightest bit of blurring can greatly hurt performance. I found that simply upscaling with bicubic or lanczos and downscaling with nearest method yielded significant gains. You can even try upscaling existing 256 predictions to 512 with lanczos and downscale again to 256 with nearest and you will likely see a significant boost in performance</p>\n<p>Incorrect pixels<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F87c08540388bfc5bbb1ec9c9bbd8837d%2FScreen%20Shot%202023-07-26%20at%2010.56.51%20PM.png?generation=1691622203802155&amp;alt=media\" alt=\"\"></p>\n<p><strong>Backbone</strong> <br>\nThis is a point I kind of regret. In the previous competition, I did so well with segformer that I stuck with that for a long time. It beat most of the other model configurations, unet, upernet, etc. But eventually scaling up to the larger backbones and finding the right setup could make a massive swing. From 0.62 with resnet26d vs .69 with convnext xl. </p>\n<p><strong>Using soft labels</strong><br>\nIn this competition, we were given the annotations from individuals as well as the aggregated labeled with their agreed upon labels. You can try to train against the final labels, but I found that training against the mean of the individual annotations was significantly better than training against the hard aggregations. These aggregated labels ended up with some very unnatural shapes not indicative of a contrail line like we expect. Doing the mean of individual annotations gave the model much better training signal showing which pixels were highly agreed upon and which ones were in contention without labeling them with hard 0's. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F97a8746635d566bb833df3c735a19acd%2FScreen%20Shot%202023-08-09%20at%203.50.38%20PM.png?generation=1691621461497517&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F94f1fe5fda377c546302f6f50d995f84%2FScreen%20Shot%202023-08-09%20at%203.50.46%20PM.png?generation=1691621475570972&amp;alt=media\" alt=\"\"></p>\n<p><strong>TTA</strong><br>\nThis didn't have a massive impact on performance but it was a very predictable improvement I got from doing flips and rotations at inference time to leaderboard performance. The model was trained against flips and rotations so some capacity is allocated to understanding these variations so it makes sense averaging between multiple representations would help some amount</p>\n<h2>Things that didn't work</h2>\n<p><strong>Pseudolabeling</strong><br>\nI tried various different flavors of pseudolabeling. My favorite was training a model on the 5th timestep and then making predictions against all of the other timesteps and then retraining on this dataset that had been expanded to all of the frames. I still dont know if there is some variant of this that would work but I couldnt make any gains from this. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Fd69c7d83568e864766beffd5ff9c14bb%2Fezgif.com-gif-maker%20(13).gif?generation=1691621924235769&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F6ad7122fe1602261bff07c580a688091%2Fezgif.com-gif-maker%20(14).gif?generation=1691621935864261&amp;alt=media\" alt=\"\"></p>\n<p><strong>More channels</strong><br>\nI tried a lot of experiments training against all channels and I always found it really surprising that I couldnt get any extra signal from any of them. I guess on a certain level it makes sense that the annotators only saw a certain representation so the labels would be limited by what could be seen in that representation. I showed pretty conclusively that the most predictive band on its own was band_16 and when I added it in it made the model converge more slowly which was surprising. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F68608469501cf449e44b39dedc59e248%2FScreen%20Shot%202023-08-09%20at%204.00.10%20PM.png?generation=1691622080639442&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F6c730d8dde0dbaa625c1df389833df74%2FScreen%20Shot%202023-08-09%20at%204.01.00%20PM.png?generation=1691622094488501&amp;alt=media\" alt=\"\"></p>\n<p><strong>More timesteps</strong><br>\nI tried various experiments trying to make models that had additional timesteps worth of information but it just seemed to muddle the signal. Maybe with enough training,I tried many different things and most of them didn't work. Wanted to summarize some of the things I figured out. What mattered and what made no difference. Some of it was kind of surprising. </p>\n<p><strong>Geographic sampling</strong><br>\nOne thing I realized was that we actually had a very uneven geographic distribution. Some lat longs we had hundreds of samples and others we might only have one or two. I tried to understand this relationship. I wanted to check if the model was simply learning the area and if it was likely to have a contrail or not and not generalizing as well. I tried training a model that had all positives and only one negative from each geo-location and found that it actually reached a similarly accurate solution while dropping thousands of images. I did not pursue this very far but it was interesting to see that you could ditch so much data and get a solution within .01 of training against all data. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F2903729af472eca94b02fff508646e1d%2FScreen%20Shot%202023-08-09%20at%203.47.49%20PM.png?generation=1691621289592984&amp;alt=media\" alt=\"\"></p>\n<h2>Ideas I wish I explored</h2>\n<p><strong>Some variant on tracking</strong><br>\nOne idea I really wish I had spent some more time pursuing was using something to predict per timestep and then do some form of tracking across time. Maybe you could make a model that predicts per frame and then train a secondary model that summarizes this to the frame of interest. Hard to say if this would be useful but I have to assume there is some value to the change over time.</p>\n<p><strong>Adjacent inference</strong><br>\nI plotted my aggregated error from all samples and it seemed like there was a pattern of my models being more incorrect along the borders. I did not evaluate thoroughly but I think there might be some way to put adjacent tiles and timesteps together so that you can do inference with them put together. It seems like the model struggled when you got just the edge of something coming into frame so maybe there is some way to mitigate that by giving it broader context. </p>\n<p>Error averaged across the whole validation set<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Ffe8b75ec045af4af74d15db59d2fa373%2FScreen%20Shot%202023-07-25%20at%202.28.53%20PM.png?generation=1691621231186415&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I tried many different things and most of them didn't work. I wanted to summarize some of the things I figured out. What mattered and what made no difference. Some of it was kind of surprising. \n\n##Things that mattered \n**Resolution**\n512 was significantly better than 256. Very unintuitive given that it is operating on exactly the same data, just upscaled. The argument in the paper is that the model is using additional capacity technically making multiple predictions per pixel, but I am not completely convinced this is the mechanism of improved performance. \n\n**Interpolation** \nI think something a lot of people probably overlooked is that even just how you interpolate and upscale and downscale images makes a big difference in the context of this competition. Bilinear, bicubic, nearest etc can move your predictions significantly because the granularity on these predictions is so fine. Just the slightest bit of blurring can greatly hurt performance. I found that simply upscaling with bicubic or lanczos and downscaling with nearest method yielded significant gains. You can even try upscaling existing 256 predictions to 512 with lanczos and downscale again to 256 with nearest and you will likely see a significant boost in performance\n\nIncorrect pixels\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F87c08540388bfc5bbb1ec9c9bbd8837d%2FScreen%20Shot%202023-07-26%20at%2010.56.51%20PM.png?generation=1691622203802155&alt=media)\n\n\n**Backbone** \nThis is a point I kind of regret. In the previous competition, I did so well with segformer that I stuck with that for a long time. It beat most of the other model configurations, unet, upernet, etc. But eventually scaling up to the larger backbones and finding the right setup could make a massive swing. From 0.62 with resnet26d vs .69 with convnext xl. \n\n**Using soft labels**\nIn this competition, we were given the annotations from individuals as well as the aggregated labeled with their agreed upon labels. You can try to train against the final labels, but I found that training against the mean of the individual annotations was significantly better than training against the hard aggregations. These aggregated labels ended up with some very unnatural shapes not indicative of a contrail line like we expect. Doing the mean of individual annotations gave the model much better training signal showing which pixels were highly agreed upon and which ones were in contention without labeling them with hard 0's. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F97a8746635d566bb833df3c735a19acd%2FScreen%20Shot%202023-08-09%20at%203.50.38%20PM.png?generation=1691621461497517&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F94f1fe5fda377c546302f6f50d995f84%2FScreen%20Shot%202023-08-09%20at%203.50.46%20PM.png?generation=1691621475570972&alt=media)\n\n**TTA**\nThis didn't have a massive impact on performance but it was a very predictable improvement I got from doing flips and rotations at inference time to leaderboard performance. The model was trained against flips and rotations so some capacity is allocated to understanding these variations so it makes sense averaging between multiple representations would help some amount\n\n## Things that didn't work\n**Pseudolabeling**\nI tried various different flavors of pseudolabeling. My favorite was training a model on the 5th timestep and then making predictions against all of the other timesteps and then retraining on this dataset that had been expanded to all of the frames. I still dont know if there is some variant of this that would work but I couldnt make any gains from this. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Fd69c7d83568e864766beffd5ff9c14bb%2Fezgif.com-gif-maker%20(13).gif?generation=1691621924235769&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F6ad7122fe1602261bff07c580a688091%2Fezgif.com-gif-maker%20(14).gif?generation=1691621935864261&alt=media)\n\n**More channels**\nI tried a lot of experiments training against all channels and I always found it really surprising that I couldnt get any extra signal from any of them. I guess on a certain level it makes sense that the annotators only saw a certain representation so the labels would be limited by what could be seen in that representation. I showed pretty conclusively that the most predictive band on its own was band_16 and when I added it in it made the model converge more slowly which was surprising. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F68608469501cf449e44b39dedc59e248%2FScreen%20Shot%202023-08-09%20at%204.00.10%20PM.png?generation=1691622080639442&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F6c730d8dde0dbaa625c1df389833df74%2FScreen%20Shot%202023-08-09%20at%204.01.00%20PM.png?generation=1691622094488501&alt=media)\n\n**More timesteps**\nI tried various experiments trying to make models that had additional timesteps worth of information but it just seemed to muddle the signal. Maybe with enough training,I tried many different things and most of them didn't work. Wanted to summarize some of the things I figured out. What mattered and what made no difference. Some of it was kind of surprising. \n\n**Geographic sampling**\nOne thing I realized was that we actually had a very uneven geographic distribution. Some lat longs we had hundreds of samples and others we might only have one or two. I tried to understand this relationship. I wanted to check if the model was simply learning the area and if it was likely to have a contrail or not and not generalizing as well. I tried training a model that had all positives and only one negative from each geo-location and found that it actually reached a similarly accurate solution while dropping thousands of images. I did not pursue this very far but it was interesting to see that you could ditch so much data and get a solution within .01 of training against all data. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F2903729af472eca94b02fff508646e1d%2FScreen%20Shot%202023-08-09%20at%203.47.49%20PM.png?generation=1691621289592984&alt=media)\n\n## Ideas I wish I explored\n**Some variant on tracking**\nOne idea I really wish I had spent some more time pursuing was using something to predict per timestep and then do some form of tracking across time. Maybe you could make a model that predicts per frame and then train a secondary model that summarizes this to the frame of interest. Hard to say if this would be useful but I have to assume there is some value to the change over time.\n\n**Adjacent inference**\nI plotted my aggregated error from all samples and it seemed like there was a pattern of my models being more incorrect along the borders. I did not evaluate thoroughly but I think there might be some way to put adjacent tiles and timesteps together so that you can do inference with them put together. It seems like the model struggled when you got just the edge of something coming into frame so maybe there is some way to mitigate that by giving it broader context. \n\nError averaged across the whole validation set\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Ffe8b75ec045af4af74d15db59d2fa373%2FScreen%20Shot%202023-07-25%20at%202.28.53%20PM.png?generation=1691621231186415&alt=media)",
      "votes": 40
    },
    {
      "id": 2382700,
      "postDate": "2023-08-10T01:02:06.027Z",
      "content": "<p>thanks for the writeup.</p>\n<p>the stack 3x3 conv of the D version (deep stem) of the resnet helps a help. This is mentioned in the datset paper and vertified in my experiments with se-resnext2d. I can get pr-auc=0.699 on validation split with  se-resnext2d + 512x512</p>\n<hr>\n<p>I search through all vision VIT transformers and forunately, i found one that use D-stem.</p>\n<p>1) <a href=\"https://github.com/bytedance/Next-ViT\" target=\"_blank\">https://github.com/bytedance/Next-ViT</a><br>\nThis has very good results. using the smallest version (83.6at top-1 image net) + 512x512, it has pr-auc=0.717 on validation split</p>\n<p>2) <a href=\"https://github.com/NVlabs/FasterViT\" target=\"_blank\">https://github.com/NVlabs/FasterViT</a><br>\ni haven't try this yet becuase of Nvida license (i will make post submission and write up later)</p>\n<p>both (1), (2) are hybrid CNN-VIT model (it uses CNN at inital stages to capture fine structure like lines, and VIT transformer at later stages for global context … i.e. long, long lines)</p>\n<p>I find them better than segformer mix-vision-transformer MIT.</p>",
      "rawMarkdown": "thanks for the writeup.\n\nthe stack 3x3 conv of the D version (deep stem) of the resnet helps a help. This is mentioned in the datset paper and vertified in my experiments with se-resnext2d. I can get pr-auc=0.699 on validation split with  se-resnext2d + 512x512\n\n---\n\nI search through all vision VIT transformers and forunately, i found one that use D-stem.\n\n1) https://github.com/bytedance/Next-ViT\nThis has very good results. using the smallest version (83.6at top-1 image net) + 512x512, it has pr-auc=0.717 on validation split\n\n2) https://github.com/NVlabs/FasterViT\ni haven't try this yet becuase of Nvida license (i will make post submission and write up later)\n\nboth (1), (2) are hybrid CNN-VIT model (it uses CNN at inital stages to capture fine structure like lines, and VIT transformer at later stages for global context ... i.e. long, long lines)\n\nI find them better than segformer mix-vision-transformer MIT.\n\n ",
      "votes": 3
    },
    {
      "id": 2382692,
      "postDate": "2023-08-10T00:44:13.560Z",
      "content": "<p>congrats and thanks for the detailed writeup . I saw that u did basic tta but were the same also used as augmentation for your training or different transforms were used there ?</p>",
      "rawMarkdown": "congrats and thanks for the detailed writeup . I saw that u did basic tta but were the same also used as augmentation for your training or different transforms were used there ?",
      "votes": 1,
      "replies": [
        {
          "id": 2382695,
          "postDate": "2023-08-10T00:46:13.203Z",
          "content": "<p>The same simple augmentations were used for training and inference. 90 degree rotates and flips</p>",
          "rawMarkdown": "The same simple augmentations were used for training and inference. 90 degree rotates and flips",
          "votes": 1
        }
      ]
    },
    {
      "id": 2382675,
      "postDate": "2023-08-10T00:28:16.060Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a>! Out of curiosity, did you only apply bicubic/Lanczos resampling at the prediction stage? Or did you also use this method to resize the inputs from 256x256 to 512x512 before feeding them through your models?</p>",
      "rawMarkdown": "Congratulations @ryches! Out of curiosity, did you only apply bicubic/Lanczos resampling at the prediction stage? Or did you also use this method to resize the inputs from 256x256 to 512x512 before feeding them through your models?",
      "votes": 1,
      "replies": [
        {
          "id": 2382676,
          "postDate": "2023-08-10T00:31:37.487Z",
          "content": "<p>My training and inference pipeline was to upscale input to 512 with bicubic interpolation. I made my models output directly at 256. It seemed better to me to have the model directly optimize in the prediction size we wanted instead of having it output something and blindly interpolate after. </p>",
          "rawMarkdown": "My training and inference pipeline was to upscale input to 512 with bicubic interpolation. I made my models output directly at 256. It seemed better to me to have the model directly optimize in the prediction size we wanted instead of having it output something and blindly interpolate after. ",
          "votes": 5,
          "replies": [
            {
              "id": 2382690,
              "postDate": "2023-08-10T00:43:35.647Z",
              "content": "<p>I see, thanks for the response! I wish I had experimented more with different resampling methods, it seems like it could have been quite helpful.</p>",
              "rawMarkdown": "I see, thanks for the response! I wish I had experimented more with different resampling methods, it seems like it could have been quite helpful.",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2382673,
      "postDate": "2023-08-10T00:19:59.423Z",
      "content": "<p>Indeed, congrats on Competition Grandmaster.  Great writeup, and thanks for including what didn't work.</p>",
      "rawMarkdown": "Indeed, congrats on Competition Grandmaster.  Great writeup, and thanks for including what didn't work.",
      "votes": 1
    },
    {
      "id": 2382681,
      "postDate": "2023-08-10T00:38:26Z",
      "content": "<p>congrats on solo gold and GM! amazing performance</p>\n<p>May I ask what was your final ensemble CV? and  if your final ensemble is based only on Segformers? </p>\n<p>ps: regarding the pseudo-labeling, a similar strategy with what you described but with additional training tricks gave a big boost in our team exp - will say more in our solution write up - but there is definitely signal in the other time slices [EDIT]</p>",
      "rawMarkdown": "congrats on solo gold and GM! amazing performance\n\nMay I ask what was your final ensemble CV? and  if your final ensemble is based only on Segformers? \n\nps: regarding the pseudo-labeling, a similar strategy with what you described but with additional training tricks gave a big boost in our team exp - will say more in our solution write up - but there is definitely signal in the other time slices [EDIT]",
      "votes": 2,
      "replies": [
        {
          "id": 2382694,
          "postDate": "2023-08-10T00:45:21.600Z",
          "content": "<p>Thank you. My final ensemble had no segformers actually. I held onto those for way too long and then eventually tried out convnext and effnetv2 and had so much better performance I ditched the segformers entirely.</p>\n<p>Interested to see what I missed on the pseudolabeling. I was positive there was something in there but couldn't quite figure it out.</p>\n<p>Final CV is a bit hard for me to state. I mostly used the validation folder but I also did full retraining with the validation folder included in training. My best individual convnext model was able to get 0.695. Effnetv2 more like 0.685. Ensembled and with my little resizing trick I was probably in the 0.71 ballpark where I landed on the leaderboard</p>",
          "rawMarkdown": "Thank you. My final ensemble had no segformers actually. I held onto those for way too long and then eventually tried out convnext and effnetv2 and had so much better performance I ditched the segformers entirely.\n\nInterested to see what I missed on the pseudolabeling. I was positive there was something in there but couldn't quite figure it out.\n\nFinal CV is a bit hard for me to state. I mostly used the validation folder but I also did full retraining with the validation folder included in training. My best individual convnext model was able to get 0.695. Effnetv2 more like 0.685. Ensembled and with my little resizing trick I was probably in the 0.71 ballpark where I landed on the leaderboard",
          "votes": 1
        }
      ]
    },
    {
      "id": 2382671,
      "postDate": "2023-08-10T00:12:48.597Z",
      "content": "<p>Thanks for the write-up, and congrats on becoming a 2x GM!</p>\n<p>How do you train w/ soft labels? Do you make each pixel a float rather than 0 or 1?</p>",
      "rawMarkdown": "Thanks for the write-up, and congrats on becoming a 2x GM!\n\nHow do you train w/ soft labels? Do you make each pixel a float rather than 0 or 1?",
      "votes": 2,
      "replies": [
        {
          "id": 2382674,
          "postDate": "2023-08-10T00:20:04.007Z",
          "content": "<p>Yes exactly. BCE works perfectly fine with float values</p>",
          "rawMarkdown": "Yes exactly. BCE works perfectly fine with float values",
          "votes": 4
        }
      ]
    },
    {
      "id": 2382662,
      "postDate": "2023-08-10T00:03:25.977Z",
      "content": "<p>Congratulations on the solo gold and becoming a Kaggle Competition GM!! Well deserved :) </p>",
      "rawMarkdown": "Congratulations on the solo gold and becoming a Kaggle Competition GM!! Well deserved :) ",
      "votes": 2
    },
    {
      "id": 2409704,
      "postDate": "2023-08-26T11:40:10.347Z",
      "content": "<p>Congratulations for your GM title and thanks for sharing! I have a couple of questions that I couldn't ask during the Discord session (thanks for that as well):</p>\n<ul>\n<li>How did you come up the idea of upscaling the 256 predictions with Lanczos and downscaling again with Nearest Neighbors? This is a very weird/neat trick!</li>\n</ul>\n<blockquote>\n  <p>I wanted to check if the model was simply learning the area and if it was likely to have a contrail or not and not generalizing as well. I tried training a model that had all positives and only one negative from each geo-location and found that it actually reached a similarly accurate solution while dropping thousands of images.</p>\n</blockquote>\n<p>You mean that you wondered whether the model could be simply dropping the mean probability of contrail or something similar that denotes that the model is learning the area, right? But if you train a model with only positive samples this could not be possible, and since the results are very similar compared with training with all samples, we can deduce that the model trained with all data is not learning the area in any way, is that correct?</p>",
      "rawMarkdown": "Congratulations for your GM title and thanks for sharing! I have a couple of questions that I couldn't ask during the Discord session (thanks for that as well):\n\n- How did you come up the idea of upscaling the 256 predictions with Lanczos and downscaling again with Nearest Neighbors? This is a very weird/neat trick!\n\n> I wanted to check if the model was simply learning the area and if it was likely to have a contrail or not and not generalizing as well. I tried training a model that had all positives and only one negative from each geo-location and found that it actually reached a similarly accurate solution while dropping thousands of images.\n\nYou mean that you wondered whether the model could be simply dropping the mean probability of contrail or something similar that denotes that the model is learning the area, right? But if you train a model with only positive samples this could not be possible, and since the results are very similar compared with training with all samples, we can deduce that the model trained with all data is not learning the area in any way, is that correct?"
    },
    {
      "id": 2386490,
      "postDate": "2023-08-12T01:50:52.023Z",
      "content": "<p>Thanks for sharing the notes!</p>\n<p>Mega Congratulations on reaching Competitions Grandmaster! </p>",
      "rawMarkdown": "Thanks for sharing the notes!\n\nMega Congratulations on reaching Competitions Grandmaster! "
    },
    {
      "id": 2382899,
      "postDate": "2023-08-10T05:09:31.570Z",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a> for your inspirational work for us to learn.</p>",
      "rawMarkdown": "Thanks @ryches for your inspirational work for us to learn."
    }
  ],
  "comments": [
    {
      "id": 2382700,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-08-10T01:02:06.027000",
      "content": "<p>thanks for the writeup.</p>\n<p>the stack 3x3 conv of the D version (deep stem) of the resnet helps a help. This is mentioned in the datset paper and vertified in my experiments with se-resnext2d. I can get pr-auc=0.699 on validation split with  se-resnext2d + 512x512</p>\n<hr>\n<p>I search through all vision VIT transformers and forunately, i found one that use D-stem.</p>\n<p>1) <a href=\"https://github.com/bytedance/Next-ViT\" target=\"_blank\">https://github.com/bytedance/Next-ViT</a><br>\nThis has very good results. using the smallest version (83.6at top-1 image net) + 512x512, it has pr-auc=0.717 on validation split</p>\n<p>2) <a href=\"https://github.com/NVlabs/FasterViT\" target=\"_blank\">https://github.com/NVlabs/FasterViT</a><br>\ni haven't try this yet becuase of Nvida license (i will make post submission and write up later)</p>\n<p>both (1), (2) are hybrid CNN-VIT model (it uses CNN at inital stages to capture fine structure like lines, and VIT transformer at later stages for global context … i.e. long, long lines)</p>\n<p>I find them better than segformer mix-vision-transformer MIT.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2382692,
      "author_name": "Gaurav Rawat",
      "author_url": "",
      "post_date": "2023-08-10T00:44:13.560000",
      "content": "<p>congrats and thanks for the detailed writeup . I saw that u did basic tta but were the same also used as augmentation for your training or different transforms were used there ?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2382695,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2023-08-10T00:46:13.203000",
          "content": "<p>The same simple augmentations were used for training and inference. 90 degree rotates and flips</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2382675,
      "author_name": "Romain Hardy",
      "author_url": "",
      "post_date": "2023-08-10T00:28:16.060000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a>! Out of curiosity, did you only apply bicubic/Lanczos resampling at the prediction stage? Or did you also use this method to resize the inputs from 256x256 to 512x512 before feeding them through your models?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2382676,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2023-08-10T00:31:37.487000",
          "content": "<p>My training and inference pipeline was to upscale input to 512 with bicubic interpolation. I made my models output directly at 256. It seemed better to me to have the model directly optimize in the prediction size we wanted instead of having it output something and blindly interpolate after. </p>",
          "votes": 5,
          "replies": [
            {
              "id": 2382690,
              "author_name": "Romain Hardy",
              "author_url": "",
              "post_date": "2023-08-10T00:43:35.647000",
              "content": "<p>I see, thanks for the response! I wish I had experimented more with different resampling methods, it seems like it could have been quite helpful.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2382673,
      "author_name": "Ted K",
      "author_url": "",
      "post_date": "2023-08-10T00:19:59.423000",
      "content": "<p>Indeed, congrats on Competition Grandmaster.  Great writeup, and thanks for including what didn't work.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2382681,
      "author_name": "Ioannis M",
      "author_url": "",
      "post_date": "2023-08-10T00:38:26",
      "content": "<p>congrats on solo gold and GM! amazing performance</p>\n<p>May I ask what was your final ensemble CV? and  if your final ensemble is based only on Segformers? </p>\n<p>ps: regarding the pseudo-labeling, a similar strategy with what you described but with additional training tricks gave a big boost in our team exp - will say more in our solution write up - but there is definitely signal in the other time slices [EDIT]</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2382694,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2023-08-10T00:45:21.600000",
          "content": "<p>Thank you. My final ensemble had no segformers actually. I held onto those for way too long and then eventually tried out convnext and effnetv2 and had so much better performance I ditched the segformers entirely.</p>\n<p>Interested to see what I missed on the pseudolabeling. I was positive there was something in there but couldn't quite figure it out.</p>\n<p>Final CV is a bit hard for me to state. I mostly used the validation folder but I also did full retraining with the validation folder included in training. My best individual convnext model was able to get 0.695. Effnetv2 more like 0.685. Ensembled and with my little resizing trick I was probably in the 0.71 ballpark where I landed on the leaderboard</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2382671,
      "author_name": "Bartley",
      "author_url": "",
      "post_date": "2023-08-10T00:12:48.597000",
      "content": "<p>Thanks for the write-up, and congrats on becoming a 2x GM!</p>\n<p>How do you train w/ soft labels? Do you make each pixel a float rather than 0 or 1?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2382674,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2023-08-10T00:20:04.007000",
          "content": "<p>Yes exactly. BCE works perfectly fine with float values</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 2382662,
      "author_name": "Trushant Kalyanpur",
      "author_url": "",
      "post_date": "2023-08-10T00:03:25.977000",
      "content": "<p>Congratulations on the solo gold and becoming a Kaggle Competition GM!! Well deserved :) </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2409704,
      "author_name": "delai50",
      "author_url": "",
      "post_date": "2023-08-26T11:40:10.347000",
      "content": "<p>Congratulations for your GM title and thanks for sharing! I have a couple of questions that I couldn't ask during the Discord session (thanks for that as well):</p>\n<ul>\n<li>How did you come up the idea of upscaling the 256 predictions with Lanczos and downscaling again with Nearest Neighbors? This is a very weird/neat trick!</li>\n</ul>\n<blockquote>\n  <p>I wanted to check if the model was simply learning the area and if it was likely to have a contrail or not and not generalizing as well. I tried training a model that had all positives and only one negative from each geo-location and found that it actually reached a similarly accurate solution while dropping thousands of images.</p>\n</blockquote>\n<p>You mean that you wondered whether the model could be simply dropping the mean probability of contrail or something similar that denotes that the model is learning the area, right? But if you train a model with only positive samples this could not be possible, and since the results are very similar compared with training with all samples, we can deduce that the model trained with all data is not learning the area in any way, is that correct?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2386490,
      "author_name": "Sanyam Bhutani",
      "author_url": "",
      "post_date": "2023-08-12T01:50:52.023000",
      "content": "<p>Thanks for sharing the notes!</p>\n<p>Mega Congratulations on reaching Competitions Grandmaster! </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2382899,
      "author_name": "Tariq Mahmood",
      "author_url": "",
      "post_date": "2023-08-10T05:09:31.570000",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a> for your inspirational work for us to learn.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2382655": "I tried many different things and most of them didn't work. I wanted to summarize some of the things I figured out. What mattered and what made no difference. Some of it was kind of surprising. \n\n##Things that mattered \n**Resolution**\n512 was significantly better than 256. Very unintuitive given that it is operating on exactly the same data, just upscaled. The argument in the paper is that the model is using additional capacity technically making multiple predictions per pixel, but I am not completely convinced this is the mechanism of improved performance. \n\n**Interpolation** \nI think something a lot of people probably overlooked is that even just how you interpolate and upscale and downscale images makes a big difference in the context of this competition. Bilinear, bicubic, nearest etc can move your predictions significantly because the granularity on these predictions is so fine. Just the slightest bit of blurring can greatly hurt performance. I found that simply upscaling with bicubic or lanczos and downscaling with nearest method yielded significant gains. You can even try upscaling existing 256 predictions to 512 with lanczos and downscale again to 256 with nearest and you will likely see a significant boost in performance\n\nIncorrect pixels\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F87c08540388bfc5bbb1ec9c9bbd8837d%2FScreen%20Shot%202023-07-26%20at%2010.56.51%20PM.png?generation=1691622203802155&alt=media)\n\n\n**Backbone** \nThis is a point I kind of regret. In the previous competition, I did so well with segformer that I stuck with that for a long time. It beat most of the other model configurations, unet, upernet, etc. But eventually scaling up to the larger backbones and finding the right setup could make a massive swing. From 0.62 with resnet26d vs .69 with convnext xl. \n\n**Using soft labels**\nIn this competition, we were given the annotations from individuals as well as the aggregated labeled with their agreed upon labels. You can try to train against the final labels, but I found that training against the mean of the individual annotations was significantly better than training against the hard aggregations. These aggregated labels ended up with some very unnatural shapes not indicative of a contrail line like we expect. Doing the mean of individual annotations gave the model much better training signal showing which pixels were highly agreed upon and which ones were in contention without labeling them with hard 0's. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F97a8746635d566bb833df3c735a19acd%2FScreen%20Shot%202023-08-09%20at%203.50.38%20PM.png?generation=1691621461497517&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F94f1fe5fda377c546302f6f50d995f84%2FScreen%20Shot%202023-08-09%20at%203.50.46%20PM.png?generation=1691621475570972&alt=media)\n\n**TTA**\nThis didn't have a massive impact on performance but it was a very predictable improvement I got from doing flips and rotations at inference time to leaderboard performance. The model was trained against flips and rotations so some capacity is allocated to understanding these variations so it makes sense averaging between multiple representations would help some amount\n\n## Things that didn't work\n**Pseudolabeling**\nI tried various different flavors of pseudolabeling. My favorite was training a model on the 5th timestep and then making predictions against all of the other timesteps and then retraining on this dataset that had been expanded to all of the frames. I still dont know if there is some variant of this that would work but I couldnt make any gains from this. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Fd69c7d83568e864766beffd5ff9c14bb%2Fezgif.com-gif-maker%20(13).gif?generation=1691621924235769&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F6ad7122fe1602261bff07c580a688091%2Fezgif.com-gif-maker%20(14).gif?generation=1691621935864261&alt=media)\n\n**More channels**\nI tried a lot of experiments training against all channels and I always found it really surprising that I couldnt get any extra signal from any of them. I guess on a certain level it makes sense that the annotators only saw a certain representation so the labels would be limited by what could be seen in that representation. I showed pretty conclusively that the most predictive band on its own was band_16 and when I added it in it made the model converge more slowly which was surprising. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F68608469501cf449e44b39dedc59e248%2FScreen%20Shot%202023-08-09%20at%204.00.10%20PM.png?generation=1691622080639442&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F6c730d8dde0dbaa625c1df389833df74%2FScreen%20Shot%202023-08-09%20at%204.01.00%20PM.png?generation=1691622094488501&alt=media)\n\n**More timesteps**\nI tried various experiments trying to make models that had additional timesteps worth of information but it just seemed to muddle the signal. Maybe with enough training,I tried many different things and most of them didn't work. Wanted to summarize some of the things I figured out. What mattered and what made no difference. Some of it was kind of surprising. \n\n**Geographic sampling**\nOne thing I realized was that we actually had a very uneven geographic distribution. Some lat longs we had hundreds of samples and others we might only have one or two. I tried to understand this relationship. I wanted to check if the model was simply learning the area and if it was likely to have a contrail or not and not generalizing as well. I tried training a model that had all positives and only one negative from each geo-location and found that it actually reached a similarly accurate solution while dropping thousands of images. I did not pursue this very far but it was interesting to see that you could ditch so much data and get a solution within .01 of training against all data. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F2903729af472eca94b02fff508646e1d%2FScreen%20Shot%202023-08-09%20at%203.47.49%20PM.png?generation=1691621289592984&alt=media)\n\n## Ideas I wish I explored\n**Some variant on tracking**\nOne idea I really wish I had spent some more time pursuing was using something to predict per timestep and then do some form of tracking across time. Maybe you could make a model that predicts per frame and then train a secondary model that summarizes this to the frame of interest. Hard to say if this would be useful but I have to assume there is some value to the change over time.\n\n**Adjacent inference**\nI plotted my aggregated error from all samples and it seemed like there was a pattern of my models being more incorrect along the borders. I did not evaluate thoroughly but I think there might be some way to put adjacent tiles and timesteps together so that you can do inference with them put together. It seems like the model struggled when you got just the edge of something coming into frame so maybe there is some way to mitigate that by giving it broader context. \n\nError averaged across the whole validation set\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Ffe8b75ec045af4af74d15db59d2fa373%2FScreen%20Shot%202023-07-25%20at%202.28.53%20PM.png?generation=1691621231186415&alt=media)",
    "2382700": "thanks for the writeup.\n\nthe stack 3x3 conv of the D version (deep stem) of the resnet helps a help. This is mentioned in the datset paper and vertified in my experiments with se-resnext2d. I can get pr-auc=0.699 on validation split with  se-resnext2d + 512x512\n\n---\n\nI search through all vision VIT transformers and forunately, i found one that use D-stem.\n\n1) https://github.com/bytedance/Next-ViT\nThis has very good results. using the smallest version (83.6at top-1 image net) + 512x512, it has pr-auc=0.717 on validation split\n\n2) https://github.com/NVlabs/FasterViT\ni haven't try this yet becuase of Nvida license (i will make post submission and write up later)\n\nboth (1), (2) are hybrid CNN-VIT model (it uses CNN at inital stages to capture fine structure like lines, and VIT transformer at later stages for global context ... i.e. long, long lines)\n\nI find them better than segformer mix-vision-transformer MIT.\n\n ",
    "2382692": "congrats and thanks for the detailed writeup . I saw that u did basic tta but were the same also used as augmentation for your training or different transforms were used there ?",
    "2382675": "Congratulations @ryches! Out of curiosity, did you only apply bicubic/Lanczos resampling at the prediction stage? Or did you also use this method to resize the inputs from 256x256 to 512x512 before feeding them through your models?",
    "2382673": "Indeed, congrats on Competition Grandmaster.  Great writeup, and thanks for including what didn't work.",
    "2382681": "congrats on solo gold and GM! amazing performance\n\nMay I ask what was your final ensemble CV? and  if your final ensemble is based only on Segformers? \n\nps: regarding the pseudo-labeling, a similar strategy with what you described but with additional training tricks gave a big boost in our team exp - will say more in our solution write up - but there is definitely signal in the other time slices [EDIT]",
    "2382671": "Thanks for the write-up, and congrats on becoming a 2x GM!\n\nHow do you train w/ soft labels? Do you make each pixel a float rather than 0 or 1?",
    "2382662": "Congratulations on the solo gold and becoming a Kaggle Competition GM!! Well deserved :) ",
    "2409704": "Congratulations for your GM title and thanks for sharing! I have a couple of questions that I couldn't ask during the Discord session (thanks for that as well):\n\n- How did you come up the idea of upscaling the 256 predictions with Lanczos and downscaling again with Nearest Neighbors? This is a very weird/neat trick!\n\n> I wanted to check if the model was simply learning the area and if it was likely to have a contrail or not and not generalizing as well. I tried training a model that had all positives and only one negative from each geo-location and found that it actually reached a similarly accurate solution while dropping thousands of images.\n\nYou mean that you wondered whether the model could be simply dropping the mean probability of contrail or something similar that denotes that the model is learning the area, right? But if you train a model with only positive samples this could not be possible, and since the results are very similar compared with training with all samples, we can deduce that the model trained with all data is not learning the area in any way, is that correct?",
    "2386490": "Thanks for sharing the notes!\n\nMega Congratulations on reaching Competitions Grandmaster! ",
    "2382899": "Thanks @ryches for your inspirational work for us to learn."
  }
}