{
  "id": 520500,
  "title": "Solutions and approaches",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/520500",
  "author_name": "wasjaip",
  "post_date": "2024-07-16T09:00:24.309000",
  "votes": 2,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Looking forward to publishing solutions. The very stable situation in the leaderboard gives reason to assume that the solutions will be interesting</p>",
  "messages": [
    {
      "id": 2924119,
      "postDate": "2024-07-16T10:45:23.977Z",
      "content": "<p>i hade a simple transformer encoder with input (N, 60, 25), only thing that might be a bit \"sophisticated\" was that i was also passing the input variable wise (N, 25, 60)  with learned embeddings into a second transformer encoder and then do cross attention with the result of the first encoder. I did not have time to study if the extra encoder + cross attention gave any value though.</p>\n<pre><code>     ():\n        \n        src2 = src1.permute(, , )  \n        \n        src1 = self.input_adapter1(src1)  \n        src2 = self.input_adapter2(src2)  \n\n        \n        src1 = self.positional_encoding_1(src1)\n        src2 = self.learned_encoding_2(src2)\n\n        \n        encoded1 = self.encoder1(src1)\n        encoded2 = self.encoder2(src2)\n\n        \n        cross_attended1 = self.cross_attention_1(encoded1, encoded2, encoded2).flatten(start_dim=)\n        cross_attended2 = self.cross_attention_2(encoded2, encoded1, encoded1).flatten(start_dim=)\n\n        \n        combined = torch.cat([cross_attended1, cross_attended2], dim=)\n\n        output = self.output_linear(combined)\n         output\n</code></pre>",
      "rawMarkdown": "i hade a simple transformer encoder with input (N, 60, 25), only thing that might be a bit \"sophisticated\" was that i was also passing the input variable wise (N, 25, 60)  with learned embeddings into a second transformer encoder and then do cross attention with the result of the first encoder. I did not have time to study if the extra encoder + cross attention gave any value though.\n\n```\n    def forward(self, src1):\n        # Reshape and reorder to N, 25, 60\n        src2 = src1.permute(0, 2, 1)  # Change to N, 25, 60\n        # Prepare inputs\n        src1 = self.input_adapter1(src1)  # N, 60, 25 -> N, 60, d_model\n        src2 = self.input_adapter2(src2)  # N, 25, 60 -> N, 25, d_model\n\n        # Add positional encoding\n        src1 = self.positional_encoding_1(src1)\n        src2 = self.learned_encoding_2(src2)\n\n        # Encode with transformers\n        encoded1 = self.encoder1(src1)\n        encoded2 = self.encoder2(src2)\n\n        # Apply cross-attention\n        cross_attended1 = self.cross_attention_1(encoded1, encoded2, encoded2).flatten(start_dim=1)\n        cross_attended2 = self.cross_attention_2(encoded2, encoded1, encoded1).flatten(start_dim=1)\n\n        # Combine outputs from both paths\n        combined = torch.cat([cross_attended1, cross_attended2], dim=1)\n\n        output = self.output_linear(combined)\n        return output\n```",
      "votes": 5,
      "replies": [
        {
          "id": 2924123,
          "postDate": "2024-07-16T10:56:14.390Z",
          "content": "<p>Thanks! Have you considered any other permutations? how much has this feature improved?</p>",
          "rawMarkdown": "Thanks! Have you considered any other permutations? how much has this feature improved?",
          "replies": [
            {
              "id": 2924200,
              "postDate": "2024-07-16T11:42:59.860Z",
              "content": "<p>hey no, i am not sure how the extra thingy with the cross attention adds, i had no time to try without it. What i did try is to have conv1d layers between the transformer encoder layers but i did not manage to make it work better in a few attempts and since there were only a few <br>\ndays left i did not try further</p>",
              "rawMarkdown": "hey no, i am not sure how the extra thingy with the cross attention adds, i had no time to try without it. What i did try is to have conv1d layers between the transformer encoder layers but i did not manage to make it work better in a few attempts and since there were only a few \ndays left i did not try further",
              "votes": 1
            }
          ]
        },
        {
          "id": 2932421,
          "postDate": "2024-07-22T23:29:34.243Z",
          "content": "<p>Neat trick! How did you come up with the idea of using cross-attention? I used a single encoder model, but the score wasn't as good as yours. </p>",
          "rawMarkdown": "Neat trick! How did you come up with the idea of using cross-attention? I used a single encoder model, but the score wasn't as good as yours. ",
          "votes": 1,
          "replies": [
            {
              "id": 2936544,
              "postDate": "2024-07-26T08:51:40.897Z",
              "content": "<p>i just was fixated that i need to do something also variable-wise and then somehow concat the result, well cross_attention seemed a better idea. but maybe you had worse score for some other reason.</p>",
              "rawMarkdown": "i just was fixated that i need to do something also variable-wise and then somehow concat the result, well cross_attention seemed a better idea. but maybe you had worse score for some other reason.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2936456,
      "postDate": "2024-07-26T06:55:19.917Z",
      "content": "<p>I was only 93rd in the competition, but I used a convolutional neural network with a big twist: it also ingested the predictions of a CatBoost decision tree model as additional input features. So the two techniques were effectively in series. (I also tried combining their predictions as various weighted sums, which didn't help, as the CNN was just better.)</p>\n<p>Apart from the actual machine learning, I found the data handling a challenge on the computer available to me, and one notable thing I did was write a class which gave me fast random access to the whole huge training file, using low-level file operations, so I could access it in arbitrary batches if I wanted to. I'd be interested to know what kind of hardware the more successful teams used for this challenge.</p>\n<p>With a quiet weekend at home with Covid, I indulged myself my writing up my adventures in a report. As I am nowhere near the top of the leaderboard, it's all pubic here: <a href=\"https://github.com/charlie-wartnaby/kaggle-leap-atmospheric-physics\" target=\"_blank\">https://github.com/charlie-wartnaby/kaggle-leap-atmospheric-physics</a></p>",
      "rawMarkdown": "I was only 93rd in the competition, but I used a convolutional neural network with a big twist: it also ingested the predictions of a CatBoost decision tree model as additional input features. So the two techniques were effectively in series. (I also tried combining their predictions as various weighted sums, which didn't help, as the CNN was just better.)\n\nApart from the actual machine learning, I found the data handling a challenge on the computer available to me, and one notable thing I did was write a class which gave me fast random access to the whole huge training file, using low-level file operations, so I could access it in arbitrary batches if I wanted to. I'd be interested to know what kind of hardware the more successful teams used for this challenge.\n\nWith a quiet weekend at home with Covid, I indulged myself my writing up my adventures in a report. As I am nowhere near the top of the leaderboard, it's all pubic here: https://github.com/charlie-wartnaby/kaggle-leap-atmospheric-physics",
      "votes": 1,
      "replies": [
        {
          "id": 2936546,
          "postDate": "2024-07-26T08:54:31.317Z",
          "content": "<p>for me the trick was to use polars and to read batches of the big csv, train then load another df which was another chunk of the csv. It is easy in polars and the time took to process these dfs was only like 2-3 seconds while the training took much more time so efficiency because of that was not an issue. <a href=\"https://github.com/VCharatsidis/kaggle_LEAP/blob/master/seq2scalar/train_weightless_last_efficient.py\" target=\"_blank\">https://github.com/VCharatsidis/kaggle_LEAP/blob/master/seq2scalar/train_weightless_last_efficient.py</a></p>",
          "rawMarkdown": "for me the trick was to use polars and to read batches of the big csv, train then load another df which was another chunk of the csv. It is easy in polars and the time took to process these dfs was only like 2-3 seconds while the training took much more time so efficiency because of that was not an issue. https://github.com/VCharatsidis/kaggle_LEAP/blob/master/seq2scalar/train_weightless_last_efficient.py",
          "votes": 3
        }
      ]
    },
    {
      "id": 2924022,
      "postDate": "2024-07-16T09:00:24.310Z",
      "content": "<p>Looking forward to publishing solutions. The very stable situation in the leaderboard gives reason to assume that the solutions will be interesting</p>",
      "rawMarkdown": "Looking forward to publishing solutions. The very stable situation in the leaderboard gives reason to assume that the solutions will be interesting",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 2924119,
      "author_name": "Vasilis",
      "author_url": "",
      "post_date": "2024-07-16T10:45:23.977000",
      "content": "<p>i hade a simple transformer encoder with input (N, 60, 25), only thing that might be a bit \"sophisticated\" was that i was also passing the input variable wise (N, 25, 60)  with learned embeddings into a second transformer encoder and then do cross attention with the result of the first encoder. I did not have time to study if the extra encoder + cross attention gave any value though.</p>\n<pre><code>     ():\n        \n        src2 = src1.permute(, , )  \n        \n        src1 = self.input_adapter1(src1)  \n        src2 = self.input_adapter2(src2)  \n\n        \n        src1 = self.positional_encoding_1(src1)\n        src2 = self.learned_encoding_2(src2)\n\n        \n        encoded1 = self.encoder1(src1)\n        encoded2 = self.encoder2(src2)\n\n        \n        cross_attended1 = self.cross_attention_1(encoded1, encoded2, encoded2).flatten(start_dim=)\n        cross_attended2 = self.cross_attention_2(encoded2, encoded1, encoded1).flatten(start_dim=)\n\n        \n        combined = torch.cat([cross_attended1, cross_attended2], dim=)\n\n        output = self.output_linear(combined)\n         output\n</code></pre>",
      "votes": 5,
      "replies": [
        {
          "id": 2924123,
          "author_name": "wasjaip",
          "author_url": "",
          "post_date": "2024-07-16T10:56:14.390000",
          "content": "<p>Thanks! Have you considered any other permutations? how much has this feature improved?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2924200,
              "author_name": "Vasilis",
              "author_url": "",
              "post_date": "2024-07-16T11:42:59.860000",
              "content": "<p>hey no, i am not sure how the extra thingy with the cross attention adds, i had no time to try without it. What i did try is to have conv1d layers between the transformer encoder layers but i did not manage to make it work better in a few attempts and since there were only a few <br>\ndays left i did not try further</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 2932421,
          "author_name": "Juan D C F",
          "author_url": "",
          "post_date": "2024-07-22T23:29:34.243000",
          "content": "<p>Neat trick! How did you come up with the idea of using cross-attention? I used a single encoder model, but the score wasn't as good as yours. </p>",
          "votes": 1,
          "replies": [
            {
              "id": 2936544,
              "author_name": "Vasilis",
              "author_url": "",
              "post_date": "2024-07-26T08:51:40.897000",
              "content": "<p>i just was fixated that i need to do something also variable-wise and then somehow concat the result, well cross_attention seemed a better idea. but maybe you had worse score for some other reason.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2936456,
      "author_name": "Charlie Wartnaby",
      "author_url": "",
      "post_date": "2024-07-26T06:55:19.917000",
      "content": "<p>I was only 93rd in the competition, but I used a convolutional neural network with a big twist: it also ingested the predictions of a CatBoost decision tree model as additional input features. So the two techniques were effectively in series. (I also tried combining their predictions as various weighted sums, which didn't help, as the CNN was just better.)</p>\n<p>Apart from the actual machine learning, I found the data handling a challenge on the computer available to me, and one notable thing I did was write a class which gave me fast random access to the whole huge training file, using low-level file operations, so I could access it in arbitrary batches if I wanted to. I'd be interested to know what kind of hardware the more successful teams used for this challenge.</p>\n<p>With a quiet weekend at home with Covid, I indulged myself my writing up my adventures in a report. As I am nowhere near the top of the leaderboard, it's all pubic here: <a href=\"https://github.com/charlie-wartnaby/kaggle-leap-atmospheric-physics\" target=\"_blank\">https://github.com/charlie-wartnaby/kaggle-leap-atmospheric-physics</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 2936546,
          "author_name": "Vasilis",
          "author_url": "",
          "post_date": "2024-07-26T08:54:31.317000",
          "content": "<p>for me the trick was to use polars and to read batches of the big csv, train then load another df which was another chunk of the csv. It is easy in polars and the time took to process these dfs was only like 2-3 seconds while the training took much more time so efficiency because of that was not an issue. <a href=\"https://github.com/VCharatsidis/kaggle_LEAP/blob/master/seq2scalar/train_weightless_last_efficient.py\" target=\"_blank\">https://github.com/VCharatsidis/kaggle_LEAP/blob/master/seq2scalar/train_weightless_last_efficient.py</a></p>",
          "votes": 3,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2924119": "i hade a simple transformer encoder with input (N, 60, 25), only thing that might be a bit \"sophisticated\" was that i was also passing the input variable wise (N, 25, 60)  with learned embeddings into a second transformer encoder and then do cross attention with the result of the first encoder. I did not have time to study if the extra encoder + cross attention gave any value though.\n\n```\n    def forward(self, src1):\n        # Reshape and reorder to N, 25, 60\n        src2 = src1.permute(0, 2, 1)  # Change to N, 25, 60\n        # Prepare inputs\n        src1 = self.input_adapter1(src1)  # N, 60, 25 -> N, 60, d_model\n        src2 = self.input_adapter2(src2)  # N, 25, 60 -> N, 25, d_model\n\n        # Add positional encoding\n        src1 = self.positional_encoding_1(src1)\n        src2 = self.learned_encoding_2(src2)\n\n        # Encode with transformers\n        encoded1 = self.encoder1(src1)\n        encoded2 = self.encoder2(src2)\n\n        # Apply cross-attention\n        cross_attended1 = self.cross_attention_1(encoded1, encoded2, encoded2).flatten(start_dim=1)\n        cross_attended2 = self.cross_attention_2(encoded2, encoded1, encoded1).flatten(start_dim=1)\n\n        # Combine outputs from both paths\n        combined = torch.cat([cross_attended1, cross_attended2], dim=1)\n\n        output = self.output_linear(combined)\n        return output\n```",
    "2936456": "I was only 93rd in the competition, but I used a convolutional neural network with a big twist: it also ingested the predictions of a CatBoost decision tree model as additional input features. So the two techniques were effectively in series. (I also tried combining their predictions as various weighted sums, which didn't help, as the CNN was just better.)\n\nApart from the actual machine learning, I found the data handling a challenge on the computer available to me, and one notable thing I did was write a class which gave me fast random access to the whole huge training file, using low-level file operations, so I could access it in arbitrary batches if I wanted to. I'd be interested to know what kind of hardware the more successful teams used for this challenge.\n\nWith a quiet weekend at home with Covid, I indulged myself my writing up my adventures in a report. As I am nowhere near the top of the leaderboard, it's all pubic here: https://github.com/charlie-wartnaby/kaggle-leap-atmospheric-physics",
    "2924022": "Looking forward to publishing solutions. The very stable situation in the leaderboard gives reason to assume that the solutions will be interesting"
  }
}