{
  "id": 609995,
  "title": "10th Place Solution – DL part",
  "url": "/competitions/ariel-data-challenge-2025/discussion/609995",
  "author_name": "Artem Veshkin",
  "post_date": "2025-10-01T00:13:21.696000",
  "votes": 6,
  "comment_count": 2,
  "views": 0,
  "content": "<p>First of all, I want to express my gratitude to the organizers for hosting this competition. I had a fantastic time working on it. </p>\n<p>And also I was a part of a great team with <a href=\"https://www.kaggle.com/solverworld\" target=\"_blank\">@solverworld</a> <a href=\"https://www.kaggle.com/vitalykudelya\" target=\"_blank\">@vitalykudelya</a> <a href=\"https://www.kaggle.com/asimandia\" target=\"_blank\">@asimandia</a>. Thank you guys, you are the best!</p>\n<p>Our final solution is composed of three entirely different parts. In this post, I will describe the Deep Learning part.</p>\n<p>It utilizes two models, which I've named <strong>TimeFormer</strong> and <strong>PlanetFormer</strong>.</p>\n<hr>\n<h2>TimeFormer</h2>\n<p>TimeFormer processes three adjacent signals (5625x3) as input and predicts three corresponding mean and sigma values.</p>\n<p>It also has a head that predicts signal zones over time. For zone labeling, I adapted an algorithm from last year's winning solution and tuned it for the new dataset.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7160377%2F0c6c734c99c85c76f40e3835532442c6%2Ftimeformer_input.png?generation=1759276778105413&amp;alt=media\" alt=\"\"></p>\n<p>To further improve the zone interpretation, <a href=\"https://www.kaggle.com/asimandia\" target=\"_blank\">@asimandia</a> spent several sleepless nights manually labeling the zones (t1, t2, t3, t4) for the entire dataset. This manual annotation was used during model training and provided a slight performance boost.</p>\n<p>We had planned to use these labels to enhance the signal zone detector for better inference quality, but unfortunately, we ran out of time.</p>\n<p>During inference, the model is applied with an overlap, and its predictions for neighboring signals are averaged.</p>\n<h3>Architectural Tricks:</h3>\n<ul>\n<li><strong>Learnable embeddings</strong> for each zone, complementing the standard positional embeddings.</li>\n<li>A hybrid approach using both <strong>convolutions with various kernel sizes and attention</strong>.</li>\n<li><strong>Zone-Aware Attention</strong>: I injected prior knowledge about statistically important signal zones directly into the attention mechanism. This was achieved by adding a learned bias to the query and key vectors to guide the model's focus: $$Q' = Q + \\alpha \\cdot \\text{proj}(E_{\\text{zone}})$$</li>\n<li><strong>Gated Residual Connections</strong></li>\n<li><strong>Mixture of Experts (MoE) Heads</strong></li>\n<li>A dedicated head that learns to <strong>predict signal zones</strong> after the transformer layer.</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7160377%2F7274145ef0c9fc1d37af22cc2bac920d%2Ftimeformer.png?generation=1759276890354426&amp;alt=media\" alt=\"\"></p>\n<p>On its own, TimeFormer can be used for prediction and achieves a public leaderboard score of approximately <strong>0.46</strong>. However, it has limitations in capturing the  global context.</p>\n<p>To analyze the entire signal, I run TimeFormer on each wave and extract the post-pooling embedding. This embedding provides a good temporal description of the signal. This process transforms a (5625x283) signal into 283 embeddings (32x283).</p>\n<p><strong>Fun observation</strong>: Here is a PCA visualization of the TimeFormer embeddings. I find it a bit suspicious that in a competition about space, the embeddings looks like the Alien :)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7160377%2Fc31d2bf22c5573e2a2a176c573a93d88%2Ftimeformer_embeddings.png?generation=1759277189238192&amp;alt=media\" alt=\"\"></p>\n<hr>\n<h2>PlanetFormer</h2>\n<p>PlanetFormer takes the 283 embeddings from TimeFormer as input and outputs 283 mean and 283 std values.</p>\n<p>The target variable exhibits a highly consistent structure, which can be leveraged to reduce the number of model parameters and mitigate overfitting. Instead of predicting 283 distinct values, the <strong>mean head</strong> predicts only 21 coefficients for a 20th-degree polynomial used to model the target.</p>\n<p>Similarly, the <strong>sigma head</strong> (which actually predicts <code>log_var</code> for numerical stability) outputs just 7 values instead of 283. The underlying idea is that uncertainty is likely concentrated in specific, adjacent groups of waves.</p>\n<p>I model this uncertainty using a set of basis functions: <code>torch.ones()</code>, <code>sin(pi * x)</code>, <code>cos(pi * x)</code>, <code>sin(2pi * x)</code>, <code>cos(2pi * x)</code>, <code>sin(3pi * x)</code>, and <code>cos(3pi * x)</code>. The output of the <code>log_var</code> head determines the weights for combining these functions.</p>\n<p>Both the mean and sigma heads also utilize MoE for their predictions.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7160377%2F53d7e19e00418f3a9f5d2d82cb53fc84%2Fplanetformer.png?generation=1759277234626437&amp;alt=media\" alt=\"\"></p>\n<p>PlanetFormer achieves a public LB score of approximately <strong>0.48-0.49</strong>.</p>\n<p>To further enhance prediction accuracy, the outputs of TimeFormer and PlanetFormer are ensembled with weights of <strong>[0.15, 0.85]</strong>.</p>\n<hr>\n<h2>Signal Preprocessing and Augmentations</h2>\n<p>As indicated earlier, the models operate with <code>binning=1</code>. Before being fed into the models, the signal is processed with a <strong>low-pass filter</strong>. Transit zone was <strong>centered</strong>. Additionally, a <strong>\"time-axis flip\" augmentation</strong> was applied during the training of both models.</p>\n<hr>\n<h2>Loss Function</h2>\n<p>Both models were trained using a custom loss function that I designed to closely approximate the competition's evaluation metric. The FGS channel is assigned a weight of 57.8641. Here is the code for the loss:</p>\n<pre><code> (nn.Module):\n     ():\n        ().__init__()\n        .naive_mean = naive_mean\n        .naive_sigma = naive_sigma\n        .fsg_sigma_true = fsg_sigma_true\n        .airs_sigma_true = airs_sigma_true\n\n     ():\n        pred_sigma = torch.sqrt(torch.exp(log_var_pred)).clamp(=)\n\n        neg_log_likelihood =  * (\n            torch.log( * np.pi * pred_sigma.()) + \n            (target - pred_mean).() / pred_sigma.()\n        )\n\n        sigma_penalty = torch.where(\n            pred_sigma &lt; .airs_sigma_true,\n            (.airs_sigma_true - pred_sigma).() * ,\n            torch.zeros_like(pred_sigma)\n        )\n\n        \n        fgs_mask = torch.where(weights &gt; )\n        sigma_penalty[fgs_mask] = torch.where(\n            pred_sigma[fgs_mask] &lt; .fsg_sigma_true,\n            (.fsg_sigma_true - pred_sigma[fgs_mask]).() * ,\n            sigma_penalty[fgs_mask]\n        )\n\n        total_loss = (neg_log_likelihood + sigma_penalty) * weights\n\n         total_loss.mean()\n</code></pre>",
  "messages": [
    {
      "id": 3296448,
      "postDate": "2025-10-01T00:13:21.697Z",
      "content": "<p>First of all, I want to express my gratitude to the organizers for hosting this competition. I had a fantastic time working on it. </p>\n<p>And also I was a part of a great team with <a href=\"https://www.kaggle.com/solverworld\" target=\"_blank\">@solverworld</a> <a href=\"https://www.kaggle.com/vitalykudelya\" target=\"_blank\">@vitalykudelya</a> <a href=\"https://www.kaggle.com/asimandia\" target=\"_blank\">@asimandia</a>. Thank you guys, you are the best!</p>\n<p>Our final solution is composed of three entirely different parts. In this post, I will describe the Deep Learning part.</p>\n<p>It utilizes two models, which I've named <strong>TimeFormer</strong> and <strong>PlanetFormer</strong>.</p>\n<hr>\n<h2>TimeFormer</h2>\n<p>TimeFormer processes three adjacent signals (5625x3) as input and predicts three corresponding mean and sigma values.</p>\n<p>It also has a head that predicts signal zones over time. For zone labeling, I adapted an algorithm from last year's winning solution and tuned it for the new dataset.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7160377%2F0c6c734c99c85c76f40e3835532442c6%2Ftimeformer_input.png?generation=1759276778105413&amp;alt=media\" alt=\"\"></p>\n<p>To further improve the zone interpretation, <a href=\"https://www.kaggle.com/asimandia\" target=\"_blank\">@asimandia</a> spent several sleepless nights manually labeling the zones (t1, t2, t3, t4) for the entire dataset. This manual annotation was used during model training and provided a slight performance boost.</p>\n<p>We had planned to use these labels to enhance the signal zone detector for better inference quality, but unfortunately, we ran out of time.</p>\n<p>During inference, the model is applied with an overlap, and its predictions for neighboring signals are averaged.</p>\n<h3>Architectural Tricks:</h3>\n<ul>\n<li><strong>Learnable embeddings</strong> for each zone, complementing the standard positional embeddings.</li>\n<li>A hybrid approach using both <strong>convolutions with various kernel sizes and attention</strong>.</li>\n<li><strong>Zone-Aware Attention</strong>: I injected prior knowledge about statistically important signal zones directly into the attention mechanism. This was achieved by adding a learned bias to the query and key vectors to guide the model's focus: $$Q' = Q + \\alpha \\cdot \\text{proj}(E_{\\text{zone}})$$</li>\n<li><strong>Gated Residual Connections</strong></li>\n<li><strong>Mixture of Experts (MoE) Heads</strong></li>\n<li>A dedicated head that learns to <strong>predict signal zones</strong> after the transformer layer.</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7160377%2F7274145ef0c9fc1d37af22cc2bac920d%2Ftimeformer.png?generation=1759276890354426&amp;alt=media\" alt=\"\"></p>\n<p>On its own, TimeFormer can be used for prediction and achieves a public leaderboard score of approximately <strong>0.46</strong>. However, it has limitations in capturing the  global context.</p>\n<p>To analyze the entire signal, I run TimeFormer on each wave and extract the post-pooling embedding. This embedding provides a good temporal description of the signal. This process transforms a (5625x283) signal into 283 embeddings (32x283).</p>\n<p><strong>Fun observation</strong>: Here is a PCA visualization of the TimeFormer embeddings. I find it a bit suspicious that in a competition about space, the embeddings looks like the Alien :)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7160377%2Fc31d2bf22c5573e2a2a176c573a93d88%2Ftimeformer_embeddings.png?generation=1759277189238192&amp;alt=media\" alt=\"\"></p>\n<hr>\n<h2>PlanetFormer</h2>\n<p>PlanetFormer takes the 283 embeddings from TimeFormer as input and outputs 283 mean and 283 std values.</p>\n<p>The target variable exhibits a highly consistent structure, which can be leveraged to reduce the number of model parameters and mitigate overfitting. Instead of predicting 283 distinct values, the <strong>mean head</strong> predicts only 21 coefficients for a 20th-degree polynomial used to model the target.</p>\n<p>Similarly, the <strong>sigma head</strong> (which actually predicts <code>log_var</code> for numerical stability) outputs just 7 values instead of 283. The underlying idea is that uncertainty is likely concentrated in specific, adjacent groups of waves.</p>\n<p>I model this uncertainty using a set of basis functions: <code>torch.ones()</code>, <code>sin(pi * x)</code>, <code>cos(pi * x)</code>, <code>sin(2pi * x)</code>, <code>cos(2pi * x)</code>, <code>sin(3pi * x)</code>, and <code>cos(3pi * x)</code>. The output of the <code>log_var</code> head determines the weights for combining these functions.</p>\n<p>Both the mean and sigma heads also utilize MoE for their predictions.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7160377%2F53d7e19e00418f3a9f5d2d82cb53fc84%2Fplanetformer.png?generation=1759277234626437&amp;alt=media\" alt=\"\"></p>\n<p>PlanetFormer achieves a public LB score of approximately <strong>0.48-0.49</strong>.</p>\n<p>To further enhance prediction accuracy, the outputs of TimeFormer and PlanetFormer are ensembled with weights of <strong>[0.15, 0.85]</strong>.</p>\n<hr>\n<h2>Signal Preprocessing and Augmentations</h2>\n<p>As indicated earlier, the models operate with <code>binning=1</code>. Before being fed into the models, the signal is processed with a <strong>low-pass filter</strong>. Transit zone was <strong>centered</strong>. Additionally, a <strong>\"time-axis flip\" augmentation</strong> was applied during the training of both models.</p>\n<hr>\n<h2>Loss Function</h2>\n<p>Both models were trained using a custom loss function that I designed to closely approximate the competition's evaluation metric. The FGS channel is assigned a weight of 57.8641. Here is the code for the loss:</p>\n<pre><code> (nn.Module):\n     ():\n        ().__init__()\n        .naive_mean = naive_mean\n        .naive_sigma = naive_sigma\n        .fsg_sigma_true = fsg_sigma_true\n        .airs_sigma_true = airs_sigma_true\n\n     ():\n        pred_sigma = torch.sqrt(torch.exp(log_var_pred)).clamp(=)\n\n        neg_log_likelihood =  * (\n            torch.log( * np.pi * pred_sigma.()) + \n            (target - pred_mean).() / pred_sigma.()\n        )\n\n        sigma_penalty = torch.where(\n            pred_sigma &lt; .airs_sigma_true,\n            (.airs_sigma_true - pred_sigma).() * ,\n            torch.zeros_like(pred_sigma)\n        )\n\n        \n        fgs_mask = torch.where(weights &gt; )\n        sigma_penalty[fgs_mask] = torch.where(\n            pred_sigma[fgs_mask] &lt; .fsg_sigma_true,\n            (.fsg_sigma_true - pred_sigma[fgs_mask]).() * ,\n            sigma_penalty[fgs_mask]\n        )\n\n        total_loss = (neg_log_likelihood + sigma_penalty) * weights\n\n         total_loss.mean()\n</code></pre>",
      "rawMarkdown": "First of all, I want to express my gratitude to the organizers for hosting this competition. I had a fantastic time working on it. \n\nAnd also I was a part of a great team with @solverworld @vitalykudelya @asimandia. Thank you guys, you are the best!\n\nOur final solution is composed of three entirely different parts. In this post, I will describe the Deep Learning part.\n\nIt utilizes two models, which I've named **TimeFormer** and **PlanetFormer**.\n\n-----\n\n## TimeFormer\n\nTimeFormer processes three adjacent signals (5625x3) as input and predicts three corresponding mean and sigma values.\n\nIt also has a head that predicts signal zones over time. For zone labeling, I adapted an algorithm from last year's winning solution and tuned it for the new dataset.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7160377%2F0c6c734c99c85c76f40e3835532442c6%2Ftimeformer_input.png?generation=1759276778105413&alt=media)\n\nTo further improve the zone interpretation, @asimandia spent several sleepless nights manually labeling the zones (t1, t2, t3, t4) for the entire dataset. This manual annotation was used during model training and provided a slight performance boost.\n\nWe had planned to use these labels to enhance the signal zone detector for better inference quality, but unfortunately, we ran out of time.\n\nDuring inference, the model is applied with an overlap, and its predictions for neighboring signals are averaged.\n\n### Architectural Tricks:\n\n  * **Learnable embeddings** for each zone, complementing the standard positional embeddings.\n  * A hybrid approach using both **convolutions with various kernel sizes and attention**.\n  * **Zone-Aware Attention**: I injected prior knowledge about statistically important signal zones directly into the attention mechanism. This was achieved by adding a learned bias to the query and key vectors to guide the model's focus: $$Q' = Q + \\alpha \\cdot \\text{proj}(E_{\\text{zone}})$$\n  * **Gated Residual Connections**\n  * **Mixture of Experts (MoE) Heads**\n  * A dedicated head that learns to **predict signal zones** after the transformer layer.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7160377%2F7274145ef0c9fc1d37af22cc2bac920d%2Ftimeformer.png?generation=1759276890354426&alt=media)\n\nOn its own, TimeFormer can be used for prediction and achieves a public leaderboard score of approximately **0.46**. However, it has limitations in capturing the  global context.\n\nTo analyze the entire signal, I run TimeFormer on each wave and extract the post-pooling embedding. This embedding provides a good temporal description of the signal. This process transforms a (5625x283) signal into 283 embeddings (32x283).\n\n**Fun observation**: Here is a PCA visualization of the TimeFormer embeddings. I find it a bit suspicious that in a competition about space, the embeddings looks like the Alien :)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7160377%2Fc31d2bf22c5573e2a2a176c573a93d88%2Ftimeformer_embeddings.png?generation=1759277189238192&alt=media)\n\n-----\n\n## PlanetFormer\n\nPlanetFormer takes the 283 embeddings from TimeFormer as input and outputs 283 mean and 283 std values.\n\nThe target variable exhibits a highly consistent structure, which can be leveraged to reduce the number of model parameters and mitigate overfitting. Instead of predicting 283 distinct values, the **mean head** predicts only 21 coefficients for a 20th-degree polynomial used to model the target.\n\nSimilarly, the **sigma head** (which actually predicts `log_var` for numerical stability) outputs just 7 values instead of 283. The underlying idea is that uncertainty is likely concentrated in specific, adjacent groups of waves.\n\nI model this uncertainty using a set of basis functions: `torch.ones()`, `sin(pi * x)`, `cos(pi * x)`, `sin(2pi * x)`, `cos(2pi * x)`, `sin(3pi * x)`, and `cos(3pi * x)`. The output of the `log_var` head determines the weights for combining these functions.\n\nBoth the mean and sigma heads also utilize MoE for their predictions.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7160377%2F53d7e19e00418f3a9f5d2d82cb53fc84%2Fplanetformer.png?generation=1759277234626437&alt=media)\n\nPlanetFormer achieves a public LB score of approximately **0.48-0.49**.\n\nTo further enhance prediction accuracy, the outputs of TimeFormer and PlanetFormer are ensembled with weights of **[0.15, 0.85]**.\n\n-----\n\n## Signal Preprocessing and Augmentations\n\nAs indicated earlier, the models operate with `binning=1`. Before being fed into the models, the signal is processed with a **low-pass filter**. Transit zone was **centered**. Additionally, a **\"time-axis flip\" augmentation** was applied during the training of both models.\n\n-----\n\n## Loss Function\n\nBoth models were trained using a custom loss function that I designed to closely approximate the competition's evaluation metric. The FGS channel is assigned a weight of 57.8641. Here is the code for the loss:\n\n```python\nclass CompetitionLoss(nn.Module):\n    def __init__(self, naive_mean=1.46890195e-02, naive_sigma=1.06613525e-02,\n                 fsg_sigma_true=1e-6, airs_sigma_true=1e-5):\n        super().__init__()\n        self.naive_mean = naive_mean\n        self.naive_sigma = naive_sigma\n        self.fsg_sigma_true = fsg_sigma_true\n        self.airs_sigma_true = airs_sigma_true\n        \n    def forward(self, pred_mean, log_var_pred, target, weights):\n        pred_sigma = torch.sqrt(torch.exp(log_var_pred)).clamp(min=1e-15)\n\n        neg_log_likelihood = 0.5 * (\n            torch.log(2 * np.pi * pred_sigma.pow(2)) + \n            (target - pred_mean).pow(2) / pred_sigma.pow(2)\n        )\n\n        sigma_penalty = torch.where(\n            pred_sigma < self.airs_sigma_true,\n            (self.airs_sigma_true - pred_sigma).pow(2) * 100,\n            torch.zeros_like(pred_sigma)\n        )\n        \n        # Extra penalty for FGS1 channel\n        fgs_mask = torch.where(weights > 1)\n        sigma_penalty[fgs_mask] = torch.where(\n            pred_sigma[fgs_mask] < self.fsg_sigma_true,\n            (self.fsg_sigma_true - pred_sigma[fgs_mask]).pow(2) * 1000,\n            sigma_penalty[fgs_mask]\n        )\n        \n        total_loss = (neg_log_likelihood + sigma_penalty) * weights\n        \n        return total_loss.mean()\n```",
      "votes": 6
    },
    {
      "id": 3296516,
      "postDate": "2025-10-01T04:42:18.423Z",
      "content": "<p>Kudos for predicting poly coefficients!<br>\nThat was one of the first things I tried but I couldn't get it working<br>\nNice to see another attention based solution</p>",
      "rawMarkdown": "Kudos for predicting poly coefficients!\nThat was one of the first things I tried but I couldn't get it working\nNice to see another attention based solution",
      "votes": 1,
      "replies": [
        {
          "id": 3296575,
          "postDate": "2025-10-01T07:44:53.043Z",
          "content": "<p>Thanks! A huge congrats to you as well!</p>",
          "rawMarkdown": "Thanks! A huge congrats to you as well!"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3296516,
      "author_name": "sroger",
      "author_url": "",
      "post_date": "2025-10-01T04:42:18.423000",
      "content": "<p>Kudos for predicting poly coefficients!<br>\nThat was one of the first things I tried but I couldn't get it working<br>\nNice to see another attention based solution</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3296575,
          "author_name": "Artem Veshkin",
          "author_url": "",
          "post_date": "2025-10-01T07:44:53.043000",
          "content": "<p>Thanks! A huge congrats to you as well!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3296448": "First of all, I want to express my gratitude to the organizers for hosting this competition. I had a fantastic time working on it. \n\nAnd also I was a part of a great team with @solverworld @vitalykudelya @asimandia. Thank you guys, you are the best!\n\nOur final solution is composed of three entirely different parts. In this post, I will describe the Deep Learning part.\n\nIt utilizes two models, which I've named **TimeFormer** and **PlanetFormer**.\n\n-----\n\n## TimeFormer\n\nTimeFormer processes three adjacent signals (5625x3) as input and predicts three corresponding mean and sigma values.\n\nIt also has a head that predicts signal zones over time. For zone labeling, I adapted an algorithm from last year's winning solution and tuned it for the new dataset.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7160377%2F0c6c734c99c85c76f40e3835532442c6%2Ftimeformer_input.png?generation=1759276778105413&alt=media)\n\nTo further improve the zone interpretation, @asimandia spent several sleepless nights manually labeling the zones (t1, t2, t3, t4) for the entire dataset. This manual annotation was used during model training and provided a slight performance boost.\n\nWe had planned to use these labels to enhance the signal zone detector for better inference quality, but unfortunately, we ran out of time.\n\nDuring inference, the model is applied with an overlap, and its predictions for neighboring signals are averaged.\n\n### Architectural Tricks:\n\n  * **Learnable embeddings** for each zone, complementing the standard positional embeddings.\n  * A hybrid approach using both **convolutions with various kernel sizes and attention**.\n  * **Zone-Aware Attention**: I injected prior knowledge about statistically important signal zones directly into the attention mechanism. This was achieved by adding a learned bias to the query and key vectors to guide the model's focus: $$Q' = Q + \\alpha \\cdot \\text{proj}(E_{\\text{zone}})$$\n  * **Gated Residual Connections**\n  * **Mixture of Experts (MoE) Heads**\n  * A dedicated head that learns to **predict signal zones** after the transformer layer.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7160377%2F7274145ef0c9fc1d37af22cc2bac920d%2Ftimeformer.png?generation=1759276890354426&alt=media)\n\nOn its own, TimeFormer can be used for prediction and achieves a public leaderboard score of approximately **0.46**. However, it has limitations in capturing the  global context.\n\nTo analyze the entire signal, I run TimeFormer on each wave and extract the post-pooling embedding. This embedding provides a good temporal description of the signal. This process transforms a (5625x283) signal into 283 embeddings (32x283).\n\n**Fun observation**: Here is a PCA visualization of the TimeFormer embeddings. I find it a bit suspicious that in a competition about space, the embeddings looks like the Alien :)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7160377%2Fc31d2bf22c5573e2a2a176c573a93d88%2Ftimeformer_embeddings.png?generation=1759277189238192&alt=media)\n\n-----\n\n## PlanetFormer\n\nPlanetFormer takes the 283 embeddings from TimeFormer as input and outputs 283 mean and 283 std values.\n\nThe target variable exhibits a highly consistent structure, which can be leveraged to reduce the number of model parameters and mitigate overfitting. Instead of predicting 283 distinct values, the **mean head** predicts only 21 coefficients for a 20th-degree polynomial used to model the target.\n\nSimilarly, the **sigma head** (which actually predicts `log_var` for numerical stability) outputs just 7 values instead of 283. The underlying idea is that uncertainty is likely concentrated in specific, adjacent groups of waves.\n\nI model this uncertainty using a set of basis functions: `torch.ones()`, `sin(pi * x)`, `cos(pi * x)`, `sin(2pi * x)`, `cos(2pi * x)`, `sin(3pi * x)`, and `cos(3pi * x)`. The output of the `log_var` head determines the weights for combining these functions.\n\nBoth the mean and sigma heads also utilize MoE for their predictions.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7160377%2F53d7e19e00418f3a9f5d2d82cb53fc84%2Fplanetformer.png?generation=1759277234626437&alt=media)\n\nPlanetFormer achieves a public LB score of approximately **0.48-0.49**.\n\nTo further enhance prediction accuracy, the outputs of TimeFormer and PlanetFormer are ensembled with weights of **[0.15, 0.85]**.\n\n-----\n\n## Signal Preprocessing and Augmentations\n\nAs indicated earlier, the models operate with `binning=1`. Before being fed into the models, the signal is processed with a **low-pass filter**. Transit zone was **centered**. Additionally, a **\"time-axis flip\" augmentation** was applied during the training of both models.\n\n-----\n\n## Loss Function\n\nBoth models were trained using a custom loss function that I designed to closely approximate the competition's evaluation metric. The FGS channel is assigned a weight of 57.8641. Here is the code for the loss:\n\n```python\nclass CompetitionLoss(nn.Module):\n    def __init__(self, naive_mean=1.46890195e-02, naive_sigma=1.06613525e-02,\n                 fsg_sigma_true=1e-6, airs_sigma_true=1e-5):\n        super().__init__()\n        self.naive_mean = naive_mean\n        self.naive_sigma = naive_sigma\n        self.fsg_sigma_true = fsg_sigma_true\n        self.airs_sigma_true = airs_sigma_true\n        \n    def forward(self, pred_mean, log_var_pred, target, weights):\n        pred_sigma = torch.sqrt(torch.exp(log_var_pred)).clamp(min=1e-15)\n\n        neg_log_likelihood = 0.5 * (\n            torch.log(2 * np.pi * pred_sigma.pow(2)) + \n            (target - pred_mean).pow(2) / pred_sigma.pow(2)\n        )\n\n        sigma_penalty = torch.where(\n            pred_sigma < self.airs_sigma_true,\n            (self.airs_sigma_true - pred_sigma).pow(2) * 100,\n            torch.zeros_like(pred_sigma)\n        )\n        \n        # Extra penalty for FGS1 channel\n        fgs_mask = torch.where(weights > 1)\n        sigma_penalty[fgs_mask] = torch.where(\n            pred_sigma[fgs_mask] < self.fsg_sigma_true,\n            (self.fsg_sigma_true - pred_sigma[fgs_mask]).pow(2) * 1000,\n            sigma_penalty[fgs_mask]\n        )\n        \n        total_loss = (neg_log_likelihood + sigma_penalty) * weights\n        \n        return total_loss.mean()\n```",
    "3296516": "Kudos for predicting poly coefficients!\nThat was one of the first things I tried but I couldn't get it working\nNice to see another attention based solution"
  }
}