{
  "id": 609252,
  "title": "3rd Place Solution",
  "url": "/competitions/ariel-data-challenge-2025/discussion/609252",
  "author_name": "moonpole",
  "post_date": "2025-09-25T09:43:22.797000",
  "votes": 32,
  "comment_count": 6,
  "views": 0,
  "content": "<h1>Introduction &amp; Summary</h1>\n<p>Huge thank you to the organizers as well as the staff for this amazing competition! I learnt a lot about data engineering and modelling techniques and I look forward to competing in the future!</p>\n<p>My approach contain those steps:</p>\n<ul>\n<li>Step 1: Preprocessing &amp; Calibration</li>\n<li>Step 2: Signal extraction</li>\n<li>Step 3: Feed into CNN for additional feature extraction</li>\n<li>Step 4: With signal features and CNN features, feed into ensemble of rational quadratic neural network</li>\n</ul>\n<h1>Preprocessing &amp; Calibration</h1>\n<p>I used the standard calibration techniques as given in the competition itself. It includes the calibration for gain/offset correction, and manages non-linearity, dark, dead and hot pixels. Additionally, at the end of the calibration step, I remove signals that are more than $8 \\sigma$ away in the time dimension. This addition alone adds ~0.05 to my CV (which is just a simple 80-20 split, turning out to be almost same as LB; LB feedback time is 5-8 hours). I couldn't spend much time on this section as rerunning this for all data, even with joblib, takes 1+ hour.</p>\n<pre><code>mean = np.nanmean(signal, )[, :, :]\nstd = np.nanstd(signal, )[, :, :]\nsignal[(signal &gt; mean +  * std) | (signal &lt; mean -  * std)] = np.nan\n</code></pre>\n<p>The above is the exact added code used in inference.</p>\n<p>Given features like <code>Rs</code> were included as well as minor feature engineering, like <code>Rs ** 2</code>, but they don't seem too important.</p>\n<h1>Signal Extraction</h1>\n<p>Phase extraction is done using extrema of gradient.</p>\n<ol>\n<li>The calibrated signal is divided by the mean of itself</li>\n<li>A simple moving average of period k is applied to the signal (k = 88 for airs and 655 for fgs1, found using optuna)</li>\n<li>A set of moving averages is then applied to the gradient of the signal, in which p,q = argmin,argmax are extracted</li>\n<li>p,q are then moved to the left,right until their gradients are within some interquartile range of all gradients so they cover all of the transit phase</li>\n</ol>\n<p>The features are then extracted via simple mean/min/max/range/min div edge (approx transit depth)/polyfit to 8th degree etc. of each of the phases. Additionally, the airs data is <code>array_split</code>'d along the frequency dimension (size 356) into 32 chunks. Each chunk then gets processed like above in the same way, and their features added to the list of features we have. This adds quite a bit of resolution and adds ~0.025 CV.</p>\n<h1>Feed into CNN</h1>\n<p>The CNN doesn't care about the phases; the input is simply the airs and fgs1 signal after the first moving average, split into 92 chunks across time with each chunk averaged. Additional steps might be adding phase signal into the CNN, as well as separating airs/fgs1 data into different subnets since they might not be time aligned, but ideas like these were not implemented</p>\n<p>There are many instances of CNN, as they are integrated into the RQ NN ensemble. To prevent overfitting, almost all layers of the CNN have the same weight and shape. No activation function other than max pool was used.</p>\n<h1>Ensemble of Rational Quadratic NN</h1>\n<h2>Preparing the features</h2>\n<p>One idea was borrowed from <a href=\"https://arxiv.org/pdf/2407.04491\" target=\"_blank\">https://arxiv.org/pdf/2407.04491</a>.</p>\n<ul>\n<li>I used same feature scaling but with $\\frac{x}{\\sqrt{18+x^2/9}}$</li>\n</ul>\n<h2>Training</h2>\n<p>The idea is deep ensembling where each instance in the ensemble is trained the exact same way, just with different weight initialization. We just straight up predict everything, the <code>pred</code> and <code>sigma</code> with our model, no sampling or quantile regression.</p>\n<ul>\n<li>Optimizer is adabelief with (b1=0.999, b2=0.9998)</li>\n<li>Cosine one cycle learning rate schedule with square root shaping</li>\n<li>36 ensembles, with 88 RQ clusters</li>\n</ul>\n<p>Intuitively, I chose soft RBF/RQ (they perform similarly, both adds ~0.15 CV) because there were many transit curve shapes (it was mentioned that many different models of limb darkening were added). So this soft RQ aims to first sort instances into soft groups, and then perform simple linear regression within. The image shows the distribution of the first RQ cluster of the first instance.</p>\n<p>The training is quite unique; instead of training on average of losses (competition metric), I spent more effort training models which are performing worse. Similar to evolution, the best models don't need to get better anymore. This avoids overfitting and improves my CV by ~0.03+.</p>\n<pre><code>m = jnp.median(losses, ) \nw = jax.nn.relu(losses - m)\nw = w / w.(, keepdims=) \nw = jax.lax.stop_gradient(w)\n (losses * w).() / w.()\n</code></pre>\n<p>The above is the exact code used in training. Only values worse than the median will get trained.</p>\n<h1>What didn't work &amp; Conclusion</h1>\n<ul>\n<li>L1, L2 regularisation.</li>\n<li>Dropout (works well in very specific circumstances; unstable).</li>\n<li>Data augmentation via random splicing of transit data (although I saw there were some instances where the phase is cut off).</li>\n<li>Fischer-information-like or entropy based regularisation for RBF/RQ, to encourage more uncertainty of groups</li>\n<li>Full trainable Mahalanobis distance</li>\n</ul>\n<p>I used both observations for the same planet if they exist during training, but only the first observation during inference. The NN inference is instant and the data reading takes up the majority of the time.</p>\n<p>In conclusion an acceptable result was achieved with little or no understanding of the underlying physics nor advanced drift/transit window denoising techniques.</p>",
  "messages": [
    {
      "id": 3294072,
      "postDate": "2025-09-25T09:43:22.797Z",
      "content": "<h1>Introduction &amp; Summary</h1>\n<p>Huge thank you to the organizers as well as the staff for this amazing competition! I learnt a lot about data engineering and modelling techniques and I look forward to competing in the future!</p>\n<p>My approach contain those steps:</p>\n<ul>\n<li>Step 1: Preprocessing &amp; Calibration</li>\n<li>Step 2: Signal extraction</li>\n<li>Step 3: Feed into CNN for additional feature extraction</li>\n<li>Step 4: With signal features and CNN features, feed into ensemble of rational quadratic neural network</li>\n</ul>\n<h1>Preprocessing &amp; Calibration</h1>\n<p>I used the standard calibration techniques as given in the competition itself. It includes the calibration for gain/offset correction, and manages non-linearity, dark, dead and hot pixels. Additionally, at the end of the calibration step, I remove signals that are more than $8 \\sigma$ away in the time dimension. This addition alone adds ~0.05 to my CV (which is just a simple 80-20 split, turning out to be almost same as LB; LB feedback time is 5-8 hours). I couldn't spend much time on this section as rerunning this for all data, even with joblib, takes 1+ hour.</p>\n<pre><code>mean = np.nanmean(signal, )[, :, :]\nstd = np.nanstd(signal, )[, :, :]\nsignal[(signal &gt; mean +  * std) | (signal &lt; mean -  * std)] = np.nan\n</code></pre>\n<p>The above is the exact added code used in inference.</p>\n<p>Given features like <code>Rs</code> were included as well as minor feature engineering, like <code>Rs ** 2</code>, but they don't seem too important.</p>\n<h1>Signal Extraction</h1>\n<p>Phase extraction is done using extrema of gradient.</p>\n<ol>\n<li>The calibrated signal is divided by the mean of itself</li>\n<li>A simple moving average of period k is applied to the signal (k = 88 for airs and 655 for fgs1, found using optuna)</li>\n<li>A set of moving averages is then applied to the gradient of the signal, in which p,q = argmin,argmax are extracted</li>\n<li>p,q are then moved to the left,right until their gradients are within some interquartile range of all gradients so they cover all of the transit phase</li>\n</ol>\n<p>The features are then extracted via simple mean/min/max/range/min div edge (approx transit depth)/polyfit to 8th degree etc. of each of the phases. Additionally, the airs data is <code>array_split</code>'d along the frequency dimension (size 356) into 32 chunks. Each chunk then gets processed like above in the same way, and their features added to the list of features we have. This adds quite a bit of resolution and adds ~0.025 CV.</p>\n<h1>Feed into CNN</h1>\n<p>The CNN doesn't care about the phases; the input is simply the airs and fgs1 signal after the first moving average, split into 92 chunks across time with each chunk averaged. Additional steps might be adding phase signal into the CNN, as well as separating airs/fgs1 data into different subnets since they might not be time aligned, but ideas like these were not implemented</p>\n<p>There are many instances of CNN, as they are integrated into the RQ NN ensemble. To prevent overfitting, almost all layers of the CNN have the same weight and shape. No activation function other than max pool was used.</p>\n<h1>Ensemble of Rational Quadratic NN</h1>\n<h2>Preparing the features</h2>\n<p>One idea was borrowed from <a href=\"https://arxiv.org/pdf/2407.04491\" target=\"_blank\">https://arxiv.org/pdf/2407.04491</a>.</p>\n<ul>\n<li>I used same feature scaling but with $\\frac{x}{\\sqrt{18+x^2/9}}$</li>\n</ul>\n<h2>Training</h2>\n<p>The idea is deep ensembling where each instance in the ensemble is trained the exact same way, just with different weight initialization. We just straight up predict everything, the <code>pred</code> and <code>sigma</code> with our model, no sampling or quantile regression.</p>\n<ul>\n<li>Optimizer is adabelief with (b1=0.999, b2=0.9998)</li>\n<li>Cosine one cycle learning rate schedule with square root shaping</li>\n<li>36 ensembles, with 88 RQ clusters</li>\n</ul>\n<p>Intuitively, I chose soft RBF/RQ (they perform similarly, both adds ~0.15 CV) because there were many transit curve shapes (it was mentioned that many different models of limb darkening were added). So this soft RQ aims to first sort instances into soft groups, and then perform simple linear regression within. The image shows the distribution of the first RQ cluster of the first instance.</p>\n<p>The training is quite unique; instead of training on average of losses (competition metric), I spent more effort training models which are performing worse. Similar to evolution, the best models don't need to get better anymore. This avoids overfitting and improves my CV by ~0.03+.</p>\n<pre><code>m = jnp.median(losses, ) \nw = jax.nn.relu(losses - m)\nw = w / w.(, keepdims=) \nw = jax.lax.stop_gradient(w)\n (losses * w).() / w.()\n</code></pre>\n<p>The above is the exact code used in training. Only values worse than the median will get trained.</p>\n<h1>What didn't work &amp; Conclusion</h1>\n<ul>\n<li>L1, L2 regularisation.</li>\n<li>Dropout (works well in very specific circumstances; unstable).</li>\n<li>Data augmentation via random splicing of transit data (although I saw there were some instances where the phase is cut off).</li>\n<li>Fischer-information-like or entropy based regularisation for RBF/RQ, to encourage more uncertainty of groups</li>\n<li>Full trainable Mahalanobis distance</li>\n</ul>\n<p>I used both observations for the same planet if they exist during training, but only the first observation during inference. The NN inference is instant and the data reading takes up the majority of the time.</p>\n<p>In conclusion an acceptable result was achieved with little or no understanding of the underlying physics nor advanced drift/transit window denoising techniques.</p>",
      "rawMarkdown": "# Introduction & Summary\n\nHuge thank you to the organizers as well as the staff for this amazing competition! I learnt a lot about data engineering and modelling techniques and I look forward to competing in the future!\n\nMy approach contain those steps:\n\n- Step 1: Preprocessing & Calibration\n- Step 2: Signal extraction\n- Step 3: Feed into CNN for additional feature extraction\n- Step 4: With signal features and CNN features, feed into ensemble of rational quadratic neural network\n\n# Preprocessing & Calibration\n\nI used the standard calibration techniques as given in the competition itself. It includes the calibration for gain/offset correction, and manages non-linearity, dark, dead and hot pixels. Additionally, at the end of the calibration step, I remove signals that are more than $8 \\sigma$ away in the time dimension. This addition alone adds ~0.05 to my CV (which is just a simple 80-20 split, turning out to be almost same as LB; LB feedback time is 5-8 hours). I couldn't spend much time on this section as rerunning this for all data, even with joblib, takes 1+ hour.\n\n```python\nmean = np.nanmean(signal, 0)[None, :, :]\nstd = np.nanstd(signal, 0)[None, :, :]\nsignal[(signal > mean + 8.0 * std) | (signal < mean - 8.0 * std)] = np.nan\n```\n\nThe above is the exact added code used in inference.\n\nGiven features like `Rs` were included as well as minor feature engineering, like `Rs ** 2`, but they don't seem too important.\n\n# Signal Extraction\n\nPhase extraction is done using extrema of gradient.\n\n1. The calibrated signal is divided by the mean of itself\n1. A simple moving average of period k is applied to the signal (k = 88 for airs and 655 for fgs1, found using optuna)\n2. A set of moving averages is then applied to the gradient of the signal, in which p,q = argmin,argmax are extracted\n3. p,q are then moved to the left,right until their gradients are within some interquartile range of all gradients so they cover all of the transit phase\n\nThe features are then extracted via simple mean/min/max/range/min div edge (approx transit depth)/polyfit to 8th degree etc. of each of the phases. Additionally, the airs data is `array_split`'d along the frequency dimension (size 356) into 32 chunks. Each chunk then gets processed like above in the same way, and their features added to the list of features we have. This adds quite a bit of resolution and adds ~0.025 CV.\n\n# Feed into CNN\n\nThe CNN doesn't care about the phases; the input is simply the airs and fgs1 signal after the first moving average, split into 92 chunks across time with each chunk averaged. Additional steps might be adding phase signal into the CNN, as well as separating airs/fgs1 data into different subnets since they might not be time aligned, but ideas like these were not implemented\n\nThere are many instances of CNN, as they are integrated into the RQ NN ensemble. To prevent overfitting, almost all layers of the CNN have the same weight and shape. No activation function other than max pool was used.\n\n# Ensemble of Rational Quadratic NN\n\n## Preparing the features\n\nOne idea was borrowed from https://arxiv.org/pdf/2407.04491.\n\n- I used same feature scaling but with $\\frac{x}{\\sqrt{18+x^2/9}}$\n\n## Training\n\nThe idea is deep ensembling where each instance in the ensemble is trained the exact same way, just with different weight initialization. We just straight up predict everything, the `pred` and `sigma` with our model, no sampling or quantile regression.\n\n- Optimizer is adabelief with (b1=0.999, b2=0.9998)\n- Cosine one cycle learning rate schedule with square root shaping\n- 36 ensembles, with 88 RQ clusters\n\nIntuitively, I chose soft RBF/RQ (they perform similarly, both adds ~0.15 CV) because there were many transit curve shapes (it was mentioned that many different models of limb darkening were added). So this soft RQ aims to first sort instances into soft groups, and then perform simple linear regression within. The image shows the distribution of the first RQ cluster of the first instance.\n\nThe training is quite unique; instead of training on average of losses (competition metric), I spent more effort training models which are performing worse. Similar to evolution, the best models don't need to get better anymore. This avoids overfitting and improves my CV by ~0.03+.\n\n```python\nm = jnp.median(losses, 0) # 0 is ensemble dimension\nw = jax.nn.relu(losses - m)\nw = w / w.sum(0, keepdims=True) # scale so every channel is treated equally\nw = jax.lax.stop_gradient(w)\nreturn (losses * w).sum() / w.sum()\n```\n\nThe above is the exact code used in training. Only values worse than the median will get trained.\n\n# What didn't work & Conclusion\n\n- L1, L2 regularisation.\n- Dropout (works well in very specific circumstances; unstable).\n- Data augmentation via random splicing of transit data (although I saw there were some instances where the phase is cut off).\n- Fischer-information-like or entropy based regularisation for RBF/RQ, to encourage more uncertainty of groups\n- Full trainable Mahalanobis distance\n\nI used both observations for the same planet if they exist during training, but only the first observation during inference. The NN inference is instant and the data reading takes up the majority of the time.\n\nIn conclusion an acceptable result was achieved with little or no understanding of the underlying physics nor advanced drift/transit window denoising techniques.",
      "votes": 32
    },
    {
      "id": 3295400,
      "postDate": "2025-09-28T16:44:41.277Z",
      "content": "<p>Thank you for sharing this it will help most of the people here to grow!</p>",
      "rawMarkdown": "Thank you for sharing this it will help most of the people here to grow!"
    },
    {
      "id": 3294898,
      "postDate": "2025-09-27T04:26:36.377Z",
      "content": "<p>Congrats! It looks solid, totally over my head 🤣</p>",
      "rawMarkdown": "Congrats! It looks solid, totally over my head 🤣"
    },
    {
      "id": 3294420,
      "postDate": "2025-09-26T06:17:43.967Z",
      "content": "<p>Interesting approach, congrats! A few questions:<br>\nDid you perform any binning during calibration or prior to the moving average?<br>\nHow did you input phase information (p, q) into the RQ networks?</p>",
      "rawMarkdown": "Interesting approach, congrats! A few questions:\nDid you perform any binning during calibration or prior to the moving average?\nHow did you input phase information (p, q) into the RQ networks?",
      "replies": [
        {
          "id": 3294424,
          "postDate": "2025-09-26T06:35:07.567Z",
          "content": "<p>Thanks! I just made my verbatim submission public, you can find all details there (except for training pipeline): <a href=\"https://www.kaggle.com/code/zayyyy/ariel-25-3rd-place-submission\" target=\"_blank\">https://www.kaggle.com/code/zayyyy/ariel-25-3rd-place-submission</a></p>\n<p>In particular,</p>\n<ul>\n<li>No binning was done at any point.</li>\n<li>Features were made from simple calculations like <code>signal[p:q].min()</code> or <code>signal[p:q].max() - signal[p:q].min()</code> which convey approx transit depth as well as shape information.</li>\n</ul>",
          "rawMarkdown": "Thanks! I just made my verbatim submission public, you can find all details there (except for training pipeline): https://www.kaggle.com/code/zayyyy/ariel-25-3rd-place-submission\n\nIn particular,\n- No binning was done at any point.\n- Features were made from simple calculations like `signal[p:q].min()` or `signal[p:q].max() - signal[p:q].min()` which convey approx transit depth as well as shape information.",
          "replies": [
            {
              "id": 3294435,
              "postDate": "2025-09-26T07:11:38.107Z",
              "content": "<p>Wow no binning! I wanted to work on raw calibrated data too but quickly realized I do not have close to enough ram.<br>\nThanks for the clarification!</p>",
              "rawMarkdown": "Wow no binning! I wanted to work on raw calibrated data too but quickly realized I do not have close to enough ram.\nThanks for the clarification!"
            }
          ]
        }
      ]
    },
    {
      "id": 3294403,
      "postDate": "2025-09-26T04:55:22.977Z",
      "content": "<p>Bravo <a href=\"https://www.kaggle.com/zayyyy\" target=\"_blank\">@zayyyy</a> </p>",
      "rawMarkdown": "Bravo @zayyyy "
    }
  ],
  "comments": [
    {
      "id": 3295400,
      "author_name": "Irakoze Ntawigenga Kelly",
      "author_url": "",
      "post_date": "2025-09-28T16:44:41.277000",
      "content": "<p>Thank you for sharing this it will help most of the people here to grow!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3294898,
      "author_name": "flyingfox",
      "author_url": "",
      "post_date": "2025-09-27T04:26:36.377000",
      "content": "<p>Congrats! It looks solid, totally over my head 🤣</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3294420,
      "author_name": "sroger",
      "author_url": "",
      "post_date": "2025-09-26T06:17:43.967000",
      "content": "<p>Interesting approach, congrats! A few questions:<br>\nDid you perform any binning during calibration or prior to the moving average?<br>\nHow did you input phase information (p, q) into the RQ networks?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3294424,
          "author_name": "moonpole",
          "author_url": "",
          "post_date": "2025-09-26T06:35:07.567000",
          "content": "<p>Thanks! I just made my verbatim submission public, you can find all details there (except for training pipeline): <a href=\"https://www.kaggle.com/code/zayyyy/ariel-25-3rd-place-submission\" target=\"_blank\">https://www.kaggle.com/code/zayyyy/ariel-25-3rd-place-submission</a></p>\n<p>In particular,</p>\n<ul>\n<li>No binning was done at any point.</li>\n<li>Features were made from simple calculations like <code>signal[p:q].min()</code> or <code>signal[p:q].max() - signal[p:q].min()</code> which convey approx transit depth as well as shape information.</li>\n</ul>",
          "votes": 0,
          "replies": [
            {
              "id": 3294435,
              "author_name": "sroger",
              "author_url": "",
              "post_date": "2025-09-26T07:11:38.107000",
              "content": "<p>Wow no binning! I wanted to work on raw calibrated data too but quickly realized I do not have close to enough ram.<br>\nThanks for the clarification!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3294403,
      "author_name": "Navneet",
      "author_url": "",
      "post_date": "2025-09-26T04:55:22.977000",
      "content": "<p>Bravo <a href=\"https://www.kaggle.com/zayyyy\" target=\"_blank\">@zayyyy</a> </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3294072": "# Introduction & Summary\n\nHuge thank you to the organizers as well as the staff for this amazing competition! I learnt a lot about data engineering and modelling techniques and I look forward to competing in the future!\n\nMy approach contain those steps:\n\n- Step 1: Preprocessing & Calibration\n- Step 2: Signal extraction\n- Step 3: Feed into CNN for additional feature extraction\n- Step 4: With signal features and CNN features, feed into ensemble of rational quadratic neural network\n\n# Preprocessing & Calibration\n\nI used the standard calibration techniques as given in the competition itself. It includes the calibration for gain/offset correction, and manages non-linearity, dark, dead and hot pixels. Additionally, at the end of the calibration step, I remove signals that are more than $8 \\sigma$ away in the time dimension. This addition alone adds ~0.05 to my CV (which is just a simple 80-20 split, turning out to be almost same as LB; LB feedback time is 5-8 hours). I couldn't spend much time on this section as rerunning this for all data, even with joblib, takes 1+ hour.\n\n```python\nmean = np.nanmean(signal, 0)[None, :, :]\nstd = np.nanstd(signal, 0)[None, :, :]\nsignal[(signal > mean + 8.0 * std) | (signal < mean - 8.0 * std)] = np.nan\n```\n\nThe above is the exact added code used in inference.\n\nGiven features like `Rs` were included as well as minor feature engineering, like `Rs ** 2`, but they don't seem too important.\n\n# Signal Extraction\n\nPhase extraction is done using extrema of gradient.\n\n1. The calibrated signal is divided by the mean of itself\n1. A simple moving average of period k is applied to the signal (k = 88 for airs and 655 for fgs1, found using optuna)\n2. A set of moving averages is then applied to the gradient of the signal, in which p,q = argmin,argmax are extracted\n3. p,q are then moved to the left,right until their gradients are within some interquartile range of all gradients so they cover all of the transit phase\n\nThe features are then extracted via simple mean/min/max/range/min div edge (approx transit depth)/polyfit to 8th degree etc. of each of the phases. Additionally, the airs data is `array_split`'d along the frequency dimension (size 356) into 32 chunks. Each chunk then gets processed like above in the same way, and their features added to the list of features we have. This adds quite a bit of resolution and adds ~0.025 CV.\n\n# Feed into CNN\n\nThe CNN doesn't care about the phases; the input is simply the airs and fgs1 signal after the first moving average, split into 92 chunks across time with each chunk averaged. Additional steps might be adding phase signal into the CNN, as well as separating airs/fgs1 data into different subnets since they might not be time aligned, but ideas like these were not implemented\n\nThere are many instances of CNN, as they are integrated into the RQ NN ensemble. To prevent overfitting, almost all layers of the CNN have the same weight and shape. No activation function other than max pool was used.\n\n# Ensemble of Rational Quadratic NN\n\n## Preparing the features\n\nOne idea was borrowed from https://arxiv.org/pdf/2407.04491.\n\n- I used same feature scaling but with $\\frac{x}{\\sqrt{18+x^2/9}}$\n\n## Training\n\nThe idea is deep ensembling where each instance in the ensemble is trained the exact same way, just with different weight initialization. We just straight up predict everything, the `pred` and `sigma` with our model, no sampling or quantile regression.\n\n- Optimizer is adabelief with (b1=0.999, b2=0.9998)\n- Cosine one cycle learning rate schedule with square root shaping\n- 36 ensembles, with 88 RQ clusters\n\nIntuitively, I chose soft RBF/RQ (they perform similarly, both adds ~0.15 CV) because there were many transit curve shapes (it was mentioned that many different models of limb darkening were added). So this soft RQ aims to first sort instances into soft groups, and then perform simple linear regression within. The image shows the distribution of the first RQ cluster of the first instance.\n\nThe training is quite unique; instead of training on average of losses (competition metric), I spent more effort training models which are performing worse. Similar to evolution, the best models don't need to get better anymore. This avoids overfitting and improves my CV by ~0.03+.\n\n```python\nm = jnp.median(losses, 0) # 0 is ensemble dimension\nw = jax.nn.relu(losses - m)\nw = w / w.sum(0, keepdims=True) # scale so every channel is treated equally\nw = jax.lax.stop_gradient(w)\nreturn (losses * w).sum() / w.sum()\n```\n\nThe above is the exact code used in training. Only values worse than the median will get trained.\n\n# What didn't work & Conclusion\n\n- L1, L2 regularisation.\n- Dropout (works well in very specific circumstances; unstable).\n- Data augmentation via random splicing of transit data (although I saw there were some instances where the phase is cut off).\n- Fischer-information-like or entropy based regularisation for RBF/RQ, to encourage more uncertainty of groups\n- Full trainable Mahalanobis distance\n\nI used both observations for the same planet if they exist during training, but only the first observation during inference. The NN inference is instant and the data reading takes up the majority of the time.\n\nIn conclusion an acceptable result was achieved with little or no understanding of the underlying physics nor advanced drift/transit window denoising techniques.",
    "3295400": "Thank you for sharing this it will help most of the people here to grow!",
    "3294898": "Congrats! It looks solid, totally over my head 🤣",
    "3294420": "Interesting approach, congrats! A few questions:\nDid you perform any binning during calibration or prior to the moving average?\nHow did you input phase information (p, q) into the RQ networks?",
    "3294403": "Bravo @zayyyy "
  }
}