{
  "id": 609629,
  "title": "5th place solution",
  "url": "/competitions/ariel-data-challenge-2025/discussion/609629",
  "author_name": "Alehandreus",
  "post_date": "2025-09-28T12:30:05.804000",
  "votes": 18,
  "comment_count": 0,
  "views": 0,
  "content": "<p>This competition was a great opportunity to take a break from deep learning and do some modeling. Huge thanks to Kaggle, <a href=\"https://www.kaggle.com/gordonyip\" target=\"_blank\">@gordonyip</a> and all the other organizers, looking forward to the 2026 edition!</p>\n<p>At the base of my solution is a physical model with parameters fit to match the noisy data. Next, I apply some post-processing to refine the results, most importantly gradient boosting.</p>\n<hr>\n<h1>Pipeline</h1>\n<ol>\n<li>Calibrate the data, remove outliers, do 5x time binning (T=1125) and 6x (W=47) spectrum binning;</li>\n<li>Fit the physical model with regularized Gauss-Newton algorithm;</li>\n<li>Refine the outputs with PCA and Boosting.</li>\n</ol>\n<hr>\n<h1>Physical Model</h1>\n<p>The model is similar to what was proposed in Ariel 2024. However, transit depth is much more complex now.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F7084217%2F87d2c3e5f8425c8052e077de83c239d4%2Fmodel2.png?generation=1759001064021533&amp;alt=media\" alt=\"\"></p>\n<p><strong>1. Star Spectrum</strong> is calculated as <code>sensor_data.mean(axis=0) / (transit_depth * poly_drift).mean(axis=0)</code>. No parameters to optimize specifically for star spectrum.</p>\n<p><strong>2. Transit Depth</strong> is the main part of the model. It takes in a bunch of parameters and outputs how much light remains for each wavelength at each moment of time. Basically, it is the same model as described in <a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a> <a href=\"https://www.kaggle.com/code/junkoda/limb-darkening\" target=\"_blank\">notebook</a>: </p>\n<p>$$\\textbf{transit depth}(t, w)= 1 - \\int_{D_{tw}} F(x)dx.$$</p>\n<p>Here D is an intersection of stellar and planetary discs at each time moment t at each wavelength w. F is the star intensity according to the limb darkening law. I used nonlinear law with six coefficients covering x^0.5 to x^3. The numerical integration was also taken from <a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a>`s notebook (it <em>was</em> necessary after all :D ). Here is the full parameter list:</p>\n<ul>\n<li><strong>Rp mean</strong> — scalar;</li>\n<li><strong>Rp variation</strong> — 15 values linearly interpolated to W values;</li>\n<li>Impact factor <strong>b</strong> — scalar in [0, 1]. Actually can be larger than 1 but it didn't help anyway;</li>\n<li><strong>limb coeffs</strong> — six coefficients for nonlinear limb darkening model;</li>\n<li><strong>ingress</strong>, <strong>egress</strong>  — two scalars for start and end of the transit window.</li>\n</ul>\n<p>I also experimented with adding <strong>orbit radius</strong> to take trajectory curvature into account. It produced better results when testing on samples manually generated with batman, but didn't work well on the training data.</p>\n<p>Notably, the transit dip is smooth and has length varying over spectrum. A planet with the biggest Rp range illustrates this well:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F7084217%2Fea8572272872ba6535ae928d3302ca0f%2Ffit3.png?generation=1759002005234323&amp;alt=media\" alt=\"\"></p>\n<p><strong>3. Polynomial Drift</strong> is 1 + f(t, w) where f is a bivariate polynomial with degree 4 in time dimension and degree 2 in wavelength dimension, resulting in 15 parameters.</p>\n<hr>\n<h1>Fitting and Implementation</h1>\n<p><strong>Optimization</strong>. The output of the model is differentiable w.r.t. all the parameters described above. I assume the noise to be gaussian and use MSE loss with wavelengths weighted according to noise deviations. I ended up using Levenberg-Marquardt method for fitting (and should thank <a href=\"https://www.kaggle.com/jeroencottaar\" target=\"_blank\">@jeroencottaar</a> whose solution of GWI made me do some digging on optimization).</p>\n<p>Optimization runs for 220 LM iterations and mainly fits all parameters simultaneously in one stage. The only exceptions are that polynomial drift is locked before 70th iteration and Rp variation is locked before 180th iteration.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F7084217%2F9924c6f3277b4d17c04f63a91652dc03%2Fanimation.gif?generation=1759054914837126&amp;alt=media\" alt=\"\"></p>\n<p><strong>Sigma prediction</strong> is constant over spectrum and is calculated simply as <code>mu.std() * 1.6</code>.</p>\n<p><strong>Cut transits</strong>. Since ingress / egress were fitted automatically, I did not have to add any special treatment for planets with one edge. </p>\n<p><strong>Performance</strong>. Fitting time for one planet depends on the transit width and usually takes around 20-25 seconds on P100 (FP64). All logic was implemented in PyTorch with autograd. The transit depth integral was estimated with 40 rings.</p>\n<p><strong>AIRS / FGS channels</strong>. Applying the model to AIRS only (and setting FGS mu as mean AIRS mu) would yield ~0.542 on Public LB. Fitting FGS too  improved the score to 0.546.</p>\n<hr>\n<h1>Refining the outputs</h1>\n<p><strong>Linear trend and instability</strong>. Somehow my method fits Rp with additional linear trend and unstable values at the right edge of the spectrum. I haven't found a better way than to simply remove this trend as a post-processing step and set <code>Rp[-50:] = Rp[-50]</code>. Here are the raw outputs for the very first planet:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F7084217%2Fadd9d8073a30782d1ab8b143527b6762%2Ffit6.png?generation=1759055112773971&amp;alt=media\" alt=\"\"></p>\n<p><strong>Planets with high impact factor</strong>. For planets with b &gt; 0.75 my model fitted mu to values ~4-6% larger than ground truth. This is another mystery I failed to solve. I added a special processing for these planets:</p>\n<pre><code> model.get_b() &gt; .:\n     = mu / .\n     = sigma * .\n</code></pre>\n<p><strong>Applying PCA</strong> on predicted Rp variation improved Public LB score by 0.007. I used 3 components.</p>\n<p><strong>Gradient Boosting</strong> worked particularly good for me and improved my score by 0.028 during the last week. I used it to refine mu and sigma for both AIRS and FGS:</p>\n<ul>\n<li>Input features are <code>star_info.csv</code> data and all fitted parameters;</li>\n<li>Outputs are constant shifts (for example, <code>mu_airs</code> becomes <code>mu_airs + dmu_airs</code> where <code>dmu_airs</code> is scalar boosting output).</li>\n</ul>\n<hr>\n<h1>Scores</h1>\n<p>The first row shows the score with trend and unstable values already removed.</p>\n<table>\n<thead>\n<tr>\n<th>Submission</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Fit model on AIRS only</td>\n<td>0.542</td>\n<td>0.550</td>\n</tr>\n<tr>\n<td>Fit model on FGS too</td>\n<td>0.546</td>\n<td>0.554</td>\n</tr>\n<tr>\n<td>Add PCA</td>\n<td>0.553</td>\n<td>0.557</td>\n</tr>\n<tr>\n<td>Add AIRS mu &amp; sigma boosting</td>\n<td>0.577</td>\n<td>0.586</td>\n</tr>\n<tr>\n<td>Add FGS mu &amp; sigma boosting</td>\n<td>0.581</td>\n<td>0.589</td>\n</tr>\n</tbody>\n</table>\n<h1>Code</h1>\n<p>GitHub: <a href=\"https://github.com/Alehandreus/ariel-2025\" target=\"_blank\">github.com/Alehandreus/ariel-2025</a><br>\nKaggle notebook: <a href=\"https://www.kaggle.com/code/alehandreus/ariel-2025-inference\" target=\"_blank\">kaggle.com/code/alehandreus/ariel-2025-inference</a></p>",
  "messages": [
    {
      "id": 3295325,
      "postDate": "2025-09-28T12:30:05.803Z",
      "content": "<p>This competition was a great opportunity to take a break from deep learning and do some modeling. Huge thanks to Kaggle, <a href=\"https://www.kaggle.com/gordonyip\" target=\"_blank\">@gordonyip</a> and all the other organizers, looking forward to the 2026 edition!</p>\n<p>At the base of my solution is a physical model with parameters fit to match the noisy data. Next, I apply some post-processing to refine the results, most importantly gradient boosting.</p>\n<hr>\n<h1>Pipeline</h1>\n<ol>\n<li>Calibrate the data, remove outliers, do 5x time binning (T=1125) and 6x (W=47) spectrum binning;</li>\n<li>Fit the physical model with regularized Gauss-Newton algorithm;</li>\n<li>Refine the outputs with PCA and Boosting.</li>\n</ol>\n<hr>\n<h1>Physical Model</h1>\n<p>The model is similar to what was proposed in Ariel 2024. However, transit depth is much more complex now.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F7084217%2F87d2c3e5f8425c8052e077de83c239d4%2Fmodel2.png?generation=1759001064021533&amp;alt=media\" alt=\"\"></p>\n<p><strong>1. Star Spectrum</strong> is calculated as <code>sensor_data.mean(axis=0) / (transit_depth * poly_drift).mean(axis=0)</code>. No parameters to optimize specifically for star spectrum.</p>\n<p><strong>2. Transit Depth</strong> is the main part of the model. It takes in a bunch of parameters and outputs how much light remains for each wavelength at each moment of time. Basically, it is the same model as described in <a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a> <a href=\"https://www.kaggle.com/code/junkoda/limb-darkening\" target=\"_blank\">notebook</a>: </p>\n<p>$$\\textbf{transit depth}(t, w)= 1 - \\int_{D_{tw}} F(x)dx.$$</p>\n<p>Here D is an intersection of stellar and planetary discs at each time moment t at each wavelength w. F is the star intensity according to the limb darkening law. I used nonlinear law with six coefficients covering x^0.5 to x^3. The numerical integration was also taken from <a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a>`s notebook (it <em>was</em> necessary after all :D ). Here is the full parameter list:</p>\n<ul>\n<li><strong>Rp mean</strong> — scalar;</li>\n<li><strong>Rp variation</strong> — 15 values linearly interpolated to W values;</li>\n<li>Impact factor <strong>b</strong> — scalar in [0, 1]. Actually can be larger than 1 but it didn't help anyway;</li>\n<li><strong>limb coeffs</strong> — six coefficients for nonlinear limb darkening model;</li>\n<li><strong>ingress</strong>, <strong>egress</strong>  — two scalars for start and end of the transit window.</li>\n</ul>\n<p>I also experimented with adding <strong>orbit radius</strong> to take trajectory curvature into account. It produced better results when testing on samples manually generated with batman, but didn't work well on the training data.</p>\n<p>Notably, the transit dip is smooth and has length varying over spectrum. A planet with the biggest Rp range illustrates this well:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F7084217%2Fea8572272872ba6535ae928d3302ca0f%2Ffit3.png?generation=1759002005234323&amp;alt=media\" alt=\"\"></p>\n<p><strong>3. Polynomial Drift</strong> is 1 + f(t, w) where f is a bivariate polynomial with degree 4 in time dimension and degree 2 in wavelength dimension, resulting in 15 parameters.</p>\n<hr>\n<h1>Fitting and Implementation</h1>\n<p><strong>Optimization</strong>. The output of the model is differentiable w.r.t. all the parameters described above. I assume the noise to be gaussian and use MSE loss with wavelengths weighted according to noise deviations. I ended up using Levenberg-Marquardt method for fitting (and should thank <a href=\"https://www.kaggle.com/jeroencottaar\" target=\"_blank\">@jeroencottaar</a> whose solution of GWI made me do some digging on optimization).</p>\n<p>Optimization runs for 220 LM iterations and mainly fits all parameters simultaneously in one stage. The only exceptions are that polynomial drift is locked before 70th iteration and Rp variation is locked before 180th iteration.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F7084217%2F9924c6f3277b4d17c04f63a91652dc03%2Fanimation.gif?generation=1759054914837126&amp;alt=media\" alt=\"\"></p>\n<p><strong>Sigma prediction</strong> is constant over spectrum and is calculated simply as <code>mu.std() * 1.6</code>.</p>\n<p><strong>Cut transits</strong>. Since ingress / egress were fitted automatically, I did not have to add any special treatment for planets with one edge. </p>\n<p><strong>Performance</strong>. Fitting time for one planet depends on the transit width and usually takes around 20-25 seconds on P100 (FP64). All logic was implemented in PyTorch with autograd. The transit depth integral was estimated with 40 rings.</p>\n<p><strong>AIRS / FGS channels</strong>. Applying the model to AIRS only (and setting FGS mu as mean AIRS mu) would yield ~0.542 on Public LB. Fitting FGS too  improved the score to 0.546.</p>\n<hr>\n<h1>Refining the outputs</h1>\n<p><strong>Linear trend and instability</strong>. Somehow my method fits Rp with additional linear trend and unstable values at the right edge of the spectrum. I haven't found a better way than to simply remove this trend as a post-processing step and set <code>Rp[-50:] = Rp[-50]</code>. Here are the raw outputs for the very first planet:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F7084217%2Fadd9d8073a30782d1ab8b143527b6762%2Ffit6.png?generation=1759055112773971&amp;alt=media\" alt=\"\"></p>\n<p><strong>Planets with high impact factor</strong>. For planets with b &gt; 0.75 my model fitted mu to values ~4-6% larger than ground truth. This is another mystery I failed to solve. I added a special processing for these planets:</p>\n<pre><code> model.get_b() &gt; .:\n     = mu / .\n     = sigma * .\n</code></pre>\n<p><strong>Applying PCA</strong> on predicted Rp variation improved Public LB score by 0.007. I used 3 components.</p>\n<p><strong>Gradient Boosting</strong> worked particularly good for me and improved my score by 0.028 during the last week. I used it to refine mu and sigma for both AIRS and FGS:</p>\n<ul>\n<li>Input features are <code>star_info.csv</code> data and all fitted parameters;</li>\n<li>Outputs are constant shifts (for example, <code>mu_airs</code> becomes <code>mu_airs + dmu_airs</code> where <code>dmu_airs</code> is scalar boosting output).</li>\n</ul>\n<hr>\n<h1>Scores</h1>\n<p>The first row shows the score with trend and unstable values already removed.</p>\n<table>\n<thead>\n<tr>\n<th>Submission</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Fit model on AIRS only</td>\n<td>0.542</td>\n<td>0.550</td>\n</tr>\n<tr>\n<td>Fit model on FGS too</td>\n<td>0.546</td>\n<td>0.554</td>\n</tr>\n<tr>\n<td>Add PCA</td>\n<td>0.553</td>\n<td>0.557</td>\n</tr>\n<tr>\n<td>Add AIRS mu &amp; sigma boosting</td>\n<td>0.577</td>\n<td>0.586</td>\n</tr>\n<tr>\n<td>Add FGS mu &amp; sigma boosting</td>\n<td>0.581</td>\n<td>0.589</td>\n</tr>\n</tbody>\n</table>\n<h1>Code</h1>\n<p>GitHub: <a href=\"https://github.com/Alehandreus/ariel-2025\" target=\"_blank\">github.com/Alehandreus/ariel-2025</a><br>\nKaggle notebook: <a href=\"https://www.kaggle.com/code/alehandreus/ariel-2025-inference\" target=\"_blank\">kaggle.com/code/alehandreus/ariel-2025-inference</a></p>",
      "rawMarkdown": "This competition was a great opportunity to take a break from deep learning and do some modeling. Huge thanks to Kaggle, @gordonyip and all the other organizers, looking forward to the 2026 edition!\n\nAt the base of my solution is a physical model with parameters fit to match the noisy data. Next, I apply some post-processing to refine the results, most importantly gradient boosting.\n\n<hr>\n\n# Pipeline\n1. Calibrate the data, remove outliers, do 5x time binning (T=1125) and 6x (W=47) spectrum binning;\n2. Fit the physical model with regularized Gauss-Newton algorithm;\n3. Refine the outputs with PCA and Boosting.\n\n<hr>\n\n# Physical Model\nThe model is similar to what was proposed in Ariel 2024. However, transit depth is much more complex now.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F7084217%2F87d2c3e5f8425c8052e077de83c239d4%2Fmodel2.png?generation=1759001064021533&alt=media)\n\n**1. Star Spectrum** is calculated as `sensor_data.mean(axis=0) / (transit_depth * poly_drift).mean(axis=0)`. No parameters to optimize specifically for star spectrum.\n\n**2. Transit Depth** is the main part of the model. It takes in a bunch of parameters and outputs how much light remains for each wavelength at each moment of time. Basically, it is the same model as described in @junkoda [notebook](https://www.kaggle.com/code/junkoda/limb-darkening): \n\n$$\\textbf{transit depth}(t, w)= 1 - \\int_{D_{tw}} F(x)dx.$$\n\nHere D is an intersection of stellar and planetary discs at each time moment t at each wavelength w. F is the star intensity according to the limb darkening law. I used nonlinear law with six coefficients covering x^0.5 to x^3. The numerical integration was also taken from @junkoda`s notebook (it *was* necessary after all :D ). Here is the full parameter list:\n- **Rp mean** — scalar;\n- **Rp variation** — 15 values linearly interpolated to W values;\n- Impact factor **b** — scalar in [0, 1]. Actually can be larger than 1 but it didn't help anyway;\n- **limb coeffs** — six coefficients for nonlinear limb darkening model;\n- **ingress**, **egress**  — two scalars for start and end of the transit window.\n\nI also experimented with adding **orbit radius** to take trajectory curvature into account. It produced better results when testing on samples manually generated with batman, but didn't work well on the training data.\n\nNotably, the transit dip is smooth and has length varying over spectrum. A planet with the biggest Rp range illustrates this well:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F7084217%2Fea8572272872ba6535ae928d3302ca0f%2Ffit3.png?generation=1759002005234323&alt=media)\n\n**3. Polynomial Drift** is 1 + f(t, w) where f is a bivariate polynomial with degree 4 in time dimension and degree 2 in wavelength dimension, resulting in 15 parameters.\n\n<hr>\n\n# Fitting and Implementation\n\n**Optimization**. The output of the model is differentiable w.r.t. all the parameters described above. I assume the noise to be gaussian and use MSE loss with wavelengths weighted according to noise deviations. I ended up using Levenberg-Marquardt method for fitting (and should thank @jeroencottaar whose solution of GWI made me do some digging on optimization).\n\nOptimization runs for 220 LM iterations and mainly fits all parameters simultaneously in one stage. The only exceptions are that polynomial drift is locked before 70th iteration and Rp variation is locked before 180th iteration.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F7084217%2F9924c6f3277b4d17c04f63a91652dc03%2Fanimation.gif?generation=1759054914837126&alt=media)\n\n**Sigma prediction** is constant over spectrum and is calculated simply as `mu.std() * 1.6`.\n\n**Cut transits**. Since ingress / egress were fitted automatically, I did not have to add any special treatment for planets with one edge. \n\n**Performance**. Fitting time for one planet depends on the transit width and usually takes around 20-25 seconds on P100 (FP64). All logic was implemented in PyTorch with autograd. The transit depth integral was estimated with 40 rings.\n\n**AIRS / FGS channels**. Applying the model to AIRS only (and setting FGS mu as mean AIRS mu) would yield ~0.542 on Public LB. Fitting FGS too  improved the score to 0.546.\n\n<hr>\n\n# Refining the outputs\n\n**Linear trend and instability**. Somehow my method fits Rp with additional linear trend and unstable values at the right edge of the spectrum. I haven't found a better way than to simply remove this trend as a post-processing step and set `Rp[-50:] = Rp[-50]`. Here are the raw outputs for the very first planet:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F7084217%2Fadd9d8073a30782d1ab8b143527b6762%2Ffit6.png?generation=1759055112773971&alt=media)\n\n**Planets with high impact factor**. For planets with b > 0.75 my model fitted mu to values ~4-6% larger than ground truth. This is another mystery I failed to solve. I added a special processing for these planets:\n```\nif model.get_b() > 0.75:\n    mu = mu / 1.05\n    sigma = sigma * 7.0\n``` \n\n**Applying PCA** on predicted Rp variation improved Public LB score by 0.007. I used 3 components.\n\n**Gradient Boosting** worked particularly good for me and improved my score by 0.028 during the last week. I used it to refine mu and sigma for both AIRS and FGS:\n- Input features are `star_info.csv` data and all fitted parameters;\n- Outputs are constant shifts (for example, `mu_airs` becomes `mu_airs + dmu_airs` where `dmu_airs` is scalar boosting output).\n\n<hr>\n\n# Scores\n\nThe first row shows the score with trend and unstable values already removed.\n\n| Submission | Public LB | Private LB |\n| --- | --- | --- |\n| Fit model on AIRS only | 0.542 | 0.550 |\n| Fit model on FGS too | 0.546 | 0.554 |\n| Add PCA | 0.553 | 0.557 |\n| Add AIRS mu & sigma boosting | 0.577 | 0.586 |\n| Add FGS mu & sigma boosting | 0.581 | 0.589 |\n\n# Code \n\nGitHub: [github.com/Alehandreus/ariel-2025](https://github.com/Alehandreus/ariel-2025)\nKaggle notebook: [kaggle.com/code/alehandreus/ariel-2025-inference](https://www.kaggle.com/code/alehandreus/ariel-2025-inference)",
      "votes": 18
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3295325": "This competition was a great opportunity to take a break from deep learning and do some modeling. Huge thanks to Kaggle, @gordonyip and all the other organizers, looking forward to the 2026 edition!\n\nAt the base of my solution is a physical model with parameters fit to match the noisy data. Next, I apply some post-processing to refine the results, most importantly gradient boosting.\n\n<hr>\n\n# Pipeline\n1. Calibrate the data, remove outliers, do 5x time binning (T=1125) and 6x (W=47) spectrum binning;\n2. Fit the physical model with regularized Gauss-Newton algorithm;\n3. Refine the outputs with PCA and Boosting.\n\n<hr>\n\n# Physical Model\nThe model is similar to what was proposed in Ariel 2024. However, transit depth is much more complex now.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F7084217%2F87d2c3e5f8425c8052e077de83c239d4%2Fmodel2.png?generation=1759001064021533&alt=media)\n\n**1. Star Spectrum** is calculated as `sensor_data.mean(axis=0) / (transit_depth * poly_drift).mean(axis=0)`. No parameters to optimize specifically for star spectrum.\n\n**2. Transit Depth** is the main part of the model. It takes in a bunch of parameters and outputs how much light remains for each wavelength at each moment of time. Basically, it is the same model as described in @junkoda [notebook](https://www.kaggle.com/code/junkoda/limb-darkening): \n\n$$\\textbf{transit depth}(t, w)= 1 - \\int_{D_{tw}} F(x)dx.$$\n\nHere D is an intersection of stellar and planetary discs at each time moment t at each wavelength w. F is the star intensity according to the limb darkening law. I used nonlinear law with six coefficients covering x^0.5 to x^3. The numerical integration was also taken from @junkoda`s notebook (it *was* necessary after all :D ). Here is the full parameter list:\n- **Rp mean** — scalar;\n- **Rp variation** — 15 values linearly interpolated to W values;\n- Impact factor **b** — scalar in [0, 1]. Actually can be larger than 1 but it didn't help anyway;\n- **limb coeffs** — six coefficients for nonlinear limb darkening model;\n- **ingress**, **egress**  — two scalars for start and end of the transit window.\n\nI also experimented with adding **orbit radius** to take trajectory curvature into account. It produced better results when testing on samples manually generated with batman, but didn't work well on the training data.\n\nNotably, the transit dip is smooth and has length varying over spectrum. A planet with the biggest Rp range illustrates this well:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F7084217%2Fea8572272872ba6535ae928d3302ca0f%2Ffit3.png?generation=1759002005234323&alt=media)\n\n**3. Polynomial Drift** is 1 + f(t, w) where f is a bivariate polynomial with degree 4 in time dimension and degree 2 in wavelength dimension, resulting in 15 parameters.\n\n<hr>\n\n# Fitting and Implementation\n\n**Optimization**. The output of the model is differentiable w.r.t. all the parameters described above. I assume the noise to be gaussian and use MSE loss with wavelengths weighted according to noise deviations. I ended up using Levenberg-Marquardt method for fitting (and should thank @jeroencottaar whose solution of GWI made me do some digging on optimization).\n\nOptimization runs for 220 LM iterations and mainly fits all parameters simultaneously in one stage. The only exceptions are that polynomial drift is locked before 70th iteration and Rp variation is locked before 180th iteration.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F7084217%2F9924c6f3277b4d17c04f63a91652dc03%2Fanimation.gif?generation=1759054914837126&alt=media)\n\n**Sigma prediction** is constant over spectrum and is calculated simply as `mu.std() * 1.6`.\n\n**Cut transits**. Since ingress / egress were fitted automatically, I did not have to add any special treatment for planets with one edge. \n\n**Performance**. Fitting time for one planet depends on the transit width and usually takes around 20-25 seconds on P100 (FP64). All logic was implemented in PyTorch with autograd. The transit depth integral was estimated with 40 rings.\n\n**AIRS / FGS channels**. Applying the model to AIRS only (and setting FGS mu as mean AIRS mu) would yield ~0.542 on Public LB. Fitting FGS too  improved the score to 0.546.\n\n<hr>\n\n# Refining the outputs\n\n**Linear trend and instability**. Somehow my method fits Rp with additional linear trend and unstable values at the right edge of the spectrum. I haven't found a better way than to simply remove this trend as a post-processing step and set `Rp[-50:] = Rp[-50]`. Here are the raw outputs for the very first planet:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F7084217%2Fadd9d8073a30782d1ab8b143527b6762%2Ffit6.png?generation=1759055112773971&alt=media)\n\n**Planets with high impact factor**. For planets with b > 0.75 my model fitted mu to values ~4-6% larger than ground truth. This is another mystery I failed to solve. I added a special processing for these planets:\n```\nif model.get_b() > 0.75:\n    mu = mu / 1.05\n    sigma = sigma * 7.0\n``` \n\n**Applying PCA** on predicted Rp variation improved Public LB score by 0.007. I used 3 components.\n\n**Gradient Boosting** worked particularly good for me and improved my score by 0.028 during the last week. I used it to refine mu and sigma for both AIRS and FGS:\n- Input features are `star_info.csv` data and all fitted parameters;\n- Outputs are constant shifts (for example, `mu_airs` becomes `mu_airs + dmu_airs` where `dmu_airs` is scalar boosting output).\n\n<hr>\n\n# Scores\n\nThe first row shows the score with trend and unstable values already removed.\n\n| Submission | Public LB | Private LB |\n| --- | --- | --- |\n| Fit model on AIRS only | 0.542 | 0.550 |\n| Fit model on FGS too | 0.546 | 0.554 |\n| Add PCA | 0.553 | 0.557 |\n| Add AIRS mu & sigma boosting | 0.577 | 0.586 |\n| Add FGS mu & sigma boosting | 0.581 | 0.589 |\n\n# Code \n\nGitHub: [github.com/Alehandreus/ariel-2025](https://github.com/Alehandreus/ariel-2025)\nKaggle notebook: [kaggle.com/code/alehandreus/ariel-2025-inference](https://www.kaggle.com/code/alehandreus/ariel-2025-inference)"
  }
}