{
  "id": 609888,
  "title": "1st place solution: Bayesian Inference, of course",
  "url": "/competitions/ariel-data-challenge-2025/discussion/609888",
  "author_name": "Jeroen Cottaar",
  "post_date": "2025-09-30T09:29:08.289000",
  "votes": 56,
  "comment_count": 11,
  "views": 0,
  "content": "<p><a href=\"https://www.kaggle.com/code/jeroencottaar/ariel2025-1st-place-bayesian-exoplanet-analysis\" target=\"_blank\">Link to code to train and submit</a><br>\nThis year I'm also releasing the full development history: <a href=\"https://github.com/jcottaar/ariel2\" target=\"_blank\">https://github.com/jcottaar/ariel2</a> <br><br>\nI'd like to start by thanking the organizers for another great competition. I had a rather frustrating start (having somehow managed to lose all my code from last year), but greatly enjoyed it overall. The Bayesian approach is not used nearly as often as it should be these days, and I'm glad to have the opportunity to show its value here.</p>\n<p>My solution is quite similar to last year's; my <a href=\"https://neurips.cc/virtual/2024/107111\" target=\"_blank\">NeurIPS talk</a> from the associated workshop can be a useful introduction. I'll briefly discuss the differences compared to last year in a reply to this post.</p>\n<h1>1. Introduction</h1>\n<p>Transit analysis is the most common method for detecting and studying exoplanets. A transit occurs when a planet crosses in front of its star relative to Earth, obscuring part of the starlight. By observing how much the light is reduced (the transit depth) as a function of wavelength, we can learn the properties of the planet. But with a tiny planet passing in front of a huge star, this problem has a very low signal to noise ratio. The challenge to us: find the transit depth, including confidence intervals, from synthetic raw spectroscopic signals as the Ariel satellite might see them in the future (For a more extensive overview, see <a href=\"https://www.kaggle.com/competitions/ariel-data-challenge-2025/overview)\" target=\"_blank\">https://www.kaggle.com/competitions/ariel-data-challenge-2025/overview)</a>.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14984949%2Fa0860242fa947c058ce0794fa0494cc4%2FScreenshot%202024-11-01%20200453.png?generation=1730488072791970&amp;alt=media\" alt=\"\"></p>\n<p>This competition is at first glance a retread of a similar competition last year, but the data is far more realistic this time around. I think I figured out some of the changes, but not all of them - as we'll see later on.<br><br>\nNote that there are two sensors in play, the single-wavelength FGS sensor and the spectroscopic AIRS sensor. In this writeup I'll focus mainly on the AIRS sensor.<br><br>\nMy general approach consists of three steps, as visualized below:</p>\n<ul>\n<li><strong>Preprocessing</strong>: starting from the raw sensor data (expressed in counts), compute the amplitude signal <em>S(λ,t)</em> as a function of wavelength <em>λ</em> and time <em>t</em>.</li>\n<li><strong>Bayesian Inference</strong>: from <em>S(λ,t)</em>,  compute the inferred transit depth <em>D(λ)</em> and an uncertainty margin on this <em>σ(λ)</em> . This step is done using a principled Bayesian approach.</li>\n<li><strong>Fudging</strong>: throw the principles out the window, and apply various fudge factors to <em>D(λ)</em> and <em>σ(λ)</em> . These fudge factors are chosen to optimize score on the training data.</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F14984949%2Fa0b169937d51e43db62a030def487f5c%2FScreenshot%202025-09-29%20212745.png?generation=1759174122624685&amp;alt=media\" alt=\"\"></p>\n<p>In sections 2, 3, and 4, I'll discuss each of these steps in turn. I'll finish with some closing thoughts in section 5 - mainly discussing why my approach isn't actually working…</p>\n<h1>2. Preprocessing</h1>\n<p>The raw data is in counts per pixel; we first need to apply several steps to come to a photon count per pixel. This is explained in <a href=\"https://www.kaggle.com/code/gordonyip/calibrating-and-binning-ariel-data\" target=\"_blank\">a notebook shared by the hosts</a>. I use a time binning of 5 frames for AIRS, and 50 frames for FGS (this puts them approximately in sync). Hot pixel invalidation is disabled. I also added a simple cosmic ray removal. <br><br>\nThe opportunity for some real improvement lies in the wavelength binning step, i.e. summing over the dispersion axis. The baseline is simple summing, but there's several effects in the data that make it possible to do better. Most notably, there's the infamous jitter, which we can see most clearly by doing a PCA (with some additional preprocessing) over all AIRS frames in the train data:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F14984949%2F6c3728e5a47c97f7371dc1f1d5fb8a27%2FScreenshot%202025-09-30%20090624.png?generation=1759216002725805&amp;alt=media\" alt=\"\"></p>\n<p>The jitter shapes presumably correspond to pointing error and defocus.</p>\n<p>The jitter shapes sum to approximately zero over the dispersion axis; this is why simple summing works OK. It also means that any more advanced preprocessing scheme is forced to deal with the jitter, so there's a high barrier to entry. Still, there are several effects in the data that such a scheme can take advantage of:</p>\n<ul>\n<li>There are invalid pixels.</li>\n<li>The noise is not quite Poisson noise (simple summing minimized noise only for Poisson noise); the noise is lower for brighter pixels, and there are occasional very noisy pixels.</li>\n<li>The jitter shapes don't quite sum to zero, presumably because detector calibration is not perfect.</li>\n<li>There is a constant background signal that we'd like to remove.</li>\n</ul>\n<p>I spent more than half of my time on this competition trying to set up a principled preprocessing flow optimally handling all these effects. Unfortunately, I haven't managed this yet - you'll find in the code several disjointed steps attempting to deal with the above. They also tend to work only in certain configurations and combinations, which I don't yet fully understand. It still leads to an improvement of ~0.01 over simple summing - I'll crack this one next year…</p>\n<h1>3. Bayesian Inference</h1>\n<p>Now that we have the photon counts as a function of wavelength and time, we are ready to infer the exoplanet transit depth using Bayesian Inference. <br><br>\n<strong>Bayesian Inference</strong> (BI) is a powerful statistical approach, based around defining a <em>prior</em> (a statistical belief about reality) and <em>observations</em> (some form of new information). Using Bayes' law, we then combine these to find the <em>posterior</em> (an updated belief about reality). In our case, this means:</p>\n<ul>\n<li><strong>Prior</strong>: a description of the physics that affect the final measured signal, describing for example detector noise, drift, and the transit behavior (including the transit depth itself) as formal distributions.</li>\n<li><strong>Observations</strong>: the provided measurements.</li>\n<li><strong>Posterior</strong>: a breakdown of the observations into the various elements defined in the prior (see figure below). From this we can simply read out the desired transit depth. Importantly, the posterior is not just a single point, it's a distribution. By taking samples from this distribution we can find the required confidence intervals (and even full covariance matrices, although that is not asked of us here).</li>\n</ul>\n<p>There are two key elements to applying BI: defining the prior (section 3.1), and applying Bayes' law to do the inference (section 3.2).</p>\n<h2>3.1 Prior definition</h2>\n<p>The signal is decomposed into four key components, as shown visually here:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14984949%2Ffbb5d6dba1e6836c600b1473cd333121%2FScreenshot%202024-11-01%20153411.png?generation=1730488097922724&amp;alt=media\" alt=\"\"></p>\n<p>This table shows a detailed description of all prior elements:</p>\n<table>\n<thead>\n<tr>\n<th>Prior element</th>\n<th>Description</th>\n<th>Tuning and hyperparameters</th>\n<th>Degrees of freedom</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Noise</td>\n<td>Uncorrelated Gaussian per time and wavelength</td>\n<td>Standard deviations found in preprocessing</td>\n<td>1350 (FGS) + 317250 (AIRS)</td>\n</tr>\n<tr>\n<td>Star spectrum</td>\n<td>Uncorrelated value per wavelength</td>\n<td>Not regularized (infinite sigma)</td>\n<td>283</td>\n</tr>\n<tr>\n<td>Drift</td>\n<td>Third order polynomial over time and wavelength</td>\n<td>Not regularized (infinite sigma)</td>\n<td>3 (FGS) + 12 (AIRS)</td>\n</tr>\n<tr>\n<td>Transit window</td>\n<td>Found using batman package, with the following free parameters:<br>- Transit mid-time t0 (separate for FGS and AIRS)<br>- Semi-major axis sma<br>- Period P<br>- Orbital inclination i<br>- Quadratic limb darkening u0 and u1 (including linear dependence on wavelength)<br>- Transit depth Rp^2/Rs^2 (see below)</td>\n<td>Covariance matrix learned on training set using maximum likelihood estimation (MLE)</td>\n<td>11 (6 limb darkening, 2 t0, sma, P, i)</td>\n</tr>\n<tr>\n<td>Transit depth: mean</td>\n<td>Single value</td>\n<td>Lightly regularized (sigma=0.01)</td>\n<td>1</td>\n</tr>\n<tr>\n<td>Transit depth: variation</td>\n<td>Decomposed into three components below</td>\n<td>Scaling factor determined per planet using MLE (this is the only hyperparameter tuned during inference)</td>\n<td>N/A</td>\n</tr>\n<tr>\n<td>Transit depth: variation FGS</td>\n<td>Single Gaussian value</td>\n<td>Standard deviation found on training set</td>\n<td>1</td>\n</tr>\n<tr>\n<td>Transit depth: variation AIRS</td>\n<td>Gaussian Process over wavelength, sum of two squared-exponential kernels</td>\n<td>Kernel hyperparameters found on training set using maximum likelihood estimation</td>\n<td>282</td>\n</tr>\n<tr>\n<td>Transit depth: variation PCA</td>\n<td>Fixed basis functions obtained from PCA analysis</td>\n<td>PCA shapes and variation magnitudes found on training data (this year there seems to be no train-test population shift for the transit depths)</td>\n<td>5</td>\n</tr>\n</tbody>\n</table>\n<h2>3.2 Solver</h2>\n<p>Having defined the prior, finding the posterior is 'just' a matter of applying Bayes' law. The prior is fully Gaussian, which means the inference should just be a single matrix operation. But as usual, there's a snag. In our case, it's that the relationship between the prior parameters and the observation is non-linear (because of the non-linear modeling in batman, and because some prior elements are multiplied rather than added). <br> <br>\nTo deal with this non-linearity, we linearize the model around the mean of the posterior. This is repeated iteratively (i.e. apply BI -&gt; linearize around the new mean of the posterior -&gt; repeat). Each step we also update the single hyperparameter (the scaling of the transit depth) using gradient descent. <br> <br>\nThis approach does mean we need a decent starting point, which is handled as follows:</p>\n<ul>\n<li><strong>Grid search</strong>: sweep over a grid of transit depth and transit mid-time to find initial values. Other transit parameters are set to the given approximation per planet (with no limb darkening).</li>\n<li><strong>BFGS</strong>: fit drift and all transit parameters, using BFGS as implemented in <code>scipy.optimize.minimize</code>. In this step we only consider the mean over wavelengths of the AIRS signal. </li>\n<li><strong>Bayesian Inference</strong>: the full non-linear BI solver as described above. My winning submission uses 8 iterations of the solver; the code linked here only does 4 iterations to fit training and inference into the 9 hour submission limit (this only costs 0.001 in the score).</li>\n</ul>\n<p>The prediction of the transit depth and the covariance matrix of its uncertainty can be read out directly in the posterior. The covariance matrix is approximated based on 200 samples; this is mainly a holdover from last year, when there were too many parameters to construct it explicitly.</p>\n<p>There are two additional tricks that I apply if the residual of the first two steps (grid search-&gt;BFGS) is suspiciously high:</p>\n<ul>\n<li>Use transit parameters for a different planet as starting point.</li>\n<li>Try chopping off the first and last few frames of the signal (these sometimes have a high residual for reasons I don't understand).</li>\n</ul>\n<h1>4. Fudging</h1>\n<p>The prediction and uncertainty obtained above can still be improved further with some post-hoc calibration. Note that this is anathema to a proper Bayesian approach - more on that below.</p>\n<p>We adapt the mean of the transit depth per sensor based on fitting:</p>\n<ul>\n<li>A constant offset</li>\n<li>A multiplicative factor</li>\n<li>Dependence on the first limb darkening parameter</li>\n</ul>\n<p>The variation of the transit depth over wavelength is not fudged, i.e. not changed in this step.</p>\n<p>We adapt the uncertainty of the transit depth per sensor based on fitting:</p>\n<ul>\n<li>A constant offset</li>\n<li>A multiplicative factor for the mean of the transit depth</li>\n<li>A multiplicative factor for the variation of the transit depth over wavelength</li>\n<li>Dependence on the magnitude of the variation of the transit depth over wavelength</li>\n</ul>\n<p>In total, the fudging has 12 parameters to fit. These are determined on the full training set, based on whichever values optimize the competition metric.</p>\n<h1>5. Final thoughts - we're missing something!</h1>\n<p>The fact that fudging helps the score is extremely worrying. And it's not subtle: last year, my solution without fudging could have gotten 1st place; my solution this year without fudging would have ended up around ~20th place. It means there's something missing in the prior, i.e. it does not describe reality (or rather the synthetic data generation process) accurately.</p>\n<p>I've spent a lot of time trying to find this gap, but haven't gotten much closer. I've come to believe that the issue lies in the transit modeling, but none of the variation I tried (such as different limb darkening models) helped. I'm convinced that finding whatever's going on could boost our scores to well over 0.700.</p>\n<p>Some more direct evidence of this issue can be found by comparing different transits for the same planet; the plot below shows the ratio of the transit profiles (with some preprocessing including low pass filtering). Nothing in the prior (and nothing else I can come up with) can explain this. See also <a href=\"https://www.kaggle.com/competitions/ariel-data-challenge-2025/discussion/602425\" target=\"_blank\">this discussion</a> and <a href=\"https://www.kaggle.com/code/jeroencottaar/demonstrate-apparent-label-issue\" target=\"_blank\">the code</a>.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14984949%2F1b6f904fd58b66870aa7febac31ea864%2FScreenshot%202025-09-30%20111015.png?generation=1759223442391305&amp;alt=media\" alt=\"\"><br>\nI hope the organizers are willing to reveal some clues about this, because I don't think we'll figure it out next year either without some help.</p>\n<p>Finally, I'll wrap it up with an overview of how much various elements of my solution contribute to the score:</p>\n<table>\n<thead>\n<tr>\n<th>Change to model</th>\n<th>Impact on private test score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Simplified loader (host's flow + cosmic ray removal), no jitter correction or background removal</td>\n<td>-0.013</td>\n</tr>\n<tr>\n<td>Don't remove first or last few frames of suspicious transits</td>\n<td>-0.002</td>\n</tr>\n<tr>\n<td>Don't use Gaussian Process for transit depth, just use PCA<br>Note: this variation does not involve any Gaussian Processes</td>\n<td>-0.007</td>\n</tr>\n<tr>\n<td>Don't use PCA for transit depth, just Gaussian Process</td>\n<td>-0.011</td>\n</tr>\n<tr>\n<td>Disable all fudging - rely purely on the outcome of the Bayesian predictions</td>\n<td>-0.152<br>(Last year: -0.000…)</td>\n</tr>\n<tr>\n<td>Replace all fudging by a single multiplicative factor for the uncertainty</td>\n<td>-0.017</td>\n</tr>\n<tr>\n<td>Disable fudging of sigma prediction based on AIRS variation</td>\n<td>-0.005</td>\n</tr>\n<tr>\n<td>Disable fudging of mean prediction based on limb darkening parameters</td>\n<td>-0.005</td>\n</tr>\n<tr>\n<td>Don't regularize the 11 transit parameters in prior</td>\n<td>-0.003</td>\n</tr>\n</tbody>\n</table>",
  "messages": [
    {
      "id": 3296150,
      "postDate": "2025-09-30T09:29:08.290Z",
      "content": "<p><a href=\"https://www.kaggle.com/code/jeroencottaar/ariel2025-1st-place-bayesian-exoplanet-analysis\" target=\"_blank\">Link to code to train and submit</a><br>\nThis year I'm also releasing the full development history: <a href=\"https://github.com/jcottaar/ariel2\" target=\"_blank\">https://github.com/jcottaar/ariel2</a> <br><br>\nI'd like to start by thanking the organizers for another great competition. I had a rather frustrating start (having somehow managed to lose all my code from last year), but greatly enjoyed it overall. The Bayesian approach is not used nearly as often as it should be these days, and I'm glad to have the opportunity to show its value here.</p>\n<p>My solution is quite similar to last year's; my <a href=\"https://neurips.cc/virtual/2024/107111\" target=\"_blank\">NeurIPS talk</a> from the associated workshop can be a useful introduction. I'll briefly discuss the differences compared to last year in a reply to this post.</p>\n<h1>1. Introduction</h1>\n<p>Transit analysis is the most common method for detecting and studying exoplanets. A transit occurs when a planet crosses in front of its star relative to Earth, obscuring part of the starlight. By observing how much the light is reduced (the transit depth) as a function of wavelength, we can learn the properties of the planet. But with a tiny planet passing in front of a huge star, this problem has a very low signal to noise ratio. The challenge to us: find the transit depth, including confidence intervals, from synthetic raw spectroscopic signals as the Ariel satellite might see them in the future (For a more extensive overview, see <a href=\"https://www.kaggle.com/competitions/ariel-data-challenge-2025/overview)\" target=\"_blank\">https://www.kaggle.com/competitions/ariel-data-challenge-2025/overview)</a>.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14984949%2Fa0860242fa947c058ce0794fa0494cc4%2FScreenshot%202024-11-01%20200453.png?generation=1730488072791970&amp;alt=media\" alt=\"\"></p>\n<p>This competition is at first glance a retread of a similar competition last year, but the data is far more realistic this time around. I think I figured out some of the changes, but not all of them - as we'll see later on.<br><br>\nNote that there are two sensors in play, the single-wavelength FGS sensor and the spectroscopic AIRS sensor. In this writeup I'll focus mainly on the AIRS sensor.<br><br>\nMy general approach consists of three steps, as visualized below:</p>\n<ul>\n<li><strong>Preprocessing</strong>: starting from the raw sensor data (expressed in counts), compute the amplitude signal <em>S(λ,t)</em> as a function of wavelength <em>λ</em> and time <em>t</em>.</li>\n<li><strong>Bayesian Inference</strong>: from <em>S(λ,t)</em>,  compute the inferred transit depth <em>D(λ)</em> and an uncertainty margin on this <em>σ(λ)</em> . This step is done using a principled Bayesian approach.</li>\n<li><strong>Fudging</strong>: throw the principles out the window, and apply various fudge factors to <em>D(λ)</em> and <em>σ(λ)</em> . These fudge factors are chosen to optimize score on the training data.</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F14984949%2Fa0b169937d51e43db62a030def487f5c%2FScreenshot%202025-09-29%20212745.png?generation=1759174122624685&amp;alt=media\" alt=\"\"></p>\n<p>In sections 2, 3, and 4, I'll discuss each of these steps in turn. I'll finish with some closing thoughts in section 5 - mainly discussing why my approach isn't actually working…</p>\n<h1>2. Preprocessing</h1>\n<p>The raw data is in counts per pixel; we first need to apply several steps to come to a photon count per pixel. This is explained in <a href=\"https://www.kaggle.com/code/gordonyip/calibrating-and-binning-ariel-data\" target=\"_blank\">a notebook shared by the hosts</a>. I use a time binning of 5 frames for AIRS, and 50 frames for FGS (this puts them approximately in sync). Hot pixel invalidation is disabled. I also added a simple cosmic ray removal. <br><br>\nThe opportunity for some real improvement lies in the wavelength binning step, i.e. summing over the dispersion axis. The baseline is simple summing, but there's several effects in the data that make it possible to do better. Most notably, there's the infamous jitter, which we can see most clearly by doing a PCA (with some additional preprocessing) over all AIRS frames in the train data:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F14984949%2F6c3728e5a47c97f7371dc1f1d5fb8a27%2FScreenshot%202025-09-30%20090624.png?generation=1759216002725805&amp;alt=media\" alt=\"\"></p>\n<p>The jitter shapes presumably correspond to pointing error and defocus.</p>\n<p>The jitter shapes sum to approximately zero over the dispersion axis; this is why simple summing works OK. It also means that any more advanced preprocessing scheme is forced to deal with the jitter, so there's a high barrier to entry. Still, there are several effects in the data that such a scheme can take advantage of:</p>\n<ul>\n<li>There are invalid pixels.</li>\n<li>The noise is not quite Poisson noise (simple summing minimized noise only for Poisson noise); the noise is lower for brighter pixels, and there are occasional very noisy pixels.</li>\n<li>The jitter shapes don't quite sum to zero, presumably because detector calibration is not perfect.</li>\n<li>There is a constant background signal that we'd like to remove.</li>\n</ul>\n<p>I spent more than half of my time on this competition trying to set up a principled preprocessing flow optimally handling all these effects. Unfortunately, I haven't managed this yet - you'll find in the code several disjointed steps attempting to deal with the above. They also tend to work only in certain configurations and combinations, which I don't yet fully understand. It still leads to an improvement of ~0.01 over simple summing - I'll crack this one next year…</p>\n<h1>3. Bayesian Inference</h1>\n<p>Now that we have the photon counts as a function of wavelength and time, we are ready to infer the exoplanet transit depth using Bayesian Inference. <br><br>\n<strong>Bayesian Inference</strong> (BI) is a powerful statistical approach, based around defining a <em>prior</em> (a statistical belief about reality) and <em>observations</em> (some form of new information). Using Bayes' law, we then combine these to find the <em>posterior</em> (an updated belief about reality). In our case, this means:</p>\n<ul>\n<li><strong>Prior</strong>: a description of the physics that affect the final measured signal, describing for example detector noise, drift, and the transit behavior (including the transit depth itself) as formal distributions.</li>\n<li><strong>Observations</strong>: the provided measurements.</li>\n<li><strong>Posterior</strong>: a breakdown of the observations into the various elements defined in the prior (see figure below). From this we can simply read out the desired transit depth. Importantly, the posterior is not just a single point, it's a distribution. By taking samples from this distribution we can find the required confidence intervals (and even full covariance matrices, although that is not asked of us here).</li>\n</ul>\n<p>There are two key elements to applying BI: defining the prior (section 3.1), and applying Bayes' law to do the inference (section 3.2).</p>\n<h2>3.1 Prior definition</h2>\n<p>The signal is decomposed into four key components, as shown visually here:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14984949%2Ffbb5d6dba1e6836c600b1473cd333121%2FScreenshot%202024-11-01%20153411.png?generation=1730488097922724&amp;alt=media\" alt=\"\"></p>\n<p>This table shows a detailed description of all prior elements:</p>\n<table>\n<thead>\n<tr>\n<th>Prior element</th>\n<th>Description</th>\n<th>Tuning and hyperparameters</th>\n<th>Degrees of freedom</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Noise</td>\n<td>Uncorrelated Gaussian per time and wavelength</td>\n<td>Standard deviations found in preprocessing</td>\n<td>1350 (FGS) + 317250 (AIRS)</td>\n</tr>\n<tr>\n<td>Star spectrum</td>\n<td>Uncorrelated value per wavelength</td>\n<td>Not regularized (infinite sigma)</td>\n<td>283</td>\n</tr>\n<tr>\n<td>Drift</td>\n<td>Third order polynomial over time and wavelength</td>\n<td>Not regularized (infinite sigma)</td>\n<td>3 (FGS) + 12 (AIRS)</td>\n</tr>\n<tr>\n<td>Transit window</td>\n<td>Found using batman package, with the following free parameters:<br>- Transit mid-time t0 (separate for FGS and AIRS)<br>- Semi-major axis sma<br>- Period P<br>- Orbital inclination i<br>- Quadratic limb darkening u0 and u1 (including linear dependence on wavelength)<br>- Transit depth Rp^2/Rs^2 (see below)</td>\n<td>Covariance matrix learned on training set using maximum likelihood estimation (MLE)</td>\n<td>11 (6 limb darkening, 2 t0, sma, P, i)</td>\n</tr>\n<tr>\n<td>Transit depth: mean</td>\n<td>Single value</td>\n<td>Lightly regularized (sigma=0.01)</td>\n<td>1</td>\n</tr>\n<tr>\n<td>Transit depth: variation</td>\n<td>Decomposed into three components below</td>\n<td>Scaling factor determined per planet using MLE (this is the only hyperparameter tuned during inference)</td>\n<td>N/A</td>\n</tr>\n<tr>\n<td>Transit depth: variation FGS</td>\n<td>Single Gaussian value</td>\n<td>Standard deviation found on training set</td>\n<td>1</td>\n</tr>\n<tr>\n<td>Transit depth: variation AIRS</td>\n<td>Gaussian Process over wavelength, sum of two squared-exponential kernels</td>\n<td>Kernel hyperparameters found on training set using maximum likelihood estimation</td>\n<td>282</td>\n</tr>\n<tr>\n<td>Transit depth: variation PCA</td>\n<td>Fixed basis functions obtained from PCA analysis</td>\n<td>PCA shapes and variation magnitudes found on training data (this year there seems to be no train-test population shift for the transit depths)</td>\n<td>5</td>\n</tr>\n</tbody>\n</table>\n<h2>3.2 Solver</h2>\n<p>Having defined the prior, finding the posterior is 'just' a matter of applying Bayes' law. The prior is fully Gaussian, which means the inference should just be a single matrix operation. But as usual, there's a snag. In our case, it's that the relationship between the prior parameters and the observation is non-linear (because of the non-linear modeling in batman, and because some prior elements are multiplied rather than added). <br> <br>\nTo deal with this non-linearity, we linearize the model around the mean of the posterior. This is repeated iteratively (i.e. apply BI -&gt; linearize around the new mean of the posterior -&gt; repeat). Each step we also update the single hyperparameter (the scaling of the transit depth) using gradient descent. <br> <br>\nThis approach does mean we need a decent starting point, which is handled as follows:</p>\n<ul>\n<li><strong>Grid search</strong>: sweep over a grid of transit depth and transit mid-time to find initial values. Other transit parameters are set to the given approximation per planet (with no limb darkening).</li>\n<li><strong>BFGS</strong>: fit drift and all transit parameters, using BFGS as implemented in <code>scipy.optimize.minimize</code>. In this step we only consider the mean over wavelengths of the AIRS signal. </li>\n<li><strong>Bayesian Inference</strong>: the full non-linear BI solver as described above. My winning submission uses 8 iterations of the solver; the code linked here only does 4 iterations to fit training and inference into the 9 hour submission limit (this only costs 0.001 in the score).</li>\n</ul>\n<p>The prediction of the transit depth and the covariance matrix of its uncertainty can be read out directly in the posterior. The covariance matrix is approximated based on 200 samples; this is mainly a holdover from last year, when there were too many parameters to construct it explicitly.</p>\n<p>There are two additional tricks that I apply if the residual of the first two steps (grid search-&gt;BFGS) is suspiciously high:</p>\n<ul>\n<li>Use transit parameters for a different planet as starting point.</li>\n<li>Try chopping off the first and last few frames of the signal (these sometimes have a high residual for reasons I don't understand).</li>\n</ul>\n<h1>4. Fudging</h1>\n<p>The prediction and uncertainty obtained above can still be improved further with some post-hoc calibration. Note that this is anathema to a proper Bayesian approach - more on that below.</p>\n<p>We adapt the mean of the transit depth per sensor based on fitting:</p>\n<ul>\n<li>A constant offset</li>\n<li>A multiplicative factor</li>\n<li>Dependence on the first limb darkening parameter</li>\n</ul>\n<p>The variation of the transit depth over wavelength is not fudged, i.e. not changed in this step.</p>\n<p>We adapt the uncertainty of the transit depth per sensor based on fitting:</p>\n<ul>\n<li>A constant offset</li>\n<li>A multiplicative factor for the mean of the transit depth</li>\n<li>A multiplicative factor for the variation of the transit depth over wavelength</li>\n<li>Dependence on the magnitude of the variation of the transit depth over wavelength</li>\n</ul>\n<p>In total, the fudging has 12 parameters to fit. These are determined on the full training set, based on whichever values optimize the competition metric.</p>\n<h1>5. Final thoughts - we're missing something!</h1>\n<p>The fact that fudging helps the score is extremely worrying. And it's not subtle: last year, my solution without fudging could have gotten 1st place; my solution this year without fudging would have ended up around ~20th place. It means there's something missing in the prior, i.e. it does not describe reality (or rather the synthetic data generation process) accurately.</p>\n<p>I've spent a lot of time trying to find this gap, but haven't gotten much closer. I've come to believe that the issue lies in the transit modeling, but none of the variation I tried (such as different limb darkening models) helped. I'm convinced that finding whatever's going on could boost our scores to well over 0.700.</p>\n<p>Some more direct evidence of this issue can be found by comparing different transits for the same planet; the plot below shows the ratio of the transit profiles (with some preprocessing including low pass filtering). Nothing in the prior (and nothing else I can come up with) can explain this. See also <a href=\"https://www.kaggle.com/competitions/ariel-data-challenge-2025/discussion/602425\" target=\"_blank\">this discussion</a> and <a href=\"https://www.kaggle.com/code/jeroencottaar/demonstrate-apparent-label-issue\" target=\"_blank\">the code</a>.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14984949%2F1b6f904fd58b66870aa7febac31ea864%2FScreenshot%202025-09-30%20111015.png?generation=1759223442391305&amp;alt=media\" alt=\"\"><br>\nI hope the organizers are willing to reveal some clues about this, because I don't think we'll figure it out next year either without some help.</p>\n<p>Finally, I'll wrap it up with an overview of how much various elements of my solution contribute to the score:</p>\n<table>\n<thead>\n<tr>\n<th>Change to model</th>\n<th>Impact on private test score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Simplified loader (host's flow + cosmic ray removal), no jitter correction or background removal</td>\n<td>-0.013</td>\n</tr>\n<tr>\n<td>Don't remove first or last few frames of suspicious transits</td>\n<td>-0.002</td>\n</tr>\n<tr>\n<td>Don't use Gaussian Process for transit depth, just use PCA<br>Note: this variation does not involve any Gaussian Processes</td>\n<td>-0.007</td>\n</tr>\n<tr>\n<td>Don't use PCA for transit depth, just Gaussian Process</td>\n<td>-0.011</td>\n</tr>\n<tr>\n<td>Disable all fudging - rely purely on the outcome of the Bayesian predictions</td>\n<td>-0.152<br>(Last year: -0.000…)</td>\n</tr>\n<tr>\n<td>Replace all fudging by a single multiplicative factor for the uncertainty</td>\n<td>-0.017</td>\n</tr>\n<tr>\n<td>Disable fudging of sigma prediction based on AIRS variation</td>\n<td>-0.005</td>\n</tr>\n<tr>\n<td>Disable fudging of mean prediction based on limb darkening parameters</td>\n<td>-0.005</td>\n</tr>\n<tr>\n<td>Don't regularize the 11 transit parameters in prior</td>\n<td>-0.003</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "[Link to code to train and submit](https://www.kaggle.com/code/jeroencottaar/ariel2025-1st-place-bayesian-exoplanet-analysis)\nThis year I'm also releasing the full development history: [https://github.com/jcottaar/ariel2](https://github.com/jcottaar/ariel2) <br>\nI'd like to start by thanking the organizers for another great competition. I had a rather frustrating start (having somehow managed to lose all my code from last year), but greatly enjoyed it overall. The Bayesian approach is not used nearly as often as it should be these days, and I'm glad to have the opportunity to show its value here.\n\nMy solution is quite similar to last year's; my [NeurIPS talk](https://neurips.cc/virtual/2024/107111) from the associated workshop can be a useful introduction. I'll briefly discuss the differences compared to last year in a reply to this post.\n# 1. Introduction\n\nTransit analysis is the most common method for detecting and studying exoplanets. A transit occurs when a planet crosses in front of its star relative to Earth, obscuring part of the starlight. By observing how much the light is reduced (the transit depth) as a function of wavelength, we can learn the properties of the planet. But with a tiny planet passing in front of a huge star, this problem has a very low signal to noise ratio. The challenge to us: find the transit depth, including confidence intervals, from synthetic raw spectroscopic signals as the Ariel satellite might see them in the future (For a more extensive overview, see https://www.kaggle.com/competitions/ariel-data-challenge-2025/overview).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14984949%2Fa0860242fa947c058ce0794fa0494cc4%2FScreenshot%202024-11-01%20200453.png?generation=1730488072791970&alt=media)\n\nThis competition is at first glance a retread of a similar competition last year, but the data is far more realistic this time around. I think I figured out some of the changes, but not all of them - as we'll see later on.<br>\nNote that there are two sensors in play, the single-wavelength FGS sensor and the spectroscopic AIRS sensor. In this writeup I'll focus mainly on the AIRS sensor.<br>\nMy general approach consists of three steps, as visualized below:\n- **Preprocessing**: starting from the raw sensor data (expressed in counts), compute the amplitude signal *S(λ,t)* as a function of wavelength *λ* and time *t*.\n- **Bayesian Inference**: from *S(λ,t)*,  compute the inferred transit depth *D(λ)* and an uncertainty margin on this *σ(λ)* . This step is done using a principled Bayesian approach.\n- **Fudging**: throw the principles out the window, and apply various fudge factors to *D(λ)* and *σ(λ)* . These fudge factors are chosen to optimize score on the training data.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F14984949%2Fa0b169937d51e43db62a030def487f5c%2FScreenshot%202025-09-29%20212745.png?generation=1759174122624685&alt=media)\n\nIn sections 2, 3, and 4, I'll discuss each of these steps in turn. I'll finish with some closing thoughts in section 5 - mainly discussing why my approach isn't actually working...\n# 2. Preprocessing\n\nThe raw data is in counts per pixel; we first need to apply several steps to come to a photon count per pixel. This is explained in [a notebook shared by the hosts](https://www.kaggle.com/code/gordonyip/calibrating-and-binning-ariel-data). I use a time binning of 5 frames for AIRS, and 50 frames for FGS (this puts them approximately in sync). Hot pixel invalidation is disabled. I also added a simple cosmic ray removal. <br>\nThe opportunity for some real improvement lies in the wavelength binning step, i.e. summing over the dispersion axis. The baseline is simple summing, but there's several effects in the data that make it possible to do better. Most notably, there's the infamous jitter, which we can see most clearly by doing a PCA (with some additional preprocessing) over all AIRS frames in the train data:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F14984949%2F6c3728e5a47c97f7371dc1f1d5fb8a27%2FScreenshot%202025-09-30%20090624.png?generation=1759216002725805&alt=media)\n\nThe jitter shapes presumably correspond to pointing error and defocus.\n\nThe jitter shapes sum to approximately zero over the dispersion axis; this is why simple summing works OK. It also means that any more advanced preprocessing scheme is forced to deal with the jitter, so there's a high barrier to entry. Still, there are several effects in the data that such a scheme can take advantage of:\n\n- There are invalid pixels.\n- The noise is not quite Poisson noise (simple summing minimized noise only for Poisson noise); the noise is lower for brighter pixels, and there are occasional very noisy pixels.\n- The jitter shapes don't quite sum to zero, presumably because detector calibration is not perfect.\n- There is a constant background signal that we'd like to remove.\n\nI spent more than half of my time on this competition trying to set up a principled preprocessing flow optimally handling all these effects. Unfortunately, I haven't managed this yet - you'll find in the code several disjointed steps attempting to deal with the above. They also tend to work only in certain configurations and combinations, which I don't yet fully understand. It still leads to an improvement of ~0.01 over simple summing - I'll crack this one next year...\n# 3. Bayesian Inference\n\nNow that we have the photon counts as a function of wavelength and time, we are ready to infer the exoplanet transit depth using Bayesian Inference. <br>\n**Bayesian Inference** (BI) is a powerful statistical approach, based around defining a *prior* (a statistical belief about reality) and *observations* (some form of new information). Using Bayes' law, we then combine these to find the *posterior* (an updated belief about reality). In our case, this means:\n- **Prior**: a description of the physics that affect the final measured signal, describing for example detector noise, drift, and the transit behavior (including the transit depth itself) as formal distributions.\n- **Observations**: the provided measurements.\n- **Posterior**: a breakdown of the observations into the various elements defined in the prior (see figure below). From this we can simply read out the desired transit depth. Importantly, the posterior is not just a single point, it's a distribution. By taking samples from this distribution we can find the required confidence intervals (and even full covariance matrices, although that is not asked of us here).\n\nThere are two key elements to applying BI: defining the prior (section 3.1), and applying Bayes' law to do the inference (section 3.2).\n## 3.1 Prior definition\n\nThe signal is decomposed into four key components, as shown visually here:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14984949%2Ffbb5d6dba1e6836c600b1473cd333121%2FScreenshot%202024-11-01%20153411.png?generation=1730488097922724&alt=media)\n\nThis table shows a detailed description of all prior elements:\n\n| Prior element                 | Description                                                                                                                                                                                                                                                                                                         | Tuning and hyperparameters                                                                                                                     | Degrees of freedom                     |\n| ----------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------- |\n| Noise                         | Uncorrelated Gaussian per time and wavelength                                                                                                                                                                                                                                                                       | Standard deviations found in preprocessing                                                                                                     | 1350 (FGS) + 317250 (AIRS)             |\n| Star spectrum                 | Uncorrelated value per wavelength                                                                                                                                                                                                                                                                                   | Not regularized (infinite sigma)                                                                                                               | 283                                    |\n| Drift                         | Third order polynomial over time and wavelength                                                                                                                                                                                                                                                                     | Not regularized (infinite sigma)                                                                                                               | 3 (FGS) + 12 (AIRS)                    |\n| Transit window                | Found using batman package, with the following free parameters:<br>- Transit mid-time t0 (separate for FGS and AIRS)<br>- Semi-major axis sma<br>- Period P<br>- Orbital inclination i<br>- Quadratic limb darkening u0 and u1 (including linear dependence on wavelength)<br>- Transit depth Rp^2/Rs^2 (see below) | Covariance matrix learned on training set using maximum likelihood estimation (MLE)                                                            | 11 (6 limb darkening, 2 t0, sma, P, i) |\n| Transit depth: mean           | Single value                                                                                                                                                                                                                                                                                                        | Lightly regularized (sigma=0.01)                                                                                                               | 1                                      |\n| Transit depth: variation      | Decomposed into three components below                                                                                                                                                                                                                                                                              | Scaling factor determined per planet using MLE (this is the only hyperparameter tuned during inference)                                        | N/A                                    |\n| Transit depth: variation FGS  | Single Gaussian value                                                                                                                                                                                                                                                                                               | Standard deviation found on training set                                                                                                       | 1                                      |\n| Transit depth: variation AIRS | Gaussian Process over wavelength, sum of two squared-exponential kernels                                                                                                                                                                                                                                            | Kernel hyperparameters found on training set using maximum likelihood estimation                                                               | 282                                    |\n| Transit depth: variation PCA  | Fixed basis functions obtained from PCA analysis                                                                                                                                                                                                                                                                    | PCA shapes and variation magnitudes found on training data (this year there seems to be no train-test population shift for the transit depths) | 5                                      |\n## 3.2 Solver\n\nHaving defined the prior, finding the posterior is 'just' a matter of applying Bayes' law. The prior is fully Gaussian, which means the inference should just be a single matrix operation. But as usual, there's a snag. In our case, it's that the relationship between the prior parameters and the observation is non-linear (because of the non-linear modeling in batman, and because some prior elements are multiplied rather than added). <br> \nTo deal with this non-linearity, we linearize the model around the mean of the posterior. This is repeated iteratively (i.e. apply BI -> linearize around the new mean of the posterior -> repeat). Each step we also update the single hyperparameter (the scaling of the transit depth) using gradient descent. <br> \nThis approach does mean we need a decent starting point, which is handled as follows:\n\n- **Grid search**: sweep over a grid of transit depth and transit mid-time to find initial values. Other transit parameters are set to the given approximation per planet (with no limb darkening).\n- **BFGS**: fit drift and all transit parameters, using BFGS as implemented in ```scipy.optimize.minimize```. In this step we only consider the mean over wavelengths of the AIRS signal. \n- **Bayesian Inference**: the full non-linear BI solver as described above. My winning submission uses 8 iterations of the solver; the code linked here only does 4 iterations to fit training and inference into the 9 hour submission limit (this only costs 0.001 in the score).\n\nThe prediction of the transit depth and the covariance matrix of its uncertainty can be read out directly in the posterior. The covariance matrix is approximated based on 200 samples; this is mainly a holdover from last year, when there were too many parameters to construct it explicitly.\n\nThere are two additional tricks that I apply if the residual of the first two steps (grid search->BFGS) is suspiciously high:\n- Use transit parameters for a different planet as starting point.\n- Try chopping off the first and last few frames of the signal (these sometimes have a high residual for reasons I don't understand).\n  \n# 4. Fudging\n\nThe prediction and uncertainty obtained above can still be improved further with some post-hoc calibration. Note that this is anathema to a proper Bayesian approach - more on that below.\n\nWe adapt the mean of the transit depth per sensor based on fitting:\n - A constant offset\n - A multiplicative factor\n - Dependence on the first limb darkening parameter\n\nThe variation of the transit depth over wavelength is not fudged, i.e. not changed in this step.\n\nWe adapt the uncertainty of the transit depth per sensor based on fitting:\n- A constant offset\n- A multiplicative factor for the mean of the transit depth\n- A multiplicative factor for the variation of the transit depth over wavelength\n- Dependence on the magnitude of the variation of the transit depth over wavelength\n\nIn total, the fudging has 12 parameters to fit. These are determined on the full training set, based on whichever values optimize the competition metric.\n# 5. Final thoughts - we're missing something!\n\nThe fact that fudging helps the score is extremely worrying. And it's not subtle: last year, my solution without fudging could have gotten 1st place; my solution this year without fudging would have ended up around ~20th place. It means there's something missing in the prior, i.e. it does not describe reality (or rather the synthetic data generation process) accurately.\n\nI've spent a lot of time trying to find this gap, but haven't gotten much closer. I've come to believe that the issue lies in the transit modeling, but none of the variation I tried (such as different limb darkening models) helped. I'm convinced that finding whatever's going on could boost our scores to well over 0.700.\n\nSome more direct evidence of this issue can be found by comparing different transits for the same planet; the plot below shows the ratio of the transit profiles (with some preprocessing including low pass filtering). Nothing in the prior (and nothing else I can come up with) can explain this. See also [this discussion](https://www.kaggle.com/competitions/ariel-data-challenge-2025/discussion/602425) and [the code](https://www.kaggle.com/code/jeroencottaar/demonstrate-apparent-label-issue).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14984949%2F1b6f904fd58b66870aa7febac31ea864%2FScreenshot%202025-09-30%20111015.png?generation=1759223442391305&alt=media)\nI hope the organizers are willing to reveal some clues about this, because I don't think we'll figure it out next year either without some help.\n\nFinally, I'll wrap it up with an overview of how much various elements of my solution contribute to the score:\n\n| Change to model                                                                                                            | Impact on private test score     |\n| -------------------------------------------------------------------------------------------------------------------------- | -------------------------------- |\n| Simplified loader (host's flow + cosmic ray removal), no jitter correction or background removal                           | -0.013                           |\n| Don't remove first or last few frames of suspicious transits                                                               | -0.002                           |\n| Don't use Gaussian Process for transit depth, just use PCA<br>Note: this variation does not involve any Gaussian Processes | -0.007                           |\n| Don't use PCA for transit depth, just Gaussian Process                                                                     | -0.011                           |\n| Disable all fudging - rely purely on the outcome of the Bayesian predictions                                               | -0.152<br>(Last year: -0.000...) |\n| Replace all fudging by a single multiplicative factor for the uncertainty                                                  | -0.017                           |\n| Disable fudging of sigma prediction based on AIRS variation                                                                | -0.005                           |\n| Disable fudging of mean prediction based on limb darkening parameters                                                      | -0.005                           |\n| Don't regularize the 11 transit parameters in prior                                                                        | -0.003                           |\n\n",
      "votes": 56
    },
    {
      "id": 3296154,
      "postDate": "2025-09-30T09:43:01.380Z",
      "content": "<p>Key differences with my <a href=\"https://www.kaggle.com/competitions/ariel-data-challenge-2024/writeups/jeroen-cottaar-2nd-place-solution-pure-bayesian-in\" target=\"_blank\">2024 solution</a>:</p>\n<ul>\n<li>Use batman to model transits.</li>\n<li>Added limb darkening to the transit model.</li>\n<li>PCA for transit depths is done on the training set, rather than on the test set (no apparent population shift this year). Also uses 5 components instead of 1.</li>\n<li>Simplified Gaussian Process for transit depth (hardly needed still - a solution without any Gaussian Processes anywhere in the prior scores 0.642).</li>\n<li>Drift is modeled as a third-order polynomial rather than a Gaussian Process; less realistic, but matches the synthetic data generation process.</li>\n<li>Initial solver (grid search and BFGS) to find a good starting point for the Bayesian solver; mainly needed because of limb darkening.</li>\n<li>Progress towards proper jitter correction and background removal in preprocessing.</li>\n<li>Extensive post-hoc calibration (i.e. fudging) to deal with gaps in the prior - no fudging was needed last year.</li>\n<li>Significantly reduced computation time and memory usage with code improvements, including moving all preprocessing to GPU.</li>\n</ul>",
      "rawMarkdown": "Key differences with my [2024 solution](https://www.kaggle.com/competitions/ariel-data-challenge-2024/writeups/jeroen-cottaar-2nd-place-solution-pure-bayesian-in):\n\n- Use batman to model transits.\n- Added limb darkening to the transit model.\n- PCA for transit depths is done on the training set, rather than on the test set (no apparent population shift this year). Also uses 5 components instead of 1.\n- Simplified Gaussian Process for transit depth (hardly needed still - a solution without any Gaussian Processes anywhere in the prior scores 0.642).\n- Drift is modeled as a third-order polynomial rather than a Gaussian Process; less realistic, but matches the synthetic data generation process.\n- Initial solver (grid search and BFGS) to find a good starting point for the Bayesian solver; mainly needed because of limb darkening.\n- Progress towards proper jitter correction and background removal in preprocessing.\n- Extensive post-hoc calibration (i.e. fudging) to deal with gaps in the prior - no fudging was needed last year.\n- Significantly reduced computation time and memory usage with code improvements, including moving all preprocessing to GPU.",
      "votes": 3
    },
    {
      "id": 3298395,
      "postDate": "2025-10-05T12:40:40.457Z",
      "content": "<p>Congrats!</p>\n<p>I wonder do you have some high-level insights on why Bayesian without NN work well in 2024 and 2025.  As our team tried end-to-end deep neural networks (DNNs) which is extremely vulnerable and still need ensemble with other methods for robustness.</p>\n<p>My insight is that <br>\n1) The actual noise functions is linearly like so that DNNs overfit it.<br>\n2) The training data is small so that DNNs overfit it.<br>\n3) The data sample is not infomation-abundent along with outliers so that DNNs overfit it.</p>",
      "rawMarkdown": "Congrats!\n\nI wonder do you have some high-level insights on why Bayesian without NN work well in 2024 and 2025.  As our team tried end-to-end deep neural networks (DNNs) which is extremely vulnerable and still need ensemble with other methods for robustness.\n\nMy insight is that \n1) The actual noise functions is linearly like so that DNNs overfit it.\n2) The training data is small so that DNNs overfit it.\n3) The data sample is not infomation-abundent along with outliers so that DNNs overfit it.",
      "replies": [
        {
          "id": 3299089,
          "postDate": "2025-10-07T08:05:31.907Z",
          "content": "<p>Bayesian methods have an edge here because the underlying phsyics are relatively simple, but at the same time challenging for neural networks to identify. You just can't get away with ignoring the domain knowledge - neural networks are not going to learn that here (except perhaps if you provide enough training data so that they can essentially memorize all solutions).</p>",
          "rawMarkdown": "Bayesian methods have an edge here because the underlying phsyics are relatively simple, but at the same time challenging for neural networks to identify. You just can't get away with ignoring the domain knowledge - neural networks are not going to learn that here (except perhaps if you provide enough training data so that they can essentially memorize all solutions).",
          "votes": 2
        }
      ]
    },
    {
      "id": 3296578,
      "postDate": "2025-10-01T08:00:16.177Z",
      "content": "<p>Thanks for the nice summary, <a href=\"https://www.kaggle.com/jeroencottaar\" target=\"_blank\">@jeroencottaar</a> !<br>\nIn your model you did not restrict the quadratic u1/u2 LD coefficients, so they could converge to Limb Brightening regime, right?</p>",
      "rawMarkdown": "Thanks for the nice summary, @jeroencottaar !\nIn your model you did not restrict the quadratic u1/u2 LD coefficients, so they could converge to Limb Brightening regime, right?",
      "replies": [
        {
          "id": 3296600,
          "postDate": "2025-10-01T08:44:26.513Z",
          "content": "<p>They are regularized but not constrained. I highly doubt they would have negative values though (that's what you mean with limb brightening right?), that'd be very far outside the training population. </p>\n<p>I'll run some submission tonight to check this explicitly.</p>",
          "rawMarkdown": "They are regularized but not constrained. I highly doubt they would have negative values though (that's what you mean with limb brightening right?), that'd be very far outside the training population. \n\nI'll run some submission tonight to check this explicitly.",
          "votes": 1,
          "replies": [
            {
              "id": 3296608,
              "postDate": "2025-10-01T09:06:58.727Z",
              "content": "<p>Thanks <a href=\"https://www.kaggle.com/jeroencottaar\" target=\"_blank\">@jeroencottaar</a> !<br>\nJust for reference - Figure 1 of <a href=\"https://arxiv.org/pdf/1308.0009\" target=\"_blank\">Kipping paper</a> illustrates physical constraints on quadratic u1, u2 values.</p>",
              "rawMarkdown": "Thanks @jeroencottaar !\nJust for reference - Figure 1 of [Kipping paper](https://arxiv.org/pdf/1308.0009) illustrates physical constraints on quadratic u1, u2 values."
            },
            {
              "id": 3296858,
              "postDate": "2025-10-01T19:18:58.860Z",
              "content": "<p>I checked, and see that I already had constraints imposed on the limb darkening parameters. No negative values occcurred in my winning submission.</p>",
              "rawMarkdown": "I checked, and see that I already had constraints imposed on the limb darkening parameters. No negative values occcurred in my winning submission."
            }
          ]
        }
      ]
    },
    {
      "id": 3296396,
      "postDate": "2025-09-30T19:26:04.147Z",
      "content": "<p>Congrats for the 1st place finish!</p>\n<p>I'm curious how well your model fits the ingress/egress period.</p>\n<p>In my model, I also had to do adjustment to the mean transit depth and I found the likely reason to be the limb darkening model not fitting well near the limbs (as well as lack of observations near the center), which greatly affects the transit depth distortion (measured depth vs (Rp/Rs)2 ) since the stellar flux is heavily dependent on the limb darkening assumption.</p>",
      "rawMarkdown": "Congrats for the 1st place finish!\n\nI'm curious how well your model fits the ingress/egress period.\n\nIn my model, I also had to do adjustment to the mean transit depth and I found the likely reason to be the limb darkening model not fitting well near the limbs (as well as lack of observations near the center), which greatly affects the transit depth distortion (measured depth vs (Rp/Rs)2 ) since the stellar flux is heavily dependent on the limb darkening assumption.",
      "replies": [
        {
          "id": 3296603,
          "postDate": "2025-10-01T08:53:09.717Z",
          "content": "<p>The main reason to adjust the mean transit depth is related to how well you do the constant background correction in preprocessing. If you don't do any, you should expect a ~0.6% correction on the AIRS mean. I think my model ends up needing ~0.1%.</p>\n<p>In general the ingress/egress period is fitted slightly worse than the rest of the transit (i.e. this range has slightly higher residuals), but not to the degree that it explains the need for fudging in my estimation.</p>",
          "rawMarkdown": "The main reason to adjust the mean transit depth is related to how well you do the constant background correction in preprocessing. If you don't do any, you should expect a ~0.6% correction on the AIRS mean. I think my model ends up needing ~0.1%.\n\nIn general the ingress/egress period is fitted slightly worse than the rest of the transit (i.e. this range has slightly higher residuals), but not to the degree that it explains the need for fudging in my estimation.",
          "replies": [
            {
              "id": 3296849,
              "postDate": "2025-10-01T18:42:00.300Z",
              "content": "<p>There is also the issue that different limb darkening models that can fit the observed data similarly well, but lead to different transit depth estimation. <br>\nWhich means there is a limit in how well we can estimate solely from the observed flux data, without incoporating  further stellar properties. </p>\n<p>To give an example of how limb darkening assumption can seriously affect transit depth estimation:<br>\nI fitted both exponential and quadratic on planet 3814020142's light curve.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13392195%2F850f9369b2a65c9bdac68d0e4a45b2dd%2F10.png?generation=1759344067867790&amp;alt=media\" alt=\"\"></p>\n<p>The residuals are similar, and the observed dip in flux is similar. However, there is a serious consequence of the choice of parametric limb darkening modeling. The best-ft quadratic model results in a disk-average intensity of 0.907 while the best-fit exponential results in a disk-average intensity of .884. This is because as you can see from the bottom graph, the exponential describes a much sharper drop in intensity towards the limbs.</p>\n<p>This means that choosing quadratic limb darkening results in 2.5% higher center-to-disk average intensity ratio. This leads to 1.8% higher (Rp/Rs)^2 estimation in the model. From my observations, polynomial limb darkening with higher linear coefficient tend to overestimate (Rp/Rs)^2.</p>",
              "rawMarkdown": "There is also the issue that different limb darkening models that can fit the observed data similarly well, but lead to different transit depth estimation. \nWhich means there is a limit in how well we can estimate solely from the observed flux data, without incoporating  further stellar properties. \n\nTo give an example of how limb darkening assumption can seriously affect transit depth estimation:\nI fitted both exponential and quadratic on planet 3814020142's light curve.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13392195%2F850f9369b2a65c9bdac68d0e4a45b2dd%2F10.png?generation=1759344067867790&alt=media)\n\nThe residuals are similar, and the observed dip in flux is similar. However, there is a serious consequence of the choice of parametric limb darkening modeling. The best-ft quadratic model results in a disk-average intensity of 0.907 while the best-fit exponential results in a disk-average intensity of .884. This is because as you can see from the bottom graph, the exponential describes a much sharper drop in intensity towards the limbs.\n\n This means that choosing quadratic limb darkening results in 2.5% higher center-to-disk average intensity ratio. This leads to 1.8% higher (Rp/Rs)^2 estimation in the model. From my observations, polynomial limb darkening with higher linear coefficient tend to overestimate (Rp/Rs)^2."
            },
            {
              "id": 3296854,
              "postDate": "2025-10-01T19:00:12.980Z",
              "content": "<p>Interesting. Perhaps the 'missing element' of my prior that leads to the need for fudging is the use of different limb darkening models for different planets (I couldn't get any improvement by switching <em>all</em> planets to a different model, but perhaps I should have been flexible - would have been a nice opportunity for some further Bayesian tricks).</p>",
              "rawMarkdown": "Interesting. Perhaps the 'missing element' of my prior that leads to the need for fudging is the use of different limb darkening models for different planets (I couldn't get any improvement by switching *all* planets to a different model, but perhaps I should have been flexible - would have been a nice opportunity for some further Bayesian tricks)."
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3296154,
      "author_name": "Jeroen Cottaar",
      "author_url": "",
      "post_date": "2025-09-30T09:43:01.380000",
      "content": "<p>Key differences with my <a href=\"https://www.kaggle.com/competitions/ariel-data-challenge-2024/writeups/jeroen-cottaar-2nd-place-solution-pure-bayesian-in\" target=\"_blank\">2024 solution</a>:</p>\n<ul>\n<li>Use batman to model transits.</li>\n<li>Added limb darkening to the transit model.</li>\n<li>PCA for transit depths is done on the training set, rather than on the test set (no apparent population shift this year). Also uses 5 components instead of 1.</li>\n<li>Simplified Gaussian Process for transit depth (hardly needed still - a solution without any Gaussian Processes anywhere in the prior scores 0.642).</li>\n<li>Drift is modeled as a third-order polynomial rather than a Gaussian Process; less realistic, but matches the synthetic data generation process.</li>\n<li>Initial solver (grid search and BFGS) to find a good starting point for the Bayesian solver; mainly needed because of limb darkening.</li>\n<li>Progress towards proper jitter correction and background removal in preprocessing.</li>\n<li>Extensive post-hoc calibration (i.e. fudging) to deal with gaps in the prior - no fudging was needed last year.</li>\n<li>Significantly reduced computation time and memory usage with code improvements, including moving all preprocessing to GPU.</li>\n</ul>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 3298395,
      "author_name": "Sakura",
      "author_url": "",
      "post_date": "2025-10-05T12:40:40.457000",
      "content": "<p>Congrats!</p>\n<p>I wonder do you have some high-level insights on why Bayesian without NN work well in 2024 and 2025.  As our team tried end-to-end deep neural networks (DNNs) which is extremely vulnerable and still need ensemble with other methods for robustness.</p>\n<p>My insight is that <br>\n1) The actual noise functions is linearly like so that DNNs overfit it.<br>\n2) The training data is small so that DNNs overfit it.<br>\n3) The data sample is not infomation-abundent along with outliers so that DNNs overfit it.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3299089,
          "author_name": "Jeroen Cottaar",
          "author_url": "",
          "post_date": "2025-10-07T08:05:31.907000",
          "content": "<p>Bayesian methods have an edge here because the underlying phsyics are relatively simple, but at the same time challenging for neural networks to identify. You just can't get away with ignoring the domain knowledge - neural networks are not going to learn that here (except perhaps if you provide enough training data so that they can essentially memorize all solutions).</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 3296578,
      "author_name": "Oleh Kivernyk",
      "author_url": "",
      "post_date": "2025-10-01T08:00:16.177000",
      "content": "<p>Thanks for the nice summary, <a href=\"https://www.kaggle.com/jeroencottaar\" target=\"_blank\">@jeroencottaar</a> !<br>\nIn your model you did not restrict the quadratic u1/u2 LD coefficients, so they could converge to Limb Brightening regime, right?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3296600,
          "author_name": "Jeroen Cottaar",
          "author_url": "",
          "post_date": "2025-10-01T08:44:26.513000",
          "content": "<p>They are regularized but not constrained. I highly doubt they would have negative values though (that's what you mean with limb brightening right?), that'd be very far outside the training population. </p>\n<p>I'll run some submission tonight to check this explicitly.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3296608,
              "author_name": "Oleh Kivernyk",
              "author_url": "",
              "post_date": "2025-10-01T09:06:58.727000",
              "content": "<p>Thanks <a href=\"https://www.kaggle.com/jeroencottaar\" target=\"_blank\">@jeroencottaar</a> !<br>\nJust for reference - Figure 1 of <a href=\"https://arxiv.org/pdf/1308.0009\" target=\"_blank\">Kipping paper</a> illustrates physical constraints on quadratic u1, u2 values.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3296858,
              "author_name": "Jeroen Cottaar",
              "author_url": "",
              "post_date": "2025-10-01T19:18:58.860000",
              "content": "<p>I checked, and see that I already had constraints imposed on the limb darkening parameters. No negative values occcurred in my winning submission.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3296396,
      "author_name": "JungleBeastDS",
      "author_url": "",
      "post_date": "2025-09-30T19:26:04.147000",
      "content": "<p>Congrats for the 1st place finish!</p>\n<p>I'm curious how well your model fits the ingress/egress period.</p>\n<p>In my model, I also had to do adjustment to the mean transit depth and I found the likely reason to be the limb darkening model not fitting well near the limbs (as well as lack of observations near the center), which greatly affects the transit depth distortion (measured depth vs (Rp/Rs)2 ) since the stellar flux is heavily dependent on the limb darkening assumption.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3296603,
          "author_name": "Jeroen Cottaar",
          "author_url": "",
          "post_date": "2025-10-01T08:53:09.717000",
          "content": "<p>The main reason to adjust the mean transit depth is related to how well you do the constant background correction in preprocessing. If you don't do any, you should expect a ~0.6% correction on the AIRS mean. I think my model ends up needing ~0.1%.</p>\n<p>In general the ingress/egress period is fitted slightly worse than the rest of the transit (i.e. this range has slightly higher residuals), but not to the degree that it explains the need for fudging in my estimation.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3296849,
              "author_name": "JungleBeastDS",
              "author_url": "",
              "post_date": "2025-10-01T18:42:00.300000",
              "content": "<p>There is also the issue that different limb darkening models that can fit the observed data similarly well, but lead to different transit depth estimation. <br>\nWhich means there is a limit in how well we can estimate solely from the observed flux data, without incoporating  further stellar properties. </p>\n<p>To give an example of how limb darkening assumption can seriously affect transit depth estimation:<br>\nI fitted both exponential and quadratic on planet 3814020142's light curve.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13392195%2F850f9369b2a65c9bdac68d0e4a45b2dd%2F10.png?generation=1759344067867790&amp;alt=media\" alt=\"\"></p>\n<p>The residuals are similar, and the observed dip in flux is similar. However, there is a serious consequence of the choice of parametric limb darkening modeling. The best-ft quadratic model results in a disk-average intensity of 0.907 while the best-fit exponential results in a disk-average intensity of .884. This is because as you can see from the bottom graph, the exponential describes a much sharper drop in intensity towards the limbs.</p>\n<p>This means that choosing quadratic limb darkening results in 2.5% higher center-to-disk average intensity ratio. This leads to 1.8% higher (Rp/Rs)^2 estimation in the model. From my observations, polynomial limb darkening with higher linear coefficient tend to overestimate (Rp/Rs)^2.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3296854,
              "author_name": "Jeroen Cottaar",
              "author_url": "",
              "post_date": "2025-10-01T19:00:12.980000",
              "content": "<p>Interesting. Perhaps the 'missing element' of my prior that leads to the need for fudging is the use of different limb darkening models for different planets (I couldn't get any improvement by switching <em>all</em> planets to a different model, but perhaps I should have been flexible - would have been a nice opportunity for some further Bayesian tricks).</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3296150": "[Link to code to train and submit](https://www.kaggle.com/code/jeroencottaar/ariel2025-1st-place-bayesian-exoplanet-analysis)\nThis year I'm also releasing the full development history: [https://github.com/jcottaar/ariel2](https://github.com/jcottaar/ariel2) <br>\nI'd like to start by thanking the organizers for another great competition. I had a rather frustrating start (having somehow managed to lose all my code from last year), but greatly enjoyed it overall. The Bayesian approach is not used nearly as often as it should be these days, and I'm glad to have the opportunity to show its value here.\n\nMy solution is quite similar to last year's; my [NeurIPS talk](https://neurips.cc/virtual/2024/107111) from the associated workshop can be a useful introduction. I'll briefly discuss the differences compared to last year in a reply to this post.\n# 1. Introduction\n\nTransit analysis is the most common method for detecting and studying exoplanets. A transit occurs when a planet crosses in front of its star relative to Earth, obscuring part of the starlight. By observing how much the light is reduced (the transit depth) as a function of wavelength, we can learn the properties of the planet. But with a tiny planet passing in front of a huge star, this problem has a very low signal to noise ratio. The challenge to us: find the transit depth, including confidence intervals, from synthetic raw spectroscopic signals as the Ariel satellite might see them in the future (For a more extensive overview, see https://www.kaggle.com/competitions/ariel-data-challenge-2025/overview).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14984949%2Fa0860242fa947c058ce0794fa0494cc4%2FScreenshot%202024-11-01%20200453.png?generation=1730488072791970&alt=media)\n\nThis competition is at first glance a retread of a similar competition last year, but the data is far more realistic this time around. I think I figured out some of the changes, but not all of them - as we'll see later on.<br>\nNote that there are two sensors in play, the single-wavelength FGS sensor and the spectroscopic AIRS sensor. In this writeup I'll focus mainly on the AIRS sensor.<br>\nMy general approach consists of three steps, as visualized below:\n- **Preprocessing**: starting from the raw sensor data (expressed in counts), compute the amplitude signal *S(λ,t)* as a function of wavelength *λ* and time *t*.\n- **Bayesian Inference**: from *S(λ,t)*,  compute the inferred transit depth *D(λ)* and an uncertainty margin on this *σ(λ)* . This step is done using a principled Bayesian approach.\n- **Fudging**: throw the principles out the window, and apply various fudge factors to *D(λ)* and *σ(λ)* . These fudge factors are chosen to optimize score on the training data.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F14984949%2Fa0b169937d51e43db62a030def487f5c%2FScreenshot%202025-09-29%20212745.png?generation=1759174122624685&alt=media)\n\nIn sections 2, 3, and 4, I'll discuss each of these steps in turn. I'll finish with some closing thoughts in section 5 - mainly discussing why my approach isn't actually working...\n# 2. Preprocessing\n\nThe raw data is in counts per pixel; we first need to apply several steps to come to a photon count per pixel. This is explained in [a notebook shared by the hosts](https://www.kaggle.com/code/gordonyip/calibrating-and-binning-ariel-data). I use a time binning of 5 frames for AIRS, and 50 frames for FGS (this puts them approximately in sync). Hot pixel invalidation is disabled. I also added a simple cosmic ray removal. <br>\nThe opportunity for some real improvement lies in the wavelength binning step, i.e. summing over the dispersion axis. The baseline is simple summing, but there's several effects in the data that make it possible to do better. Most notably, there's the infamous jitter, which we can see most clearly by doing a PCA (with some additional preprocessing) over all AIRS frames in the train data:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F14984949%2F6c3728e5a47c97f7371dc1f1d5fb8a27%2FScreenshot%202025-09-30%20090624.png?generation=1759216002725805&alt=media)\n\nThe jitter shapes presumably correspond to pointing error and defocus.\n\nThe jitter shapes sum to approximately zero over the dispersion axis; this is why simple summing works OK. It also means that any more advanced preprocessing scheme is forced to deal with the jitter, so there's a high barrier to entry. Still, there are several effects in the data that such a scheme can take advantage of:\n\n- There are invalid pixels.\n- The noise is not quite Poisson noise (simple summing minimized noise only for Poisson noise); the noise is lower for brighter pixels, and there are occasional very noisy pixels.\n- The jitter shapes don't quite sum to zero, presumably because detector calibration is not perfect.\n- There is a constant background signal that we'd like to remove.\n\nI spent more than half of my time on this competition trying to set up a principled preprocessing flow optimally handling all these effects. Unfortunately, I haven't managed this yet - you'll find in the code several disjointed steps attempting to deal with the above. They also tend to work only in certain configurations and combinations, which I don't yet fully understand. It still leads to an improvement of ~0.01 over simple summing - I'll crack this one next year...\n# 3. Bayesian Inference\n\nNow that we have the photon counts as a function of wavelength and time, we are ready to infer the exoplanet transit depth using Bayesian Inference. <br>\n**Bayesian Inference** (BI) is a powerful statistical approach, based around defining a *prior* (a statistical belief about reality) and *observations* (some form of new information). Using Bayes' law, we then combine these to find the *posterior* (an updated belief about reality). In our case, this means:\n- **Prior**: a description of the physics that affect the final measured signal, describing for example detector noise, drift, and the transit behavior (including the transit depth itself) as formal distributions.\n- **Observations**: the provided measurements.\n- **Posterior**: a breakdown of the observations into the various elements defined in the prior (see figure below). From this we can simply read out the desired transit depth. Importantly, the posterior is not just a single point, it's a distribution. By taking samples from this distribution we can find the required confidence intervals (and even full covariance matrices, although that is not asked of us here).\n\nThere are two key elements to applying BI: defining the prior (section 3.1), and applying Bayes' law to do the inference (section 3.2).\n## 3.1 Prior definition\n\nThe signal is decomposed into four key components, as shown visually here:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14984949%2Ffbb5d6dba1e6836c600b1473cd333121%2FScreenshot%202024-11-01%20153411.png?generation=1730488097922724&alt=media)\n\nThis table shows a detailed description of all prior elements:\n\n| Prior element                 | Description                                                                                                                                                                                                                                                                                                         | Tuning and hyperparameters                                                                                                                     | Degrees of freedom                     |\n| ----------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------- |\n| Noise                         | Uncorrelated Gaussian per time and wavelength                                                                                                                                                                                                                                                                       | Standard deviations found in preprocessing                                                                                                     | 1350 (FGS) + 317250 (AIRS)             |\n| Star spectrum                 | Uncorrelated value per wavelength                                                                                                                                                                                                                                                                                   | Not regularized (infinite sigma)                                                                                                               | 283                                    |\n| Drift                         | Third order polynomial over time and wavelength                                                                                                                                                                                                                                                                     | Not regularized (infinite sigma)                                                                                                               | 3 (FGS) + 12 (AIRS)                    |\n| Transit window                | Found using batman package, with the following free parameters:<br>- Transit mid-time t0 (separate for FGS and AIRS)<br>- Semi-major axis sma<br>- Period P<br>- Orbital inclination i<br>- Quadratic limb darkening u0 and u1 (including linear dependence on wavelength)<br>- Transit depth Rp^2/Rs^2 (see below) | Covariance matrix learned on training set using maximum likelihood estimation (MLE)                                                            | 11 (6 limb darkening, 2 t0, sma, P, i) |\n| Transit depth: mean           | Single value                                                                                                                                                                                                                                                                                                        | Lightly regularized (sigma=0.01)                                                                                                               | 1                                      |\n| Transit depth: variation      | Decomposed into three components below                                                                                                                                                                                                                                                                              | Scaling factor determined per planet using MLE (this is the only hyperparameter tuned during inference)                                        | N/A                                    |\n| Transit depth: variation FGS  | Single Gaussian value                                                                                                                                                                                                                                                                                               | Standard deviation found on training set                                                                                                       | 1                                      |\n| Transit depth: variation AIRS | Gaussian Process over wavelength, sum of two squared-exponential kernels                                                                                                                                                                                                                                            | Kernel hyperparameters found on training set using maximum likelihood estimation                                                               | 282                                    |\n| Transit depth: variation PCA  | Fixed basis functions obtained from PCA analysis                                                                                                                                                                                                                                                                    | PCA shapes and variation magnitudes found on training data (this year there seems to be no train-test population shift for the transit depths) | 5                                      |\n## 3.2 Solver\n\nHaving defined the prior, finding the posterior is 'just' a matter of applying Bayes' law. The prior is fully Gaussian, which means the inference should just be a single matrix operation. But as usual, there's a snag. In our case, it's that the relationship between the prior parameters and the observation is non-linear (because of the non-linear modeling in batman, and because some prior elements are multiplied rather than added). <br> \nTo deal with this non-linearity, we linearize the model around the mean of the posterior. This is repeated iteratively (i.e. apply BI -> linearize around the new mean of the posterior -> repeat). Each step we also update the single hyperparameter (the scaling of the transit depth) using gradient descent. <br> \nThis approach does mean we need a decent starting point, which is handled as follows:\n\n- **Grid search**: sweep over a grid of transit depth and transit mid-time to find initial values. Other transit parameters are set to the given approximation per planet (with no limb darkening).\n- **BFGS**: fit drift and all transit parameters, using BFGS as implemented in ```scipy.optimize.minimize```. In this step we only consider the mean over wavelengths of the AIRS signal. \n- **Bayesian Inference**: the full non-linear BI solver as described above. My winning submission uses 8 iterations of the solver; the code linked here only does 4 iterations to fit training and inference into the 9 hour submission limit (this only costs 0.001 in the score).\n\nThe prediction of the transit depth and the covariance matrix of its uncertainty can be read out directly in the posterior. The covariance matrix is approximated based on 200 samples; this is mainly a holdover from last year, when there were too many parameters to construct it explicitly.\n\nThere are two additional tricks that I apply if the residual of the first two steps (grid search->BFGS) is suspiciously high:\n- Use transit parameters for a different planet as starting point.\n- Try chopping off the first and last few frames of the signal (these sometimes have a high residual for reasons I don't understand).\n  \n# 4. Fudging\n\nThe prediction and uncertainty obtained above can still be improved further with some post-hoc calibration. Note that this is anathema to a proper Bayesian approach - more on that below.\n\nWe adapt the mean of the transit depth per sensor based on fitting:\n - A constant offset\n - A multiplicative factor\n - Dependence on the first limb darkening parameter\n\nThe variation of the transit depth over wavelength is not fudged, i.e. not changed in this step.\n\nWe adapt the uncertainty of the transit depth per sensor based on fitting:\n- A constant offset\n- A multiplicative factor for the mean of the transit depth\n- A multiplicative factor for the variation of the transit depth over wavelength\n- Dependence on the magnitude of the variation of the transit depth over wavelength\n\nIn total, the fudging has 12 parameters to fit. These are determined on the full training set, based on whichever values optimize the competition metric.\n# 5. Final thoughts - we're missing something!\n\nThe fact that fudging helps the score is extremely worrying. And it's not subtle: last year, my solution without fudging could have gotten 1st place; my solution this year without fudging would have ended up around ~20th place. It means there's something missing in the prior, i.e. it does not describe reality (or rather the synthetic data generation process) accurately.\n\nI've spent a lot of time trying to find this gap, but haven't gotten much closer. I've come to believe that the issue lies in the transit modeling, but none of the variation I tried (such as different limb darkening models) helped. I'm convinced that finding whatever's going on could boost our scores to well over 0.700.\n\nSome more direct evidence of this issue can be found by comparing different transits for the same planet; the plot below shows the ratio of the transit profiles (with some preprocessing including low pass filtering). Nothing in the prior (and nothing else I can come up with) can explain this. See also [this discussion](https://www.kaggle.com/competitions/ariel-data-challenge-2025/discussion/602425) and [the code](https://www.kaggle.com/code/jeroencottaar/demonstrate-apparent-label-issue).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14984949%2F1b6f904fd58b66870aa7febac31ea864%2FScreenshot%202025-09-30%20111015.png?generation=1759223442391305&alt=media)\nI hope the organizers are willing to reveal some clues about this, because I don't think we'll figure it out next year either without some help.\n\nFinally, I'll wrap it up with an overview of how much various elements of my solution contribute to the score:\n\n| Change to model                                                                                                            | Impact on private test score     |\n| -------------------------------------------------------------------------------------------------------------------------- | -------------------------------- |\n| Simplified loader (host's flow + cosmic ray removal), no jitter correction or background removal                           | -0.013                           |\n| Don't remove first or last few frames of suspicious transits                                                               | -0.002                           |\n| Don't use Gaussian Process for transit depth, just use PCA<br>Note: this variation does not involve any Gaussian Processes | -0.007                           |\n| Don't use PCA for transit depth, just Gaussian Process                                                                     | -0.011                           |\n| Disable all fudging - rely purely on the outcome of the Bayesian predictions                                               | -0.152<br>(Last year: -0.000...) |\n| Replace all fudging by a single multiplicative factor for the uncertainty                                                  | -0.017                           |\n| Disable fudging of sigma prediction based on AIRS variation                                                                | -0.005                           |\n| Disable fudging of mean prediction based on limb darkening parameters                                                      | -0.005                           |\n| Don't regularize the 11 transit parameters in prior                                                                        | -0.003                           |\n\n",
    "3296154": "Key differences with my [2024 solution](https://www.kaggle.com/competitions/ariel-data-challenge-2024/writeups/jeroen-cottaar-2nd-place-solution-pure-bayesian-in):\n\n- Use batman to model transits.\n- Added limb darkening to the transit model.\n- PCA for transit depths is done on the training set, rather than on the test set (no apparent population shift this year). Also uses 5 components instead of 1.\n- Simplified Gaussian Process for transit depth (hardly needed still - a solution without any Gaussian Processes anywhere in the prior scores 0.642).\n- Drift is modeled as a third-order polynomial rather than a Gaussian Process; less realistic, but matches the synthetic data generation process.\n- Initial solver (grid search and BFGS) to find a good starting point for the Bayesian solver; mainly needed because of limb darkening.\n- Progress towards proper jitter correction and background removal in preprocessing.\n- Extensive post-hoc calibration (i.e. fudging) to deal with gaps in the prior - no fudging was needed last year.\n- Significantly reduced computation time and memory usage with code improvements, including moving all preprocessing to GPU.",
    "3298395": "Congrats!\n\nI wonder do you have some high-level insights on why Bayesian without NN work well in 2024 and 2025.  As our team tried end-to-end deep neural networks (DNNs) which is extremely vulnerable and still need ensemble with other methods for robustness.\n\nMy insight is that \n1) The actual noise functions is linearly like so that DNNs overfit it.\n2) The training data is small so that DNNs overfit it.\n3) The data sample is not infomation-abundent along with outliers so that DNNs overfit it.",
    "3296578": "Thanks for the nice summary, @jeroencottaar !\nIn your model you did not restrict the quadratic u1/u2 LD coefficients, so they could converge to Limb Brightening regime, right?",
    "3296396": "Congrats for the 1st place finish!\n\nI'm curious how well your model fits the ingress/egress period.\n\nIn my model, I also had to do adjustment to the mean transit depth and I found the likely reason to be the limb darkening model not fitting well near the limbs (as well as lack of observations near the center), which greatly affects the transit depth distortion (measured depth vs (Rp/Rs)2 ) since the stellar flux is heavily dependent on the limb darkening assumption."
  }
}