{
  "id": 609700,
  "title": "19th Place Solution",
  "url": "/competitions/ariel-data-challenge-2025/discussion/609700",
  "author_name": "Olly Powell",
  "post_date": "2025-09-29T04:48:10.158000",
  "votes": 2,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Thanks very much to the organisers, and my fellow competitors for an interesting and challenging competition.  I really appreciate these challenges, even when they're not directly in my field of research.  The cross disciplinary problem solving, and more general lessons are invaluable to me.</p>\n<p>My final solution was a mixture of careful curve fitting, binning in both wavelength and time, then interpolation in wavelength and using nested LGBM models to improve the fitted estimates, and also to estimate the uncertainty.</p>\n<p>The full training and inference notebook is on GitHub  <a href=\"https://github.com/Wologman/Kaggle_Ariel_Data_Challenge_25/blob/main/ariel25-batman-lgbm.ipynb\" target=\"_blank\">here</a></p>\n<p>The first stage of the curve fitting is shown below.  Here I fitted a 'baseline' curve to a time-binned and wavelength-averaged signal for each planet.  I relied on only one breakpoint.  Both the out-of-transit part, and the fully occluded part share the same linear term.  I figured the fit didn't need to be perfect at this stage, just consistent, and handle edge cases gracefully, as I would later try to correct for any shortcomings with my boosting models.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F5224109%2F0e53600c7fc7fbaf1ed4b283a6d07fed%2Fexample_0.png?generation=1759120384269436&amp;alt=media\" alt=\"\"></p>\n<p>I then used the centres, breakpoints and linear correction term to produce 13 more wavelength-specific fits.  These were later interpolated to fill out all 283 predictions.  I used the features from these 13 fits to produce an LGBM based correction, and a second LGBM model with a nested data split to estimate the uncertainty.    I also used the <a href=\"https://lkreidberg.github.io/batman/docs/html/index.html\" target=\"_blank\">batman model</a> for curve fit only to generate a additional features for the LGBM model.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F5224109%2Fbe7aec64eac11f6079e5adb0ed476703%2FScreenshot%202025-09-26%20205639.png?generation=1759120640858346&amp;alt=media\" alt=\"\"> </p>\n<p>Thanks again, and I hope to see everyone next time.</p>",
  "messages": [
    {
      "id": 3295547,
      "postDate": "2025-09-29T04:48:10.160Z",
      "content": "<p>Thanks very much to the organisers, and my fellow competitors for an interesting and challenging competition.  I really appreciate these challenges, even when they're not directly in my field of research.  The cross disciplinary problem solving, and more general lessons are invaluable to me.</p>\n<p>My final solution was a mixture of careful curve fitting, binning in both wavelength and time, then interpolation in wavelength and using nested LGBM models to improve the fitted estimates, and also to estimate the uncertainty.</p>\n<p>The full training and inference notebook is on GitHub  <a href=\"https://github.com/Wologman/Kaggle_Ariel_Data_Challenge_25/blob/main/ariel25-batman-lgbm.ipynb\" target=\"_blank\">here</a></p>\n<p>The first stage of the curve fitting is shown below.  Here I fitted a 'baseline' curve to a time-binned and wavelength-averaged signal for each planet.  I relied on only one breakpoint.  Both the out-of-transit part, and the fully occluded part share the same linear term.  I figured the fit didn't need to be perfect at this stage, just consistent, and handle edge cases gracefully, as I would later try to correct for any shortcomings with my boosting models.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F5224109%2F0e53600c7fc7fbaf1ed4b283a6d07fed%2Fexample_0.png?generation=1759120384269436&amp;alt=media\" alt=\"\"></p>\n<p>I then used the centres, breakpoints and linear correction term to produce 13 more wavelength-specific fits.  These were later interpolated to fill out all 283 predictions.  I used the features from these 13 fits to produce an LGBM based correction, and a second LGBM model with a nested data split to estimate the uncertainty.    I also used the <a href=\"https://lkreidberg.github.io/batman/docs/html/index.html\" target=\"_blank\">batman model</a> for curve fit only to generate a additional features for the LGBM model.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F5224109%2Fbe7aec64eac11f6079e5adb0ed476703%2FScreenshot%202025-09-26%20205639.png?generation=1759120640858346&amp;alt=media\" alt=\"\"> </p>\n<p>Thanks again, and I hope to see everyone next time.</p>",
      "rawMarkdown": "Thanks very much to the organisers, and my fellow competitors for an interesting and challenging competition.  I really appreciate these challenges, even when they're not directly in my field of research.  The cross disciplinary problem solving, and more general lessons are invaluable to me.\n\nMy final solution was a mixture of careful curve fitting, binning in both wavelength and time, then interpolation in wavelength and using nested LGBM models to improve the fitted estimates, and also to estimate the uncertainty.\n\nThe full training and inference notebook is on GitHub  [here](https://github.com/Wologman/Kaggle_Ariel_Data_Challenge_25/blob/main/ariel25-batman-lgbm.ipynb)\n\nThe first stage of the curve fitting is shown below.  Here I fitted a 'baseline' curve to a time-binned and wavelength-averaged signal for each planet.  I relied on only one breakpoint.  Both the out-of-transit part, and the fully occluded part share the same linear term.  I figured the fit didn't need to be perfect at this stage, just consistent, and handle edge cases gracefully, as I would later try to correct for any shortcomings with my boosting models.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F5224109%2F0e53600c7fc7fbaf1ed4b283a6d07fed%2Fexample_0.png?generation=1759120384269436&alt=media)\n\nI then used the centres, breakpoints and linear correction term to produce 13 more wavelength-specific fits.  These were later interpolated to fill out all 283 predictions.  I used the features from these 13 fits to produce an LGBM based correction, and a second LGBM model with a nested data split to estimate the uncertainty.    I also used the [batman model](https://lkreidberg.github.io/batman/docs/html/index.html) for curve fit only to generate a additional features for the LGBM model.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F5224109%2Fbe7aec64eac11f6079e5adb0ed476703%2FScreenshot%202025-09-26%20205639.png?generation=1759120640858346&alt=media) \n\nThanks again, and I hope to see everyone next time.",
      "votes": 2
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3295547": "Thanks very much to the organisers, and my fellow competitors for an interesting and challenging competition.  I really appreciate these challenges, even when they're not directly in my field of research.  The cross disciplinary problem solving, and more general lessons are invaluable to me.\n\nMy final solution was a mixture of careful curve fitting, binning in both wavelength and time, then interpolation in wavelength and using nested LGBM models to improve the fitted estimates, and also to estimate the uncertainty.\n\nThe full training and inference notebook is on GitHub  [here](https://github.com/Wologman/Kaggle_Ariel_Data_Challenge_25/blob/main/ariel25-batman-lgbm.ipynb)\n\nThe first stage of the curve fitting is shown below.  Here I fitted a 'baseline' curve to a time-binned and wavelength-averaged signal for each planet.  I relied on only one breakpoint.  Both the out-of-transit part, and the fully occluded part share the same linear term.  I figured the fit didn't need to be perfect at this stage, just consistent, and handle edge cases gracefully, as I would later try to correct for any shortcomings with my boosting models.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F5224109%2F0e53600c7fc7fbaf1ed4b283a6d07fed%2Fexample_0.png?generation=1759120384269436&alt=media)\n\nI then used the centres, breakpoints and linear correction term to produce 13 more wavelength-specific fits.  These were later interpolated to fill out all 283 predictions.  I used the features from these 13 fits to produce an LGBM based correction, and a second LGBM model with a nested data split to estimate the uncertainty.    I also used the [batman model](https://lkreidberg.github.io/batman/docs/html/index.html) for curve fit only to generate a additional features for the LGBM model.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F5224109%2Fbe7aec64eac11f6079e5adb0ed476703%2FScreenshot%202025-09-26%20205639.png?generation=1759120640858346&alt=media) \n\nThanks again, and I hope to see everyone next time."
  }
}