{
  "id": 609420,
  "title": "33rd place solution",
  "url": "/competitions/ariel-data-challenge-2025/discussion/609420",
  "author_name": "D. DiMonte",
  "post_date": "2025-09-26T14:19:21.667000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<p><strong>Data exploration</strong></p>\n<p><a href=\"https://www.kaggle.com/code/ddimonte/understanding-the-data\" target=\"_blank\">This notebook</a> provides detailed exploration of the data format and applies processing steps borrowed from the <a href=\"https://www.kaggle.com/code/gordonyip/update-calibrating-and-binning-astronomical-data\" target=\"_blank\">binning and processing notebook</a>. Key changes in this notebook include minimal time binning (just enough to standardize AIRS-CH0 and FGS1 lengths), updated masking behavior, the addition of inpainting for missing data, and extensive before/after visualizations.</p>\n<p>In <a href=\"https://www.kaggle.com/code/ddimonte/process-data\" target=\"_blank\">this notebook</a> we fit a polynomial baseline to the out-of-transit regions of each light curve, using smoothing and change point detection to identify the transit onset and offset. However, we found that some transits in our dataset are incomplete or contain somewhat complex background changes, making baseline normalization unreliable for all cases without more careful consideration. To address this, we briefly attempted to fit a BATMAN transit model to the data, but ultimately decided not to pursue this further due to time constraints and the need for robust validation.</p>\n<p>We then investigated the ratio between the provided ground truth transit depth and the maximum observed transit depth from the light curve. We found that the ground truth value is not always at the minimum of the light curve or a fixed fraction of the fitted model, and can even be lower than the observed minimum for some planets. This suggests that other parameters besides for the visible transit depth in the light curve alone play a significant role in determining the true transit depth such as planet and star data or the shape of the transit curve. This is likely due to transit characteristics and limb darkening.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4184177%2Febf65ba60085550dba27cb4cde82f708%2FIllustration3.png?generation=1759267472752844&amp;alt=media\" alt=\"Light curve examples including group truth\"></p>\n<p><em>Figure 1:</em> Light curve examples including ground truth estimate.</p>\n<p><strong>Preprocessing</strong> </p>\n<p><a href=\"https://www.kaggle.com/code/ddimonte/submission-code\" target=\"_blank\">This notebook</a> contains the final submission code with all the necessary preprocessing on the test data. Much of the code was taken from the <a href=\"https://www.kaggle.com/code/gordonyip/update-calibrating-and-binning-astronomical-data\" target=\"_blank\">binning and processing notebook</a>. We also use code from the ruptures library.\nPreprocessing steps:</p>\n<ul>\n<li>Load and Calibrate Detector Signals</li>\n<li>Load Calibration Files and Apply Detector Cleaning<ul>\n<li>Mask hot/dead pixels: Replace unusable data using mask maps and dark files.</li>\n<li>Non-linearity correction</li>\n<li>Flat Field Correction: Correct for pixel-to-pixel sensitivity variations using loaded flat field maps.</li>\n<li>Dark subtraction: Subtract the appropriate dark current background.</li>\n<li>Correlated Double Sampling</li></ul></li>\n<li>Time binning 12 to FGS1 to align time dimensions.</li>\n<li>Inpainting: Fill remaining missing or masked regions using biharmonic inpainting treating time as channels.</li>\n<li>Spatial Summation: Sum along spatial axes to produce light curves and prepare input shapes appropriate for the model.</li>\n<li>Median filtering kernel size 101</li>\n<li>Downsampling stride 10</li>\n</ul>\n<p><strong>Model</strong></p>\n<p>The overall model architecture is shown in Figure 2. The preprocessed data [wavelength x time] is normalized per wavelength and passed through a Time-Reducing Residual Stack. The resulting feature map is pooled and then concatenated with the wavelength means z-scores, the wavelength standard deviations z-scores, and the normalized planet features. These scores are calculated using summary statistics from the entire dataset in order to incorporate information about the sample's placement in the population distribution into the model after our earlier observations shown in Figure 1. The combined vector is passed through a fully connected layer, whose output feeds three separate fully connected output heads. The first two heads are used to compute the per-wavelength prediction, and the third head is used for the uncertainty (sigma) values. For the prediction, we calculate the relative change between the first two outputs. For the third output, we train the model to output log(sigma) so we take the exponential of the third output head for our final sigma value.</p>\n<p>The Time-Reducing Residual Block details are shown in Figure 3. In the figure N represents the block number from 1 to 6 since 6 blocks are stacked in the final model. The wavelength dimension uses circular padding and increasing kernel dilation 3 so that all of the wavelengths' features are convolved with each other within 6 stacked residual blocks. The time dimension uses 0 padding and a kernel dilation of 1 so the time field of view remains within the nearest times from small increments to large increments as the time dimension is pooled in successive blocks. In our model we used a time dimension pooling kernel of 4 for the first 4 blocks and then eased off down to 2 to maintain the desired data size for a 6 block stack and to prioritize faster compression of the time dimension earlier.</p>\n<p>We trained the model with a combined loss utilizing both MSE as well as the GLL error used for scoring.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4184177%2Fdb9296173f2451ecbdddb8173866cffb%2Fmodellight.drawio.svg?generation=1773942873498721&amp;alt=media\" alt=\"Model flowchart\"></p>\n<p><em>Figure 2:</em> Model flowchart. The input is the pre-processed data array x, and the outputs are the predictions for the requested wavelengths and the associated sigma values.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4184177%2F6be839924011f903d4b8284dabf35127%2Fdiagram-Page-2_light.drawio.svg?generation=1772150272461207&amp;alt=media\" alt=\"Detailed view of the Time Reducing Residual Block N for N as 1 through 6\"></p>\n<p><em>Figure 3:</em> Detailed view of the Time Reducing Residual Block N for N as 1 through 6.</p>\n<p><strong>Augmentation</strong></p>\n<p>To improve model generalization and robustness and decrease overfitting, we implemented several data augmentation techniques during training:</p>\n<ul>\n<li><p>Time Flipping: Each light curve was flipped along the time axis to expose the model to both forward and reversed transit scenarios.</p></li>\n<li><p>Variable Downsampling and Offsets: We used strides of 8, 9, or 10 and multiple random offsets, simulating the effect of time running at different speeds and producing light curves with varying sampling densities. This approach also increased the representation of incomplete transits, helping the model learn to handle edge cases and irregular data spans. Because the signals were median-filtered (kernel size 101), the incremental effect of offset variation is modest, but still introduces some diversity.</p></li>\n<li><p>Additive Linear Trend: For each wavelength channel, we injected a random linear signal with a maximum slope set equal to the channel’s own range. The maximum was set per wavelength, but the randomization was done per planet observation.</p></li>\n<li><p>Additive Sinusoidal Trend: For each wavelength channel we injected a random sinusoidal signal with low frequency.</p></li>\n<li><p>Planet Parameter Noise: Small, zero-mean Gaussian noise (σ = 0.1) was added to each set of normalized planet parameters fed into the model. This reduces overfitting and allows the network to be more robust to small errors or uncertainty in planet/star parameters.</p></li>\n</ul>\n<p><strong>Ensembling methods</strong></p>\n<p>To aggregate the model predictions for each planet, we explored several ensembling approaches seen in our <a href=\"https://www.kaggle.com/code/ddimonte/submission-code\" target=\"_blank\">submission notebook</a>. We found that the basic average provided the best results due to how we were predicting our values and sigmas. Each planet had predictions from (potentially) multiple observations, 5 different offsets of stride 10, and multiple model predictions. Methods we tried:</p>\n<ul>\n<li>Basic Average: A simple average of all predictions for each wavelength channel for each planet, both for the value and the sigma.</li>\n<li>Best Sigma Selection: Selecting only the single set of predictions (per planet) with the lowest mean predicted uncertainty across all wavelengths.</li>\n<li>Best N Average: A simple average of the N rows of data per planet that had the lowest mean sigma values.</li>\n<li>Weighted Averaging: Weighted each prediction by the inverse of its predicted uncertainty, computing weighted means and associated uncertainties for each wavelength.</li>\n</ul>\n<p><strong>Other notes:</strong></p>\n<p><em>Stratification</em></p>\n<p>We explored stratification based on star and planet features [<a href=\"https://www.kaggle.com/code/ddimonte/stratification\" target=\"_blank\">notebook</a>], hoping to create folds that better represent different subclasses or difficulty levels. However, the lack of strong natural clusters, weak correlations between features and prediction error, and overlapping distributions showed that meaningful strata did not exist in the data for the features we analyzed. As a result, simple random folds were used.</p>\n<p><em>Other models</em></p>\n<p><strong>Citations</strong></p>\n<ul>\n<li>C. Truong, L. Oudre, N. Vayatis. Selective review of offline change point detection methods. Signal Processing, 167:107299, 2020.</li>\n</ul>",
  "messages": [
    {
      "id": 3294676,
      "postDate": "2025-09-26T14:19:21.667Z",
      "content": "<p><strong>Data exploration</strong></p>\n<p><a href=\"https://www.kaggle.com/code/ddimonte/understanding-the-data\" target=\"_blank\">This notebook</a> provides detailed exploration of the data format and applies processing steps borrowed from the <a href=\"https://www.kaggle.com/code/gordonyip/update-calibrating-and-binning-astronomical-data\" target=\"_blank\">binning and processing notebook</a>. Key changes in this notebook include minimal time binning (just enough to standardize AIRS-CH0 and FGS1 lengths), updated masking behavior, the addition of inpainting for missing data, and extensive before/after visualizations.</p>\n<p>In <a href=\"https://www.kaggle.com/code/ddimonte/process-data\" target=\"_blank\">this notebook</a> we fit a polynomial baseline to the out-of-transit regions of each light curve, using smoothing and change point detection to identify the transit onset and offset. However, we found that some transits in our dataset are incomplete or contain somewhat complex background changes, making baseline normalization unreliable for all cases without more careful consideration. To address this, we briefly attempted to fit a BATMAN transit model to the data, but ultimately decided not to pursue this further due to time constraints and the need for robust validation.</p>\n<p>We then investigated the ratio between the provided ground truth transit depth and the maximum observed transit depth from the light curve. We found that the ground truth value is not always at the minimum of the light curve or a fixed fraction of the fitted model, and can even be lower than the observed minimum for some planets. This suggests that other parameters besides for the visible transit depth in the light curve alone play a significant role in determining the true transit depth such as planet and star data or the shape of the transit curve. This is likely due to transit characteristics and limb darkening.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4184177%2Febf65ba60085550dba27cb4cde82f708%2FIllustration3.png?generation=1759267472752844&amp;alt=media\" alt=\"Light curve examples including group truth\"></p>\n<p><em>Figure 1:</em> Light curve examples including ground truth estimate.</p>\n<p><strong>Preprocessing</strong> </p>\n<p><a href=\"https://www.kaggle.com/code/ddimonte/submission-code\" target=\"_blank\">This notebook</a> contains the final submission code with all the necessary preprocessing on the test data. Much of the code was taken from the <a href=\"https://www.kaggle.com/code/gordonyip/update-calibrating-and-binning-astronomical-data\" target=\"_blank\">binning and processing notebook</a>. We also use code from the ruptures library.\nPreprocessing steps:</p>\n<ul>\n<li>Load and Calibrate Detector Signals</li>\n<li>Load Calibration Files and Apply Detector Cleaning<ul>\n<li>Mask hot/dead pixels: Replace unusable data using mask maps and dark files.</li>\n<li>Non-linearity correction</li>\n<li>Flat Field Correction: Correct for pixel-to-pixel sensitivity variations using loaded flat field maps.</li>\n<li>Dark subtraction: Subtract the appropriate dark current background.</li>\n<li>Correlated Double Sampling</li></ul></li>\n<li>Time binning 12 to FGS1 to align time dimensions.</li>\n<li>Inpainting: Fill remaining missing or masked regions using biharmonic inpainting treating time as channels.</li>\n<li>Spatial Summation: Sum along spatial axes to produce light curves and prepare input shapes appropriate for the model.</li>\n<li>Median filtering kernel size 101</li>\n<li>Downsampling stride 10</li>\n</ul>\n<p><strong>Model</strong></p>\n<p>The overall model architecture is shown in Figure 2. The preprocessed data [wavelength x time] is normalized per wavelength and passed through a Time-Reducing Residual Stack. The resulting feature map is pooled and then concatenated with the wavelength means z-scores, the wavelength standard deviations z-scores, and the normalized planet features. These scores are calculated using summary statistics from the entire dataset in order to incorporate information about the sample's placement in the population distribution into the model after our earlier observations shown in Figure 1. The combined vector is passed through a fully connected layer, whose output feeds three separate fully connected output heads. The first two heads are used to compute the per-wavelength prediction, and the third head is used for the uncertainty (sigma) values. For the prediction, we calculate the relative change between the first two outputs. For the third output, we train the model to output log(sigma) so we take the exponential of the third output head for our final sigma value.</p>\n<p>The Time-Reducing Residual Block details are shown in Figure 3. In the figure N represents the block number from 1 to 6 since 6 blocks are stacked in the final model. The wavelength dimension uses circular padding and increasing kernel dilation 3 so that all of the wavelengths' features are convolved with each other within 6 stacked residual blocks. The time dimension uses 0 padding and a kernel dilation of 1 so the time field of view remains within the nearest times from small increments to large increments as the time dimension is pooled in successive blocks. In our model we used a time dimension pooling kernel of 4 for the first 4 blocks and then eased off down to 2 to maintain the desired data size for a 6 block stack and to prioritize faster compression of the time dimension earlier.</p>\n<p>We trained the model with a combined loss utilizing both MSE as well as the GLL error used for scoring.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4184177%2Fdb9296173f2451ecbdddb8173866cffb%2Fmodellight.drawio.svg?generation=1773942873498721&amp;alt=media\" alt=\"Model flowchart\"></p>\n<p><em>Figure 2:</em> Model flowchart. The input is the pre-processed data array x, and the outputs are the predictions for the requested wavelengths and the associated sigma values.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4184177%2F6be839924011f903d4b8284dabf35127%2Fdiagram-Page-2_light.drawio.svg?generation=1772150272461207&amp;alt=media\" alt=\"Detailed view of the Time Reducing Residual Block N for N as 1 through 6\"></p>\n<p><em>Figure 3:</em> Detailed view of the Time Reducing Residual Block N for N as 1 through 6.</p>\n<p><strong>Augmentation</strong></p>\n<p>To improve model generalization and robustness and decrease overfitting, we implemented several data augmentation techniques during training:</p>\n<ul>\n<li><p>Time Flipping: Each light curve was flipped along the time axis to expose the model to both forward and reversed transit scenarios.</p></li>\n<li><p>Variable Downsampling and Offsets: We used strides of 8, 9, or 10 and multiple random offsets, simulating the effect of time running at different speeds and producing light curves with varying sampling densities. This approach also increased the representation of incomplete transits, helping the model learn to handle edge cases and irregular data spans. Because the signals were median-filtered (kernel size 101), the incremental effect of offset variation is modest, but still introduces some diversity.</p></li>\n<li><p>Additive Linear Trend: For each wavelength channel, we injected a random linear signal with a maximum slope set equal to the channel’s own range. The maximum was set per wavelength, but the randomization was done per planet observation.</p></li>\n<li><p>Additive Sinusoidal Trend: For each wavelength channel we injected a random sinusoidal signal with low frequency.</p></li>\n<li><p>Planet Parameter Noise: Small, zero-mean Gaussian noise (σ = 0.1) was added to each set of normalized planet parameters fed into the model. This reduces overfitting and allows the network to be more robust to small errors or uncertainty in planet/star parameters.</p></li>\n</ul>\n<p><strong>Ensembling methods</strong></p>\n<p>To aggregate the model predictions for each planet, we explored several ensembling approaches seen in our <a href=\"https://www.kaggle.com/code/ddimonte/submission-code\" target=\"_blank\">submission notebook</a>. We found that the basic average provided the best results due to how we were predicting our values and sigmas. Each planet had predictions from (potentially) multiple observations, 5 different offsets of stride 10, and multiple model predictions. Methods we tried:</p>\n<ul>\n<li>Basic Average: A simple average of all predictions for each wavelength channel for each planet, both for the value and the sigma.</li>\n<li>Best Sigma Selection: Selecting only the single set of predictions (per planet) with the lowest mean predicted uncertainty across all wavelengths.</li>\n<li>Best N Average: A simple average of the N rows of data per planet that had the lowest mean sigma values.</li>\n<li>Weighted Averaging: Weighted each prediction by the inverse of its predicted uncertainty, computing weighted means and associated uncertainties for each wavelength.</li>\n</ul>\n<p><strong>Other notes:</strong></p>\n<p><em>Stratification</em></p>\n<p>We explored stratification based on star and planet features [<a href=\"https://www.kaggle.com/code/ddimonte/stratification\" target=\"_blank\">notebook</a>], hoping to create folds that better represent different subclasses or difficulty levels. However, the lack of strong natural clusters, weak correlations between features and prediction error, and overlapping distributions showed that meaningful strata did not exist in the data for the features we analyzed. As a result, simple random folds were used.</p>\n<p><em>Other models</em></p>\n<p><strong>Citations</strong></p>\n<ul>\n<li>C. Truong, L. Oudre, N. Vayatis. Selective review of offline change point detection methods. Signal Processing, 167:107299, 2020.</li>\n</ul>",
      "rawMarkdown": "**Data exploration**\n\n[This notebook](https://www.kaggle.com/code/ddimonte/understanding-the-data) provides detailed exploration of the data format and applies processing steps borrowed from the [binning and processing notebook](https://www.kaggle.com/code/gordonyip/update-calibrating-and-binning-astronomical-data). Key changes in this notebook include minimal time binning (just enough to standardize AIRS-CH0 and FGS1 lengths), updated masking behavior, the addition of inpainting for missing data, and extensive before/after visualizations.\n\nIn [this notebook](https://www.kaggle.com/code/ddimonte/process-data) we fit a polynomial baseline to the out-of-transit regions of each light curve, using smoothing and change point detection to identify the transit onset and offset. However, we found that some transits in our dataset are incomplete or contain somewhat complex background changes, making baseline normalization unreliable for all cases without more careful consideration. To address this, we briefly attempted to fit a BATMAN transit model to the data, but ultimately decided not to pursue this further due to time constraints and the need for robust validation.\n\nWe then investigated the ratio between the provided ground truth transit depth and the maximum observed transit depth from the light curve. We found that the ground truth value is not always at the minimum of the light curve or a fixed fraction of the fitted model, and can even be lower than the observed minimum for some planets. This suggests that other parameters besides for the visible transit depth in the light curve alone play a significant role in determining the true transit depth such as planet and star data or the shape of the transit curve. This is likely due to transit characteristics and limb darkening.\n\n![Light curve examples including group truth](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4184177%2Febf65ba60085550dba27cb4cde82f708%2FIllustration3.png?generation=1759267472752844&alt=media)\n\n*Figure 1:* Light curve examples including ground truth estimate.\n\n**Preprocessing** \n\n[This notebook](https://www.kaggle.com/code/ddimonte/submission-code) contains the final submission code with all the necessary preprocessing on the test data. Much of the code was taken from the [binning and processing notebook](https://www.kaggle.com/code/gordonyip/update-calibrating-and-binning-astronomical-data). We also use code from the ruptures library.\nPreprocessing steps:\n- Load and Calibrate Detector Signals\n- Load Calibration Files and Apply Detector Cleaning\n    - Mask hot/dead pixels: Replace unusable data using mask maps and dark files.\n    - Non-linearity correction\n    - Flat Field Correction: Correct for pixel-to-pixel sensitivity variations using loaded flat field maps.\n    - Dark subtraction: Subtract the appropriate dark current background.\n    - Correlated Double Sampling\n- Time binning 12 to FGS1 to align time dimensions.\n- Inpainting: Fill remaining missing or masked regions using biharmonic inpainting treating time as channels.\n- Spatial Summation: Sum along spatial axes to produce light curves and prepare input shapes appropriate for the model.\n- Median filtering kernel size 101\n- Downsampling stride 10\n\n**Model**\n\nThe overall model architecture is shown in Figure 2. The preprocessed data [wavelength x time] is normalized per wavelength and passed through a Time-Reducing Residual Stack. The resulting feature map is pooled and then concatenated with the wavelength means z-scores, the wavelength standard deviations z-scores, and the normalized planet features. These scores are calculated using summary statistics from the entire dataset in order to incorporate information about the sample's placement in the population distribution into the model after our earlier observations shown in Figure 1. The combined vector is passed through a fully connected layer, whose output feeds three separate fully connected output heads. The first two heads are used to compute the per-wavelength prediction, and the third head is used for the uncertainty (sigma) values. For the prediction, we calculate the relative change between the first two outputs. For the third output, we train the model to output log(sigma) so we take the exponential of the third output head for our final sigma value.\n\nThe Time-Reducing Residual Block details are shown in Figure 3. In the figure N represents the block number from 1 to 6 since 6 blocks are stacked in the final model. The wavelength dimension uses circular padding and increasing kernel dilation 3<sup>N-1</sup> so that all of the wavelengths' features are convolved with each other within 6 stacked residual blocks. The time dimension uses 0 padding and a kernel dilation of 1 so the time field of view remains within the nearest times from small increments to large increments as the time dimension is pooled in successive blocks. In our model we used a time dimension pooling kernel of 4 for the first 4 blocks and then eased off down to 2 to maintain the desired data size for a 6 block stack and to prioritize faster compression of the time dimension earlier.\n\nWe trained the model with a combined loss utilizing both MSE as well as the GLL error used for scoring.\n\n![Model flowchart](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4184177%2Fdb9296173f2451ecbdddb8173866cffb%2Fmodellight.drawio.svg?generation=1773942873498721&alt=media)\n\n*Figure 2:* Model flowchart. The input is the pre-processed data array x, and the outputs are the predictions for the requested wavelengths and the associated sigma values.\n\n![Detailed view of the Time Reducing Residual Block N for N as 1 through 6](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4184177%2F6be839924011f903d4b8284dabf35127%2Fdiagram-Page-2_light.drawio.svg?generation=1772150272461207&alt=media)\n\n*Figure 3:* Detailed view of the Time Reducing Residual Block N for N as 1 through 6.\n\n**Augmentation**\n\nTo improve model generalization and robustness and decrease overfitting, we implemented several data augmentation techniques during training:\n\n- Time Flipping: Each light curve was flipped along the time axis to expose the model to both forward and reversed transit scenarios.\n\n- Variable Downsampling and Offsets: We used strides of 8, 9, or 10 and multiple random offsets, simulating the effect of time running at different speeds and producing light curves with varying sampling densities. This approach also increased the representation of incomplete transits, helping the model learn to handle edge cases and irregular data spans. Because the signals were median-filtered (kernel size 101), the incremental effect of offset variation is modest, but still introduces some diversity.\n\n- Additive Linear Trend: For each wavelength channel, we injected a random linear signal with a maximum slope set equal to the channel’s own range. The maximum was set per wavelength, but the randomization was done per planet observation.\n\n- Additive Sinusoidal Trend: For each wavelength channel we injected a random sinusoidal signal with low frequency.\n\n- Planet Parameter Noise: Small, zero-mean Gaussian noise (σ = 0.1) was added to each set of normalized planet parameters fed into the model. This reduces overfitting and allows the network to be more robust to small errors or uncertainty in planet/star parameters.\n\n**Ensembling methods**\n\nTo aggregate the model predictions for each planet, we explored several ensembling approaches seen in our [submission notebook](https://www.kaggle.com/code/ddimonte/submission-code). We found that the basic average provided the best results due to how we were predicting our values and sigmas. Each planet had predictions from (potentially) multiple observations, 5 different offsets of stride 10, and multiple model predictions. Methods we tried:\n\n- Basic Average: A simple average of all predictions for each wavelength channel for each planet, both for the value and the sigma.\n- Best Sigma Selection: Selecting only the single set of predictions (per planet) with the lowest mean predicted uncertainty across all wavelengths.\n- Best N Average: A simple average of the N rows of data per planet that had the lowest mean sigma values.\n- Weighted Averaging: Weighted each prediction by the inverse of its predicted uncertainty, computing weighted means and associated uncertainties for each wavelength.\n\n**Other notes:**\n\n*Stratification*\n\nWe explored stratification based on star and planet features [[notebook](https://www.kaggle.com/code/ddimonte/stratification)], hoping to create folds that better represent different subclasses or difficulty levels. However, the lack of strong natural clusters, weak correlations between features and prediction error, and overlapping distributions showed that meaningful strata did not exist in the data for the features we analyzed. As a result, simple random folds were used.\n\n*Other models*\n\n**Citations**\n\n- C. Truong, L. Oudre, N. Vayatis. Selective review of offline change point detection methods. Signal Processing, 167:107299, 2020.",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3294676": "**Data exploration**\n\n[This notebook](https://www.kaggle.com/code/ddimonte/understanding-the-data) provides detailed exploration of the data format and applies processing steps borrowed from the [binning and processing notebook](https://www.kaggle.com/code/gordonyip/update-calibrating-and-binning-astronomical-data). Key changes in this notebook include minimal time binning (just enough to standardize AIRS-CH0 and FGS1 lengths), updated masking behavior, the addition of inpainting for missing data, and extensive before/after visualizations.\n\nIn [this notebook](https://www.kaggle.com/code/ddimonte/process-data) we fit a polynomial baseline to the out-of-transit regions of each light curve, using smoothing and change point detection to identify the transit onset and offset. However, we found that some transits in our dataset are incomplete or contain somewhat complex background changes, making baseline normalization unreliable for all cases without more careful consideration. To address this, we briefly attempted to fit a BATMAN transit model to the data, but ultimately decided not to pursue this further due to time constraints and the need for robust validation.\n\nWe then investigated the ratio between the provided ground truth transit depth and the maximum observed transit depth from the light curve. We found that the ground truth value is not always at the minimum of the light curve or a fixed fraction of the fitted model, and can even be lower than the observed minimum for some planets. This suggests that other parameters besides for the visible transit depth in the light curve alone play a significant role in determining the true transit depth such as planet and star data or the shape of the transit curve. This is likely due to transit characteristics and limb darkening.\n\n![Light curve examples including group truth](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4184177%2Febf65ba60085550dba27cb4cde82f708%2FIllustration3.png?generation=1759267472752844&alt=media)\n\n*Figure 1:* Light curve examples including ground truth estimate.\n\n**Preprocessing** \n\n[This notebook](https://www.kaggle.com/code/ddimonte/submission-code) contains the final submission code with all the necessary preprocessing on the test data. Much of the code was taken from the [binning and processing notebook](https://www.kaggle.com/code/gordonyip/update-calibrating-and-binning-astronomical-data). We also use code from the ruptures library.\nPreprocessing steps:\n- Load and Calibrate Detector Signals\n- Load Calibration Files and Apply Detector Cleaning\n    - Mask hot/dead pixels: Replace unusable data using mask maps and dark files.\n    - Non-linearity correction\n    - Flat Field Correction: Correct for pixel-to-pixel sensitivity variations using loaded flat field maps.\n    - Dark subtraction: Subtract the appropriate dark current background.\n    - Correlated Double Sampling\n- Time binning 12 to FGS1 to align time dimensions.\n- Inpainting: Fill remaining missing or masked regions using biharmonic inpainting treating time as channels.\n- Spatial Summation: Sum along spatial axes to produce light curves and prepare input shapes appropriate for the model.\n- Median filtering kernel size 101\n- Downsampling stride 10\n\n**Model**\n\nThe overall model architecture is shown in Figure 2. The preprocessed data [wavelength x time] is normalized per wavelength and passed through a Time-Reducing Residual Stack. The resulting feature map is pooled and then concatenated with the wavelength means z-scores, the wavelength standard deviations z-scores, and the normalized planet features. These scores are calculated using summary statistics from the entire dataset in order to incorporate information about the sample's placement in the population distribution into the model after our earlier observations shown in Figure 1. The combined vector is passed through a fully connected layer, whose output feeds three separate fully connected output heads. The first two heads are used to compute the per-wavelength prediction, and the third head is used for the uncertainty (sigma) values. For the prediction, we calculate the relative change between the first two outputs. For the third output, we train the model to output log(sigma) so we take the exponential of the third output head for our final sigma value.\n\nThe Time-Reducing Residual Block details are shown in Figure 3. In the figure N represents the block number from 1 to 6 since 6 blocks are stacked in the final model. The wavelength dimension uses circular padding and increasing kernel dilation 3<sup>N-1</sup> so that all of the wavelengths' features are convolved with each other within 6 stacked residual blocks. The time dimension uses 0 padding and a kernel dilation of 1 so the time field of view remains within the nearest times from small increments to large increments as the time dimension is pooled in successive blocks. In our model we used a time dimension pooling kernel of 4 for the first 4 blocks and then eased off down to 2 to maintain the desired data size for a 6 block stack and to prioritize faster compression of the time dimension earlier.\n\nWe trained the model with a combined loss utilizing both MSE as well as the GLL error used for scoring.\n\n![Model flowchart](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4184177%2Fdb9296173f2451ecbdddb8173866cffb%2Fmodellight.drawio.svg?generation=1773942873498721&alt=media)\n\n*Figure 2:* Model flowchart. The input is the pre-processed data array x, and the outputs are the predictions for the requested wavelengths and the associated sigma values.\n\n![Detailed view of the Time Reducing Residual Block N for N as 1 through 6](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4184177%2F6be839924011f903d4b8284dabf35127%2Fdiagram-Page-2_light.drawio.svg?generation=1772150272461207&alt=media)\n\n*Figure 3:* Detailed view of the Time Reducing Residual Block N for N as 1 through 6.\n\n**Augmentation**\n\nTo improve model generalization and robustness and decrease overfitting, we implemented several data augmentation techniques during training:\n\n- Time Flipping: Each light curve was flipped along the time axis to expose the model to both forward and reversed transit scenarios.\n\n- Variable Downsampling and Offsets: We used strides of 8, 9, or 10 and multiple random offsets, simulating the effect of time running at different speeds and producing light curves with varying sampling densities. This approach also increased the representation of incomplete transits, helping the model learn to handle edge cases and irregular data spans. Because the signals were median-filtered (kernel size 101), the incremental effect of offset variation is modest, but still introduces some diversity.\n\n- Additive Linear Trend: For each wavelength channel, we injected a random linear signal with a maximum slope set equal to the channel’s own range. The maximum was set per wavelength, but the randomization was done per planet observation.\n\n- Additive Sinusoidal Trend: For each wavelength channel we injected a random sinusoidal signal with low frequency.\n\n- Planet Parameter Noise: Small, zero-mean Gaussian noise (σ = 0.1) was added to each set of normalized planet parameters fed into the model. This reduces overfitting and allows the network to be more robust to small errors or uncertainty in planet/star parameters.\n\n**Ensembling methods**\n\nTo aggregate the model predictions for each planet, we explored several ensembling approaches seen in our [submission notebook](https://www.kaggle.com/code/ddimonte/submission-code). We found that the basic average provided the best results due to how we were predicting our values and sigmas. Each planet had predictions from (potentially) multiple observations, 5 different offsets of stride 10, and multiple model predictions. Methods we tried:\n\n- Basic Average: A simple average of all predictions for each wavelength channel for each planet, both for the value and the sigma.\n- Best Sigma Selection: Selecting only the single set of predictions (per planet) with the lowest mean predicted uncertainty across all wavelengths.\n- Best N Average: A simple average of the N rows of data per planet that had the lowest mean sigma values.\n- Weighted Averaging: Weighted each prediction by the inverse of its predicted uncertainty, computing weighted means and associated uncertainties for each wavelength.\n\n**Other notes:**\n\n*Stratification*\n\nWe explored stratification based on star and planet features [[notebook](https://www.kaggle.com/code/ddimonte/stratification)], hoping to create folds that better represent different subclasses or difficulty levels. However, the lack of strong natural clusters, weak correlations between features and prediction error, and overlapping distributions showed that meaningful strata did not exist in the data for the features we analyzed. As a result, simple random folds were used.\n\n*Other models*\n\n**Citations**\n\n- C. Truong, L. Oudre, N. Vayatis. Selective review of offline change point detection methods. Signal Processing, 167:107299, 2020."
  }
}