{
  "id": 609372,
  "title": "16th Place Solution - Probablistic Regression",
  "url": "/competitions/ariel-data-challenge-2025/discussion/609372",
  "author_name": "Viji",
  "post_date": "2025-09-26T09:26:07.198000",
  "votes": 6,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Thanks to the organizers and my fellow Kagglers for a really interesting competition. My goal from the start was to improve on my Ariel 2024 neural network solution (which earned bronze). I didn’t go down the physics-based modeling route (where many others excelled) — instead I focused on refining a probabilistic neural network approach.</p>\n<p><strong>Model</strong></p>\n<p>The core model is a two-input Conv1D neural network built in TensorFlow/Keras. One branch processes the FGS1 signal, the other processes the AIRS signal:</p>\n<p>Conv1D (ReLU)<br>\nMaxPooling<br>\nFlatten<br>\nThe two branches are concatenated into a shared representation, combining complementary information. The key difference from a standard regression model is the probabilistic output layer: instead of predicting a single value per wavelength, the model outputs both a mean and an uncertainty across all 283 wavelengths.</p>\n<p><strong>Mixture Distribution</strong></p>\n<p>The output is defined as a Laplace-based mixture distribution with 15 components × 283 wavelengths. Parameters from the dense layer are split into:</p>\n<p>mixture logits (weights)<br>\nper-wavelength means (locs)<br>\nper-wavelength scales (uncertainties)</p>\n<p>Stability tricks:</p>\n<p>Scales clipped using variance-based bounds<br>\nSoftplus applied to enforce positivity</p>\n<p>TensorFlow Probability’s DistributionLambda then builds the full mixture model. This setup lets the network learn both expected values and uncertainties, which was central to the challenge.</p>\n<p><strong>Data Preprocessing</strong></p>\n<p>The preprocessing pipeline was fairly straightforward:</p>\n<p>Load raw data<br>\nWinsorize<br>\nSegment signals into transient shapes</p>\n<p>For each segment, the mean is compared with the “unobscured” parts (outside the segment). A surprise boost came when I added the inverse of the segmented transients (domes) alongside the original dips — the model benefited from explicitly seeing both contrasting patterns.</p>\n<p>Further gains came from combining segments of different lengths, so the final model used multi-segment data (coarse + fine).</p>\n<p><strong>Why This Helped</strong></p>\n<p>Segments give compact but informative context for transits.<br>\nAdding domes (inverse of dips) enriched the feature space and helped the network learn contrasting shapes.<br>\nMulti-segments allowed the model to generalize better across fine and coarse structures.</p>\n<p><strong>Training Loss</strong></p>\n<p>The loss was Negative Log Likelihood (NLL) with wavelength-based weighting:</p>\n<p>Weights were split between FGS1 and AIRS<br>\nf_share = 0.1702 gave FGS1 its relative importance<br>\nCombined into a single weight array so both channels influenced training consistently</p>\n<p>**Other Tweaks</p>\n<p>Adding a small VSN gate (Conv1D + sigmoid) gave a small but consistent lift.<br>\nEnsembling different mixtures (Laplace, Normal, Logistic) across folds and seeds helped stabilize scores.</p>\n<p><strong>Results &amp; Takeaways</strong></p>\n<p>The mixture distribution was the most impactful part, though it’s still not perfect — per-wavelength variation can sometimes collapse toward the mean.<br>\nWith 10-fold CV, validation–public gaps were usually around 0.018–0.02 (val higher) and validation–private gaps were &lt;0.01, so generalization was stable.<br>\nFull training time was about 1 hour per model, which made experimentation manageable.</p>\n<p>Overall, it was a satisfying to develop an approach which gave competitive scores, and most importantly, can be adapted to other datasets where uncertainty estimation is important. Big thanks again to the organizers and the community — for this great learning experience.</p>",
  "messages": [
    {
      "id": 3294498,
      "postDate": "2025-09-26T09:26:07.197Z",
      "content": "<p>Thanks to the organizers and my fellow Kagglers for a really interesting competition. My goal from the start was to improve on my Ariel 2024 neural network solution (which earned bronze). I didn’t go down the physics-based modeling route (where many others excelled) — instead I focused on refining a probabilistic neural network approach.</p>\n<p><strong>Model</strong></p>\n<p>The core model is a two-input Conv1D neural network built in TensorFlow/Keras. One branch processes the FGS1 signal, the other processes the AIRS signal:</p>\n<p>Conv1D (ReLU)<br>\nMaxPooling<br>\nFlatten<br>\nThe two branches are concatenated into a shared representation, combining complementary information. The key difference from a standard regression model is the probabilistic output layer: instead of predicting a single value per wavelength, the model outputs both a mean and an uncertainty across all 283 wavelengths.</p>\n<p><strong>Mixture Distribution</strong></p>\n<p>The output is defined as a Laplace-based mixture distribution with 15 components × 283 wavelengths. Parameters from the dense layer are split into:</p>\n<p>mixture logits (weights)<br>\nper-wavelength means (locs)<br>\nper-wavelength scales (uncertainties)</p>\n<p>Stability tricks:</p>\n<p>Scales clipped using variance-based bounds<br>\nSoftplus applied to enforce positivity</p>\n<p>TensorFlow Probability’s DistributionLambda then builds the full mixture model. This setup lets the network learn both expected values and uncertainties, which was central to the challenge.</p>\n<p><strong>Data Preprocessing</strong></p>\n<p>The preprocessing pipeline was fairly straightforward:</p>\n<p>Load raw data<br>\nWinsorize<br>\nSegment signals into transient shapes</p>\n<p>For each segment, the mean is compared with the “unobscured” parts (outside the segment). A surprise boost came when I added the inverse of the segmented transients (domes) alongside the original dips — the model benefited from explicitly seeing both contrasting patterns.</p>\n<p>Further gains came from combining segments of different lengths, so the final model used multi-segment data (coarse + fine).</p>\n<p><strong>Why This Helped</strong></p>\n<p>Segments give compact but informative context for transits.<br>\nAdding domes (inverse of dips) enriched the feature space and helped the network learn contrasting shapes.<br>\nMulti-segments allowed the model to generalize better across fine and coarse structures.</p>\n<p><strong>Training Loss</strong></p>\n<p>The loss was Negative Log Likelihood (NLL) with wavelength-based weighting:</p>\n<p>Weights were split between FGS1 and AIRS<br>\nf_share = 0.1702 gave FGS1 its relative importance<br>\nCombined into a single weight array so both channels influenced training consistently</p>\n<p>**Other Tweaks</p>\n<p>Adding a small VSN gate (Conv1D + sigmoid) gave a small but consistent lift.<br>\nEnsembling different mixtures (Laplace, Normal, Logistic) across folds and seeds helped stabilize scores.</p>\n<p><strong>Results &amp; Takeaways</strong></p>\n<p>The mixture distribution was the most impactful part, though it’s still not perfect — per-wavelength variation can sometimes collapse toward the mean.<br>\nWith 10-fold CV, validation–public gaps were usually around 0.018–0.02 (val higher) and validation–private gaps were &lt;0.01, so generalization was stable.<br>\nFull training time was about 1 hour per model, which made experimentation manageable.</p>\n<p>Overall, it was a satisfying to develop an approach which gave competitive scores, and most importantly, can be adapted to other datasets where uncertainty estimation is important. Big thanks again to the organizers and the community — for this great learning experience.</p>",
      "rawMarkdown": "Thanks to the organizers and my fellow Kagglers for a really interesting competition. My goal from the start was to improve on my Ariel 2024 neural network solution (which earned bronze). I didn’t go down the physics-based modeling route (where many others excelled) — instead I focused on refining a probabilistic neural network approach.\n\n**Model**\n\nThe core model is a two-input Conv1D neural network built in TensorFlow/Keras. One branch processes the FGS1 signal, the other processes the AIRS signal:\n\nConv1D (ReLU)\nMaxPooling\nFlatten\nThe two branches are concatenated into a shared representation, combining complementary information. The key difference from a standard regression model is the probabilistic output layer: instead of predicting a single value per wavelength, the model outputs both a mean and an uncertainty across all 283 wavelengths.\n\n**Mixture Distribution**\n\nThe output is defined as a Laplace-based mixture distribution with 15 components × 283 wavelengths. Parameters from the dense layer are split into:\n\nmixture logits (weights)\nper-wavelength means (locs)\nper-wavelength scales (uncertainties)\n\nStability tricks:\n\nScales clipped using variance-based bounds\nSoftplus applied to enforce positivity\n\nTensorFlow Probability’s DistributionLambda then builds the full mixture model. This setup lets the network learn both expected values and uncertainties, which was central to the challenge.\n\n**Data Preprocessing**\n\nThe preprocessing pipeline was fairly straightforward:\n\nLoad raw data\nWinsorize\nSegment signals into transient shapes\n\nFor each segment, the mean is compared with the “unobscured” parts (outside the segment). A surprise boost came when I added the inverse of the segmented transients (domes) alongside the original dips — the model benefited from explicitly seeing both contrasting patterns.\n\nFurther gains came from combining segments of different lengths, so the final model used multi-segment data (coarse + fine).\n\n**Why This Helped**\n\nSegments give compact but informative context for transits.\nAdding domes (inverse of dips) enriched the feature space and helped the network learn contrasting shapes.\nMulti-segments allowed the model to generalize better across fine and coarse structures.\n\n**Training Loss**\n\nThe loss was Negative Log Likelihood (NLL) with wavelength-based weighting:\n\nWeights were split between FGS1 and AIRS\nf_share = 0.1702 gave FGS1 its relative importance\nCombined into a single weight array so both channels influenced training consistently\n\n**Other Tweaks\n\nAdding a small VSN gate (Conv1D + sigmoid) gave a small but consistent lift.\nEnsembling different mixtures (Laplace, Normal, Logistic) across folds and seeds helped stabilize scores.\n\n**Results & Takeaways**\n\nThe mixture distribution was the most impactful part, though it’s still not perfect — per-wavelength variation can sometimes collapse toward the mean.\nWith 10-fold CV, validation–public gaps were usually around 0.018–0.02 (val higher) and validation–private gaps were <0.01, so generalization was stable.\nFull training time was about 1 hour per model, which made experimentation manageable.\n\nOverall, it was a satisfying to develop an approach which gave competitive scores, and most importantly, can be adapted to other datasets where uncertainty estimation is important. Big thanks again to the organizers and the community — for this great learning experience.\n",
      "votes": 6
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3294498": "Thanks to the organizers and my fellow Kagglers for a really interesting competition. My goal from the start was to improve on my Ariel 2024 neural network solution (which earned bronze). I didn’t go down the physics-based modeling route (where many others excelled) — instead I focused on refining a probabilistic neural network approach.\n\n**Model**\n\nThe core model is a two-input Conv1D neural network built in TensorFlow/Keras. One branch processes the FGS1 signal, the other processes the AIRS signal:\n\nConv1D (ReLU)\nMaxPooling\nFlatten\nThe two branches are concatenated into a shared representation, combining complementary information. The key difference from a standard regression model is the probabilistic output layer: instead of predicting a single value per wavelength, the model outputs both a mean and an uncertainty across all 283 wavelengths.\n\n**Mixture Distribution**\n\nThe output is defined as a Laplace-based mixture distribution with 15 components × 283 wavelengths. Parameters from the dense layer are split into:\n\nmixture logits (weights)\nper-wavelength means (locs)\nper-wavelength scales (uncertainties)\n\nStability tricks:\n\nScales clipped using variance-based bounds\nSoftplus applied to enforce positivity\n\nTensorFlow Probability’s DistributionLambda then builds the full mixture model. This setup lets the network learn both expected values and uncertainties, which was central to the challenge.\n\n**Data Preprocessing**\n\nThe preprocessing pipeline was fairly straightforward:\n\nLoad raw data\nWinsorize\nSegment signals into transient shapes\n\nFor each segment, the mean is compared with the “unobscured” parts (outside the segment). A surprise boost came when I added the inverse of the segmented transients (domes) alongside the original dips — the model benefited from explicitly seeing both contrasting patterns.\n\nFurther gains came from combining segments of different lengths, so the final model used multi-segment data (coarse + fine).\n\n**Why This Helped**\n\nSegments give compact but informative context for transits.\nAdding domes (inverse of dips) enriched the feature space and helped the network learn contrasting shapes.\nMulti-segments allowed the model to generalize better across fine and coarse structures.\n\n**Training Loss**\n\nThe loss was Negative Log Likelihood (NLL) with wavelength-based weighting:\n\nWeights were split between FGS1 and AIRS\nf_share = 0.1702 gave FGS1 its relative importance\nCombined into a single weight array so both channels influenced training consistently\n\n**Other Tweaks\n\nAdding a small VSN gate (Conv1D + sigmoid) gave a small but consistent lift.\nEnsembling different mixtures (Laplace, Normal, Logistic) across folds and seeds helped stabilize scores.\n\n**Results & Takeaways**\n\nThe mixture distribution was the most impactful part, though it’s still not perfect — per-wavelength variation can sometimes collapse toward the mean.\nWith 10-fold CV, validation–public gaps were usually around 0.018–0.02 (val higher) and validation–private gaps were <0.01, so generalization was stable.\nFull training time was about 1 hour per model, which made experimentation manageable.\n\nOverall, it was a satisfying to develop an approach which gave competitive scores, and most importantly, can be adapted to other datasets where uncertainty estimation is important. Big thanks again to the organizers and the community — for this great learning experience.\n"
  }
}