{
  "id": 609532,
  "title": "25th Place - Polynomnial Fitting / Nelder-Mead",
  "url": "/competitions/ariel-data-challenge-2025/discussion/609532",
  "author_name": "Adrian Wiśniewski",
  "post_date": "2025-09-27T14:58:42.391000",
  "votes": 2,
  "comment_count": 0,
  "views": 0,
  "content": "<h2>1. Data Preprocessing</h2>\n<p>Data calibration was done in accordance with the organizers' suggested solution. The sole modification was to utilize a subset of the sensor data's pixels. Consequently, the array slices [:, 10:22, 39:321], [:, 10:22, :] were utilized for the AIRS-CH0 and FGS1 sensors.</p>\n<h2>2. Detecting Ingress / Egress.</h2>\n<p>Ingress and egress were computed using an algorithm based on the gradient approach. Windows of <code>phase_a_start</code> till <code>phase_a_end</code> and <code>phase_b_start</code> till <code>phase_b_end</code> were excluded  prior to polynomnial fitting. In case of invalid detection of <code>phase_a</code> i used only egress part to estimate depth and vice versa. In case of invalid detection of <code>phase_a</code> and <code>phase_b</code> i did not make any predictions.</p>\n<pre><code> ():\n    \n\n    \n    flux = flux.mean(axis=-)\n\n    \n    min_index = np.argmin(flux)\n    in_a = flux[:min_index]\n    in_b = flux[min_index:]\n\n    \n    gradient_a = np.gradient(in_a, edge_order=)\n\n    gradient_2a = np.gradient(gradient_a, edge_order=)\n    gradient_2b = np.gradient(gradient_b, edge_order=)\n\n    \n    phase_a = gradient_a[:].argmin() + \n    phase_b = gradient_b[:-].argmax() + (gradient_a)\n\n    phase_a_start = gradient_2a.argmin()\n    phase_a_end = gradient_2a.argmax()\n\n    phase_b_start = gradient_2b.argmax() + (gradient_a)\n    phase_b_end = gradient_2b.argmin() + (gradient_a)\n\n    breakpoints = {: min_index,\n                   : phase_a_start - ,\n                   : phase_a,\n                   : phase_a_end + ,\n                   : phase_b_start - ,\n                   : phase_b,\n                   : phase_b_end + }\n\n     breakpoints\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4869926%2Fb077caac2a7fc0c03089516802b8f816%2Fgradient%20(2).png?generation=1758973537444006&amp;alt=media\" alt=\"\"></p>\n<h2>3. Transit depths estimation.</h2>\n<p>The transit depths were determined using Polynomnial Fitting and Nelder-Mead optimization. </p>\n<pre><code>def F(, phase_a_data, phase_x_data, phase_b_data, polyorder):\n     = np.concatenate((\n        phase_a_data,\n        phase_x_data * ,\n        phase_b_data\n    ))\n\n    X = np.arange(len())\n\n    z = np.polyfit(X, , polyorder)\n    p = np.poly1d(z)\n     = np.(p(X) - ).()\n     \n</code></pre>\n<h3>3.1  Dynamic Spectrum Estimation</h3>\n<p>The transit depth were determined using Polynomnial Fitting (up to 3rd polynomnial grade)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4869926%2Fea57ee1308faa20c5f22032365781f7c%2F116369715.solution.airs.png?generation=1758972658570624&amp;alt=media\" alt=\"\"></p>\n<h3>3.2 Flat Spectrum</h3>\n<p>I began the approach by forecasting solely the flat outputs because the majority of the spectra in the data have relatively minor oscillations.</p>\n<p>The transit depth were determined at first using Polynomnial Fitting (up to 18th polynomnial grade) just once acrossed averaged wavelenghts. <br>\nEmploying polynomial fitting up to degree 18 may appear counterintuitive; however, it yielded superior results compared to fitting up to degree 3 / 4.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4869926%2F5ac83fdd90ad07835ee0ea7c0e093ca9%2Fpol11.png?generation=1758975603846509&amp;alt=media\" alt=\"\"> <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4869926%2Fc9123af73bbb7c04f7a08e542788641b%2Fpoly33.png?generation=1758975614650218&amp;alt=media\" alt=\"\"> <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4869926%2Fb78123c4df9aeacffc8e9f53f9f114c4%2Fpol28.png?generation=1758975630123179&amp;alt=media\" alt=\"\"></p>\n<h3>3.3 Transit Depths Fluctuations</h3>\n<p>To accurately forecast the final spectrum, the amount of dynamics in depth measurements was determined. Measurements were computed on non-smoothed time-series.</p>\n<p>Zero Crossings Rate - (ZCR) is a feature used to measure the number of times a signal crosses the zero level (i.e., changes its sign).</p>\n<pre><code>   ():\n       librosa.zero_crossings(signal - signal.mean())[].sum()\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4869926%2Fc43da980c84b9c59fb5fc721c9631717%2Fzcr2_post.png?generation=1758973240612006&amp;alt=media\" alt=\"\"> <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4869926%2F0d8e0c874c5f29eeb208c986e9858e1b%2Fzcr1_post.png?generation=1758973002578826&amp;alt=media\" alt=\"\"></p>\n<h2>4. Final Solution - Combining - Flat Spectrum / Dynamic Spectrum and Hyperparameters - Zero-Crossings-Rate / Orbital Period / Transit Wide  / Transit Correctness.</h2>\n<h3>To put it briefly:</h3>\n<ul>\n<li><p>A flat spectrum with low sigma was employed as a solution to the problem if the computed transit depths had minimal fluctuations (High ZCR) because there was little chance of a prediction error.</p></li>\n<li><p>Since there was a significant chance of prediction error, Non-Flattened (dynamic) spectra with larger sigma were employed as a solution if the computed transit depths exhibited significant swings (Low ZCR).</p></li>\n<li><p>The predicted transit depth is modeled as a cubic polynomial function of the scaled signal, using coefficients a0, a1, a2, and a3.<br>\nsignal = a0 + a1 * signal + a2 * signal ** 2 + a3 * signal ** 3</p></li>\n</ul>\n<p>The Nelder-Mead algorithm and the following code were used to determine the values of sigma / signal coefficients and the amount of non-flattening depths to employ.</p>\n<pre><code> ():\n    \n\n    \n    a0, a1, a2, a3, cross_sigma, sigma, residual_coef, post_sigma = args[:]\n    feature_coeffs = args[:]  \n\n    \n    sigma = sigma + cross_sigma * zcr\n\n    \n     i, (feature_name, feature_value)  (additional_features.items()):\n        sigma += feature_coeffs[i] * feature_value\n        signal += feature_coeffs[i] * feature_value\n\n    \n     ScalerSigma.mode == :\n        \n        residual = dynamic_depths - np.mean(dynamic_depths, axis=-, keepdims=)\n        residual = savgol_filter(residual, window_length=, polyorder=)\n\n        \n        signal += zcr[:, ] * residual\n\n    \n    scaled_signal = a0 + a1 * signal + a2 * signal **  + a3 * signal ** \n\n     scaled_signal, sigma\n</code></pre>\n<h3>4.1 Example Predictions</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4869926%2Fdaac5a1ffd1be514fa6f281173615337%2F2805775794.solution.airs.png?generation=1758980816516371&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4869926%2Fef4157b485b3f748d8aa85c3fcfac9d5%2F1031303815.solution.airs.png?generation=1758980801042084&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 3295063,
      "postDate": "2025-09-27T14:58:42.390Z",
      "content": "<h2>1. Data Preprocessing</h2>\n<p>Data calibration was done in accordance with the organizers' suggested solution. The sole modification was to utilize a subset of the sensor data's pixels. Consequently, the array slices [:, 10:22, 39:321], [:, 10:22, :] were utilized for the AIRS-CH0 and FGS1 sensors.</p>\n<h2>2. Detecting Ingress / Egress.</h2>\n<p>Ingress and egress were computed using an algorithm based on the gradient approach. Windows of <code>phase_a_start</code> till <code>phase_a_end</code> and <code>phase_b_start</code> till <code>phase_b_end</code> were excluded  prior to polynomnial fitting. In case of invalid detection of <code>phase_a</code> i used only egress part to estimate depth and vice versa. In case of invalid detection of <code>phase_a</code> and <code>phase_b</code> i did not make any predictions.</p>\n<pre><code> ():\n    \n\n    \n    flux = flux.mean(axis=-)\n\n    \n    min_index = np.argmin(flux)\n    in_a = flux[:min_index]\n    in_b = flux[min_index:]\n\n    \n    gradient_a = np.gradient(in_a, edge_order=)\n\n    gradient_2a = np.gradient(gradient_a, edge_order=)\n    gradient_2b = np.gradient(gradient_b, edge_order=)\n\n    \n    phase_a = gradient_a[:].argmin() + \n    phase_b = gradient_b[:-].argmax() + (gradient_a)\n\n    phase_a_start = gradient_2a.argmin()\n    phase_a_end = gradient_2a.argmax()\n\n    phase_b_start = gradient_2b.argmax() + (gradient_a)\n    phase_b_end = gradient_2b.argmin() + (gradient_a)\n\n    breakpoints = {: min_index,\n                   : phase_a_start - ,\n                   : phase_a,\n                   : phase_a_end + ,\n                   : phase_b_start - ,\n                   : phase_b,\n                   : phase_b_end + }\n\n     breakpoints\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4869926%2Fb077caac2a7fc0c03089516802b8f816%2Fgradient%20(2).png?generation=1758973537444006&amp;alt=media\" alt=\"\"></p>\n<h2>3. Transit depths estimation.</h2>\n<p>The transit depths were determined using Polynomnial Fitting and Nelder-Mead optimization. </p>\n<pre><code>def F(, phase_a_data, phase_x_data, phase_b_data, polyorder):\n     = np.concatenate((\n        phase_a_data,\n        phase_x_data * ,\n        phase_b_data\n    ))\n\n    X = np.arange(len())\n\n    z = np.polyfit(X, , polyorder)\n    p = np.poly1d(z)\n     = np.(p(X) - ).()\n     \n</code></pre>\n<h3>3.1  Dynamic Spectrum Estimation</h3>\n<p>The transit depth were determined using Polynomnial Fitting (up to 3rd polynomnial grade)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4869926%2Fea57ee1308faa20c5f22032365781f7c%2F116369715.solution.airs.png?generation=1758972658570624&amp;alt=media\" alt=\"\"></p>\n<h3>3.2 Flat Spectrum</h3>\n<p>I began the approach by forecasting solely the flat outputs because the majority of the spectra in the data have relatively minor oscillations.</p>\n<p>The transit depth were determined at first using Polynomnial Fitting (up to 18th polynomnial grade) just once acrossed averaged wavelenghts. <br>\nEmploying polynomial fitting up to degree 18 may appear counterintuitive; however, it yielded superior results compared to fitting up to degree 3 / 4.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4869926%2F5ac83fdd90ad07835ee0ea7c0e093ca9%2Fpol11.png?generation=1758975603846509&amp;alt=media\" alt=\"\"> <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4869926%2Fc9123af73bbb7c04f7a08e542788641b%2Fpoly33.png?generation=1758975614650218&amp;alt=media\" alt=\"\"> <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4869926%2Fb78123c4df9aeacffc8e9f53f9f114c4%2Fpol28.png?generation=1758975630123179&amp;alt=media\" alt=\"\"></p>\n<h3>3.3 Transit Depths Fluctuations</h3>\n<p>To accurately forecast the final spectrum, the amount of dynamics in depth measurements was determined. Measurements were computed on non-smoothed time-series.</p>\n<p>Zero Crossings Rate - (ZCR) is a feature used to measure the number of times a signal crosses the zero level (i.e., changes its sign).</p>\n<pre><code>   ():\n       librosa.zero_crossings(signal - signal.mean())[].sum()\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4869926%2Fc43da980c84b9c59fb5fc721c9631717%2Fzcr2_post.png?generation=1758973240612006&amp;alt=media\" alt=\"\"> <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4869926%2F0d8e0c874c5f29eeb208c986e9858e1b%2Fzcr1_post.png?generation=1758973002578826&amp;alt=media\" alt=\"\"></p>\n<h2>4. Final Solution - Combining - Flat Spectrum / Dynamic Spectrum and Hyperparameters - Zero-Crossings-Rate / Orbital Period / Transit Wide  / Transit Correctness.</h2>\n<h3>To put it briefly:</h3>\n<ul>\n<li><p>A flat spectrum with low sigma was employed as a solution to the problem if the computed transit depths had minimal fluctuations (High ZCR) because there was little chance of a prediction error.</p></li>\n<li><p>Since there was a significant chance of prediction error, Non-Flattened (dynamic) spectra with larger sigma were employed as a solution if the computed transit depths exhibited significant swings (Low ZCR).</p></li>\n<li><p>The predicted transit depth is modeled as a cubic polynomial function of the scaled signal, using coefficients a0, a1, a2, and a3.<br>\nsignal = a0 + a1 * signal + a2 * signal ** 2 + a3 * signal ** 3</p></li>\n</ul>\n<p>The Nelder-Mead algorithm and the following code were used to determine the values of sigma / signal coefficients and the amount of non-flattening depths to employ.</p>\n<pre><code> ():\n    \n\n    \n    a0, a1, a2, a3, cross_sigma, sigma, residual_coef, post_sigma = args[:]\n    feature_coeffs = args[:]  \n\n    \n    sigma = sigma + cross_sigma * zcr\n\n    \n     i, (feature_name, feature_value)  (additional_features.items()):\n        sigma += feature_coeffs[i] * feature_value\n        signal += feature_coeffs[i] * feature_value\n\n    \n     ScalerSigma.mode == :\n        \n        residual = dynamic_depths - np.mean(dynamic_depths, axis=-, keepdims=)\n        residual = savgol_filter(residual, window_length=, polyorder=)\n\n        \n        signal += zcr[:, ] * residual\n\n    \n    scaled_signal = a0 + a1 * signal + a2 * signal **  + a3 * signal ** \n\n     scaled_signal, sigma\n</code></pre>\n<h3>4.1 Example Predictions</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4869926%2Fdaac5a1ffd1be514fa6f281173615337%2F2805775794.solution.airs.png?generation=1758980816516371&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4869926%2Fef4157b485b3f748d8aa85c3fcfac9d5%2F1031303815.solution.airs.png?generation=1758980801042084&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "## 1. Data Preprocessing\n\n\nData calibration was done in accordance with the organizers' suggested solution. The sole modification was to utilize a subset of the sensor data's pixels. Consequently, the array slices [:, 10:22, 39:321], [:, 10:22, :] were utilized for the AIRS-CH0 and FGS1 sensors.\n\n\n## 2. Detecting Ingress / Egress.\n\n\nIngress and egress were computed using an algorithm based on the gradient approach. Windows of `phase_a_start` till `phase_a_end` and `phase_b_start` till `phase_b_end` were excluded  prior to polynomnial fitting. In case of invalid detection of `phase_a` i used only egress part to estimate depth and vice versa. In case of invalid detection of `phase_a` and `phase_b` i did not make any predictions.\n\n\n    def detect_breakpoints(flux, verbose=False):\n        \"\"\"\n        Detect the phases in the airs flux signal by calculating the gradient of the detrended signal.\n        \"\"\"\n    \n        # Mean and filter signal\n        flux = flux.mean(axis=-1)\n\n        # Find middle of transit\n        min_index = np.argmin(flux)\n        in_a = flux[:min_index]\n        in_b = flux[min_index:]\n\n        # Compute the gradients of both halves\n        gradient_a = np.gradient(in_a, edge_order=1)\n\n        gradient_2a = np.gradient(gradient_a, edge_order=1)\n        gradient_2b = np.gradient(gradient_b, edge_order=1)\n\n        # Identify the phase indices based on the gradients\n        phase_a = gradient_a[1:].argmin() + 1\n        phase_b = gradient_b[:-1].argmax() + len(gradient_a)\n\n        phase_a_start = gradient_2a.argmin()\n        phase_a_end = gradient_2a.argmax()\n    \n        phase_b_start = gradient_2b.argmax() + len(gradient_a)\n        phase_b_end = gradient_2b.argmin() + len(gradient_a)\n    \n        breakpoints = {'center': min_index,\n                       'phase_a_start': phase_a_start - 1,\n                       'phase_a': phase_a,\n                       'phase_a_end': phase_a_end + 1,\n                       'phase_b_start': phase_b_start - 1,\n                       'phase_b': phase_b,\n                       'phase_b_end': phase_b_end + 1}\n\n        return breakpoints\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4869926%2Fb077caac2a7fc0c03089516802b8f816%2Fgradient%20(2).png?generation=1758973537444006&alt=media)\n\n## 3. Transit depths estimation.\nThe transit depths were determined using Polynomnial Fitting and Nelder-Mead optimization. \n\n\n    def F(depth, phase_a_data, phase_x_data, phase_b_data, polyorder):\n        y = np.concatenate((\n            phase_a_data,\n            phase_x_data * depth,\n            phase_b_data\n        ))\n\n        X = np.arange(len(y))\n\n        z = np.polyfit(X, y, polyorder)\n        p = np.poly1d(z)\n        score = np.abs(p(X) - y).mean()\n        return score\n\n### 3.1  Dynamic Spectrum Estimation\nThe transit depth were determined using Polynomnial Fitting (up to 3rd polynomnial grade)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4869926%2Fea57ee1308faa20c5f22032365781f7c%2F116369715.solution.airs.png?generation=1758972658570624&alt=media)\n\n### 3.2 Flat Spectrum\n\nI began the approach by forecasting solely the flat outputs because the majority of the spectra in the data have relatively minor oscillations.\n\nThe transit depth were determined at first using Polynomnial Fitting (up to 18th polynomnial grade) just once acrossed averaged wavelenghts. \nEmploying polynomial fitting up to degree 18 may appear counterintuitive; however, it yielded superior results compared to fitting up to degree 3 / 4.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4869926%2F5ac83fdd90ad07835ee0ea7c0e093ca9%2Fpol11.png?generation=1758975603846509&alt=media) ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4869926%2Fc9123af73bbb7c04f7a08e542788641b%2Fpoly33.png?generation=1758975614650218&alt=media) ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4869926%2Fb78123c4df9aeacffc8e9f53f9f114c4%2Fpol28.png?generation=1758975630123179&alt=media)\n\n\n### 3.3 Transit Depths Fluctuations\n\nTo accurately forecast the final spectrum, the amount of dynamics in depth measurements was determined. Measurements were computed on non-smoothed time-series.\n\nZero Crossings Rate - (ZCR) is a feature used to measure the number of times a signal crosses the zero level (i.e., changes its sign).\n\n      def compute_zcr(signal, burn_in=20):\n          return librosa.zero_crossings(signal - signal.mean())[:-burn_in].sum()\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4869926%2Fc43da980c84b9c59fb5fc721c9631717%2Fzcr2_post.png?generation=1758973240612006&alt=media) ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4869926%2F0d8e0c874c5f29eeb208c986e9858e1b%2Fzcr1_post.png?generation=1758973002578826&alt=media)\n\n## 4. Final Solution - Combining - Flat Spectrum / Dynamic Spectrum and Hyperparameters - Zero-Crossings-Rate / Orbital Period / Transit Wide  / Transit Correctness.\n\n### To put it briefly:\n\n* A flat spectrum with low sigma was employed as a solution to the problem if the computed transit depths had minimal fluctuations (High ZCR) because there was little chance of a prediction error.\n\n* Since there was a significant chance of prediction error, Non-Flattened (dynamic) spectra with larger sigma were employed as a solution if the computed transit depths exhibited significant swings (Low ZCR).\n\n* The predicted transit depth is modeled as a cubic polynomial function of the scaled signal, using coefficients a0, a1, a2, and a3.\n   signal = a0 + a1 * signal + a2 * signal ** 2 + a3 * signal ** 3\n\nThe Nelder-Mead algorithm and the following code were used to determine the values of sigma / signal coefficients and the amount of non-flattening depths to employ.\n\n    def operator(signal, zcr, additional_features, dynamic_depths, args):\n        \"\"\"\n        Simplified version:\n        - Combine base parameters with ZCR and features to modify sigma.\n        - Adjust the signal based on features.\n        - Add residual depths to flatten based on zcr if in 'airs' mode.\n        - Finally, scale the signal with polynomial coefficients.\n        \"\"\"\n\n        # Unpack parameters\n        a0, a1, a2, a3, cross_sigma, sigma, residual_coef, post_sigma = args[:8]\n        feature_coeffs = args[8:]  # Coefficients for additional features\n\n        # Initialize sigma with base and ZCR influence\n        sigma = sigma + cross_sigma * zcr\n\n        # Incorporate effects of additional features into sigma and signal\n        for i, (feature_name, feature_value) in enumerate(additional_features.items()):\n            sigma += feature_coeffs[i] * feature_value\n            signal += feature_coeffs[i] * feature_value\n\n        # Additional correction if mode is 'airs'\n        if ScalerSigma.mode == 'airs':\n            # Compute residual (difference from mean) and flatten it\n            residual = dynamic_depths - np.mean(dynamic_depths, axis=-1, keepdims=True)\n            residual = savgol_filter(residual, window_length=10, polyorder=1)\n\n            # Add residual to the signal scaled by zcr\n            signal += zcr[:, None] * residual\n\n        # Final scaling of the signal\n        scaled_signal = a0 + a1 * signal + a2 * signal ** 2 + a3 * signal ** 3\n\n        return scaled_signal, sigma\n\n### 4.1 Example Predictions\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4869926%2Fdaac5a1ffd1be514fa6f281173615337%2F2805775794.solution.airs.png?generation=1758980816516371&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4869926%2Fef4157b485b3f748d8aa85c3fcfac9d5%2F1031303815.solution.airs.png?generation=1758980801042084&alt=media)",
      "votes": 2
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3295063": "## 1. Data Preprocessing\n\n\nData calibration was done in accordance with the organizers' suggested solution. The sole modification was to utilize a subset of the sensor data's pixels. Consequently, the array slices [:, 10:22, 39:321], [:, 10:22, :] were utilized for the AIRS-CH0 and FGS1 sensors.\n\n\n## 2. Detecting Ingress / Egress.\n\n\nIngress and egress were computed using an algorithm based on the gradient approach. Windows of `phase_a_start` till `phase_a_end` and `phase_b_start` till `phase_b_end` were excluded  prior to polynomnial fitting. In case of invalid detection of `phase_a` i used only egress part to estimate depth and vice versa. In case of invalid detection of `phase_a` and `phase_b` i did not make any predictions.\n\n\n    def detect_breakpoints(flux, verbose=False):\n        \"\"\"\n        Detect the phases in the airs flux signal by calculating the gradient of the detrended signal.\n        \"\"\"\n    \n        # Mean and filter signal\n        flux = flux.mean(axis=-1)\n\n        # Find middle of transit\n        min_index = np.argmin(flux)\n        in_a = flux[:min_index]\n        in_b = flux[min_index:]\n\n        # Compute the gradients of both halves\n        gradient_a = np.gradient(in_a, edge_order=1)\n\n        gradient_2a = np.gradient(gradient_a, edge_order=1)\n        gradient_2b = np.gradient(gradient_b, edge_order=1)\n\n        # Identify the phase indices based on the gradients\n        phase_a = gradient_a[1:].argmin() + 1\n        phase_b = gradient_b[:-1].argmax() + len(gradient_a)\n\n        phase_a_start = gradient_2a.argmin()\n        phase_a_end = gradient_2a.argmax()\n    \n        phase_b_start = gradient_2b.argmax() + len(gradient_a)\n        phase_b_end = gradient_2b.argmin() + len(gradient_a)\n    \n        breakpoints = {'center': min_index,\n                       'phase_a_start': phase_a_start - 1,\n                       'phase_a': phase_a,\n                       'phase_a_end': phase_a_end + 1,\n                       'phase_b_start': phase_b_start - 1,\n                       'phase_b': phase_b,\n                       'phase_b_end': phase_b_end + 1}\n\n        return breakpoints\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4869926%2Fb077caac2a7fc0c03089516802b8f816%2Fgradient%20(2).png?generation=1758973537444006&alt=media)\n\n## 3. Transit depths estimation.\nThe transit depths were determined using Polynomnial Fitting and Nelder-Mead optimization. \n\n\n    def F(depth, phase_a_data, phase_x_data, phase_b_data, polyorder):\n        y = np.concatenate((\n            phase_a_data,\n            phase_x_data * depth,\n            phase_b_data\n        ))\n\n        X = np.arange(len(y))\n\n        z = np.polyfit(X, y, polyorder)\n        p = np.poly1d(z)\n        score = np.abs(p(X) - y).mean()\n        return score\n\n### 3.1  Dynamic Spectrum Estimation\nThe transit depth were determined using Polynomnial Fitting (up to 3rd polynomnial grade)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4869926%2Fea57ee1308faa20c5f22032365781f7c%2F116369715.solution.airs.png?generation=1758972658570624&alt=media)\n\n### 3.2 Flat Spectrum\n\nI began the approach by forecasting solely the flat outputs because the majority of the spectra in the data have relatively minor oscillations.\n\nThe transit depth were determined at first using Polynomnial Fitting (up to 18th polynomnial grade) just once acrossed averaged wavelenghts. \nEmploying polynomial fitting up to degree 18 may appear counterintuitive; however, it yielded superior results compared to fitting up to degree 3 / 4.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4869926%2F5ac83fdd90ad07835ee0ea7c0e093ca9%2Fpol11.png?generation=1758975603846509&alt=media) ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4869926%2Fc9123af73bbb7c04f7a08e542788641b%2Fpoly33.png?generation=1758975614650218&alt=media) ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4869926%2Fb78123c4df9aeacffc8e9f53f9f114c4%2Fpol28.png?generation=1758975630123179&alt=media)\n\n\n### 3.3 Transit Depths Fluctuations\n\nTo accurately forecast the final spectrum, the amount of dynamics in depth measurements was determined. Measurements were computed on non-smoothed time-series.\n\nZero Crossings Rate - (ZCR) is a feature used to measure the number of times a signal crosses the zero level (i.e., changes its sign).\n\n      def compute_zcr(signal, burn_in=20):\n          return librosa.zero_crossings(signal - signal.mean())[:-burn_in].sum()\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4869926%2Fc43da980c84b9c59fb5fc721c9631717%2Fzcr2_post.png?generation=1758973240612006&alt=media) ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4869926%2F0d8e0c874c5f29eeb208c986e9858e1b%2Fzcr1_post.png?generation=1758973002578826&alt=media)\n\n## 4. Final Solution - Combining - Flat Spectrum / Dynamic Spectrum and Hyperparameters - Zero-Crossings-Rate / Orbital Period / Transit Wide  / Transit Correctness.\n\n### To put it briefly:\n\n* A flat spectrum with low sigma was employed as a solution to the problem if the computed transit depths had minimal fluctuations (High ZCR) because there was little chance of a prediction error.\n\n* Since there was a significant chance of prediction error, Non-Flattened (dynamic) spectra with larger sigma were employed as a solution if the computed transit depths exhibited significant swings (Low ZCR).\n\n* The predicted transit depth is modeled as a cubic polynomial function of the scaled signal, using coefficients a0, a1, a2, and a3.\n   signal = a0 + a1 * signal + a2 * signal ** 2 + a3 * signal ** 3\n\nThe Nelder-Mead algorithm and the following code were used to determine the values of sigma / signal coefficients and the amount of non-flattening depths to employ.\n\n    def operator(signal, zcr, additional_features, dynamic_depths, args):\n        \"\"\"\n        Simplified version:\n        - Combine base parameters with ZCR and features to modify sigma.\n        - Adjust the signal based on features.\n        - Add residual depths to flatten based on zcr if in 'airs' mode.\n        - Finally, scale the signal with polynomial coefficients.\n        \"\"\"\n\n        # Unpack parameters\n        a0, a1, a2, a3, cross_sigma, sigma, residual_coef, post_sigma = args[:8]\n        feature_coeffs = args[8:]  # Coefficients for additional features\n\n        # Initialize sigma with base and ZCR influence\n        sigma = sigma + cross_sigma * zcr\n\n        # Incorporate effects of additional features into sigma and signal\n        for i, (feature_name, feature_value) in enumerate(additional_features.items()):\n            sigma += feature_coeffs[i] * feature_value\n            signal += feature_coeffs[i] * feature_value\n\n        # Additional correction if mode is 'airs'\n        if ScalerSigma.mode == 'airs':\n            # Compute residual (difference from mean) and flatten it\n            residual = dynamic_depths - np.mean(dynamic_depths, axis=-1, keepdims=True)\n            residual = savgol_filter(residual, window_length=10, polyorder=1)\n\n            # Add residual to the signal scaled by zcr\n            signal += zcr[:, None] * residual\n\n        # Final scaling of the signal\n        scaled_signal = a0 + a1 * signal + a2 * signal ** 2 + a3 * signal ** 3\n\n        return scaled_signal, sigma\n\n### 4.1 Example Predictions\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4869926%2Fdaac5a1ffd1be514fa6f281173615337%2F2805775794.solution.airs.png?generation=1758980816516371&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4869926%2Fef4157b485b3f748d8aa85c3fcfac9d5%2F1031303815.solution.airs.png?generation=1758980801042084&alt=media)"
  }
}