{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":101849,"databundleVersionId":13093295,"sourceType":"competition"}],"dockerImageVersionId":31089,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# **Signal Processing and Spectral Visualization**","metadata":{}},{"cell_type":"markdown","source":"This notebook demonstrates loading and preprocessing of Ariel exoplanet spectral data, including calibration, signal normalization, and light curve extraction. It also visualizes raw and processed frames as well as planetary spectra to support further analysis.","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nfrom pathlib import Path\nimport pyarrow.parquet as pq\n\n# Set the data directory\nDATA_DIR = Path(\"/kaggle/input/ariel-data-challenge-2025\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-25T02:18:17.800336Z","iopub.execute_input":"2025-07-25T02:18:17.800664Z","iopub.status.idle":"2025-07-25T02:18:17.805825Z","shell.execute_reply.started":"2025-07-25T02:18:17.800641Z","shell.execute_reply":"2025-07-25T02:18:17.804866Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"This step loads all necessary metadata from CSV files, including training labels, wavelength values, star properties, and ADC (Analog-to-Digital Converter) parameters. It then extracts the gain and offset values specifically for the FGS1 instrument, which are later used to decode the raw signal data.","metadata":{}},{"cell_type":"code","source":"# 1. Load metadata\ntrain_df = pd.read_csv(DATA_DIR / \"train.csv\")\nwavelengths = pd.read_csv(DATA_DIR / \"wavelengths.csv\").values.flatten()\nstar_info = pd.read_csv(DATA_DIR / \"train_star_info.csv\")\nadc_info = pd.read_csv(DATA_DIR / \"adc_info.csv\")\n\ndisplay(adc_info)\n\n# Retrieve ADC parameters\ngain = adc_info['FGS1_adc_gain'].values[0]\noffset = adc_info['FGS1_adc_offset'].values[0]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-25T02:18:17.807246Z","iopub.execute_input":"2025-07-25T02:18:17.807604Z","iopub.status.idle":"2025-07-25T02:18:17.98056Z","shell.execute_reply.started":"2025-07-25T02:18:17.807572Z","shell.execute_reply":"2025-07-25T02:18:17.979579Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"This function loads the signal data for a given planet and instrument from a Parquet file. It applies an ADC (analog-to-digital converter) transformation to restore the dynamic range of the raw signal values, converts the data to 64-bit floating point, and reshapes it into an image-like format. The shape depends on the instrument type: AIRS-CH0 results in 32×356 frames, while FGS1 gives 32×32 frames. An example call loads the AIRS-CH0 data for the first planet in the training set.","metadata":{}},{"cell_type":"code","source":"# 2. Load and preprocess signal data\ndef load_signal_data(planet_id, instrument=\"AIRS-CH0\", obs_count=0):\n    \"\"\"Load and preprocess signal data\"\"\"\n    file_path = DATA_DIR / f\"train/{planet_id}/{instrument}_signal_{obs_count}.parquet\"\n    data = pq.read_table(file_path).to_pandas().values\n\n    # Reconstruct dynamic range via ADC conversion\n    data = (data / gain) + offset\n    data = data.astype(np.float64)\n\n    # Reshape into image format (frames, height, width)\n    if instrument == \"AIRS-CH0\":\n        data = data.reshape(-1, 32, 356)\n    else:  # FGS1\n        data = data.reshape(-1, 32, 32)\n\n    return data\n\n# Example: Load AIRS-CH0 data for the first planet\nplanet_id = train_df['planet_id'].values[0]\nairs_data = load_signal_data(planet_id, \"AIRS-CH0\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-25T02:18:17.982038Z","iopub.execute_input":"2025-07-25T02:18:17.982371Z","iopub.status.idle":"2025-07-25T02:18:20.132175Z","shell.execute_reply.started":"2025-07-25T02:18:17.982342Z","shell.execute_reply":"2025-07-25T02:18:20.130898Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"This part loads calibration data (such as dark frames) for a given planet and instrument from a Parquet file. The calib_type parameter specifies the type of calibration (e.g., \"dark\"). The data is returned as a NumPy array. In the example, dark frames for the AIRS-CH0 instrument of the first planet are loaded.","metadata":{}},{"cell_type":"code","source":"# 3. Load calibration data\ndef load_calibration(planet_id, instrument=\"AIRS-CH0\", calib_type=\"dark\"):\n    \"\"\"Load calibration data\"\"\"\n    file_path = DATA_DIR / f\"train/{planet_id}/{instrument}_calibration_0/{calib_type}.parquet\"\n    return pq.read_table(file_path).to_pandas().values\n\n# Load dark frames\ndark_frames = load_calibration(planet_id, \"AIRS-CH0\", \"dark\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-25T02:18:20.13309Z","iopub.execute_input":"2025-07-25T02:18:20.133353Z","iopub.status.idle":"2025-07-25T02:18:20.169463Z","shell.execute_reply.started":"2025-07-25T02:18:20.133332Z","shell.execute_reply":"2025-07-25T02:18:20.168398Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"This function displays a few sample frames from the signal data as images. It plots the first n_samples frames using a color map to show pixel intensities. In the example, the first 3 frames of AIRS-CH0 data are visualized to inspect raw input quality and structure.","metadata":{}},{"cell_type":"code","source":"# 4. Visualize data\ndef plot_sample_frames(data, title, n_samples=3):\n    \"\"\"Visualize sample frames\"\"\"\n    plt.figure(figsize=(15, 5))\n    for i in range(n_samples):\n        plt.subplot(1, n_samples, i+1)\n        plt.imshow(data[i], cmap='viridis')\n        plt.title(f\"{title} - Frame {i}\")\n        plt.colorbar()\n    plt.tight_layout()\n    plt.show()\n\n# Visualize the first 3 frames of AIRS-CH0 data\nplot_sample_frames(airs_data, \"AIRS-CH0 Raw Data\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-25T02:18:20.171706Z","iopub.execute_input":"2025-07-25T02:18:20.171998Z","iopub.status.idle":"2025-07-25T02:18:20.899715Z","shell.execute_reply.started":"2025-07-25T02:18:20.171976Z","shell.execute_reply":"2025-07-25T02:18:20.898743Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"This preprocessing step improves signal quality by first subtracting the average dark frame to remove sensor bias. Then, it normalizes each pixel’s intensity over time to stabilize brightness variations. The process is applied to the first 1000 frames for memory efficiency, and the results are visualized.","metadata":{}},{"cell_type":"code","source":"# 5. Basic preprocessing pipeline\ndef preprocess_data(raw_data, dark_frames):\n    \"\"\"Basic preprocessing: dark subtraction and normalization\"\"\"\n    # Subtract average dark frame\n    avg_dark = np.mean(dark_frames, axis=0)\n    processed = raw_data - avg_dark\n\n    # Normalize along the time axis for each pixel\n    processed = (processed - np.mean(processed, axis=0)) / np.std(processed, axis=0)\n\n    return processed\n\n# Apply preprocessing (first 1000 frames only to save memory)\nprocessed_data = preprocess_data(airs_data[:1000], dark_frames)\nplot_sample_frames(processed_data, \"Processed Data\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-25T02:18:20.900701Z","iopub.execute_input":"2025-07-25T02:18:20.90099Z","iopub.status.idle":"2025-07-25T02:18:21.836448Z","shell.execute_reply.started":"2025-07-25T02:18:20.900971Z","shell.execute_reply":"2025-07-25T02:18:21.835388Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"This step extracts a light curve by averaging pixel values over a small region (by default, centered in the image) across all frames. It handles missing values by ignoring NaNs during averaging and replaces the result with zeros if all values are NaN. The light curve shows how brightness changes over time, which is important for detecting transit signals. A summary of the processed data and the light curve plot is also displayed.","metadata":{}},{"cell_type":"code","source":"# 6. Extract time series signal (light curve from center pixel region)\ndef extract_light_curve(data, x=None, y=None, radius=5):\n    \"\"\"Extract brightness variations from a specified region with NaN handling\"\"\"\n    if x is None or y is None:\n        # Automatically compute the center coordinates\n        height, width = data.shape[1], data.shape[2]\n        x, y = width // 2, height // 2\n\n    y_slice = slice(max(0, y - radius), min(data.shape[1], y + radius + 1))\n    x_slice = slice(max(0, x - radius), min(data.shape[2], x + radius + 1))\n\n    # Compute the mean while ignoring NaN values\n    light_curve = np.nanmean(data[:, y_slice, x_slice], axis=(1, 2))\n\n    # If all values are NaN, fill with zeros\n    if np.all(np.isnan(light_curve)):\n        light_curve = np.zeros_like(light_curve)\n\n    return light_curve\n\n# Check statistics of the processed data\nprint(\"Processed data stats:\")\nprint(f\"Min: {np.nanmin(processed_data)}, Max: {np.nanmax(processed_data)}\")\nprint(f\"NaN ratio: {np.mean(np.isnan(processed_data)):.2%}\")\n\n# Re-extract light curve\nlc = extract_light_curve(processed_data)\nprint(\"Light curve samples:\", lc[:10])  # Show the first 10 points\n\n# Visualization\nplt.figure(figsize=(12, 4))\nplt.plot(lc)\nplt.title(\"Light Curve (with NaN handling)\")\nplt.xlabel(\"Frame Number\")\nplt.ylabel(\"Normalized Flux\")\nplt.grid()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-25T02:18:21.837518Z","iopub.execute_input":"2025-07-25T02:18:21.837838Z","iopub.status.idle":"2025-07-25T02:18:22.184526Z","shell.execute_reply.started":"2025-07-25T02:18:21.837811Z","shell.execute_reply":"2025-07-25T02:18:22.183079Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"This function plots the spectrum of a specific planet, showing how its observed flux varies with wavelength. It retrieves the spectral data from the training DataFrame and uses the corresponding wavelength values for the x-axis. This visualization helps analyze the planet's atmospheric composition or thermal properties.","metadata":{}},{"cell_type":"code","source":"# 7. Visualize spectral data\ndef plot_spectrum(planet_id, train_df):\n    \"\"\"Visualize the spectrum of a given planet\"\"\"\n    spectrum = train_df[train_df['planet_id'] == planet_id].iloc[:, 1:].values.flatten()\n\n    plt.figure(figsize=(12, 4))\n    plt.plot(wavelengths, spectrum, 'b-')\n    plt.title(f\"Spectrum of Planet {planet_id}\")\n    plt.xlabel(\"Wavelength (µm)\")\n    plt.ylabel(\"Flux\")\n    plt.grid()\n    plt.show()\n\nplot_spectrum(planet_id, train_df)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-25T02:18:22.185749Z","iopub.execute_input":"2025-07-25T02:18:22.1861Z","iopub.status.idle":"2025-07-25T02:18:22.397216Z","shell.execute_reply.started":"2025-07-25T02:18:22.186073Z","shell.execute_reply":"2025-07-25T02:18:22.396306Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Summary\n\nIn this notebook, we successfully loaded and preprocessed the Ariel exoplanet data, including applying ADC corrections and dark frame subtraction. We extracted light curves from the calibrated signal and visualized both raw and processed data frames. Additionally, planetary spectra were plotted to facilitate atmospheric analysis. These steps lay a solid foundation for further modeling and interpretation of the observations.","metadata":{}},{"cell_type":"markdown","source":"## Next Steps\n\n* Implement advanced noise reduction and outlier detection techniques to improve light curve quality.\n* Explore feature extraction methods for better characterization of planetary atmospheres.\n* Develop machine learning models to predict planetary properties from spectral and time-series data.\n* Incorporate additional calibration data and multi-instrument fusion to enhance robustness.","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}