{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.11","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":101849,"databundleVersionId":12846694,"sourceType":"competition"}],"dockerImageVersionId":31040,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"### In this Notebook, We :\n* Understand the Domain of Competition\n* Diffrential Statistics\n* Infrential Statistics\n* EDA\n* Understand Important Points/Analysis required for Training\n\n### Viualization of Data/Guide is already done here : [here](https://www.kaggle.com/code/samasiayushman/complete-visualization-guide)","metadata":{}},{"cell_type":"markdown","source":"![pic](https://i.ibb.co/xKbBFjjS/Copy-of-FB-3.png)","metadata":{}},{"cell_type":"markdown","source":"# A. Goal Of Competition","metadata":{}},{"cell_type":"markdown","source":"The competition tackles a critical bottleneck in exoplanet research: moving from raw data to interpretable atmospheric models. Successful solutions will directly influence how we:\n\n### 🔭 Process data from ARIEL/JWST\n### 🌌 Prioritize targets for follow-up observations\n### 🧪 Design future instruments with better noise mitigation","metadata":{}},{"cell_type":"markdown","source":"## Develop machine learning/statistical methods to:\n\n* Extract exoplanet atmospheric signals from noisy transit spectroscopy data\n\n* Separate true planetary spectra from:\n\n* Instrumental noise (telescope/system artifacts)\n\n* Stellar activity (starspots, flares)\n\n* Limb-darkening effects (non-uniform stellar brightness)\n\n* Generalize across diverse planetary systems with varying:\n\n* Stellar types (hot/cold, large/small stars)\n\n* Orbital configurations (close/far orbits, inclined paths)","metadata":{}},{"cell_type":"markdown","source":"## 0. All Imports","metadata":{}},{"cell_type":"code","source":"!pip install -q scipy statsmodels\n!pip install -q scipy pingouin ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-27T07:44:55.605335Z","iopub.execute_input":"2025-06-27T07:44:55.605733Z","iopub.status.idle":"2025-06-27T07:45:02.993244Z","shell.execute_reply.started":"2025-06-27T07:44:55.60568Z","shell.execute_reply":"2025-06-27T07:45:02.991954Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport scipy.stats as stats\nfrom statsmodels.stats.multicomp import pairwise_tukeyhsd\nimport pingouin as pg\n\n# Load the data\ntrain = pd.read_csv('/kaggle/input/ariel-data-challenge-2025/train.csv')\nstar_info = pd.read_csv('/kaggle/input/ariel-data-challenge-2025/train_star_info.csv')\nmerged_data = pd.merge(train, star_info, on='planet_id')\n\n# Create wavelength summary statistics\nwl_columns = [col for col in train.columns if col.startswith('wl_')]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-27T07:45:02.995227Z","iopub.execute_input":"2025-06-27T07:45:02.995542Z","iopub.status.idle":"2025-06-27T07:45:03.118118Z","shell.execute_reply.started":"2025-06-27T07:45:02.995513Z","shell.execute_reply":"2025-06-27T07:45:03.117247Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# B. Descriptive Statistics","metadata":{}},{"cell_type":"markdown","source":" ## 1. Basic Statistics for System Parameters","metadata":{}},{"cell_type":"code","source":"desc_stats = merged_data[['Rs', 'Ms', 'Ts', 'Mp', 'P', 'sma', 'i']].describe().T\ndesc_stats['skewness'] = merged_data[['Rs', 'Ms', 'Ts', 'Mp', 'P', 'sma', 'i']].skew()\ndesc_stats['kurtosis'] = merged_data[['Rs', 'Ms', 'Ts', 'Mp', 'P', 'sma', 'i']].kurt()\ndisplay(desc_stats)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-27T07:45:03.118978Z","iopub.execute_input":"2025-06-27T07:45:03.119215Z","iopub.status.idle":"2025-06-27T07:45:03.151672Z","shell.execute_reply.started":"2025-06-27T07:45:03.119197Z","shell.execute_reply":"2025-06-27T07:45:03.150881Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Analysis :\n**count**: Number of observations (all variables have 1100 non-missing values).\n\n**mean**: The average value of the variable.\n\n**std**: Standard deviation (measures how spread out the data is).\n\n**min**: The smallest observed value.\n\n**25%** (First Quartile): 25% of the data is below this value.\n\n**50%** (Median): The middle value; 50% of the data is below this.\n\n**75%** (Third Quartile): 75% of the data is below this value.\n\n**max**: The largest observed value.\n\n**skewness**: Measures asymmetry in the data distribution.\n\n* **Positive skew**: Right-tailed (longer tail on the right).\n\n* **Negative skew**: Left-tailed (longer tail on the left).\n\n**kurtosis**: Measures \"tailedness\" (how extreme outliers are).\n\n* **Positive kurtosis**: Heavy tails (more outliers).\n\n* **Negative kurtosis**: Light tails (fewer outliers).\n\n","metadata":{}},{"cell_type":"markdown","source":"### Conclusion : \n**Mp** have HIGH Extreme Outliers\n\n**Ts** only have Left Skewness","metadata":{}},{"cell_type":"markdown","source":"## 2. Wavelength Statistics","metadata":{}},{"cell_type":"code","source":"wl_stats = train[wl_columns].agg(['mean', 'std', 'min', 'max', 'skew', 'kurtosis']).T\nwl_stats['cv'] = wl_stats['std'] / wl_stats['mean']  # Coefficient of variation\n\nplt.figure(figsize=(12, 6))\nplt.plot(wl_stats.index, wl_stats['mean'], label='Mean')\nplt.fill_between(wl_stats.index, \n                wl_stats['mean'] - wl_stats['std'], \n                wl_stats['mean'] + wl_stats['std'], \n                alpha=0.2, label='±1 Std Dev')\nplt.title('Mean Wavelength Flux with Standard Deviation')\nplt.xlabel('Wavelength Index')\nplt.ylabel('Normalized Flux')\nplt.legend()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-27T07:45:03.153589Z","iopub.execute_input":"2025-06-27T07:45:03.153876Z","iopub.status.idle":"2025-06-27T07:45:04.833634Z","shell.execute_reply.started":"2025-06-27T07:45:03.153855Z","shell.execute_reply":"2025-06-27T07:45:04.832753Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Conlusion :\n\nMean is **CONSTANT**","metadata":{}},{"cell_type":"markdown","source":"## 3. Correlation Analysis","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(10, 8))\ncorr_matrix = merged_data[['Rs', 'Ms', 'Ts', 'Mp', 'P', 'sma', 'i']].corr()\nsns.heatmap(corr_matrix, annot=True, cmap='coolwarm', center=0)\nplt.title('Correlation Matrix of System Parameters')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-27T07:45:04.834521Z","iopub.execute_input":"2025-06-27T07:45:04.834837Z","iopub.status.idle":"2025-06-27T07:45:05.170575Z","shell.execute_reply.started":"2025-06-27T07:45:04.834809Z","shell.execute_reply":"2025-06-27T07:45:05.169762Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Conlusion :\n\n**sma** and **P** have High CORELATION \n\nwhile **sma** and **Rs** and **Ms** have high NEGATIVE CORelation","metadata":{}},{"cell_type":"markdown","source":"# C. Inferential Statistics","metadata":{}},{"cell_type":"markdown","source":"## 1. Confidence Intervals for Key Parameters","metadata":{}},{"cell_type":"code","source":"def get_ci(data, confidence=0.95):\n    n = len(data)\n    mean = np.mean(data)\n    std_err = stats.sem(data)\n    h = std_err * stats.t.ppf((1 + confidence) / 2, n-1)\n    return mean, mean-h, mean+h\n\nparams = ['Ts', 'Mp', 'P']\nci_results = {}\n\nfor param in params:\n    ci_results[param] = get_ci(merged_data[param])\n    \nci_df = pd.DataFrame(ci_results, index=['mean', 'lower_ci', 'upper_ci']).T\ndisplay(ci_df)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-27T07:45:05.171694Z","iopub.execute_input":"2025-06-27T07:45:05.172613Z","iopub.status.idle":"2025-06-27T07:45:05.1881Z","shell.execute_reply.started":"2025-06-27T07:45:05.172579Z","shell.execute_reply":"2025-06-27T07:45:05.187328Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**mean**: The average value of the variable.\n\n**lower_ci**: The lower bound of the 95% confidence interval for the mean.\n\n**upper_ci**: The upper bound of the 95% confidence interval for the mean.","metadata":{}},{"cell_type":"markdown","source":"## Conclusion:\n\n### A 95% confidence interval (CI) means we are 95% confident that the true population mean lies within this range.\n\n### Ts has a wide CI due to its large scale (mean ~5839).\n\n### Mp has a relatively wide CI compared to its mean, likely due to high variance and skewness.\n\n### P has the narrowest CI, suggesting more stable estimates.","metadata":{}},{"cell_type":"markdown","source":"## 2. Comparing Hot vs Cool Stars","metadata":{}},{"cell_type":"code","source":"# Create temperature groups\nmerged_data['temp_group'] = pd.qcut(merged_data['Ts'], q=3, labels=['Cool', 'Medium', 'Hot'])\n\n# Compare planet masses\nhot_cool = merged_data[merged_data['temp_group'].isin(['Hot', 'Cool'])]\nprint(stats.ttest_ind(\n    hot_cool[hot_cool['temp_group'] == 'Hot']['Mp'],\n    hot_cool[hot_cool['temp_group'] == 'Cool']['Mp'],\n    equal_var=False\n))\n\n# Visual comparison\nplt.figure(figsize=(10, 6))\nsns.boxplot(x='temp_group', y='Mp', data=hot_cool)\nplt.title('Planet Mass Distribution by Star Temperature Group')\nplt.ylabel('Planet Mass (Mj)')\nplt.xlabel('Star Temperature Group')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-27T07:45:05.188955Z","iopub.execute_input":"2025-06-27T07:45:05.189211Z","iopub.status.idle":"2025-06-27T07:45:05.37193Z","shell.execute_reply.started":"2025-06-27T07:45:05.189192Z","shell.execute_reply":"2025-06-27T07:45:05.371218Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"p > 0.05 → Fail to reject the null hypothesis (no statistically significant difference).\n\np ≤ 0.05 → Reject the null (significant difference).\n\n### Here, 0.136 > 0.05, so no significant difference between groups.","metadata":{}},{"cell_type":"markdown","source":"### Conclusion: There is no strong evidence that the two group means are different (*p = 0.136*).","metadata":{}},{"cell_type":"markdown","source":"# D. Hypothesis Testing","metadata":{}},{"cell_type":"markdown","source":"## 1. Normality Tests","metadata":{}},{"cell_type":"code","source":"normality_results = {}\nfor param in ['Ts', 'Mp', 'P', 'sma']:\n    stat, p = stats.shapiro(merged_data[param])\n    normality_results[param] = {'statistic': stat, 'p-value': p}\n    \nnormality_df = pd.DataFrame(normality_results).T\ndisplay(normality_df)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-27T07:45:05.37279Z","iopub.execute_input":"2025-06-27T07:45:05.373047Z","iopub.status.idle":"2025-06-27T07:45:05.386221Z","shell.execute_reply.started":"2025-06-27T07:45:05.373022Z","shell.execute_reply":"2025-06-27T07:45:05.385342Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Results Breakdown:\n\nVariable\tW-statistic\tp-value\tConclusion\n\nTs\t          0.970 \t      1.99 × 10⁻¹⁴\t    Non-normal (p ≪ 0.05)\n\nMp\t           0.519\t     1.92 × 10⁻⁴⁷\t    Extremely non-normal\n\nP\t          0.955\t         5.90 × 10⁻¹⁸\t    Non-normal\n\nsma\t          0.984\t         1.79 × 10⁻⁹\t     Non-normal (but close)","metadata":{}},{"cell_type":"markdown","source":"### Conclusion :\n\n### All variables are significantly non-normal ","metadata":{}},{"cell_type":"markdown","source":"## 2. ANOVA for Multiple Groups","metadata":{}},{"cell_type":"markdown","source":"This is a one-way ANOVA table, which tests **whether there are statistically significant differences between the means of three or more groups**","metadata":{}},{"cell_type":"code","source":"# One-way ANOVA for planet mass across temperature groups\nanova_result = pg.anova(data=merged_data, dv='Mp', between='temp_group', detailed=True)\ndisplay(anova_result)\n\n# Post-hoc test if ANOVA is significant\nif anova_result['p-unc'][0] < 0.05:\n    print(\"\\nPost-hoc Tukey HSD:\")\n    tukey = pairwise_tukeyhsd(endog=merged_data['Mp'],\n                             groups=merged_data['temp_group'],\n                             alpha=0.05)\n    print(tukey)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-27T07:45:05.387079Z","iopub.execute_input":"2025-06-27T07:45:05.38734Z","iopub.status.idle":"2025-06-27T07:45:05.421524Z","shell.execute_reply.started":"2025-06-27T07:45:05.387321Z","shell.execute_reply":"2025-06-27T07:45:05.420583Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Conclusion\n\n### No significant differences exist between the groups (p = 0.368)\n\n### The effect size (η² = 0.0018) is negligible, meaning temp_group has almost no practical impact on the outcome variable.","metadata":{}},{"cell_type":"markdown","source":"## 3. Non-parametric Tests","metadata":{}},{"cell_type":"markdown","source":"#### Kruskal-Wallis Test (Overall Group Differences)\nTests whether at least one group (e.g., Hot, Cool, Neutral) differs significantly from the others.\n\np = 0.055 → Borderline non-significance (just above the typical α = 0.05 threshold)\n\n#### Mann-Whitney U Test (Hot vs. Cool Only)\n\nDirectly compares Hot vs. Cool groups.\n\np = 0.039 → Statistically significant difference (p < 0.05).\n\nImplication: The distributions of Hot and Cool groups are not identical.","metadata":{}},{"cell_type":"code","source":"# Kruskal-Wallis test (non-parametric alternative to ANOVA)\nprint(\"Kruskal-Wallis Test:\")\nprint(stats.kruskal(\n    merged_data[merged_data['temp_group'] == 'Cool']['Mp'],\n    merged_data[merged_data['temp_group'] == 'Medium']['Mp'],\n    merged_data[merged_data['temp_group'] == 'Hot']['Mp']\n))\n\n# Mann-Whitney U test for two groups\nprint(\"\\nMann-Whitney U Test (Hot vs Cool):\")\nprint(stats.mannwhitneyu(\n    merged_data[merged_data['temp_group'] == 'Hot']['Mp'],\n    merged_data[merged_data['temp_group'] == 'Cool']['Mp']\n))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-27T07:45:05.423791Z","iopub.execute_input":"2025-06-27T07:45:05.424065Z","iopub.status.idle":"2025-06-27T07:45:05.438269Z","shell.execute_reply.started":"2025-06-27T07:45:05.424043Z","shell.execute_reply":"2025-06-27T07:45:05.437428Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Conlusion :\n\n#### Overall (Kruskal-Wallis): No strong evidence of differences across all groups (*p = 0.055*).\n\n#### Pairwise (Mann-Whitney): Hot and Cool differ significantly (*p = 0.039*)","metadata":{}},{"cell_type":"markdown","source":"# E. Advanced Statistical Analysis","metadata":{}},{"cell_type":"markdown","source":"## 1. Principal Component Analysis (PCA) of Wavelengths","metadata":{}},{"cell_type":"code","source":"from sklearn.decomposition import PCA\nfrom sklearn.preprocessing import StandardScaler\n\n# Standardize wavelength data\nscaler = StandardScaler()\nwl_scaled = scaler.fit_transform(train[wl_columns])\n\n# Perform PCA\npca = PCA(n_components=3)\npca_results = pca.fit_transform(wl_scaled)\n\n# Plot explained variance\nplt.figure(figsize=(10, 5))\nplt.bar(range(3), pca.explained_variance_ratio_)\nplt.title('PCA Explained Variance Ratio')\nplt.xlabel('Principal Component')\nplt.ylabel('Variance Explained')\nplt.show()\n\n# Plot PC1 vs PC2\nplt.figure(figsize=(10, 8))\nsc = plt.scatter(pca_results[:, 0], pca_results[:, 1], \n                c=merged_data['Ts'], cmap='viridis')\nplt.colorbar(sc, label='Star Temperature (K)')\nplt.title('PCA of Wavelength Data (PC1 vs PC2)')\nplt.xlabel('Principal Component 1')\nplt.ylabel('Principal Component 2')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-27T07:45:05.439373Z","iopub.execute_input":"2025-06-27T07:45:05.439671Z","iopub.status.idle":"2025-06-27T07:45:05.957782Z","shell.execute_reply.started":"2025-06-27T07:45:05.439651Z","shell.execute_reply":"2025-06-27T07:45:05.956753Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 2. Regression Analysis","metadata":{}},{"cell_type":"code","source":"from sklearn.linear_model import LinearRegression\nfrom sklearn.metrics import r2_score\n\n# Prepare data\nX = merged_data[['Rs', 'Ms', 'Ts']]\ny = merged_data['Mp']\n\n# Fit model\nmodel = LinearRegression().fit(X, y)\n\n# Print results\nprint(f\"Intercept: {model.intercept_:.4f}\")\nprint(\"Coefficients:\")\nfor name, coef in zip(X.columns, model.coef_):\n    print(f\"{name}: {coef:.4f}\")\n\n# Calculate R-squared\nr2 = r2_score(y, model.predict(X))\nprint(f\"\\nR-squared: {r2:.4f}\")\n\n# Plot results\nplt.figure(figsize=(8, 8))\nplt.scatter(y, model.predict(X))\nplt.plot([y.min(), y.max()], [y.min(), y.max()], 'r--')\nplt.xlabel('Actual Planet Mass')\nplt.ylabel('Predicted Planet Mass')\nplt.title('Linear Regression (scikit-learn)')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-27T07:45:05.958858Z","iopub.execute_input":"2025-06-27T07:45:05.959084Z","iopub.status.idle":"2025-06-27T07:45:06.172737Z","shell.execute_reply.started":"2025-06-27T07:45:05.959067Z","shell.execute_reply":"2025-06-27T07:45:06.171944Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# G. Analyze wavelength as time series for a planet","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport matplotlib.pyplot as plt\nfrom scipy.stats import norm\nfrom scipy.signal import correlate\n\ndef custom_acf(x, nlags=40):\n    \"\"\"Manual implementation of autocorrelation function\"\"\"\n    x = np.asarray(x) - np.mean(x)\n    corr = correlate(x, x, mode='full')[-len(x):-len(x)+nlags+1]\n    return corr / corr[0]\n\ndef custom_pacf(x, nlags=40):\n    \"\"\"Manual implementation of partial autocorrelation using Yule-Walker\"\"\"\n    x = np.asarray(x) - np.mean(x)\n    pacf = [1.0]\n    for k in range(1, nlags+1):\n        r = custom_acf(x, nlags=k)\n        # Solve Yule-Walker equations\n        R = np.array([r[:k]])\n        r_k = r[1:k+1]\n        phi = np.linalg.solve(np.toeplitz(R), r_k)\n        pacf.append(phi[-1])\n    return np.array(pacf)\n\ndef robust_wavelength_analysis(planet_index=0):\n    \"\"\"Completely self-contained time series analysis\"\"\"\n    try:\n        # Get sample planet data\n        sample_planet = train.iloc[planet_index][wl_columns].values\n        planet_id = train.iloc[planet_index][\"planet_id\"]\n        \n        # Basic statistics\n        print(f\"\\nAnalysis for Planet {planet_id}\")\n        print(f\"Mean flux: {np.mean(sample_planet):.4f}\")\n        print(f\"Std dev: {np.std(sample_planet):.4f}\")\n        \n        # Manual ACF/PACF\n        nlags = min(40, len(sample_planet)//3)  # Conservative lag value\n        acf_vals = custom_acf(sample_planet, nlags=nlags)\n        pacf_vals = custom_pacf(sample_planet, nlags=nlags)\n        \n        # Plotting\n        fig, (ax1, ax2) = plt.subplots(2, 1, figsize=(12, 8))\n        \n        # ACF plot\n        ax1.stem(acf_vals)\n        conf = norm.ppf(1 - 0.05/2) / np.sqrt(len(sample_planet))\n        ax1.axhspan(-conf, conf, alpha=0.2, color='blue')\n        ax1.set_title(f'Autocorrelation (Planet {planet_id})')\n        \n        # PACF plot\n        ax2.stem(pacf_vals)\n        ax2.axhspan(-conf, conf, alpha=0.2, color='blue')\n        ax2.set_title(f'Partial Autocorrelation (Planet {planet_id})')\n        \n        plt.tight_layout()\n        plt.show()\n        \n        # Basic stationarity check\n        diff = np.diff(sample_planet)\n        print(\"\\nStationarity indicators:\")\n        print(f\"Original variance: {np.var(sample_planet):.4f}\")\n        print(f\"1st difference variance: {np.var(diff):.4f}\")\n        \n    except Exception as e:\n        print(f\"Error analyzing planet {planet_index}: {str(e)}\")\n\n# Run analysis on first 3 planets\nfor i in range(3):\n    robust_wavelength_analysis(i)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-27T07:45:06.17381Z","iopub.execute_input":"2025-06-27T07:45:06.174542Z","iopub.status.idle":"2025-06-27T07:45:06.191809Z","shell.execute_reply.started":"2025-06-27T07:45:06.174518Z","shell.execute_reply":"2025-06-27T07:45:06.19106Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# F. Basic periodogram analysis","metadata":{}},{"cell_type":"code","source":"def spectral_analysis(planet_index=0):\n    \"\"\"Basic periodogram analysis\"\"\"\n    from scipy.signal import periodogram\n    \n    sample_planet = train.iloc[planet_index][wl_columns].values\n    freqs, psd = periodogram(sample_planet)\n    \n    plt.figure(figsize=(12,4))\n    plt.semilogy(freqs, psd)\n    plt.title('Power Spectral Density')\n    plt.xlabel('Frequency')\n    plt.ylabel('Power')\n    plt.show()\n\nspectral_analysis(0)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-27T07:45:06.192655Z","iopub.execute_input":"2025-06-27T07:45:06.192932Z","iopub.status.idle":"2025-06-27T07:45:06.482947Z","shell.execute_reply.started":"2025-06-27T07:45:06.192914Z","shell.execute_reply":"2025-06-27T07:45:06.482047Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# H. Plot functional principal components","metadata":{}},{"cell_type":"code","source":"!pip install -q scikit-fda\n\nfrom skfda import FDataGrid\nfrom skfda.preprocessing.dim_reduction import FPCA\n\n# Convert wavelength data to functional objects\nfd = FDataGrid(\n    data_matrix=train[wl_columns].values,\n    grid_points=np.arange(len(wl_columns))\n)\n\n# Functional PCA\nfpca = FPCA(n_components=3)\nfpca.fit(fd)\nscores = fpca.transform(fd)\n\n\nfig = plt.figure(figsize=(12, 6))\nfor i in range(3):\n    fd.mean().plot(label='Mean')\n    (fd.mean() + fpca.components_[i]*np.sqrt(fpca.explained_variance_[i])).plot(\n        label=f'FPC {i+1}')\n    plt.title(f'Functional Principal Component {i+1}')\n    plt.legend()\n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-27T07:45:21.39349Z","iopub.execute_input":"2025-06-27T07:45:21.393829Z","iopub.status.idle":"2025-06-27T07:45:29.155508Z","shell.execute_reply.started":"2025-06-27T07:45:21.393805Z","shell.execute_reply":"2025-06-27T07:45:29.154533Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# I. Planetary System Similarity Network","metadata":{}},{"cell_type":"code","source":"import networkx as nx\nfrom community import community_louvain\n\n# Create similarity network\ncorr_matrix = train[wl_columns].T.corr()\nG = nx.Graph()\n\nthreshold = 0.9\nfor i in range(len(corr_matrix)):\n    for j in range(i+1, len(corr_matrix)):\n        if corr_matrix.iloc[i,j] > threshold:\n            G.add_edge(i, j, weight=corr_matrix.iloc[i,j])\n\n# Community detection\npartition = community_louvain.best_partition(G)\n\n# Visualize\npos = nx.spring_layout(G)\nplt.figure(figsize=(12, 12))\nnx.draw_networkx_nodes(G, pos, node_size=50, \n                      cmap=plt.cm.viridis,\n                      node_color=list(partition.values()))\nnx.draw_networkx_edges(G, pos, alpha=0.1)\nplt.title('Planetary System Similarity Network')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-27T07:45:33.314826Z","iopub.execute_input":"2025-06-27T07:45:33.315371Z","iopub.status.idle":"2025-06-27T07:45:55.778475Z","shell.execute_reply.started":"2025-06-27T07:45:33.315339Z","shell.execute_reply":"2025-06-27T07:45:55.777534Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<div style=\"border: 2px solid #4CAF50; padding: 15px; border-radius: 10px; background-color: #ffffff; box-shadow: 2px 2px 8px rgba(0, 0, 0, 0.1);\">\n\n<h2 style=\"color: #4CAF50; text-align: center;\">Thank You! 🎉</h2>\n\n<p style=\"font-size: 16px; text-align: justify;\">\nThank you for exploring my notebook! I hope you found it insightful and helpful.\nIf you have any feedback, I’d love to hear from you!\nYour support and comments mean a lot to me. 😊\n</p>\n\n<div style=\"text-align: left; margin: 10px 0;\">\n<img src=\"https://i.pinimg.com/236x/90/e8/ff/90e8ff78f7bef8490e0ab9cb5b83ee0f.jpg\" alt=\"Thank You Image\" style=\"border-radius: 0%; border: 2px solid #ffffff;\">\n</div>\n\n<h3 style=\"color: #4CAF50;\">🌐 Useful Links:</h3>\n<ul style=\"font-size: 16px;\">\n<li><a href=\"https://www.kaggle.com/samasiayushman\" target=\"_blank\" style=\"color: #2196F3;\">My Kaggle Profile</a></li>\n<li><a href=\"https://github.com/Hariswar8018\" target=\"_blank\" style=\"color: #2196F3;\">Notebook Code Repository</a></li>\n</ul>\n\n<p style=\"font-size: 16px; text-align: center; font-weight: bold;\">\n✨ Don’t forget to upvote if you found this notebook helpful! ✨\n</p>\n\n</div>\n","metadata":{}}]}