{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":56537,"databundleVersionId":8877088,"sourceType":"competition"}],"dockerImageVersionId":30732,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# CAPSTONE Project: Atmospheric Physics using AI (ClimSim)\n## Samanyu Parvathaneni","metadata":{}},{"cell_type":"markdown","source":"## Abstract\nClimate models are crucial for understanding Earth's climate system. Due to the complexity of Earth's climate, these models use parameterizations to approximate the effects of physical processes that occur at scales smaller than the size of their grid cells. These approximations, however, are imperfect and contribute significantly to uncertainties in predicted warming, changing precipitation patterns, and the frequency and severity of extreme events. The Multi-scale Modeling Framework (MMF) approach, however, represents these subgrid processes more explicitly but at a computational cost too high for operational climate prediction.\n\nThis project aims to develop machine learning (ML) models that emulate subgrid atmospheric processes—such as storms, clouds, turbulence, rainfall, and radiation—within E3SM-MMF, a multi-scale climate model supported by the U.S. Department of Energy. ML emulators are significantly cheaper to run than MMF, and advancements in this area could enable high-resolution, physically credible long-term climate projections to become broadly accessible. This would enhance our understanding of climate-related hazards and empower policymakers with the knowledge needed to mitigate them.","metadata":{"execution":{"iopub.status.busy":"2024-07-21T16:33:50.341578Z","iopub.execute_input":"2024-07-21T16:33:50.342534Z","iopub.status.idle":"2024-07-21T16:33:50.385699Z","shell.execute_reply.started":"2024-07-21T16:33:50.342493Z","shell.execute_reply":"2024-07-21T16:33:50.384149Z"}}},{"cell_type":"markdown","source":"## Research Question\n\nMotivation: With the increasing frequency and severity of climate disasters, governments and other institutions are investing more time and money into projects that will be able to predict these catastrophies. We aim to model these weather events and figue out what features are the most predictive of natural disasters.\n\nResearch Question: What relationships can we find between observable weather data and climate events?","metadata":{}},{"cell_type":"markdown","source":"## Hypothesis\n\nWe believe that there will be many highly correlated variables, particularly the ones with the same number of dimensions.","metadata":{}},{"cell_type":"markdown","source":"## Data Preparation","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport dask as dd\nimport polars as p1\nimport ray\nimport gc  # Garbage Collector interface\nimport random\nimport time\n# import tensorflow as tf\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.preprocessing import StandardScaler\nfrom tensorflow.keras import layers, regularizers, models\nfrom tensorflow.keras.callbacks import EarlyStopping\nfrom sklearn.metrics import r2_score\n\n# Plots\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2024-08-25T05:37:33.934543Z","iopub.execute_input":"2024-08-25T05:37:33.934946Z","iopub.status.idle":"2024-08-25T05:37:49.253576Z","shell.execute_reply.started":"2024-08-25T05:37:33.934914Z","shell.execute_reply":"2024-08-25T05:37:49.252428Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Exploratory Data Analysis (EDA)","metadata":{}},{"cell_type":"code","source":"train_df = p1.read_csv('../input/leap-atmospheric-physics-ai-climsim/train.csv', n_rows = 100000).to_pandas()\ntest_df = p1.read_csv('../input/leap-atmospheric-physics-ai-climsim/test.csv', n_rows = 100000).to_pandas()","metadata":{"execution":{"iopub.status.busy":"2024-08-25T05:38:23.894013Z","iopub.execute_input":"2024-08-25T05:38:23.894745Z","iopub.status.idle":"2024-08-25T05:38:39.34968Z","shell.execute_reply.started":"2024-08-25T05:38:23.894709Z","shell.execute_reply":"2024-08-25T05:38:39.348841Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.head(10)","metadata":{"execution":{"iopub.status.busy":"2024-08-24T17:40:05.011924Z","iopub.execute_input":"2024-08-24T17:40:05.012523Z","iopub.status.idle":"2024-08-24T17:40:05.066904Z","shell.execute_reply.started":"2024-08-24T17:40:05.012478Z","shell.execute_reply":"2024-08-24T17:40:05.065714Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df.head(10)","metadata":{"execution":{"iopub.status.busy":"2024-08-24T17:40:07.044242Z","iopub.execute_input":"2024-08-24T17:40:07.044697Z","iopub.status.idle":"2024-08-24T17:40:07.096173Z","shell.execute_reply.started":"2024-08-24T17:40:07.044662Z","shell.execute_reply":"2024-08-24T17:40:07.094881Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# targets (extract from submission file)\ntargets = [x for x in submission_df.columns if x not in ['sample_id']]\n\n# numerical features\nnumerical_features = [x for x in train_df.columns if x not in ['sample_id']+targets]\n\n# categorical features\ncategorical_features = []\n\n# all features combined\nfeatures_all = numerical_features + categorical_features","metadata":{"execution":{"iopub.status.busy":"2024-08-24T17:40:12.198006Z","iopub.execute_input":"2024-08-24T17:40:12.198457Z","iopub.status.idle":"2024-08-24T17:40:12.216874Z","shell.execute_reply.started":"2024-08-24T17:40:12.198425Z","shell.execute_reply":"2024-08-24T17:40:12.215324Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Output of dimensions\nprint('# of numerical features: ', len(numerical_features))\nprint('# of categorical features: ', len(categorical_features))\nprint('# of targets: ', len(targets))","metadata":{"execution":{"iopub.status.busy":"2024-08-24T17:40:13.190416Z","iopub.execute_input":"2024-08-24T17:40:13.191203Z","iopub.status.idle":"2024-08-24T17:40:13.199047Z","shell.execute_reply.started":"2024-08-24T17:40:13.191167Z","shell.execute_reply":"2024-08-24T17:40:13.197611Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df[targets].describe()","metadata":{"execution":{"iopub.status.busy":"2024-08-24T17:40:43.494796Z","iopub.execute_input":"2024-08-24T17:40:43.496112Z","iopub.status.idle":"2024-08-24T17:40:46.0291Z","shell.execute_reply.started":"2024-08-24T17:40:43.496057Z","shell.execute_reply":"2024-08-24T17:40:46.027885Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plot target distributions in compact matrix form\nfig, axs = plt.subplots(92, 4, figsize=(16, 350))\ni = 0\nfor t in targets:\n    current_ax = axs.flat[i]\n    current_ax.hist(train_df[t], bins=100, color='darkgreen')\n    current_ax.set_title('Target' + str(t))\n    current_ax.grid()\n    i += 1","metadata":{"_kg_hide-output":true,"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-08-24T19:24:42.709334Z","iopub.execute_input":"2024-08-24T19:24:42.710311Z","iopub.status.idle":"2024-08-24T19:27:21.860166Z","shell.execute_reply.started":"2024-08-24T19:24:42.710265Z","shell.execute_reply":"2024-08-24T19:27:21.858329Z"},"scrolled":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Basic stats - Train\ntrain_df[numerical_features].describe()","metadata":{"execution":{"iopub.status.busy":"2024-08-24T17:47:34.214475Z","iopub.execute_input":"2024-08-24T17:47:34.215042Z","iopub.status.idle":"2024-08-24T17:47:40.024909Z","shell.execute_reply.started":"2024-08-24T17:47:34.215007Z","shell.execute_reply":"2024-08-24T17:47:40.023665Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Basic stats - Test\ntest_df[numerical_features].describe()","metadata":{"execution":{"iopub.status.busy":"2024-08-24T17:48:03.355017Z","iopub.execute_input":"2024-08-24T17:48:03.356241Z","iopub.status.idle":"2024-08-24T17:48:06.970386Z","shell.execute_reply.started":"2024-08-24T17:48:03.356198Z","shell.execute_reply":"2024-08-24T17:48:06.969117Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plot histograms for numerical features (train and test)\nfor f in numerical_features:\n    plt.figure(figsize=(12,2))\n    ax1 = plt.subplot(1,2,1)\n    train_df[f].plot(kind='hist', bins=100, color='darkblue')\n    plt.title(f + ' - Train')\n    plt.grid()\n    ax2 = plt.subplot(1,2,2, sharex=ax1)\n    test_df[f].plot(kind='hist', bins=100, color='darkred')\n    plt.title(f + ' - Test')\n    plt.grid()\n    # plt.show()","metadata":{"execution":{"iopub.status.busy":"2024-08-24T17:50:28.131995Z","iopub.execute_input":"2024-08-24T17:50:28.132503Z","iopub.status.idle":"2024-08-24T18:01:07.562284Z","shell.execute_reply.started":"2024-08-24T17:50:28.132464Z","shell.execute_reply":"2024-08-24T18:01:07.561022Z"},"_kg_hide-input":false,"_kg_hide-output":true,"scrolled":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# compact boxplot of all features - train only\nn_plot_rows = 10\nn_plot_cols = 60\nn = len(numerical_features)\nfor i in range(n_plot_rows):\n    a = n_plot_cols*i+1\n    b = min(n_plot_cols*i+n_plot_cols, n)\n    print('Columns', a, 'to', b)\n    train_df.iloc[:,a:(b+1)].plot(kind='box', figsize=(15,5))\n    plt.xticks(rotation=90)\n    plt.grid()\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2024-08-24T18:03:48.824329Z","iopub.execute_input":"2024-08-24T18:03:48.824911Z","iopub.status.idle":"2024-08-24T18:04:06.756408Z","shell.execute_reply.started":"2024-08-24T18:03:48.824875Z","shell.execute_reply":"2024-08-24T18:04:06.7549Z"},"scrolled":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for f in numerical_features:\n    plt.figure(figsize=(14,0.5))\n    ax1 = plt.subplot(1,2,1)\n    temp_df = train_df[f].dropna() # boxplot does not like missings...\n    plt.boxplot(temp_df, vert=False)\n    plt.title(f + ' - Train')\n    plt.grid()\n    ax2 = plt.subplot(1,2,2, sharex=ax1)\n    temp_df = test_df[f].dropna()\n    plt.boxplot(temp_df, vert=False)\n    plt.title(f + ' - Test')\n    plt.grid()\n    plt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-08-24T18:09:46.233216Z","iopub.execute_input":"2024-08-24T18:09:46.234066Z","iopub.status.idle":"2024-08-24T18:13:55.82105Z","shell.execute_reply.started":"2024-08-24T18:09:46.234025Z","shell.execute_reply":"2024-08-24T18:13:55.819717Z"},"scrolled":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Calc and plot correlation matrix\ncor_p_target = train_df[targets].corr(method='pearson')\nplt.figure(figsize=(14,12))\nsns.heatmap(cor_p_target, annot=False, cmap='RdYlGn',\n            vmin=-1, vmax=+1)\nplt.title('Targets - Pearson Correlation')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-08-24T18:15:40.864088Z","iopub.execute_input":"2024-08-24T18:15:40.864575Z","iopub.status.idle":"2024-08-24T18:16:19.028193Z","shell.execute_reply.started":"2024-08-24T18:15:40.864538Z","shell.execute_reply":"2024-08-24T18:16:19.02693Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Calc and plot correlation matrix\ncor_p_train = train_df[numerical_features].corr(method='pearson')\nplt.figure(figsize=(14,12))\nsns.heatmap(cor_p_train, annot=False, cmap='RdYlGn',\n            vmin=-1, vmax=+1)\nplt.title('Features - Pearson Correlation (train)')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-08-24T18:16:40.013657Z","iopub.execute_input":"2024-08-24T18:16:40.014792Z","iopub.status.idle":"2024-08-24T18:18:05.058924Z","shell.execute_reply.started":"2024-08-24T18:16:40.014751Z","shell.execute_reply":"2024-08-24T18:18:05.057703Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Calc and plot correlation matrix\ncor_p_test = test_df[numerical_features].corr(method='pearson')\nplt.figure(figsize=(14,12))\nsns.heatmap(cor_p_test, annot=False, cmap='RdYlGn',\n            vmin=-1, vmax=+1)\nplt.title('Features - Pearson Correlation (test)')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-08-24T18:18:56.621947Z","iopub.execute_input":"2024-08-24T18:18:56.622404Z","iopub.status.idle":"2024-08-24T18:20:21.832381Z","shell.execute_reply.started":"2024-08-24T18:18:56.622372Z","shell.execute_reply":"2024-08-24T18:20:21.831082Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plot target values\nfor t in targets:\n    plt.figure(figsize=(14,2))\n    plt.scatter(train_df.index, train_df[t], color='darkblue',\n                alpha=0.25, s=1)\n    plt.title(t)\n    plt.grid()\n    plt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-08-24T18:20:53.601766Z","iopub.execute_input":"2024-08-24T18:20:53.602215Z","iopub.status.idle":"2024-08-24T18:22:51.621926Z","shell.execute_reply.started":"2024-08-24T18:20:53.602181Z","shell.execute_reply":"2024-08-24T18:22:51.620523Z"},"scrolled":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Define columns (has to be a list)\nn_max = 20 # columns with index 0..n_max\ncols_select = ['state_t_' + str(t) for t in range(0,n_max+1)]\nprint(cols_select)","metadata":{"execution":{"iopub.status.busy":"2024-08-24T18:40:01.237158Z","iopub.execute_input":"2024-08-24T18:40:01.238209Z","iopub.status.idle":"2024-08-24T18:40:01.245763Z","shell.execute_reply.started":"2024-08-24T18:40:01.238162Z","shell.execute_reply":"2024-08-24T18:40:01.244508Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load only selected column\nt1 = time.time()\ndf_col = p1.read_csv('../input/leap-atmospheric-physics-ai-climsim/train.csv', columns=cols_select).to_pandas()\nt2 = time.time()\nprint('Elapsed time [s]: ', np.round(t2-t1,2))\nprint('Number of rows: ', df_col.shape[0])\nprint('Number of cols: ', df_col.shape[1])","metadata":{"execution":{"iopub.status.busy":"2024-08-24T18:42:14.37037Z","iopub.execute_input":"2024-08-24T18:42:14.370902Z","iopub.status.idle":"2024-08-24T18:54:40.562649Z","shell.execute_reply.started":"2024-08-24T18:42:14.370858Z","shell.execute_reply":"2024-08-24T18:54:40.560981Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Basic stats\ndf_col.describe(percentiles=[0.01,0.1,0.25,0.5,0.75,0.9,0.99])","metadata":{"execution":{"iopub.status.busy":"2024-08-24T18:55:17.187087Z","iopub.execute_input":"2024-08-24T18:55:17.187568Z","iopub.status.idle":"2024-08-24T18:55:30.126824Z","shell.execute_reply.started":"2024-08-24T18:55:17.18753Z","shell.execute_reply":"2024-08-24T18:55:30.125407Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# plot distributions\nfor f in cols_select:\n    plt.figure(figsize=(10,3))\n    plt.hist(df_col[f], bins=1000, color='darkred')\n    plt.title(f + ' - full data')\n    plt.grid()\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2024-08-24T19:01:27.587747Z","iopub.execute_input":"2024-08-24T19:01:27.588612Z","iopub.status.idle":"2024-08-24T19:02:13.337861Z","shell.execute_reply.started":"2024-08-24T19:01:27.588514Z","shell.execute_reply":"2024-08-24T19:02:13.336673Z"},"scrolled":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Boxplot\nplt.figure(figsize=(10,0.5))\nplt.boxplot(df_col.state_t_0, vert=False)\nplt.title('state_t_0')\nplt.grid()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-08-24T19:02:41.836256Z","iopub.execute_input":"2024-08-24T19:02:41.836742Z","iopub.status.idle":"2024-08-24T19:02:42.873015Z","shell.execute_reply.started":"2024-08-24T19:02:41.836702Z","shell.execute_reply":"2024-08-24T19:02:42.871698Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Let's check the most extreme outliers\ndf_col[df_col.state_t_0>400]","metadata":{"execution":{"iopub.status.busy":"2024-08-24T19:03:25.200609Z","iopub.execute_input":"2024-08-24T19:03:25.201514Z","iopub.status.idle":"2024-08-24T19:03:25.252494Z","shell.execute_reply.started":"2024-08-24T19:03:25.201473Z","shell.execute_reply":"2024-08-24T19:03:25.251213Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Calc and plot correlation matrix\ncor_p_train_few_cols = train_df[cols_select].corr(method='pearson')\nplt.figure(figsize=(14,10))\nsns.heatmap(cor_p_train_few_cols, annot=True, cmap='RdYlGn',\n            fmt='.2f', linecolor='black', linewidth=.5,\n            vmin=-1, vmax=+1)\nplt.title('Features - Pearson Correlation (train/selected columns)')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-08-24T19:03:39.622681Z","iopub.execute_input":"2024-08-24T19:03:39.623138Z","iopub.status.idle":"2024-08-24T19:03:41.458584Z","shell.execute_reply.started":"2024-08-24T19:03:39.623103Z","shell.execute_reply":"2024-08-24T19:03:41.457259Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Model Training","metadata":{}},{"cell_type":"markdown","source":"#### Neural Network (Keras) with ReLU activation, batch normalization, and dropout","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.preprocessing import StandardScaler\nfrom tensorflow.keras import layers, regularizers, models\nfrom tensorflow.keras.callbacks import EarlyStopping\nfrom sklearn.metrics import r2_score\nimport matplotlib.pyplot as plt\n\n# Assume train_df and test_df are already loaded with the specified columns\n\n# Separate features and targets\nfeature_columns = ['state_t_0', 'state_t_1', 'state_t_2', 'state_t_3', 'state_t_4', 'state_t_5', 'state_t_6', 'state_t_7', 'state_t_8','state_t_9', \n                   'state_t_10', 'state_t_11', 'state_t_12', 'state_t_13', 'state_t_14', 'state_t_15', 'state_t_16', 'state_t_17', 'state_t_18', 'state_t_19', \n                   'state_t_20','state_t_21', 'state_t_22', 'state_t_23', 'state_t_24', 'state_t_25', 'state_t_26', 'state_t_27', 'state_t_28', 'state_t_29', \n                   'state_t_30', 'state_t_31', 'state_t_32', 'state_t_33', 'state_t_34', 'state_t_35', 'state_t_36', 'state_t_37', 'state_t_38', 'state_t_39',\n                   'state_t_40', 'state_t_41', 'state_t_42', 'state_t_43', 'state_t_44', 'state_t_45', 'state_t_46', 'state_t_47', 'state_t_48', 'state_t_49', \n                   'state_t_50', 'state_t_51', 'state_t_52', 'state_t_53', 'state_t_54', 'state_t_55', 'state_t_56', 'state_t_57', 'state_t_58','state_v_59', \n                   'state_q0001_0', 'state_q0001_1', 'state_q0001_2', 'state_q0001_3', 'state_q0001_4', 'state_q0001_5', 'state_q0001_6', 'state_q0001_7', 'state_q0001_8','state_q0001_9', \n                   'state_q0001_10', 'state_q0001_11', 'state_q0001_12', 'state_q0001_13', 'state_q0001_14', 'state_q0001_15', 'state_q0001_16', 'state_q0001_17', 'state_q0001_18', 'state_q0001_19', \n                   'state_q0001_20','state_q0001_21', 'state_q0001_22', 'state_q0001_23', 'state_q0001_24', 'state_q0001_25', 'state_q0001_26', 'state_q0001_27', 'state_q0001_28', 'state_q0001_29', \n                   'state_q0001_30', 'state_q0001_31', 'state_q0001_32', 'state_q0001_33', 'state_q0001_34', 'state_q0001_35', 'state_q0001_36', 'state_q0001_37', 'state_q0001_38', 'state_q0001_39',\n                   'state_q0001_40', 'state_q0001_41', 'state_q0001_42', 'state_q0001_43', 'state_q0001_44', 'state_q0001_45', 'state_q0001_46', 'state_q0001_47', 'state_q0001_48', 'state_q0001_49', \n                   'state_q0001_50', 'state_q0001_51', 'state_q0001_52', 'state_q0001_53', 'state_q0001_54', 'state_q0001_55', 'state_q0001_56', 'state_q0001_57', 'state_q0001_58','state_v_59',\n                   'state_q0002_0', 'state_q0002_1', 'state_q0002_2', 'state_q0002_3', 'state_q0002_4', 'state_q0002_5', 'state_q0002_6', 'state_q0002_7', 'state_q0002_8','state_q0002_9', \n                   'state_q0002_10', 'state_q0002_11', 'state_q0002_12', 'state_q0002_13', 'state_q0002_14', 'state_q0002_15', 'state_q0002_16', 'state_q0002_17', 'state_q0002_18', 'state_q0002_19', \n                   'state_q0002_20','state_q0002_21', 'state_q0002_22', 'state_q0002_23', 'state_q0002_24', 'state_q0002_25', 'state_q0002_26', 'state_q0002_27', 'state_q0002_28', 'state_q0002_29', \n                   'state_q0002_30', 'state_q0002_31', 'state_q0002_32', 'state_q0002_33', 'state_q0002_34', 'state_q0002_35', 'state_q0002_36', 'state_q0002_37', 'state_q0002_38', 'state_q0002_39',\n                   'state_q0002_40', 'state_q0002_41', 'state_q0002_42', 'state_q0002_43', 'state_q0002_44', 'state_q0002_45', 'state_q0002_46', 'state_q0002_47', 'state_q0002_48', 'state_q0002_49', \n                   'state_q0002_50', 'state_q0002_51', 'state_q0002_52', 'state_q0002_53', 'state_q0002_54', 'state_q0002_55', 'state_q0002_56', 'state_q0002_57', 'state_q0002_58','state_v_59',\n                   'state_q0003_0', 'state_q0003_1', 'state_q0003_2', 'state_q0003_3', 'state_q0003_4', 'state_q0003_5', 'state_q0003_6', 'state_q0003_7', 'state_q0003_8','state_q0003_9', \n                   'state_q0003_10', 'state_q0003_11', 'state_q0003_12', 'state_q0003_13', 'state_q0003_14', 'state_q0003_15', 'state_q0003_16', 'state_q0003_17', 'state_q0003_18', 'state_q0003_19', \n                   'state_q0003_20','state_q0003_21', 'state_q0003_22', 'state_q0003_23', 'state_q0003_24', 'state_q0003_25', 'state_q0003_26', 'state_q0003_27', 'state_q0003_28', 'state_q0003_29', \n                   'state_q0003_30', 'state_q0003_31', 'state_q0003_32', 'state_q0003_33', 'state_q0003_34', 'state_q0003_35', 'state_q0003_36', 'state_q0003_37', 'state_q0003_38', 'state_q0003_39',\n                   'state_q0003_40', 'state_q0003_41', 'state_q0003_42', 'state_q0003_43', 'state_q0003_44', 'state_q0003_45', 'state_q0003_46', 'state_q0003_47', 'state_q0003_48', 'state_q0003_49', \n                   'state_q0003_50', 'state_q0003_51', 'state_q0003_52', 'state_q0003_53', 'state_q0003_54', 'state_q0003_55', 'state_q0003_56', 'state_q0003_57', 'state_q0003_58','state_v_59',\n                   'state_u_0', 'state_u_1', 'state_u_2', 'state_u_3', 'state_u_4', 'state_u_5', 'state_u_6', 'state_u_7', 'state_u_8','state_u_9', \n                   'state_u_10', 'state_u_11', 'state_u_12', 'state_u_13', 'state_u_14', 'state_u_15', 'state_u_16', 'state_u_17', 'state_u_18', 'state_u_19', \n                   'state_u_20','state_u_21', 'state_u_22', 'state_u_23', 'state_u_24', 'state_u_25', 'state_u_26', 'state_u_27', 'state_u_28', 'state_u_29', \n                   'state_u_30', 'state_u_31', 'state_u_32', 'state_u_33', 'state_u_34', 'state_u_35', 'state_u_36', 'state_u_37', 'state_u_38', 'state_u_39',\n                   'state_u_40', 'state_u_41', 'state_u_42', 'state_u_43', 'state_u_44', 'state_u_45', 'state_u_46', 'state_u_47', 'state_u_48', 'state_u_49', \n                   'state_u_50', 'state_u_51', 'state_u_52', 'state_u_53', 'state_u_54', 'state_u_55', 'state_u_56', 'state_u_57', 'state_u_58','state_v_59',\n                   'state_v_0', 'state_v_1', 'state_v_2', 'state_v_3', 'state_v_4', 'state_v_5', 'state_v_6', 'state_v_7', 'state_v_8','state_v_9', \n                   'state_v_10', 'state_v_11', 'state_v_12', 'state_v_13', 'state_v_14', 'state_v_15', 'state_v_16', 'state_v_17', 'state_v_18', 'state_v_19', \n                   'state_v_20','state_v_21', 'state_v_22', 'state_v_23', 'state_v_24', 'state_v_25', 'state_v_26', 'state_v_27', 'state_v_28', 'state_v_29', \n                   'state_v_30', 'state_v_31', 'state_v_32', 'state_v_33', 'state_v_34', 'state_v_35', 'state_v_36', 'state_v_37', 'state_v_38', 'state_v_39',\n                   'state_v_40', 'state_v_41', 'state_v_42', 'state_v_43', 'state_v_44', 'state_v_45', 'state_v_46', 'state_v_47', 'state_v_48', 'state_v_49', \n                   'state_v_50', 'state_v_51', 'state_v_52', 'state_v_53', 'state_v_54', 'state_v_55', 'state_v_56', 'state_v_57', 'state_v_58','state_v_59',\n                   'state_ps', 'pbuf_SOLIN', 'pbuf_LHFLX', 'pbuf_TAUX', 'pbuf_TAUY', 'pbuf_COSZRS', 'cam_in_ALDIF', 'cam_in_ASDIF', 'cam_in_LWUP', 'cam_in_ICEFRAC', 'cam_in_LANDFRAC', 'cam_in_OCNFRAC', 'cam_in_SNOWHLAND',\n                   'pbuf_ozone_0', 'pbuf_ozone_1', 'pbuf_ozone_2', 'pbuf_ozone_3', 'pbuf_ozone_4', 'pbuf_ozone_5', 'pbuf_ozone_6', 'pbuf_ozone_7', 'pbuf_ozone_8','pbuf_ozone_9', \n                   'pbuf_ozone_10', 'pbuf_ozone_11', 'pbuf_ozone_12', 'pbuf_ozone_13', 'pbuf_ozone_14', 'pbuf_ozone_15', 'pbuf_ozone_16', 'pbuf_ozone_17', 'pbuf_ozone_18', 'pbuf_ozone_19', \n                   'pbuf_ozone_20','pbuf_ozone_21', 'pbuf_ozone_22', 'pbuf_ozone_23', 'pbuf_ozone_24', 'pbuf_ozone_25', 'pbuf_ozone_26', 'pbuf_ozone_27', 'pbuf_ozone_28', 'pbuf_ozone_29', \n                   'pbuf_ozone_30', 'pbuf_ozone_31', 'pbuf_ozone_32', 'pbuf_ozone_33', 'pbuf_ozone_34', 'pbuf_ozone_35', 'pbuf_ozone_36', 'pbuf_ozone_37', 'pbuf_ozone_38', 'pbuf_ozone_39',\n                   'pbuf_ozone_40', 'pbuf_ozone_41', 'pbuf_ozone_42', 'pbuf_ozone_43', 'pbuf_ozone_44', 'pbuf_ozone_45', 'pbuf_ozone_46', 'pbuf_ozone_47', 'pbuf_ozone_48', 'pbuf_ozone_49', \n                   'pbuf_ozone_50', 'pbuf_ozone_51', 'pbuf_ozone_52', 'pbuf_ozone_53', 'pbuf_ozone_54', 'pbuf_ozone_55', 'pbuf_ozone_56', 'pbuf_ozone_57', 'pbuf_ozone_58','pbuf_ozone_59',\n                   'pbuf_CH4_0', 'pbuf_CH4_1', 'pbuf_CH4_2', 'pbuf_CH4_3', 'pbuf_CH4_4', 'pbuf_CH4_5', 'pbuf_CH4_6', 'pbuf_CH4_7', 'pbuf_CH4_8','pbuf_CH4_9', \n                   'pbuf_CH4_10', 'pbuf_CH4_11', 'pbuf_CH4_12', 'pbuf_CH4_13', 'pbuf_CH4_14', 'pbuf_CH4_15', 'pbuf_CH4_16', 'pbuf_CH4_17', 'pbuf_CH4_18', 'pbuf_CH4_19', \n                   'pbuf_CH4_20','pbuf_CH4_21', 'pbuf_CH4_22', 'pbuf_CH4_23', 'pbuf_CH4_24', 'pbuf_CH4_25', 'pbuf_CH4_26', 'pbuf_CH4_27', 'pbuf_CH4_28', 'pbuf_CH4_29', \n                   'pbuf_CH4_30', 'pbuf_CH4_31', 'pbuf_CH4_32', 'pbuf_CH4_33', 'pbuf_CH4_34', 'pbuf_CH4_35', 'pbuf_CH4_36', 'pbuf_CH4_37', 'pbuf_CH4_38', 'pbuf_CH4_39',\n                   'pbuf_CH4_40', 'pbuf_CH4_41', 'pbuf_CH4_42', 'pbuf_CH4_43', 'pbuf_CH4_44', 'pbuf_CH4_45', 'pbuf_CH4_46', 'pbuf_CH4_47', 'pbuf_CH4_48', 'pbuf_CH4_49', \n                   'pbuf_CH4_50', 'pbuf_CH4_51', 'pbuf_CH4_52', 'pbuf_CH4_53', 'pbuf_CH4_54', 'pbuf_CH4_55', 'pbuf_CH4_56', 'pbuf_CH4_57', 'pbuf_CH4_58','pbuf_CH4_59',\n                   'pbuf_N2O_0', 'pbuf_N2O_1', 'pbuf_N2O_2', 'pbuf_N2O_3', 'pbuf_N2O_4', 'pbuf_N2O_5', 'pbuf_N2O_6', 'pbuf_N2O_7', 'pbuf_N2O_8','pbuf_N2O_9', \n                   'pbuf_N2O_10', 'pbuf_N2O_11', 'pbuf_N2O_12', 'pbuf_N2O_13', 'pbuf_N2O_14', 'pbuf_N2O_15', 'pbuf_N2O_16', 'pbuf_N2O_17', 'pbuf_N2O_18', 'pbuf_N2O_19', \n                   'pbuf_N2O_20','pbuf_N2O_21', 'pbuf_N2O_22', 'pbuf_N2O_23', 'pbuf_N2O_24', 'pbuf_N2O_25', 'pbuf_N2O_26', 'pbuf_N2O_27', 'pbuf_N2O_28', 'pbuf_N2O_29', \n                   'pbuf_N2O_30', 'pbuf_N2O_31', 'pbuf_N2O_32', 'pbuf_N2O_33', 'pbuf_N2O_34', 'pbuf_N2O_35', 'pbuf_N2O_36', 'pbuf_N2O_37', 'pbuf_N2O_38', 'pbuf_N2O_39',\n                   'pbuf_N2O_40', 'pbuf_N2O_41', 'pbuf_N2O_42', 'pbuf_N2O_43', 'pbuf_N2O_44', 'pbuf_N2O_45', 'pbuf_N2O_46', 'pbuf_N2O_47', 'pbuf_N2O_48', 'pbuf_N2O_49', \n                   'pbuf_N2O_50', 'pbuf_N2O_51', 'pbuf_N2O_52', 'pbuf_N2O_53', 'pbuf_N2O_54', 'pbuf_N2O_55', 'pbuf_N2O_56', 'pbuf_N2O_57', 'pbuf_N2O_58','pbuf_N2O_59']\n\ntarget_columns = ['ptend_t_0', 'ptend_t_1', 'ptend_t_2', 'ptend_t_3', 'ptend_t_4', 'ptend_t_5', 'ptend_t_6', 'ptend_t_7', 'ptend_t_8','ptend_t_9', \n                   'ptend_t_10', 'ptend_t_11', 'ptend_t_12', 'ptend_t_13', 'ptend_t_14', 'ptend_t_15', 'ptend_t_16', 'ptend_t_17', 'ptend_t_18', 'ptend_t_19', \n                   'ptend_t_20','ptend_t_21', 'ptend_t_22', 'ptend_t_23', 'ptend_t_24', 'ptend_t_25', 'ptend_t_26', 'ptend_t_27', 'ptend_t_28', 'ptend_t_29', \n                   'ptend_t_30', 'ptend_t_31', 'ptend_t_32', 'ptend_t_33', 'ptend_t_34', 'ptend_t_35', 'ptend_t_36', 'ptend_t_37', 'ptend_t_38', 'ptend_t_39',\n                   'ptend_t_40', 'ptend_t_41', 'ptend_t_42', 'ptend_t_43', 'ptend_t_44', 'ptend_t_45', 'ptend_t_46', 'ptend_t_47', 'ptend_t_48', 'ptend_t_49', \n                   'ptend_t_50', 'ptend_t_51', 'ptend_t_52', 'ptend_t_53', 'ptend_t_54', 'ptend_t_55', 'ptend_t_56', 'ptend_t_57', 'ptend_t_58','ptend_t_59',\n                  'ptend_q0001_0', 'ptend_q0001_1', 'ptend_q0001_2', 'ptend_q0001_3', 'ptend_q0001_4', 'ptend_q0001_5', 'ptend_q0001_6', 'ptend_q0001_7', 'ptend_q0001_8','ptend_q0001_9', \n                   'ptend_q0001_10', 'ptend_q0001_11', 'ptend_q0001_12', 'ptend_q0001_13', 'ptend_q0001_14', 'ptend_q0001_15', 'ptend_q0001_16', 'ptend_q0001_17', 'ptend_q0001_18', 'ptend_q0001_19', \n                   'ptend_q0001_20','ptend_q0001_21', 'ptend_q0001_22', 'ptend_q0001_23', 'ptend_q0001_24', 'ptend_q0001_25', 'ptend_q0001_26', 'ptend_q0001_27', 'ptend_q0001_28', 'ptend_q0001_29', \n                   'ptend_q0001_30', 'ptend_q0001_31', 'ptend_q0001_32', 'ptend_q0001_33', 'ptend_q0001_34', 'ptend_q0001_35', 'ptend_q0001_36', 'ptend_q0001_37', 'ptend_q0001_38', 'ptend_q0001_39',\n                   'ptend_q0001_40', 'ptend_q0001_41', 'ptend_q0001_42', 'ptend_q0001_43', 'ptend_q0001_44', 'ptend_q0001_45', 'ptend_q0001_46', 'ptend_q0001_47', 'ptend_q0001_48', 'ptend_q0001_49', \n                   'ptend_q0001_50', 'ptend_q0001_51', 'ptend_q0001_52', 'ptend_q0001_53', 'ptend_q0001_54', 'ptend_q0001_55', 'ptend_q0001_56', 'ptend_q0001_57', 'ptend_q0001_58','ptend_q0001_59',\n                  'ptend_q0002_0', 'ptend_q0002_1', 'ptend_q0002_2', 'ptend_q0002_3', 'ptend_q0002_4', 'ptend_q0002_5', 'ptend_q0002_6', 'ptend_q0002_7', 'ptend_q0002_8','ptend_q0002_9', \n                   'ptend_q0002_10', 'ptend_q0002_11', 'ptend_q0002_12', 'ptend_q0002_13', 'ptend_q0002_14', 'ptend_q0002_15', 'ptend_q0002_16', 'ptend_q0002_17', 'ptend_q0002_18', 'ptend_q0002_19', \n                   'ptend_q0002_20','ptend_q0002_21', 'ptend_q0002_22', 'ptend_q0002_23', 'ptend_q0002_24', 'ptend_q0002_25', 'ptend_q0002_26', 'ptend_q0002_27', 'ptend_q0002_28', 'ptend_q0002_29', \n                   'ptend_q0002_30', 'ptend_q0002_31', 'ptend_q0002_32', 'ptend_q0002_33', 'ptend_q0002_34', 'ptend_q0002_35', 'ptend_q0002_36', 'ptend_q0002_37', 'ptend_q0002_38', 'ptend_q0002_39',\n                   'ptend_q0002_40', 'ptend_q0002_41', 'ptend_q0002_42', 'ptend_q0002_43', 'ptend_q0002_44', 'ptend_q0002_45', 'ptend_q0002_46', 'ptend_q0002_47', 'ptend_q0002_48', 'ptend_q0002_49', \n                   'ptend_q0002_50', 'ptend_q0002_51', 'ptend_q0002_52', 'ptend_q0002_53', 'ptend_q0002_54', 'ptend_q0002_55', 'ptend_q0002_56', 'ptend_q0002_57', 'ptend_q0002_58','ptend_q0002_59',\n                  'ptend_q0003_0', 'ptend_q0003_1', 'ptend_q0003_2', 'ptend_q0003_3', 'ptend_q0003_4', 'ptend_q0003_5', 'ptend_q0003_6', 'ptend_q0003_7', 'ptend_q0003_8','ptend_q0003_9', \n                   'ptend_q0003_10', 'ptend_q0003_11', 'ptend_q0003_12', 'ptend_q0003_13', 'ptend_q0003_14', 'ptend_q0003_15', 'ptend_q0003_16', 'ptend_q0003_17', 'ptend_q0003_18', 'ptend_q0003_19', \n                   'ptend_q0003_20','ptend_q0003_21', 'ptend_q0003_22', 'ptend_q0003_23', 'ptend_q0003_24', 'ptend_q0003_25', 'ptend_q0003_26', 'ptend_q0003_27', 'ptend_q0003_28', 'ptend_q0003_29', \n                   'ptend_q0003_30', 'ptend_q0003_31', 'ptend_q0003_32', 'ptend_q0003_33', 'ptend_q0003_34', 'ptend_q0003_35', 'ptend_q0003_36', 'ptend_q0003_37', 'ptend_q0003_38', 'ptend_q0003_39',\n                   'ptend_q0003_40', 'ptend_q0003_41', 'ptend_q0003_42', 'ptend_q0003_43', 'ptend_q0003_44', 'ptend_q0003_45', 'ptend_q0003_46', 'ptend_q0003_47', 'ptend_q0003_48', 'ptend_q0003_49', \n                   'ptend_q0003_50', 'ptend_q0003_51', 'ptend_q0003_52', 'ptend_q0003_53', 'ptend_q0003_54', 'ptend_q0003_55', 'ptend_q0003_56', 'ptend_q0003_57', 'ptend_q0003_58','ptend_q0003_59',\n                  'ptend_u_0', 'ptend_u_1', 'ptend_u_2', 'ptend_u_3', 'ptend_u_4', 'ptend_u_5', 'ptend_u_6', 'ptend_u_7', 'ptend_u_8','ptend_u_9', \n                   'ptend_u_10', 'ptend_u_11', 'ptend_u_12', 'ptend_u_13', 'ptend_u_14', 'ptend_u_15', 'ptend_u_16', 'ptend_u_17', 'ptend_u_18', 'ptend_u_19', \n                   'ptend_u_20','ptend_u_21', 'ptend_u_22', 'ptend_u_23', 'ptend_u_24', 'ptend_u_25', 'ptend_u_26', 'ptend_u_27', 'ptend_u_28', 'ptend_u_29', \n                   'ptend_u_30', 'ptend_u_31', 'ptend_u_32', 'ptend_u_33', 'ptend_u_34', 'ptend_u_35', 'ptend_u_36', 'ptend_u_37', 'ptend_u_38', 'ptend_u_39',\n                   'ptend_u_40', 'ptend_u_41', 'ptend_u_42', 'ptend_u_43', 'ptend_u_44', 'ptend_u_45', 'ptend_u_46', 'ptend_u_47', 'ptend_u_48', 'ptend_u_49', \n                   'ptend_u_50', 'ptend_u_51', 'ptend_u_52', 'ptend_u_53', 'ptend_u_54', 'ptend_u_55', 'ptend_u_56', 'ptend_u_57', 'ptend_u_58','ptend_u_59',\n                  'ptend_v_0', 'ptend_v_1', 'ptend_v_2', 'ptend_v_3', 'ptend_v_4', 'ptend_v_5', 'ptend_v_6', 'ptend_v_7', 'ptend_v_8','ptend_v_9', \n                   'ptend_v_10', 'ptend_v_11', 'ptend_v_12', 'ptend_v_13', 'ptend_v_14', 'ptend_v_15', 'ptend_v_16', 'ptend_v_17', 'ptend_v_18', 'ptend_v_19', \n                   'ptend_v_20','ptend_v_21', 'ptend_v_22', 'ptend_v_23', 'ptend_v_24', 'ptend_v_25', 'ptend_v_26', 'ptend_v_27', 'ptend_v_28', 'ptend_v_29', \n                   'ptend_v_30', 'ptend_v_31', 'ptend_v_32', 'ptend_v_33', 'ptend_v_34', 'ptend_v_35', 'ptend_v_36', 'ptend_v_37', 'ptend_v_38', 'ptend_v_39',\n                   'ptend_v_40', 'ptend_v_41', 'ptend_v_42', 'ptend_v_43', 'ptend_v_44', 'ptend_v_45', 'ptend_v_46', 'ptend_v_47', 'ptend_v_48', 'ptend_v_49', \n                   'ptend_v_50', 'ptend_v_51', 'ptend_v_52', 'ptend_v_53', 'ptend_v_54', 'ptend_v_55', 'ptend_v_56', 'ptend_v_57', 'ptend_v_58','ptend_v_59',\n                  'cam_out_NETSW', 'cam_out_FLWDS', 'cam_out_PRECSC', 'cam_out_PRECC', 'cam_out_SOLS', 'cam_out_SOLL', 'cam_out_SOLSD', 'cam_out_SOLLD']\n\nX_train = train_df[feature_columns][:70000]\ny_train = train_df[target_columns][:70000]\nX_test = train_df[feature_columns][70001:]\ny_test = train_df[target_columns][70001:]\n\n# Normalize the data\nscaler_X = StandardScaler()\nscaler_y = StandardScaler()\n\nX_train_scaled = scaler_X.fit_transform(X_train)\ny_train_scaled = scaler_y.fit_transform(y_train)\nX_test_scaled = scaler_X.transform(X_test)\ny_test_scaled = scaler_y.transform(y_test)","metadata":{"execution":{"iopub.status.busy":"2024-08-24T20:40:53.725953Z","iopub.execute_input":"2024-08-24T20:40:53.726447Z","iopub.status.idle":"2024-08-24T20:40:56.285202Z","shell.execute_reply.started":"2024-08-24T20:40:53.726413Z","shell.execute_reply":"2024-08-24T20:40:56.28374Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Define the model\ndef build_model(input_shape, output_shape):\n    inputs = layers.Input(shape=(input_shape,))\n    x = layers.Dense(512, activation='relu', kernel_regularizer=regularizers.l2(1e-4))(inputs)\n    x = layers.BatchNormalization()(x)\n    x = layers.Dropout(0.3)(x)\n    x = layers.Dense(256, activation='relu', kernel_regularizer=regularizers.l2(1e-4))(x)\n    x = layers.BatchNormalization()(x)\n    x = layers.Dropout(0.3)(x)\n    outputs = layers.Dense(output_shape, activation='linear')(x)\n    \n    model = models.Model(inputs=inputs, outputs=outputs)\n    model.compile(optimizer='adam', loss='mse', metrics=['mse'])\n    return model\n\n# Build the model\nmodel = build_model(X_train_scaled.shape[1], y_train_scaled.shape[1])\n\n# Define callbacks\nearly_stopping = EarlyStopping(monitor='val_loss', patience=10, restore_best_weights=True)\n\n# Custom R2 score callback\nclass R2ScoreCallback(tf.keras.callbacks.Callback):\n    def __init__(self, validation_data):\n        super().__init__()\n        self.validation_data = validation_data\n\n    def on_epoch_end(self, epoch, logs=None):\n        val_pred = self.model.predict(self.validation_data[0])\n        val_true = self.validation_data[1]\n        r2 = r2_score(val_true, val_pred)\n        print(f\" - val_r2: {r2:.4f}\")\n        logs['val_r2'] = r2\n\n# Train the model\nhistory = model.fit(\n    X_train_scaled, y_train_scaled,\n    validation_split=0.2,\n    epochs=100,\n    batch_size=2048,\n    callbacks=[early_stopping, R2ScoreCallback((X_test_scaled, y_test_scaled))],\n    verbose=1\n)\n\n# Save the trained model\nmodel.save('/path/to/best_model_dataframe.h5')","metadata":{"execution":{"iopub.status.busy":"2024-08-24T20:43:19.474001Z","iopub.execute_input":"2024-08-24T20:43:19.474479Z","iopub.status.idle":"2024-08-24T20:48:29.416225Z","shell.execute_reply.started":"2024-08-24T20:43:19.474441Z","shell.execute_reply":"2024-08-24T20:48:29.414584Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plotting learning curves\nplt.figure(figsize=(12, 4))\nplt.subplot(1, 2, 1)\nplt.plot(history.history['loss'], label='Training Loss')\nplt.plot(history.history['val_loss'], label='Validation Loss')\nplt.title('Model Loss')\nplt.xlabel('Epoch')\nplt.ylabel('Loss')\nplt.legend()\n\nplt.subplot(1, 2, 2)\nplt.plot(history.history['val_r2'], label='Validation R2 Score')\nplt.title('Model R2 Score')\nplt.xlabel('Epoch')\nplt.ylabel('R2 Score')\nplt.legend()\n\nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-08-24T21:09:35.73576Z","iopub.execute_input":"2024-08-24T21:09:35.736332Z","iopub.status.idle":"2024-08-24T21:09:36.467634Z","shell.execute_reply.started":"2024-08-24T21:09:35.73628Z","shell.execute_reply":"2024-08-24T21:09:36.466242Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Evaluate the model on the test set\ntest_loss, test_mse = model.evaluate(X_test_scaled, y_test_scaled)\nprint(f\"Test Loss: {test_loss:.4f}\")\nprint(f\"Test MSE: {test_mse:.4f}\")\n\n# Make predictions on the test set\ny_pred_scaled = model.predict(X_test_scaled)\ny_pred = scaler_y.inverse_transform(y_pred_scaled)\n\n# Evaluate the model on the test set\ntest_loss, test_mse = model.evaluate(X_test_scaled, y_test_scaled)\nprint(f\"Test Loss: {test_loss:.4f}\")\nprint(f\"Test MSE: {test_mse:.4f}\")\n\n# Make predictions on the test set\ny_pred_scaled = model.predict(X_test_scaled)\ny_pred = scaler_y.inverse_transform(y_pred_scaled)\n\n# Calculate R2 score for the test set\ntest_r2 = r2_score(y_test, y_pred)\nprint(f\"Test R2 Score: {test_r2:.4f}\")","metadata":{"execution":{"iopub.status.busy":"2024-08-24T21:21:49.522955Z","iopub.execute_input":"2024-08-24T21:21:49.523532Z","iopub.status.idle":"2024-08-24T21:22:05.436921Z","shell.execute_reply.started":"2024-08-24T21:21:49.523492Z","shell.execute_reply":"2024-08-24T21:22:05.435701Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### XGBoost","metadata":{}},{"cell_type":"code","source":"import torch\nfrom torch.utils.data import DataLoader, TensorDataset\nfrom statistics import mean\nfrom xgboost import XGBRegressor\nfrom sklearn.metrics import r2_score\nfrom sklearn.preprocessing import StandardScaler\nfrom sklearn.metrics import mean_squared_error\n\ntrain_df = p1.scan_csv(f'/kaggle/input/leap-atmospheric-physics-ai-climsim/train.csv')\nfetch_size =1000\ninput_columns = train_df.columns[1:557]\noutput_columns = train_df.columns[557:]\nX = train_df.select(p1.col(input_columns)).fetch(fetch_size)\nY = train_df.select(p1.col(output_columns)).fetch(fetch_size)\nbig = train_df.select(p1.col(\"*\")).fetch(fetch_size)","metadata":{"execution":{"iopub.status.busy":"2024-08-24T21:51:58.958858Z","iopub.execute_input":"2024-08-24T21:51:58.959317Z","iopub.status.idle":"2024-08-24T21:51:59.392805Z","shell.execute_reply.started":"2024-08-24T21:51:58.959282Z","shell.execute_reply":"2024-08-24T21:51:59.390115Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def adicionar_agua_v_erro_info(big: p1.DataFrame, margen = 0.00005) ->p1.DataFrame:\n    gravidade = 9.80665\n    ptend_vapor = big.select(p1.col(\"^ptend_q0001.*$\")).to_numpy()\n    ptend_liquido = big.select(p1.col(\"^ptend_q0002.*$\")).to_numpy()\n    ptend_gelo = big.select(p1.col(\"^ptend_q0003.*$\")).to_numpy()\n    \n    state_vapor = big.select(p1.col(\"^state_q0001.*$\")).to_numpy()\n    state_liquido = big.select(p1.col(\"^state_q0002.*$\")).to_numpy()\n    state_gelo = big.select(p1.col(\"^state_q0003.*$\")).to_numpy()\n    \n    precipitacao_agua = big.select(p1.col(\"^cam_out_PRECC.*$\")).to_numpy().flatten()\n    precipitacao_neve = big.select(p1.col(\"^cam_out_PRECSC.*$\")).to_numpy().flatten()\n    \n    \n    pressao = big.select(p1.col(\"^state_ps.*$\")).to_numpy().flatten()\n    volume_inicial = np.sum(state_vapor + state_liquido + state_gelo,axis=1)*(pressao/gravidade)\n    volume_predito = np.sum(ptend_vapor + ptend_liquido + ptend_gelo,axis=1)*(pressao/gravidade) + precipitacao_agua + precipitacao_neve\n    \n    \n    erro_volumetrico = abs((volume_predito - volume_inicial)/volume_inicial +1)\n    flag_erro_volumetrico = (erro_volumetrico >margen).astype(np.int8)\n    big = big.with_columns([\n        p1.Series(\"erro_volumetrico\",erro_volumetrico),\n        p1.Series(\"flag_erro_volumetrico\",flag_erro_volumetrico)\n    ])\n    return big","metadata":{"execution":{"iopub.status.busy":"2024-08-24T21:54:27.449056Z","iopub.execute_input":"2024-08-24T21:54:27.449596Z","iopub.status.idle":"2024-08-24T21:54:27.464176Z","shell.execute_reply.started":"2024-08-24T21:54:27.449558Z","shell.execute_reply":"2024-08-24T21:54:27.462784Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def custom_log_transform(column):\n    \"\"\"Applies a log1p transformation to a given column with an adjustment for negative values.\"\"\"\n    min_value = column.min()\n    transformed_column = np.log1p(column - min_value + 1)  # log1p is log(1 + x) which is stable for small x\n    return transformed_column\n\ndef log_transformation(data):\n    \"\"\"Applies custom log transformation to each column in the data.\"\"\"\n    data_log_transformed = np.array(data, dtype=\"float64\")\n    for i in range(data.shape[1]):\n        data_log_transformed[:, i] = custom_log_transform(data[:, i])\n    return data_log_transformed\n\ndef preprocess_data(X, Y, transform=True, log=False):\n    \"\"\"Preprocesses the data by scaling and optionally applying log transformation.\"\"\"\n    X_np = X.to_numpy()\n    Y_np = Y.to_numpy()\n\n    if transform:\n        # Standardize features\n        X_scaler = StandardScaler().fit(X_np)\n        Y_scaler = StandardScaler().fit(Y_np)\n\n        X_transformed = X_scaler.transform(X_np)\n        Y_transformed = Y_scaler.transform(Y_np)\n\n        if log:\n            X_transformed = log_transformation(X_transformed)\n            Y_transformed = log_transformation(Y_transformed)\n        \n        # Clean up\n        del X, Y, X_np, Y_np\n\n    else:\n        # Split data into training and validation sets without transformation\n        #X_train, X_val, Y_train, Y_val = train_test_split(X_np, Y_np, test_size=0.2, random_state=42)\n        \n        # Clean up\n        del X, Y, X_np, Y_np\n\n    return X_transformed, Y_transformed\n\n# Example usage\n# Assuming X and Y are defined and are Polars DataFrames\nX_process, Y_process = preprocess_data(X, Y, transform=True, log=False)\nX_train, X_val, Y_train, Y_val = train_test_split(X_process, Y_process, test_size=0.2, random_state=42)","metadata":{"execution":{"iopub.status.busy":"2024-08-24T21:55:01.284957Z","iopub.execute_input":"2024-08-24T21:55:01.285379Z","iopub.status.idle":"2024-08-24T21:55:01.354729Z","shell.execute_reply.started":"2024-08-24T21:55:01.285348Z","shell.execute_reply":"2024-08-24T21:55:01.353523Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"big=adicionar_agua_v_erro_info(big)","metadata":{"execution":{"iopub.status.busy":"2024-08-24T21:55:27.646215Z","iopub.execute_input":"2024-08-24T21:55:27.646678Z","iopub.status.idle":"2024-08-24T21:55:27.679183Z","shell.execute_reply.started":"2024-08-24T21:55:27.646643Z","shell.execute_reply":"2024-08-24T21:55:27.678068Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.boxplot(big[\"erro_volumetrico\"])\n","metadata":{"execution":{"iopub.status.busy":"2024-08-24T21:56:17.328268Z","iopub.execute_input":"2024-08-24T21:56:17.328802Z","iopub.status.idle":"2024-08-24T21:56:17.550711Z","shell.execute_reply.started":"2024-08-24T21:56:17.328744Z","shell.execute_reply":"2024-08-24T21:56:17.549352Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sum(big[\"flag_erro_volumetrico\"]>0)","metadata":{"execution":{"iopub.status.busy":"2024-08-24T21:56:23.858218Z","iopub.execute_input":"2024-08-24T21:56:23.858663Z","iopub.status.idle":"2024-08-24T21:56:23.881366Z","shell.execute_reply.started":"2024-08-24T21:56:23.858614Z","shell.execute_reply":"2024-08-24T21:56:23.88007Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.boxplot(big[\"erro_volumetrico\"])","metadata":{"execution":{"iopub.status.busy":"2024-08-24T21:56:41.034824Z","iopub.execute_input":"2024-08-24T21:56:41.03526Z","iopub.status.idle":"2024-08-24T21:56:41.299412Z","shell.execute_reply.started":"2024-08-24T21:56:41.035223Z","shell.execute_reply":"2024-08-24T21:56:41.298179Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Initialize model\nmodel = XGBRegressor(\n    n_estimators= 1150,\n    max_depth=56,\n    learning_rate=0.003,\n    colsample_bytree=0.9,\n    subsample=0.8,\n    min_child_weight=1,\n    random_state=42,\n    objective='reg:squarederror',\n    n_jobs=-1,  # Use all available cores\n    tree_method='hist',\n    device ='cuda',\n)\n\nr2_vector = []\nmse_vector = []\n\ny_pred_df = pd.DataFrame()\nout_i =360\nout_f=361\nfor i in range(out_i, out_f):\n    \n    model.fit(X_train, Y_train[:, i]) \n    y_pred = model.predict(X_val)\n    y_pred_df[f'y_pred_{i}'] = y_pred\n    \n    # Calculate r\n    r2 = r2_score(Y_val[:, i], y_pred)\n    mse = mean_squared_error(Y_val[:, i], y_pred)\n    print(f\"{r2},{mse}\")\n    # Append scores to the lists\n    r2_vector.append(r2)\n    mse_vector.append(mse)\n\n# Optionally convert lists to DataFrame for r2 and mse scores\nr2_df = pd.DataFrame({'r2_score': r2_vector})\nmse_df = pd.DataFrame({'mse_score': mse_vector})\n\n# Now y_pred_df contains all the predictions for each target variable","metadata":{"execution":{"iopub.status.busy":"2024-08-24T21:57:14.245306Z","iopub.execute_input":"2024-08-24T21:57:14.245821Z","iopub.status.idle":"2024-08-24T22:00:07.195732Z","shell.execute_reply.started":"2024-08-24T21:57:14.245781Z","shell.execute_reply":"2024-08-24T22:00:07.194182Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"columns = output_columns[out_i:out_f]\nplt.figure(figsize=(20, 6))\nplt.plot(columns, mse_vector, label='MSE', color='blue', marker='o')\nplt.plot(columns, r2_vector, label='R²', color='red', marker='x')\n\nplt.xlabel('Column Number')\nplt.ylabel('Score')\nplt.title('MSE and R² for each column')\nplt.legend()\nplt.grid(True)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-08-25T00:11:31.944727Z","iopub.execute_input":"2024-08-25T00:11:31.945377Z","iopub.status.idle":"2024-08-25T00:11:32.348206Z","shell.execute_reply.started":"2024-08-25T00:11:31.945222Z","shell.execute_reply":"2024-08-25T00:11:32.346763Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Multi-Layer Perceptron","metadata":{}},{"cell_type":"code","source":"leap_raw_df = p1.scan_csv(f'/kaggle/input/leap-atmospheric-physics-ai-climsim/train.csv')\nfetch_size =100000\ninput_columns = leap_raw_df.columns[1:557]\noutput_columns = leap_raw_df.columns[557:]\n\nX = leap_raw_df.select(p1.col(input_columns)).fetch(fetch_size)\nY = leap_raw_df.select(p1.col(output_columns)).fetch(fetch_size)","metadata":{"execution":{"iopub.status.busy":"2024-08-25T00:30:05.808782Z","iopub.execute_input":"2024-08-25T00:30:05.812074Z","iopub.status.idle":"2024-08-25T00:30:15.469108Z","shell.execute_reply.started":"2024-08-25T00:30:05.812017Z","shell.execute_reply":"2024-08-25T00:30:15.46775Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.preprocessing import StandardScaler\nfrom sklearn.preprocessing import FunctionTransformer\nconstant_epson = 1\nconstant_epson_r = 1\ndef log_transform(x):\n    return (np.log1p(x+constant_epson) +constant_epson_r)\n\n# Convert Polars DataFrame to NumPy array\nX_np = X.to_numpy()\nY_np = Y.to_numpy()\n\ntransformer_X = StandardScaler().fit(X_np)\ntransformer_Y = StandardScaler().fit(Y_np)\n\nX_transformed = transformer_X.transform(X_np)\nY_transformed = transformer_Y.transform(Y_np)\n\n\n\n# Log\n#transformer =FunctionTransformer(log_transform)\n#X_transformado=transformer.transform(X_np)\n#Y_transformado=transformer.transform(Y_np)","metadata":{"execution":{"iopub.status.busy":"2024-08-25T00:30:55.263776Z","iopub.execute_input":"2024-08-25T00:30:55.264217Z","iopub.status.idle":"2024-08-25T00:30:57.639684Z","shell.execute_reply.started":"2024-08-25T00:30:55.264185Z","shell.execute_reply":"2024-08-25T00:30:57.6384Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(X_transformed.dtype)\nprint(Y_transformed.dtype)","metadata":{"execution":{"iopub.status.busy":"2024-08-25T00:31:06.096522Z","iopub.execute_input":"2024-08-25T00:31:06.097013Z","iopub.status.idle":"2024-08-25T00:31:06.106184Z","shell.execute_reply.started":"2024-08-25T00:31:06.09698Z","shell.execute_reply":"2024-08-25T00:31:06.104491Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import torch\nimport torch.nn as nn\nimport torch.optim as optim\n\nX_tensor = torch.tensor(X_transformed, dtype=torch.float64)\nY_tensor = torch.tensor(Y_transformed, dtype=torch.float64)","metadata":{"execution":{"iopub.status.busy":"2024-08-25T00:31:10.290878Z","iopub.execute_input":"2024-08-25T00:31:10.291358Z","iopub.status.idle":"2024-08-25T00:31:11.061572Z","shell.execute_reply.started":"2024-08-25T00:31:10.29132Z","shell.execute_reply":"2024-08-25T00:31:11.060274Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from torch.utils.data import DataLoader, TensorDataset\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import r2_score","metadata":{"execution":{"iopub.status.busy":"2024-08-25T00:31:20.619213Z","iopub.execute_input":"2024-08-25T00:31:20.62004Z","iopub.status.idle":"2024-08-25T00:31:20.627386Z","shell.execute_reply.started":"2024-08-25T00:31:20.62Z","shell.execute_reply":"2024-08-25T00:31:20.62587Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train, X_val, Y_train,Y_val = train_test_split(X_tensor, Y_tensor, test_size=0.2, random_state=42)\ntrain_dataset = TensorDataset(X_train, Y_train)\nval_dataset = TensorDataset(X_val, Y_val)\n\ndef create_dataloader(dataset, batch_size):\n    return DataLoader(dataset, batch_size=batch_size, shuffle=True)","metadata":{"execution":{"iopub.status.busy":"2024-08-25T00:31:30.110652Z","iopub.execute_input":"2024-08-25T00:31:30.111564Z","iopub.status.idle":"2024-08-25T00:31:32.863604Z","shell.execute_reply.started":"2024-08-25T00:31:30.111525Z","shell.execute_reply":"2024-08-25T00:31:32.862315Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import optuna\nimport torch\nimport torch.nn as nn\nimport torch.optim as optim\nfrom sklearn.metrics import mean_squared_error, r2_score\nfrom sklearn.preprocessing import StandardScaler\nimport polars as pl\n\ndef objective(trial):\n    # Define the actual input size and output size based on your data\n    input_size = 556  # Number of features in X\n    output_size = 368  # Number of targets in Y\n\n    # Tune the number of hidden layers and units per layer\n    hidden_layers = trial.suggest_int('hidden_layers', 3, 8)\n    hidden_units = trial.suggest_int('hidden_units', 512, 1024)\n    batch_size = trial.suggest_int('batch_size', 1536, 4608)\n    learning_rate = trial.suggest_loguniform('learning_rate', 2.5e-4, 2.5e-2)\n\n    X_train_tensor = torch.tensor(X_train, dtype=torch.float64)\n    Y_train_tensor = torch.tensor(Y_train, dtype=torch.float64)\n    X_val_tensor = torch.tensor(X_val, dtype=torch.float64)\n    Y_val_tensor = torch.tensor(Y_val, dtype=torch.float64)\n\n    # Define the architecture of the MLP\n    class MLPRegressor(nn.Module):\n        def __init__(self, input_size, output_size, hidden_layers, hidden_units):\n            super(MLPRegressor, self).__init__()\n            layers = []\n            layers.append(nn.Linear(input_size, hidden_units, dtype=torch.float64))\n            layers.append(nn.LeakyReLU(0.15))\n            for _ in range(hidden_layers - 1):\n                layers.append(nn.Linear(hidden_units, hidden_units, dtype=torch.float64))\n                layers.append(nn.LeakyReLU(0.15))\n            layers.append(nn.Linear(hidden_units, output_size, dtype=torch.float64))\n            self.model = nn.Sequential(*layers)\n        \n        def forward(self, x):\n            return self.model(x)\n\n    model = MLPRegressor(input_size, output_size, hidden_layers, hidden_units).double()\n    criterion = nn.MSELoss()\n    optimizer = optim.Adam(model.parameters(), lr=learning_rate)\n\n    def create_dataloader(X, Y, batch_size):\n        dataset = torch.utils.data.TensorDataset(X, Y)\n        return torch.utils.data.DataLoader(dataset, batch_size=batch_size, shuffle=True)\n\n    train_loader = create_dataloader(X_train_tensor, Y_train_tensor, batch_size)\n    val_loader = create_dataloader(X_val_tensor, Y_val_tensor, batch_size)\n\n    # Train the model\n    def train(model, criterion, optimizer, train_loader, val_loader, epochs=4):\n        model.train()\n        for epoch in range(epochs):\n            for inputs, targets in train_loader:\n                optimizer.zero_grad()\n                outputs = model(inputs)\n                loss = criterion(outputs, targets)\n                loss.backward()\n                optimizer.step()\n        \n        # Evaluate on the validation set\n        model.eval()\n        val_preds = []\n        val_targets = []\n        with torch.no_grad():\n            for val_inputs, val_targets_batch in val_loader:\n                val_outputs = model(val_inputs)\n                val_preds.append(val_outputs)\n                val_targets.append(val_targets_batch)\n        \n        val_preds = torch.cat(val_preds).cpu().numpy()\n        val_targets = torch.cat(val_targets).cpu().numpy()\n        val_preds[:, 61] = 0\n        mse = mean_squared_error(val_targets, val_preds)\n        \n        r2_values = []\n        for i in range(output_size):\n            r2 = r2_score(val_targets[:, i], val_preds[:, i])\n            r2_values.append(r2)\n        print(f\"R² for column 62: {r2_values[61]}\")\n\n        # Calculate the average R² excluding the value at index 61\n        r2_values_excluding_61 = [r2 for idx, r2 in enumerate(r2_values) if idx != 61]\n        avg_r2_excluding_61 = sum(r2_values_excluding_61) / len(r2_values_excluding_61)\n        print(\"Average R² excluding column 61:\", avg_r2_excluding_61)\n        return avg_r2_excluding_61\n\n    avg_r2_excluding_61 = train(model, criterion, optimizer, train_loader, val_loader)\n    return avg_r2_excluding_61\n\nstudy = optuna.create_study(direction='maximize')\nstudy.optimize(objective, n_trials=100)\n\nprint(\"Best hyperparameters: \", study.best_params)\nprint(\"Best value (MSE): \", study.best_value)","metadata":{"execution":{"iopub.status.busy":"2024-08-25T00:34:03.841461Z","iopub.execute_input":"2024-08-25T00:34:03.842448Z","iopub.status.idle":"2024-08-25T03:45:56.274151Z","shell.execute_reply.started":"2024-08-25T00:34:03.842405Z","shell.execute_reply":"2024-08-25T03:45:56.270152Z"},"scrolled":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import optuna\nimport warnings\nwarnings.filterwarnings(\"ignore\")","metadata":{"execution":{"iopub.status.busy":"2024-08-25T03:46:31.962206Z","iopub.execute_input":"2024-08-25T03:46:31.962754Z","iopub.status.idle":"2024-08-25T03:46:31.970184Z","shell.execute_reply.started":"2024-08-25T03:46:31.962714Z","shell.execute_reply":"2024-08-25T03:46:31.968527Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import optuna.visualization as vis\nimport matplotlib.pyplot as plt","metadata":{"execution":{"iopub.status.busy":"2024-08-25T03:46:34.655276Z","iopub.execute_input":"2024-08-25T03:46:34.655787Z","iopub.status.idle":"2024-08-25T03:46:35.315955Z","shell.execute_reply.started":"2024-08-25T03:46:34.655731Z","shell.execute_reply":"2024-08-25T03:46:35.314611Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"importance_plot = vis.plot_param_importances(study)\nimportance_plot.show()","metadata":{"execution":{"iopub.status.busy":"2024-08-25T03:46:45.781955Z","iopub.execute_input":"2024-08-25T03:46:45.782409Z","iopub.status.idle":"2024-08-25T03:46:48.186319Z","shell.execute_reply.started":"2024-08-25T03:46:45.782373Z","shell.execute_reply":"2024-08-25T03:46:48.185015Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"opt_history_plot = vis.plot_optimization_history(study)\nopt_history_plot.show()","metadata":{"execution":{"iopub.status.busy":"2024-08-25T03:46:58.655313Z","iopub.execute_input":"2024-08-25T03:46:58.655795Z","iopub.status.idle":"2024-08-25T03:46:58.727912Z","shell.execute_reply.started":"2024-08-25T03:46:58.655755Z","shell.execute_reply":"2024-08-25T03:46:58.726578Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"parallel_plot = vis.plot_parallel_coordinate(study)\nparallel_plot.show()","metadata":{"execution":{"iopub.status.busy":"2024-08-25T03:47:01.067901Z","iopub.execute_input":"2024-08-25T03:47:01.0689Z","iopub.status.idle":"2024-08-25T03:47:01.170993Z","shell.execute_reply.started":"2024-08-25T03:47:01.068856Z","shell.execute_reply":"2024-08-25T03:47:01.169706Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"contour_plot = vis.plot_contour(study)\ncontour_plot.show()","metadata":{"execution":{"iopub.status.busy":"2024-08-25T03:47:03.314435Z","iopub.execute_input":"2024-08-25T03:47:03.315335Z","iopub.status.idle":"2024-08-25T03:47:04.60961Z","shell.execute_reply.started":"2024-08-25T03:47:03.315296Z","shell.execute_reply":"2024-08-25T03:47:04.608056Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"conda install git","metadata":{"execution":{"iopub.status.busy":"2024-08-25T04:12:39.376526Z","iopub.execute_input":"2024-08-25T04:12:39.377172Z","iopub.status.idle":"2024-08-25T04:14:04.054137Z","shell.execute_reply.started":"2024-08-25T04:12:39.377124Z","shell.execute_reply":"2024-08-25T04:14:04.052528Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"conda install boto3","metadata":{"execution":{"iopub.status.busy":"2024-08-25T05:22:42.543746Z","iopub.execute_input":"2024-08-25T05:22:42.544479Z","iopub.status.idle":"2024-08-25T05:24:00.504439Z","shell.execute_reply.started":"2024-08-25T05:22:42.544369Z","shell.execute_reply":"2024-08-25T05:24:00.50257Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import boto3\nimport os\nimport os\n\nos.environ['AWS_ACCESS_KEY_ID'] = 'AKIAS6J7P6YLAF2TYLCH'\nos.environ['AWS_SECRET_ACCESS_KEY'] = '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'\nos.environ['AWS_DEFAULT_REGION'] = 'us-east-2'\n\ndef upload_to_s3(file_path, bucket_name, s3_key):\n    # Create an S3 client\n    s3 = boto3.client('s3')\n\n    try:\n        # Upload the file\n        s3.upload_file(file_path, bucket_name, s3_key)\n        print(f\"Successfully uploaded {file_path} to {bucket_name}/{s3_key}\")\n    except Exception as e:\n        print(f\"Error uploading file: {str(e)}\")\n\n# Example usage\nlocal_file_path = '../input/leap-atmospheric-physics-ai-climsim/train.csv'\ns3_bucket_name = 'capstone-project-ucsd-mle-bootcamp'\ns3_file_key = 'data.csv'\n\nupload_to_s3(local_file_path, s3_bucket_name, s3_file_key)","metadata":{"execution":{"iopub.status.busy":"2024-08-25T05:51:45.609572Z","iopub.execute_input":"2024-08-25T05:51:45.609944Z","iopub.status.idle":"2024-08-25T06:20:16.242971Z","shell.execute_reply.started":"2024-08-25T05:51:45.609918Z","shell.execute_reply":"2024-08-25T06:20:16.239484Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}