{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":56537,"databundleVersionId":8015876,"sourceType":"competition"},{"sourceId":8174297,"sourceType":"datasetVersion","datasetId":4838367}],"dockerImageVersionId":30698,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2024-04-20T10:45:43.71577Z","iopub.execute_input":"2024-04-20T10:45:43.716428Z","iopub.status.idle":"2024-04-20T10:45:43.730005Z","shell.execute_reply.started":"2024-04-20T10:45:43.716361Z","shell.execute_reply":"2024-04-20T10:45:43.728355Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<span style=\"display:block;color:white;font-size:18px;background-color:#5499C7;font-weight:bold;\">  1 | Introduction: Problem  Definition </span>","metadata":{}},{"cell_type":"markdown","source":"Climate models are essential to understanding Earth’s climate system. Because of the complexity of Earth’s climate, these models rely on parameterizations to approximate the effects of physical processes that occur at scales smaller than the size of their grid cells. These approximations are imperfect, however, and their imperfections are a leading source of uncertainty in expected warming, changing precipitation patterns, and the frequency and severity of extreme events. The Multi-scale Modeling Framework (MMF) approach, by contrast, more explicitly represents these subgrid processes, but at a cost too high to be used for operational climate prediction.\n\nYour task is to develop ML models that emulate subgrid atmospheric processes–such as storms, clouds, turbulence, rainfall, and radiation–within E3SM-MMF, a multi-scale climate model backed by the U.S. Department of Energy. Because ML emulators are significantly cheaper to inference than MMF, progress on this front can help scientists realize a future in which high-resolution and physically credible long-term climate projections are broadly accessible, bringing greater clarity to the hazards associated with climate change and empowering policymakers with the knowledge necessary to mitigate them.\n\nThis competition accompanies an upcoming 2024 ICML Machine Learning for Earth System Modeling (ML4ESM) Workshop and is based on the ClimSim paper and dataset which won the Outstanding Datasets and Benchmarks Paper award at NeurIPS 2023. Winning submissions will be highlighted at the upcoming ML4ESM ICML workshop, and participants in this Kaggle competition are also encouraged to submit workshop papers.\n\nTask is to develop a machine learning models that accurately emulate subgrid-scale atmospheric physics in an operational climate model—an important step in improving climate projections and reducing uncertainty surrounding future climate trends.","metadata":{}},{"cell_type":"code","source":"!pip install dask","metadata":{"execution":{"iopub.status.busy":"2024-04-20T10:45:47.404332Z","iopub.execute_input":"2024-04-20T10:45:47.404986Z","iopub.status.idle":"2024-04-20T10:46:04.263904Z","shell.execute_reply.started":"2024-04-20T10:45:47.404933Z","shell.execute_reply":"2024-04-20T10:46:04.262144Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<span style=\"display:block;color:white;font-size:18px;background-color:#5499C7;font-weight:bold;\">  2 | Load Libraries </span>","metadata":{}},{"cell_type":"code","source":"import dask.dataframe as dataf\nimport seaborn as sb\nsb.color_palette(\"Set2\")\nsb.set_style(\"whitegrid\")\nimport matplotlib.pyplot as plt\nfrom colorama import Fore","metadata":{"execution":{"iopub.status.busy":"2024-04-20T11:34:06.139217Z","iopub.execute_input":"2024-04-20T11:34:06.13982Z","iopub.status.idle":"2024-04-20T11:34:06.149213Z","shell.execute_reply.started":"2024-04-20T11:34:06.139776Z","shell.execute_reply":"2024-04-20T11:34:06.147718Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font color='blue'> Take about 190 MB dataset for initial calculation </font>","metadata":{}},{"cell_type":"markdown","source":"Using dask to read the 187 GB dataset. Dask is performing better than Modin.","metadata":{}},{"cell_type":"code","source":"%%time\nPATH = \"/kaggle/input/leap-atmospheric-physics-ai-climsim/train.csv\"\ndf = dataf.read_csv(PATH,blocksize=\"50MB\")\n\n    ","metadata":{"execution":{"iopub.status.busy":"2024-04-20T10:46:17.32389Z","iopub.execute_input":"2024-04-20T10:46:17.324844Z","iopub.status.idle":"2024-04-20T10:46:17.334367Z","shell.execute_reply.started":"2024-04-20T10:46:17.324788Z","shell.execute_reply":"2024-04-20T10:46:17.332941Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.npartitions","metadata":{"execution":{"iopub.status.busy":"2024-04-20T10:46:26.708285Z","iopub.execute_input":"2024-04-20T10:46:26.708814Z","iopub.status.idle":"2024-04-20T10:46:26.718841Z","shell.execute_reply.started":"2024-04-20T10:46:26.708773Z","shell.execute_reply":"2024-04-20T10:46:26.71707Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df1 = df.partitions[0]\ndf2 = df.partitions[1]\ndf3 = df.partitions[2]\ndf4 = df.partitions[3]","metadata":{"execution":{"iopub.status.busy":"2024-04-20T10:47:36.774143Z","iopub.execute_input":"2024-04-20T10:47:36.774648Z","iopub.status.idle":"2024-04-20T10:47:36.783791Z","shell.execute_reply.started":"2024-04-20T10:47:36.774613Z","shell.execute_reply":"2024-04-20T10:47:36.782405Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df1 = df1.compute()\ndf2 = df2.compute()\ndf3 = df3.compute()\ndf4 = df4.compute()","metadata":{"execution":{"iopub.status.busy":"2024-04-20T10:48:28.54391Z","iopub.execute_input":"2024-04-20T10:48:28.544441Z","iopub.status.idle":"2024-04-20T10:48:38.238574Z","shell.execute_reply.started":"2024-04-20T10:48:28.544398Z","shell.execute_reply":"2024-04-20T10:48:38.237388Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df1.to_csv(f\"/kaggle/working/df_1.csv\",sep=',',index=False)\ndf2.to_csv(f\"/kaggle/working/df_2.csv\",sep=',',index=False)\ndf3.to_csv(f\"/kaggle/working/df_3.csv\",sep=',',index=False)\ndf4.to_csv(f\"/kaggle/working/df_4.csv\",sep=',',index=False)","metadata":{"execution":{"iopub.status.busy":"2024-04-20T10:49:18.044466Z","iopub.execute_input":"2024-04-20T10:49:18.04505Z","iopub.status.idle":"2024-04-20T10:49:43.69672Z","shell.execute_reply.started":"2024-04-20T10:49:18.045013Z","shell.execute_reply":"2024-04-20T10:49:43.695246Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<font color='blue'> Above 4 datasets will be used for initial calculation.</font>","metadata":{}},{"cell_type":"code","source":"train_df1 = pd.read_csv(\"/kaggle/input/smalldat/data/df_1.csv\")\ntrain_df1.isnull().sum().sum()","metadata":{"execution":{"iopub.status.busy":"2024-04-20T11:03:50.104032Z","iopub.execute_input":"2024-04-20T11:03:50.104552Z","iopub.status.idle":"2024-04-20T11:03:51.045168Z","shell.execute_reply.started":"2024-04-20T11:03:50.104518Z","shell.execute_reply":"2024-04-20T11:03:51.043887Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df2 = pd.read_csv(\"/kaggle/input/smalldat/data/df_2.csv\")\ntrain_df2.isnull().sum().sum()","metadata":{"execution":{"iopub.status.busy":"2024-04-20T11:04:16.111238Z","iopub.execute_input":"2024-04-20T11:04:16.111755Z","iopub.status.idle":"2024-04-20T11:04:17.566036Z","shell.execute_reply.started":"2024-04-20T11:04:16.111716Z","shell.execute_reply":"2024-04-20T11:04:17.564618Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df3 = pd.read_csv(\"/kaggle/input/smalldat/data/df_3.csv\")\ntrain_df3.isnull().sum().sum()","metadata":{"execution":{"iopub.status.busy":"2024-04-20T11:04:45.800722Z","iopub.execute_input":"2024-04-20T11:04:45.801264Z","iopub.status.idle":"2024-04-20T11:04:47.232285Z","shell.execute_reply.started":"2024-04-20T11:04:45.80123Z","shell.execute_reply":"2024-04-20T11:04:47.230747Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df4 = pd.read_csv(\"/kaggle/input/smalldat/data/df_4.csv\")\ntrain_df4.isnull().sum().sum()","metadata":{"execution":{"iopub.status.busy":"2024-04-20T11:05:03.949209Z","iopub.execute_input":"2024-04-20T11:05:03.949889Z","iopub.status.idle":"2024-04-20T11:05:05.438167Z","shell.execute_reply.started":"2024-04-20T11:05:03.949832Z","shell.execute_reply":"2024-04-20T11:05:05.43677Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\ntest_df = pd.read_csv(\"/kaggle/input/leap-atmospheric-physics-ai-climsim/sample_submission.csv\")\ntarget_cols = list(test_df.columns)","metadata":{"execution":{"iopub.status.busy":"2024-04-20T11:07:11.904855Z","iopub.execute_input":"2024-04-20T11:07:11.905949Z","iopub.status.idle":"2024-04-20T11:09:07.194423Z","shell.execute_reply.started":"2024-04-20T11:07:11.905882Z","shell.execute_reply":"2024-04-20T11:09:07.193025Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df.head()","metadata":{"execution":{"iopub.status.busy":"2024-04-20T11:10:00.8913Z","iopub.execute_input":"2024-04-20T11:10:00.892807Z","iopub.status.idle":"2024-04-20T11:10:00.926157Z","shell.execute_reply.started":"2024-04-20T11:10:00.892749Z","shell.execute_reply":"2024-04-20T11:10:00.924708Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"target_cols.remove('sample_id')\nlen(target_cols)","metadata":{"execution":{"iopub.status.busy":"2024-04-20T11:10:53.44965Z","iopub.execute_input":"2024-04-20T11:10:53.450175Z","iopub.status.idle":"2024-04-20T11:10:53.461289Z","shell.execute_reply.started":"2024-04-20T11:10:53.450139Z","shell.execute_reply":"2024-04-20T11:10:53.459733Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df1.drop('sample_id',axis=1,inplace=True)\ntrain_df2.drop('sample_id',axis=1,inplace=True)\ntrain_df3.drop('sample_id',axis=1,inplace=True)\ntrain_df4.drop('sample_id',axis=1,inplace=True)","metadata":{"execution":{"iopub.status.busy":"2024-04-20T11:12:18.865767Z","iopub.execute_input":"2024-04-20T11:12:18.866258Z","iopub.status.idle":"2024-04-20T11:12:18.923104Z","shell.execute_reply.started":"2024-04-20T11:12:18.866222Z","shell.execute_reply":"2024-04-20T11:12:18.921552Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"features = list(set(train_df1.columns)-set(target_cols))\nlen(features)","metadata":{"execution":{"iopub.status.busy":"2024-04-20T11:12:23.285866Z","iopub.execute_input":"2024-04-20T11:12:23.286369Z","iopub.status.idle":"2024-04-20T11:12:23.296165Z","shell.execute_reply.started":"2024-04-20T11:12:23.286338Z","shell.execute_reply":"2024-04-20T11:12:23.294642Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def box_plots(ind1:int,ind2:int):\n     plt.figure(figsize=(15,5)) \n     sb.boxplot(data=train_df1.iloc[:,ind1:ind2], palette=\"deep\")\n     sb.despine(left=True)\n     plt.title(f\"col_{ind1} to col_{ind2}\")\n     plt.xticks(rotation=90)   \n     plt.show()","metadata":{"execution":{"iopub.status.busy":"2024-04-20T11:41:18.875889Z","iopub.execute_input":"2024-04-20T11:41:18.876441Z","iopub.status.idle":"2024-04-20T11:41:18.886393Z","shell.execute_reply.started":"2024-04-20T11:41:18.876404Z","shell.execute_reply":"2024-04-20T11:41:18.884856Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"box_plots(0,60)","metadata":{"execution":{"iopub.status.busy":"2024-04-20T11:41:22.647286Z","iopub.execute_input":"2024-04-20T11:41:22.647845Z","iopub.status.idle":"2024-04-20T11:41:27.561391Z","shell.execute_reply.started":"2024-04-20T11:41:22.647805Z","shell.execute_reply":"2024-04-20T11:41:27.559969Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"box_plots(61,121)","metadata":{"execution":{"iopub.status.busy":"2024-04-20T11:41:48.08697Z","iopub.execute_input":"2024-04-20T11:41:48.087486Z","iopub.status.idle":"2024-04-20T11:41:53.070729Z","shell.execute_reply.started":"2024-04-20T11:41:48.087449Z","shell.execute_reply":"2024-04-20T11:41:53.069425Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"features += ['ptend_t_0']","metadata":{"execution":{"iopub.status.busy":"2024-04-20T11:55:47.360754Z","iopub.execute_input":"2024-04-20T11:55:47.361323Z","iopub.status.idle":"2024-04-20T11:55:47.368164Z","shell.execute_reply.started":"2024-04-20T11:55:47.361289Z","shell.execute_reply":"2024-04-20T11:55:47.366612Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"correlations = train_df1[features].corr(method='pearson')\n","metadata":{"execution":{"iopub.status.busy":"2024-04-20T11:56:18.247929Z","iopub.execute_input":"2024-04-20T11:56:18.248441Z","iopub.status.idle":"2024-04-20T11:56:20.592648Z","shell.execute_reply.started":"2024-04-20T11:56:18.248407Z","shell.execute_reply":"2024-04-20T11:56:20.591589Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ptend_t_0_corr = correlations.loc[features,'ptend_t_0']","metadata":{"execution":{"iopub.status.busy":"2024-04-20T12:03:26.709568Z","iopub.execute_input":"2024-04-20T12:03:26.710524Z","iopub.status.idle":"2024-04-20T12:03:26.719098Z","shell.execute_reply.started":"2024-04-20T12:03:26.710474Z","shell.execute_reply":"2024-04-20T12:03:26.717743Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ptend_t0_corr_df = pd.DataFrame({'Corr':ptend_t_0_corr})\nptend_t0_corr_df.sort_values(by='Corr',ascending=False,inplace=True)","metadata":{"execution":{"iopub.status.busy":"2024-04-20T12:14:54.097342Z","iopub.execute_input":"2024-04-20T12:14:54.097865Z","iopub.status.idle":"2024-04-20T12:14:54.106963Z","shell.execute_reply.started":"2024-04-20T12:14:54.09783Z","shell.execute_reply":"2024-04-20T12:14:54.105435Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ptend_t0_corr_df","metadata":{"execution":{"iopub.status.busy":"2024-04-20T12:14:58.18942Z","iopub.execute_input":"2024-04-20T12:14:58.189957Z","iopub.status.idle":"2024-04-20T12:14:58.20554Z","shell.execute_reply.started":"2024-04-20T12:14:58.18992Z","shell.execute_reply":"2024-04-20T12:14:58.204026Z"},"trusted":true},"execution_count":null,"outputs":[]}]}