{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":56537,"databundleVersionId":8015876,"sourceType":"competition"},{"sourceId":8182922,"sourceType":"datasetVersion","datasetId":4844949}],"dockerImageVersionId":30698,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2024-04-27T03:11:03.858041Z","iopub.execute_input":"2024-04-27T03:11:03.858486Z","iopub.status.idle":"2024-04-27T03:11:03.869285Z","shell.execute_reply.started":"2024-04-27T03:11:03.858454Z","shell.execute_reply":"2024-04-27T03:11:03.868047Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Learning the Earth with Artificial Intelligence and Physics (LEAP)¶\n","metadata":{}},{"cell_type":"markdown","source":"# ClimSim: The largest-ever dataset designed for hybrid ML-physics research¶\nClimSim: A large multi-scale dataset for hybrid physics-ML climate emulation\n\nAuthors: Sungduk Yu1, Walter M. Hannah,Liran Peng,Jerry Lin, Mohamed Aziz Bhouri, Ritwik Gupta, Björn Lütjens, Justus C. Will, Gunnar Behrens,Julius J. M. Busecke, Nora Loose, Charles Stern, Tom Beucler, Bryce E. Harrop, Benjamin R. Hillman, Andrea M. Jenney,Savannah L. Ferretti, Nana Liu, Anima Anandkumar,Noah D. Brenowitz, Veronika Eyring,Nicholas Geneva, Pierre Gentine, Stephan Mandt, Jaideep Pathak, Akshay Subramaniam, Carl Vondrick, Rose Yu,Laure Zanna, Tian Zhen,Ryan P. Abernathey, Fiaz Ahmed, David C. Bader, Pierre Baldi, Elizabeth A. Barnes, Christopher S. Bretherton, Peter M. Caldwell, Wayne Chuang, Yilun Han, Yu Huang, Fernando Iglesias-Suarez, Sanket Jantre, Karthik Kashinath, Marat Khairoutdinov,Thorsten Kurth, Nicholas J. Lutsko, Po-Lun Ma, Griffin Mooers, J. David Neelin, David A. Randall, Sara Shamekh, Mark A. Taylor, Nathan M. Urban, Janni Yuval, Guang J. Zhang, Michael S. Pritchard\n\n\"The authors presented ClimSim, the largest-ever dataset designed for hybrid ML-physics research. It comprises multi-scale climate simulations, developed by a consortium of climate scientists and ML researchers. It consists of 5.7 billion pairs of multivariate input and output vectors that isolate the influence of locally-nested, high-resolution, high-fidelity physics on a host climate simulator’s macro-scale physical state.\"\n\n\"ClimSim, the most physically comprehensive dataset yet published for training ML emulators of atmospheric storms, clouds, turbulence, rainfall, and radiation for use in hybrid-ML climate simulation. It contains all inputs and outputs necessary for downstream coupling in a fullcomplexity multi-scale climate simulator. The authors conducted a series of experiments on a subset of these variables that demonstrate the degree to which climate data scientists have been able to fit their deterministic and stochastic components.\"\n\n\"They hope ML community engagement in ClimSim will advance fundamental ML methodology and clarify the path to producing increasingly skillful sub-grid physics emulators that can be reliably used for operational climate simulation. To facilitate two-way commications between ML practitioners and climate scientists, the authors incorporated many desired characteristics for an ideal benchmark dataset.\"\n\n\"Such interdisciplinary collaboration will open up an exciting future in which the computational limits that currently constrain climate simulation can be reconsidered. The authors plan to soon extend ClimSim to include, first, a sampling of multiple future climate states. Second, they aim to provide a protocol for downstream hybrid simulation testing. We hope lessons learned in their chosen limit of multi-scale atmospheric simulation will have applicability in other sub-fields of Earth System Science where computational constraints are currently a barrier to including explicit representations of more systems of nested complexity.\"\n\nPhysics-Informed Guidance to Improve Generalizability and Coupled Performance\n\n\"Physical Constraints: Mass and energy conservation are important criteria for Earth system simulation. If these terms are not conserved, errors in estimating sea level rise or temperature change over time may become as large as the signals we hope to measure. Enforcing conservation on emulated results helps constrain results to be physically plausible and reduce the potential for errors accumulating over long time scales. We discuss how to do this and enforce additional constraints, such as non-negativity for precipitation, condensate, and moisture variables in the Supporting Information.\"\n\n\"Stochasticity and Memory: The results of the embedded convection calculations regulating do are chaotic, and thus worthy of stochastic architectures, as in our RPN, HSR, and cVAE baselines. These solutions are likewise sensitive to sub-grid initial state variables from an interior nested spatial dimension that has not been included in their data.\"\n\n\"Temporal Locality: Incorporating the previous timesteps’ target or feature in the input vector inflation could be beneficial as it captures some information about this convective memory and utilizes temporal autocorrelations present in atmospheric data.\"\n\n\"Causal Pruning: A systematic and quantitative pruning of the input vector based on objectively assessed causal relationships to subsets of the target vector has been proposed as an attractive preprocessing strategy, as it helps remove spurious correlations due to confounding variables and optimize the ML algorithm.\"\n\n\"Normalization: Normalization that goes beyond removing vertical structure could be strategic, such as removing the geographic mean (e.g., latitudinal, land/sea structure) or composite seasonal variances (e.g., local smoothed annual cycle) present in the data. For variables exhibiting exponential variation and approaching zero at the highest level (e.g., metrics of moisture), log-normalization might be beneficial.\"\n","metadata":{}},{"cell_type":"markdown","source":"# Data overview¶\n# Inputs:\n* Arrays (dimension 60):\n* state_t: air temperature\n* state_q0001: specific humidity\n* state_q0002: cloud liquid mixing ratio\n* state_q0003: cloud ice mixing ratio\n* state_u: zonal wind speed\n* state_v: meridional wind speed\n* pbuf_ozone: ozone volume mixing ratio\n* pbuf_CH4: methane volume mixing ratio\n* pbuf_N2O: nitrous oxide volume mixing ratio\n# Scalars:\n* state_ps: surface pressure\n* pbuf_SOLIN: solar insolation\n* pbuf_LHFLX: surface latent heat flux\n* pbuf_SHFLX: surface sensible heat flux\n* pbuf_TAUX: zonal surface stress\n* pbuf_TAUY: meridional surface stress\n* pbuf_COSZRS: cosine of solar zenith angle\n* cam_in_ALDIF: albedo for diffuse longwave radiation\n* cam_in_ALDIR: albedo for direct longwave radiation\n* cam_in_ASDIF: albedo for diffuse shortwave radiation\n* cam_in_ASDIR: albedo for direct shortwave radiation\n* cam_in_LWUP: upward longwave flux\n* cam_in_ICEFRAC: sea-ice areal fraction\n* cam_in_LANDFRAC: land areal fraction\n* cam_in_OCNFRAC: ocean areal fraction\n* cam_in_SNOWHLAND: snow depth over land\n# Targets:\n* Arrays (dimension 60):\n* ptend_t: heating tendency\n* ptend_q0001: moistening tendency\n* ptend_q0002: cloud liquid mixing ratio change over time\n* ptend_q0003: cloud ice mixing ratio change over time\n* ptend_u: zonal wind acceleration\n* ptend_v: meridional wind acceleration\n# Scalars:\n* cam_out_NETSW: net shortwave flux at surface\n* cam_out_FLWDS: downward longwave flux at surface\n* cam_out_PRECSC: snow rate (liquid water equivalent)\n* cam_out_PRECC: rain rate\n* cam_out_SOLS: downward visible direct solar flux to surface\n* cam_out_SOLL: downward near-infrared direct solar flux to surface\n* cam_out_SOLSD: downward diffuse solar flux to surface\n* cam_out_SOLLD: downward diffuse near-infrared solar flux to surface\n\n# Init","metadata":{}},{"cell_type":"code","source":"# packages\n\n# standard\nimport numpy as np\nimport pandas as pd\nimport time\n\n# plots\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n# faster alternative to pandas\nimport polars as pl\n","metadata":{"execution":{"iopub.status.busy":"2024-04-27T03:11:03.871133Z","iopub.execute_input":"2024-04-27T03:11:03.871513Z","iopub.status.idle":"2024-04-27T03:11:03.879035Z","shell.execute_reply.started":"2024-04-27T03:11:03.871482Z","shell.execute_reply":"2024-04-27T03:11:03.87789Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# configs\npd.set_option('display.max_columns', None) # we want to display all columns in this notebook\n\n# aesthetics\ndefault_color_1 = 'darkblue'\ndefault_color_2 = 'darkgreen'\ndefault_color_3 = 'darkred'\n\n# random seed\nmy_random_seed = 111\n","metadata":{"execution":{"iopub.status.busy":"2024-04-27T03:11:03.880469Z","iopub.execute_input":"2024-04-27T03:11:03.88096Z","iopub.status.idle":"2024-04-27T03:11:03.891037Z","shell.execute_reply.started":"2024-04-27T03:11:03.880918Z","shell.execute_reply":"2024-04-27T03:11:03.889874Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n## Import and Overview (subset)\n","metadata":{}},{"cell_type":"code","source":"# file overview\n!ls -l '../input/leap-atmospheric-physics-ai-climsim/'\n","metadata":{"execution":{"iopub.status.busy":"2024-04-27T03:11:03.89325Z","iopub.execute_input":"2024-04-27T03:11:03.893684Z","iopub.status.idle":"2024-04-27T03:11:05.109034Z","shell.execute_reply.started":"2024-04-27T03:11:03.893646Z","shell.execute_reply":"2024-04-27T03:11:05.107638Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Datasets are HUGE, let's start with a small subset:¶\n","metadata":{}},{"cell_type":"code","source":"# import SUBSET of data\nn_rows = 50000\nfolder = 'leap-atmospheric-physics-ai-climsim'\nt1 = time.time()\ndf_train = pl.read_csv('../input/'+folder+'/train.csv', n_rows=n_rows).to_pandas()\ndf_test = pl.read_csv('../input/'+folder+'/test.csv', n_rows=n_rows).to_pandas()\ndf_sub = pl.read_csv('../input/'+folder+'/sample_submission.csv', n_rows=n_rows).to_pandas()\nt2 = time.time()\nprint('Elapsed time [s]: ', np.round(t2-t1,2))","metadata":{"execution":{"iopub.status.busy":"2024-04-27T03:11:05.111109Z","iopub.execute_input":"2024-04-27T03:11:05.111502Z","iopub.status.idle":"2024-04-27T03:11:09.050525Z","shell.execute_reply.started":"2024-04-27T03:11:05.111467Z","shell.execute_reply":"2024-04-27T03:11:09.049692Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# preview - train\ndf_train.head(10)\n","metadata":{"execution":{"iopub.status.busy":"2024-04-27T03:11:09.052457Z","iopub.execute_input":"2024-04-27T03:11:09.053031Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# train set overview\ndf_train.info(verbose=True, show_counts=True)\n","metadata":{"_kg_hide-output":false,"_kg_hide-input":false,"execution":{"iopub.status.busy":"2024-04-27T03:11:10.049446Z","iopub.execute_input":"2024-04-27T03:11:10.049893Z","iopub.status.idle":"2024-04-27T03:11:10.203418Z","shell.execute_reply.started":"2024-04-27T03:11:10.049834Z","shell.execute_reply":"2024-04-27T03:11:10.202179Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# preview - test\ndf_test.head(10)\n","metadata":{"execution":{"iopub.status.busy":"2024-04-27T03:11:10.204701Z","iopub.execute_input":"2024-04-27T03:11:10.205068Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# test set overview\ndf_test.info(verbose=True, show_counts=True)\n","metadata":{"execution":{"iopub.execute_input":"2024-04-27T03:11:10.814058Z","iopub.status.idle":"2024-04-27T03:11:10.917893Z","shell.execute_reply.started":"2024-04-27T03:11:10.814006Z","shell.execute_reply":"2024-04-27T03:11:10.916765Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Targets and Features¶\n","metadata":{}},{"cell_type":"code","source":"# targets (extract from submission file)\ntargets = [x for x in df_sub.columns.tolist() if x not in ['sample_id']]\n\n# numerical features\nfeatures_num = [x for x in df_train.columns.tolist() if x not in ['sample_id']+targets]\n\n# categorical features\nfeatures_cat = []\n\n","metadata":{"execution":{"iopub.status.busy":"2024-04-27T03:11:10.919211Z","iopub.execute_input":"2024-04-27T03:11:10.91953Z","iopub.status.idle":"2024-04-27T03:11:10.933074Z","shell.execute_reply.started":"2024-04-27T03:11:10.919505Z","shell.execute_reply":"2024-04-27T03:11:10.931996Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# all features combined\nfeatures = features_num + features_cat\n# ouput of dimensions\nprint(\"Number of numerical features:\", len(features_num))\nprint(\"Number of categorical features:\", len(features_cat))\nprint(\"Number of targets:\", len(targets))\nprint(\"Size of subset:\", n_rows)\n","metadata":{"execution":{"iopub.status.busy":"2024-04-27T03:11:10.936749Z","iopub.execute_input":"2024-04-27T03:11:10.93769Z","iopub.status.idle":"2024-04-27T03:11:10.94573Z","shell.execute_reply.started":"2024-04-27T03:11:10.937648Z","shell.execute_reply":"2024-04-27T03:11:10.944693Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def box_plots(ind1:int,ind2:int):\n     plt.figure(figsize=(15,5)) \n     sns.boxplot(data=df_train.iloc[:,ind1:ind2], palette=\"deep\")\n     sns.despine(left=True)\n     plt.title(f\"col_{ind1} to col_{ind2}\")\n     plt.xticks(rotation=90)   \n     plt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-04-27T03:11:10.947103Z","iopub.execute_input":"2024-04-27T03:11:10.947403Z","iopub.status.idle":"2024-04-27T03:11:10.958544Z","shell.execute_reply.started":"2024-04-27T03:11:10.94738Z","shell.execute_reply":"2024-04-27T03:11:10.957459Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"box_plots(0,60)\n","metadata":{"execution":{"iopub.status.busy":"2024-04-27T03:11:10.960162Z","iopub.execute_input":"2024-04-27T03:11:10.960471Z","iopub.status.idle":"2024-04-27T03:11:12.292128Z","shell.execute_reply.started":"2024-04-27T03:11:10.960446Z","shell.execute_reply":"2024-04-27T03:11:12.290994Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"box_plots(61,121)\n","metadata":{"execution":{"iopub.status.busy":"2024-04-27T03:11:12.293433Z","iopub.execute_input":"2024-04-27T03:11:12.293763Z","iopub.status.idle":"2024-04-27T03:11:13.655696Z","shell.execute_reply.started":"2024-04-27T03:11:12.293734Z","shell.execute_reply":"2024-04-27T03:11:13.654457Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Targets","metadata":{}},{"cell_type":"code","source":"# basic stats - targets\ndf_train[targets].describe()\n","metadata":{"execution":{"iopub.status.busy":"2024-04-27T03:11:13.657316Z","iopub.execute_input":"2024-04-27T03:11:13.657766Z","iopub.status.idle":"2024-04-27T03:11:15.47392Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"'''#plot target distributions in compact matrix form\nfig, axs = plt.subplots(92, 4, figsize=(16,350))\ni = 0\nfor t in targets:\n    current_ax = axs.flat[i]\n    current_ax.hist(df_train[t], bins=100, color=default_color_3)\n    current_ax.set_title('Target ' + str(t))\n    current_ax.grid()\n    i = i + 1'''\n","metadata":{"execution":{"iopub.status.busy":"2024-04-27T03:11:15.475267Z","iopub.execute_input":"2024-04-27T03:11:15.475606Z","iopub.status.idle":"2024-04-27T03:11:15.483138Z","shell.execute_reply.started":"2024-04-27T03:11:15.475575Z","shell.execute_reply":"2024-04-27T03:11:15.481889Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Features¶\n","metadata":{}},{"cell_type":"code","source":"# basic stats - train\ndf_train[features_num].describe()\n","metadata":{"execution":{"iopub.status.busy":"2024-04-27T03:11:15.48468Z","iopub.execute_input":"2024-04-27T03:11:15.485299Z","iopub.status.idle":"2024-04-27T03:11:18.039495Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# basic stats - test\ndf_test[features_num].describe()\n","metadata":{"execution":{"iopub.status.busy":"2024-04-27T03:11:18.04102Z","iopub.execute_input":"2024-04-27T03:11:18.041429Z","iopub.status.idle":"2024-04-27T03:11:20.548447Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"'''# plot histograms for numerical features (train and test)\nfor f in features_num:\n    plt.figure(figsize=(12,2))\n    ax1 = plt.subplot(1,2,1)\n    df_train[f].plot(kind='hist', bins=100, color=default_color_1)\n    plt.title(f + ' - Train')\n    plt.grid()\n    ax2 = plt.subplot(1,2,2, sharex=ax1)\n    df_test[f].plot(kind='hist', bins=100, color=default_color_2)\n    plt.title(f + ' - Test')\n    plt.grid()\n    plt.show()'''\n","metadata":{"execution":{"iopub.status.busy":"2024-04-27T03:11:20.5502Z","iopub.execute_input":"2024-04-27T03:11:20.551131Z","iopub.status.idle":"2024-04-27T03:11:20.558322Z","shell.execute_reply.started":"2024-04-27T03:11:20.551082Z","shell.execute_reply":"2024-04-27T03:11:20.557161Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"'''# compact boxplot of all features - train only\nn_plot_rows = 10\nn_plot_cols = 60\nn = len(features_num)\nfor i in range(n_plot_rows):\n    a = n_plot_cols*i+1\n    b = min(n_plot_cols*i+n_plot_cols, n)\n    print('Columns', a, 'to', b)\n    plt.figure(figsize=(14,3))\n    df_train.iloc[:,a:(b+1)].plot(kind='box', figsize=(15,5))\n    plt.xticks(rotation=90)\n    plt.grid()\n    plt.show()'''","metadata":{"execution":{"iopub.status.busy":"2024-04-27T03:11:20.559796Z","iopub.execute_input":"2024-04-27T03:11:20.560162Z","iopub.status.idle":"2024-04-27T03:11:20.579861Z","shell.execute_reply.started":"2024-04-27T03:11:20.560134Z","shell.execute_reply":"2024-04-27T03:11:20.578737Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"'''# boxplots (train and test)\nfor f in features_num:\n    plt.figure(figsize=(14,0.5))\n    ax1 = plt.subplot(1,2,1)\n    df_temp = df_train[f].dropna() # boxplot does not like missings...\n    plt.boxplot(df_temp, vert=False)\n    plt.title(f + ' - Train')\n    plt.grid()\n    ax2 = plt.subplot(1,2,2, sharex=ax1)\n    df_temp = df_test[f].dropna()\n    plt.boxplot(df_temp, vert=False)\n    plt.title(f + ' - Test')\n    plt.grid()\n    plt.show()\n'''","metadata":{"execution":{"iopub.status.busy":"2024-04-27T03:11:20.58113Z","iopub.execute_input":"2024-04-27T03:11:20.581451Z","iopub.status.idle":"2024-04-27T03:11:20.590559Z","shell.execute_reply.started":"2024-04-27T03:11:20.581425Z","shell.execute_reply":"2024-04-27T03:11:20.589628Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Correlations¶\n# Targets\n","metadata":{}},{"cell_type":"code","source":"'''# calc and plot correlation matrix\ncor_p_target = df_train[targets].corr(method='pearson')\nplt.figure(figsize=(14,12))\nsns.heatmap(cor_p_target, annot=False, cmap='RdYlGn',\n            vmin=-1, vmax=+1)\nplt.title('Targets - Pearson Correlation')\nplt.show()'''","metadata":{"execution":{"iopub.status.busy":"2024-04-27T03:11:20.591802Z","iopub.execute_input":"2024-04-27T03:11:20.592152Z","iopub.status.idle":"2024-04-27T03:11:20.602097Z","shell.execute_reply.started":"2024-04-27T03:11:20.592125Z","shell.execute_reply":"2024-04-27T03:11:20.601134Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Features (train)¶\n","metadata":{}},{"cell_type":"code","source":"'''# calc and plot correlation matrix\ncor_p_train = df_train[features_num].corr(method='pearson')\nplt.figure(figsize=(14,12))\nsns.heatmap(cor_p_train, annot=False, cmap='RdYlGn',\n            vmin=-1, vmax=+1)\nplt.title('Features - Pearson Correlation (train)')\nplt.show()'''\n","metadata":{"execution":{"iopub.status.busy":"2024-04-27T03:11:20.60336Z","iopub.execute_input":"2024-04-27T03:11:20.603671Z","iopub.status.idle":"2024-04-27T03:11:20.613564Z","shell.execute_reply.started":"2024-04-27T03:11:20.603644Z","shell.execute_reply":"2024-04-27T03:11:20.612529Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Features (test)¶\n","metadata":{}},{"cell_type":"code","source":"'''# calc and plot correlation matrix\ncor_p_test = df_test[features_num].corr(method='pearson')\nplt.figure(figsize=(14,12))\nsns.heatmap(cor_p_test, annot=False, cmap='RdYlGn',\n            vmin=-1, vmax=+1)\nplt.title('Features - Pearson Correlation (test)')\nplt.show()'''\n","metadata":{"execution":{"iopub.status.busy":"2024-04-27T03:11:20.615074Z","iopub.execute_input":"2024-04-27T03:11:20.615795Z","iopub.status.idle":"2024-04-27T03:11:20.627428Z","shell.execute_reply.started":"2024-04-27T03:11:20.615763Z","shell.execute_reply":"2024-04-27T03:11:20.626377Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"'''# export results\ndf_train.to_csv('df_train_subset.csv')\ncor_p_target.to_csv('cor_p_target.csv')\ncor_p_train.to_csv('cor_p_train.csv')\ncor_p_test.to_csv('cor_p_test.csv')'''","metadata":{"execution":{"iopub.status.busy":"2024-04-27T03:11:20.628922Z","iopub.execute_input":"2024-04-27T03:11:20.629388Z","iopub.status.idle":"2024-04-27T03:11:20.644707Z","shell.execute_reply.started":"2024-04-27T03:11:20.629348Z","shell.execute_reply":"2024-04-27T03:11:20.643123Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Other Explorations¶\n","metadata":{}},{"cell_type":"markdown","source":"Target vs row index\n","metadata":{}},{"cell_type":"code","source":"# plot target values\n'''for t in targets:\n    plt.figure(figsize=(14,2))\n    plt.scatter(df_train.index, df_train[t], color=default_color_3,\n                alpha=0.25, s=1)\n    plt.title(t)\n    plt.grid()\n    plt.show()\n'''","metadata":{"execution":{"iopub.status.busy":"2024-04-27T03:11:20.65195Z","iopub.execute_input":"2024-04-27T03:11:20.652405Z","iopub.status.idle":"2024-04-27T03:11:20.661863Z","shell.execute_reply.started":"2024-04-27T03:11:20.652361Z","shell.execute_reply":"2024-04-27T03:11:20.660764Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Features vs row index\n","metadata":{}},{"cell_type":"code","source":"'''# plot target values\nfor f in features_num:\n    plt.figure(figsize=(14,2))\n    plt.scatter(df_train.index, df_train[f], color=default_color_1,\n                alpha=0.25, s=1)\n    plt.title(f)\n    plt.grid()\n    plt.show()\n'''\n","metadata":{"execution":{"iopub.status.busy":"2024-04-27T03:11:20.665233Z","iopub.execute_input":"2024-04-27T03:11:20.665929Z","iopub.status.idle":"2024-04-27T03:11:20.674299Z","shell.execute_reply.started":"2024-04-27T03:11:20.665883Z","shell.execute_reply":"2024-04-27T03:11:20.672421Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Import individual columns (full data)\n","metadata":{}},{"cell_type":"markdown","source":"In order to approach the full dataset we could try to import just a subset of columns. This is done in the following section.\n","metadata":{}},{"cell_type":"code","source":"# define columns (has to be a list)\nn_max = 20 # columns with index 0..n_max\ncols_select = ['state_t_' + str(t) for t in range(0,n_max+1)]\nprint(cols_select)","metadata":{"execution":{"iopub.status.busy":"2024-04-27T03:11:20.676306Z","iopub.execute_input":"2024-04-27T03:11:20.676708Z","iopub.status.idle":"2024-04-27T03:11:20.68506Z","shell.execute_reply.started":"2024-04-27T03:11:20.676674Z","shell.execute_reply":"2024-04-27T03:11:20.683758Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"'''# load only selected column\nt1 = time.time()\ndf_col = pl.read_csv('../input/'+folder+'/train.csv', columns=cols_select).to_pandas()\nt2 = time.time()\nprint('Elapsed time [s]: ', np.round(t2-t1,2))\nprint('Number of rows: ', df_col.shape[0])\nprint('Number of cols: ', df_col.shape[1])'''","metadata":{"execution":{"iopub.status.busy":"2024-04-27T03:11:20.687411Z","iopub.execute_input":"2024-04-27T03:11:20.687892Z","iopub.status.idle":"2024-04-27T03:11:20.701695Z","shell.execute_reply.started":"2024-04-27T03:11:20.68785Z","shell.execute_reply":"2024-04-27T03:11:20.698391Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"'''# basic stats\ndf_col.describe(percentiles=[0.01,0.1,0.25,0.5,0.75,0.9,0.99])\n'''","metadata":{"execution":{"iopub.status.busy":"2024-04-27T03:11:20.703464Z","iopub.execute_input":"2024-04-27T03:11:20.704628Z","iopub.status.idle":"2024-04-27T03:11:20.714664Z","shell.execute_reply.started":"2024-04-27T03:11:20.70458Z","shell.execute_reply":"2024-04-27T03:11:20.713098Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"'''# plot distributions\nfor f in cols_select:\n    plt.figure(figsize=(10,3))\n    plt.hist(df_col[f], bins=1000, color=default_color_1)\n    plt.title(f + ' - full data')\n    plt.grid()\n    plt.show()'''\n","metadata":{"execution":{"iopub.status.busy":"2024-04-27T03:11:20.717404Z","iopub.execute_input":"2024-04-27T03:11:20.717936Z","iopub.status.idle":"2024-04-27T03:11:20.725394Z","shell.execute_reply.started":"2024-04-27T03:11:20.717889Z","shell.execute_reply":"2024-04-27T03:11:20.724384Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"💡 state_t_0 shows some unusually high values, let's have a closer look:¶\n","metadata":{}},{"cell_type":"code","source":"'''# boxplot\nplt.figure(figsize=(10,0.5))\nplt.boxplot(df_col.state_t_0, vert=False)\nplt.title('state_t_0')\nplt.grid()\nplt.show()'''\n","metadata":{"execution":{"iopub.status.busy":"2024-04-27T03:11:20.727109Z","iopub.execute_input":"2024-04-27T03:11:20.7275Z","iopub.status.idle":"2024-04-27T03:11:20.736899Z","shell.execute_reply.started":"2024-04-27T03:11:20.72747Z","shell.execute_reply":"2024-04-27T03:11:20.735888Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"'''# let's check the most extreme outliers\ndf_col[df_col.state_t_0>400]\n'''","metadata":{"execution":{"iopub.status.busy":"2024-04-27T03:11:20.738176Z","iopub.execute_input":"2024-04-27T03:11:20.738533Z","iopub.status.idle":"2024-04-27T03:11:20.749505Z","shell.execute_reply.started":"2024-04-27T03:11:20.738505Z","shell.execute_reply":"2024-04-27T03:11:20.748562Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Correlations:¶\n","metadata":{}},{"cell_type":"code","source":"'''# calc and plot correlation matrix\ncor_p_train_few_cols = df_train[cols_select].corr(method='pearson')\nplt.figure(figsize=(14,10))\nsns.heatmap(cor_p_train_few_cols, annot=True, cmap='RdYlGn',\n            fmt='.2f', linecolor='black', linewidth=.5,\n            vmin=-1, vmax=+1)\nplt.title('Features - Pearson Correlation (train/selected columns)')\nplt.show()'''","metadata":{"execution":{"iopub.status.busy":"2024-04-27T03:11:20.750924Z","iopub.execute_input":"2024-04-27T03:11:20.751249Z","iopub.status.idle":"2024-04-27T03:11:20.762881Z","shell.execute_reply.started":"2024-04-27T03:11:20.751221Z","shell.execute_reply":"2024-04-27T03:11:20.761857Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# chunk wise training trial","metadata":{}},{"cell_type":"code","source":"data_path='/kaggle/input/leap-atmospheric-physics-ai-climsim/'\n","metadata":{"execution":{"iopub.status.busy":"2024-04-27T03:11:20.764205Z","iopub.execute_input":"2024-04-27T03:11:20.765054Z","iopub.status.idle":"2024-04-27T03:11:20.77525Z","shell.execute_reply.started":"2024-04-27T03:11:20.765021Z","shell.execute_reply":"2024-04-27T03:11:20.774182Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cols_13_pbuf_N2O = [\"pbuf_N2O_\"+str(i) for i in range(27,60)]\ncols_13_pbuf_CH4 = [\"pbuf_CH4_\"+str(i) for i in range(27,60)]\n\ncols_to_delete = cols_13_pbuf_N2O + cols_13_pbuf_CH4\n\ncols_to_delete.append('sample_id')\nlen(cols_to_delete)\n","metadata":{"execution":{"iopub.status.busy":"2024-04-27T03:11:20.776703Z","iopub.execute_input":"2024-04-27T03:11:20.77715Z","iopub.status.idle":"2024-04-27T03:11:20.789902Z","shell.execute_reply.started":"2024-04-27T03:11:20.777112Z","shell.execute_reply":"2024-04-27T03:11:20.788876Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub_df=pd.read_csv(data_path+'sample_submission.csv',nrows=2)\n","metadata":{"execution":{"iopub.status.busy":"2024-04-27T03:11:20.791178Z","iopub.execute_input":"2024-04-27T03:11:20.791614Z","iopub.status.idle":"2024-04-27T03:11:20.817003Z","shell.execute_reply.started":"2024-04-27T03:11:20.791584Z","shell.execute_reply":"2024-04-27T03:11:20.815894Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub_df.drop(columns='sample_id',inplace=True)\n","metadata":{"execution":{"iopub.status.busy":"2024-04-27T03:11:20.818284Z","iopub.execute_input":"2024-04-27T03:11:20.81862Z","iopub.status.idle":"2024-04-27T03:11:20.825163Z","shell.execute_reply.started":"2024-04-27T03:11:20.818591Z","shell.execute_reply":"2024-04-27T03:11:20.823898Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"target_cols=sub_df.columns\n","metadata":{"execution":{"iopub.status.busy":"2024-04-27T03:11:20.826797Z","iopub.execute_input":"2024-04-27T03:11:20.827183Z","iopub.status.idle":"2024-04-27T03:11:20.837469Z","shell.execute_reply.started":"2024-04-27T03:11:20.827143Z","shell.execute_reply":"2024-04-27T03:11:20.836295Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_chunk=pd.read_csv(data_path+'train.csv',chunksize=100000,usecols=lambda col: col not in cols_to_delete)\n","metadata":{"execution":{"iopub.status.busy":"2024-04-27T03:11:20.839128Z","iopub.execute_input":"2024-04-27T03:11:20.839576Z","iopub.status.idle":"2024-04-27T03:11:20.852874Z","shell.execute_reply.started":"2024-04-27T03:11:20.839537Z","shell.execute_reply":"2024-04-27T03:11:20.85166Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"'''import pandas as pd\nfrom sklearn.linear_model import LinearRegression\nfrom sklearn.metrics import r2_score\nfrom sklearn.decomposition import PCA\nimport pickle\n\n# Initialize list to store R^2 scores\nr2_scores = []\n\n# Iterate through each chunk of data\nfor chunk_idx, chunk in enumerate(data_chunk, 1):\n    # Split chunk into features (X_train) and target (y_train)\n    y_train = chunk[target_cols]\n    X_train = chunk.drop(columns=target_cols)\n    \n    # Apply PCA to X_train\n    pca = PCA()\n    X_train_pca = pca.fit_transform(X_train)\n    \n    # Train a linear regression model on PCA-transformed data\n    model_pca = LinearRegression()\n    model_pca.fit(X_train_pca, y_train)\n    \n    # Calculate R^2 score\n    y_pred_pca = model_pca.predict(X_train_pca)\n    r2_pca = r2_score(y_train, y_pred_pca)\n    \n    # Print and store R^2 score\n    print(f\"Chunk {chunk_idx} - R^2 Score: {r2_pca}\")\n    r2_scores.append(r2_pca)\n    \n    # Write R^2 score to pickle file\n    r2_filename = f'r2_score_chunk_{chunk_idx}.pkl'\n    with open(r2_filename, 'wb') as r2_file:\n        pickle.dump(r2_pca, r2_file)\n'''","metadata":{"execution":{"iopub.status.busy":"2024-04-27T03:11:20.854427Z","iopub.execute_input":"2024-04-27T03:11:20.854941Z","iopub.status.idle":"2024-04-27T03:11:20.863357Z","shell.execute_reply.started":"2024-04-27T03:11:20.854899Z","shell.execute_reply":"2024-04-27T03:11:20.862231Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%pip install HROCH\n","metadata":{"execution":{"iopub.status.busy":"2024-04-27T03:13:07.879285Z","iopub.execute_input":"2024-04-27T03:13:07.879711Z","iopub.status.idle":"2024-04-27T03:13:21.320358Z","shell.execute_reply.started":"2024-04-27T03:13:07.879678Z","shell.execute_reply":"2024-04-27T03:13:21.318688Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os, gc\nimport dask.dataframe as dd\nfrom HROCH import SymbolicRegressor\nfrom sklearn.linear_model import LinearRegression\nfrom sklearn.tree import DecisionTreeRegressor\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import r2_score\n","metadata":{"execution":{"iopub.status.busy":"2024-04-27T03:13:23.182364Z","iopub.execute_input":"2024-04-27T03:13:23.183344Z","iopub.status.idle":"2024-04-27T03:13:24.075393Z","shell.execute_reply.started":"2024-04-27T03:13:23.183302Z","shell.execute_reply":"2024-04-27T03:13:24.074277Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\ndf_sample = pd.read_csv('/kaggle/input/leap-sample-1-perc/train_sample_1.csv')\n#forgoten index\ndf_sample.drop(columns=df_sample.columns[0], axis=1, inplace=True)","metadata":{"execution":{"iopub.status.busy":"2024-04-27T03:21:00.557135Z","iopub.execute_input":"2024-04-27T03:21:00.558138Z","iopub.status.idle":"2024-04-27T03:21:45.492419Z","shell.execute_reply.started":"2024-04-27T03:21:00.558096Z","shell.execute_reply":"2024-04-27T03:21:45.491094Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_sub_tmp = pd.read_csv('/kaggle/input/leap-atmospheric-physics-ai-climsim/sample_submission.csv', nrows=10)\ntargets = [x for x in df_sub_tmp.columns.tolist() if x != 'sample_id']\nfeatures = [x for x in df_train.columns.tolist() if x not in ['sample_id']+targets]\nlen(targets)","metadata":{"execution":{"iopub.status.busy":"2024-04-27T03:21:47.713113Z","iopub.execute_input":"2024-04-27T03:21:47.714232Z","iopub.status.idle":"2024-04-27T03:21:47.744162Z","shell.execute_reply.started":"2024-04-27T03:21:47.714194Z","shell.execute_reply":"2024-04-27T03:21:47.743083Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nDEBUG = False\nX = df_sample.drop(targets + ['sample_id'], axis = 1)\n\nlr_models = {}\ndt_models = {}\nsr_models = {}\nfor i, target in enumerate(targets):\n    print(f'[{i}.] Fit target {target}')\n    y = df_sample[target]\n    X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.9, random_state=42)\n    mean_y = np.mean(y)\n    mean_train = np.mean(y_train)\n    score_mean = r2_score(y_test, [mean_train]*len(y_test))\n    print(f'MEAN score: test {score_mean}')\n    lr_est = LinearRegression()\n    lr_est.fit(X_train, y_train)\n    lr_score_train, lr_score_test = lr_est.score(X_train, y_train), lr_est.score(X_test, y_test)\n    print(f'LR score: train {lr_score_train} test {lr_score_test}')\n    lr_models[target] = (lr_score_test, lr_est, np.mean(y_train))\n    if lr_score_test > 0.95 or DEBUG:\n        continue\n    \n    dt_est = DecisionTreeRegressor()\n    dt_est.fit(X_train, y_train)\n    dt_score_train, dt_score_test = dt_est.score(X_train, y_train), dt_est.score(X_test, y_test)\n    print(f'DT score: train {dt_score_train} test {dt_score_test}')\n    dt_models[target] = (dt_score_test, dt_est)\n    \n    sr_est = SymbolicRegressor(time_limit=1.0 if DEBUG else 5.0,num_threads=4, precision='f64', feature_probs = dt_est.feature_importances_)\n    sr_est.fit(X_train, y_train)\n    sr_score_train, sr_score_test = sr_est.score(X_train, y_train), sr_est.score(X_test, y_test)\n    print(f'SR score: train {sr_score_train} test {sr_score_test}')\n    print(sr_est.sexpr_)\n    sr_models[target] = (sr_score_test, sr_est)\n","metadata":{"execution":{"iopub.status.busy":"2024-04-27T03:21:48.476536Z","iopub.execute_input":"2024-04-27T03:21:48.476984Z","iopub.status.idle":"2024-04-27T04:30:13.246667Z","shell.execute_reply.started":"2024-04-27T03:21:48.476949Z","shell.execute_reply":"2024-04-27T04:30:13.245066Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\ndf_train = dd.read_csv('/kaggle/input/leap-atmospheric-physics-ai-climsim/train.csv', blocksize = '10mb')\ndf_train.head()","metadata":{"execution":{"iopub.status.busy":"2024-04-27T04:30:13.249362Z","iopub.execute_input":"2024-04-27T04:30:13.249768Z","iopub.status.idle":"2024-04-27T04:30:17.921745Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\npart = df_train.get_partition(0)\npart_size = len(part)\npart_size, df_train.npartitions, part_size*df_train.npartitions\n","metadata":{"execution":{"iopub.status.busy":"2024-04-27T04:30:17.923921Z","iopub.execute_input":"2024-04-27T04:30:17.924424Z","iopub.status.idle":"2024-04-27T04:30:19.393295Z","shell.execute_reply.started":"2024-04-27T04:30:17.924385Z","shell.execute_reply":"2024-04-27T04:30:19.392047Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# %%time\n #import warnings\n #warnings.filterwarnings('ignore')\n #avg_score = [0.0]*len(lr_models)\n #for part in range(df_train.npartitions):\n #    gc.collect()\n #    df_test_ = df_train.get_partition(part)\n #    df_test = df_test_.compute()\n #    print(part)\n #    for target, model in lr_models.items():\n  #       if model[0] < 1.0:\n   #          continue\n    #     X = df_test.drop(targets + ['sample_id'], axis = 1).to_numpy()\n     #    y = df_test[target]\n      #   est = model[1]\n    #     score = est.score(X, y)\n    #     if score < 1.0:\n    #         print(f'{target} train score {model[0]} test score {score}')\n  # \n \n","metadata":{"execution":{"iopub.status.busy":"2024-04-27T04:46:37.850549Z","iopub.status.idle":"2024-04-27T04:46:37.850936Z","shell.execute_reply.started":"2024-04-27T04:46:37.850735Z","shell.execute_reply":"2024-04-27T04:46:37.85075Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nsample_submission = pd.read_csv(\"/kaggle/input/leap-atmospheric-physics-ai-climsim/sample_submission.csv\")\n","metadata":{"execution":{"iopub.status.busy":"2024-04-27T04:46:46.926995Z","iopub.execute_input":"2024-04-27T04:46:46.927721Z","iopub.status.idle":"2024-04-27T04:47:35.83351Z","shell.execute_reply.started":"2024-04-27T04:46:46.927683Z","shell.execute_reply":"2024-04-27T04:47:35.831223Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\ndf_test = pd.read_csv(\"/kaggle/input/leap-atmospheric-physics-ai-climsim/test.csv\")\n","metadata":{"execution":{"iopub.status.busy":"2024-04-27T04:48:30.249311Z","iopub.execute_input":"2024-04-27T04:48:30.249795Z","iopub.status.idle":"2024-04-27T04:50:39.738262Z","shell.execute_reply.started":"2024-04-27T04:48:30.24976Z","shell.execute_reply":"2024-04-27T04:50:39.736962Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sum = 0.0\nfor target, val in lr_models.items():\n    sum += val[0] if val[0] == 1.0 else 0.0\n    \nsum\n","metadata":{"execution":{"iopub.status.busy":"2024-04-27T04:50:39.741146Z","iopub.execute_input":"2024-04-27T04:50:39.741637Z","iopub.status.idle":"2024-04-27T04:50:39.752953Z","shell.execute_reply.started":"2024-04-27T04:50:39.741595Z","shell.execute_reply":"2024-04-27T04:50:39.751788Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sum = 0.0\nfor target, val in lr_models.items():\n    sum += 1.0 if val[0] > 0.999 else 0.0\n    \nsum\n","metadata":{"execution":{"iopub.status.busy":"2024-04-27T04:51:06.21232Z","iopub.execute_input":"2024-04-27T04:51:06.213286Z","iopub.status.idle":"2024-04-27T04:51:06.221708Z","shell.execute_reply.started":"2024-04-27T04:51:06.213246Z","shell.execute_reply":"2024-04-27T04:51:06.220418Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfor target, val in lr_models.items():\n    est = None\n    _est = 'M'\n    if val[0] > 0.05:\n        est = val[1]\n        _est = 'L'\n    if target in sr_models:\n        v2 = sr_models[target]\n        if v2[0] > val[0] and v2[0] > 0.05:\n            est = v2[1]\n            _est = 'S'\n    \n    if est is None:   \n        sample_submission[target] = sample_submission[target]*val[2]\n    else:\n        X = df_test.drop('sample_id', axis=1)\n        preds = est.predict(X)\n        sample_submission[target] = sample_submission[target]*preds\n    print(f'{target} done. {_est}')\n","metadata":{"execution":{"iopub.status.busy":"2024-04-27T04:51:07.055477Z","iopub.execute_input":"2024-04-27T04:51:07.055959Z","iopub.status.idle":"2024-04-27T04:58:07.799669Z","shell.execute_reply.started":"2024-04-27T04:51:07.055924Z","shell.execute_reply":"2024-04-27T04:58:07.798227Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nsample_submission.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2024-04-27T05:05:59.279479Z","iopub.execute_input":"2024-04-27T05:05:59.279953Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_submission.head()\n","metadata":{"execution":{"iopub.status.busy":"2024-04-27T05:03:02.641303Z","iopub.execute_input":"2024-04-27T05:03:02.641715Z","iopub.status.idle":"2024-04-27T05:03:02.930383Z","shell.execute_reply.started":"2024-04-27T05:03:02.641685Z","shell.execute_reply":"2024-04-27T05:03:02.92908Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}