{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":56537,"databundleVersionId":8877088,"sourceType":"competition"}],"dockerImageVersionId":30732,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"Inspired from Chris x.  \nLearning notebook.  \nReference to CHRIS X EDA link: https://www.kaggle.com/code/docxian/leap-climsim-visual-eda","metadata":{}},{"cell_type":"markdown","source":"## Data Description\n### Inputs:\n#### Arrays (dimension 60):\n- state_t : air temperature\n- state_q0001 : specific humidity\n- state_q0002 : cloud liquid mixing ratio\n- state_q0003 : cloud ice mixing ratio\n- state_u : zonal wind speed\n- state_v : meridional wind speed\n- pbuf_ozone : ozone volume mixing ratio\n- pbuf_CH4 : methane volume mixing ratio\n- pbuf_N2O : nitros oxide volume mixing ratio\n\n#### Scalars:\n- state_ps : surface pressure\n- pbuf_SOLIN : solar insolation\n- pbuf_LHFLX : surface latent heat flux\n- pbuf_SHFLX : surface sensible heat flux\n- pbuf_TAUX : zonal surface stress\n- pbuf_TAUY : meridional surface stress\n- pbuf_COSZRS: cosine of solar zenith angle\n- cam_in_ALDIF: albedo for diffuse longwave radiation\n- cam_in_ALDIR: albedo for direct longwave radiation\n- cam_in_ASDIF: albedo for diffuse shortwave radiation\n- cam_in_ASDIR: albedo for direct shortwave radiation\n- cam_in_LWUP: upward longwave flux\n- cam_in_ICEFRAC: sea-ice areal fraction\n- cam_in_LANDFRAC: land areal fraction\n- cam_in_OCNFRAC: ocean areal fraction\n- cam_in_SNOWHLAND: snow depth over land\n\n### Targets:\n#### Arrays (dimension 60):\n- ptend_t: heating tendency\n- ptend_q0001: moistening tendency\n- ptend_q0002: cloud liquid mixing ratio change over time\n- ptend_q0003: cloud ice mixing ratio change over time\n- ptend_u: zonal wind acceleration\n- ptend_v: meridional wind acceleration\n\n#### Scalars:\n- cam_out_NETSW: net shortwave flux at surface\n- cam_out_FLWDS: downward longwave flux at surface\n- cam_out_PRECSC: snow rate (liquid water equivalent)\n- cam_out_PRECC: rain rate\n- cam_out_SOLS: downward visible direct solar flux to surface\n- cam_out_SOLL: downward near-infrared direct solar flux to surface\n- cam_out_SOLSD: downward diffuse solar flux to surface\n- cam_out_SOLLD: downward diffuse near-infrared solar flux to surface","metadata":{}},{"cell_type":"markdown","source":"### Init","metadata":{}},{"cell_type":"code","source":"print(\"start\")","metadata":{"execution":{"iopub.status.busy":"2024-06-25T08:17:39.748225Z","iopub.execute_input":"2024-06-25T08:17:39.748775Z","iopub.status.idle":"2024-06-25T08:17:39.756741Z","shell.execute_reply.started":"2024-06-25T08:17:39.748736Z","shell.execute_reply":"2024-06-25T08:17:39.755168Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# packages\n\n# standard\nimport numpy as np\nimport pandas as pd\nimport time\n\n# plots\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n# faster alternative to pandas\nimport polars as pl","metadata":{"execution":{"iopub.status.busy":"2024-06-25T08:17:39.759578Z","iopub.execute_input":"2024-06-25T08:17:39.760002Z","iopub.status.idle":"2024-06-25T08:17:39.771545Z","shell.execute_reply.started":"2024-06-25T08:17:39.75997Z","shell.execute_reply":"2024-06-25T08:17:39.769978Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# configs\n\n# to display all columns\npd.set_option('display.max_columns', None)\n\n# aesthetics\ndefault_color_1 = 'darkblue'\ndefault_color_2 = 'darkgreen'\ndefault_color_3 = 'darkred'\n\n# random seed\nmy_random_seed = 42","metadata":{"execution":{"iopub.status.busy":"2024-06-25T08:17:39.773482Z","iopub.execute_input":"2024-06-25T08:17:39.774247Z","iopub.status.idle":"2024-06-25T08:17:39.792051Z","shell.execute_reply.started":"2024-06-25T08:17:39.774193Z","shell.execute_reply":"2024-06-25T08:17:39.790455Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Import and Overview (subset)","metadata":{}},{"cell_type":"code","source":"!ls -l '../input/leap-atmospheric-physics-ai-climsim/'","metadata":{"execution":{"iopub.status.busy":"2024-06-25T08:17:39.793667Z","iopub.execute_input":"2024-06-25T08:17:39.79415Z","iopub.status.idle":"2024-06-25T08:17:41.249929Z","shell.execute_reply.started":"2024-06-25T08:17:39.794107Z","shell.execute_reply":"2024-06-25T08:17:41.24823Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Huge dataset, starting with small subset**","metadata":{}},{"cell_type":"code","source":"# import SUBSET of data\n\nn_rows = 100000 # 1 Lakh\nfolder = 'leap-atmospheric-physics-ai-climsim'\nt1 = time.time()\ndf_train = pl.read_csv('../input/' + folder + '/train.csv', n_rows=n_rows).to_pandas()\ndf_test = pl.read_csv('../input/' + folder + '/test_old.csv', n_rows=n_rows).to_pandas()\ndf_submission = pl.read_csv('../input/' + folder + '/sample_submission_old.csv', n_rows=n_rows).to_pandas()\nt2 = time.time()\nprint('Elapsed time [s]: ', np.round(t2-t1, 2))","metadata":{"execution":{"iopub.status.busy":"2024-06-25T08:17:41.255518Z","iopub.execute_input":"2024-06-25T08:17:41.256075Z","iopub.status.idle":"2024-06-25T08:17:52.285892Z","shell.execute_reply.started":"2024-06-25T08:17:41.25603Z","shell.execute_reply":"2024-06-25T08:17:52.284383Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.head(10)","metadata":{"execution":{"iopub.status.busy":"2024-06-25T08:17:52.287547Z","iopub.execute_input":"2024-06-25T08:17:52.287918Z","iopub.status.idle":"2024-06-25T08:17:53.70215Z","shell.execute_reply.started":"2024-06-25T08:17:52.287889Z","shell.execute_reply":"2024-06-25T08:17:53.70063Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.info(verbose=True, show_counts=True)","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-06-25T08:17:53.704042Z","iopub.execute_input":"2024-06-25T08:17:53.704528Z","iopub.status.idle":"2024-06-25T08:17:53.999162Z","shell.execute_reply.started":"2024-06-25T08:17:53.704484Z","shell.execute_reply":"2024-06-25T08:17:53.997854Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# preview - test\ndf_test.head(10)","metadata":{"execution":{"iopub.status.busy":"2024-06-25T08:17:54.000882Z","iopub.execute_input":"2024-06-25T08:17:54.001364Z","iopub.status.idle":"2024-06-25T08:17:54.795221Z","shell.execute_reply.started":"2024-06-25T08:17:54.00132Z","shell.execute_reply":"2024-06-25T08:17:54.793651Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# test set overview\ndf_test.info(verbose=True, show_counts=True)","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-06-25T08:17:54.797376Z","iopub.execute_input":"2024-06-25T08:17:54.798104Z","iopub.status.idle":"2024-06-25T08:17:55.013306Z","shell.execute_reply.started":"2024-06-25T08:17:54.798049Z","shell.execute_reply":"2024-06-25T08:17:55.011757Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Targets and Features","metadata":{}},{"cell_type":"code","source":"# targets (extract from submission file)\ntargets = [x for x in df_submission.columns.tolist() if x not in ['sample_id']]\n\n# numerical features\nfeatures_numerical = [x for x in df_train.columns.tolist() if x not in ['sample_id']+targets]\n\n# categorical features\nfeatures_categorical = []\n\n# all features combined\nfeatures_all = features_numerical + features_categorical","metadata":{"execution":{"iopub.status.busy":"2024-06-25T08:17:55.014676Z","iopub.execute_input":"2024-06-25T08:17:55.015002Z","iopub.status.idle":"2024-06-25T08:17:55.03042Z","shell.execute_reply.started":"2024-06-25T08:17:55.014973Z","shell.execute_reply":"2024-06-25T08:17:55.029156Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# output of dimensions\nprint('Number of numerical features: ', len(features_numerical))\nprint('Number of categorical features: ', len(features_categorical))\nprint('Number of targets: ', len(targets))\nprint('Size of subset: ', n_rows)","metadata":{"execution":{"iopub.status.busy":"2024-06-25T08:17:55.032155Z","iopub.execute_input":"2024-06-25T08:17:55.032586Z","iopub.status.idle":"2024-06-25T08:17:55.044182Z","shell.execute_reply.started":"2024-06-25T08:17:55.032527Z","shell.execute_reply":"2024-06-25T08:17:55.042876Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Targets","metadata":{}},{"cell_type":"code","source":"# basic stats - targets\ndf_train[targets].describe()","metadata":{"execution":{"iopub.status.busy":"2024-06-25T08:17:55.046082Z","iopub.execute_input":"2024-06-25T08:17:55.046639Z","iopub.status.idle":"2024-06-25T08:17:58.289863Z","shell.execute_reply.started":"2024-06-25T08:17:55.046596Z","shell.execute_reply":"2024-06-25T08:17:58.288642Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# plot target distributions in compact matrix form\nfig, axs = plt.subplots(92, 4, figsize=(16, 350))\ni = 0\nfor t in targets:\n    current_ax = axs.flat[i]\n    current_ax.hist(df_train[t], bins=100, color=default_color_3)\n    current_ax.set_title('Target' + str(t))\n    current_ax.grid()\n    i += 1","metadata":{"execution":{"iopub.status.busy":"2024-06-25T08:17:58.291317Z","iopub.execute_input":"2024-06-25T08:17:58.291703Z","iopub.status.idle":"2024-06-25T08:20:43.37741Z","shell.execute_reply.started":"2024-06-25T08:17:58.29167Z","shell.execute_reply":"2024-06-25T08:20:43.37526Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Features","metadata":{}},{"cell_type":"code","source":"# basic stats - train\ndf_train[features_numerical].describe()","metadata":{"execution":{"iopub.status.busy":"2024-06-25T08:20:43.384614Z","iopub.execute_input":"2024-06-25T08:20:43.385136Z","iopub.status.idle":"2024-06-25T08:20:48.158945Z","shell.execute_reply.started":"2024-06-25T08:20:43.385095Z","shell.execute_reply":"2024-06-25T08:20:48.15758Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# basic stats - test\ndf_test[features_numerical].describe()","metadata":{"execution":{"iopub.status.busy":"2024-06-25T08:20:48.160655Z","iopub.execute_input":"2024-06-25T08:20:48.161088Z","iopub.status.idle":"2024-06-25T08:20:52.585037Z","shell.execute_reply.started":"2024-06-25T08:20:48.161049Z","shell.execute_reply":"2024-06-25T08:20:52.583744Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# plot histograms for numerical features (train and test)\nfor f in features_numerical:\n    plt.figure(figsize=(12,2))\n    ax1 = plt.subplot(1,2,1)\n    df_train[f].plot(kind='hist', bins=100, color=default_color_1)\n    plt.title(f + ' - Train')\n    plt.grid()\n    ax2 = plt.subplot(1,2,2, sharex=ax1)\n    df_test[f].plot(kind='hist', bins=100, color=default_color_2)\n    plt.title(f + ' - Test')\n    plt.grid()\n#     plt.show()","metadata":{"execution":{"iopub.status.busy":"2024-06-25T08:20:52.586885Z","iopub.execute_input":"2024-06-25T08:20:52.587379Z","iopub.status.idle":"2024-06-25T08:29:22.412517Z","shell.execute_reply.started":"2024-06-25T08:20:52.587327Z","shell.execute_reply":"2024-06-25T08:29:22.410769Z"},"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# compact boxplot of all features - train only\nn_plot_rows = 10\nn_plot_cols = 60\nn = len(features_numerical)\nfor i in range(n_plot_rows):\n    a = n_plot_cols*i+1\n    b = min(n_plot_cols*i+n_plot_cols, n)\n    print('Columns', a, 'to', b)\n    df_train.iloc[:,a:(b+1)].plot(kind='box', figsize=(15,5))\n    plt.xticks(rotation=90)\n    plt.grid()\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2024-06-25T08:29:22.414624Z","iopub.execute_input":"2024-06-25T08:29:22.415141Z","iopub.status.idle":"2024-06-25T08:29:42.682491Z","shell.execute_reply.started":"2024-06-25T08:29:22.415092Z","shell.execute_reply":"2024-06-25T08:29:42.681198Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for f in features_numerical:\n    plt.figure(figsize=(14,0.5))\n    ax1 = plt.subplot(1,2,1)\n    df_temp = df_train[f].dropna() # boxplot does not like missings...\n    plt.boxplot(df_temp, vert=False)\n    plt.title(f + ' - Train')\n    plt.grid()\n    ax2 = plt.subplot(1,2,2, sharex=ax1)\n    df_temp = df_test[f].dropna()\n    plt.boxplot(df_temp, vert=False)\n    plt.title(f + ' - Test')\n    plt.grid()\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2024-06-25T08:29:42.683987Z","iopub.execute_input":"2024-06-25T08:29:42.684341Z","iopub.status.idle":"2024-06-25T08:33:23.493829Z","shell.execute_reply.started":"2024-06-25T08:29:42.684297Z","shell.execute_reply":"2024-06-25T08:33:23.492369Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Correlations\n#### Targets","metadata":{}},{"cell_type":"code","source":"# calc and plot correlation matrix\ncor_p_target = df_train[targets].corr(method='pearson')\nplt.figure(figsize=(14,12))\nsns.heatmap(cor_p_target, annot=False, cmap='RdYlGn',\n            vmin=-1, vmax=+1)\nplt.title('Targets - Pearson Correlation')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-06-25T08:33:23.495805Z","iopub.execute_input":"2024-06-25T08:33:23.496312Z","iopub.status.idle":"2024-06-25T08:34:02.778761Z","shell.execute_reply.started":"2024-06-25T08:33:23.496274Z","shell.execute_reply":"2024-06-25T08:34:02.777006Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Features (train)","metadata":{}},{"cell_type":"code","source":"# calc and plot correlation matrix\ncor_p_train = df_train[features_numerical].corr(method='pearson')\nplt.figure(figsize=(14,12))\nsns.heatmap(cor_p_train, annot=False, cmap='RdYlGn',\n            vmin=-1, vmax=+1)\nplt.title('Features - Pearson Correlation (train)')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-06-25T08:34:02.780701Z","iopub.execute_input":"2024-06-25T08:34:02.781139Z","iopub.status.idle":"2024-06-25T08:35:31.019372Z","shell.execute_reply.started":"2024-06-25T08:34:02.781102Z","shell.execute_reply":"2024-06-25T08:35:31.01806Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Features (test)","metadata":{}},{"cell_type":"code","source":"# calc and plot correlation matrix\ncor_p_test = df_test[features_numerical].corr(method='pearson')\nplt.figure(figsize=(14,12))\nsns.heatmap(cor_p_test, annot=False, cmap='RdYlGn',\n            vmin=-1, vmax=+1)\nplt.title('Features - Pearson Correlation (test)')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-06-25T08:35:31.021396Z","iopub.execute_input":"2024-06-25T08:35:31.021885Z","iopub.status.idle":"2024-06-25T08:37:03.816869Z","shell.execute_reply.started":"2024-06-25T08:35:31.021845Z","shell.execute_reply":"2024-06-25T08:37:03.815423Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# # export results\n# df_train.to_csv('df_train_subset.csv')\n# cor_p_target.to_csv('cor_p_target.csv')\n# cor_p_train.to_csv('cor_p_train.csv')\n# cor_p_test.to_csv('cor_p_test.csv')","metadata":{"execution":{"iopub.status.busy":"2024-06-25T08:37:03.818879Z","iopub.execute_input":"2024-06-25T08:37:03.819353Z","iopub.status.idle":"2024-06-25T08:37:03.825193Z","shell.execute_reply.started":"2024-06-25T08:37:03.819312Z","shell.execute_reply":"2024-06-25T08:37:03.8238Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Other Explorations\n#### Target vs row index","metadata":{}},{"cell_type":"code","source":"# plot target values\nfor t in targets:\n    plt.figure(figsize=(14,2))\n    plt.scatter(df_train.index, df_train[t], color=default_color_3,\n                alpha=0.25, s=1)\n    plt.title(t)\n    plt.grid()\n#     plt.show()","metadata":{"execution":{"iopub.status.busy":"2024-06-25T08:37:03.82702Z","iopub.execute_input":"2024-06-25T08:37:03.827829Z","iopub.status.idle":"2024-06-25T08:37:03.842116Z","shell.execute_reply.started":"2024-06-25T08:37:03.827787Z","shell.execute_reply":"2024-06-25T08:37:03.84073Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Features vs row index","metadata":{}},{"cell_type":"code","source":"# plot feature values\nfor f in features_numerical:\n    plt.figure(figsize=(14,2))\n    plt.scatter(df_train.index, df_train[f], color=default_color_1,\n                alpha=0.25, s=1)\n    plt.title(f)\n    plt.grid()\n#     plt.show()","metadata":{"execution":{"iopub.status.busy":"2024-06-25T08:37:03.843524Z","iopub.execute_input":"2024-06-25T08:37:03.843908Z","iopub.status.idle":"2024-06-25T08:37:03.855487Z","shell.execute_reply.started":"2024-06-25T08:37:03.843878Z","shell.execute_reply":"2024-06-25T08:37:03.854053Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Import individual columns (full data)\nIn order to approach the full dataset we could try to import just a subset of columns. This is shown in the following section.","metadata":{}},{"cell_type":"code","source":"# define columns (has to be a list)\nn_max = 20 # columns with index 0..n_max\ncols_select = ['state_t_' + str(t) for t in range(0,n_max+1)]\nprint(cols_select)","metadata":{"execution":{"iopub.status.busy":"2024-06-25T08:37:03.85743Z","iopub.execute_input":"2024-06-25T08:37:03.85795Z","iopub.status.idle":"2024-06-25T08:37:03.87015Z","shell.execute_reply.started":"2024-06-25T08:37:03.857904Z","shell.execute_reply":"2024-06-25T08:37:03.868629Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# load only selected column\nt1 = time.time()\ndf_col = pl.read_csv('../input/'+folder+'/train.csv', columns=cols_select).to_pandas()\nt2 = time.time()\nprint('Elapsed time [s]: ', np.round(t2-t1,2))\nprint('Number of rows: ', df_col.shape[0])\nprint('Number of cols: ', df_col.shape[1])","metadata":{"execution":{"iopub.status.busy":"2024-06-25T08:37:03.872064Z","iopub.execute_input":"2024-06-25T08:37:03.872574Z","iopub.status.idle":"2024-06-25T08:47:57.448954Z","shell.execute_reply.started":"2024-06-25T08:37:03.872505Z","shell.execute_reply":"2024-06-25T08:47:57.447182Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# basic stats\ndf_col.describe(percentiles=[0.01,0.1,0.25,0.5,0.75,0.9,0.99])","metadata":{"execution":{"iopub.status.busy":"2024-06-25T08:47:57.451052Z","iopub.execute_input":"2024-06-25T08:47:57.451624Z","iopub.status.idle":"2024-06-25T08:48:09.377118Z","shell.execute_reply.started":"2024-06-25T08:47:57.451546Z","shell.execute_reply":"2024-06-25T08:48:09.375205Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# plot distributions\nfor f in cols_select:\n    plt.figure(figsize=(10,3))\n    plt.hist(df_col[f], bins=1000, color=default_color_1)\n    plt.title(f + ' - full data')\n    plt.grid()\n#     plt.show()","metadata":{"execution":{"iopub.status.busy":"2024-06-25T08:48:09.37911Z","iopub.execute_input":"2024-06-25T08:48:09.37974Z","iopub.status.idle":"2024-06-25T08:48:57.217371Z","shell.execute_reply.started":"2024-06-25T08:48:09.379687Z","shell.execute_reply":"2024-06-25T08:48:57.215962Z"},"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"💡 state_t_0 shows some unusually high values, let's have a closer look:\n","metadata":{}},{"cell_type":"code","source":"# boxplot\nplt.figure(figsize=(10,0.5))\nplt.boxplot(df_col.state_t_0, vert=False)\nplt.title('state_t_0')\nplt.grid()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-06-25T08:48:57.219084Z","iopub.execute_input":"2024-06-25T08:48:57.219488Z","iopub.status.idle":"2024-06-25T08:48:58.290339Z","shell.execute_reply.started":"2024-06-25T08:48:57.219453Z","shell.execute_reply":"2024-06-25T08:48:58.28894Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# let's check the most extreme outliers\ndf_col[df_col.state_t_0>400]","metadata":{"execution":{"iopub.status.busy":"2024-06-25T08:48:58.291841Z","iopub.execute_input":"2024-06-25T08:48:58.292189Z","iopub.status.idle":"2024-06-25T08:48:58.339043Z","shell.execute_reply.started":"2024-06-25T08:48:58.29216Z","shell.execute_reply":"2024-06-25T08:48:58.337762Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Correlations","metadata":{}},{"cell_type":"code","source":"# calc and plot correlation matrix\ncor_p_train_few_cols = df_train[cols_select].corr(method='pearson')\nplt.figure(figsize=(14,10))\nsns.heatmap(cor_p_train_few_cols, annot=True, cmap='RdYlGn',\n            fmt='.2f', linecolor='black', linewidth=.5,\n            vmin=-1, vmax=+1)\nplt.title('Features - Pearson Correlation (train/selected columns)')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-06-25T08:48:58.340808Z","iopub.execute_input":"2024-06-25T08:48:58.341234Z","iopub.status.idle":"2024-06-25T08:49:00.195031Z","shell.execute_reply.started":"2024-06-25T08:48:58.341199Z","shell.execute_reply":"2024-06-25T08:49:00.193671Z"},"trusted":true},"execution_count":null,"outputs":[]}]}