{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":56537,"databundleVersionId":8015876,"sourceType":"competition"}],"dockerImageVersionId":30698,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Relationships between Input and Output Variables in E3SM-MMF Climate Model\n\nUnderstanding the relationships between input and output columns in the context of emulating small-scale atmospheric processes within the E3SM-MMF climate model involves grasping how changes in the state variables (inputs) influence the tendencies or changes in atmospheric properties (outputs) over time. Let's explore the potential relationships between these variables:\n\n## Air Temperature (State_t) and Heating Tendency (Ptend_t):\nHigher air temperatures typically lead to increased heating tendencies, as warmer air has a greater capacity to absorb solar radiation and transfer heat to the surrounding environment. Conversely, cooler temperatures may result in decreased heating tendencies.\n\n## Specific Humidity (State_q0001) and Moistening Tendency (Ptend_q0001):\nHigher specific humidity levels indicate greater moisture content in the atmosphere, which can lead to increased moistening tendencies. Conversely, lower specific humidity levels may result in decreased moistening tendencies.\n\n## Cloud Liquid/Ice Mixing Ratios (State_q0002/q0003) and Cloud Formation/Change Tendencies (Ptend_q0002/q0003):\nHigher mixing ratios of cloud liquid and ice indicate the presence of clouds in the atmosphere. Changes in these mixing ratios over time reflect processes such as cloud formation, growth, or dissipation. Positive tendencies indicate increasing cloud content, while negative tendencies suggest cloud dissipation.\n\n## Wind Speeds (State_u, State_v) and Wind Acceleration Tendencies (Ptend_u, Ptend_v):\nChanges in wind speeds over time can result from various atmospheric dynamics, including pressure gradients, temperature gradients, and surface friction. Acceleration tendencies represent changes in wind speeds with respect to time, reflecting the influence of atmospheric forces on wind motion.\n\n## Surface Pressure (State_ps) and Atmospheric Dynamics:\nSurface pressure variations influence atmospheric circulation patterns and weather systems. Changes in surface pressure may affect wind patterns, temperature distributions, and precipitation regimes.\n\n## Surface Fluxes (Solar Insolation, Latent Heat Flux, Sensible Heat Flux) and Surface Energy Budget:\nSurface fluxes play a crucial role in the exchange of energy between the Earth's surface and the atmosphere. Changes in solar insolation, latent heat flux (from evaporation), and sensible heat flux (from temperature gradients) influence the surface energy budget, which in turn impacts atmospheric dynamics and climate.\n\n## Atmospheric Gases Mixing Ratios (Ozone, Methane, Nitrous Oxide) and Radiative Forcing:\nChanges in the concentrations of greenhouse gases such as ozone, methane, and nitrous oxide can affect the Earth's radiative balance and contribute to climate change. Alterations in these mixing ratios may lead to changes in radiative forcing and atmospheric temperatures over time.\n\nUnderstanding these relationships allows for the development of machine learning models that can effectively capture the complex interactions between input state variables and output tendencies, enabling accurate emulation of small-scale atmospheric processes within the E3SM-MMF climate model.\n","metadata":{}},{"cell_type":"markdown","source":"| Input Variables (State Variables)              | Symbol       | Output Variables (Tendencies)                     | Symbol          |\n|-----------------------------------------------|--------------|----------------------------------------------------|-----------------|\n| Air Temperature                               | state_t      | Heating Tendency                                   | ptend_t         |\n| Specific Humidity                             | state_q0001 | Moistening Tendency                                | ptend_q0001    |\n| Cloud Liquid Mixing Ratio                     | state_q0002 | Cloud Liquid Mixing Ratio Change                   | ptend_q0002    |\n| Cloud Ice Mixing Ratio                        | state_q0003 | Cloud Ice Mixing Ratio Change                      | ptend_q0003    |\n| Zonal Wind Speed                              | state_u      | Zonal Wind Acceleration                            | ptend_u         |\n| Meridional Wind Speed                         | state_v      | Meridional Wind Acceleration                       | ptend_v         |\n| Surface Pressure                              | state_ps     | Net Shortwave Flux at Surface                      | cam_out_NETSW  |\n| Solar Insolation                              | pbuf_SOLIN  | Downward Longwave Flux at Surface                  | cam_out_FLWDS  |\n| Surface Latent Heat Flux                      | pbuf_LHFLX  | Snow Rate (Liquid Water Equivalent)                | cam_out_PRECSC |\n| Surface Sensible Heat Flux                    | pbuf_SHFLX  | Rain Rate                                          | cam_out_PRECC  |\n| Zonal Surface Stress                          | pbuf_TAUX    | Downward Visible Direct Solar Flux to Surface      | cam_out_SOLS   |\n| Meridional Surface Stress                     | pbuf_TAUY    | Downward Near-Infrared Direct Solar Flux to Surface| cam_out_SOLL   |\n| Cosine of Solar Zenith Angle                  | pbuf_COSZRS  | Downward Diffuse Solar Flux to Surface             | cam_out_SOLSD  |\n| Albedo for Diffuse Longwave Radiation         | cam_in_ALDIF | Downward Diffuse Near-Infrared Solar Flux to Surface| cam_out_SOLLD  |\n| Albedo for Direct Longwave Radiation          | cam_in_ALDIR |                                                    |                 |\n| Albedo for Diffuse Shortwave Radiation        | cam_in_ASDIF |                                                    |                 |\n| Albedo for Direct Shortwave Radiation         | cam_in_ASDIR |                                                    |                 |\n| Upward Longwave Flux                          | cam_in_LWUP  |                                                    |                 |\n| Sea-Ice Areal Fraction                        | cam_in_ICEFRAC |                                                  |                 |\n| Land Areal Fraction                           | cam_in_LANDFRAC |                                                 |                 |\n| Ocean Areal Fraction                          | cam_in_OCNFRAC |                                                  |                 |\n| Snow Depth over Land                          | cam_in_SNOWHLAND |                                                 |                 |\n| Ozone Volume Mixing Ratio                     | pbuf_ozone   |                                                  |                 |\n| Methane Volume Mixing Ratio                   | pbuf_CH4     |                                                  |                 |\n| Nitrous Oxide Volume Mixing Ratio             | pbuf_N2O     |                                                  |                 |\n","metadata":{}},{"cell_type":"markdown","source":"\n\n## Input Variables (State Variables)\n- **State_t (Air Temperature):** Represents air temperature at different vertical levels within the atmosphere (60 levels). Measured in Kelvin (K).\n- **State_q0001 (Specific Humidity):** Indicates specific humidity at different vertical levels (60 levels). Measured in kilograms of water vapor per kilogram of air (kg/kg).\n- **State_q0002 (Cloud Liquid Mixing Ratio):** Represents the mixing ratio of cloud liquid water content to total air mass at different vertical levels (60 levels). Measured in kg/kg.\n- **State_q0003 (Cloud Ice Mixing Ratio):** Similar to cloud liquid mixing ratio but for ice clouds. Measured in kg/kg.\n- **State_u (Zonal Wind Speed):** Represents the zonal (east-west) component of wind speed at different vertical levels (60 levels). Measured in meters per second (m/s).\n- **State_v (Meridional Wind Speed):** Represents the meridional (north-south) component of wind speed at different vertical levels (60 levels). Measured in meters per second (m/s).\n- **State_ps (Surface Pressure):** Indicates surface pressure at the Earth's surface. Measured in Pascals (Pa).\n- **Other state variables:** Various surface fluxes, angles, albedo, upward longwave flux, sea-ice fraction, land fraction, ocean fraction, snow depth over land, and atmospheric gases mixing ratios.\n\n## Output Variables (Target Variables)\n- **Ptend_t (Heating Tendency):** Represents the rate of change of air temperature with respect to time at different vertical levels (60 levels). Measured in Kelvin per second (K/s).\n- **Ptend_q0001 (Moistening Tendency):** Indicates the rate of change of specific humidity with respect to time at different vertical levels (60 levels). Measured in kg/kg per second (kg/kg/s).\n- **Ptend_q0002 and Ptend_q0003:** Represent the rate of change of cloud liquid and ice mixing ratios with respect to time, respectively, at different vertical levels (60 levels). Measured in kg/kg per second (kg/kg/s).\n- **Ptend_u and Ptend_v:** Represent the acceleration of zonal and meridional wind speeds with respect to time, respectively, at different vertical levels (60 levels). Measured in meters per second squared (m/s^2).\n- **Other output variables:** Net shortwave flux, downward longwave flux, precipitation rates (snow and rain), downward solar fluxes (visible and near-infrared, direct and diffuse), and other relevant atmospheric fluxes.\n\n\n","metadata":{}},{"cell_type":"markdown","source":"## Important\n\nFor the given dataset, you have to predict the output columns for each level of the input columns. Since many of the input and output variables are vertically resolved, meaning they have different values at different vertical levels within the atmosphere, your machine learning model needs to predict the corresponding tendencies or changes in these variables at each vertical level.","metadata":{}},{"cell_type":"code","source":"import polars as pl","metadata":{"execution":{"iopub.status.busy":"2024-04-25T06:06:54.795693Z","iopub.execute_input":"2024-04-25T06:06:54.796084Z","iopub.status.idle":"2024-04-25T06:06:55.144392Z","shell.execute_reply.started":"2024-04-25T06:06:54.796018Z","shell.execute_reply":"2024-04-25T06:06:55.143312Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"NUM_TRAIN_SAMPLES = 10_000","metadata":{"execution":{"iopub.status.busy":"2024-04-25T06:06:56.674176Z","iopub.execute_input":"2024-04-25T06:06:56.674549Z","iopub.status.idle":"2024-04-25T06:06:56.679627Z","shell.execute_reply.started":"2024-04-25T06:06:56.67452Z","shell.execute_reply":"2024-04-25T06:06:56.678114Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"lf_train = pl.scan_csv(\n    '/kaggle/input/leap-atmospheric-physics-ai-climsim/train.csv',\n).head(NUM_TRAIN_SAMPLES).drop('sample_id')","metadata":{"execution":{"iopub.status.busy":"2024-04-25T06:06:58.55721Z","iopub.execute_input":"2024-04-25T06:06:58.557569Z","iopub.status.idle":"2024-04-25T06:06:58.647302Z","shell.execute_reply.started":"2024-04-25T06:06:58.557543Z","shell.execute_reply":"2024-04-25T06:06:58.646118Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"lf_train.columns","metadata":{"execution":{"iopub.status.busy":"2024-04-25T06:07:13.639269Z","iopub.execute_input":"2024-04-25T06:07:13.639651Z","iopub.status.idle":"2024-04-25T06:07:13.660906Z","shell.execute_reply.started":"2024-04-25T06:07:13.639619Z","shell.execute_reply":"2024-04-25T06:07:13.659694Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Get the input column names\ninput_columns = ['state_t', 'state_q0001', 'state_q0002', 'state_q0003', 'state_u', 'state_v', 'state_ps',\n                 'pbuf_SOLIN', 'pbuf_LHFLX', 'pbuf_SHFLX', 'pbuf_TAUX', 'pbuf_TAUY', 'pbuf_COSZRS',\n                 'cam_in_ALDIF', 'cam_in_ALDIR', 'cam_in_ASDIF', 'cam_in_ASDIR', 'cam_in_LWUP',\n                 'cam_in_ICEFRAC', 'cam_in_LANDFRAC', 'cam_in_OCNFRAC', 'cam_in_SNOWHLAND',\n                 'pbuf_ozone', 'pbuf_CH4', 'pbuf_N2O']\n\n# Dictionary to store columns related to each input column\nrelated_columns_dict = {}\n\n# Loop through each input column\nfor input_col in input_columns:\n    # Find columns related to the current input column\n    related_columns = [col for col in lf_train.columns if input_col in col]\n    # Add the related columns to the dictionary\n    related_columns_dict[input_col] = related_columns\n\ntotal_no_of_distinct = 0\n# Display the dictionary\nfor input_col, related_columns in related_columns_dict.items():\n    print(f\"{input_col}: {len(related_columns)}\")\n    total_no_of_distinct += len(related_columns)\n    \nprint(f\"Total no of distinct columns in test dataset : {total_no_of_distinct}\")","metadata":{"execution":{"iopub.status.busy":"2024-04-24T12:29:46.199485Z","iopub.execute_input":"2024-04-24T12:29:46.199913Z","iopub.status.idle":"2024-04-24T12:29:46.214292Z","shell.execute_reply.started":"2024-04-24T12:29:46.199885Z","shell.execute_reply":"2024-04-24T12:29:46.21325Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"lf_test = pl.scan_csv(\n    '/kaggle/input/leap-atmospheric-physics-ai-climsim/train.csv',\n).head(NUM_TRAIN_SAMPLES).drop('sample_id')","metadata":{"execution":{"iopub.status.busy":"2024-04-24T12:27:12.725953Z","iopub.execute_input":"2024-04-24T12:27:12.726425Z","iopub.status.idle":"2024-04-24T12:27:12.752688Z","shell.execute_reply.started":"2024-04-24T12:27:12.726391Z","shell.execute_reply":"2024-04-24T12:27:12.751525Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Get the input column names\ninput_columns = ['state_t', 'state_q0001', 'state_q0002', 'state_q0003', 'state_u', 'state_v', 'state_ps',\n                 'pbuf_SOLIN', 'pbuf_LHFLX', 'pbuf_SHFLX', 'pbuf_TAUX', 'pbuf_TAUY', 'pbuf_COSZRS',\n                 'cam_in_ALDIF', 'cam_in_ALDIR', 'cam_in_ASDIF', 'cam_in_ASDIR', 'cam_in_LWUP',\n                 'cam_in_ICEFRAC', 'cam_in_LANDFRAC', 'cam_in_OCNFRAC', 'cam_in_SNOWHLAND',\n                 'pbuf_ozone', 'pbuf_CH4', 'pbuf_N2O']\n\n# Dictionary to store columns related to each input column\nrelated_columns_dict = {}\n\n# Loop through each input column\nfor input_col in input_columns:\n    # Find columns related to the current input column\n    related_columns = [col for col in lf_test.columns if input_col in col]\n    # Add the related columns to the dictionary\n    related_columns_dict[input_col] = related_columns\n\ntotal_no_of_distinct = 0\n# Display the dictionary\nfor input_col, related_columns in related_columns_dict.items():\n    print(f\"{input_col}: {len(related_columns)}\")\n    total_no_of_distinct += len(related_columns)\n\nprint(f\"Total no of distinct columns in test dataset : {total_no_of_distinct}\")","metadata":{"execution":{"iopub.status.busy":"2024-04-24T12:29:52.927397Z","iopub.execute_input":"2024-04-24T12:29:52.927801Z","iopub.status.idle":"2024-04-24T12:29:52.941295Z","shell.execute_reply.started":"2024-04-24T12:29:52.927773Z","shell.execute_reply":"2024-04-24T12:29:52.940074Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_submission = pl.scan_csv(\n    '/kaggle/input/leap-atmospheric-physics-ai-climsim/sample_submission.csv',\n).head(NUM_TRAIN_SAMPLES).drop('sample_id')","metadata":{"execution":{"iopub.status.busy":"2024-04-24T12:32:15.122764Z","iopub.execute_input":"2024-04-24T12:32:15.123186Z","iopub.status.idle":"2024-04-24T12:32:15.147977Z","shell.execute_reply.started":"2024-04-24T12:32:15.123159Z","shell.execute_reply":"2024-04-24T12:32:15.146616Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Dictionary mapping each output column to its related columns\noutput_columns_dict = {\n    \"ptend_t\": [],\n    \"ptend_q0001\": [],\n    \"ptend_q0002\": [],\n    \"ptend_q0003\": [],\n    \"ptend_u\": [],\n    \"ptend_v\": [],\n    \"cam_out_NETSW\": [],\n    \"cam_out_FLWDS\": [],\n    \"cam_out_PRECSC\": [],\n    \"cam_out_PRECC\": [],\n    \"cam_out_SOLS\": [],\n    \"cam_out_SOLL\": [],\n    \"cam_out_SOLSD\": [],\n    \"cam_out_SOLLD\": []\n}\n\n# Get the related columns for each output column\nfor output_col in output_columns_dict.keys():\n    related_columns = [col for col in sample_submission.columns if output_col in col]\n    output_columns_dict[output_col] = related_columns\n\n# Display the dictionary\nfor output_col, related_columns in output_columns_dict.items():\n    print(f\"{output_col}: {len(related_columns)}\")","metadata":{"execution":{"iopub.status.busy":"2024-04-24T12:35:07.177001Z","iopub.execute_input":"2024-04-24T12:35:07.177505Z","iopub.status.idle":"2024-04-24T12:35:07.188832Z","shell.execute_reply.started":"2024-04-24T12:35:07.177471Z","shell.execute_reply":"2024-04-24T12:35:07.187562Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_schema = pl.DataFrame(\n    list(lf_train.schema.items()),\n    schema = [('feature', pl.String), ('dtype', pl.Object)],\n)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display(df_schema)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display(df_schema.select('dtype').to_series().value_counts())","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"FEATURE_GROUPS_60 = [\n    'state_t',\n    'state_q0001',\n    'state_q0002',\n    'state_q0003',\n    'state_u',\n    'state_v',\n    'pbuf_ozone',\n    'pbuf_CH4',\n    'pbuf_N2O',\n]\nFEATURE_GROUPS_1 = [\n    'state_ps',\n    'pbuf_SOLIN',\n    'pbuf_LHFLX',\n    'pbuf_SHFLX',\n    'pbuf_TAUX',\n    'pbuf_TAUY',\n    'pbuf_COSZRS',\n    'cam_in_ALDIF',\n    'cam_in_ALDIR',\n    'cam_in_ASDIF',\n    'cam_in_ASDIR',\n    'cam_in_LWUP',\n    'cam_in_ICEFRAC',\n    'cam_in_LANDFRAC',\n    'cam_in_OCNFRAC',\n    'cam_in_SNOWHLAND',\n]\nTARGET_GROUPS_60 = [\n    'ptend_t',\n    'ptend_q0001',\n    'ptend_q0002',\n    'ptend_q0003',\n    'ptend_u',\n    'ptend_v',\n]\nTARGET_GROUPS_1 = [\n    'cam_out_NETSW',\n    'cam_out_FLWDS',\n    'cam_out_PRECSC',\n    'cam_out_PRECC',\n    'cam_out_SOLS',\n    'cam_out_SOLL',\n    'cam_out_SOLSD',\n    'cam_out_SOLLD',\n]","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display(lf_train.select(pl.concat_list(pl.all().null_count()).list.sum().alias('total_null_count')).collect())","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Exlporation","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns\nimport polars.selectors as cs\nimport warnings","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display(lf_train.select(cs.starts_with('state_t')).describe())","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def plot_feature_60(lf, feature):\n    _, axes = plt.subplots(10, 6, figsize = (30, 36), layout = 'tight')\n    with warnings.catch_warnings():\n        warnings.simplefilter('ignore')\n        for i in range(10):\n            for j in range(6):\n                n = 6 * i + j\n                sns.histplot(x = lf_train.select(f'{feature}_{n}').collect().to_series().to_pandas(),\n                             ax = axes[i, j],\n                             bins = 'sturges',\n                             kde = True)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def plot_corr_60(lf, feature):\n    _, ax = plt.subplots(figsize = (36, 32))\n    sns.heatmap(lf_train.select(cs.starts_with(feature)).collect().to_pandas().corr(),\n                ax = ax,\n                vmin = -1, vmax = 1, cmap = 'coolwarm', annot = True)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_feature_60(lf_train, 'state_t')","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_corr_60(lf_train, 'state_t')","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display(lf_train.select(FEATURE_GROUPS_1).describe())","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"n = len(FEATURE_GROUPS_1)\n_, axes = plt.subplots(n, figsize = (6, 3 * n), layout = 'tight')\nwith warnings.catch_warnings():\n    warnings.simplefilter('ignore')\n    for ax, c in zip(axes, FEATURE_GROUPS_1):\n        sns.histplot(x = lf_train.select(c).collect().to_series().to_pandas(),\n                     ax = ax,\n                     bins = 'sturges',\n                     kde = True)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display(lf_train.select(TARGET_GROUPS_1).describe())","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"n = len(TARGET_GROUPS_1)\n_, axes = plt.subplots(n, figsize = (6, 3 * n), layout = 'tight')\nwith warnings.catch_warnings():\n    warnings.simplefilter('ignore')\n    for ax, c in zip(axes, TARGET_GROUPS_1):\n        sns.histplot(x = lf_train.select(c).collect().to_series().to_pandas(),\n                     ax = ax,\n                     bins = 'sturges', kde = True)","metadata":{},"execution_count":null,"outputs":[]}]}