{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":56537,"databundleVersionId":8015876,"sourceType":"competition"}],"dockerImageVersionId":30698,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"Published on April 18, 2024. By Marília Prata, mpwolke","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nimport plotly.graph_objs as go\nimport plotly.offline as py\nimport plotly.express as px\n\n#Ignore warnings\nimport warnings\nwarnings.filterwarnings('ignore')\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-04-19T02:34:13.200317Z","iopub.execute_input":"2024-04-19T02:34:13.201358Z","iopub.status.idle":"2024-04-19T02:34:17.117315Z","shell.execute_reply.started":"2024-04-19T02:34:13.201318Z","shell.execute_reply":"2024-04-19T02:34:17.115642Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Learning the Earth with Artificial Intelligence and Physics (LEAP)\n\n![](https://encrypted-tbn0.gstatic.com/images?q=tbn:ANd9GcSPula2TcLHFlQL64XGTbeSV_dtywqkIEljVGES0iy0Pw&s)https://github.com/leap-stc/ClimSim","metadata":{}},{"cell_type":"markdown","source":"#ClimSim: The largest-ever dataset designed for hybrid ML-physics research \n\nClimSim: A large multi-scale dataset for hybrid physics-ML climate emulation\n\nAuthors: Sungduk Yu1, Walter M. Hannah,Liran Peng,Jerry Lin, Mohamed Aziz Bhouri, Ritwik Gupta, Björn Lütjens, Justus C. Will, Gunnar Behrens,Julius J. M. Busecke, Nora Loose, Charles Stern, Tom Beucler,\nBryce E. Harrop, Benjamin R. Hillman, Andrea M. Jenney,Savannah L. Ferretti, Nana Liu, Anima Anandkumar,Noah D. Brenowitz, Veronika Eyring,Nicholas Geneva, Pierre Gentine, Stephan Mandt, Jaideep Pathak, Akshay Subramaniam, Carl Vondrick, Rose Yu,Laure Zanna, Tian Zhen,Ryan P. Abernathey, Fiaz Ahmed, David C. Bader, Pierre Baldi, Elizabeth A. Barnes, Christopher S. Bretherton, Peter M. Caldwell,\nWayne Chuang, Yilun Han, Yu Huang, Fernando Iglesias-Suarez, Sanket Jantre, Karthik Kashinath, Marat Khairoutdinov,Thorsten Kurth, Nicholas J. Lutsko, Po-Lun Ma, Griffin Mooers, J. David Neelin, David A. Randall, Sara Shamekh, Mark A. Taylor, Nathan M. Urban, Janni Yuval, Guang J. Zhang, Michael S. Pritchard\n\n\"The authors presented ClimSim, the largest-ever dataset designed for hybrid ML-physics research. It comprises multi-scale climate simulations, developed by a consortium of climate scientists and ML researchers. It consists of 5.7 billion pairs of multivariate input and output vectors that isolate the influence of locally-nested, high-resolution, high-fidelity physics on a host climate simulator’s macro-scale physical state.\"\n\n\"ClimSim, the most physically comprehensive dataset yet published for training ML emulators of atmospheric storms, clouds, turbulence, rainfall, and radiation for use in hybrid-ML\nclimate simulation. It contains all inputs and outputs necessary for downstream coupling in a fullcomplexity multi-scale climate simulator. The authors conducted a series of experiments on a subset of these variables that demonstrate the degree to which climate data scientists have been able to fit their deterministic and stochastic components.\"\n\n\"They hope ML community engagement in ClimSim will advance fundamental ML methodology and\nclarify the path to producing increasingly skillful sub-grid physics emulators that can be reliably used for operational climate simulation. To facilitate two-way commications between ML practitioners\nand climate scientists, the authors incorporated many desired characteristics for an ideal benchmark dataset.\"\n\n\"Such interdisciplinary collaboration will open up an exciting future in which the computational limits that currently constrain climate simulation can be reconsidered. The authors plan to soon extend ClimSim to include, first, a sampling of multiple future climate states. Second, they aim to provide a protocol for downstream hybrid simulation testing. We hope lessons learned in their chosen limit of multi-scale atmospheric simulation will have applicability in other sub-fields of Earth System Science where computational constraints are currently a barrier to including explicit representations of more systems of nested complexity.\"\n\nPhysics-Informed Guidance to Improve Generalizability and Coupled Performance\n\n\"Physical Constraints: Mass and energy conservation are important criteria for Earth system simulation. If these terms are not conserved, errors in estimating sea level rise or temperature change over\ntime may become as large as the signals we hope to measure. Enforcing conservation on emulated\nresults helps constrain results to be physically plausible and reduce the potential for errors accumulating over long time scales. We discuss how to do this and enforce additional constraints, such as non-negativity for precipitation, condensate, and moisture variables in the Supporting Information.\"\n\n\"Stochasticity and Memory: The results of the embedded convection calculations regulating do\nare chaotic, and thus worthy of stochastic architectures, as in our RPN, HSR, and cVAE baselines.\nThese solutions are likewise sensitive to sub-grid initial state variables from an interior nested spatial dimension that has not been included in their data.\"\n\n\"Temporal Locality: Incorporating the previous timesteps’ target or feature in the input vector\ninflation could be beneficial as it captures some information about this convective memory and\nutilizes temporal autocorrelations present in atmospheric data.\"\n\n\"Causal Pruning: A systematic and quantitative pruning of the input vector based on objectively\nassessed causal relationships to subsets of the target vector has been proposed as an attractive\npreprocessing strategy, as it helps remove spurious correlations due to confounding variables and\noptimize the ML algorithm.\"\n\n\"Normalization: Normalization that goes beyond removing vertical structure could be strategic,\nsuch as removing the geographic mean (e.g., latitudinal, land/sea structure) or composite seasonal\nvariances (e.g., local smoothed annual cycle) present in the data. For variables exhibiting exponential\nvariation and approaching zero at the highest level (e.g., metrics of moisture), log-normalization\nmight be beneficial.\"\n\nhttps://arxiv.org/pdf/2306.08754.pdf\n\nhttps://huggingface.co/datasets/LEAP/ClimSim_high-res\n\nhttps://leap-stc.github.io/ClimSim/README.html","metadata":{}},{"cell_type":"markdown","source":"#Started 22:13 to load till 23:35. After 1h:26 it has alocated more memory than is available. It has restarted.","metadata":{}},{"cell_type":"markdown","source":"@misc{leap-atmospheric-physics-ai-climsim,\n\n    author = {Jerry Lin, Zeyuan Hu, Sungduk Yu, Mike Pritchard, Ritwik Gupta, Tian Zheng, Walter Hannah, Laura Mansfield, Yongquan Qu, Margarita Geleta, Molly Lopez, Maja Rudolph, Ashley Chow, Walter Reade},\n    \n    title = {LEAP - Atmospheric Physics using AI (ClimSim)},\n    publisher = {Kaggle},\n    \n    year = {2024},\n    url = {https://kaggle.com/competitions/leap-atmospheric-physics-ai-climsim}\n}","metadata":{}},{"cell_type":"code","source":"#Code by Luca Massaron  https://www.kaggle.com/lucamassaron/training-data-to-feather-python-r-low-mem\n\n# defining data types\n\ntraining_path = '../input/leap-atmospheric-physics-ai-climsim/train.csv'\n\ndtypes = {\n    'sample_id': 'str',\n    #'ptend_t': 'float32',\n    #'ptend_q0001': 'float32',\n    #'ptend_q0002': 'float32',\n    #'ptend_q0003': 'float32',\n    #'ptend_u': 'float32',\n    #'ptend_v': 'float32',\n    'cam_out_NETSW': 'float32',\n    'cam_out_FLWDS': 'float32',\n    'cam_out_PRECSC': 'float32',\n    'cam_out_PRECC': 'float32',\n    'cam_out_SOLS': 'float32',\n    'cam_out_SOLL': 'float32',\n    'cam_out_SOLSD': 'float32',\n    'cam_out_SOLLD': 'float32',\n    'cam_out_SOLL': 'float32',\n    'state_t_0': 'float32',\n    'state_q0001_0': 'float32',\n    'state_q0002_0': 'float32',\n    'state_q0003_0': 'float32',\n    'state_u_0': 'float32',\n    'state_v_0': 'float32',\n    #'pbuf_ozone': 'float32',\n    #'pbuf_CH4': 'float32',\n    #'pbuf_N2O': 'float32',\n    \n}\n\ndtypes.update({f'state_t_{i}': 'float32' for i in range(60)})#We have 60 Dimension\ndtypes.update({f'state_q0001_{i}': 'float32' for i in range(60)})\ndtypes.update({f'state_q0002_{i}': 'float32' for i in range(60)})\ndtypes.update({f'state_q0003_{i}': 'float32' for i in range(60)})\ndtypes.update({f'state_u_{i}': 'float32' for i in range(60)})\ndtypes.update({f'state_v_{i}': 'float32' for i in range(60)})# I added this because we have Dimension 60","metadata":{"execution":{"iopub.status.busy":"2024-04-19T03:04:03.42254Z","iopub.execute_input":"2024-04-19T03:04:03.423069Z","iopub.status.idle":"2024-04-19T03:04:03.432141Z","shell.execute_reply.started":"2024-04-19T03:04:03.423031Z","shell.execute_reply":"2024-04-19T03:04:03.430796Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I think dtypes.updates should have other lines for the other features. In fact, I'm lost with this data\n\nThat below is taking so Looooong! Besides, it isn't correct.","metadata":{}},{"cell_type":"code","source":"#Code by Luca Massaron  https://www.kaggle.com/lucamassaron/training-data-to-feather-python-r-low-mem\n\ntrain = pd.read_csv(\n        training_path,\n        usecols=list(dtypes.keys()),\n        dtype=dtypes\n    )","metadata":{"execution":{"iopub.status.busy":"2024-04-19T03:04:08.748495Z","iopub.execute_input":"2024-04-19T03:04:08.748922Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Your notebook tried to allocate more memory than is available. It has restarted.","metadata":{}},{"cell_type":"code","source":"#Code by Luca Massaron  https://www.kaggle.com/lucamassaron/training-data-to-feather-python-r-low-mem\n\ntrain.to_feather(\"train.feather\")","metadata":{"execution":{"iopub.status.busy":"2024-04-19T03:00:10.681191Z","iopub.execute_input":"2024-04-19T03:00:10.681666Z","iopub.status.idle":"2024-04-19T03:00:10.719092Z","shell.execute_reply.started":"2024-04-19T03:00:10.68163Z","shell.execute_reply":"2024-04-19T03:00:10.717339Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Code by Luca Massaron  https://www.kaggle.com/lucamassaron/training-data-to-feather-python-r-low-mem\n\ntrain = pd.read_feather(\"train.feather\")","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Code by Luca Massaron  https://www.kaggle.com/lucamassaron/training-data-to-feather-python-r-low-mem\n\ntrain.info(memory_usage=True)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Acknowledgement\n\nLuca Massaron https://www.kaggle.com/lucamassaron/training-data-to-feather-python-r-low-mem","metadata":{}}]}