{
  "id": 494928,
  "title": "Dealing with Large Datasets. Kaggle Notebooks, Topics and Competitions.",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/494928",
  "author_name": "Marília Prata",
  "post_date": "2024-04-19T01:05:59.910000",
  "votes": 37,
  "comment_count": 4,
  "views": 0,
  "content": "<h1>ClimSim, the largest-ever dataset designed for hybrid ML-physics</h1>\n<p>\"The authors presented ClimSim, the largest-ever dataset designed for hybrid ML-physics research. It comprises multi-scale climate simulations, developed by a consortium of climate scientists and ML researchers. It consists of 5.7 billion pairs of multivariate input and output vectors that isolate the influence of locally-nested, high-resolution, high-fidelity physics on a host climate simulator’s macro-scale physical state.\"<br>\n<a href=\"https://arxiv.org/pdf/2306.08754.pdf\" target=\"_blank\">https://arxiv.org/pdf/2306.08754.pdf</a></p>\n<h1>Handling Large Datasets</h1>\n<p>Methods and Formats: Rapids, Dask, Datatable, Feather, HDF5, Jay, Parquet, Pickle, Polars</p>\n<p>METHODS: Dask, Datatable, Pandas, Rapids.</p>\n<h1>Kaggle Notebooks</h1>\n<p><a href=\"https://www.kaggle.com/code/rohanrao/tutorial-on-reading-large-datasets/notebook\" target=\"_blank\">Tutorial on reading large datasets</a> By Vopani</p>\n<p>Dask</p>\n<p><a href=\"https://www.kaggle.com/aakashnain/can-we-read-faster\" target=\"_blank\">Can we read faster?</a> By Nain</p>\n<p>DATATABLE</p>\n<p><a href=\"https://www.kaggle.com/udbhavpangotra/reading-the-data-datatable-works-like-a-charm\" target=\"_blank\">Reading the data, Datatable works like a charm!</a> By Udbhav Pangotra</p>\n<p>However, Datatable has some missing functionalities. Which means that there are some functions in pandas that do not have an equivalent in datatable yet, and are likely to be implemented.</p>\n<p><a href=\"https://datatable.readthedocs.io/en/latest/manual/comparison_with_pandas.html\" target=\"_blank\">https://datatable.readthedocs.io/en/latest/manual/comparison_with_pandas.html</a></p>\n<h1>FORMATS: Feather, hdf5, Parquet, Pickle, Jay.</h1>\n<p>FEATHER<br>\nStore data in feather (binary) format specifically for pandas. It significantly improves reading speed of datasets.</p>\n<p>Format: hdf5</p>\n<p>\"HDF5 is a high-performance data management suite to store, manage and process large and complex data.\"</p>\n<p>Format: Jay<br>\n\"Datatable uses .jay (binary) format which makes reading datasets blazing fast.\"</p>\n<p>Format: Parquet<br>\n\"Parquet now extensively used with Spark.\"</p>\n<p><a href=\"https://www.kaggle.com/gauravbrills/10-folds-stratified-parquet-feather\" target=\"_blank\">10_folds_stratified_parquet&amp;feather</a> By Gaurav Rawat </p>\n<p>Format: Pickle<br>\n\"Python objects can be stored in the form of pickle files and pandas has inbuilt functions to read and write dataframes as pickle objects.\"</p>\n<h1>Kaggle Topics:</h1>\n<p><a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/495128\" target=\"_blank\">Useful code and comparison result of save speeds between Polars and Pandas</a> By Chumajin</p>\n<p>Targets extraction: (By Chumajin)</p>\n<blockquote>\n  <p>import polars as pl<br>\n  sample = pl.read_csv(\"/kaggle/input/leap-atmospheric-physics-ai-climsim/sample_submission.csv\",n_rows=1)</p>\n  <h1>read only 1 row</h1>\n  <p>targets = sample.columns[1:]</p>\n</blockquote>\n<p>Save submission.csv (By Chumajin)</p>\n<blockquote>\n  <h2>change pandas to polars if you use pandas</h2>\n  <p>test_polars = pl.from_pandas(test[[\"sample_id\"]+targets])<br>\n  test_polars.write_csv(\"submission.csv\")</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364721\" target=\"_blank\">Recommendation Systems for Large Datasets</a> By Ravi Shah</p>\n<p><a href=\"https://www.kaggle.com/discussions/getting-started/9512\" target=\"_blank\">Large Datasets</a> 10 Years ago By Danny Diaz</p>\n<p><a href=\"https://www.kaggle.com/competitions/feedback-prize-2021/discussion/304706\" target=\"_blank\">Training Large Models Effectively with DeepSpeed</a> By Mr. KnowNothing</p>\n<h1>Large Data Kaggle Competitions</h1>\n<p><a href=\"https://www.kaggle.com/competitions/riiid-test-answer-prediction\" target=\"_blank\">Riiid Answer Correctness Prediction</a></p>\n<p><a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/data\" target=\"_blank\">Stanford Ribonanza RNA Folding</a></p>\n<p><a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/data\" target=\"_blank\">IceCube - Neutrinos in Deep Ice</a> </p>\n<p><a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/491328\" target=\"_blank\">Leash Bio - Predict New Medicines with BELKA</a></p>\n<p><a href=\"https://www.kaggle.com/competitions/geolifeclef-2024/data\" target=\"_blank\">GeoLifeCLEF 2024 @ LifeCLEF &amp; CVPR-FGVC</a></p>\n<p><a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/data\" target=\"_blank\">LEAP - Atmospheric Physics using AI - ClimSim</a></p>\n<h1>Reduce Memory Usage</h1>\n<p><a href=\"https://www.kaggle.com/code/arjanso/reducing-dataframe-memory-size-by-65\" target=\"_blank\">Reducing DataFrame memory size by ~65%</a> By ArjanGroen</p>\n<p><a href=\"https://www.kaggle.com/code/gemartin/load-data-reduce-memory-usage/notebook\" target=\"_blank\">load data -reduce memory usage</a> By Guillaume Martin</p>\n<p><a href=\"https://www.kaggle.com/competitions/widsdatathon2023/discussion/376649\" target=\"_blank\">df_shrink from fastai </a> Tip By Carl MacBride Ellis </p>",
  "messages": [
    {
      "id": 2759880,
      "postDate": "2024-04-19T01:05:59.910Z",
      "content": "<h1>ClimSim, the largest-ever dataset designed for hybrid ML-physics</h1>\n<p>\"The authors presented ClimSim, the largest-ever dataset designed for hybrid ML-physics research. It comprises multi-scale climate simulations, developed by a consortium of climate scientists and ML researchers. It consists of 5.7 billion pairs of multivariate input and output vectors that isolate the influence of locally-nested, high-resolution, high-fidelity physics on a host climate simulator’s macro-scale physical state.\"<br>\n<a href=\"https://arxiv.org/pdf/2306.08754.pdf\" target=\"_blank\">https://arxiv.org/pdf/2306.08754.pdf</a></p>\n<h1>Handling Large Datasets</h1>\n<p>Methods and Formats: Rapids, Dask, Datatable, Feather, HDF5, Jay, Parquet, Pickle, Polars</p>\n<p>METHODS: Dask, Datatable, Pandas, Rapids.</p>\n<h1>Kaggle Notebooks</h1>\n<p><a href=\"https://www.kaggle.com/code/rohanrao/tutorial-on-reading-large-datasets/notebook\" target=\"_blank\">Tutorial on reading large datasets</a> By Vopani</p>\n<p>Dask</p>\n<p><a href=\"https://www.kaggle.com/aakashnain/can-we-read-faster\" target=\"_blank\">Can we read faster?</a> By Nain</p>\n<p>DATATABLE</p>\n<p><a href=\"https://www.kaggle.com/udbhavpangotra/reading-the-data-datatable-works-like-a-charm\" target=\"_blank\">Reading the data, Datatable works like a charm!</a> By Udbhav Pangotra</p>\n<p>However, Datatable has some missing functionalities. Which means that there are some functions in pandas that do not have an equivalent in datatable yet, and are likely to be implemented.</p>\n<p><a href=\"https://datatable.readthedocs.io/en/latest/manual/comparison_with_pandas.html\" target=\"_blank\">https://datatable.readthedocs.io/en/latest/manual/comparison_with_pandas.html</a></p>\n<h1>FORMATS: Feather, hdf5, Parquet, Pickle, Jay.</h1>\n<p>FEATHER<br>\nStore data in feather (binary) format specifically for pandas. It significantly improves reading speed of datasets.</p>\n<p>Format: hdf5</p>\n<p>\"HDF5 is a high-performance data management suite to store, manage and process large and complex data.\"</p>\n<p>Format: Jay<br>\n\"Datatable uses .jay (binary) format which makes reading datasets blazing fast.\"</p>\n<p>Format: Parquet<br>\n\"Parquet now extensively used with Spark.\"</p>\n<p><a href=\"https://www.kaggle.com/gauravbrills/10-folds-stratified-parquet-feather\" target=\"_blank\">10_folds_stratified_parquet&amp;feather</a> By Gaurav Rawat </p>\n<p>Format: Pickle<br>\n\"Python objects can be stored in the form of pickle files and pandas has inbuilt functions to read and write dataframes as pickle objects.\"</p>\n<h1>Kaggle Topics:</h1>\n<p><a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/495128\" target=\"_blank\">Useful code and comparison result of save speeds between Polars and Pandas</a> By Chumajin</p>\n<p>Targets extraction: (By Chumajin)</p>\n<blockquote>\n  <p>import polars as pl<br>\n  sample = pl.read_csv(\"/kaggle/input/leap-atmospheric-physics-ai-climsim/sample_submission.csv\",n_rows=1)</p>\n  <h1>read only 1 row</h1>\n  <p>targets = sample.columns[1:]</p>\n</blockquote>\n<p>Save submission.csv (By Chumajin)</p>\n<blockquote>\n  <h2>change pandas to polars if you use pandas</h2>\n  <p>test_polars = pl.from_pandas(test[[\"sample_id\"]+targets])<br>\n  test_polars.write_csv(\"submission.csv\")</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364721\" target=\"_blank\">Recommendation Systems for Large Datasets</a> By Ravi Shah</p>\n<p><a href=\"https://www.kaggle.com/discussions/getting-started/9512\" target=\"_blank\">Large Datasets</a> 10 Years ago By Danny Diaz</p>\n<p><a href=\"https://www.kaggle.com/competitions/feedback-prize-2021/discussion/304706\" target=\"_blank\">Training Large Models Effectively with DeepSpeed</a> By Mr. KnowNothing</p>\n<h1>Large Data Kaggle Competitions</h1>\n<p><a href=\"https://www.kaggle.com/competitions/riiid-test-answer-prediction\" target=\"_blank\">Riiid Answer Correctness Prediction</a></p>\n<p><a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/data\" target=\"_blank\">Stanford Ribonanza RNA Folding</a></p>\n<p><a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/data\" target=\"_blank\">IceCube - Neutrinos in Deep Ice</a> </p>\n<p><a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/491328\" target=\"_blank\">Leash Bio - Predict New Medicines with BELKA</a></p>\n<p><a href=\"https://www.kaggle.com/competitions/geolifeclef-2024/data\" target=\"_blank\">GeoLifeCLEF 2024 @ LifeCLEF &amp; CVPR-FGVC</a></p>\n<p><a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/data\" target=\"_blank\">LEAP - Atmospheric Physics using AI - ClimSim</a></p>\n<h1>Reduce Memory Usage</h1>\n<p><a href=\"https://www.kaggle.com/code/arjanso/reducing-dataframe-memory-size-by-65\" target=\"_blank\">Reducing DataFrame memory size by ~65%</a> By ArjanGroen</p>\n<p><a href=\"https://www.kaggle.com/code/gemartin/load-data-reduce-memory-usage/notebook\" target=\"_blank\">load data -reduce memory usage</a> By Guillaume Martin</p>\n<p><a href=\"https://www.kaggle.com/competitions/widsdatathon2023/discussion/376649\" target=\"_blank\">df_shrink from fastai </a> Tip By Carl MacBride Ellis </p>",
      "rawMarkdown": "#ClimSim, the largest-ever dataset designed for hybrid ML-physics\n\n\"The authors presented ClimSim, the largest-ever dataset designed for hybrid ML-physics research. It comprises multi-scale climate simulations, developed by a consortium of climate scientists and ML researchers. It consists of 5.7 billion pairs of multivariate input and output vectors that isolate the influence of locally-nested, high-resolution, high-fidelity physics on a host climate simulator’s macro-scale physical state.\"\nhttps://arxiv.org/pdf/2306.08754.pdf\n\n#Handling Large Datasets \n\nMethods and Formats: Rapids, Dask, Datatable, Feather, HDF5, Jay, Parquet, Pickle, Polars\n\nMETHODS: Dask, Datatable, Pandas, Rapids.\n\n#Kaggle Notebooks\n\n[Tutorial on reading large datasets](https://www.kaggle.com/code/rohanrao/tutorial-on-reading-large-datasets/notebook) By Vopani\n\nDask\n\n[Can we read faster?](https://www.kaggle.com/aakashnain/can-we-read-faster) By Nain\n\nDATATABLE\n\n[Reading the data, Datatable works like a charm!](https://www.kaggle.com/udbhavpangotra/reading-the-data-datatable-works-like-a-charm) By Udbhav Pangotra\n\nHowever, Datatable has some missing functionalities. Which means that there are some functions in pandas that do not have an equivalent in datatable yet, and are likely to be implemented.\n\nhttps://datatable.readthedocs.io/en/latest/manual/comparison_with_pandas.html\n\n\n#FORMATS: Feather, hdf5, Parquet, Pickle, Jay.\n\nFEATHER\nStore data in feather (binary) format specifically for pandas. It significantly improves reading speed of datasets.\n\nFormat: hdf5\n\n\"HDF5 is a high-performance data management suite to store, manage and process large and complex data.\"\n\nFormat: Jay\n\"Datatable uses .jay (binary) format which makes reading datasets blazing fast.\"\n\nFormat: Parquet\n\"Parquet now extensively used with Spark.\"\n\n[10_folds_stratified_parquet&feather](https://www.kaggle.com/gauravbrills/10-folds-stratified-parquet-feather) By Gaurav Rawat \n\nFormat: Pickle\n\"Python objects can be stored in the form of pickle files and pandas has inbuilt functions to read and write dataframes as pickle objects.\"\n\n#Kaggle Topics:\n\n[Useful code and comparison result of save speeds between Polars and Pandas](https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/495128) By Chumajin\n\nTargets extraction: (By Chumajin)\n\n>import polars as pl\nsample = pl.read_csv(\"/kaggle/input/leap-atmospheric-physics-ai-climsim/sample_submission.csv\",n_rows=1)\n# read only 1 row\ntargets = sample.columns[1:]\n\nSave submission.csv (By Chumajin)\n\n>## change pandas to polars if you use pandas\ntest_polars = pl.from_pandas(test[[\"sample_id\"]+targets])\ntest_polars.write_csv(\"submission.csv\")\n\n[Recommendation Systems for Large Datasets](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364721) By Ravi Shah\n\n[Large Datasets](https://www.kaggle.com/discussions/getting-started/9512) 10 Years ago By Danny Diaz\n\n[Training Large Models Effectively with DeepSpeed](https://www.kaggle.com/competitions/feedback-prize-2021/discussion/304706) By Mr. KnowNothing\n\n#Large Data Kaggle Competitions \n\n[Riiid Answer Correctness Prediction](https://www.kaggle.com/competitions/riiid-test-answer-prediction)\n\n[Stanford Ribonanza RNA Folding](https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/data)\n\n[IceCube - Neutrinos in Deep Ice](https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/data) \n\n[Leash Bio - Predict New Medicines with BELKA](https://www.kaggle.com/competitions/leash-BELKA/discussion/491328)\n\n[GeoLifeCLEF 2024 @ LifeCLEF & CVPR-FGVC](https://www.kaggle.com/competitions/geolifeclef-2024/data)\n\n[LEAP - Atmospheric Physics using AI - ClimSim](https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/data)\n\n#Reduce Memory Usage\n\n[Reducing DataFrame memory size by ~65%](https://www.kaggle.com/code/arjanso/reducing-dataframe-memory-size-by-65) By ArjanGroen\n\n[load data -reduce memory usage](https://www.kaggle.com/code/gemartin/load-data-reduce-memory-usage/notebook) By Guillaume Martin\n\n[df_shrink from fastai ](https://www.kaggle.com/competitions/widsdatathon2023/discussion/376649) Tip By Carl MacBride Ellis ",
      "votes": 37
    },
    {
      "id": 2788151,
      "postDate": "2024-05-02T05:30:12.097Z",
      "content": "<p>Thank you for the great guide. It's real help to me.</p>",
      "rawMarkdown": "Thank you for the great guide. It's real help to me.",
      "votes": 1,
      "replies": [
        {
          "id": 2788899,
          "postDate": "2024-05-02T13:24:16.883Z",
          "content": "<p>Hi WongsangK<br>\nI'm glad that you found it helpful. I can only thanks for those guys that shared their knowledge (codes and topics) on Kaggle.  </p>",
          "rawMarkdown": "Hi WongsangK\nI'm glad that you found it helpful. I can only thanks for those guys that shared their knowledge (codes and topics) on Kaggle.  "
        }
      ]
    },
    {
      "id": 2760106,
      "postDate": "2024-04-19T05:44:01.023Z",
      "content": "<p>Wow, this dataset is quite substantial. It's the largest one I've encountered thus far. My poor system is crying in the corner 😀</p>",
      "rawMarkdown": "Wow, this dataset is quite substantial. It's the largest one I've encountered thus far. My poor system is crying in the corner 😀",
      "votes": 1,
      "replies": [
        {
          "id": 2760153,
          "postDate": "2024-04-19T06:20:04.357Z",
          "content": "<p>Hi Dee Dee,</p>\n<p>I spent so long and the Notebook failed.  It has alocated more memory than is available. It has restarted.. Not funny, I hope to learn to work with such amount of data. I' m feeling I was ran over by a Huge amount of data. Now, I need to recover 🤣 </p>",
          "rawMarkdown": "Hi Dee Dee,\n\nI spent so long and the Notebook failed.  It has alocated more memory than is available. It has restarted.. Not funny, I hope to learn to work with such amount of data. I' m feeling I was ran over by a Huge amount of data. Now, I need to recover 🤣 ",
          "votes": 2
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2788151,
      "author_name": "WonsangK",
      "author_url": "",
      "post_date": "2024-05-02T05:30:12.097000",
      "content": "<p>Thank you for the great guide. It's real help to me.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2788899,
          "author_name": "Marília Prata",
          "author_url": "",
          "post_date": "2024-05-02T13:24:16.883000",
          "content": "<p>Hi WongsangK<br>\nI'm glad that you found it helpful. I can only thanks for those guys that shared their knowledge (codes and topics) on Kaggle.  </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2760106,
      "author_name": "Dhruv D",
      "author_url": "",
      "post_date": "2024-04-19T05:44:01.023000",
      "content": "<p>Wow, this dataset is quite substantial. It's the largest one I've encountered thus far. My poor system is crying in the corner 😀</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2760153,
          "author_name": "Marília Prata",
          "author_url": "",
          "post_date": "2024-04-19T06:20:04.357000",
          "content": "<p>Hi Dee Dee,</p>\n<p>I spent so long and the Notebook failed.  It has alocated more memory than is available. It has restarted.. Not funny, I hope to learn to work with such amount of data. I' m feeling I was ran over by a Huge amount of data. Now, I need to recover 🤣 </p>",
          "votes": 2,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2759880": "#ClimSim, the largest-ever dataset designed for hybrid ML-physics\n\n\"The authors presented ClimSim, the largest-ever dataset designed for hybrid ML-physics research. It comprises multi-scale climate simulations, developed by a consortium of climate scientists and ML researchers. It consists of 5.7 billion pairs of multivariate input and output vectors that isolate the influence of locally-nested, high-resolution, high-fidelity physics on a host climate simulator’s macro-scale physical state.\"\nhttps://arxiv.org/pdf/2306.08754.pdf\n\n#Handling Large Datasets \n\nMethods and Formats: Rapids, Dask, Datatable, Feather, HDF5, Jay, Parquet, Pickle, Polars\n\nMETHODS: Dask, Datatable, Pandas, Rapids.\n\n#Kaggle Notebooks\n\n[Tutorial on reading large datasets](https://www.kaggle.com/code/rohanrao/tutorial-on-reading-large-datasets/notebook) By Vopani\n\nDask\n\n[Can we read faster?](https://www.kaggle.com/aakashnain/can-we-read-faster) By Nain\n\nDATATABLE\n\n[Reading the data, Datatable works like a charm!](https://www.kaggle.com/udbhavpangotra/reading-the-data-datatable-works-like-a-charm) By Udbhav Pangotra\n\nHowever, Datatable has some missing functionalities. Which means that there are some functions in pandas that do not have an equivalent in datatable yet, and are likely to be implemented.\n\nhttps://datatable.readthedocs.io/en/latest/manual/comparison_with_pandas.html\n\n\n#FORMATS: Feather, hdf5, Parquet, Pickle, Jay.\n\nFEATHER\nStore data in feather (binary) format specifically for pandas. It significantly improves reading speed of datasets.\n\nFormat: hdf5\n\n\"HDF5 is a high-performance data management suite to store, manage and process large and complex data.\"\n\nFormat: Jay\n\"Datatable uses .jay (binary) format which makes reading datasets blazing fast.\"\n\nFormat: Parquet\n\"Parquet now extensively used with Spark.\"\n\n[10_folds_stratified_parquet&feather](https://www.kaggle.com/gauravbrills/10-folds-stratified-parquet-feather) By Gaurav Rawat \n\nFormat: Pickle\n\"Python objects can be stored in the form of pickle files and pandas has inbuilt functions to read and write dataframes as pickle objects.\"\n\n#Kaggle Topics:\n\n[Useful code and comparison result of save speeds between Polars and Pandas](https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/discussion/495128) By Chumajin\n\nTargets extraction: (By Chumajin)\n\n>import polars as pl\nsample = pl.read_csv(\"/kaggle/input/leap-atmospheric-physics-ai-climsim/sample_submission.csv\",n_rows=1)\n# read only 1 row\ntargets = sample.columns[1:]\n\nSave submission.csv (By Chumajin)\n\n>## change pandas to polars if you use pandas\ntest_polars = pl.from_pandas(test[[\"sample_id\"]+targets])\ntest_polars.write_csv(\"submission.csv\")\n\n[Recommendation Systems for Large Datasets](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364721) By Ravi Shah\n\n[Large Datasets](https://www.kaggle.com/discussions/getting-started/9512) 10 Years ago By Danny Diaz\n\n[Training Large Models Effectively with DeepSpeed](https://www.kaggle.com/competitions/feedback-prize-2021/discussion/304706) By Mr. KnowNothing\n\n#Large Data Kaggle Competitions \n\n[Riiid Answer Correctness Prediction](https://www.kaggle.com/competitions/riiid-test-answer-prediction)\n\n[Stanford Ribonanza RNA Folding](https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/data)\n\n[IceCube - Neutrinos in Deep Ice](https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/data) \n\n[Leash Bio - Predict New Medicines with BELKA](https://www.kaggle.com/competitions/leash-BELKA/discussion/491328)\n\n[GeoLifeCLEF 2024 @ LifeCLEF & CVPR-FGVC](https://www.kaggle.com/competitions/geolifeclef-2024/data)\n\n[LEAP - Atmospheric Physics using AI - ClimSim](https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/data)\n\n#Reduce Memory Usage\n\n[Reducing DataFrame memory size by ~65%](https://www.kaggle.com/code/arjanso/reducing-dataframe-memory-size-by-65) By ArjanGroen\n\n[load data -reduce memory usage](https://www.kaggle.com/code/gemartin/load-data-reduce-memory-usage/notebook) By Guillaume Martin\n\n[df_shrink from fastai ](https://www.kaggle.com/competitions/widsdatathon2023/discussion/376649) Tip By Carl MacBride Ellis ",
    "2788151": "Thank you for the great guide. It's real help to me.",
    "2760106": "Wow, this dataset is quite substantial. It's the largest one I've encountered thus far. My poor system is crying in the corner 😀"
  }
}