{
  "id": 506305,
  "title": "PyArrow: The best python library for reading big datasets ",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/506305",
  "author_name": "MTP",
  "post_date": "2024-05-21T11:08:07.390000",
  "votes": 4,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Dear all, as per our investigations (consult this notebook: <a href=\"https://www.kaggle.com/code/tariqcp/pyarrow-vs-duckdb-pandas-and-polars)\" target=\"_blank\">https://www.kaggle.com/code/tariqcp/pyarrow-vs-duckdb-pandas-and-polars)</a>, the best library to read big datasets is PyArrow. It can read data in chunks/ batches efficiently without any memory issue. Pandas is on the second number in the list. DuckDB consumed the maximum time especially reading data in chunks. Whereas Polars library could not read the file.  For example, I used following code for reading train.csv file of this competition but it crashes Kaggle RAM after some time. Polars has some memory management issue while it reads data in batches. If you have to use this library, please be remember to keep batch size smaller. for example, 5000.</p>\n<hr>\n<p>import polars as pl<br>\nimport gc<br>\n%%time </p>\n<h1>Initialize an empty DataFrame</h1>\n<p>reader = pl.read_csv_batched(<br>\n    base_path+'train.csv',<br>\n    batch_size = 10000, <br>\n    separator=\",\",<br>\n )  <br>\nfor batch in reader.next_batches(10000):<br>\n    # process the batch (which is a DataFrame)<br>\n    del batch<br>\n    gc.collect()</p>\n<hr>\n<h2>On the hand when I used following code of PyArrow. It successfully read the entire train.csv file in just 38 minutes.</h2>\n<p>from pyarrow.csv import CSVStreamingReader as csv_sr<br>\n%%time </p>\n<h1>return csv stream reader</h1>\n<p>csv_sr =csv.open_csv(base_path + 'train.csv')<br>\nfor batch in csv_sr:<br>\n    # Read the next batch<br>\n    table = pa.Table.from_batches([batch])<br>\n    #df = table.to_pandas()</p>",
  "messages": [
    {
      "id": 2827275,
      "postDate": "2024-05-21T11:08:07.390Z",
      "content": "<p>Dear all, as per our investigations (consult this notebook: <a href=\"https://www.kaggle.com/code/tariqcp/pyarrow-vs-duckdb-pandas-and-polars)\" target=\"_blank\">https://www.kaggle.com/code/tariqcp/pyarrow-vs-duckdb-pandas-and-polars)</a>, the best library to read big datasets is PyArrow. It can read data in chunks/ batches efficiently without any memory issue. Pandas is on the second number in the list. DuckDB consumed the maximum time especially reading data in chunks. Whereas Polars library could not read the file.  For example, I used following code for reading train.csv file of this competition but it crashes Kaggle RAM after some time. Polars has some memory management issue while it reads data in batches. If you have to use this library, please be remember to keep batch size smaller. for example, 5000.</p>\n<hr>\n<p>import polars as pl<br>\nimport gc<br>\n%%time </p>\n<h1>Initialize an empty DataFrame</h1>\n<p>reader = pl.read_csv_batched(<br>\n    base_path+'train.csv',<br>\n    batch_size = 10000, <br>\n    separator=\",\",<br>\n )  <br>\nfor batch in reader.next_batches(10000):<br>\n    # process the batch (which is a DataFrame)<br>\n    del batch<br>\n    gc.collect()</p>\n<hr>\n<h2>On the hand when I used following code of PyArrow. It successfully read the entire train.csv file in just 38 minutes.</h2>\n<p>from pyarrow.csv import CSVStreamingReader as csv_sr<br>\n%%time </p>\n<h1>return csv stream reader</h1>\n<p>csv_sr =csv.open_csv(base_path + 'train.csv')<br>\nfor batch in csv_sr:<br>\n    # Read the next batch<br>\n    table = pa.Table.from_batches([batch])<br>\n    #df = table.to_pandas()</p>",
      "rawMarkdown": "Dear all, as per our investigations (consult this notebook: https://www.kaggle.com/code/tariqcp/pyarrow-vs-duckdb-pandas-and-polars), the best library to read big datasets is PyArrow. It can read data in chunks/ batches efficiently without any memory issue. Pandas is on the second number in the list. DuckDB consumed the maximum time especially reading data in chunks. Whereas Polars library could not read the file.  For example, I used following code for reading train.csv file of this competition but it crashes Kaggle RAM after some time. Polars has some memory management issue while it reads data in batches. If you have to use this library, please be remember to keep batch size smaller. for example, 5000.\n\n-------------------------------------------\nimport polars as pl\nimport gc\n%%time \n\n# Initialize an empty DataFrame\nreader = pl.read_csv_batched(\n    base_path+'train.csv',\n    batch_size = 10000, \n    separator=\",\",\n )  \nfor batch in reader.next_batches(10000):\n    # process the batch (which is a DataFrame)\n    del batch\n    gc.collect()\n\n-----------------------------------------------\nOn the hand when I used following code of PyArrow. It successfully read the entire train.csv file in just 38 minutes.\n------------------------------------------------\nfrom pyarrow.csv import CSVStreamingReader as csv_sr\n%%time \n#return csv stream reader\ncsv_sr =csv.open_csv(base_path + 'train.csv')\nfor batch in csv_sr:\n    # Read the next batch\n    table = pa.Table.from_batches([batch])\n    #df = table.to_pandas()",
      "votes": 4
    },
    {
      "id": 2841931,
      "postDate": "2024-05-28T18:56:31.400Z",
      "content": "<p>Great work in pyhon library <a href=\"https://www.kaggle.com/tariqcp\" target=\"_blank\">@tariqcp</a> </p>",
      "rawMarkdown": "Great work in pyhon library @tariqcp "
    }
  ],
  "comments": [
    {
      "id": 2841931,
      "author_name": "Sheema Zain",
      "author_url": "",
      "post_date": "2024-05-28T18:56:31.400000",
      "content": "<p>Great work in pyhon library <a href=\"https://www.kaggle.com/tariqcp\" target=\"_blank\">@tariqcp</a> </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2827275": "Dear all, as per our investigations (consult this notebook: https://www.kaggle.com/code/tariqcp/pyarrow-vs-duckdb-pandas-and-polars), the best library to read big datasets is PyArrow. It can read data in chunks/ batches efficiently without any memory issue. Pandas is on the second number in the list. DuckDB consumed the maximum time especially reading data in chunks. Whereas Polars library could not read the file.  For example, I used following code for reading train.csv file of this competition but it crashes Kaggle RAM after some time. Polars has some memory management issue while it reads data in batches. If you have to use this library, please be remember to keep batch size smaller. for example, 5000.\n\n-------------------------------------------\nimport polars as pl\nimport gc\n%%time \n\n# Initialize an empty DataFrame\nreader = pl.read_csv_batched(\n    base_path+'train.csv',\n    batch_size = 10000, \n    separator=\",\",\n )  \nfor batch in reader.next_batches(10000):\n    # process the batch (which is a DataFrame)\n    del batch\n    gc.collect()\n\n-----------------------------------------------\nOn the hand when I used following code of PyArrow. It successfully read the entire train.csv file in just 38 minutes.\n------------------------------------------------\nfrom pyarrow.csv import CSVStreamingReader as csv_sr\n%%time \n#return csv stream reader\ncsv_sr =csv.open_csv(base_path + 'train.csv')\nfor batch in csv_sr:\n    # Read the next batch\n    table = pa.Table.from_batches([batch])\n    #df = table.to_pandas()",
    "2841931": "Great work in pyhon library @tariqcp "
  }
}