{
  "id": 497217,
  "title": "How to read in the train dataset, not enough RAM?",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/497217",
  "author_name": "Rehan Daya",
  "post_date": "2024-04-24T02:58:23.479000",
  "votes": 2,
  "comment_count": 4,
  "views": 0,
  "content": "<p>The train dataset is 180GB, how are we supposed to read it in, to then convert to parquet and load in as polar or dask, if we only have 20gb of ram? It fully saturates the ram and crashes before it can convert to parquet. I am using Kaggle notebook here, my personal comp is too weak for this, and I don't wan to pay for cloud computing.</p>",
  "messages": [
    {
      "id": 2772324,
      "postDate": "2024-04-24T16:51:35.593Z",
      "content": "<p>I processed the huge dataset and saved it in a more friendly format <a href=\"https://www.kaggle.com/datasets/titericz/leap-dataset-giba\" target=\"_blank\">here</a>. It was split in 17 parts with 625k rows each. So you can easily load a subset and play with it.</p>",
      "rawMarkdown": "I processed the huge dataset and saved it in a more friendly format [here](https://www.kaggle.com/datasets/titericz/leap-dataset-giba). It was split in 17 parts with 625k rows each. So you can easily load a subset and play with it.",
      "votes": 5,
      "replies": [
        {
          "id": 2781855,
          "postDate": "2024-04-29T03:21:17.130Z",
          "content": "<p><code>OSError: Could not open Parquet input source '&lt;Buffer&gt;': Metadata contains Thrift LogicalType that is not recognized</code></p>\n<p>could not load the data</p>",
          "rawMarkdown": "`OSError: Could not open Parquet input source '<Buffer>': Metadata contains Thrift LogicalType that is not recognized`\n\ncould not load the data",
          "replies": [
            {
              "id": 2783143,
              "postDate": "2024-04-29T15:38:33.707Z",
              "content": "<p>Of course you can. I'm loading it here: <a href=\"https://www.kaggle.com/code/titericz/giba-baseline-xgboost\" target=\"_blank\">https://www.kaggle.com/code/titericz/giba-baseline-xgboost</a></p>",
              "rawMarkdown": "Of course you can. I'm loading it here: https://www.kaggle.com/code/titericz/giba-baseline-xgboost",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2770787,
      "postDate": "2024-04-24T03:05:17.057Z",
      "content": "<p><a href=\"https://www.kaggle.com/rehankhandaya\" target=\"_blank\">@rehankhandaya</a> perhaps reading in chunks will help you, both pandas and polars offer chunked reading in of files.<br>\nI agree with you that using a 196GB csv file as a single training set was sub-optimal. I hope such issues could be handled better by hosts and Kaggle going ahead. <br>\nFor this purpose, you can also use the TPU, but be mindful of the weekly quotas herewith!</p>",
      "rawMarkdown": "@rehankhandaya perhaps reading in chunks will help you, both pandas and polars offer chunked reading in of files.\nI agree with you that using a 196GB csv file as a single training set was sub-optimal. I hope such issues could be handled better by hosts and Kaggle going ahead. \nFor this purpose, you can also use the TPU, but be mindful of the weekly quotas herewith!",
      "votes": 3
    },
    {
      "id": 2770776,
      "postDate": "2024-04-24T02:58:23.480Z",
      "content": "<p>The train dataset is 180GB, how are we supposed to read it in, to then convert to parquet and load in as polar or dask, if we only have 20gb of ram? It fully saturates the ram and crashes before it can convert to parquet. I am using Kaggle notebook here, my personal comp is too weak for this, and I don't wan to pay for cloud computing.</p>",
      "rawMarkdown": "The train dataset is 180GB, how are we supposed to read it in, to then convert to parquet and load in as polar or dask, if we only have 20gb of ram? It fully saturates the ram and crashes before it can convert to parquet. I am using Kaggle notebook here, my personal comp is too weak for this, and I don't wan to pay for cloud computing.",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 2772324,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2024-04-24T16:51:35.593000",
      "content": "<p>I processed the huge dataset and saved it in a more friendly format <a href=\"https://www.kaggle.com/datasets/titericz/leap-dataset-giba\" target=\"_blank\">here</a>. It was split in 17 parts with 625k rows each. So you can easily load a subset and play with it.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 2781855,
          "author_name": "jewinv",
          "author_url": "",
          "post_date": "2024-04-29T03:21:17.130000",
          "content": "<p><code>OSError: Could not open Parquet input source '&lt;Buffer&gt;': Metadata contains Thrift LogicalType that is not recognized</code></p>\n<p>could not load the data</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2783143,
              "author_name": "Giba",
              "author_url": "",
              "post_date": "2024-04-29T15:38:33.707000",
              "content": "<p>Of course you can. I'm loading it here: <a href=\"https://www.kaggle.com/code/titericz/giba-baseline-xgboost\" target=\"_blank\">https://www.kaggle.com/code/titericz/giba-baseline-xgboost</a></p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2770787,
      "author_name": "Ravi Ramakrishnan",
      "author_url": "",
      "post_date": "2024-04-24T03:05:17.057000",
      "content": "<p><a href=\"https://www.kaggle.com/rehankhandaya\" target=\"_blank\">@rehankhandaya</a> perhaps reading in chunks will help you, both pandas and polars offer chunked reading in of files.<br>\nI agree with you that using a 196GB csv file as a single training set was sub-optimal. I hope such issues could be handled better by hosts and Kaggle going ahead. <br>\nFor this purpose, you can also use the TPU, but be mindful of the weekly quotas herewith!</p>",
      "votes": 3,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2772324": "I processed the huge dataset and saved it in a more friendly format [here](https://www.kaggle.com/datasets/titericz/leap-dataset-giba). It was split in 17 parts with 625k rows each. So you can easily load a subset and play with it.",
    "2770787": "@rehankhandaya perhaps reading in chunks will help you, both pandas and polars offer chunked reading in of files.\nI agree with you that using a 196GB csv file as a single training set was sub-optimal. I hope such issues could be handled better by hosts and Kaggle going ahead. \nFor this purpose, you can also use the TPU, but be mindful of the weekly quotas herewith!",
    "2770776": "The train dataset is 180GB, how are we supposed to read it in, to then convert to parquet and load in as polar or dask, if we only have 20gb of ram? It fully saturates the ram and crashes before it can convert to parquet. I am using Kaggle notebook here, my personal comp is too weak for this, and I don't wan to pay for cloud computing."
  }
}