{
  "id": 497304,
  "title": "HDF5 training data",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/497304",
  "author_name": "Taichi Uemura",
  "post_date": "2024-04-24T08:00:15.373000",
  "votes": 16,
  "comment_count": 2,
  "views": 0,
  "content": "<p>HI,</p>\n<p>I created an HDF5 version of the training data <a href=\"https://www.kaggle.com/datasets/taichiuemura/leap-climsim-hdf5-training-data/\" target=\"_blank\">https://www.kaggle.com/datasets/taichiuemura/leap-climsim-hdf5-training-data/</a>.</p>\n<p>My motivation is to make random row access easier, but I believe 37GB HDF5 file is easier to work with than 180GB CSV file for any purpose.</p>\n<p>I also made a public notebook using the dataset <a href=\"https://www.kaggle.com/code/taichiuemura/leap-mlp/\" target=\"_blank\">https://www.kaggle.com/code/taichiuemura/leap-mlp/</a>.</p>",
  "messages": [
    {
      "id": 2771364,
      "postDate": "2024-04-24T08:00:15.373Z",
      "content": "<p>HI,</p>\n<p>I created an HDF5 version of the training data <a href=\"https://www.kaggle.com/datasets/taichiuemura/leap-climsim-hdf5-training-data/\" target=\"_blank\">https://www.kaggle.com/datasets/taichiuemura/leap-climsim-hdf5-training-data/</a>.</p>\n<p>My motivation is to make random row access easier, but I believe 37GB HDF5 file is easier to work with than 180GB CSV file for any purpose.</p>\n<p>I also made a public notebook using the dataset <a href=\"https://www.kaggle.com/code/taichiuemura/leap-mlp/\" target=\"_blank\">https://www.kaggle.com/code/taichiuemura/leap-mlp/</a>.</p>",
      "rawMarkdown": "HI,\n\nI created an HDF5 version of the training data <https://www.kaggle.com/datasets/taichiuemura/leap-climsim-hdf5-training-data/>.\n\nMy motivation is to make random row access easier, but I believe 37GB HDF5 file is easier to work with than 180GB CSV file for any purpose.\n\nI also made a public notebook using the dataset <https://www.kaggle.com/code/taichiuemura/leap-mlp/>.",
      "votes": 16
    },
    {
      "id": 2778261,
      "postDate": "2024-04-27T04:32:53.747Z",
      "content": "<p><a href=\"https://www.kaggle.com/taichiuemura\" target=\"_blank\">@taichiuemura</a> , the dataset looks great. However, I noticed that indexing trough the data produces a tuple of (sample_id, sample_data) which makes data slicing a two step approach. </p>\n<p>Can you look into this? According to the data description, the sample id has been dropped. So, It will be nice that the tuple just gives the data and not the id as a tuple </p>",
      "rawMarkdown": "@taichiuemura , the dataset looks great. However, I noticed that indexing trough the data produces a tuple of (sample_id, sample_data) which makes data slicing a two step approach. \n\nCan you look into this? According to the data description, the sample id has been dropped. So, It will be nice that the tuple just gives the data and not the id as a tuple ",
      "replies": [
        {
          "id": 2778388,
          "postDate": "2024-04-27T06:23:44.233Z",
          "content": "<p>Right. The index column was inserted by Pandas.</p>\n<p>Fortunately, field access is possible. The following code reads selected rows of the dataset without indexes.</p>\n<pre><code> h5py.File(PATH_TO_H5FILE)  hf:\n    array = hf[][].fields()[ROW_INDEXES]\n</code></pre>",
          "rawMarkdown": "Right. The index column was inserted by Pandas.\n\nFortunately, field access is possible. The following code reads selected rows of the dataset without indexes.\n\n```python\nwith h5py.File(PATH_TO_H5FILE) as hf:\n    array = hf['data']['table'].fields('values_block_0')[ROW_INDEXES]\n```\n",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2778261,
      "author_name": "HungryLearner",
      "author_url": "",
      "post_date": "2024-04-27T04:32:53.747000",
      "content": "<p><a href=\"https://www.kaggle.com/taichiuemura\" target=\"_blank\">@taichiuemura</a> , the dataset looks great. However, I noticed that indexing trough the data produces a tuple of (sample_id, sample_data) which makes data slicing a two step approach. </p>\n<p>Can you look into this? According to the data description, the sample id has been dropped. So, It will be nice that the tuple just gives the data and not the id as a tuple </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2778388,
          "author_name": "Taichi Uemura",
          "author_url": "",
          "post_date": "2024-04-27T06:23:44.233000",
          "content": "<p>Right. The index column was inserted by Pandas.</p>\n<p>Fortunately, field access is possible. The following code reads selected rows of the dataset without indexes.</p>\n<pre><code> h5py.File(PATH_TO_H5FILE)  hf:\n    array = hf[][].fields()[ROW_INDEXES]\n</code></pre>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2771364": "HI,\n\nI created an HDF5 version of the training data <https://www.kaggle.com/datasets/taichiuemura/leap-climsim-hdf5-training-data/>.\n\nMy motivation is to make random row access easier, but I believe 37GB HDF5 file is easier to work with than 180GB CSV file for any purpose.\n\nI also made a public notebook using the dataset <https://www.kaggle.com/code/taichiuemura/leap-mlp/>.",
    "2778261": "@taichiuemura , the dataset looks great. However, I noticed that indexing trough the data produces a tuple of (sample_id, sample_data) which makes data slicing a two step approach. \n\nCan you look into this? According to the data description, the sample id has been dropped. So, It will be nice that the tuple just gives the data and not the id as a tuple "
  }
}