{
  "id": 409332,
  "title": "How to handle big datasets",
  "url": "/competitions/google-research-identify-contrails-reduce-global-warming/discussion/409332",
  "author_name": "Beubeu4L",
  "post_date": "2023-05-10T15:25:49.291000",
  "votes": 4,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p>The dataset is big (450 Gb). How can we deal with such big dataset with limited resources? Please share any valuable resources that can help :) </p>",
  "messages": [
    {
      "id": 2253998,
      "postDate": "2023-05-10T15:25:49.290Z",
      "content": "<p>Hi everyone,</p>\n<p>The dataset is big (450 Gb). How can we deal with such big dataset with limited resources? Please share any valuable resources that can help :) </p>",
      "rawMarkdown": "Hi everyone,\n\nThe dataset is big (450 Gb). How can we deal with such big dataset with limited resources? Please share any valuable resources that can help :) ",
      "votes": 3
    },
    {
      "id": 2254014,
      "postDate": "2023-05-10T15:40:46.630Z",
      "content": "<p>You don't need to load the full dataset in memory.  It is enough to load it in batches. Additionally, the dataset consists of *.npy files and therefore it will be efficiently handled by standard dataloaders.</p>",
      "rawMarkdown": "You don't need to load the full dataset in memory.  It is enough to load it in batches. Additionally, the dataset consists of *.npy files and therefore it will be efficiently handled by standard dataloaders.",
      "votes": 4
    },
    {
      "id": 2255449,
      "postDate": "2023-05-11T18:06:37.183Z",
      "content": "<p>I wrote a dataloader for Pytorch which loads one batch at a time.</p>\n<p>It might interest you:<br>\n<a href=\"https://www.kaggle.com/code/thomasrochefort/pytorch-dataloader-example\" target=\"_blank\">https://www.kaggle.com/code/thomasrochefort/pytorch-dataloader-example</a></p>\n<p>Not sure if its an efficient way to do this, I now you can define parallel workers in the dataloader to have parallel data fetching. Might be worth looking into!</p>",
      "rawMarkdown": "I wrote a dataloader for Pytorch which loads one batch at a time.\n\nIt might interest you:\nhttps://www.kaggle.com/code/thomasrochefort/pytorch-dataloader-example\n\nNot sure if its an efficient way to do this, I now you can define parallel workers in the dataloader to have parallel data fetching. Might be worth looking into!\n",
      "votes": 1,
      "replies": [
        {
          "id": 2256530,
          "postDate": "2023-05-12T14:51:06.370Z",
          "content": "<p>Can I use this to load only a smaller section of the data? </p>",
          "rawMarkdown": "Can I use this to load only a smaller section of the data? ",
          "replies": [
            {
              "id": 2261699,
              "postDate": "2023-05-16T14:02:56.267Z",
              "content": "<p><a href=\"https://www.kaggle.com/snowclipsed\" target=\"_blank\">@snowclipsed</a> you can use any batch size you'd like, but you would have to modify the Dataset class to not read the entire directory of the dataset.</p>",
              "rawMarkdown": "@snowclipsed you can use any batch size you'd like, but you would have to modify the Dataset class to not read the entire directory of the dataset."
            }
          ]
        }
      ]
    },
    {
      "id": 2254367,
      "postDate": "2023-05-10T22:05:05.623Z",
      "content": "<p>For big data projects, it's always a good idea to process data as chunks.</p>",
      "rawMarkdown": "For big data projects, it's always a good idea to process data as chunks.",
      "votes": 1
    },
    {
      "id": 2257107,
      "postDate": "2023-05-13T03:43:00.847Z",
      "content": "<p>Capture your environment…</p>",
      "rawMarkdown": "Capture your environment..."
    },
    {
      "id": 2255364,
      "postDate": "2023-05-11T17:14:51.950Z",
      "content": "<p>Is there a way to download just a part of the training data?  It would be nice if there were a 5 or 10 GB download for testing before committing to the full download.</p>",
      "rawMarkdown": "Is there a way to download just a part of the training data?  It would be nice if there were a 5 or 10 GB download for testing before committing to the full download.",
      "replies": [
        {
          "id": 2256676,
          "postDate": "2023-05-12T16:44:22.453Z",
          "content": "<p>you can use kaggle notebook</p>",
          "rawMarkdown": "you can use kaggle notebook",
          "votes": 1,
          "replies": [
            {
              "id": 2256874,
              "postDate": "2023-05-12T19:46:11.353Z",
              "content": "<p>That sounds like the best option.</p>",
              "rawMarkdown": "That sounds like the best option."
            }
          ]
        }
      ]
    },
    {
      "id": 2254275,
      "postDate": "2023-05-10T19:30:59.870Z",
      "content": "<p>A python package that I use to handle big datasets is h5py, which is a wrapper for handling data in HDF5 format. You can find more information about this package <a href=\"https://github.com/h5py/h5py\" target=\"_blank\">here</a>. </p>",
      "rawMarkdown": "A python package that I use to handle big datasets is h5py, which is a wrapper for handling data in HDF5 format. You can find more information about this package [here](https://github.com/h5py/h5py). ",
      "votes": 2,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2254014,
      "author_name": "nymfree",
      "author_url": "",
      "post_date": "2023-05-10T15:40:46.630000",
      "content": "<p>You don't need to load the full dataset in memory.  It is enough to load it in batches. Additionally, the dataset consists of *.npy files and therefore it will be efficiently handled by standard dataloaders.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 2255449,
      "author_name": "Thomas Rochefort-Beaudoin",
      "author_url": "",
      "post_date": "2023-05-11T18:06:37.183000",
      "content": "<p>I wrote a dataloader for Pytorch which loads one batch at a time.</p>\n<p>It might interest you:<br>\n<a href=\"https://www.kaggle.com/code/thomasrochefort/pytorch-dataloader-example\" target=\"_blank\">https://www.kaggle.com/code/thomasrochefort/pytorch-dataloader-example</a></p>\n<p>Not sure if its an efficient way to do this, I now you can define parallel workers in the dataloader to have parallel data fetching. Might be worth looking into!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2256530,
          "author_name": "snowclipsed",
          "author_url": "",
          "post_date": "2023-05-12T14:51:06.370000",
          "content": "<p>Can I use this to load only a smaller section of the data? </p>",
          "votes": 0,
          "replies": [
            {
              "id": 2261699,
              "author_name": "Thomas Rochefort-Beaudoin",
              "author_url": "",
              "post_date": "2023-05-16T14:02:56.267000",
              "content": "<p><a href=\"https://www.kaggle.com/snowclipsed\" target=\"_blank\">@snowclipsed</a> you can use any batch size you'd like, but you would have to modify the Dataset class to not read the entire directory of the dataset.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2254367,
      "author_name": "Francisco Segura",
      "author_url": "",
      "post_date": "2023-05-10T22:05:05.623000",
      "content": "<p>For big data projects, it's always a good idea to process data as chunks.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2257107,
      "author_name": "Aisuluu Ulan kyzy",
      "author_url": "",
      "post_date": "2023-05-13T03:43:00.847000",
      "content": "<p>Capture your environment…</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2255364,
      "author_name": "gardn999",
      "author_url": "",
      "post_date": "2023-05-11T17:14:51.950000",
      "content": "<p>Is there a way to download just a part of the training data?  It would be nice if there were a 5 or 10 GB download for testing before committing to the full download.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2256676,
          "author_name": "Rounak Kumbhakar",
          "author_url": "",
          "post_date": "2023-05-12T16:44:22.453000",
          "content": "<p>you can use kaggle notebook</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2256874,
              "author_name": "gardn999",
              "author_url": "",
              "post_date": "2023-05-12T19:46:11.353000",
              "content": "<p>That sounds like the best option.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2254275,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-05-10T19:30:59.870000",
      "content": "<p>A python package that I use to handle big datasets is h5py, which is a wrapper for handling data in HDF5 format. You can find more information about this package <a href=\"https://github.com/h5py/h5py\" target=\"_blank\">here</a>. </p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2253998": "Hi everyone,\n\nThe dataset is big (450 Gb). How can we deal with such big dataset with limited resources? Please share any valuable resources that can help :) ",
    "2254014": "You don't need to load the full dataset in memory.  It is enough to load it in batches. Additionally, the dataset consists of *.npy files and therefore it will be efficiently handled by standard dataloaders.",
    "2255449": "I wrote a dataloader for Pytorch which loads one batch at a time.\n\nIt might interest you:\nhttps://www.kaggle.com/code/thomasrochefort/pytorch-dataloader-example\n\nNot sure if its an efficient way to do this, I now you can define parallel workers in the dataloader to have parallel data fetching. Might be worth looking into!\n",
    "2254367": "For big data projects, it's always a good idea to process data as chunks.",
    "2257107": "Capture your environment...",
    "2255364": "Is there a way to download just a part of the training data?  It would be nice if there were a 5 or 10 GB download for testing before committing to the full download.",
    "2254275": "A python package that I use to handle big datasets is h5py, which is a wrapper for handling data in HDF5 format. You can find more information about this package [here](https://github.com/h5py/h5py). "
  }
}