{
  "id": 428681,
  "title": "Working with large datasets",
  "url": "/competitions/google-research-identify-contrails-reduce-global-warming/discussion/428681",
  "author_name": "Stephen Mulligan",
  "post_date": "2023-08-02T10:59:21.069000",
  "votes": 0,
  "comment_count": 1,
  "views": 0,
  "content": "<p>This dataset is pretty huge, and I'm not really sure what the best approach for working with it. I usually use google colab but due to memory limits downloading the training dataset to there is not an option. Is there a way to batch the download with the kaggle api? That way I could download part, train on this, and repeat the rest of the process for the rest of the dataset. Any help or other suggestions would be much appreciated</p>",
  "messages": [
    {
      "id": 2372185,
      "postDate": "2023-08-03T14:36:38.783Z",
      "content": "<p>Hi Stephen, what it would help you so much is this <a href=\"https://www.kaggle.com/datasets/shashwatraman/contrails-images-ash-color\" target=\"_blank\">dataset</a> from <a href=\"https://www.kaggle.com/shashwatraman\" target=\"_blank\">@shashwatraman</a> that most of the people is currently using. Here you will find all images already processed in ash-color (which makes contrails better perceived in images) and only taking 3 channels out of the 8 bands, the 5th frame which is the target out of 8 time frames and these images are converted to float16 taking half of the memory but keeping almost all the info. Every image also contains its label btw, and there are all train and validation images. </p>\n<p>So it's going <strong>from 450GB</strong> that takes the original entire dataset to <strong>only 12GB</strong></p>",
      "rawMarkdown": "Hi Stephen, what it would help you so much is this [dataset](https://www.kaggle.com/datasets/shashwatraman/contrails-images-ash-color) from @shashwatraman that most of the people is currently using. Here you will find all images already processed in ash-color (which makes contrails better perceived in images) and only taking 3 channels out of the 8 bands, the 5th frame which is the target out of 8 time frames and these images are converted to float16 taking half of the memory but keeping almost all the info. Every image also contains its label btw, and there are all train and validation images. \n\nSo it's going **from 450GB** that takes the original entire dataset to **only 12GB**"
    },
    {
      "id": 2370364,
      "postDate": "2023-08-02T10:59:21.070Z",
      "content": "<p>This dataset is pretty huge, and I'm not really sure what the best approach for working with it. I usually use google colab but due to memory limits downloading the training dataset to there is not an option. Is there a way to batch the download with the kaggle api? That way I could download part, train on this, and repeat the rest of the process for the rest of the dataset. Any help or other suggestions would be much appreciated</p>",
      "rawMarkdown": "This dataset is pretty huge, and I'm not really sure what the best approach for working with it. I usually use google colab but due to memory limits downloading the training dataset to there is not an option. Is there a way to batch the download with the kaggle api? That way I could download part, train on this, and repeat the rest of the process for the rest of the dataset. Any help or other suggestions would be much appreciated"
    }
  ],
  "comments": [
    {
      "id": 2372185,
      "author_name": "Enric Domingo",
      "author_url": "",
      "post_date": "2023-08-03T14:36:38.783000",
      "content": "<p>Hi Stephen, what it would help you so much is this <a href=\"https://www.kaggle.com/datasets/shashwatraman/contrails-images-ash-color\" target=\"_blank\">dataset</a> from <a href=\"https://www.kaggle.com/shashwatraman\" target=\"_blank\">@shashwatraman</a> that most of the people is currently using. Here you will find all images already processed in ash-color (which makes contrails better perceived in images) and only taking 3 channels out of the 8 bands, the 5th frame which is the target out of 8 time frames and these images are converted to float16 taking half of the memory but keeping almost all the info. Every image also contains its label btw, and there are all train and validation images. </p>\n<p>So it's going <strong>from 450GB</strong> that takes the original entire dataset to <strong>only 12GB</strong></p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2372185": "Hi Stephen, what it would help you so much is this [dataset](https://www.kaggle.com/datasets/shashwatraman/contrails-images-ash-color) from @shashwatraman that most of the people is currently using. Here you will find all images already processed in ash-color (which makes contrails better perceived in images) and only taking 3 channels out of the 8 bands, the 5th frame which is the target out of 8 time frames and these images are converted to float16 taking half of the memory but keeping almost all the info. Every image also contains its label btw, and there are all train and validation images. \n\nSo it's going **from 450GB** that takes the original entire dataset to **only 12GB**",
    "2370364": "This dataset is pretty huge, and I'm not really sure what the best approach for working with it. I usually use google colab but due to memory limits downloading the training dataset to there is not an option. Is there a way to batch the download with the kaggle api? That way I could download part, train on this, and repeat the rest of the process for the rest of the dataset. Any help or other suggestions would be much appreciated"
  }
}