{
  "id": 410316,
  "title": "Super Fast Data Loading but with some tradeoff",
  "url": "/competitions/google-research-identify-contrails-reduce-global-warming/discussion/410316",
  "author_name": "Soumyadeep Khandual",
  "post_date": "2023-05-14T21:06:32.553000",
  "votes": 4,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi,<br>\nData loading for this dataset has been slow. I think the primary cause is, for each example - target pair we have to read 10+ npy files. This in my opinion is the bottle-neck. To reduce this high number of IO operations, I have created a fork of the dataset, that reduces the IO by a factor of 10.</p>\n<p>forked dataset <br>\npart1 - <a href=\"https://www.kaggle.com/datasets/soumyadeepkhandual/google-contrails-normalized-float16-part-1\" target=\"_blank\">https://www.kaggle.com/datasets/soumyadeepkhandual/google-contrails-normalized-float16-part-1</a><br>\npart2 - <a href=\"https://www.kaggle.com/datasets/soumyadeepkhandual/google-contrails-normalized-float16-part-2\" target=\"_blank\">https://www.kaggle.com/datasets/soumyadeepkhandual/google-contrails-normalized-float16-part-2</a><br>\npart3 - <a href=\"https://www.kaggle.com/datasets/soumyadeepkhandual/google-contrails-normalized-float16-part-3\" target=\"_blank\">https://www.kaggle.com/datasets/soumyadeepkhandual/google-contrails-normalized-float16-part-3</a><br>\npart4 - <a href=\"https://www.kaggle.com/datasets/soumyadeepkhandual/google-contrails-normalized-float16-part-4\" target=\"_blank\">https://www.kaggle.com/datasets/soumyadeepkhandual/google-contrails-normalized-float16-part-4</a><br>\npart5 - <a href=\"https://www.kaggle.com/datasets/soumyadeepkhandual/google-contrails-normalized-float16-part-5\" target=\"_blank\">https://www.kaggle.com/datasets/soumyadeepkhandual/google-contrails-normalized-float16-part-5</a></p>\n<p>Features of this dataset:</p>\n<ul>\n<li>about 55% of the original dataset did not contain contrails which have been removed here.</li>\n<li>each example target pair is saved in a single npz file, instead of 10 npy files, making it load much faster with a lot less cpu usage, since number of IO operation are 10 times less.</li>\n<li>Examples stored in float16. This saves a lot of space. The loss of information from using float16 instead of float32 (mean absolute reconstruction error) was less than 0.019%.</li>\n<li>All Examples have been normalized as per channel mean and std.</li>\n</ul>\n<p>Disadvantages:</p>\n<ul>\n<li>Individual human mask have been omitted.</li>\n<li>The examples without contrails have been omitted</li>\n<li>Information lost (less than 0.019%) due to the use of float16 </li>\n</ul>\n<p>How to use - <a href=\"https://www.kaggle.com/code/soumyadeepkhandual/superfast-dataloading\" target=\"_blank\">https://www.kaggle.com/code/soumyadeepkhandual/superfast-dataloading</a></p>\n<p>** The dataset is split over 5 parts because of the 20GB limit, each part contains 2000 example-target pairs saved in npz file. All parts have been added.</p>\n<p>If you find any mistakes, please let me know. If the dataset was helpful, upvote, so that people can find this out.</p>",
  "messages": [
    {
      "id": 2259314,
      "postDate": "2023-05-14T21:06:32.553Z",
      "content": "<p>Hi,<br>\nData loading for this dataset has been slow. I think the primary cause is, for each example - target pair we have to read 10+ npy files. This in my opinion is the bottle-neck. To reduce this high number of IO operations, I have created a fork of the dataset, that reduces the IO by a factor of 10.</p>\n<p>forked dataset <br>\npart1 - <a href=\"https://www.kaggle.com/datasets/soumyadeepkhandual/google-contrails-normalized-float16-part-1\" target=\"_blank\">https://www.kaggle.com/datasets/soumyadeepkhandual/google-contrails-normalized-float16-part-1</a><br>\npart2 - <a href=\"https://www.kaggle.com/datasets/soumyadeepkhandual/google-contrails-normalized-float16-part-2\" target=\"_blank\">https://www.kaggle.com/datasets/soumyadeepkhandual/google-contrails-normalized-float16-part-2</a><br>\npart3 - <a href=\"https://www.kaggle.com/datasets/soumyadeepkhandual/google-contrails-normalized-float16-part-3\" target=\"_blank\">https://www.kaggle.com/datasets/soumyadeepkhandual/google-contrails-normalized-float16-part-3</a><br>\npart4 - <a href=\"https://www.kaggle.com/datasets/soumyadeepkhandual/google-contrails-normalized-float16-part-4\" target=\"_blank\">https://www.kaggle.com/datasets/soumyadeepkhandual/google-contrails-normalized-float16-part-4</a><br>\npart5 - <a href=\"https://www.kaggle.com/datasets/soumyadeepkhandual/google-contrails-normalized-float16-part-5\" target=\"_blank\">https://www.kaggle.com/datasets/soumyadeepkhandual/google-contrails-normalized-float16-part-5</a></p>\n<p>Features of this dataset:</p>\n<ul>\n<li>about 55% of the original dataset did not contain contrails which have been removed here.</li>\n<li>each example target pair is saved in a single npz file, instead of 10 npy files, making it load much faster with a lot less cpu usage, since number of IO operation are 10 times less.</li>\n<li>Examples stored in float16. This saves a lot of space. The loss of information from using float16 instead of float32 (mean absolute reconstruction error) was less than 0.019%.</li>\n<li>All Examples have been normalized as per channel mean and std.</li>\n</ul>\n<p>Disadvantages:</p>\n<ul>\n<li>Individual human mask have been omitted.</li>\n<li>The examples without contrails have been omitted</li>\n<li>Information lost (less than 0.019%) due to the use of float16 </li>\n</ul>\n<p>How to use - <a href=\"https://www.kaggle.com/code/soumyadeepkhandual/superfast-dataloading\" target=\"_blank\">https://www.kaggle.com/code/soumyadeepkhandual/superfast-dataloading</a></p>\n<p>** The dataset is split over 5 parts because of the 20GB limit, each part contains 2000 example-target pairs saved in npz file. All parts have been added.</p>\n<p>If you find any mistakes, please let me know. If the dataset was helpful, upvote, so that people can find this out.</p>",
      "rawMarkdown": "Hi,\nData loading for this dataset has been slow. I think the primary cause is, for each example - target pair we have to read 10+ npy files. This in my opinion is the bottle-neck. To reduce this high number of IO operations, I have created a fork of the dataset, that reduces the IO by a factor of 10.\n\nforked dataset \npart1 - https://www.kaggle.com/datasets/soumyadeepkhandual/google-contrails-normalized-float16-part-1\npart2 - https://www.kaggle.com/datasets/soumyadeepkhandual/google-contrails-normalized-float16-part-2\npart3 - https://www.kaggle.com/datasets/soumyadeepkhandual/google-contrails-normalized-float16-part-3\npart4 - https://www.kaggle.com/datasets/soumyadeepkhandual/google-contrails-normalized-float16-part-4\npart5 - https://www.kaggle.com/datasets/soumyadeepkhandual/google-contrails-normalized-float16-part-5\n\nFeatures of this dataset:\n- about 55% of the original dataset did not contain contrails which have been removed here.\n- each example target pair is saved in a single npz file, instead of 10 npy files, making it load much faster with a lot less cpu usage, since number of IO operation are 10 times less.\n- Examples stored in float16. This saves a lot of space. The loss of information from using float16 instead of float32 (mean absolute reconstruction error) was less than 0.019%.\n- All Examples have been normalized as per channel mean and std.\n\nDisadvantages:\n- Individual human mask have been omitted.\n- The examples without contrails have been omitted\n- Information lost (less than 0.019%) due to the use of float16 \n\nHow to use - https://www.kaggle.com/code/soumyadeepkhandual/superfast-dataloading\n\n** The dataset is split over 5 parts because of the 20GB limit, each part contains 2000 example-target pairs saved in npz file. All parts have been added.\n\nIf you find any mistakes, please let me know. If the dataset was helpful, upvote, so that people can find this out.",
      "votes": 3
    },
    {
      "id": 2260135,
      "postDate": "2023-05-15T13:35:36.873Z",
      "content": "<p>Concatenating the arrays and changing the dtype is in my opinion very usefull. But i don't think that it's a good idea to remove images without contrails. These images still contain information (that this is no contrail) and the test dataset will have the same label distribution</p>",
      "rawMarkdown": "Concatenating the arrays and changing the dtype is in my opinion very usefull. But i don't think that it's a good idea to remove images without contrails. These images still contain information (that this is no contrail) and the test dataset will have the same label distribution"
    },
    {
      "id": 2259605,
      "postDate": "2023-05-15T05:47:47.917Z",
      "content": "<p>The shape of the tensor is [1, 8, 9, 256, 256]. How can you use this dataset to train a model? </p>",
      "rawMarkdown": "The shape of the tensor is [1, 8, 9, 256, 256]. How can you use this dataset to train a model? ",
      "replies": [
        {
          "id": 2259615,
          "postDate": "2023-05-15T06:01:40.277Z",
          "content": "<p>Output of the dataloader is of shape (Batch_size, TimeFrame, Bands, H, W). So i my case i used batch size 1 for demonstration purpose, you can change as per your choice. There are 8 time frames for each band. There are total 9 bands and the image shape is 256 x 256. This makes the shape (1, 8, 9, 256, 256).</p>",
          "rawMarkdown": "Output of the dataloader is of shape (Batch_size, TimeFrame, Bands, H, W). So i my case i used batch size 1 for demonstration purpose, you can change as per your choice. There are 8 time frames for each band. There are total 9 bands and the image shape is 256 x 256. This makes the shape (1, 8, 9, 256, 256)."
        },
        {
          "id": 2260131,
          "postDate": "2023-05-15T13:32:51.863Z",
          "content": "<p>CNN's work with 4D Datasets (3D for image channels and one dimension for batches). the 5th dimension is the time dimension, so you probably have to train a RNN-CNN combination</p>",
          "rawMarkdown": "CNN's work with 4D Datasets (3D for image channels and one dimension for batches). the 5th dimension is the time dimension, so you probably have to train a RNN-CNN combination"
        }
      ]
    },
    {
      "id": 2259498,
      "postDate": "2023-05-15T03:30:16.067Z",
      "content": "<p>Waiting for you updating and thanks for you sharing.</p>",
      "rawMarkdown": "Waiting for you updating and thanks for you sharing.",
      "replies": [
        {
          "id": 2259654,
          "postDate": "2023-05-15T06:42:36.763Z",
          "content": "<p>Hey, All parts have been added. How to use : check this <a href=\"https://www.kaggle.com/code/soumyadeepkhandual/superfast-dataloading\" target=\"_blank\">notebook</a> </p>",
          "rawMarkdown": "Hey, All parts have been added. How to use : check this [notebook](https://www.kaggle.com/code/soumyadeepkhandual/superfast-dataloading) "
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2260135,
      "author_name": "Patchef",
      "author_url": "",
      "post_date": "2023-05-15T13:35:36.873000",
      "content": "<p>Concatenating the arrays and changing the dtype is in my opinion very usefull. But i don't think that it's a good idea to remove images without contrails. These images still contain information (that this is no contrail) and the test dataset will have the same label distribution</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2259605,
      "author_name": "Harsha Arya",
      "author_url": "",
      "post_date": "2023-05-15T05:47:47.917000",
      "content": "<p>The shape of the tensor is [1, 8, 9, 256, 256]. How can you use this dataset to train a model? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2259615,
          "author_name": "Soumyadeep Khandual",
          "author_url": "",
          "post_date": "2023-05-15T06:01:40.277000",
          "content": "<p>Output of the dataloader is of shape (Batch_size, TimeFrame, Bands, H, W). So i my case i used batch size 1 for demonstration purpose, you can change as per your choice. There are 8 time frames for each band. There are total 9 bands and the image shape is 256 x 256. This makes the shape (1, 8, 9, 256, 256).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2260131,
          "author_name": "Patchef",
          "author_url": "",
          "post_date": "2023-05-15T13:32:51.863000",
          "content": "<p>CNN's work with 4D Datasets (3D for image channels and one dimension for batches). the 5th dimension is the time dimension, so you probably have to train a RNN-CNN combination</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2259498,
      "author_name": "Dewei Chen",
      "author_url": "",
      "post_date": "2023-05-15T03:30:16.067000",
      "content": "<p>Waiting for you updating and thanks for you sharing.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2259654,
          "author_name": "Soumyadeep Khandual",
          "author_url": "",
          "post_date": "2023-05-15T06:42:36.763000",
          "content": "<p>Hey, All parts have been added. How to use : check this <a href=\"https://www.kaggle.com/code/soumyadeepkhandual/superfast-dataloading\" target=\"_blank\">notebook</a> </p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2259314": "Hi,\nData loading for this dataset has been slow. I think the primary cause is, for each example - target pair we have to read 10+ npy files. This in my opinion is the bottle-neck. To reduce this high number of IO operations, I have created a fork of the dataset, that reduces the IO by a factor of 10.\n\nforked dataset \npart1 - https://www.kaggle.com/datasets/soumyadeepkhandual/google-contrails-normalized-float16-part-1\npart2 - https://www.kaggle.com/datasets/soumyadeepkhandual/google-contrails-normalized-float16-part-2\npart3 - https://www.kaggle.com/datasets/soumyadeepkhandual/google-contrails-normalized-float16-part-3\npart4 - https://www.kaggle.com/datasets/soumyadeepkhandual/google-contrails-normalized-float16-part-4\npart5 - https://www.kaggle.com/datasets/soumyadeepkhandual/google-contrails-normalized-float16-part-5\n\nFeatures of this dataset:\n- about 55% of the original dataset did not contain contrails which have been removed here.\n- each example target pair is saved in a single npz file, instead of 10 npy files, making it load much faster with a lot less cpu usage, since number of IO operation are 10 times less.\n- Examples stored in float16. This saves a lot of space. The loss of information from using float16 instead of float32 (mean absolute reconstruction error) was less than 0.019%.\n- All Examples have been normalized as per channel mean and std.\n\nDisadvantages:\n- Individual human mask have been omitted.\n- The examples without contrails have been omitted\n- Information lost (less than 0.019%) due to the use of float16 \n\nHow to use - https://www.kaggle.com/code/soumyadeepkhandual/superfast-dataloading\n\n** The dataset is split over 5 parts because of the 20GB limit, each part contains 2000 example-target pairs saved in npz file. All parts have been added.\n\nIf you find any mistakes, please let me know. If the dataset was helpful, upvote, so that people can find this out.",
    "2260135": "Concatenating the arrays and changing the dtype is in my opinion very usefull. But i don't think that it's a good idea to remove images without contrails. These images still contain information (that this is no contrail) and the test dataset will have the same label distribution",
    "2259605": "The shape of the tensor is [1, 8, 9, 256, 256]. How can you use this dataset to train a model? ",
    "2259498": "Waiting for you updating and thanks for you sharing."
  }
}