{
  "id": 417759,
  "title": "Data doesn't load properly",
  "url": "/competitions/google-research-identify-contrails-reduce-global-warming/discussion/417759",
  "author_name": "Cartid",
  "post_date": "2023-06-17T03:19:48.449000",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I'm somewhat new to Kaggle so apologies if the question seems stupid. </p>\n<p>I've tried the following code in slight variations for a while now, and for some reason the \"run\" button next to the cell is forever stuck at \"loading.\"</p>\n<p>Clearly there is some kind of memory issue, and definitely 28000 npy files is a lot, but how does one preprocess the data then? Using np.load seems the fastest way but is there a caveat I'm not seeing?</p>\n<p>Thanks in advance for any help.</p>\n<p>import numpy as np # linear algebra<br>\nimport os</p>\n<p>dir = '/kaggle/input/google-research-identify-contrails-reduce-global-warming/train'</p>\n<p>for folder in os.listdir(dir):<br>\n    foldpath = os.path.join(dir, folder) <br>\n    np_array = [np.load(os.path.join(foldpath, file)) for file in os.listdir(foldpath)]</p>",
  "messages": [
    {
      "id": 2307642,
      "postDate": "2023-06-18T10:55:11.093Z",
      "content": "<p>You are downloading all the bands of all the timeframes for each scene.<br>\nThats too miuch data and that makes your NB crash. The solution is to download 3 bands for the 5th timeframe only and turn it into a ASHRGB image and then work with that.<br>\nSee public notebooks such as <a href=\"https://www.kaggle.com/shashwatraman\" target=\"_blank\">@shashwatraman</a>'s baseline notebook which should clarify a lot of things for you.</p>\n<ul>\n<li><strong>making the dataset :</strong> <a href=\"https://www.kaggle.com/code/shashwatraman/contrails-dataset-ash-color\" target=\"_blank\">https://www.kaggle.com/code/shashwatraman/contrails-dataset-ash-color</a></li>\n<li><strong>training the model :</strong> <a href=\"https://www.kaggle.com/code/shashwatraman/simple-unet-baseline-infer-lb-0-580\" target=\"_blank\">https://www.kaggle.com/code/shashwatraman/simple-unet-baseline-infer-lb-0-580</a></li>\n<li><strong>submitting the model :</strong> <a href=\"https://www.kaggle.com/code/shashwatraman/simple-unet-baseline-train-lb-0-580\" target=\"_blank\">https://www.kaggle.com/code/shashwatraman/simple-unet-baseline-train-lb-0-580</a></li>\n</ul>",
      "rawMarkdown": "You are downloading all the bands of all the timeframes for each scene.\nThats too miuch data and that makes your NB crash. The solution is to download 3 bands for the 5th timeframe only and turn it into a ASHRGB image and then work with that.\nSee public notebooks such as @shashwatraman's baseline notebook which should clarify a lot of things for you.\n- **making the dataset :** https://www.kaggle.com/code/shashwatraman/contrails-dataset-ash-color\n- **training the model :** https://www.kaggle.com/code/shashwatraman/simple-unet-baseline-infer-lb-0-580\n- **submitting the model :** https://www.kaggle.com/code/shashwatraman/simple-unet-baseline-train-lb-0-580",
      "votes": 1
    },
    {
      "id": 2306729,
      "postDate": "2023-06-17T14:13:33.063Z",
      "content": "<p>It is looks like you are trying to load 300 GB data to 15 GB memory of notebook. This is impossible. Try to use some dataloader </p>",
      "rawMarkdown": "It is looks like you are trying to load 300 GB data to 15 GB memory of notebook. This is impossible. Try to use some dataloader ",
      "replies": [
        {
          "id": 2306857,
          "postDate": "2023-06-17T16:46:26.283Z",
          "content": "<p>So, does anyone use np.load? If not, why is it there? </p>\n<p>Thanks for the advice, much appreciated. </p>",
          "rawMarkdown": "So, does anyone use np.load? If not, why is it there? \n\nThanks for the advice, much appreciated. ",
          "replies": [
            {
              "id": 2309864,
              "postDate": "2023-06-20T00:16:40.553Z",
              "content": "<p>Everyone uses np.load but in different way.</p>\n<p>There is a concept of iterator/generator, when you go through all the elements of the list, but do not keep them in memory at the same time.</p>\n<pre><code> p  paths:\n    entry = np.load(p)\n    process(entry) \n</code></pre>\n<p>When you write something like <code>[np.load(p) for p in paths]</code> you store ALL the data into memory in the same time.</p>",
              "rawMarkdown": "Everyone uses np.load but in different way.\n\nThere is a concept of iterator/generator, when you go through all the elements of the list, but do not keep them in memory at the same time.\n\n```python\nfor p in paths:\n    entry = np.load(p)\n    process(entry) # do not store in memory heavy things here, just process, mb write on disk and forget\n```\n\nWhen you write something like `[np.load(p) for p in paths]` you store ALL the data into memory in the same time."
            }
          ]
        }
      ]
    },
    {
      "id": 2305970,
      "postDate": "2023-06-17T03:19:48.450Z",
      "content": "<p>I'm somewhat new to Kaggle so apologies if the question seems stupid. </p>\n<p>I've tried the following code in slight variations for a while now, and for some reason the \"run\" button next to the cell is forever stuck at \"loading.\"</p>\n<p>Clearly there is some kind of memory issue, and definitely 28000 npy files is a lot, but how does one preprocess the data then? Using np.load seems the fastest way but is there a caveat I'm not seeing?</p>\n<p>Thanks in advance for any help.</p>\n<p>import numpy as np # linear algebra<br>\nimport os</p>\n<p>dir = '/kaggle/input/google-research-identify-contrails-reduce-global-warming/train'</p>\n<p>for folder in os.listdir(dir):<br>\n    foldpath = os.path.join(dir, folder) <br>\n    np_array = [np.load(os.path.join(foldpath, file)) for file in os.listdir(foldpath)]</p>",
      "rawMarkdown": "I'm somewhat new to Kaggle so apologies if the question seems stupid. \n\nI've tried the following code in slight variations for a while now, and for some reason the \"run\" button next to the cell is forever stuck at \"loading.\"\n\nClearly there is some kind of memory issue, and definitely 28000 npy files is a lot, but how does one preprocess the data then? Using np.load seems the fastest way but is there a caveat I'm not seeing?\n\nThanks in advance for any help.\n\n\n\n\nimport numpy as np # linear algebra\nimport os\n\ndir = '/kaggle/input/google-research-identify-contrails-reduce-global-warming/train'\n\n\n\nfor folder in os.listdir(dir):\n    foldpath = os.path.join(dir, folder) \n    np_array = [np.load(os.path.join(foldpath, file)) for file in os.listdir(foldpath)]\n       \n    \n\n"
    }
  ],
  "comments": [
    {
      "id": 2307642,
      "author_name": "JEANMPIA",
      "author_url": "",
      "post_date": "2023-06-18T10:55:11.093000",
      "content": "<p>You are downloading all the bands of all the timeframes for each scene.<br>\nThats too miuch data and that makes your NB crash. The solution is to download 3 bands for the 5th timeframe only and turn it into a ASHRGB image and then work with that.<br>\nSee public notebooks such as <a href=\"https://www.kaggle.com/shashwatraman\" target=\"_blank\">@shashwatraman</a>'s baseline notebook which should clarify a lot of things for you.</p>\n<ul>\n<li><strong>making the dataset :</strong> <a href=\"https://www.kaggle.com/code/shashwatraman/contrails-dataset-ash-color\" target=\"_blank\">https://www.kaggle.com/code/shashwatraman/contrails-dataset-ash-color</a></li>\n<li><strong>training the model :</strong> <a href=\"https://www.kaggle.com/code/shashwatraman/simple-unet-baseline-infer-lb-0-580\" target=\"_blank\">https://www.kaggle.com/code/shashwatraman/simple-unet-baseline-infer-lb-0-580</a></li>\n<li><strong>submitting the model :</strong> <a href=\"https://www.kaggle.com/code/shashwatraman/simple-unet-baseline-train-lb-0-580\" target=\"_blank\">https://www.kaggle.com/code/shashwatraman/simple-unet-baseline-train-lb-0-580</a></li>\n</ul>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2306729,
      "author_name": "Aleksandr Lavrikov",
      "author_url": "",
      "post_date": "2023-06-17T14:13:33.063000",
      "content": "<p>It is looks like you are trying to load 300 GB data to 15 GB memory of notebook. This is impossible. Try to use some dataloader </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2306857,
          "author_name": "Cartid",
          "author_url": "",
          "post_date": "2023-06-17T16:46:26.283000",
          "content": "<p>So, does anyone use np.load? If not, why is it there? </p>\n<p>Thanks for the advice, much appreciated. </p>",
          "votes": 0,
          "replies": [
            {
              "id": 2309864,
              "author_name": "Alex Ozerin",
              "author_url": "",
              "post_date": "2023-06-20T00:16:40.553000",
              "content": "<p>Everyone uses np.load but in different way.</p>\n<p>There is a concept of iterator/generator, when you go through all the elements of the list, but do not keep them in memory at the same time.</p>\n<pre><code> p  paths:\n    entry = np.load(p)\n    process(entry) \n</code></pre>\n<p>When you write something like <code>[np.load(p) for p in paths]</code> you store ALL the data into memory in the same time.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2307642": "You are downloading all the bands of all the timeframes for each scene.\nThats too miuch data and that makes your NB crash. The solution is to download 3 bands for the 5th timeframe only and turn it into a ASHRGB image and then work with that.\nSee public notebooks such as @shashwatraman's baseline notebook which should clarify a lot of things for you.\n- **making the dataset :** https://www.kaggle.com/code/shashwatraman/contrails-dataset-ash-color\n- **training the model :** https://www.kaggle.com/code/shashwatraman/simple-unet-baseline-infer-lb-0-580\n- **submitting the model :** https://www.kaggle.com/code/shashwatraman/simple-unet-baseline-train-lb-0-580",
    "2306729": "It is looks like you are trying to load 300 GB data to 15 GB memory of notebook. This is impossible. Try to use some dataloader ",
    "2305970": "I'm somewhat new to Kaggle so apologies if the question seems stupid. \n\nI've tried the following code in slight variations for a while now, and for some reason the \"run\" button next to the cell is forever stuck at \"loading.\"\n\nClearly there is some kind of memory issue, and definitely 28000 npy files is a lot, but how does one preprocess the data then? Using np.load seems the fastest way but is there a caveat I'm not seeing?\n\nThanks in advance for any help.\n\n\n\n\nimport numpy as np # linear algebra\nimport os\n\ndir = '/kaggle/input/google-research-identify-contrails-reduce-global-warming/train'\n\n\n\nfor folder in os.listdir(dir):\n    foldpath = os.path.join(dir, folder) \n    np_array = [np.load(os.path.join(foldpath, file)) for file in os.listdir(foldpath)]\n       \n    \n\n"
  }
}