{
  "id": 181785,
  "title": "How to handle such huge data?",
  "url": "/competitions/rsna-str-pulmonary-embolism-detection/discussion/181785",
  "author_name": "Akshay Chavan",
  "post_date": "2020-09-10T06:47:29.319000",
  "votes": 5,
  "comment_count": 18,
  "views": 0,
  "content": "<p>Hello,<br>\nI am a newbie and want to work on such a huge amount of data for the very first time. I want to ask you how you used to manage the resources for this much data. I mean Kaggle kernel will show out of memory error while training the model and the same thing goes with the google colab.<br>\nSo I want to ask is there any free cloud platform available for such huge data with GPU system.<br>\nAny kind of suggestions and thoughts would be appreciated for this topic. </p>\n<p>Thank you</p>",
  "messages": [
    {
      "id": 1004952,
      "postDate": "2020-09-10T06:47:29.320Z",
      "content": "<p>Hello,<br>\nI am a newbie and want to work on such a huge amount of data for the very first time. I want to ask you how you used to manage the resources for this much data. I mean Kaggle kernel will show out of memory error while training the model and the same thing goes with the google colab.<br>\nSo I want to ask is there any free cloud platform available for such huge data with GPU system.<br>\nAny kind of suggestions and thoughts would be appreciated for this topic. </p>\n<p>Thank you</p>",
      "rawMarkdown": "Hello,\nI am a newbie and want to work on such a huge amount of data for the very first time. I want to ask you how you used to manage the resources for this much data. I mean Kaggle kernel will show out of memory error while training the model and the same thing goes with the google colab.\nSo I want to ask is there any free cloud platform available for such huge data with GPU system.\nAny kind of suggestions and thoughts would be appreciated for this topic. \n\nThank you",
      "votes": 5
    },
    {
      "id": 1010331,
      "postDate": "2020-09-14T17:14:44.860Z",
      "content": "<p>Personally, I am preprocessing the dcm files into jpg files on kaggle notebooks. I had to split the image preprocessing into 10 separate notebooks for 128x128 - ~3GB per notebooks (5GB is notebook write limit). As for modeling if you are using a dataloader or generator, you shouldn't have memory issues if you choose the right batchsize.</p>",
      "rawMarkdown": "Personally, I am preprocessing the dcm files into jpg files on kaggle notebooks. I had to split the image preprocessing into 10 separate notebooks for 128x128 - ~3GB per notebooks (5GB is notebook write limit). As for modeling if you are using a dataloader or generator, you shouldn't have memory issues if you choose the right batchsize.",
      "votes": 4,
      "replies": [
        {
          "id": 1010343,
          "postDate": "2020-09-14T17:22:51.747Z",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/teeyee314\" target=\"_blank\">@teeyee314</a> for your response.<br>\nI will definitely try your approach as this is a good point to start working on the large datasets.<br>\nBut for long-running notebooks after every 40-minute kernel stops automatically if we aren't active on it.<br>\nSo keeping eye on the notebook to check its execution is difficult if we have multiple notebooks running simultaneously. So how you handle these scenarios?  </p>",
          "rawMarkdown": "Thank you @teeyee314 for your response.\nI will definitely try your approach as this is a good point to start working on the large datasets.\nBut for long-running notebooks after every 40-minute kernel stops automatically if we aren't active on it.\nSo keeping eye on the notebook to check its execution is difficult if we have multiple notebooks running simultaneously. So how you handle these scenarios?  ",
          "replies": [
            {
              "id": 1010351,
              "postDate": "2020-09-14T17:27:09.153Z",
              "content": "<p>You can commit your notebook and it will run unattended for up to 9 hours (3 hours with a TPU). The output of a committed notebook can be put into a Dataset without downloading.</p>",
              "rawMarkdown": "You can commit your notebook and it will run unattended for up to 9 hours (3 hours with a TPU). The output of a committed notebook can be put into a Dataset without downloading."
            }
          ]
        },
        {
          "id": 1010357,
          "postDate": "2020-09-14T17:29:15.893Z",
          "content": "<p>you can save and commit the notebook so it will continue to run. just make sure your script doesn't have bugs. you can run a small test to find bugs before committing the notebook. Also make sure your notebook will run within the time limits (9hr for CPU and GPU, 3hr for TPU)</p>",
          "rawMarkdown": "you can save and commit the notebook so it will continue to run. just make sure your script doesn't have bugs. you can run a small test to find bugs before committing the notebook. Also make sure your notebook will run within the time limits (9hr for CPU and GPU, 3hr for TPU)",
          "votes": 1
        },
        {
          "id": 1010757,
          "postDate": "2020-09-15T03:48:21.630Z",
          "content": "<p>Okay <a href=\"https://www.kaggle.com/teeyee314\" target=\"_blank\">@teeyee314</a> <br>\nThese tips will definitely help me with this problem. </p>",
          "rawMarkdown": "Okay @teeyee314 \nThese tips will definitely help me with this problem. "
        }
      ]
    },
    {
      "id": 1009118,
      "postDate": "2020-09-13T16:58:22.963Z",
      "content": "<p>Can we discard some of the data and train partially by somehow analyzing the metadata in the <code>train.csv</code> file? I know more data means better generalization, however, in order to reduce training time and have a quick jumpstart, maybe it can be good starting point.. </p>",
      "rawMarkdown": "Can we discard some of the data and train partially by somehow analyzing the metadata in the `train.csv` file? I know more data means better generalization, however, in order to reduce training time and have a quick jumpstart, maybe it can be good starting point.. \n",
      "votes": 1,
      "replies": [
        {
          "id": 1009129,
          "postDate": "2020-09-13T17:08:29.480Z",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/redwankarimsony\" target=\"_blank\">@redwankarimsony</a>.<br>\nFor the base models and to start with the problem we can take a sample of data. But if we have do get the better and generalize model then at that time we have to increase the size of the data. So in that anyhow we have to handle this huge data. </p>",
          "rawMarkdown": "Thank you @redwankarimsony.\nFor the base models and to start with the problem we can take a sample of data. But if we have do get the better and generalize model then at that time we have to increase the size of the data. So in that anyhow we have to handle this huge data. ",
          "votes": 1
        },
        {
          "id": 1009144,
          "postDate": "2020-09-13T17:16:09.107Z",
          "content": "<p>Thanks, <a href=\"https://www.kaggle.com/akshaychavan123\" target=\"_blank\">@akshaychavan123</a> <br>\nIf we make all the preprocessing (resizing, normalization, sorting of slices etc) and make that another dataset, I hope it will save everyone a lot of trouble. However, what's the point of being it as a competition? 😃😃</p>\n<p>I hope to see someone will rise up and make the preprocessing and create a separate dataset accordingly of smaller size.. :D </p>",
          "rawMarkdown": "Thanks, @akshaychavan123 \nIf we make all the preprocessing (resizing, normalization, sorting of slices etc) and make that another dataset, I hope it will save everyone a lot of trouble. However, what's the point of being it as a competition? 😃😃\n\nI hope to see someone will rise up and make the preprocessing and create a separate dataset accordingly of smaller size.. :D ",
          "votes": 1
        },
        {
          "id": 1009155,
          "postDate": "2020-09-13T17:21:28.747Z",
          "content": "<p>Yes, <a href=\"https://www.kaggle.com/redwankarimsony\" target=\"_blank\">@redwankarimsony</a> let's hope in the future for those things 😃😃😃.</p>\n<p>So can we say that for such competition if we do not have enough resources then we cannot move forward? I mean not all can afford the paid resources. </p>",
          "rawMarkdown": "Yes, @redwankarimsony let's hope in the future for those things 😃😃😃.\n\nSo can we say that for such competition if we do not have enough resources then we cannot move forward? I mean not all can afford the paid resources. ",
          "votes": 1
        },
        {
          "id": 1011651,
          "postDate": "2020-09-15T15:57:29.990Z",
          "content": "<p>In past experience - RSNA Intracranial, GCP credits were given out and that was just enough to stay competitive for me. Here, I would like to think that Kaggle TPU notebooks can stand a chance. Ultimately I don't think Kaggle GPU's will be able to get very far.</p>",
          "rawMarkdown": "In past experience - RSNA Intracranial, GCP credits were given out and that was just enough to stay competitive for me. Here, I would like to think that Kaggle TPU notebooks can stand a chance. Ultimately I don't think Kaggle GPU's will be able to get very far.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1007176,
      "postDate": "2020-09-11T21:58:09.380Z",
      "content": "<p>How about discard-by-train, means loading only minibatch to GPU memory from disk?</p>",
      "rawMarkdown": "How about discard-by-train, means loading only minibatch to GPU memory from disk?",
      "votes": 1,
      "replies": [
        {
          "id": 1007353,
          "postDate": "2020-09-12T05:23:49.143Z",
          "content": "<p>By doing this it takes a lot of time to train if the size of the batch is small. </p>",
          "rawMarkdown": "By doing this it takes a lot of time to train if the size of the batch is small. "
        }
      ]
    },
    {
      "id": 1004967,
      "postDate": "2020-09-10T07:05:27.503Z",
      "content": "<p>I think TPU is the way to go. </p>",
      "rawMarkdown": "I think TPU is the way to go. ",
      "votes": 2,
      "replies": [
        {
          "id": 1004971,
          "postDate": "2020-09-10T07:07:45.230Z",
          "content": "<p><a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a> Thank you for your response. </p>",
          "rawMarkdown": "@serigne Thank you for your response. "
        }
      ]
    },
    {
      "id": 1004965,
      "postDate": "2020-09-10T07:04:11.517Z",
      "content": "<p>I would recommand you to train with small batches (8-16) therefor you will not run out of memory.</p>",
      "rawMarkdown": "I would recommand you to train with small batches (8-16) therefor you will not run out of memory.",
      "votes": 2,
      "replies": [
        {
          "id": 1004972,
          "postDate": "2020-09-10T07:08:07.843Z",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/docdocb\" target=\"_blank\">@docdocb</a> </p>",
          "rawMarkdown": "Thank you @docdocb "
        },
        {
          "id": 1005001,
          "postDate": "2020-09-10T07:41:55.027Z",
          "content": "<p>Nice idea by Arnaud, but just keep in mind that it would take much more time for training with small batch size than larger batch size. So, adjust your batch size accordingly.</p>",
          "rawMarkdown": "Nice idea by Arnaud, but just keep in mind that it would take much more time for training with small batch size than larger batch size. So, adjust your batch size accordingly.",
          "votes": 1
        },
        {
          "id": 1005018,
          "postDate": "2020-09-10T07:59:37.357Z",
          "content": "<p>Yes, <a href=\"https://www.kaggle.com/oneplustricks\" target=\"_blank\">@oneplustricks</a>. That is why I was asking is there any alternatives for this. Because kernels also stop every 40 minutes if we are not active on it. And with such a small batch size, it takes a lot of time to train the model. </p>",
          "rawMarkdown": "Yes, @oneplustricks. That is why I was asking is there any alternatives for this. Because kernels also stop every 40 minutes if we are not active on it. And with such a small batch size, it takes a lot of time to train the model. ",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1010331,
      "author_name": "Tim Yee",
      "author_url": "",
      "post_date": "2020-09-14T17:14:44.860000",
      "content": "<p>Personally, I am preprocessing the dcm files into jpg files on kaggle notebooks. I had to split the image preprocessing into 10 separate notebooks for 128x128 - ~3GB per notebooks (5GB is notebook write limit). As for modeling if you are using a dataloader or generator, you shouldn't have memory issues if you choose the right batchsize.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1010343,
          "author_name": "Akshay Chavan",
          "author_url": "",
          "post_date": "2020-09-14T17:22:51.747000",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/teeyee314\" target=\"_blank\">@teeyee314</a> for your response.<br>\nI will definitely try your approach as this is a good point to start working on the large datasets.<br>\nBut for long-running notebooks after every 40-minute kernel stops automatically if we aren't active on it.<br>\nSo keeping eye on the notebook to check its execution is difficult if we have multiple notebooks running simultaneously. So how you handle these scenarios?  </p>",
          "votes": 0,
          "replies": [
            {
              "id": 1010351,
              "author_name": "quadcore/Richard Epstein",
              "author_url": "",
              "post_date": "2020-09-14T17:27:09.153000",
              "content": "<p>You can commit your notebook and it will run unattended for up to 9 hours (3 hours with a TPU). The output of a committed notebook can be put into a Dataset without downloading.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 1010357,
          "author_name": "Tim Yee",
          "author_url": "",
          "post_date": "2020-09-14T17:29:15.893000",
          "content": "<p>you can save and commit the notebook so it will continue to run. just make sure your script doesn't have bugs. you can run a small test to find bugs before committing the notebook. Also make sure your notebook will run within the time limits (9hr for CPU and GPU, 3hr for TPU)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1010757,
          "author_name": "Akshay Chavan",
          "author_url": "",
          "post_date": "2020-09-15T03:48:21.630000",
          "content": "<p>Okay <a href=\"https://www.kaggle.com/teeyee314\" target=\"_blank\">@teeyee314</a> <br>\nThese tips will definitely help me with this problem. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1009118,
      "author_name": "Redwan Sony",
      "author_url": "",
      "post_date": "2020-09-13T16:58:22.963000",
      "content": "<p>Can we discard some of the data and train partially by somehow analyzing the metadata in the <code>train.csv</code> file? I know more data means better generalization, however, in order to reduce training time and have a quick jumpstart, maybe it can be good starting point.. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1009129,
          "author_name": "Akshay Chavan",
          "author_url": "",
          "post_date": "2020-09-13T17:08:29.480000",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/redwankarimsony\" target=\"_blank\">@redwankarimsony</a>.<br>\nFor the base models and to start with the problem we can take a sample of data. But if we have do get the better and generalize model then at that time we have to increase the size of the data. So in that anyhow we have to handle this huge data. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1009144,
          "author_name": "Redwan Sony",
          "author_url": "",
          "post_date": "2020-09-13T17:16:09.107000",
          "content": "<p>Thanks, <a href=\"https://www.kaggle.com/akshaychavan123\" target=\"_blank\">@akshaychavan123</a> <br>\nIf we make all the preprocessing (resizing, normalization, sorting of slices etc) and make that another dataset, I hope it will save everyone a lot of trouble. However, what's the point of being it as a competition? 😃😃</p>\n<p>I hope to see someone will rise up and make the preprocessing and create a separate dataset accordingly of smaller size.. :D </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1009155,
          "author_name": "Akshay Chavan",
          "author_url": "",
          "post_date": "2020-09-13T17:21:28.747000",
          "content": "<p>Yes, <a href=\"https://www.kaggle.com/redwankarimsony\" target=\"_blank\">@redwankarimsony</a> let's hope in the future for those things 😃😃😃.</p>\n<p>So can we say that for such competition if we do not have enough resources then we cannot move forward? I mean not all can afford the paid resources. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1011651,
          "author_name": "Tim Yee",
          "author_url": "",
          "post_date": "2020-09-15T15:57:29.990000",
          "content": "<p>In past experience - RSNA Intracranial, GCP credits were given out and that was just enough to stay competitive for me. Here, I would like to think that Kaggle TPU notebooks can stand a chance. Ultimately I don't think Kaggle GPU's will be able to get very far.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1007176,
      "author_name": "Gandom",
      "author_url": "",
      "post_date": "2020-09-11T21:58:09.380000",
      "content": "<p>How about discard-by-train, means loading only minibatch to GPU memory from disk?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1007353,
          "author_name": "Akshay Chavan",
          "author_url": "",
          "post_date": "2020-09-12T05:23:49.143000",
          "content": "<p>By doing this it takes a lot of time to train if the size of the batch is small. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1004967,
      "author_name": "Serigne ",
      "author_url": "",
      "post_date": "2020-09-10T07:05:27.503000",
      "content": "<p>I think TPU is the way to go. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 1004971,
          "author_name": "Akshay Chavan",
          "author_url": "",
          "post_date": "2020-09-10T07:07:45.230000",
          "content": "<p><a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a> Thank you for your response. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1004965,
      "author_name": "Arnaud Berenbaum",
      "author_url": "",
      "post_date": "2020-09-10T07:04:11.517000",
      "content": "<p>I would recommand you to train with small batches (8-16) therefor you will not run out of memory.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1004972,
          "author_name": "Akshay Chavan",
          "author_url": "",
          "post_date": "2020-09-10T07:08:07.843000",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/docdocb\" target=\"_blank\">@docdocb</a> </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1005001,
          "author_name": "Parth Chhabra",
          "author_url": "",
          "post_date": "2020-09-10T07:41:55.027000",
          "content": "<p>Nice idea by Arnaud, but just keep in mind that it would take much more time for training with small batch size than larger batch size. So, adjust your batch size accordingly.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1005018,
          "author_name": "Akshay Chavan",
          "author_url": "",
          "post_date": "2020-09-10T07:59:37.357000",
          "content": "<p>Yes, <a href=\"https://www.kaggle.com/oneplustricks\" target=\"_blank\">@oneplustricks</a>. That is why I was asking is there any alternatives for this. Because kernels also stop every 40 minutes if we are not active on it. And with such a small batch size, it takes a lot of time to train the model. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1004952": "Hello,\nI am a newbie and want to work on such a huge amount of data for the very first time. I want to ask you how you used to manage the resources for this much data. I mean Kaggle kernel will show out of memory error while training the model and the same thing goes with the google colab.\nSo I want to ask is there any free cloud platform available for such huge data with GPU system.\nAny kind of suggestions and thoughts would be appreciated for this topic. \n\nThank you",
    "1010331": "Personally, I am preprocessing the dcm files into jpg files on kaggle notebooks. I had to split the image preprocessing into 10 separate notebooks for 128x128 - ~3GB per notebooks (5GB is notebook write limit). As for modeling if you are using a dataloader or generator, you shouldn't have memory issues if you choose the right batchsize.",
    "1009118": "Can we discard some of the data and train partially by somehow analyzing the metadata in the `train.csv` file? I know more data means better generalization, however, in order to reduce training time and have a quick jumpstart, maybe it can be good starting point.. \n",
    "1007176": "How about discard-by-train, means loading only minibatch to GPU memory from disk?",
    "1004967": "I think TPU is the way to go. ",
    "1004965": "I would recommand you to train with small batches (8-16) therefor you will not run out of memory."
  }
}