{
  "id": 191542,
  "title": "Storing preprocessed data",
  "url": "/competitions/rsna-str-pulmonary-embolism-detection/discussion/191542",
  "author_name": "Jayasuriya Senthilvelan",
  "post_date": "2020-10-17T03:45:33.445000",
  "votes": 2,
  "comment_count": 16,
  "views": 0,
  "content": "<p>Hi,</p>\n<p>The original input data is nearly 1 TB. I found a way to preprocess the datasets, but I'm having trouble storing the data for subsequent use by my model. </p>\n<p>My preprocessed data is ~100 GB. I realize I have to store it in Kaggle/working/ so I can access it in another kernel. Unfortunately, this dir. is limited to 5 G in space. If anyone could let me know how they stored their preprocessed data, I would greatly appreciate it.</p>\n<p>Thanks!</p>",
  "messages": [
    {
      "id": 1056844,
      "postDate": "2020-10-22T06:17:20.707Z",
      "content": "<p>A follow up question -- Can we upgrade to a Google Cloud Platform notebook for our final submission so we can have a little more compute? cc: <a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a> <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a></p>",
      "rawMarkdown": "A follow up question -- Can we upgrade to a Google Cloud Platform notebook for our final submission so we can have a little more compute? cc: @juliaelliott @philculliton",
      "votes": 1,
      "replies": [
        {
          "id": 1058380,
          "postDate": "2020-10-23T16:25:45.343Z",
          "content": "<p><a href=\"https://www.kaggle.com/aksg87\" target=\"_blank\">@aksg87</a> Good question! You're certainly welcome to upgrade via our <a href=\"https://www.kaggle.com/product-feedback/159602\" target=\"_blank\">Cloud AI Platform Notebooks integration</a> for more self-paid compute. But when bringing your notebook back into Kaggle, you'll still be expected to meet the competition submission constraints. So it may be a great option for enhancing your training power, and subsequently bringing that trained model into Kaggle for submission.</p>",
          "rawMarkdown": "@aksg87 Good question! You're certainly welcome to upgrade via our [Cloud AI Platform Notebooks integration](https://www.kaggle.com/product-feedback/159602) for more self-paid compute. But when bringing your notebook back into Kaggle, you'll still be expected to meet the competition submission constraints. So it may be a great option for enhancing your training power, and subsequently bringing that trained model into Kaggle for submission.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1056654,
      "postDate": "2020-10-22T00:32:59.187Z",
      "content": "<p>When running inference, can we preprocess the data into another format and store it to disk while running the notebook interactively? </p>\n<p>I.e. let's say I was processing the dicom data into a resized numpy arrays and wanted to use 50-100 GIG prior to inference. </p>\n<p>The other option is to build this into the inference pipeline and not save studies which is fine as well. I'm just curious what the computation constraints are.  cc: <a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a> <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> </p>\n<p>I don't see storage discussed in: <a href=\"https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/overview/code-requirements\" target=\"_blank\">https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/overview/code-requirements</a></p>",
      "rawMarkdown": "When running inference, can we preprocess the data into another format and store it to disk while running the notebook interactively? \n\nI.e. let's say I was processing the dicom data into a resized numpy arrays and wanted to use 50-100 GIG prior to inference. \n\nThe other option is to build this into the inference pipeline and not save studies which is fine as well. I'm just curious what the computation constraints are.  cc: @juliaelliott @philculliton \n\nI don't see storage discussed in: https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/overview/code-requirements",
      "votes": 1,
      "replies": [
        {
          "id": 1056659,
          "postDate": "2020-10-22T00:43:33.743Z",
          "content": "<p><a href=\"https://www.kaggle.com/aksg87\" target=\"_blank\">@aksg87</a> Memory limitations follow global notebook constraints. So if you click on the session stats bar, you will see what memory is available, depending on whether you're using an accelerator.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F603584%2F9d646951c314d9eab36d9429c5594a6a%2FScreen%20Shot%202020-10-21%20at%205.39.04%20PM.png?generation=1603327196668983&amp;alt=media\" alt=\"session stats bar\"></p>",
          "rawMarkdown": "@aksg87 Memory limitations follow global notebook constraints. So if you click on the session stats bar, you will see what memory is available, depending on whether you're using an accelerator.\n![session stats bar](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F603584%2F9d646951c314d9eab36d9429c5594a6a%2FScreen%20Shot%202020-10-21%20at%205.39.04%20PM.png?generation=1603327196668983&alt=media)\n\n",
          "votes": 2
        },
        {
          "id": 1056660,
          "postDate": "2020-10-22T00:43:59.720Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1056665,
          "postDate": "2020-10-22T00:51:49.627Z",
          "content": "<p>Thanks! I am not really concerned about memory but rather the HDD space. </p>\n<p>So to confirm, we cannot preprocess the data and produce anything that takes up over 5 GB (so our inference script should basically delete preprocessed files as we run it on all studies).</p>\n<p>I was under the impression that the 5gb was for saving files for AFTER the notebook session was terminated, but I thought there might be more space available during the execution phase.  </p>\n<p><a href=\"https://www.kaggle.com/richardepstein\" target=\"_blank\">@richardepstein</a> mentioned \"I think you can use more than 5 GB interactively\"</p>",
          "rawMarkdown": "Thanks! I am not really concerned about memory but rather the HDD space. \n\nSo to confirm, we cannot preprocess the data and produce anything that takes up over 5 GB (so our inference script should basically delete preprocessed files as we run it on all studies).\n\nI was under the impression that the 5gb was for saving files for AFTER the notebook session was terminated, but I thought there might be more space available during the execution phase.  \n\n@richardepstein mentioned \"I think you can use more than 5 GB interactively\""
        },
        {
          "id": 1056669,
          "postDate": "2020-10-22T00:57:44.520Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1056670,
          "postDate": "2020-10-22T01:01:01.757Z",
          "content": "<p><a href=\"https://www.kaggle.com/stanleyjzheng\" target=\"_blank\">@stanleyjzheng</a> thanks! I just wanted to make sure it was over 50-100 GB.</p>\n<p>Given the large dataset size, it's cleaner to pre-process it all in one step for inference.</p>",
          "rawMarkdown": "@stanleyjzheng thanks! I just wanted to make sure it was over 50-100 GB.\n\nGiven the large dataset size, it's cleaner to pre-process it all in one step for inference.\n"
        }
      ]
    },
    {
      "id": 1051909,
      "postDate": "2020-10-17T03:45:33.447Z",
      "content": "<p>Hi,</p>\n<p>The original input data is nearly 1 TB. I found a way to preprocess the datasets, but I'm having trouble storing the data for subsequent use by my model. </p>\n<p>My preprocessed data is ~100 GB. I realize I have to store it in Kaggle/working/ so I can access it in another kernel. Unfortunately, this dir. is limited to 5 G in space. If anyone could let me know how they stored their preprocessed data, I would greatly appreciate it.</p>\n<p>Thanks!</p>",
      "rawMarkdown": "Hi,\n\nThe original input data is nearly 1 TB. I found a way to preprocess the datasets, but I'm having trouble storing the data for subsequent use by my model. \n\nMy preprocessed data is ~100 GB. I realize I have to store it in Kaggle/working/ so I can access it in another kernel. Unfortunately, this dir. is limited to 5 G in space. If anyone could let me know how they stored their preprocessed data, I would greatly appreciate it.\n\nThanks!",
      "votes": 2
    },
    {
      "id": 1056970,
      "postDate": "2020-10-22T08:24:17.540Z",
      "content": "<p>I did a test and after 5 GB of HDD the instance automatically stops suggesting GCP. Can we use GCP for our submission notebook or does it need to run in the normal Kaggle environment. Thanks for any clarifications here!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1847649%2F32d4206d29c785d8124df162781f2e42%2Fchrome_mSEMc85Rz2.png?generation=1603354960668574&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I did a test and after 5 GB of HDD the instance automatically stops suggesting GCP. Can we use GCP for our submission notebook or does it need to run in the normal Kaggle environment. Thanks for any clarifications here!\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1847649%2F32d4206d29c785d8124df162781f2e42%2Fchrome_mSEMc85Rz2.png?generation=1603354960668574&alt=media)",
      "votes": -1
    },
    {
      "id": 1057109,
      "postDate": "2020-10-22T11:26:28.110Z",
      "content": "<p>I think I realized the answer with the temp folder. Thanks!</p>",
      "rawMarkdown": "I think I realized the answer with the temp folder. Thanks!"
    },
    {
      "id": 1052434,
      "postDate": "2020-10-17T17:58:34.347Z",
      "content": "<p>public datasets in Kaggle have higher limits (might not be limited). Just make it public, if you are comfortable doing that.</p>",
      "rawMarkdown": "public datasets in Kaggle have higher limits (might not be limited). Just make it public, if you are comfortable doing that.",
      "replies": [
        {
          "id": 1052438,
          "postDate": "2020-10-17T18:11:40.007Z",
          "content": "<p>Do I make the notebook public? Or how do I make the file public and then save it?</p>",
          "rawMarkdown": "Do I make the notebook public? Or how do I make the file public and then save it?",
          "replies": [
            {
              "id": 1052457,
              "postDate": "2020-10-17T18:37:20.587Z",
              "content": "<p>You don't need to make the notebook public. Just a Dataset.</p>\n<p>I think if you commit your notebook, then the output can be added to a Dataset directly.</p>\n<p>If you run interactively, you can download your data and then upload it to a dataset.</p>\n<p>You can add to an existing dataset, so if your preprocessed data is too large for one notebook, you can do it in stages.</p>",
              "rawMarkdown": "You don't need to make the notebook public. Just a Dataset.\n\nI think if you commit your notebook, then the output can be added to a Dataset directly.\n\nIf you run interactively, you can download your data and then upload it to a dataset.\n\nYou can add to an existing dataset, so if your preprocessed data is too large for one notebook, you can do it in stages."
            },
            {
              "id": 1052461,
              "postDate": "2020-10-17T18:46:05.483Z",
              "content": "<p>I know I can commit, but at max I can commit and save 5 G per notebook. Are you suggesting I create like 20 notebooks, and commit 5 G each?</p>\n<p>How do I add to to an existing dataset? I'm unfamiliar with this.</p>",
              "rawMarkdown": "I know I can commit, but at max I can commit and save 5 G per notebook. Are you suggesting I create like 20 notebooks, and commit 5 G each?\n\nHow do I add to to an existing dataset? I'm unfamiliar with this."
            },
            {
              "id": 1052468,
              "postDate": "2020-10-17T18:53:40.957Z",
              "content": "<p>I believe kaggle/working allows more than 5 GB, although it won't save it at the end of a commit. I think you can use more than 5 GB interactively. Then download it. And then upload to a dataset.</p>\n<p>After you create a dataset, there is an option \"New Version\". This lets you add to an existing dataset. </p>\n<p>I have public datasets up to 56 GB.</p>\n<p>It is slow moving these large datasets, but especially for TFRecords and TPU processing, it is worth it.</p>",
              "rawMarkdown": "I believe kaggle/working allows more than 5 GB, although it won't save it at the end of a commit. I think you can use more than 5 GB interactively. Then download it. And then upload to a dataset.\n\nAfter you create a dataset, there is an option \"New Version\". This lets you add to an existing dataset. \n\nI have public datasets up to 56 GB.\n\nIt is slow moving these large datasets, but especially for TFRecords and TPU processing, it is worth it."
            },
            {
              "id": 1056662,
              "postDate": "2020-10-22T00:45:35.993Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1056844,
      "author_name": "aksg87",
      "author_url": "",
      "post_date": "2020-10-22T06:17:20.707000",
      "content": "<p>A follow up question -- Can we upgrade to a Google Cloud Platform notebook for our final submission so we can have a little more compute? cc: <a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a> <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 1058380,
          "author_name": "Julia Elliott",
          "author_url": "",
          "post_date": "2020-10-23T16:25:45.343000",
          "content": "<p><a href=\"https://www.kaggle.com/aksg87\" target=\"_blank\">@aksg87</a> Good question! You're certainly welcome to upgrade via our <a href=\"https://www.kaggle.com/product-feedback/159602\" target=\"_blank\">Cloud AI Platform Notebooks integration</a> for more self-paid compute. But when bringing your notebook back into Kaggle, you'll still be expected to meet the competition submission constraints. So it may be a great option for enhancing your training power, and subsequently bringing that trained model into Kaggle for submission.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1056654,
      "author_name": "aksg87",
      "author_url": "",
      "post_date": "2020-10-22T00:32:59.187000",
      "content": "<p>When running inference, can we preprocess the data into another format and store it to disk while running the notebook interactively? </p>\n<p>I.e. let's say I was processing the dicom data into a resized numpy arrays and wanted to use 50-100 GIG prior to inference. </p>\n<p>The other option is to build this into the inference pipeline and not save studies which is fine as well. I'm just curious what the computation constraints are.  cc: <a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a> <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> </p>\n<p>I don't see storage discussed in: <a href=\"https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/overview/code-requirements\" target=\"_blank\">https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/overview/code-requirements</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 1056659,
          "author_name": "Julia Elliott",
          "author_url": "",
          "post_date": "2020-10-22T00:43:33.743000",
          "content": "<p><a href=\"https://www.kaggle.com/aksg87\" target=\"_blank\">@aksg87</a> Memory limitations follow global notebook constraints. So if you click on the session stats bar, you will see what memory is available, depending on whether you're using an accelerator.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F603584%2F9d646951c314d9eab36d9429c5594a6a%2FScreen%20Shot%202020-10-21%20at%205.39.04%20PM.png?generation=1603327196668983&amp;alt=media\" alt=\"session stats bar\"></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1056660,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-10-22T00:43:59.720000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1056665,
          "author_name": "aksg87",
          "author_url": "",
          "post_date": "2020-10-22T00:51:49.627000",
          "content": "<p>Thanks! I am not really concerned about memory but rather the HDD space. </p>\n<p>So to confirm, we cannot preprocess the data and produce anything that takes up over 5 GB (so our inference script should basically delete preprocessed files as we run it on all studies).</p>\n<p>I was under the impression that the 5gb was for saving files for AFTER the notebook session was terminated, but I thought there might be more space available during the execution phase.  </p>\n<p><a href=\"https://www.kaggle.com/richardepstein\" target=\"_blank\">@richardepstein</a> mentioned \"I think you can use more than 5 GB interactively\"</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1056669,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-10-22T00:57:44.520000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1056670,
          "author_name": "aksg87",
          "author_url": "",
          "post_date": "2020-10-22T01:01:01.757000",
          "content": "<p><a href=\"https://www.kaggle.com/stanleyjzheng\" target=\"_blank\">@stanleyjzheng</a> thanks! I just wanted to make sure it was over 50-100 GB.</p>\n<p>Given the large dataset size, it's cleaner to pre-process it all in one step for inference.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1056970,
      "author_name": "aksg87",
      "author_url": "",
      "post_date": "2020-10-22T08:24:17.540000",
      "content": "<p>I did a test and after 5 GB of HDD the instance automatically stops suggesting GCP. Can we use GCP for our submission notebook or does it need to run in the normal Kaggle environment. Thanks for any clarifications here!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1847649%2F32d4206d29c785d8124df162781f2e42%2Fchrome_mSEMc85Rz2.png?generation=1603354960668574&amp;alt=media\" alt=\"\"></p>",
      "votes": -1,
      "replies": []
    },
    {
      "id": 1057109,
      "author_name": "aksg87",
      "author_url": "",
      "post_date": "2020-10-22T11:26:28.110000",
      "content": "<p>I think I realized the answer with the temp folder. Thanks!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1052434,
      "author_name": "quadcore/Richard Epstein",
      "author_url": "",
      "post_date": "2020-10-17T17:58:34.347000",
      "content": "<p>public datasets in Kaggle have higher limits (might not be limited). Just make it public, if you are comfortable doing that.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1052438,
          "author_name": "Jayasuriya Senthilvelan",
          "author_url": "",
          "post_date": "2020-10-17T18:11:40.007000",
          "content": "<p>Do I make the notebook public? Or how do I make the file public and then save it?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 1052457,
              "author_name": "quadcore/Richard Epstein",
              "author_url": "",
              "post_date": "2020-10-17T18:37:20.587000",
              "content": "<p>You don't need to make the notebook public. Just a Dataset.</p>\n<p>I think if you commit your notebook, then the output can be added to a Dataset directly.</p>\n<p>If you run interactively, you can download your data and then upload it to a dataset.</p>\n<p>You can add to an existing dataset, so if your preprocessed data is too large for one notebook, you can do it in stages.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 1052461,
              "author_name": "Jayasuriya Senthilvelan",
              "author_url": "",
              "post_date": "2020-10-17T18:46:05.483000",
              "content": "<p>I know I can commit, but at max I can commit and save 5 G per notebook. Are you suggesting I create like 20 notebooks, and commit 5 G each?</p>\n<p>How do I add to to an existing dataset? I'm unfamiliar with this.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 1052468,
              "author_name": "quadcore/Richard Epstein",
              "author_url": "",
              "post_date": "2020-10-17T18:53:40.957000",
              "content": "<p>I believe kaggle/working allows more than 5 GB, although it won't save it at the end of a commit. I think you can use more than 5 GB interactively. Then download it. And then upload to a dataset.</p>\n<p>After you create a dataset, there is an option \"New Version\". This lets you add to an existing dataset. </p>\n<p>I have public datasets up to 56 GB.</p>\n<p>It is slow moving these large datasets, but especially for TFRecords and TPU processing, it is worth it.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 1056662,
              "author_name": "",
              "author_url": "",
              "post_date": "2020-10-22T00:45:35.993000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1056844": "A follow up question -- Can we upgrade to a Google Cloud Platform notebook for our final submission so we can have a little more compute? cc: @juliaelliott @philculliton",
    "1056654": "When running inference, can we preprocess the data into another format and store it to disk while running the notebook interactively? \n\nI.e. let's say I was processing the dicom data into a resized numpy arrays and wanted to use 50-100 GIG prior to inference. \n\nThe other option is to build this into the inference pipeline and not save studies which is fine as well. I'm just curious what the computation constraints are.  cc: @juliaelliott @philculliton \n\nI don't see storage discussed in: https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/overview/code-requirements",
    "1051909": "Hi,\n\nThe original input data is nearly 1 TB. I found a way to preprocess the datasets, but I'm having trouble storing the data for subsequent use by my model. \n\nMy preprocessed data is ~100 GB. I realize I have to store it in Kaggle/working/ so I can access it in another kernel. Unfortunately, this dir. is limited to 5 G in space. If anyone could let me know how they stored their preprocessed data, I would greatly appreciate it.\n\nThanks!",
    "1056970": "I did a test and after 5 GB of HDD the instance automatically stops suggesting GCP. Can we use GCP for our submission notebook or does it need to run in the normal Kaggle environment. Thanks for any clarifications here!\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1847649%2F32d4206d29c785d8124df162781f2e42%2Fchrome_mSEMc85Rz2.png?generation=1603354960668574&alt=media)",
    "1057109": "I think I realized the answer with the temp folder. Thanks!",
    "1052434": "public datasets in Kaggle have higher limits (might not be limited). Just make it public, if you are comfortable doing that."
  }
}