{
  "id": 182801,
  "title": "Large DataSet - can get started in Notebooks",
  "url": "/competitions/rsna-str-pulmonary-embolism-detection/discussion/182801",
  "author_name": "quadcore/Richard Epstein",
  "post_date": "2020-09-14T11:34:12.376000",
  "votes": 16,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Kagglers,</p>\n<p>Based on the leaderboard number of submissions and number of competitors, I suspect many are having trouble with the large dataset. I know it took me 2 days to download the data (one failed attempt half-way through). And a lot of time to clear off enough disk space and unzip the file (same size as zipped, by the way). For those with a poor internet connection or lack of 1.6 TB disk space, it will be impossible.</p>\n<p>While I was waiting, I started working in the Notebooks and I think you can make some progress without local resources.</p>\n<p>I'm a big TPU/TFRecord fan. Even though you cannot use the TPU for the test data submission, you can use them for training. And TFRecords can be used with GPUs.</p>\n<p>Break the problem into steps:</p>\n<ol>\n<li><p>Create a metadata_train.csv file from the training data.</p></li>\n<li><p>Create TFRecords and store in a Dataset. This can be broken into 100,000 record chunks (about 9 hours processing). Can copy your notebook and run multiple times simultaneously. Don't bother trying to do all the images; your needs will change before you are done and you will be recreating new TFRecords in a different format. Crop or downsize the images. Pulmonary embolisms are not in the air outside the patient or the edges of the patient. You probably won't be able to detect a really tiny pulmonary embolism on the edge of the lungs anyway. Right now I'm working at the image level only, so I only build records for patients with pulmonary embolism. Their images have plenty of slices without pulmonary embolism on them. [EDIT: I'm concerned I'll train my model too strongly on slices in the middle of the chest matching with Pulmonary Embolism. So I'm adding negative patients also.]</p></li>\n<li><p>While creating all those TFRecords, do some training on a small subset of the data. 100,000 records is larger than most datasets we work with, and you can test out some models.</p></li>\n<li><p>Save the trained model to a Dataset.</p></li>\n<li><p>Build your test submission pipeline. Here, you need to put all the steps in one notebook (utility scripts might work, haven't tried) and rerun it each time start to finish. I check the size of the test.csv file so my program knows if it is running on the sample test.csv file or the entire (committed) file. If it's just on the sample, I restrict myself to a few hundred images, so I can make sure my code works, but I don't waste time. I use the same TFRecord structure/processing as I do for Training. Keeps things standardized. Might be faster to skip the TFRecords in testing since I cannot save them for use another time anyway. Eventually, if I run into processing time limits, I'll rewrite to not use TFRecords for prediction.</p>\n<p>Create metadata_test.csv<br>\nCreate TFRecords<br>\nRun against trained model in external dataset<br>\nPost Process into correct Submission.csv format.</p></li>\n</ol>\n<p>Here, you might run into processing limits. So far, I have only run a subset of the test data. Your score won't improve as much as it would with the full test data, but it should move in the right direction, letting you know if you are making progress.</p>\n<ol>\n<li><p>So far, I've built the pipeline, but I'm just starting to scratch the surface on the actual image analysis, and I haven't tackled the 3D nature of the data yet.</p></li>\n<li><p>Since we are so limited in running Test data and getting a leaderboard score, cross-validation is very important. You cannot easily use the leaderboard as your cross-validation (which never works out well anyway).</p></li>\n<li><p>A lot of good ideas in the recent Melanoma contest (doesn't have the 3D aspect, but has DICOM, large datasets and TPU/TFRecords). Also current Pulmonary Fibrosis contest deals with DICOM Data for chest CTs.</p></li>\n</ol>\n<p>My very transient Leaderboard score is 99% from some smart defaulting, not some amazing model. It's taken a while to build the infrastructure, but hopefully now I can get into real modeling.</p>\n<p>Good luck to all!</p>\n<p>-Rich</p>",
  "messages": [
    {
      "id": 1009938,
      "postDate": "2020-09-14T11:34:12.377Z",
      "content": "<p>Kagglers,</p>\n<p>Based on the leaderboard number of submissions and number of competitors, I suspect many are having trouble with the large dataset. I know it took me 2 days to download the data (one failed attempt half-way through). And a lot of time to clear off enough disk space and unzip the file (same size as zipped, by the way). For those with a poor internet connection or lack of 1.6 TB disk space, it will be impossible.</p>\n<p>While I was waiting, I started working in the Notebooks and I think you can make some progress without local resources.</p>\n<p>I'm a big TPU/TFRecord fan. Even though you cannot use the TPU for the test data submission, you can use them for training. And TFRecords can be used with GPUs.</p>\n<p>Break the problem into steps:</p>\n<ol>\n<li><p>Create a metadata_train.csv file from the training data.</p></li>\n<li><p>Create TFRecords and store in a Dataset. This can be broken into 100,000 record chunks (about 9 hours processing). Can copy your notebook and run multiple times simultaneously. Don't bother trying to do all the images; your needs will change before you are done and you will be recreating new TFRecords in a different format. Crop or downsize the images. Pulmonary embolisms are not in the air outside the patient or the edges of the patient. You probably won't be able to detect a really tiny pulmonary embolism on the edge of the lungs anyway. Right now I'm working at the image level only, so I only build records for patients with pulmonary embolism. Their images have plenty of slices without pulmonary embolism on them. [EDIT: I'm concerned I'll train my model too strongly on slices in the middle of the chest matching with Pulmonary Embolism. So I'm adding negative patients also.]</p></li>\n<li><p>While creating all those TFRecords, do some training on a small subset of the data. 100,000 records is larger than most datasets we work with, and you can test out some models.</p></li>\n<li><p>Save the trained model to a Dataset.</p></li>\n<li><p>Build your test submission pipeline. Here, you need to put all the steps in one notebook (utility scripts might work, haven't tried) and rerun it each time start to finish. I check the size of the test.csv file so my program knows if it is running on the sample test.csv file or the entire (committed) file. If it's just on the sample, I restrict myself to a few hundred images, so I can make sure my code works, but I don't waste time. I use the same TFRecord structure/processing as I do for Training. Keeps things standardized. Might be faster to skip the TFRecords in testing since I cannot save them for use another time anyway. Eventually, if I run into processing time limits, I'll rewrite to not use TFRecords for prediction.</p>\n<p>Create metadata_test.csv<br>\nCreate TFRecords<br>\nRun against trained model in external dataset<br>\nPost Process into correct Submission.csv format.</p></li>\n</ol>\n<p>Here, you might run into processing limits. So far, I have only run a subset of the test data. Your score won't improve as much as it would with the full test data, but it should move in the right direction, letting you know if you are making progress.</p>\n<ol>\n<li><p>So far, I've built the pipeline, but I'm just starting to scratch the surface on the actual image analysis, and I haven't tackled the 3D nature of the data yet.</p></li>\n<li><p>Since we are so limited in running Test data and getting a leaderboard score, cross-validation is very important. You cannot easily use the leaderboard as your cross-validation (which never works out well anyway).</p></li>\n<li><p>A lot of good ideas in the recent Melanoma contest (doesn't have the 3D aspect, but has DICOM, large datasets and TPU/TFRecords). Also current Pulmonary Fibrosis contest deals with DICOM Data for chest CTs.</p></li>\n</ol>\n<p>My very transient Leaderboard score is 99% from some smart defaulting, not some amazing model. It's taken a while to build the infrastructure, but hopefully now I can get into real modeling.</p>\n<p>Good luck to all!</p>\n<p>-Rich</p>",
      "rawMarkdown": "Kagglers,\n\nBased on the leaderboard number of submissions and number of competitors, I suspect many are having trouble with the large dataset. I know it took me 2 days to download the data (one failed attempt half-way through). And a lot of time to clear off enough disk space and unzip the file (same size as zipped, by the way). For those with a poor internet connection or lack of 1.6 TB disk space, it will be impossible.\n\nWhile I was waiting, I started working in the Notebooks and I think you can make some progress without local resources.\n\nI'm a big TPU/TFRecord fan. Even though you cannot use the TPU for the test data submission, you can use them for training. And TFRecords can be used with GPUs.\n\nBreak the problem into steps:\n\n1. Create a metadata_train.csv file from the training data.\n\n2. Create TFRecords and store in a Dataset. This can be broken into 100,000 record chunks (about 9 hours processing). Can copy your notebook and run multiple times simultaneously. Don't bother trying to do all the images; your needs will change before you are done and you will be recreating new TFRecords in a different format. Crop or downsize the images. Pulmonary embolisms are not in the air outside the patient or the edges of the patient. You probably won't be able to detect a really tiny pulmonary embolism on the edge of the lungs anyway. Right now I'm working at the image level only, so I only build records for patients with pulmonary embolism. Their images have plenty of slices without pulmonary embolism on them. [EDIT: I'm concerned I'll train my model too strongly on slices in the middle of the chest matching with Pulmonary Embolism. So I'm adding negative patients also.]\n\n3. While creating all those TFRecords, do some training on a small subset of the data. 100,000 records is larger than most datasets we work with, and you can test out some models.\n\n4. Save the trained model to a Dataset.\n\n6. Build your test submission pipeline. Here, you need to put all the steps in one notebook (utility scripts might work, haven't tried) and rerun it each time start to finish. I check the size of the test.csv file so my program knows if it is running on the sample test.csv file or the entire (committed) file. If it's just on the sample, I restrict myself to a few hundred images, so I can make sure my code works, but I don't waste time. I use the same TFRecord structure/processing as I do for Training. Keeps things standardized. Might be faster to skip the TFRecords in testing since I cannot save them for use another time anyway. Eventually, if I run into processing time limits, I'll rewrite to not use TFRecords for prediction.\n\n    Create metadata_test.csv\n    Create TFRecords\n    Run against trained model in external dataset\n    Post Process into correct Submission.csv format.\n\nHere, you might run into processing limits. So far, I have only run a subset of the test data. Your score won't improve as much as it would with the full test data, but it should move in the right direction, letting you know if you are making progress.\n\n7. So far, I've built the pipeline, but I'm just starting to scratch the surface on the actual image analysis, and I haven't tackled the 3D nature of the data yet.\n\n8. Since we are so limited in running Test data and getting a leaderboard score, cross-validation is very important. You cannot easily use the leaderboard as your cross-validation (which never works out well anyway).\n\n9. A lot of good ideas in the recent Melanoma contest (doesn't have the 3D aspect, but has DICOM, large datasets and TPU/TFRecords). Also current Pulmonary Fibrosis contest deals with DICOM Data for chest CTs.\n\nMy very transient Leaderboard score is 99% from some smart defaulting, not some amazing model. It's taken a while to build the infrastructure, but hopefully now I can get into real modeling.\n\nGood luck to all!\n\n-Rich",
      "votes": 16
    },
    {
      "id": 1010311,
      "postDate": "2020-09-14T16:55:32.867Z",
      "content": "<blockquote>\n  <p>It's taken a while to build the infrastructure, but hopefully now I can get into real modeling.</p>\n</blockquote>\n<p>I spent the last 3-4 days just getting an image pre-processing script going on kaggle notebooks. I had to break the training images up into 10 notebooks for 128x128 (currently totaling ~30GB) and will be increasingly more notebooks for larger resolutions. So far. I think the biggest challenge (and fun) will be getting a TPU training and GPU inference pipeline going.</p>",
      "rawMarkdown": "> It's taken a while to build the infrastructure, but hopefully now I can get into real modeling.\n\nI spent the last 3-4 days just getting an image pre-processing script going on kaggle notebooks. I had to break the training images up into 10 notebooks for 128x128 (currently totaling ~30GB) and will be increasingly more notebooks for larger resolutions. So far. I think the biggest challenge (and fun) will be getting a TPU training and GPU inference pipeline going."
    },
    {
      "id": 1010160,
      "postDate": "2020-09-14T14:46:29.953Z",
      "content": "<p>Thanks a lot <a href=\"/richardepstein\">@richardepstein</a> for sharing your these steps with us. </p>",
      "rawMarkdown": "Thanks a lot @richardepstein for sharing your these steps with us. "
    },
    {
      "id": 1011783,
      "postDate": "2020-09-15T17:17:04.173Z",
      "content": "<p>Following you. Thanks.</p>",
      "rawMarkdown": "Following you. Thanks."
    }
  ],
  "comments": [
    {
      "id": 1010311,
      "author_name": "Tim Yee",
      "author_url": "",
      "post_date": "2020-09-14T16:55:32.867000",
      "content": "<blockquote>\n  <p>It's taken a while to build the infrastructure, but hopefully now I can get into real modeling.</p>\n</blockquote>\n<p>I spent the last 3-4 days just getting an image pre-processing script going on kaggle notebooks. I had to break the training images up into 10 notebooks for 128x128 (currently totaling ~30GB) and will be increasingly more notebooks for larger resolutions. So far. I think the biggest challenge (and fun) will be getting a TPU training and GPU inference pipeline going.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1010160,
      "author_name": "Redwan Sony",
      "author_url": "",
      "post_date": "2020-09-14T14:46:29.953000",
      "content": "<p>Thanks a lot <a href=\"/richardepstein\">@richardepstein</a> for sharing your these steps with us. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1011783,
      "author_name": "Mobasshir Bhuiya Shagor",
      "author_url": "",
      "post_date": "2020-09-15T17:17:04.173000",
      "content": "<p>Following you. Thanks.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1009938": "Kagglers,\n\nBased on the leaderboard number of submissions and number of competitors, I suspect many are having trouble with the large dataset. I know it took me 2 days to download the data (one failed attempt half-way through). And a lot of time to clear off enough disk space and unzip the file (same size as zipped, by the way). For those with a poor internet connection or lack of 1.6 TB disk space, it will be impossible.\n\nWhile I was waiting, I started working in the Notebooks and I think you can make some progress without local resources.\n\nI'm a big TPU/TFRecord fan. Even though you cannot use the TPU for the test data submission, you can use them for training. And TFRecords can be used with GPUs.\n\nBreak the problem into steps:\n\n1. Create a metadata_train.csv file from the training data.\n\n2. Create TFRecords and store in a Dataset. This can be broken into 100,000 record chunks (about 9 hours processing). Can copy your notebook and run multiple times simultaneously. Don't bother trying to do all the images; your needs will change before you are done and you will be recreating new TFRecords in a different format. Crop or downsize the images. Pulmonary embolisms are not in the air outside the patient or the edges of the patient. You probably won't be able to detect a really tiny pulmonary embolism on the edge of the lungs anyway. Right now I'm working at the image level only, so I only build records for patients with pulmonary embolism. Their images have plenty of slices without pulmonary embolism on them. [EDIT: I'm concerned I'll train my model too strongly on slices in the middle of the chest matching with Pulmonary Embolism. So I'm adding negative patients also.]\n\n3. While creating all those TFRecords, do some training on a small subset of the data. 100,000 records is larger than most datasets we work with, and you can test out some models.\n\n4. Save the trained model to a Dataset.\n\n6. Build your test submission pipeline. Here, you need to put all the steps in one notebook (utility scripts might work, haven't tried) and rerun it each time start to finish. I check the size of the test.csv file so my program knows if it is running on the sample test.csv file or the entire (committed) file. If it's just on the sample, I restrict myself to a few hundred images, so I can make sure my code works, but I don't waste time. I use the same TFRecord structure/processing as I do for Training. Keeps things standardized. Might be faster to skip the TFRecords in testing since I cannot save them for use another time anyway. Eventually, if I run into processing time limits, I'll rewrite to not use TFRecords for prediction.\n\n    Create metadata_test.csv\n    Create TFRecords\n    Run against trained model in external dataset\n    Post Process into correct Submission.csv format.\n\nHere, you might run into processing limits. So far, I have only run a subset of the test data. Your score won't improve as much as it would with the full test data, but it should move in the right direction, letting you know if you are making progress.\n\n7. So far, I've built the pipeline, but I'm just starting to scratch the surface on the actual image analysis, and I haven't tackled the 3D nature of the data yet.\n\n8. Since we are so limited in running Test data and getting a leaderboard score, cross-validation is very important. You cannot easily use the leaderboard as your cross-validation (which never works out well anyway).\n\n9. A lot of good ideas in the recent Melanoma contest (doesn't have the 3D aspect, but has DICOM, large datasets and TPU/TFRecords). Also current Pulmonary Fibrosis contest deals with DICOM Data for chest CTs.\n\nMy very transient Leaderboard score is 99% from some smart defaulting, not some amazing model. It's taken a while to build the infrastructure, but hopefully now I can get into real modeling.\n\nGood luck to all!\n\n-Rich",
    "1010311": "> It's taken a while to build the infrastructure, but hopefully now I can get into real modeling.\n\nI spent the last 3-4 days just getting an image pre-processing script going on kaggle notebooks. I had to break the training images up into 10 notebooks for 128x128 (currently totaling ~30GB) and will be increasingly more notebooks for larger resolutions. So far. I think the biggest challenge (and fun) will be getting a TPU training and GPU inference pipeline going.",
    "1010160": "Thanks a lot @richardepstein for sharing your these steps with us. ",
    "1011783": "Following you. Thanks."
  }
}