{
  "id": 183409,
  "title": "How long does it take to train for one epoch and inference?",
  "url": "/competitions/rsna-str-pulmonary-embolism-detection/discussion/183409",
  "author_name": "Salaryman",
  "post_date": "2020-09-16T14:56:03.021000",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi can anyone share how long does it take to train for one epoch and inference(and img size)?<br>\nThere are several comps ongoing and Im thinking should I join this one.</p>\n<p>Thanks so much </p>",
  "messages": [
    {
      "id": 1013308,
      "postDate": "2020-09-16T16:20:21.740Z",
      "content": "<p>I trained a DenseNet121 on 256x256 images. I used only 700,00 images (out of 1,790,594 training images).</p>\n<p>TFRec and TPU -  6 minutes per epoch</p>\n<p>So, full training set looks like 15 minutes per epoch</p>\n<p>Had to create TFRecs beforehand. That takes like a day, running multiple processes.</p>\n<p>Also had preprocessed metadata from DICOM images.</p>\n<p>Inference is tricky. I'm currently creating metadata and TFRecs on the fly, since we cannot preprocess the test set. Only using about 100,000 out of 342,732 test images for my testing. Still takes 9 hours of GPU time. Might be doing something wrong. Using TFRecs to maintain consistency with training process, but might be faster to skip them since I cannot reuse them in inference.</p>\n<p>-Rich</p>",
      "rawMarkdown": "I trained a DenseNet121 on 256x256 images. I used only 700,00 images (out of 1,790,594 training images).\n\nTFRec and TPU -  6 minutes per epoch\n\nSo, full training set looks like 15 minutes per epoch\n\nHad to create TFRecs beforehand. That takes like a day, running multiple processes.\n\nAlso had preprocessed metadata from DICOM images.\n\nInference is tricky. I'm currently creating metadata and TFRecs on the fly, since we cannot preprocess the test set. Only using about 100,000 out of 342,732 test images for my testing. Still takes 9 hours of GPU time. Might be doing something wrong. Using TFRecs to maintain consistency with training process, but might be faster to skip them since I cannot reuse them in inference.\n\n-Rich",
      "votes": 3,
      "replies": [
        {
          "id": 1013339,
          "postDate": "2020-09-16T16:44:11.380Z",
          "content": "<p>Thanks for your insight Richard. Love to hear more in the coming days if you figured out inference for TFRecords. Thanks for the share!</p>",
          "rawMarkdown": "Thanks for your insight Richard. Love to hear more in the coming days if you figured out inference for TFRecords. Thanks for the share!"
        },
        {
          "id": 1013376,
          "postDate": "2020-09-16T17:05:49.913Z",
          "content": "<p><a href=\"https://www.kaggle.com/richardepstein\" target=\"_blank\">@richardepstein</a> I had an idea of inferencing the given test set ahead of time with TFRecords and then making sure the script doesn't re-inference on those images during the re-run. So perhaps you can just import the predictions of the public test saved as a csv merge that onto the re-run kernel somehow and then for every prediction made for public test, just ignore it for the re-run</p>",
          "rawMarkdown": "@richardepstein I had an idea of inferencing the given test set ahead of time with TFRecords and then making sure the script doesn't re-inference on those images during the re-run. So perhaps you can just import the predictions of the public test saved as a csv merge that onto the re-run kernel somehow and then for every prediction made for public test, just ignore it for the re-run"
        },
        {
          "id": 1013397,
          "postDate": "2020-09-16T17:18:57.377Z",
          "content": "<p>Thanks for your info Richard!<br>\nWill work on it from Saturday!🔥</p>",
          "rawMarkdown": "Thanks for your info Richard!\nWill work on it from Saturday!🔥"
        }
      ]
    },
    {
      "id": 1013191,
      "postDate": "2020-09-16T14:56:03.023Z",
      "content": "<p>Hi can anyone share how long does it take to train for one epoch and inference(and img size)?<br>\nThere are several comps ongoing and Im thinking should I join this one.</p>\n<p>Thanks so much </p>",
      "rawMarkdown": "Hi can anyone share how long does it take to train for one epoch and inference(and img size)?\nThere are several comps ongoing and Im thinking should I join this one.\n\nThanks so much ",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 1013308,
      "author_name": "quadcore/Richard Epstein",
      "author_url": "",
      "post_date": "2020-09-16T16:20:21.740000",
      "content": "<p>I trained a DenseNet121 on 256x256 images. I used only 700,00 images (out of 1,790,594 training images).</p>\n<p>TFRec and TPU -  6 minutes per epoch</p>\n<p>So, full training set looks like 15 minutes per epoch</p>\n<p>Had to create TFRecs beforehand. That takes like a day, running multiple processes.</p>\n<p>Also had preprocessed metadata from DICOM images.</p>\n<p>Inference is tricky. I'm currently creating metadata and TFRecs on the fly, since we cannot preprocess the test set. Only using about 100,000 out of 342,732 test images for my testing. Still takes 9 hours of GPU time. Might be doing something wrong. Using TFRecs to maintain consistency with training process, but might be faster to skip them since I cannot reuse them in inference.</p>\n<p>-Rich</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1013339,
          "author_name": "Bo Peng",
          "author_url": "",
          "post_date": "2020-09-16T16:44:11.380000",
          "content": "<p>Thanks for your insight Richard. Love to hear more in the coming days if you figured out inference for TFRecords. Thanks for the share!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1013376,
          "author_name": "Tim Yee",
          "author_url": "",
          "post_date": "2020-09-16T17:05:49.913000",
          "content": "<p><a href=\"https://www.kaggle.com/richardepstein\" target=\"_blank\">@richardepstein</a> I had an idea of inferencing the given test set ahead of time with TFRecords and then making sure the script doesn't re-inference on those images during the re-run. So perhaps you can just import the predictions of the public test saved as a csv merge that onto the re-run kernel somehow and then for every prediction made for public test, just ignore it for the re-run</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1013397,
          "author_name": "Salaryman",
          "author_url": "",
          "post_date": "2020-09-16T17:18:57.377000",
          "content": "<p>Thanks for your info Richard!<br>\nWill work on it from Saturday!🔥</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1013308": "I trained a DenseNet121 on 256x256 images. I used only 700,00 images (out of 1,790,594 training images).\n\nTFRec and TPU -  6 minutes per epoch\n\nSo, full training set looks like 15 minutes per epoch\n\nHad to create TFRecs beforehand. That takes like a day, running multiple processes.\n\nAlso had preprocessed metadata from DICOM images.\n\nInference is tricky. I'm currently creating metadata and TFRecs on the fly, since we cannot preprocess the test set. Only using about 100,000 out of 342,732 test images for my testing. Still takes 9 hours of GPU time. Might be doing something wrong. Using TFRecs to maintain consistency with training process, but might be faster to skip them since I cannot reuse them in inference.\n\n-Rich",
    "1013191": "Hi can anyone share how long does it take to train for one epoch and inference(and img size)?\nThere are several comps ongoing and Im thinking should I join this one.\n\nThanks so much "
  }
}