{
  "id": 110671,
  "title": "Keras is slow",
  "url": "/competitions/rsna-intracranial-hemorrhage-detection/discussion/110671",
  "author_name": "Alex",
  "post_date": "2019-09-30T08:05:14.569000",
  "votes": 9,
  "comment_count": 20,
  "views": 0,
  "content": "<p>To quote @taindow from his kernel: \"Using mixed precision along with efficientnet-b0 and a little bit of pre-processing, a single pass of the entire 670k image dataset should take approx. 45m (at 224x224 resolution).\"<br></p>\n\n<p>I'm using Keras (with tensorflow backend) and it take's about 2-2.5 hours per epoch, using efficientnetb0 and 224x224 input (full train set). It's also ran on a Kaggle kernel. I would like to know if anyone experience the same? And do you have any idea of how to solve this for keras/tensorflow? Or is just that keras/tensorflow hasn't optimized for EfficientNets yet? I actually also find most of the keras_applications models relatively slow (InceptionResNetV2 takes ~6 hours per epoch etc.).<br>\n<em>Briefly what I've tried:</em><br>\n* 'channels_first' and 'channels_last' (So far I only made it work for EfficietNets)\n* Different implementations (all in keras/tensorflow)\n* Adding this to my script (from <a href=\"https://www.ibm.com/support/knowledgecenter/en/SS5SF7_1.6.1/navigation/wmlce_getstarted_tflmsv2.html#wmlce_getstarted_tflmsv2__usage_tips\">IBM's website</a>):\n<code>config = tf.ConfigProto()\nconfig.graph_options.rewrite_options.layout_optimizer=rewriter_config_pb2.RewriterConfig.OFF\nsess = tf.Session(config=config)\nkeras.backend.tensorflow_backend.set_session(sess)  # set new TensorFlow session as the default</code>\n* Locally (RTX 2070) or Kaggle kernel</p>\n\n<p>Doing 'channels_first' or adding that code to the script improved the training time by ~25% for EfficientNets, which is nice, but It's still relatively slow. It would be nice if you guys could share your training run-times! Thank you in advance.</p>",
  "messages": [
    {
      "id": 636801,
      "postDate": "2019-09-30T08:05:14.570Z",
      "content": "<p>To quote @taindow from his kernel: \"Using mixed precision along with efficientnet-b0 and a little bit of pre-processing, a single pass of the entire 670k image dataset should take approx. 45m (at 224x224 resolution).\"<br></p>\n\n<p>I'm using Keras (with tensorflow backend) and it take's about 2-2.5 hours per epoch, using efficientnetb0 and 224x224 input (full train set). It's also ran on a Kaggle kernel. I would like to know if anyone experience the same? And do you have any idea of how to solve this for keras/tensorflow? Or is just that keras/tensorflow hasn't optimized for EfficientNets yet? I actually also find most of the keras_applications models relatively slow (InceptionResNetV2 takes ~6 hours per epoch etc.).<br>\n<em>Briefly what I've tried:</em><br>\n* 'channels_first' and 'channels_last' (So far I only made it work for EfficietNets)\n* Different implementations (all in keras/tensorflow)\n* Adding this to my script (from <a href=\"https://www.ibm.com/support/knowledgecenter/en/SS5SF7_1.6.1/navigation/wmlce_getstarted_tflmsv2.html#wmlce_getstarted_tflmsv2__usage_tips\">IBM's website</a>):\n<code>config = tf.ConfigProto()\nconfig.graph_options.rewrite_options.layout_optimizer=rewriter_config_pb2.RewriterConfig.OFF\nsess = tf.Session(config=config)\nkeras.backend.tensorflow_backend.set_session(sess)  # set new TensorFlow session as the default</code>\n* Locally (RTX 2070) or Kaggle kernel</p>\n\n<p>Doing 'channels_first' or adding that code to the script improved the training time by ~25% for EfficientNets, which is nice, but It's still relatively slow. It would be nice if you guys could share your training run-times! Thank you in advance.</p>",
      "rawMarkdown": "To quote @taindow from his kernel: \"Using mixed precision along with efficientnet-b0 and a little bit of pre-processing, a single pass of the entire 670k image dataset should take approx. 45m (at 224x224 resolution).\"<br>\n\nI'm using Keras (with tensorflow backend) and it take's about 2-2.5 hours per epoch, using efficientnetb0 and 224x224 input (full train set). It's also ran on a Kaggle kernel. I would like to know if anyone experience the same? And do you have any idea of how to solve this for keras/tensorflow? Or is just that keras/tensorflow hasn't optimized for EfficientNets yet? I actually also find most of the keras_applications models relatively slow (InceptionResNetV2 takes ~6 hours per epoch etc.).<br>\n*Briefly what I've tried:*<br>\n* 'channels\\_first' and 'channels\\_last' (So far I only made it work for EfficietNets)\n* Different implementations (all in keras/tensorflow)\n* Adding this to my script (from [IBM's website](https://www.ibm.com/support/knowledgecenter/en/SS5SF7_1.6.1/navigation/wmlce_getstarted_tflmsv2.html#wmlce_getstarted_tflmsv2__usage_tips)):\n`config = tf.ConfigProto()\nconfig.graph_options.rewrite_options.layout_optimizer=rewriter_config_pb2.RewriterConfig.OFF\nsess = tf.Session(config=config)\nkeras.backend.tensorflow_backend.set_session(sess)  # set new TensorFlow session as the default`\n* Locally (RTX 2070) or Kaggle kernel\n\nDoing 'channels_first' or adding that code to the script improved the training time by ~25% for EfficientNets, which is nice, but It's still relatively slow. It would be nice if you guys could share your training run-times! Thank you in advance.\n\n\n",
      "votes": 9
    },
    {
      "id": 636875,
      "postDate": "2019-09-30T09:46:28.853Z",
      "content": "<p>pytorch can give up to 50% training speed boost over keras\nand you can also you apex, which further 2x your pytorch training speed</p>",
      "rawMarkdown": "pytorch can give up to 50% training speed boost over keras\nand you can also you apex, which further 2x your pytorch training speed",
      "votes": 3,
      "replies": [
        {
          "id": 636884,
          "postDate": "2019-09-30T10:02:25.847Z",
          "content": "<p>That is great, Thank you for the info! What a huge difference.. I will have to learn pytorch :-) </p>",
          "rawMarkdown": "That is great, Thank you for the info! What a huge difference.. I will have to learn pytorch :-) "
        },
        {
          "id": 636946,
          "postDate": "2019-09-30T12:11:13.213Z",
          "content": "<p><a href=\"/moewie94\">@moewie94</a>  sorry I never use apex before. Will using apex affect the training accuracy compared to normal pytorch only? thanks so much</p>",
          "rawMarkdown": "@moewie94  sorry I never use apex before. Will using apex affect the training accuracy compared to normal pytorch only? thanks so much"
        },
        {
          "id": 637051,
          "postDate": "2019-09-30T15:32:55.193Z",
          "content": "<p>I just implemented a Pytorch + apex version of my public kernel (not public/commited). I quickly tested ResNet50, ResNext50, InceptionV3 and EfficientNets. Training time of ResNet50 is pretty similar to before, but the other 3 are much faster now. <br></p>\n\n<p>It's still not 45min/epoch for EfficientNetB0, which probably means that my DataGenerator/DataLoader is the bottleneck (I'm using the dcms and not resized png files)</p>",
          "rawMarkdown": "I just implemented a Pytorch + apex version of my public kernel (not public/commited). I quickly tested ResNet50, ResNext50, InceptionV3 and EfficientNets. Training time of ResNet50 is pretty similar to before, but the other 3 are much faster now. <br>\n\nIt's still not 45min/epoch for EfficientNetB0, which probably means that my DataGenerator/DataLoader is the bottleneck (I'm using the dcms and not resized png files)",
          "votes": 1
        },
        {
          "id": 637053,
          "postDate": "2019-09-30T15:43:10.343Z",
          "content": "<p>You can use pytorch profiler to find out exactly what is bottleneck... </p>\n\n<p><a href=\"https://pytorch.org/docs/stable/autograd.html?highlight=autograd%20profiler#torch.autograd.profiler.profile\">https://pytorch.org/docs/stable/autograd.html?highlight=autograd%20profiler#torch.autograd.profiler.profile</a></p>\n\n<p>or you can use this nice package\n<a href=\"https://github.com/rkern/line_profiler\">https://github.com/rkern/line_profiler</a></p>",
          "rawMarkdown": "You can use pytorch profiler to find out exactly what is bottleneck... \n\nhttps://pytorch.org/docs/stable/autograd.html?highlight=autograd%20profiler#torch.autograd.profiler.profile\n\nor you can use this nice package\nhttps://github.com/rkern/line_profiler",
          "votes": 2
        },
        {
          "id": 637588,
          "postDate": "2019-10-01T06:26:53.057Z",
          "content": "<p><a href=\"/fiyeroleung\">@fiyeroleung</a> yeah it will affect the performance a little bit but the tradeoff is worth it</p>",
          "rawMarkdown": "@fiyeroleung yeah it will affect the performance a little bit but the tradeoff is worth it"
        },
        {
          "id": 638673,
          "postDate": "2019-10-02T08:59:05.620Z",
          "content": "<p>thanks so much. Sadly I only own a 1080ti which wont be benefited from FP16 calculation.🤕 </p>",
          "rawMarkdown": "thanks so much. Sadly I only own a 1080ti which wont be benefited from FP16 calculation.🤕 "
        }
      ]
    },
    {
      "id": 636826,
      "postDate": "2019-09-30T08:51:00.050Z",
      "content": "<p>size=224x224 batch_size=16 efficientnet b4 7hrs/epoch. Only uses ~35% gpu. No cpu/io bottleneck. I’ll try your suggestion.</p>",
      "rawMarkdown": "size=224x224 batch_size=16 efficientnet b4 7hrs/epoch. Only uses ~35% gpu. No cpu/io bottleneck. I’ll try your suggestion.",
      "votes": 1,
      "replies": [
        {
          "id": 636835,
          "postDate": "2019-09-30T08:59:00.190Z",
          "content": "<p>for that size you can use batch_size=64 and it will work fast(note : don't use full data)</p>",
          "rawMarkdown": "for that size you can use batch_size=64 and it will work fast(note : don't use full data)",
          "votes": 1
        },
        {
          "id": 636849,
          "postDate": "2019-09-30T09:08:57.687Z",
          "content": "<p>With full dataset that’s the best batch size I can afford. I’ll try use 10k positive and 10k negative subset later. Thanks for your suggestion!</p>",
          "rawMarkdown": "With full dataset that’s the best batch size I can afford. I’ll try use 10k positive and 10k negative subset later. Thanks for your suggestion!",
          "votes": 1
        },
        {
          "id": 636990,
          "postDate": "2019-09-30T13:32:47.887Z",
          "content": "<p>You dont have to go for <code>b4</code> or complex model. The best strategy is use smaller model and do bunch of experiments (e.g tfms, training schedule and etc).  After you are satisfied you can scale up.. My current LB standing is model trained on <code>b0</code> with 224 image size... </p>",
          "rawMarkdown": "You dont have to go for `b4` or complex model. The best strategy is use smaller model and do bunch of experiments (e.g tfms, training schedule and etc).  After you are satisfied you can scale up.. My current LB standing is model trained on `b0` with 224 image size... ",
          "votes": 8
        },
        {
          "id": 637022,
          "postDate": "2019-09-30T14:44:54.327Z",
          "content": "<p>This is really solid advice. Start small and experiment. More parameters and bigger images can come later.. </p>",
          "rawMarkdown": "This is really solid advice. Start small and experiment. More parameters and bigger images can come later.. ",
          "votes": 2
        },
        {
          "id": 637032,
          "postDate": "2019-09-30T15:05:35.173Z",
          "content": "<p>Thanks for the pro tips. 👍  It’s the first time I try EfficientNet so maybe I got a bit too excited that I forget the importance of fast iteration.</p>",
          "rawMarkdown": "Thanks for the pro tips. 👍  It’s the first time I try EfficientNet so maybe I got a bit too excited that I forget the importance of fast iteration.",
          "votes": 1
        },
        {
          "id": 637033,
          "postDate": "2019-09-30T15:10:29.517Z",
          "content": "<p>Its all good =) I have to learn this stuff hard way as well =) </p>",
          "rawMarkdown": "Its all good =) I have to learn this stuff hard way as well =) ",
          "votes": 1
        }
      ]
    },
    {
      "id": 636807,
      "postDate": "2019-09-30T08:16:15.753Z",
      "content": "<p>Absolutely <a href=\"/akensert\">@akensert</a>  ,no doubt about it,i have wasted so many gpu hours during working on<a href=\"https://www.kaggle.com/mobassir/keras-efficientnetb4-for-intracranial-hemorrhage\"> this kernel</a> from my quota dealing with kernel errors,it doesn't finish for even 3 epochs for 256x256 using only 340000 images within due time(9 hours in kaggle),i think keras team need to boost their model performance,keras is having lot of c/c++ down there,i don't know why it is still not as fast as expected :(</p>\n\n<p>How can i know in advance that my model will finish training within 9 hours? can you tell me? i don't want to waste gpu hours from my kaggle quota anymore,thanks in advance,specially thanks a lot for this post</p>",
      "rawMarkdown": "Absolutely @akensert  ,no doubt about it,i have wasted so many gpu hours during working on[ this kernel](https://www.kaggle.com/mobassir/keras-efficientnetb4-for-intracranial-hemorrhage) from my quota dealing with kernel errors,it doesn't finish for even 3 epochs for 256x256 using only 340000 images within due time(9 hours in kaggle),i think keras team need to boost their model performance,keras is having lot of c/c++ down there,i don't know why it is still not as fast as expected :(\n\nHow can i know in advance that my model will finish training within 9 hours? can you tell me? i don't want to waste gpu hours from my kaggle quota anymore,thanks in advance,specially thanks a lot for this post",
      "votes": 1,
      "replies": [
        {
          "id": 636815,
          "postDate": "2019-09-30T08:30:12.210Z",
          "content": "<p>Yeah, I also expect it to perform relatively good. I've been sitting with Keras for quite some time now (and I really like it), but right now I'm considering going to pytorch. Or perhaps try a different backend. <br></p>\n\n<p>To answer your question: I would run the model training in interactive mode just for a brief moment (with <code>verbose=1</code>), to approximate the total run time. And be a bit pessimistic about the run-time :-)</p>",
          "rawMarkdown": "Yeah, I also expect it to perform relatively good. I've been sitting with Keras for quite some time now (and I really like it), but right now I'm considering going to pytorch. Or perhaps try a different backend. <br>\n\nTo answer your question: I would run the model training in interactive mode just for a brief moment (with `verbose=1`), to approximate the total run time. And be a bit pessimistic about the run-time :-)",
          "votes": 1
        },
        {
          "id": 636843,
          "postDate": "2019-09-30T09:03:58.573Z",
          "content": "<p><a href=\"/akensert\">@akensert</a>  thank you,one more question if you don't mind?\nhow slow or fast gtx 1080  8gb and gtx 1080 16gb in our local machine compared to kaggle kernels gpu?\ni want to make pc for deep learning that will work as good as kaggle kernels or just slightly better than that.\nso, is gtx 1080 8gb with 16gb ddr4 and i7 is enough?\nor even gtx 1080 16gb with 16gb ddr4 and i7 is not better than kaggle kernels?\ni want to make pc for deep learning within range 1000-1200 dollar,if you could recommend me any laptop or desktop within this range then it will be highly appreciated,,,kaggle gpu quota hours limit sucks :(</p>",
          "rawMarkdown": "@akensert  thank you,one more question if you don't mind?\nhow slow or fast gtx 1080  8gb and gtx 1080 16gb in our local machine compared to kaggle kernels gpu?\ni want to make pc for deep learning that will work as good as kaggle kernels or just slightly better than that.\nso, is gtx 1080 8gb with 16gb ddr4 and i7 is enough?\nor even gtx 1080 16gb with 16gb ddr4 and i7 is not better than kaggle kernels?\ni want to make pc for deep learning within range 1000-1200 dollar,if you could recommend me any laptop or desktop within this range then it will be highly appreciated,,,kaggle gpu quota hours limit sucks :("
        },
        {
          "id": 636858,
          "postDate": "2019-09-30T09:23:09.003Z",
          "content": "<p><a href=\"/mobassir\">@mobassir</a> I would do a google search on this for sure! If you ask for my opinion, I would think that if you went for a gtx 1080 you would get similar or slightly faster model training than the Kaggle kernels'. For big datasets, which can't be read directly to memory, a fast CPU with multiple cores (perhaps hexacore) could be good. The GPU memory, 8 GB (which I have on my RTX 2070) I sometimes find a bit limited (but it's OK).</p>",
          "rawMarkdown": "@mobassir I would do a google search on this for sure! If you ask for my opinion, I would think that if you went for a gtx 1080 you would get similar or slightly faster model training than the Kaggle kernels'. For big datasets, which can't be read directly to memory, a fast CPU with multiple cores (perhaps hexacore) could be good. The GPU memory, 8 GB (which I have on my RTX 2070) I sometimes find a bit limited (but it's OK).",
          "votes": 1
        }
      ]
    },
    {
      "id": 637876,
      "postDate": "2019-10-01T10:35:12.100Z",
      "content": "<p>I tried to use <code>tf.train.experimental.enable_mixed_precision_graph_rewrite</code> with slightly larger batch size. Now EfficientNetB0 takes &lt;2hr per epoch instead of ~3.5hr. It should have been even faster because I read DCM files and resize the image on the fly and there's CPU bottleneck with 2 GPUs running simultaneously.</p>",
      "rawMarkdown": "I tried to use ```tf.train.experimental.enable_mixed_precision_graph_rewrite``` with slightly larger batch size. Now EfficientNetB0 takes &lt;2hr per epoch instead of ~3.5hr. It should have been even faster because I read DCM files and resize the image on the fly and there's CPU bottleneck with 2 GPUs running simultaneously.",
      "votes": 2
    },
    {
      "id": 638850,
      "postDate": "2019-10-02T13:39:13.227Z",
      "content": "<p>The interesting topic! I'm struggling now with slowness. So I subscribe here. :)</p>",
      "rawMarkdown": "The interesting topic! I'm struggling now with slowness. So I subscribe here. :)"
    }
  ],
  "comments": [
    {
      "id": 636875,
      "author_name": "DatNT",
      "author_url": "",
      "post_date": "2019-09-30T09:46:28.853000",
      "content": "<p>pytorch can give up to 50% training speed boost over keras\nand you can also you apex, which further 2x your pytorch training speed</p>",
      "votes": 3,
      "replies": [
        {
          "id": 636884,
          "author_name": "Alex",
          "author_url": "",
          "post_date": "2019-09-30T10:02:25.847000",
          "content": "<p>That is great, Thank you for the info! What a huge difference.. I will have to learn pytorch :-) </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 636946,
          "author_name": "Salaryman",
          "author_url": "",
          "post_date": "2019-09-30T12:11:13.213000",
          "content": "<p><a href=\"/moewie94\">@moewie94</a>  sorry I never use apex before. Will using apex affect the training accuracy compared to normal pytorch only? thanks so much</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 637051,
          "author_name": "Alex",
          "author_url": "",
          "post_date": "2019-09-30T15:32:55.193000",
          "content": "<p>I just implemented a Pytorch + apex version of my public kernel (not public/commited). I quickly tested ResNet50, ResNext50, InceptionV3 and EfficientNets. Training time of ResNet50 is pretty similar to before, but the other 3 are much faster now. <br></p>\n\n<p>It's still not 45min/epoch for EfficientNetB0, which probably means that my DataGenerator/DataLoader is the bottleneck (I'm using the dcms and not resized png files)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 637053,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-09-30T15:43:10.343000",
          "content": "<p>You can use pytorch profiler to find out exactly what is bottleneck... </p>\n\n<p><a href=\"https://pytorch.org/docs/stable/autograd.html?highlight=autograd%20profiler#torch.autograd.profiler.profile\">https://pytorch.org/docs/stable/autograd.html?highlight=autograd%20profiler#torch.autograd.profiler.profile</a></p>\n\n<p>or you can use this nice package\n<a href=\"https://github.com/rkern/line_profiler\">https://github.com/rkern/line_profiler</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 637588,
          "author_name": "DatNT",
          "author_url": "",
          "post_date": "2019-10-01T06:26:53.057000",
          "content": "<p><a href=\"/fiyeroleung\">@fiyeroleung</a> yeah it will affect the performance a little bit but the tradeoff is worth it</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 638673,
          "author_name": "Salaryman",
          "author_url": "",
          "post_date": "2019-10-02T08:59:05.620000",
          "content": "<p>thanks so much. Sadly I only own a 1080ti which wont be benefited from FP16 calculation.🤕 </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 636826,
      "author_name": "Yifeng (Ethan) Zou",
      "author_url": "",
      "post_date": "2019-09-30T08:51:00.050000",
      "content": "<p>size=224x224 batch_size=16 efficientnet b4 7hrs/epoch. Only uses ~35% gpu. No cpu/io bottleneck. I’ll try your suggestion.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 636835,
          "author_name": "Mobassir",
          "author_url": "",
          "post_date": "2019-09-30T08:59:00.190000",
          "content": "<p>for that size you can use batch_size=64 and it will work fast(note : don't use full data)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 636849,
          "author_name": "Yifeng (Ethan) Zou",
          "author_url": "",
          "post_date": "2019-09-30T09:08:57.687000",
          "content": "<p>With full dataset that’s the best batch size I can afford. I’ll try use 10k positive and 10k negative subset later. Thanks for your suggestion!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 636990,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-09-30T13:32:47.887000",
          "content": "<p>You dont have to go for <code>b4</code> or complex model. The best strategy is use smaller model and do bunch of experiments (e.g tfms, training schedule and etc).  After you are satisfied you can scale up.. My current LB standing is model trained on <code>b0</code> with 224 image size... </p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 637022,
          "author_name": "Tom Aindow",
          "author_url": "",
          "post_date": "2019-09-30T14:44:54.327000",
          "content": "<p>This is really solid advice. Start small and experiment. More parameters and bigger images can come later.. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 637032,
          "author_name": "Yifeng (Ethan) Zou",
          "author_url": "",
          "post_date": "2019-09-30T15:05:35.173000",
          "content": "<p>Thanks for the pro tips. 👍  It’s the first time I try EfficientNet so maybe I got a bit too excited that I forget the importance of fast iteration.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 637033,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-09-30T15:10:29.517000",
          "content": "<p>Its all good =) I have to learn this stuff hard way as well =) </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 636807,
      "author_name": "Mobassir",
      "author_url": "",
      "post_date": "2019-09-30T08:16:15.753000",
      "content": "<p>Absolutely <a href=\"/akensert\">@akensert</a>  ,no doubt about it,i have wasted so many gpu hours during working on<a href=\"https://www.kaggle.com/mobassir/keras-efficientnetb4-for-intracranial-hemorrhage\"> this kernel</a> from my quota dealing with kernel errors,it doesn't finish for even 3 epochs for 256x256 using only 340000 images within due time(9 hours in kaggle),i think keras team need to boost their model performance,keras is having lot of c/c++ down there,i don't know why it is still not as fast as expected :(</p>\n\n<p>How can i know in advance that my model will finish training within 9 hours? can you tell me? i don't want to waste gpu hours from my kaggle quota anymore,thanks in advance,specially thanks a lot for this post</p>",
      "votes": 1,
      "replies": [
        {
          "id": 636815,
          "author_name": "Alex",
          "author_url": "",
          "post_date": "2019-09-30T08:30:12.210000",
          "content": "<p>Yeah, I also expect it to perform relatively good. I've been sitting with Keras for quite some time now (and I really like it), but right now I'm considering going to pytorch. Or perhaps try a different backend. <br></p>\n\n<p>To answer your question: I would run the model training in interactive mode just for a brief moment (with <code>verbose=1</code>), to approximate the total run time. And be a bit pessimistic about the run-time :-)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 636843,
          "author_name": "Mobassir",
          "author_url": "",
          "post_date": "2019-09-30T09:03:58.573000",
          "content": "<p><a href=\"/akensert\">@akensert</a>  thank you,one more question if you don't mind?\nhow slow or fast gtx 1080  8gb and gtx 1080 16gb in our local machine compared to kaggle kernels gpu?\ni want to make pc for deep learning that will work as good as kaggle kernels or just slightly better than that.\nso, is gtx 1080 8gb with 16gb ddr4 and i7 is enough?\nor even gtx 1080 16gb with 16gb ddr4 and i7 is not better than kaggle kernels?\ni want to make pc for deep learning within range 1000-1200 dollar,if you could recommend me any laptop or desktop within this range then it will be highly appreciated,,,kaggle gpu quota hours limit sucks :(</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 636858,
          "author_name": "Alex",
          "author_url": "",
          "post_date": "2019-09-30T09:23:09.003000",
          "content": "<p><a href=\"/mobassir\">@mobassir</a> I would do a google search on this for sure! If you ask for my opinion, I would think that if you went for a gtx 1080 you would get similar or slightly faster model training than the Kaggle kernels'. For big datasets, which can't be read directly to memory, a fast CPU with multiple cores (perhaps hexacore) could be good. The GPU memory, 8 GB (which I have on my RTX 2070) I sometimes find a bit limited (but it's OK).</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 637876,
      "author_name": "Yifeng (Ethan) Zou",
      "author_url": "",
      "post_date": "2019-10-01T10:35:12.100000",
      "content": "<p>I tried to use <code>tf.train.experimental.enable_mixed_precision_graph_rewrite</code> with slightly larger batch size. Now EfficientNetB0 takes &lt;2hr per epoch instead of ~3.5hr. It should have been even faster because I read DCM files and resize the image on the fly and there's CPU bottleneck with 2 GPUs running simultaneously.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 638850,
      "author_name": "Sergey Zlobin",
      "author_url": "",
      "post_date": "2019-10-02T13:39:13.227000",
      "content": "<p>The interesting topic! I'm struggling now with slowness. So I subscribe here. :)</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "636801": "To quote @taindow from his kernel: \"Using mixed precision along with efficientnet-b0 and a little bit of pre-processing, a single pass of the entire 670k image dataset should take approx. 45m (at 224x224 resolution).\"<br>\n\nI'm using Keras (with tensorflow backend) and it take's about 2-2.5 hours per epoch, using efficientnetb0 and 224x224 input (full train set). It's also ran on a Kaggle kernel. I would like to know if anyone experience the same? And do you have any idea of how to solve this for keras/tensorflow? Or is just that keras/tensorflow hasn't optimized for EfficientNets yet? I actually also find most of the keras_applications models relatively slow (InceptionResNetV2 takes ~6 hours per epoch etc.).<br>\n*Briefly what I've tried:*<br>\n* 'channels\\_first' and 'channels\\_last' (So far I only made it work for EfficietNets)\n* Different implementations (all in keras/tensorflow)\n* Adding this to my script (from [IBM's website](https://www.ibm.com/support/knowledgecenter/en/SS5SF7_1.6.1/navigation/wmlce_getstarted_tflmsv2.html#wmlce_getstarted_tflmsv2__usage_tips)):\n`config = tf.ConfigProto()\nconfig.graph_options.rewrite_options.layout_optimizer=rewriter_config_pb2.RewriterConfig.OFF\nsess = tf.Session(config=config)\nkeras.backend.tensorflow_backend.set_session(sess)  # set new TensorFlow session as the default`\n* Locally (RTX 2070) or Kaggle kernel\n\nDoing 'channels_first' or adding that code to the script improved the training time by ~25% for EfficientNets, which is nice, but It's still relatively slow. It would be nice if you guys could share your training run-times! Thank you in advance.\n\n\n",
    "636875": "pytorch can give up to 50% training speed boost over keras\nand you can also you apex, which further 2x your pytorch training speed",
    "636826": "size=224x224 batch_size=16 efficientnet b4 7hrs/epoch. Only uses ~35% gpu. No cpu/io bottleneck. I’ll try your suggestion.",
    "636807": "Absolutely @akensert  ,no doubt about it,i have wasted so many gpu hours during working on[ this kernel](https://www.kaggle.com/mobassir/keras-efficientnetb4-for-intracranial-hemorrhage) from my quota dealing with kernel errors,it doesn't finish for even 3 epochs for 256x256 using only 340000 images within due time(9 hours in kaggle),i think keras team need to boost their model performance,keras is having lot of c/c++ down there,i don't know why it is still not as fast as expected :(\n\nHow can i know in advance that my model will finish training within 9 hours? can you tell me? i don't want to waste gpu hours from my kaggle quota anymore,thanks in advance,specially thanks a lot for this post",
    "637876": "I tried to use ```tf.train.experimental.enable_mixed_precision_graph_rewrite``` with slightly larger batch size. Now EfficientNetB0 takes &lt;2hr per epoch instead of ~3.5hr. It should have been even faster because I read DCM files and resize the image on the fly and there's CPU bottleneck with 2 GPUs running simultaneously.",
    "638850": "The interesting topic! I'm struggling now with slowness. So I subscribe here. :)"
  }
}