{
  "id": 375881,
  "title": "1.4x-1.6x inference speedup with torch_tensorrt ",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/375881",
  "author_name": "David Austin",
  "post_date": "2023-01-03T23:08:55.769000",
  "votes": 35,
  "comment_count": 23,
  "views": 0,
  "content": "<p>Speeding up model inference in this competition could end up being very important.  <a href=\"https://github.com/pytorch/TensorRT\" target=\"_blank\">torch_tensorrt </a> is an integration for PyTorch that leverages inference optimizations of TensorRT on NVIDIA GPUs.  I've created a notebook and dataset that will allow you install all the dependencies (internet disabled) and compile a torch_tensorrt model from your pytorch model, either in fp32 or fp16.</p>\n<p><a href=\"https://www.kaggle.com/code/tivfrvqhs5/torch-tensorrt-infer-fp16-and-fp32-benchmarks/notebook\" target=\"_blank\">Notebook link</a></p>\n<p>Speedup results for EFN-B2 using 1024x1024x3 image size (images/sec).</p>\n<table>\n  <tbody><tr>\n    <th></th>\n    <th>pytorch</th>\n    <th>torch_tensorrt</th>\n  </tr>\n  <tr>\n    <th>FP32</th>\n    <td>47.91</td>\n    <td>66.07</td>\n  </tr>\n  <tr>\n    <th>FP16</th>\n    <td>54.72</td>\n    <td>77.86</td>\n  </tr>\n</tbody></table>",
  "messages": [
    {
      "id": 2085061,
      "postDate": "2023-01-03T23:08:55.770Z",
      "content": "<p>Speeding up model inference in this competition could end up being very important.  <a href=\"https://github.com/pytorch/TensorRT\" target=\"_blank\">torch_tensorrt </a> is an integration for PyTorch that leverages inference optimizations of TensorRT on NVIDIA GPUs.  I've created a notebook and dataset that will allow you install all the dependencies (internet disabled) and compile a torch_tensorrt model from your pytorch model, either in fp32 or fp16.</p>\n<p><a href=\"https://www.kaggle.com/code/tivfrvqhs5/torch-tensorrt-infer-fp16-and-fp32-benchmarks/notebook\" target=\"_blank\">Notebook link</a></p>\n<p>Speedup results for EFN-B2 using 1024x1024x3 image size (images/sec).</p>\n<table>\n  <tbody><tr>\n    <th></th>\n    <th>pytorch</th>\n    <th>torch_tensorrt</th>\n  </tr>\n  <tr>\n    <th>FP32</th>\n    <td>47.91</td>\n    <td>66.07</td>\n  </tr>\n  <tr>\n    <th>FP16</th>\n    <td>54.72</td>\n    <td>77.86</td>\n  </tr>\n</tbody></table>",
      "rawMarkdown": "Speeding up model inference in this competition could end up being very important.  [torch_tensorrt ](https://github.com/pytorch/TensorRT) is an integration for PyTorch that leverages inference optimizations of TensorRT on NVIDIA GPUs.  I've created a notebook and dataset that will allow you install all the dependencies (internet disabled) and compile a torch_tensorrt model from your pytorch model, either in fp32 or fp16.\n\n[Notebook link](https://www.kaggle.com/code/tivfrvqhs5/torch-tensorrt-infer-fp16-and-fp32-benchmarks/notebook)\n\nSpeedup results for EFN-B2 using 1024x1024x3 image size (images/sec).\n<table>\n  <tr>\n    <th></th>\n    <th>pytorch</th>\n    <th>torch_tensorrt</th>\n  </tr>\n  <tr>\n    <th>FP32</th>\n    <td>47.91</td>\n    <td>66.07</td>\n  </tr>\n  <tr>\n    <th>FP16</th>\n    <td>54.72</td>\n    <td>77.86</td>\n  </tr>\n</table>\n\n",
      "votes": 35
    },
    {
      "id": 2100212,
      "postDate": "2023-01-15T01:59:58.823Z",
      "content": "<p>i think you should not use model.half() when compile 16fp in your tensortRT example:</p>\n<p>my experimental results</p>\n<pre><code>model effnet-b4 on 1536x960\n\n***************ok!\npytorch autocast:\n\n10935 / 10935   7 min 09 sec\n\nauc 0.8820045604670901\nf1score.max() 0.48506685839497704\n@threshold 0.3469387755102041\n\n***************ok!\ntensorrt\n\n10935 / 10935   4 min 01 sec\n\nauc 0.8824084583371677\nf1score.max() 0.48506685839497704\n@threshold 0.3469387755102041\n\n***************ok!\ntensorrt compiled with model.half()\n\n10935 / 10935   4 min 06 sec\n\nauc 0.8296021350348171\nf1score.max() 0.39217532383668313\n@threshold 0.04081632653061224\n\n***************ok!\n</code></pre>\n<pre><code>       f = torch.load(checkpoint, map_location=lambda storage, loc: storage)\n    model.load_state_dict(f['state_dict'], strict=False)\n    model.eval().cuda()\n    model.half()  ### should not be used !!!!!\n\n\n    print('torch_tensorrt.compile() ...')\n    with torch_tensorrt.logging.debug():\n        trt_model_fp16 = torch_tensorrt.compile(  \n            model,\n            inputs=[\n                torch_tensorrt.Input(\n                    [batch_size, 1, image_height, image_width],\n                    dtype=torch.half\n                )],\n            enabled_precisions={torch.half},  \n            workspace_size=1 &lt;&lt; 32,\n            require_full_compilation=True,\n            #debug=True,\n        )\n    torch.jit.save(trt_model_fp16, tft_file)\n</code></pre>",
      "rawMarkdown": "i think you should not use model.half() when compile 16fp in your tensortRT example:\n\nmy experimental results\n```\nmodel effnet-b4 on 1536x960\n\n***************ok!\npytorch autocast:\n\n10935 / 10935   7 min 09 sec\n \nauc 0.8820045604670901\nf1score.max() 0.48506685839497704\n@threshold 0.3469387755102041\n\n***************ok!\ntensorrt\n\n10935 / 10935   4 min 01 sec\n\nauc 0.8824084583371677\nf1score.max() 0.48506685839497704\n@threshold 0.3469387755102041\n\n***************ok!\ntensorrt compiled with model.half()\n\n10935 / 10935   4 min 06 sec\n\nauc 0.8296021350348171\nf1score.max() 0.39217532383668313\n@threshold 0.04081632653061224\n\n***************ok!\n\n```\n\n```\n       f = torch.load(checkpoint, map_location=lambda storage, loc: storage)\n\tmodel.load_state_dict(f['state_dict'], strict=False)\n\tmodel.eval().cuda()\n\tmodel.half()  ### should not be used !!!!!\n\n\n\tprint('torch_tensorrt.compile() ...')\n\twith torch_tensorrt.logging.debug():\n\t\ttrt_model_fp16 = torch_tensorrt.compile(  \n\t\t\tmodel,\n\t\t\tinputs=[\n\t\t\t\ttorch_tensorrt.Input(\n\t\t\t\t\t[batch_size, 1, image_height, image_width],\n\t\t\t\t\tdtype=torch.half\n\t\t\t\t)],\n\t\t\tenabled_precisions={torch.half},  \n\t\t\tworkspace_size=1 << 32,\n\t\t\trequire_full_compilation=True,\n\t\t\t#debug=True,\n\t\t)\n\ttorch.jit.save(trt_model_fp16, tft_file)\n\n```\n",
      "votes": 1
    },
    {
      "id": 2098825,
      "postDate": "2023-01-13T21:51:06.163Z",
      "content": "<p>my timing for nextVIT transformer with and without tensorRT on P100.<br>\nIt is about 4x speed improvement.</p>\n<p>speed is about the same as claim in paper (i.e. about efficientnet speed)</p>\n<pre><code>** tesorRT **\n - public LB submission tensorRT 4 hr : LB 0.56\n - local cv 10939 images: \ntensorRT fp16 nextVIT-B (1539x960) \n\nCPU utilisation 120%,  13/13 GB\nGPU utilisation 99% , 4.5/15 GB\n23 min 21 sec (7.80799 images per sec)\n\n** without tesorRT **\n- submission  9hr : LB 0.56\n- local cv 10939 images: \nmerged_bn fp16 (1539x960)  \n\nCPU utilisation 108%,  13/13 GB\nGPU utilisation 100% , 8.5/15 GB\n95 min 32 sec (1.90819 images per sec)\n</code></pre>\n<p>more details breakdown of timing at <br>\n<a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/370333#2098821\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/370333#2098821</a> </p>",
      "rawMarkdown": "my timing for nextVIT transformer with and without tensorRT on P100.\nIt is about 4x speed improvement.\n\nspeed is about the same as claim in paper (i.e. about efficientnet speed)\n\n```\n** tesorRT **\n - public LB submission tensorRT 4 hr : LB 0.56\n - local cv 10939 images: \ntensorRT fp16 nextVIT-B (1539x960) \n\nCPU utilisation 120%,  13/13 GB\nGPU utilisation 99% , 4.5/15 GB\n23 min 21 sec (7.80799 images per sec)\n\n** without tesorRT **\n- submission  9hr : LB 0.56\n- local cv 10939 images: \nmerged_bn fp16 (1539x960)  \n\nCPU utilisation 108%,  13/13 GB\nGPU utilisation 100% , 8.5/15 GB\n95 min 32 sec (1.90819 images per sec)\n\n```\n\nmore details breakdown of timing at \nhttps://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/370333#2098821 \n",
      "votes": 1,
      "replies": [
        {
          "id": 2099205,
          "postDate": "2023-01-14T08:50:05.130Z",
          "content": "<p>there seems to be some unknown bug in my implementation.<br>\nthe bug is in kaggle notebook or tensorrt lib or the nextvit code?<br>\n<a href=\"https://www.kaggle.com/code/hengck23/3hr-tensorrt-nextvit-example\" target=\"_blank\">https://www.kaggle.com/code/hengck23/3hr-tensorrt-nextvit-example</a></p>\n<p>when i call torch_tensorrt.compile(), it seems to stall and nothing happens after 10 min.<br>\nthen click the stop run button at the side, the process is terminated with usual verbal messages flushed out.<br>\nif i run again, this produces the tft engine file finally.</p>\n<p>i tried a few times from fresh start, and it seems this is always the case.</p>\n<p>(this doesn't happens on my local machine with the same code, but it uses pycharm ide instead of notebook)</p>",
          "rawMarkdown": "there seems to be some unknown bug in my implementation.\nthe bug is in kaggle notebook or tensorrt lib or the nextvit code?\nhttps://www.kaggle.com/code/hengck23/3hr-tensorrt-nextvit-example\n\nwhen i call torch_tensorrt.compile(), it seems to stall and nothing happens after 10 min.\nthen click the stop run button at the side, the process is terminated with usual verbal messages flushed out.\nif i run again, this produces the tft engine file finally.\n\ni tried a few times from fresh start, and it seems this is always the case.\n\n(this doesn't happens on my local machine with the same code, but it uses pycharm ide instead of notebook)",
          "votes": 1,
          "replies": [
            {
              "id": 2100125,
              "postDate": "2023-01-15T00:17:01.800Z",
              "content": "<p>I've seen this several times when running a jupyter notebook, but it doesn't happen when running from a .py file.  There are github issues reporting the same</p>",
              "rawMarkdown": "I've seen this several times when running a jupyter notebook, but it doesn't happen when running from a .py file.  There are github issues reporting the same",
              "votes": 2
            },
            {
              "id": 2110681,
              "postDate": "2023-01-22T10:50:05.207Z",
              "content": "<p>I've faced this issue too and it's fixed with Torch Tensor RT 1.3.0</p>",
              "rawMarkdown": "I've faced this issue too and it's fixed with Torch Tensor RT 1.3.0"
            },
            {
              "id": 2112897,
              "postDate": "2023-01-23T23:26:26.857Z",
              "content": "<p>thanks.</p>\n<p>i solved the issue by using the following in kaggle notebook</p>\n<pre><code>batch_size = 4\nfrom model_nextvit_multi import *\nif 1:\n    ! python /kaggle/input/.../convert_model_nextvit_multi.py --checkpoint_file ...\n</code></pre>",
              "rawMarkdown": "thanks.\n\ni solved the issue by using the following in kaggle notebook\n\n```\nbatch_size = 4\nfrom model_nextvit_multi import *\nif 1:\n    ! python /kaggle/input/.../convert_model_nextvit_multi.py --checkpoint_file ...\n\n```",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2087644,
      "postDate": "2023-01-05T18:24:36.023Z",
      "content": "<p>Update:  Running the same code on a 2x T4 notebook shows marginally faster times for torch_tensorrt but slower times for pytorch inference (single GPU mode).  The reason for this is likely that T4 GPU's have fp32 and fp16 tensor cores which torch_tensorrt utilizes but P100 GPU's don't have them.  Even faster speedup could be achieved in dual GPU inference mode assuming your dataloader isn't a bottleneck.</p>\n<p>Speedup results for EFN B-2 using 1024x1024x3 image size on T4 GPU (single GPU mode).</p>\n<table>\n  <tbody><tr>\n    <th></th>\n    <th>pytorch</th>\n    <th>torch_tensorrt</th>\n  </tr>\n  <tr>\n    <th>FP32</th>\n    <td>29.1</td>\n    <td>37.63</td>\n  </tr>\n  <tr>\n    <th>FP16</th>\n    <td>56.87</td>\n    <td>81.64</td>\n  </tr>\n</tbody></table>",
      "rawMarkdown": "Update:  Running the same code on a 2x T4 notebook shows marginally faster times for torch_tensorrt but slower times for pytorch inference (single GPU mode).  The reason for this is likely that T4 GPU's have fp32 and fp16 tensor cores which torch_tensorrt utilizes but P100 GPU's don't have them.  Even faster speedup could be achieved in dual GPU inference mode assuming your dataloader isn't a bottleneck.\n\nSpeedup results for EFN B-2 using 1024x1024x3 image size on T4 GPU (single GPU mode).\n\n<table>\n  <tr>\n    <th></th>\n    <th>pytorch</th>\n    <th>torch_tensorrt</th>\n  </tr>\n  <tr>\n    <th>FP32</th>\n    <td>29.1</td>\n    <td>37.63</td>\n  </tr>\n  <tr>\n    <th>FP16</th>\n    <td>56.87</td>\n    <td>81.64</td>\n  </tr>\n</table>",
      "votes": 1
    },
    {
      "id": 2098401,
      "postDate": "2023-01-13T14:10:01.740Z",
      "content": "<p>ironically, the speed gain by tensorrt model is reduce by the extra long time to install tensorrt</p>",
      "rawMarkdown": "ironically, the speed gain by tensorrt model is reduce by the extra long time to install tensorrt",
      "votes": 2,
      "replies": [
        {
          "id": 2098676,
          "postDate": "2023-01-13T18:19:21.303Z",
          "content": "<p>Yes it's a 3-4 minute install time, it would be nice if tensorrt was included in the base docker build since it helps reduce the compute load on the system.</p>",
          "rawMarkdown": "Yes it's a 3-4 minute install time, it would be nice if tensorrt was included in the base docker build since it helps reduce the compute load on the system.",
          "votes": 3,
          "replies": [
            {
              "id": 2098725,
              "postDate": "2023-01-13T19:02:10.893Z",
              "content": "<p>We're always accepting new package proposals here: <a href=\"https://github.com/Kaggle/docker-python#requesting-new-packages\" target=\"_blank\">https://github.com/Kaggle/docker-python#requesting-new-packages</a></p>",
              "rawMarkdown": "We're always accepting new package proposals here: https://github.com/Kaggle/docker-python#requesting-new-packages",
              "votes": 4
            },
            {
              "id": 2110709,
              "postDate": "2023-01-22T11:07:35.763Z",
              "content": "<p>Request done (open issue)</p>",
              "rawMarkdown": "Request done (open issue)",
              "votes": 5
            },
            {
              "id": 2112620,
              "postDate": "2023-01-23T18:07:19.550Z",
              "content": "<p>Great, thanks! I'll let the engineer who handles that repo know just how much use this library is getting.</p>",
              "rawMarkdown": "Great, thanks! I'll let the engineer who handles that repo know just how much use this library is getting.",
              "votes": 3
            },
            {
              "id": 2112678,
              "postDate": "2023-01-23T18:45:35.850Z",
              "content": "<p>I've requested for TensorRT 1.3 which is better than 1.2 already demonstrated here. I'm able to install and make it works on current Kaggle environement but I have to patch torch_tensorrt source code. I believe your guys will face similar issues. It will be great if it could be available. </p>",
              "rawMarkdown": "I've requested for TensorRT 1.3 which is better than 1.2 already demonstrated here. I'm able to install and make it works on current Kaggle environement but I have to patch torch_tensorrt source code. I believe your guys will face similar issues. It will be great if it could be available. ",
              "votes": 1
            },
            {
              "id": 2114162,
              "postDate": "2023-01-24T18:38:20.920Z",
              "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> it looks like there are some version conflicts with your proposal: <a href=\"https://github.com/Kaggle/docker-python/issues/1210\" target=\"_blank\">https://github.com/Kaggle/docker-python/issues/1210</a></p>",
              "rawMarkdown": "@mpware it looks like there are some version conflicts with your proposal: https://github.com/Kaggle/docker-python/issues/1210",
              "votes": 1
            },
            {
              "id": 2114262,
              "postDate": "2023-01-24T20:18:15.937Z",
              "content": "<p>Yes, I've seen the answer and I understand the conflict. We've to wait to have it in a next docker image. In the mean time, as suggested by the answer, we can install it ourself (works for me). </p>",
              "rawMarkdown": "Yes, I've seen the answer and I understand the conflict. We've to wait to have it in a next docker image. In the mean time, as suggested by the answer, we can install it ourself (works for me). ",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2088500,
      "postDate": "2023-01-06T11:33:10.510Z",
      "content": "<p><a href=\"https://github.com/bytedance/Next-ViT\" target=\"_blank\">https://github.com/bytedance/Next-ViT</a><br>\nif you are interested, you can check tensorRT friendly transformer.<br>\nfrom my experiments, Next-ViT small has same performance as efficientb2 to b4</p>\n<p><img src=\"https://raw.githubusercontent.com/bytedance/Next-ViT/main/images/result.png\" alt=\"https://raw.githubusercontent.com/bytedance/Next-ViT/main/images/result.png\"></p>\n<p>Quote from paper:<br>\n\"Table 4 is uniformly measured based on the TensorRT-8.0.3 framework with a T4 GPU (batch size=8)\" </p>",
      "rawMarkdown": "https://github.com/bytedance/Next-ViT\nif you are interested, you can check tensorRT friendly transformer.\nfrom my experiments, Next-ViT small has same performance as efficientb2 to b4\n\n![https://raw.githubusercontent.com/bytedance/Next-ViT/main/images/result.png](https://raw.githubusercontent.com/bytedance/Next-ViT/main/images/result.png)\n\nQuote from paper:\n\"Table 4 is uniformly measured based on the TensorRT-8.0.3 framework with a T4 GPU (batch size=8)\" \n\n",
      "votes": 2
    },
    {
      "id": 2085091,
      "postDate": "2023-01-03T23:35:05.617Z",
      "content": "<p>This is great <a href=\"https://www.kaggle.com/tivfrvqhs5\" target=\"_blank\">@tivfrvqhs5</a>! That is a huge difference, wow! Really cool to see <code>tensorrt</code> used in the wild! 🙂 Thx for sharing!</p>",
      "rawMarkdown": "This is great @tivfrvqhs5! That is a huge difference, wow! Really cool to see `tensorrt` used in the wild! 🙂 Thx for sharing!",
      "votes": 2
    },
    {
      "id": 2159628,
      "postDate": "2023-02-25T22:07:12.877Z",
      "content": "<p>By the way, I run the <a href=\"https://www.kaggle.com/code/tivfrvqhs5/torch-tensorrt-infer-fp16-and-fp32-benchmarks/notebook\" target=\"_blank\">notebook</a> on a P100 and on a GeForce 1080Ti GPUs and the inference speedup of Torch TensorRT models is greater in 1080Ti GPU than in a P100. Furthermore, curiously the Torch TensorRT fp32 model in the 1080Ti GPU is ~11% faster (images/sec) than in the P100 GPU. On another hand, and unlike the P100, in the 1080Ti GPUs the inference time using Torch TensorRT models with fp32 and fp16 models is almost the same:</p>\n<p><strong>Inference time (images/sec):</strong></p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>1080Ti</th>\n<th>P100</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>pytorch (FP32)</strong></td>\n<td>31.96</td>\n<td>47.86</td>\n</tr>\n<tr>\n<td><strong>torch_tensorrt (FP32)</strong></td>\n<td>73.41</td>\n<td>65.92</td>\n</tr>\n<tr>\n<td><strong>pytorch (FP16)</strong></td>\n<td>37.19</td>\n<td>54.62</td>\n</tr>\n<tr>\n<td><strong>torch_tensorrt (FP16)</strong></td>\n<td>73.65</td>\n<td>77.56</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "By the way, I run the [notebook](https://www.kaggle.com/code/tivfrvqhs5/torch-tensorrt-infer-fp16-and-fp32-benchmarks/notebook) on a P100 and on a GeForce 1080Ti GPUs and the inference speedup of Torch TensorRT models is greater in 1080Ti GPU than in a P100. Furthermore, curiously the Torch TensorRT fp32 model in the 1080Ti GPU is ~11% faster (images/sec) than in the P100 GPU. On another hand, and unlike the P100, in the 1080Ti GPUs the inference time using Torch TensorRT models with fp32 and fp16 models is almost the same:\n\n**Inference time (images/sec):**\n| | 1080Ti  | P100 |\n| --- | --- | --- |\n| **pytorch (FP32)** | 31.96 | 47.86  |\n| **torch_tensorrt (FP32)** | 73.41  | 65.92  |\n| **pytorch (FP16)** | 37.19  | 54.62  |\n| **torch_tensorrt (FP16)** | 73.65 | 77.56 |",
      "replies": [
        {
          "id": 2159629,
          "postDate": "2023-02-25T22:07:47.963Z",
          "content": "<p>Torch TensorRT inference speedup:</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>1080Ti</th>\n<th>P100</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>FP32</td>\n<td>2.3x</td>\n<td>1.4x</td>\n</tr>\n<tr>\n<td>FP16</td>\n<td>2.0x</td>\n<td>1.4x</td>\n</tr>\n</tbody>\n</table>",
          "rawMarkdown": "Torch TensorRT inference speedup:\n| | 1080Ti | P100 |\n| --- | --- | --- |\n| FP32 | 2.3x | 1.4x |\n| FP16 | 2.0x | 1.4x | "
        }
      ]
    },
    {
      "id": 2109479,
      "postDate": "2023-01-21T12:20:06.830Z",
      "content": "<p>using onnx + trtexec pipeline i can view the tensortRT engine graph like this:<br>\n<img src=\"https://i.ibb.co/zPy9M7X/Selection-614.png\" alt=\"https://i.ibb.co/zPy9M7X/Selection-614.png\"></p>\n<p>How can i do the same thing in torch_tensorrt???</p>",
      "rawMarkdown": "using onnx + trtexec pipeline i can view the tensortRT engine graph like this:\n![https://i.ibb.co/zPy9M7X/Selection-614.png](https://i.ibb.co/zPy9M7X/Selection-614.png)\n\nHow can i do the same thing in torch_tensorrt???",
      "replies": [
        {
          "id": 2109716,
          "postDate": "2023-01-21T15:46:09.697Z",
          "content": "<p>maybe worth having a look at <a href=\"https://developer.nvidia.com/blog/exploring-tensorrt-engines-with-trex/\" target=\"_blank\">https://developer.nvidia.com/blog/exploring-tensorrt-engines-with-trex/</a></p>",
          "rawMarkdown": "maybe worth having a look at https://developer.nvidia.com/blog/exploring-tensorrt-engines-with-trex/"
        }
      ]
    },
    {
      "id": 2087215,
      "postDate": "2023-01-05T12:34:14.397Z",
      "content": "<p>Thanks for you sharing:). A little change can consequence a great impact on inference time.</p>",
      "rawMarkdown": "Thanks for you sharing:). A little change can consequence a great impact on inference time."
    },
    {
      "id": 2099199,
      "postDate": "2023-01-14T08:46:18.450Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2100212,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-01-15T01:59:58.823000",
      "content": "<p>i think you should not use model.half() when compile 16fp in your tensortRT example:</p>\n<p>my experimental results</p>\n<pre><code>model effnet-b4 on 1536x960\n\n***************ok!\npytorch autocast:\n\n10935 / 10935   7 min 09 sec\n\nauc 0.8820045604670901\nf1score.max() 0.48506685839497704\n@threshold 0.3469387755102041\n\n***************ok!\ntensorrt\n\n10935 / 10935   4 min 01 sec\n\nauc 0.8824084583371677\nf1score.max() 0.48506685839497704\n@threshold 0.3469387755102041\n\n***************ok!\ntensorrt compiled with model.half()\n\n10935 / 10935   4 min 06 sec\n\nauc 0.8296021350348171\nf1score.max() 0.39217532383668313\n@threshold 0.04081632653061224\n\n***************ok!\n</code></pre>\n<pre><code>       f = torch.load(checkpoint, map_location=lambda storage, loc: storage)\n    model.load_state_dict(f['state_dict'], strict=False)\n    model.eval().cuda()\n    model.half()  ### should not be used !!!!!\n\n\n    print('torch_tensorrt.compile() ...')\n    with torch_tensorrt.logging.debug():\n        trt_model_fp16 = torch_tensorrt.compile(  \n            model,\n            inputs=[\n                torch_tensorrt.Input(\n                    [batch_size, 1, image_height, image_width],\n                    dtype=torch.half\n                )],\n            enabled_precisions={torch.half},  \n            workspace_size=1 &lt;&lt; 32,\n            require_full_compilation=True,\n            #debug=True,\n        )\n    torch.jit.save(trt_model_fp16, tft_file)\n</code></pre>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2098825,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-01-13T21:51:06.163000",
      "content": "<p>my timing for nextVIT transformer with and without tensorRT on P100.<br>\nIt is about 4x speed improvement.</p>\n<p>speed is about the same as claim in paper (i.e. about efficientnet speed)</p>\n<pre><code>** tesorRT **\n - public LB submission tensorRT 4 hr : LB 0.56\n - local cv 10939 images: \ntensorRT fp16 nextVIT-B (1539x960) \n\nCPU utilisation 120%,  13/13 GB\nGPU utilisation 99% , 4.5/15 GB\n23 min 21 sec (7.80799 images per sec)\n\n** without tesorRT **\n- submission  9hr : LB 0.56\n- local cv 10939 images: \nmerged_bn fp16 (1539x960)  \n\nCPU utilisation 108%,  13/13 GB\nGPU utilisation 100% , 8.5/15 GB\n95 min 32 sec (1.90819 images per sec)\n</code></pre>\n<p>more details breakdown of timing at <br>\n<a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/370333#2098821\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/370333#2098821</a> </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2099205,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2023-01-14T08:50:05.130000",
          "content": "<p>there seems to be some unknown bug in my implementation.<br>\nthe bug is in kaggle notebook or tensorrt lib or the nextvit code?<br>\n<a href=\"https://www.kaggle.com/code/hengck23/3hr-tensorrt-nextvit-example\" target=\"_blank\">https://www.kaggle.com/code/hengck23/3hr-tensorrt-nextvit-example</a></p>\n<p>when i call torch_tensorrt.compile(), it seems to stall and nothing happens after 10 min.<br>\nthen click the stop run button at the side, the process is terminated with usual verbal messages flushed out.<br>\nif i run again, this produces the tft engine file finally.</p>\n<p>i tried a few times from fresh start, and it seems this is always the case.</p>\n<p>(this doesn't happens on my local machine with the same code, but it uses pycharm ide instead of notebook)</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2100125,
              "author_name": "David Austin",
              "author_url": "",
              "post_date": "2023-01-15T00:17:01.800000",
              "content": "<p>I've seen this several times when running a jupyter notebook, but it doesn't happen when running from a .py file.  There are github issues reporting the same</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2110681,
              "author_name": "MPWARE",
              "author_url": "",
              "post_date": "2023-01-22T10:50:05.207000",
              "content": "<p>I've faced this issue too and it's fixed with Torch Tensor RT 1.3.0</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2112897,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-01-23T23:26:26.857000",
              "content": "<p>thanks.</p>\n<p>i solved the issue by using the following in kaggle notebook</p>\n<pre><code>batch_size = 4\nfrom model_nextvit_multi import *\nif 1:\n    ! python /kaggle/input/.../convert_model_nextvit_multi.py --checkpoint_file ...\n</code></pre>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2087644,
      "author_name": "David Austin",
      "author_url": "",
      "post_date": "2023-01-05T18:24:36.023000",
      "content": "<p>Update:  Running the same code on a 2x T4 notebook shows marginally faster times for torch_tensorrt but slower times for pytorch inference (single GPU mode).  The reason for this is likely that T4 GPU's have fp32 and fp16 tensor cores which torch_tensorrt utilizes but P100 GPU's don't have them.  Even faster speedup could be achieved in dual GPU inference mode assuming your dataloader isn't a bottleneck.</p>\n<p>Speedup results for EFN B-2 using 1024x1024x3 image size on T4 GPU (single GPU mode).</p>\n<table>\n  <tbody><tr>\n    <th></th>\n    <th>pytorch</th>\n    <th>torch_tensorrt</th>\n  </tr>\n  <tr>\n    <th>FP32</th>\n    <td>29.1</td>\n    <td>37.63</td>\n  </tr>\n  <tr>\n    <th>FP16</th>\n    <td>56.87</td>\n    <td>81.64</td>\n  </tr>\n</tbody></table>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2098401,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-01-13T14:10:01.740000",
      "content": "<p>ironically, the speed gain by tensorrt model is reduce by the extra long time to install tensorrt</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2098676,
          "author_name": "David Austin",
          "author_url": "",
          "post_date": "2023-01-13T18:19:21.303000",
          "content": "<p>Yes it's a 3-4 minute install time, it would be nice if tensorrt was included in the base docker build since it helps reduce the compute load on the system.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2098725,
              "author_name": "Sohier Dane",
              "author_url": "",
              "post_date": "2023-01-13T19:02:10.893000",
              "content": "<p>We're always accepting new package proposals here: <a href=\"https://github.com/Kaggle/docker-python#requesting-new-packages\" target=\"_blank\">https://github.com/Kaggle/docker-python#requesting-new-packages</a></p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 2110709,
              "author_name": "MPWARE",
              "author_url": "",
              "post_date": "2023-01-22T11:07:35.763000",
              "content": "<p>Request done (open issue)</p>",
              "votes": 5,
              "replies": []
            },
            {
              "id": 2112620,
              "author_name": "Sohier Dane",
              "author_url": "",
              "post_date": "2023-01-23T18:07:19.550000",
              "content": "<p>Great, thanks! I'll let the engineer who handles that repo know just how much use this library is getting.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2112678,
              "author_name": "MPWARE",
              "author_url": "",
              "post_date": "2023-01-23T18:45:35.850000",
              "content": "<p>I've requested for TensorRT 1.3 which is better than 1.2 already demonstrated here. I'm able to install and make it works on current Kaggle environement but I have to patch torch_tensorrt source code. I believe your guys will face similar issues. It will be great if it could be available. </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2114162,
              "author_name": "Sohier Dane",
              "author_url": "",
              "post_date": "2023-01-24T18:38:20.920000",
              "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> it looks like there are some version conflicts with your proposal: <a href=\"https://github.com/Kaggle/docker-python/issues/1210\" target=\"_blank\">https://github.com/Kaggle/docker-python/issues/1210</a></p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2114262,
              "author_name": "MPWARE",
              "author_url": "",
              "post_date": "2023-01-24T20:18:15.937000",
              "content": "<p>Yes, I've seen the answer and I understand the conflict. We've to wait to have it in a next docker image. In the mean time, as suggested by the answer, we can install it ourself (works for me). </p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2088500,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-01-06T11:33:10.510000",
      "content": "<p><a href=\"https://github.com/bytedance/Next-ViT\" target=\"_blank\">https://github.com/bytedance/Next-ViT</a><br>\nif you are interested, you can check tensorRT friendly transformer.<br>\nfrom my experiments, Next-ViT small has same performance as efficientb2 to b4</p>\n<p><img src=\"https://raw.githubusercontent.com/bytedance/Next-ViT/main/images/result.png\" alt=\"https://raw.githubusercontent.com/bytedance/Next-ViT/main/images/result.png\"></p>\n<p>Quote from paper:<br>\n\"Table 4 is uniformly measured based on the TensorRT-8.0.3 framework with a T4 GPU (batch size=8)\" </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2085091,
      "author_name": "Radek Osmulski",
      "author_url": "",
      "post_date": "2023-01-03T23:35:05.617000",
      "content": "<p>This is great <a href=\"https://www.kaggle.com/tivfrvqhs5\" target=\"_blank\">@tivfrvqhs5</a>! That is a huge difference, wow! Really cool to see <code>tensorrt</code> used in the wild! 🙂 Thx for sharing!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2159628,
      "author_name": "Pablo Rios",
      "author_url": "",
      "post_date": "2023-02-25T22:07:12.877000",
      "content": "<p>By the way, I run the <a href=\"https://www.kaggle.com/code/tivfrvqhs5/torch-tensorrt-infer-fp16-and-fp32-benchmarks/notebook\" target=\"_blank\">notebook</a> on a P100 and on a GeForce 1080Ti GPUs and the inference speedup of Torch TensorRT models is greater in 1080Ti GPU than in a P100. Furthermore, curiously the Torch TensorRT fp32 model in the 1080Ti GPU is ~11% faster (images/sec) than in the P100 GPU. On another hand, and unlike the P100, in the 1080Ti GPUs the inference time using Torch TensorRT models with fp32 and fp16 models is almost the same:</p>\n<p><strong>Inference time (images/sec):</strong></p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>1080Ti</th>\n<th>P100</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>pytorch (FP32)</strong></td>\n<td>31.96</td>\n<td>47.86</td>\n</tr>\n<tr>\n<td><strong>torch_tensorrt (FP32)</strong></td>\n<td>73.41</td>\n<td>65.92</td>\n</tr>\n<tr>\n<td><strong>pytorch (FP16)</strong></td>\n<td>37.19</td>\n<td>54.62</td>\n</tr>\n<tr>\n<td><strong>torch_tensorrt (FP16)</strong></td>\n<td>73.65</td>\n<td>77.56</td>\n</tr>\n</tbody>\n</table>",
      "votes": 0,
      "replies": [
        {
          "id": 2159629,
          "author_name": "Pablo Rios",
          "author_url": "",
          "post_date": "2023-02-25T22:07:47.963000",
          "content": "<p>Torch TensorRT inference speedup:</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>1080Ti</th>\n<th>P100</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>FP32</td>\n<td>2.3x</td>\n<td>1.4x</td>\n</tr>\n<tr>\n<td>FP16</td>\n<td>2.0x</td>\n<td>1.4x</td>\n</tr>\n</tbody>\n</table>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2109479,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-01-21T12:20:06.830000",
      "content": "<p>using onnx + trtexec pipeline i can view the tensortRT engine graph like this:<br>\n<img src=\"https://i.ibb.co/zPy9M7X/Selection-614.png\" alt=\"https://i.ibb.co/zPy9M7X/Selection-614.png\"></p>\n<p>How can i do the same thing in torch_tensorrt???</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2109716,
          "author_name": "Eleftherios Fanioudakis",
          "author_url": "",
          "post_date": "2023-01-21T15:46:09.697000",
          "content": "<p>maybe worth having a look at <a href=\"https://developer.nvidia.com/blog/exploring-tensorrt-engines-with-trex/\" target=\"_blank\">https://developer.nvidia.com/blog/exploring-tensorrt-engines-with-trex/</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2087215,
      "author_name": "Armin Azhdehnia",
      "author_url": "",
      "post_date": "2023-01-05T12:34:14.397000",
      "content": "<p>Thanks for you sharing:). A little change can consequence a great impact on inference time.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2099199,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-01-14T08:46:18.450000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2085061": "Speeding up model inference in this competition could end up being very important.  [torch_tensorrt ](https://github.com/pytorch/TensorRT) is an integration for PyTorch that leverages inference optimizations of TensorRT on NVIDIA GPUs.  I've created a notebook and dataset that will allow you install all the dependencies (internet disabled) and compile a torch_tensorrt model from your pytorch model, either in fp32 or fp16.\n\n[Notebook link](https://www.kaggle.com/code/tivfrvqhs5/torch-tensorrt-infer-fp16-and-fp32-benchmarks/notebook)\n\nSpeedup results for EFN-B2 using 1024x1024x3 image size (images/sec).\n<table>\n  <tr>\n    <th></th>\n    <th>pytorch</th>\n    <th>torch_tensorrt</th>\n  </tr>\n  <tr>\n    <th>FP32</th>\n    <td>47.91</td>\n    <td>66.07</td>\n  </tr>\n  <tr>\n    <th>FP16</th>\n    <td>54.72</td>\n    <td>77.86</td>\n  </tr>\n</table>\n\n",
    "2100212": "i think you should not use model.half() when compile 16fp in your tensortRT example:\n\nmy experimental results\n```\nmodel effnet-b4 on 1536x960\n\n***************ok!\npytorch autocast:\n\n10935 / 10935   7 min 09 sec\n \nauc 0.8820045604670901\nf1score.max() 0.48506685839497704\n@threshold 0.3469387755102041\n\n***************ok!\ntensorrt\n\n10935 / 10935   4 min 01 sec\n\nauc 0.8824084583371677\nf1score.max() 0.48506685839497704\n@threshold 0.3469387755102041\n\n***************ok!\ntensorrt compiled with model.half()\n\n10935 / 10935   4 min 06 sec\n\nauc 0.8296021350348171\nf1score.max() 0.39217532383668313\n@threshold 0.04081632653061224\n\n***************ok!\n\n```\n\n```\n       f = torch.load(checkpoint, map_location=lambda storage, loc: storage)\n\tmodel.load_state_dict(f['state_dict'], strict=False)\n\tmodel.eval().cuda()\n\tmodel.half()  ### should not be used !!!!!\n\n\n\tprint('torch_tensorrt.compile() ...')\n\twith torch_tensorrt.logging.debug():\n\t\ttrt_model_fp16 = torch_tensorrt.compile(  \n\t\t\tmodel,\n\t\t\tinputs=[\n\t\t\t\ttorch_tensorrt.Input(\n\t\t\t\t\t[batch_size, 1, image_height, image_width],\n\t\t\t\t\tdtype=torch.half\n\t\t\t\t)],\n\t\t\tenabled_precisions={torch.half},  \n\t\t\tworkspace_size=1 << 32,\n\t\t\trequire_full_compilation=True,\n\t\t\t#debug=True,\n\t\t)\n\ttorch.jit.save(trt_model_fp16, tft_file)\n\n```\n",
    "2098825": "my timing for nextVIT transformer with and without tensorRT on P100.\nIt is about 4x speed improvement.\n\nspeed is about the same as claim in paper (i.e. about efficientnet speed)\n\n```\n** tesorRT **\n - public LB submission tensorRT 4 hr : LB 0.56\n - local cv 10939 images: \ntensorRT fp16 nextVIT-B (1539x960) \n\nCPU utilisation 120%,  13/13 GB\nGPU utilisation 99% , 4.5/15 GB\n23 min 21 sec (7.80799 images per sec)\n\n** without tesorRT **\n- submission  9hr : LB 0.56\n- local cv 10939 images: \nmerged_bn fp16 (1539x960)  \n\nCPU utilisation 108%,  13/13 GB\nGPU utilisation 100% , 8.5/15 GB\n95 min 32 sec (1.90819 images per sec)\n\n```\n\nmore details breakdown of timing at \nhttps://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/370333#2098821 \n",
    "2087644": "Update:  Running the same code on a 2x T4 notebook shows marginally faster times for torch_tensorrt but slower times for pytorch inference (single GPU mode).  The reason for this is likely that T4 GPU's have fp32 and fp16 tensor cores which torch_tensorrt utilizes but P100 GPU's don't have them.  Even faster speedup could be achieved in dual GPU inference mode assuming your dataloader isn't a bottleneck.\n\nSpeedup results for EFN B-2 using 1024x1024x3 image size on T4 GPU (single GPU mode).\n\n<table>\n  <tr>\n    <th></th>\n    <th>pytorch</th>\n    <th>torch_tensorrt</th>\n  </tr>\n  <tr>\n    <th>FP32</th>\n    <td>29.1</td>\n    <td>37.63</td>\n  </tr>\n  <tr>\n    <th>FP16</th>\n    <td>56.87</td>\n    <td>81.64</td>\n  </tr>\n</table>",
    "2098401": "ironically, the speed gain by tensorrt model is reduce by the extra long time to install tensorrt",
    "2088500": "https://github.com/bytedance/Next-ViT\nif you are interested, you can check tensorRT friendly transformer.\nfrom my experiments, Next-ViT small has same performance as efficientb2 to b4\n\n![https://raw.githubusercontent.com/bytedance/Next-ViT/main/images/result.png](https://raw.githubusercontent.com/bytedance/Next-ViT/main/images/result.png)\n\nQuote from paper:\n\"Table 4 is uniformly measured based on the TensorRT-8.0.3 framework with a T4 GPU (batch size=8)\" \n\n",
    "2085091": "This is great @tivfrvqhs5! That is a huge difference, wow! Really cool to see `tensorrt` used in the wild! 🙂 Thx for sharing!",
    "2159628": "By the way, I run the [notebook](https://www.kaggle.com/code/tivfrvqhs5/torch-tensorrt-infer-fp16-and-fp32-benchmarks/notebook) on a P100 and on a GeForce 1080Ti GPUs and the inference speedup of Torch TensorRT models is greater in 1080Ti GPU than in a P100. Furthermore, curiously the Torch TensorRT fp32 model in the 1080Ti GPU is ~11% faster (images/sec) than in the P100 GPU. On another hand, and unlike the P100, in the 1080Ti GPUs the inference time using Torch TensorRT models with fp32 and fp16 models is almost the same:\n\n**Inference time (images/sec):**\n| | 1080Ti  | P100 |\n| --- | --- | --- |\n| **pytorch (FP32)** | 31.96 | 47.86  |\n| **torch_tensorrt (FP32)** | 73.41  | 65.92  |\n| **pytorch (FP16)** | 37.19  | 54.62  |\n| **torch_tensorrt (FP16)** | 73.65 | 77.56 |",
    "2109479": "using onnx + trtexec pipeline i can view the tensortRT engine graph like this:\n![https://i.ibb.co/zPy9M7X/Selection-614.png](https://i.ibb.co/zPy9M7X/Selection-614.png)\n\nHow can i do the same thing in torch_tensorrt???",
    "2087215": "Thanks for you sharing:). A little change can consequence a great impact on inference time.",
    "2099199": ""
  }
}