{
  "id": 377986,
  "title": "Has anyone compared performance between P100 and T4 notebooks?",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/377986",
  "author_name": "hengck23",
  "post_date": "2023-01-13T20:30:21.484000",
  "votes": 6,
  "comment_count": 7,
  "views": 0,
  "content": "<p><br>\n</p>\n<p>T4 CPU also has 2 core, see below</p>\n<p>So is there an overall gain for using 2xT4?<br>\nI would like to hear your views before trying it, thanks!</p>\n<p>please report time to convert dicom images using dicomsdl, etc. </p>",
  "messages": [
    {
      "id": 2098776,
      "postDate": "2023-01-13T20:30:21.483Z",
      "content": "<p><br>\n</p>\n<p>T4 CPU also has 2 core, see below</p>\n<p>So is there an overall gain for using 2xT4?<br>\nI would like to hear your views before trying it, thanks!</p>\n<p>please report time to convert dicom images using dicomsdl, etc. </p>",
      "rawMarkdown": "~~while T4 has 2 GPU, but its CPU has only one core.~~\n~~This means that the decoding of dicom images will be slower.~~\n\nT4 CPU also has 2 core, see below\n\n\nSo is there an overall gain for using 2xT4?\nI would like to hear your views before trying it, thanks!\n\nplease report time to convert dicom images using dicomsdl, etc. ",
      "votes": 6
    },
    {
      "id": 2098819,
      "postDate": "2023-01-13T21:35:38.503Z",
      "content": "<p>Why do you think 2xT4 has only one CPU Core?</p>\n<p>Both P100 and 2XT4 kernels have same CPU.</p>\n<pre><code>CPU(s):                          \nOn-line CPU(s) :             ,\nThread(s) per core:              \nCore(s) per socket:              \nSocket(s):                       \nNUMA node(s):                    \nVendor ID:                       GenuineIntel\nCPU family:                      \nModel:                           \nModel name:                      Intel(R) Xeon(R) CPU @ GHz\n</code></pre>",
      "rawMarkdown": "Why do you think 2xT4 has only one CPU Core?\n\nBoth P100 and 2XT4 kernels have same CPU.\n\n```python\nCPU(s):                          2\nOn-line CPU(s) list:             0,1\nThread(s) per core:              2\nCore(s) per socket:              1\nSocket(s):                       1\nNUMA node(s):                    1\nVendor ID:                       GenuineIntel\nCPU family:                      6\nModel:                           85\nModel name:                      Intel(R) Xeon(R) CPU @ 2.00GHz\n```",
      "votes": 3,
      "replies": [
        {
          "id": 2098822,
          "postDate": "2023-01-13T21:43:41.507Z",
          "content": "<p>thanks, it is my mistake. then there will definitely be a gain in speed!</p>",
          "rawMarkdown": "thanks, it is my mistake. then there will definitely be a gain in speed!"
        }
      ]
    },
    {
      "id": 2114308,
      "postDate": "2023-01-24T21:17:45.140Z",
      "content": "<p>In general i think, 1xT4 is about the same speed as 1xP100 when using full precision for both but T4 will benefit from using mixed precision whereas P100 will not. So 1xT4 in mixed precision is roughly 2x faster than 1x P100 in full precision. This is because T4 have special tensor cores and P100 do not. Also 2xT4 should be roughly 2x faster than 1x P100.</p>",
      "rawMarkdown": "In general i think, 1xT4 is about the same speed as 1xP100 when using full precision for both but T4 will benefit from using mixed precision whereas P100 will not. So 1xT4 in mixed precision is roughly 2x faster than 1x P100 in full precision. This is because T4 have special tensor cores and P100 do not. Also 2xT4 should be roughly 2x faster than 1x P100.",
      "votes": 1
    },
    {
      "id": 2112900,
      "postDate": "2023-01-23T23:34:16.163Z",
      "content": "<p>simplest (not necessarily the best) way to use 2 gpu</p>\n<pre><code>#one gpu code\noutput = net(batch) \n</code></pre>\n<pre><code>#one or two gpu code\nimport os \nos.environ['CUDA_VISIBLE_DEVICES'] = '0,1'\nfrom torch.nn.parallel.data_parallel import data_parallel\n\noutput = data_parallel(net, batch)  # only need to change one line of code\n</code></pre>\n<p>you need not have to worry about save or load model, etc</p>",
      "rawMarkdown": "simplest (not necessarily the best) way to use 2 gpu\n\n```\n#one gpu code\noutput = net(batch) \n```\n\n```\n#one or two gpu code\nimport os \nos.environ['CUDA_VISIBLE_DEVICES'] = '0,1'\nfrom torch.nn.parallel.data_parallel import data_parallel\n\noutput = data_parallel(net, batch)  # only need to change one line of code\n```\nyou need not have to worry about save or load model, etc",
      "votes": 1
    },
    {
      "id": 2114972,
      "postDate": "2023-01-25T11:20:13.823Z",
      "content": "<p>2xT4 gave me nearly 1.7 times speed up. But it's all about tiny models</p>",
      "rawMarkdown": "2xT4 gave me nearly 1.7 times speed up. But it's all about tiny models"
    },
    {
      "id": 2102021,
      "postDate": "2023-01-16T10:52:04.413Z",
      "content": "<p>From my experience, P100 was slightly faster than T4 ×2 for DICOM + DALI conversion and resize : 184.9&nbsp;s vs 219.6&nbsp;s (full notebook time) for converting 50 images to 256&nbsp;px large, respectively.</p>\n<p>(I didn’t check how many of the 50 images were dicomsdl-converted and how many where DALI-converted, though.)</p>",
      "rawMarkdown": "From my experience, P100 was slightly faster than T4 ×2 for DICOM + DALI conversion and resize : 184.9 s vs 219.6 s (full notebook time) for converting 50 images to 256 px large, respectively.\n\n(I didn’t check how many of the 50 images were dicomsdl-converted and how many where DALI-converted, though.)"
    },
    {
      "id": 2099059,
      "postDate": "2023-01-14T06:05:44.897Z",
      "content": "<p>For submission, they take almost same time, over 8hrs.  I forked public notebooks .</p>",
      "rawMarkdown": "For submission, they take almost same time, over 8hrs.  I forked public notebooks ."
    }
  ],
  "comments": [
    {
      "id": 2098819,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2023-01-13T21:35:38.503000",
      "content": "<p>Why do you think 2xT4 has only one CPU Core?</p>\n<p>Both P100 and 2XT4 kernels have same CPU.</p>\n<pre><code>CPU(s):                          \nOn-line CPU(s) :             ,\nThread(s) per core:              \nCore(s) per socket:              \nSocket(s):                       \nNUMA node(s):                    \nVendor ID:                       GenuineIntel\nCPU family:                      \nModel:                           \nModel name:                      Intel(R) Xeon(R) CPU @ GHz\n</code></pre>",
      "votes": 3,
      "replies": [
        {
          "id": 2098822,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2023-01-13T21:43:41.507000",
          "content": "<p>thanks, it is my mistake. then there will definitely be a gain in speed!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2114308,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2023-01-24T21:17:45.140000",
      "content": "<p>In general i think, 1xT4 is about the same speed as 1xP100 when using full precision for both but T4 will benefit from using mixed precision whereas P100 will not. So 1xT4 in mixed precision is roughly 2x faster than 1x P100 in full precision. This is because T4 have special tensor cores and P100 do not. Also 2xT4 should be roughly 2x faster than 1x P100.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2112900,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-01-23T23:34:16.163000",
      "content": "<p>simplest (not necessarily the best) way to use 2 gpu</p>\n<pre><code>#one gpu code\noutput = net(batch) \n</code></pre>\n<pre><code>#one or two gpu code\nimport os \nos.environ['CUDA_VISIBLE_DEVICES'] = '0,1'\nfrom torch.nn.parallel.data_parallel import data_parallel\n\noutput = data_parallel(net, batch)  # only need to change one line of code\n</code></pre>\n<p>you need not have to worry about save or load model, etc</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2114972,
      "author_name": "Pizzaboi",
      "author_url": "",
      "post_date": "2023-01-25T11:20:13.823000",
      "content": "<p>2xT4 gave me nearly 1.7 times speed up. But it's all about tiny models</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2102021,
      "author_name": "gguillard",
      "author_url": "",
      "post_date": "2023-01-16T10:52:04.413000",
      "content": "<p>From my experience, P100 was slightly faster than T4 ×2 for DICOM + DALI conversion and resize : 184.9&nbsp;s vs 219.6&nbsp;s (full notebook time) for converting 50 images to 256&nbsp;px large, respectively.</p>\n<p>(I didn’t check how many of the 50 images were dicomsdl-converted and how many where DALI-converted, though.)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2099059,
      "author_name": "dragon zhang",
      "author_url": "",
      "post_date": "2023-01-14T06:05:44.897000",
      "content": "<p>For submission, they take almost same time, over 8hrs.  I forked public notebooks .</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2098776": "~~while T4 has 2 GPU, but its CPU has only one core.~~\n~~This means that the decoding of dicom images will be slower.~~\n\nT4 CPU also has 2 core, see below\n\n\nSo is there an overall gain for using 2xT4?\nI would like to hear your views before trying it, thanks!\n\nplease report time to convert dicom images using dicomsdl, etc. ",
    "2098819": "Why do you think 2xT4 has only one CPU Core?\n\nBoth P100 and 2XT4 kernels have same CPU.\n\n```python\nCPU(s):                          2\nOn-line CPU(s) list:             0,1\nThread(s) per core:              2\nCore(s) per socket:              1\nSocket(s):                       1\nNUMA node(s):                    1\nVendor ID:                       GenuineIntel\nCPU family:                      6\nModel:                           85\nModel name:                      Intel(R) Xeon(R) CPU @ 2.00GHz\n```",
    "2114308": "In general i think, 1xT4 is about the same speed as 1xP100 when using full precision for both but T4 will benefit from using mixed precision whereas P100 will not. So 1xT4 in mixed precision is roughly 2x faster than 1x P100 in full precision. This is because T4 have special tensor cores and P100 do not. Also 2xT4 should be roughly 2x faster than 1x P100.",
    "2112900": "simplest (not necessarily the best) way to use 2 gpu\n\n```\n#one gpu code\noutput = net(batch) \n```\n\n```\n#one or two gpu code\nimport os \nos.environ['CUDA_VISIBLE_DEVICES'] = '0,1'\nfrom torch.nn.parallel.data_parallel import data_parallel\n\noutput = data_parallel(net, batch)  # only need to change one line of code\n```\nyou need not have to worry about save or load model, etc",
    "2114972": "2xT4 gave me nearly 1.7 times speed up. But it's all about tiny models",
    "2102021": "From my experience, P100 was slightly faster than T4 ×2 for DICOM + DALI conversion and resize : 184.9 s vs 219.6 s (full notebook time) for converting 50 images to 256 px large, respectively.\n\n(I didn’t check how many of the 50 images were dicomsdl-converted and how many where DALI-converted, though.)",
    "2099059": "For submission, they take almost same time, over 8hrs.  I forked public notebooks ."
  }
}