{
  "id": 191312,
  "title": "Am I wasting my time trying pytorch on Jpegs?",
  "url": "/competitions/rsna-str-pulmonary-embolism-detection/discussion/191312",
  "author_name": "Mark P",
  "post_date": "2020-10-15T19:19:46.706000",
  "votes": 0,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hi all,</p>\n<p>I've been trying to train an efficient b0 on Ian Pan's jpegs and trying all the tricks I can to speed up training but its still very slow - the CPU is proving to be a bottleneck - during training it hits 150% whilst GPU is only about 25%. Probably this is in the dataloader?</p>\n<p>Changing pin_memory=True and increasing num_workers to 16 improved time per epoch from 2hr15 to about 1hr30 and then using turbojeg thanks to this kind fellow <a href=\"https://www.kaggle.com/kozodoi\" target=\"_blank\">@kozodoi</a> reduced it to 1hr15 per epoch but this is still too long…</p>\n<p>But is it all just a waste of time?  I really want to stick with pytorch as a learning experience (though I reckon a TPU with TFrecords would be a lot easier and quicker).  Am I wasting my time or can someone encourage me that this might bear fruit?</p>\n<p>My notebook is here:</p>\n<p><a href=\"url\" target=\"_blank\">https://www.kaggle.com/cascadenite/pytorch-effnet-jpegs</a></p>",
  "messages": [
    {
      "id": 1050961,
      "postDate": "2020-10-16T00:58:58.437Z",
      "content": "<p>In my experiments, eff-b0 took about 15 min with Ian's dataset with bs=32.</p>",
      "rawMarkdown": "In my experiments, eff-b0 took about 15 min with Ian's dataset with bs=32.\n",
      "votes": 1,
      "replies": [
        {
          "id": 1050969,
          "postDate": "2020-10-16T01:22:33.517Z",
          "content": "<p>You're saying it takes 15 minutes to load (much less train on) 1.8 million images?</p>",
          "rawMarkdown": "You're saying it takes 15 minutes to load (much less train on) 1.8 million images?"
        },
        {
          "id": 1050971,
          "postDate": "2020-10-16T01:45:02.877Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1050972,
          "postDate": "2020-10-16T01:49:29.083Z",
          "content": "<p>Let me elaborate: I actually downsample negative images, so the actual image trained is much less (about 100k)</p>",
          "rawMarkdown": "Let me elaborate: I actually downsample negative images, so the actual image trained is much less (about 100k)",
          "votes": 1
        },
        {
          "id": 1051655,
          "postDate": "2020-10-16T17:52:37.070Z",
          "content": "<p><a href=\"https://www.kaggle.com/kyoshioka47\" target=\"_blank\">@kyoshioka47</a>  how would one take qi  for loss here ,if we do downsample for postive cts ?</p>",
          "rawMarkdown": "@kyoshioka47  how would one take qi  for loss here ,if we do downsample for postive cts ?"
        }
      ]
    },
    {
      "id": 1050872,
      "postDate": "2020-10-15T20:06:34.013Z",
      "content": "<p>Don't be discouraged. This is a very resource intensive competition - RAM is a big constraint if you are using Kaggle GPUs. Yes, TPUs are the way to go if you have a pipeline ready. </p>",
      "rawMarkdown": "Don't be discouraged. This is a very resource intensive competition - RAM is a big constraint if you are using Kaggle GPUs. Yes, TPUs are the way to go if you have a pipeline ready. ",
      "votes": 1
    },
    {
      "id": 1050868,
      "postDate": "2020-10-15T19:58:51.113Z",
      "content": "<p>Try using without pin_memory=True and instead using mixed precision, then try to double the batch size. Try and see if it helps.</p>",
      "rawMarkdown": "Try using without pin_memory=True and instead using mixed precision, then try to double the batch size. Try and see if it helps."
    },
    {
      "id": 1050838,
      "postDate": "2020-10-15T19:19:46.707Z",
      "content": "<p>Hi all,</p>\n<p>I've been trying to train an efficient b0 on Ian Pan's jpegs and trying all the tricks I can to speed up training but its still very slow - the CPU is proving to be a bottleneck - during training it hits 150% whilst GPU is only about 25%. Probably this is in the dataloader?</p>\n<p>Changing pin_memory=True and increasing num_workers to 16 improved time per epoch from 2hr15 to about 1hr30 and then using turbojeg thanks to this kind fellow <a href=\"https://www.kaggle.com/kozodoi\" target=\"_blank\">@kozodoi</a> reduced it to 1hr15 per epoch but this is still too long…</p>\n<p>But is it all just a waste of time?  I really want to stick with pytorch as a learning experience (though I reckon a TPU with TFrecords would be a lot easier and quicker).  Am I wasting my time or can someone encourage me that this might bear fruit?</p>\n<p>My notebook is here:</p>\n<p><a href=\"url\" target=\"_blank\">https://www.kaggle.com/cascadenite/pytorch-effnet-jpegs</a></p>",
      "rawMarkdown": "Hi all,\n\nI've been trying to train an efficient b0 on Ian Pan's jpegs and trying all the tricks I can to speed up training but its still very slow - the CPU is proving to be a bottleneck - during training it hits 150% whilst GPU is only about 25%. Probably this is in the dataloader?\n\nChanging pin_memory=True and increasing num_workers to 16 improved time per epoch from 2hr15 to about 1hr30 and then using turbojeg thanks to this kind fellow @kozodoi reduced it to 1hr15 per epoch but this is still too long...\n\nBut is it all just a waste of time?  I really want to stick with pytorch as a learning experience (though I reckon a TPU with TFrecords would be a lot easier and quicker).  Am I wasting my time or can someone encourage me that this might bear fruit?\n\nMy notebook is here:\n\n[https://www.kaggle.com/cascadenite/pytorch-effnet-jpegs](url)\n"
    }
  ],
  "comments": [
    {
      "id": 1050961,
      "author_name": "arutema47",
      "author_url": "",
      "post_date": "2020-10-16T00:58:58.437000",
      "content": "<p>In my experiments, eff-b0 took about 15 min with Ian's dataset with bs=32.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1050969,
          "author_name": "holste1",
          "author_url": "",
          "post_date": "2020-10-16T01:22:33.517000",
          "content": "<p>You're saying it takes 15 minutes to load (much less train on) 1.8 million images?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1050971,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-10-16T01:45:02.877000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1050972,
          "author_name": "arutema47",
          "author_url": "",
          "post_date": "2020-10-16T01:49:29.083000",
          "content": "<p>Let me elaborate: I actually downsample negative images, so the actual image trained is much less (about 100k)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1051655,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-10-16T17:52:37.070000",
          "content": "<p><a href=\"https://www.kaggle.com/kyoshioka47\" target=\"_blank\">@kyoshioka47</a>  how would one take qi  for loss here ,if we do downsample for postive cts ?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1050872,
      "author_name": "Vee",
      "author_url": "",
      "post_date": "2020-10-15T20:06:34.013000",
      "content": "<p>Don't be discouraged. This is a very resource intensive competition - RAM is a big constraint if you are using Kaggle GPUs. Yes, TPUs are the way to go if you have a pipeline ready. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1050868,
      "author_name": "Kirderf",
      "author_url": "",
      "post_date": "2020-10-15T19:58:51.113000",
      "content": "<p>Try using without pin_memory=True and instead using mixed precision, then try to double the batch size. Try and see if it helps.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1050961": "In my experiments, eff-b0 took about 15 min with Ian's dataset with bs=32.\n",
    "1050872": "Don't be discouraged. This is a very resource intensive competition - RAM is a big constraint if you are using Kaggle GPUs. Yes, TPUs are the way to go if you have a pipeline ready. ",
    "1050868": "Try using without pin_memory=True and instead using mixed precision, then try to double the batch size. Try and see if it helps.",
    "1050838": "Hi all,\n\nI've been trying to train an efficient b0 on Ian Pan's jpegs and trying all the tricks I can to speed up training but its still very slow - the CPU is proving to be a bottleneck - during training it hits 150% whilst GPU is only about 25%. Probably this is in the dataloader?\n\nChanging pin_memory=True and increasing num_workers to 16 improved time per epoch from 2hr15 to about 1hr30 and then using turbojeg thanks to this kind fellow @kozodoi reduced it to 1hr15 per epoch but this is still too long...\n\nBut is it all just a waste of time?  I really want to stick with pytorch as a learning experience (though I reckon a TPU with TFrecords would be a lot easier and quicker).  Am I wasting my time or can someone encourage me that this might bear fruit?\n\nMy notebook is here:\n\n[https://www.kaggle.com/cascadenite/pytorch-effnet-jpegs](url)\n"
  }
}