{
  "id": 388403,
  "title": "TPU training strange behavior",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/388403",
  "author_name": "Igor Litvin",
  "post_date": "2023-02-17T09:45:47.993000",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Good day.</p>\n<p>I have learned how to use TPU (TF) and used it recently for this competition.</p>\n<p>It work extremally faster but I can not receive even the same score which I got by GPU (for exactly the same dataset and resolution).</p>\n<p>The validation and LB scores are much lower in comparison to GPU training. </p>\n<p>Possible somebody got such issue. Does TPU training is the same as GPU in terms of backpropagation (or it needs some special knowledge)? </p>\n<p>I use the same fit function model was compiled in strategy.scope():</p>\n<p>model.fit(images, labels,shuffle = True, epochs=1, verbose=1, batch_size =batch_size, steps_per_epoch=n_batches, callbacks=callbacks_list,class_weight=class_weight)</p>\n<p>Speed is great but validation losses much higher than for GPU.</p>\n<p>Thanks.</p>\n<p>Igor.</p>",
  "messages": [
    {
      "id": 2148289,
      "postDate": "2023-02-17T10:16:19.617Z",
      "content": "<p>The Kaggle TPUs have 8 compute units, hence the name TPUV3-8, which could be interpreted as 8 GPUs working in parallel.<br>\nTry scaling the batch size by a factor 8 and the learning rate as well, you should get similar results.</p>",
      "rawMarkdown": "The Kaggle TPUs have 8 compute units, hence the name TPUV3-8, which could be interpreted as 8 GPUs working in parallel.\nTry scaling the batch size by a factor 8 and the learning rate as well, you should get similar results.",
      "votes": 4
    },
    {
      "id": 2148269,
      "postDate": "2023-02-17T09:45:47.993Z",
      "content": "<p>Good day.</p>\n<p>I have learned how to use TPU (TF) and used it recently for this competition.</p>\n<p>It work extremally faster but I can not receive even the same score which I got by GPU (for exactly the same dataset and resolution).</p>\n<p>The validation and LB scores are much lower in comparison to GPU training. </p>\n<p>Possible somebody got such issue. Does TPU training is the same as GPU in terms of backpropagation (or it needs some special knowledge)? </p>\n<p>I use the same fit function model was compiled in strategy.scope():</p>\n<p>model.fit(images, labels,shuffle = True, epochs=1, verbose=1, batch_size =batch_size, steps_per_epoch=n_batches, callbacks=callbacks_list,class_weight=class_weight)</p>\n<p>Speed is great but validation losses much higher than for GPU.</p>\n<p>Thanks.</p>\n<p>Igor.</p>",
      "rawMarkdown": "Good day.\n\nI have learned how to use TPU (TF) and used it recently for this competition.\n\nIt work extremally faster but I can not receive even the same score which I got by GPU (for exactly the same dataset and resolution).\n\nThe validation and LB scores are much lower in comparison to GPU training. \n\nPossible somebody got such issue. Does TPU training is the same as GPU in terms of backpropagation (or it needs some special knowledge)? \n\nI use the same fit function model was compiled in strategy.scope():\n\nmodel.fit(images, labels,shuffle = True, epochs=1, verbose=1, batch_size =batch_size, steps_per_epoch=n_batches, callbacks=callbacks_list,class_weight=class_weight)\n\nSpeed is great but validation losses much higher than for GPU.\n\nThanks.\n\nIgor.\n\n",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 2148289,
      "author_name": "Mark Wijkhuizen",
      "author_url": "",
      "post_date": "2023-02-17T10:16:19.617000",
      "content": "<p>The Kaggle TPUs have 8 compute units, hence the name TPUV3-8, which could be interpreted as 8 GPUs working in parallel.<br>\nTry scaling the batch size by a factor 8 and the learning rate as well, you should get similar results.</p>",
      "votes": 4,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2148289": "The Kaggle TPUs have 8 compute units, hence the name TPUV3-8, which could be interpreted as 8 GPUs working in parallel.\nTry scaling the batch size by a factor 8 and the learning rate as well, you should get similar results.",
    "2148269": "Good day.\n\nI have learned how to use TPU (TF) and used it recently for this competition.\n\nIt work extremally faster but I can not receive even the same score which I got by GPU (for exactly the same dataset and resolution).\n\nThe validation and LB scores are much lower in comparison to GPU training. \n\nPossible somebody got such issue. Does TPU training is the same as GPU in terms of backpropagation (or it needs some special knowledge)? \n\nI use the same fit function model was compiled in strategy.scope():\n\nmodel.fit(images, labels,shuffle = True, epochs=1, verbose=1, batch_size =batch_size, steps_per_epoch=n_batches, callbacks=callbacks_list,class_weight=class_weight)\n\nSpeed is great but validation losses much higher than for GPU.\n\nThanks.\n\nIgor.\n\n"
  }
}