{
  "id": 512787,
  "title": "TPU's accuracy is lower than GPU?",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/512787",
  "author_name": "xsong2020",
  "post_date": "2024-06-17T01:05:10.292000",
  "votes": 3,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Hi,</p>\n<p>I try to train the same model with the exactly same data and strategy (e.g. same epochs, same scheduler, etc…) on both TPU and GPU. And find that TPU's CV and LB are both worse than GPU results. <br>\nMay I ask if anyone also met the similar situation?</p>",
  "messages": [
    {
      "id": 2879996,
      "postDate": "2024-06-20T01:05:31.617Z",
      "content": "<p>Been my observation over lots of time that change in batch size leads to different results - for sure when I run TPU I make a large batch size vs gpu batch - any chance you seeing batch size change difference?</p>",
      "rawMarkdown": "Been my observation over lots of time that change in batch size leads to different results - for sure when I run TPU I make a large batch size vs gpu batch - any chance you seeing batch size change difference?",
      "votes": 1
    },
    {
      "id": 2875241,
      "postDate": "2024-06-17T01:05:10.293Z",
      "content": "<p>Hi,</p>\n<p>I try to train the same model with the exactly same data and strategy (e.g. same epochs, same scheduler, etc…) on both TPU and GPU. And find that TPU's CV and LB are both worse than GPU results. <br>\nMay I ask if anyone also met the similar situation?</p>",
      "rawMarkdown": "Hi,\n\nI try to train the same model with the exactly same data and strategy (e.g. same epochs, same scheduler, etc...) on both TPU and GPU. And find that TPU's CV and LB are both worse than GPU results. \nMay I ask if anyone also met the similar situation?",
      "votes": 2
    },
    {
      "id": 2881210,
      "postDate": "2024-06-20T15:25:55.500Z",
      "content": "<p>From some experiments of mine, distributed training (on multi-cpu, but should be the same on TPUs) using tensorflow backend for keras3 is bugged: the copies on different devices are not updated as often as they should (or at all) so you need REEAALLY big batches to make everything work. JAX backend seems to work fine and you get almost identical results independently of the number of devices you use for training.</p>\n<p>I will investigate how to use the JAX backend on TPU, if it is possible at all</p>",
      "rawMarkdown": "From some experiments of mine, distributed training (on multi-cpu, but should be the same on TPUs) using tensorflow backend for keras3 is bugged: the copies on different devices are not updated as often as they should (or at all) so you need REEAALLY big batches to make everything work. JAX backend seems to work fine and you get almost identical results independently of the number of devices you use for training.\n\nI will investigate how to use the JAX backend on TPU, if it is possible at all",
      "replies": [
        {
          "id": 2881368,
          "postDate": "2024-06-20T16:58:00.563Z",
          "content": "<p>Easier solution then: just downgrade to keras 2. I do it anyway since keras 3 has some other bugs/changes that makes life harder. It's really bad right now. I recommend using Keras 2.15 or whatever is compatible with your TF version.</p>",
          "rawMarkdown": "Easier solution then: just downgrade to keras 2. I do it anyway since keras 3 has some other bugs/changes that makes life harder. It's really bad right now. I recommend using Keras 2.15 or whatever is compatible with your TF version."
        },
        {
          "id": 2881610,
          "postDate": "2024-06-20T20:55:22.677Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        }
      ]
    },
    {
      "id": 2876703,
      "postDate": "2024-06-18T00:16:27.020Z",
      "content": "<p>This is very interesting… I haven't used the TPU yet, so far my models are small enough that GPU is enough. Are you using Keras?  From comments below, it seems the problem has to do with fp16/bf16 matrix multiplication and not having a large enough model. Maybe setting the dtype policy to float32 would help. </p>",
      "rawMarkdown": "This is very interesting... I haven't used the TPU yet, so far my models are small enough that GPU is enough. Are you using Keras?  From comments below, it seems the problem has to do with fp16/bf16 matrix multiplication and not having a large enough model. Maybe setting the dtype policy to float32 would help. ",
      "replies": [
        {
          "id": 2876704,
          "postDate": "2024-06-18T00:16:42.040Z",
          "content": "<p><a href=\"https://keras.io/api/mixed_precision/policy/#dtypepolicy-class\" target=\"_blank\">https://keras.io/api/mixed_precision/policy/#dtypepolicy-class</a></p>",
          "rawMarkdown": "https://keras.io/api/mixed_precision/policy/#dtypepolicy-class"
        }
      ]
    },
    {
      "id": 2876212,
      "postDate": "2024-06-17T15:53:08.007Z",
      "content": "<p>I have the same issue even using bfloat16 on GPU the results on TPU are 2% worse. Will investigate more into the details in the next few days</p>",
      "rawMarkdown": "I have the same issue even using bfloat16 on GPU the results on TPU are 2% worse. Will investigate more into the details in the next few days",
      "replies": [
        {
          "id": 2876303,
          "postDate": "2024-06-17T17:24:58.343Z",
          "content": "<p>2% is a <em>lot</em>! Do you mind sharing how much parameters is said model? Also, if you use a multihead attention (usually as part of a transformer), do you use tensorflow's implementation or a custom one? I remember I had a significant drop in accuracy with a custom implementation once that was solved by switching to the official implementation</p>",
          "rawMarkdown": "2% is a *lot*! Do you mind sharing how much parameters is said model? Also, if you use a multihead attention (usually as part of a transformer), do you use tensorflow's implementation or a custom one? I remember I had a significant drop in accuracy with a custom implementation once that was solved by switching to the official implementation",
          "replies": [
            {
              "id": 2877103,
              "postDate": "2024-06-18T07:13:24.410Z",
              "content": "<p>I tested models from 2M to 20M parameters, both CNN and Tranformer (custom transformer but keras MHAttention). Even when using the exact same parameters (including batch_size and precisione policy) I get a 2% drop even after 10k updates. I have a few ideas (different gradient aggregation, different details in the mixed_bfloat16 policy, error in the dataset sharding) but now solution yet</p>",
              "rawMarkdown": "I tested models from 2M to 20M parameters, both CNN and Tranformer (custom transformer but keras MHAttention). Even when using the exact same parameters (including batch_size and precisione policy) I get a 2% drop even after 10k updates. I have a few ideas (different gradient aggregation, different details in the mixed_bfloat16 policy, error in the dataset sharding) but now solution yet",
              "votes": 1
            }
          ]
        },
        {
          "id": 2876309,
          "postDate": "2024-06-17T17:40:37.747Z",
          "content": "<p>Yes, my TPU LB score is also about 2.2% worse than my GPU LB score.</p>",
          "rawMarkdown": "Yes, my TPU LB score is also about 2.2% worse than my GPU LB score."
        }
      ]
    },
    {
      "id": 2875405,
      "postDate": "2024-06-17T05:56:38.757Z",
      "content": "<p>How much worse? A small drop in accuracy is OK since TPU use bfloat16 for matrix multiplication.  However for large enough models the difference is very small and basically you overcome it by making your TPU model larger than what you would be able to train on GPU. If there is a large difference something is probably wrong…</p>",
      "rawMarkdown": "How much worse? A small drop in accuracy is OK since TPU use bfloat16 for matrix multiplication.  However for large enough models the difference is very small and basically you overcome it by making your TPU model larger than what you would be able to train on GPU. If there is a large difference something is probably wrong...",
      "replies": [
        {
          "id": 2875492,
          "postDate": "2024-06-17T07:23:31.657Z",
          "content": "<p>Also, for the same reason, inference on GPU would have slightly better accuracy, even for a model trained on TPU. In this competition, to put things in context, 'slightly better accuracy' is ~0.0005</p>",
          "rawMarkdown": "Also, for the same reason, inference on GPU would have slightly better accuracy, even for a model trained on TPU. In this competition, to put things in context, 'slightly better accuracy' is ~0.0005",
          "votes": 3
        }
      ]
    },
    {
      "id": 2875968,
      "postDate": "2024-06-17T13:14:13.370Z",
      "content": "<p>Same here regarding the TPU.<br>\nAlso, running with GPU gets different results than with CPU. Not very different, but have to tune some parameters differently to get almost same results as with CPU. This is not possible with TPU (for me).</p>",
      "rawMarkdown": "Same here regarding the TPU.\nAlso, running with GPU gets different results than with CPU. Not very different, but have to tune some parameters differently to get almost same results as with CPU. This is not possible with TPU (for me).",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2879996,
      "author_name": "PC Jimmmy",
      "author_url": "",
      "post_date": "2024-06-20T01:05:31.617000",
      "content": "<p>Been my observation over lots of time that change in batch size leads to different results - for sure when I run TPU I make a large batch size vs gpu batch - any chance you seeing batch size change difference?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2881210,
      "author_name": "Amedeo Biolatti",
      "author_url": "",
      "post_date": "2024-06-20T15:25:55.500000",
      "content": "<p>From some experiments of mine, distributed training (on multi-cpu, but should be the same on TPUs) using tensorflow backend for keras3 is bugged: the copies on different devices are not updated as often as they should (or at all) so you need REEAALLY big batches to make everything work. JAX backend seems to work fine and you get almost identical results independently of the number of devices you use for training.</p>\n<p>I will investigate how to use the JAX backend on TPU, if it is possible at all</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2881368,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2024-06-20T16:58:00.563000",
          "content": "<p>Easier solution then: just downgrade to keras 2. I do it anyway since keras 3 has some other bugs/changes that makes life harder. It's really bad right now. I recommend using Keras 2.15 or whatever is compatible with your TF version.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2881610,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-06-20T20:55:22.677000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2876703,
      "author_name": "Juan D C F",
      "author_url": "",
      "post_date": "2024-06-18T00:16:27.020000",
      "content": "<p>This is very interesting… I haven't used the TPU yet, so far my models are small enough that GPU is enough. Are you using Keras?  From comments below, it seems the problem has to do with fp16/bf16 matrix multiplication and not having a large enough model. Maybe setting the dtype policy to float32 would help. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2876704,
          "author_name": "Juan D C F",
          "author_url": "",
          "post_date": "2024-06-18T00:16:42.040000",
          "content": "<p><a href=\"https://keras.io/api/mixed_precision/policy/#dtypepolicy-class\" target=\"_blank\">https://keras.io/api/mixed_precision/policy/#dtypepolicy-class</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2876212,
      "author_name": "Amedeo Biolatti",
      "author_url": "",
      "post_date": "2024-06-17T15:53:08.007000",
      "content": "<p>I have the same issue even using bfloat16 on GPU the results on TPU are 2% worse. Will investigate more into the details in the next few days</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2876303,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2024-06-17T17:24:58.343000",
          "content": "<p>2% is a <em>lot</em>! Do you mind sharing how much parameters is said model? Also, if you use a multihead attention (usually as part of a transformer), do you use tensorflow's implementation or a custom one? I remember I had a significant drop in accuracy with a custom implementation once that was solved by switching to the official implementation</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2877103,
              "author_name": "Amedeo Biolatti",
              "author_url": "",
              "post_date": "2024-06-18T07:13:24.410000",
              "content": "<p>I tested models from 2M to 20M parameters, both CNN and Tranformer (custom transformer but keras MHAttention). Even when using the exact same parameters (including batch_size and precisione policy) I get a 2% drop even after 10k updates. I have a few ideas (different gradient aggregation, different details in the mixed_bfloat16 policy, error in the dataset sharding) but now solution yet</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 2876309,
          "author_name": "xsong2020",
          "author_url": "",
          "post_date": "2024-06-17T17:40:37.747000",
          "content": "<p>Yes, my TPU LB score is also about 2.2% worse than my GPU LB score.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2875405,
      "author_name": "greySnow",
      "author_url": "",
      "post_date": "2024-06-17T05:56:38.757000",
      "content": "<p>How much worse? A small drop in accuracy is OK since TPU use bfloat16 for matrix multiplication.  However for large enough models the difference is very small and basically you overcome it by making your TPU model larger than what you would be able to train on GPU. If there is a large difference something is probably wrong…</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2875492,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2024-06-17T07:23:31.657000",
          "content": "<p>Also, for the same reason, inference on GPU would have slightly better accuracy, even for a model trained on TPU. In this competition, to put things in context, 'slightly better accuracy' is ~0.0005</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2875968,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-06-17T13:14:13.370000",
      "content": "<p>Same here regarding the TPU.<br>\nAlso, running with GPU gets different results than with CPU. Not very different, but have to tune some parameters differently to get almost same results as with CPU. This is not possible with TPU (for me).</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2879996": "Been my observation over lots of time that change in batch size leads to different results - for sure when I run TPU I make a large batch size vs gpu batch - any chance you seeing batch size change difference?",
    "2875241": "Hi,\n\nI try to train the same model with the exactly same data and strategy (e.g. same epochs, same scheduler, etc...) on both TPU and GPU. And find that TPU's CV and LB are both worse than GPU results. \nMay I ask if anyone also met the similar situation?",
    "2881210": "From some experiments of mine, distributed training (on multi-cpu, but should be the same on TPUs) using tensorflow backend for keras3 is bugged: the copies on different devices are not updated as often as they should (or at all) so you need REEAALLY big batches to make everything work. JAX backend seems to work fine and you get almost identical results independently of the number of devices you use for training.\n\nI will investigate how to use the JAX backend on TPU, if it is possible at all",
    "2876703": "This is very interesting... I haven't used the TPU yet, so far my models are small enough that GPU is enough. Are you using Keras?  From comments below, it seems the problem has to do with fp16/bf16 matrix multiplication and not having a large enough model. Maybe setting the dtype policy to float32 would help. ",
    "2876212": "I have the same issue even using bfloat16 on GPU the results on TPU are 2% worse. Will investigate more into the details in the next few days",
    "2875405": "How much worse? A small drop in accuracy is OK since TPU use bfloat16 for matrix multiplication.  However for large enough models the difference is very small and basically you overcome it by making your TPU model larger than what you would be able to train on GPU. If there is a large difference something is probably wrong...",
    "2875968": "Same here regarding the TPU.\nAlso, running with GPU gets different results than with CPU. Not very different, but have to tune some parameters differently to get almost same results as with CPU. This is not possible with TPU (for me)."
  }
}