{
  "id": 421923,
  "title": "Tensorflow InvalidArgumentError when training VisionTransformer with GPU, works with CPU.",
  "url": "/competitions/google-research-identify-contrails-reduce-global-warming/discussion/421923",
  "author_name": "Sashi",
  "post_date": "2023-07-07T13:58:06.689000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hello,<br>\nI am trying to incorporate a VisionTransformer and got the following error. The same model works with CPU.</p>\n<pre><code>All the layers of TFViTModel were initialized  the model checkpoint at google/vit-base-patch16--in21k.\n your  is similar to the  the model of the checkpoint was trained on, you can already use TFViTModel  predictions without further training.\n---------------------------------------------------------------------------\nInvalidArgumentError                      Traceback (most recent  last)\nCell In[], line \n         loss_value = loss_fn(labels, logits)\n     # Backward pass and optimization\n--&gt;  grads = tape.gradient(loss_value, model.trainable_variables)\n     optimizer.apply_gradients(zip(grads, model.trainable_variables))\n     #  training loss   epoch\n\n condapython3.tensorfloweager/backprop.py:, in GradientTape.gradient(self, target, sources, output_gradients, unconnected_gradients)\n      output_gradients = (\n          composite_tensor_gradient.get_flat_tensors_for_gradients(\n              output_gradients))\n      output_gradients = [None  x is None  ops.convert_to_tensor(x)\n                           x in output_gradients]\n-&gt;  flat_grad = imperative_grad.imperative_grad(\n        self._tape,\n        flat_targets,\n        flat_sources,\n        output_gradients=output_gradients,\n        sources_raw=flat_sources_raw,\n        unconnected_gradients=unconnected_gradients)\n     not self._persistent:\n      # Keep track of watched variables before setting tape to None\n      self._watched_variables = self._tape.watched_variables()\n\n condapython3.tensorfloweager/imperative_grad.py:, in imperative_grad(tape, target, sources, output_gradients, sources_raw, unconnected_gradients)\n      except ValueError:\n        raise ValueError(\n             % unconnected_gradients)\n---&gt;   pywrap_tfe.TFE_Py_TapeGradient(\n          tape._tape,  # pylint: disable=-access\n          target,\n          sources,\n          output_gradients,\n          sources_raw,\n          compat.as_str(unconnected_gradients.value))\n\n condapython3.tensorfloweager/backprop.py:, in _gradient_function(op_name, attr_tuple, num_inputs, inputs, outputs, out_grads, skip_input_indices, forward_pass_name_scope)\n         gradient_name_scope += forward_pass_name_scope + \n       with ops.name_scope(gradient_name_scope):\n--&gt;       grad_fn(mock_op, *out_grads)\n     :\n        grad_fn(mock_op, *out_grads)\n\n condapython3.tensorflowops/array_grad.py:, in _TransposeGrad(op, grad)\n     \n     p = op.inputs[]\n--&gt;   [array_ops.transpose(grad, array_ops.invert_permutation(p)), None]\n\n condapython3.tensorflowops/gen_array_ops.py:, in invert_permutation(x, name)\n       _result\n    except _core._NotOkStatusException as e:\n-&gt;    _ops.raise_from_not_ok_status(e, name)\n    except _core._FallbackException:\n      pass\n\n condapython3.tensorflowframework/ops.py:, in raise_from_not_ok_status(e, name)\n     raise_from_not_ok_status(e, name):\n      e.message += ( + name  name is not None  )\n-&gt;    raise core._status_to_exception(e)  None\n\nInvalidArgumentError: {{function_node __wrapped__InvertPermutation_device_replica:device:GPU:}} invert_permutation expects a D vector. [Op:InvertPermutation]\n</code></pre>\n<p>Any ideas? Thank you.</p>",
  "messages": [
    {
      "id": 2334193,
      "postDate": "2023-07-07T13:58:06.690Z",
      "content": "<p>Hello,<br>\nI am trying to incorporate a VisionTransformer and got the following error. The same model works with CPU.</p>\n<pre><code>All the layers of TFViTModel were initialized  the model checkpoint at google/vit-base-patch16--in21k.\n your  is similar to the  the model of the checkpoint was trained on, you can already use TFViTModel  predictions without further training.\n---------------------------------------------------------------------------\nInvalidArgumentError                      Traceback (most recent  last)\nCell In[], line \n         loss_value = loss_fn(labels, logits)\n     # Backward pass and optimization\n--&gt;  grads = tape.gradient(loss_value, model.trainable_variables)\n     optimizer.apply_gradients(zip(grads, model.trainable_variables))\n     #  training loss   epoch\n\n condapython3.tensorfloweager/backprop.py:, in GradientTape.gradient(self, target, sources, output_gradients, unconnected_gradients)\n      output_gradients = (\n          composite_tensor_gradient.get_flat_tensors_for_gradients(\n              output_gradients))\n      output_gradients = [None  x is None  ops.convert_to_tensor(x)\n                           x in output_gradients]\n-&gt;  flat_grad = imperative_grad.imperative_grad(\n        self._tape,\n        flat_targets,\n        flat_sources,\n        output_gradients=output_gradients,\n        sources_raw=flat_sources_raw,\n        unconnected_gradients=unconnected_gradients)\n     not self._persistent:\n      # Keep track of watched variables before setting tape to None\n      self._watched_variables = self._tape.watched_variables()\n\n condapython3.tensorfloweager/imperative_grad.py:, in imperative_grad(tape, target, sources, output_gradients, sources_raw, unconnected_gradients)\n      except ValueError:\n        raise ValueError(\n             % unconnected_gradients)\n---&gt;   pywrap_tfe.TFE_Py_TapeGradient(\n          tape._tape,  # pylint: disable=-access\n          target,\n          sources,\n          output_gradients,\n          sources_raw,\n          compat.as_str(unconnected_gradients.value))\n\n condapython3.tensorfloweager/backprop.py:, in _gradient_function(op_name, attr_tuple, num_inputs, inputs, outputs, out_grads, skip_input_indices, forward_pass_name_scope)\n         gradient_name_scope += forward_pass_name_scope + \n       with ops.name_scope(gradient_name_scope):\n--&gt;       grad_fn(mock_op, *out_grads)\n     :\n        grad_fn(mock_op, *out_grads)\n\n condapython3.tensorflowops/array_grad.py:, in _TransposeGrad(op, grad)\n     \n     p = op.inputs[]\n--&gt;   [array_ops.transpose(grad, array_ops.invert_permutation(p)), None]\n\n condapython3.tensorflowops/gen_array_ops.py:, in invert_permutation(x, name)\n       _result\n    except _core._NotOkStatusException as e:\n-&gt;    _ops.raise_from_not_ok_status(e, name)\n    except _core._FallbackException:\n      pass\n\n condapython3.tensorflowframework/ops.py:, in raise_from_not_ok_status(e, name)\n     raise_from_not_ok_status(e, name):\n      e.message += ( + name  name is not None  )\n-&gt;    raise core._status_to_exception(e)  None\n\nInvalidArgumentError: {{function_node __wrapped__InvertPermutation_device_replica:device:GPU:}} invert_permutation expects a D vector. [Op:InvertPermutation]\n</code></pre>\n<p>Any ideas? Thank you.</p>",
      "rawMarkdown": "Hello,\nI am trying to incorporate a VisionTransformer and got the following error. The same model works with CPU.\n\n```\nAll the layers of TFViTModel were initialized from the model checkpoint at google/vit-base-patch16-224-in21k.\nIf your task is similar to the task the model of the checkpoint was trained on, you can already use TFViTModel for predictions without further training.\n---------------------------------------------------------------------------\nInvalidArgumentError                      Traceback (most recent call last)\nCell In[1], line 146\n    143     loss_value = loss_fn(labels, logits)\n    145 # Backward pass and optimization\n--> 146 grads = tape.gradient(loss_value, model.trainable_variables)\n    147 optimizer.apply_gradients(zip(grads, model.trainable_variables))\n    149 # Print training loss for each epoch\n\nFile /opt/conda/lib/python3.10/site-packages/tensorflow/python/eager/backprop.py:1063, in GradientTape.gradient(self, target, sources, output_gradients, unconnected_gradients)\n   1057   output_gradients = (\n   1058       composite_tensor_gradient.get_flat_tensors_for_gradients(\n   1059           output_gradients))\n   1060   output_gradients = [None if x is None else ops.convert_to_tensor(x)\n   1061                       for x in output_gradients]\n-> 1063 flat_grad = imperative_grad.imperative_grad(\n   1064     self._tape,\n   1065     flat_targets,\n   1066     flat_sources,\n   1067     output_gradients=output_gradients,\n   1068     sources_raw=flat_sources_raw,\n   1069     unconnected_gradients=unconnected_gradients)\n   1071 if not self._persistent:\n   1072   # Keep track of watched variables before setting tape to None\n   1073   self._watched_variables = self._tape.watched_variables()\n\nFile /opt/conda/lib/python3.10/site-packages/tensorflow/python/eager/imperative_grad.py:67, in imperative_grad(tape, target, sources, output_gradients, sources_raw, unconnected_gradients)\n     63 except ValueError:\n     64   raise ValueError(\n     65       \"Unknown value for unconnected_gradients: %r\" % unconnected_gradients)\n---> 67 return pywrap_tfe.TFE_Py_TapeGradient(\n     68     tape._tape,  # pylint: disable=protected-access\n     69     target,\n     70     sources,\n     71     output_gradients,\n     72     sources_raw,\n     73     compat.as_str(unconnected_gradients.value))\n\nFile /opt/conda/lib/python3.10/site-packages/tensorflow/python/eager/backprop.py:146, in _gradient_function(op_name, attr_tuple, num_inputs, inputs, outputs, out_grads, skip_input_indices, forward_pass_name_scope)\n    144     gradient_name_scope += forward_pass_name_scope + \"/\"\n    145   with ops.name_scope(gradient_name_scope):\n--> 146     return grad_fn(mock_op, *out_grads)\n    147 else:\n    148   return grad_fn(mock_op, *out_grads)\n\nFile /opt/conda/lib/python3.10/site-packages/tensorflow/python/ops/array_grad.py:832, in _TransposeGrad(op, grad)\n    830 \"\"\"Returns unshuffle(grad).\"\"\"\n    831 p = op.inputs[1]\n--> 832 return [array_ops.transpose(grad, array_ops.invert_permutation(p)), None]\n\nFile /opt/conda/lib/python3.10/site-packages/tensorflow/python/ops/gen_array_ops.py:4527, in invert_permutation(x, name)\n   4525   return _result\n   4526 except _core._NotOkStatusException as e:\n-> 4527   _ops.raise_from_not_ok_status(e, name)\n   4528 except _core._FallbackException:\n   4529   pass\n\nFile /opt/conda/lib/python3.10/site-packages/tensorflow/python/framework/ops.py:7262, in raise_from_not_ok_status(e, name)\n   7260 def raise_from_not_ok_status(e, name):\n   7261   e.message += (\" name: \" + name if name is not None else \"\")\n-> 7262   raise core._status_to_exception(e) from None\n\nInvalidArgumentError: {{function_node __wrapped__InvertPermutation_device_/job:localhost/replica:0/task:0/device:GPU:0}} invert_permutation expects a 1D vector. [Op:InvertPermutation]\n```\n\nAny ideas? Thank you."
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2334193": "Hello,\nI am trying to incorporate a VisionTransformer and got the following error. The same model works with CPU.\n\n```\nAll the layers of TFViTModel were initialized from the model checkpoint at google/vit-base-patch16-224-in21k.\nIf your task is similar to the task the model of the checkpoint was trained on, you can already use TFViTModel for predictions without further training.\n---------------------------------------------------------------------------\nInvalidArgumentError                      Traceback (most recent call last)\nCell In[1], line 146\n    143     loss_value = loss_fn(labels, logits)\n    145 # Backward pass and optimization\n--> 146 grads = tape.gradient(loss_value, model.trainable_variables)\n    147 optimizer.apply_gradients(zip(grads, model.trainable_variables))\n    149 # Print training loss for each epoch\n\nFile /opt/conda/lib/python3.10/site-packages/tensorflow/python/eager/backprop.py:1063, in GradientTape.gradient(self, target, sources, output_gradients, unconnected_gradients)\n   1057   output_gradients = (\n   1058       composite_tensor_gradient.get_flat_tensors_for_gradients(\n   1059           output_gradients))\n   1060   output_gradients = [None if x is None else ops.convert_to_tensor(x)\n   1061                       for x in output_gradients]\n-> 1063 flat_grad = imperative_grad.imperative_grad(\n   1064     self._tape,\n   1065     flat_targets,\n   1066     flat_sources,\n   1067     output_gradients=output_gradients,\n   1068     sources_raw=flat_sources_raw,\n   1069     unconnected_gradients=unconnected_gradients)\n   1071 if not self._persistent:\n   1072   # Keep track of watched variables before setting tape to None\n   1073   self._watched_variables = self._tape.watched_variables()\n\nFile /opt/conda/lib/python3.10/site-packages/tensorflow/python/eager/imperative_grad.py:67, in imperative_grad(tape, target, sources, output_gradients, sources_raw, unconnected_gradients)\n     63 except ValueError:\n     64   raise ValueError(\n     65       \"Unknown value for unconnected_gradients: %r\" % unconnected_gradients)\n---> 67 return pywrap_tfe.TFE_Py_TapeGradient(\n     68     tape._tape,  # pylint: disable=protected-access\n     69     target,\n     70     sources,\n     71     output_gradients,\n     72     sources_raw,\n     73     compat.as_str(unconnected_gradients.value))\n\nFile /opt/conda/lib/python3.10/site-packages/tensorflow/python/eager/backprop.py:146, in _gradient_function(op_name, attr_tuple, num_inputs, inputs, outputs, out_grads, skip_input_indices, forward_pass_name_scope)\n    144     gradient_name_scope += forward_pass_name_scope + \"/\"\n    145   with ops.name_scope(gradient_name_scope):\n--> 146     return grad_fn(mock_op, *out_grads)\n    147 else:\n    148   return grad_fn(mock_op, *out_grads)\n\nFile /opt/conda/lib/python3.10/site-packages/tensorflow/python/ops/array_grad.py:832, in _TransposeGrad(op, grad)\n    830 \"\"\"Returns unshuffle(grad).\"\"\"\n    831 p = op.inputs[1]\n--> 832 return [array_ops.transpose(grad, array_ops.invert_permutation(p)), None]\n\nFile /opt/conda/lib/python3.10/site-packages/tensorflow/python/ops/gen_array_ops.py:4527, in invert_permutation(x, name)\n   4525   return _result\n   4526 except _core._NotOkStatusException as e:\n-> 4527   _ops.raise_from_not_ok_status(e, name)\n   4528 except _core._FallbackException:\n   4529   pass\n\nFile /opt/conda/lib/python3.10/site-packages/tensorflow/python/framework/ops.py:7262, in raise_from_not_ok_status(e, name)\n   7260 def raise_from_not_ok_status(e, name):\n   7261   e.message += (\" name: \" + name if name is not None else \"\")\n-> 7262   raise core._status_to_exception(e) from None\n\nInvalidArgumentError: {{function_node __wrapped__InvertPermutation_device_/job:localhost/replica:0/task:0/device:GPU:0}} invert_permutation expects a 1D vector. [Op:InvertPermutation]\n```\n\nAny ideas? Thank you."
  }
}