{
  "id": 458151,
  "title": "Problem with xFormers installation via utility script",
  "url": "/competitions/UBC-OCEAN/discussion/458151",
  "author_name": "Patrick Robitaille",
  "post_date": "2023-11-28T14:08:49.486000",
  "votes": 3,
  "comment_count": 7,
  "views": 0,
  "content": "<p>As per normal Kaggle practice when you need packages not available in Kaggle's regular environment, I tried to set up a notebook containing the packages I need for this competition in order to use that notebook in my submission notebook as a utility script. However, one package is giving me some grief: xFormers.</p>\n<p>As recommended in their Github repo and in line with the specs of the P100 GPU, I have installed xFormers using the following pip command: <code>pip3 install -U xformers --index-url https://download.pytorch.org/whl/cu118</code>. However, when I try to import any component from xFormers in my submission notebook, I get hit with a Errno 13 message as xFormers looks at importing Triton components (I will only show the relevant part of the error message):</p>\n<pre><code>File /kaggle/usr/lib/ubc_ocean_packages/triton/language/math.py:\n        functools\n        os\n----&gt;   .  core\n       @functools.lru_cache()\n        ():\n            torch\n\nFile /kaggle/usr/lib/ubc_ocean_packages/triton/language/core.py:\n        rvalue, rindices = reduce((, index), axis, combine_fn,\n                                  _builder=_builder, _generator=_generator)\n         rvalue, rindices\n    @jit\n-&gt;   ():\n        \n         where(x &lt; y, x, y)\n\nFile /kaggle/usr/lib/ubc_ocean_packages/triton/runtime/jit.py:,  jit(fn, version, do_not_specialize, debug, noinline, interpret)\n              JITFunction(\n                 fn,\n                 version=version,\n   (...)\n                 noinline=noinline,\n             )\n      fn   :\n--&gt;       decorator(fn)\n     :\n          decorator\n\nFile /kaggle/usr/lib/ubc_ocean_packages/triton/runtime/jit.py:,  jit.&lt;&gt;.decorator(fn)\n          GridSelector(fn)\n     :\n--&gt;       JITFunction(\n             fn,\n             version=version,\n             do_not_specialize=do_not_specialize,\n             debug=debug,\n             noinline=noinline,\n         )\n\nFile /kaggle/usr/lib/ubc_ocean_packages/triton/runtime/jit.py:,  JITFunction.__init__(self, fn, version, do_not_specialize, debug, noinline)\n     self.constexprs = [self.arg_names.index(name)  name, ty  self.__annotations__.items()    ty]\n     \n--&gt;  self.run = self._make_launcher()\n     \n     self.__doc__ = fn.__doc__\n\nFile /kaggle/usr/lib/ubc_ocean_packages/triton/runtime/jit.py:,  JITFunction._make_launcher(self)\n             args_signature = .join(name  dflt == inspect._empty    name, dflt  (self.arg_names, self.arg_defaults))\n             src = \n--&gt;          scope = {: version_key(),\n                      : get_cuda_stream,\n                      : self,\n                      : self._spec_of,\n                      : self._key_of,\n                      : self._device_of,\n                      : self._pinned_memory_of,\n                      : self.cache,\n                      : __spec__,\n                      : get_backend,\n                      : get_current_device,\n                      : set_current_device}\n             (src, scope)\n              scope[self.fn.__name__]\n\nFile /kaggle/usr/lib/ubc_ocean_packages/triton/runtime/jit.py:,  version_key()\n             contents += [hashlib.md5(f.read()).hexdigest()]\n     \n--&gt;  ptxas = path_to_ptxas()[]\n     ptxas_version = hashlib.md5(subprocess.check_output([ptxas, ])).hexdigest()\n      .join(TRITON_VERSION) +  + ptxas_version +  + .join(contents)\n\nFile /kaggle/usr/lib/ubc_ocean_packages/triton/common/backend.py:,  path_to_ptxas()\n     ptxas_bin = ptxas.split()[]\n      os.path.exists(ptxas_bin)  os.path.isfile(ptxas_bin):\n--&gt;      result = subprocess.check_output([ptxas_bin, ], stderr=subprocess.STDOUT)\n          result   :\n             version = re.search(, result.decode(), flags=re.MULTILINE)\n\nFile /opt/conda/lib/python3/subprocess.py:,  check_output(timeout, *popenargs, **kwargs)\n             empty = \n         kwargs[] = empty\n--&gt;   run(*popenargs, stdout=PIPE, timeout=timeout, check=,\n                **kwargs).stdout\n\nFile /opt/conda/lib/python3/subprocess.py:,  run(, capture_output, timeout, check, *popenargs, **kwargs)\n         kwargs[] = PIPE\n         kwargs[] = PIPE\n--&gt;   Popen(*popenargs, **kwargs)  process:\n         :\n             stdout, stderr = process.communicate(, timeout=timeout)\n\nFile /opt/conda/lib/python3/subprocess.py:,  Popen.__init__(self, args, bufsize, executable, stdin, stdout, stderr, preexec_fn, close_fds, shell, cwd, env, universal_newlines, startupinfo, creationflags, restore_signals, start_new_session, pass_fds, user, group, extra_groups, encoding, errors, text, umask, pipesize)\n              self.text_mode:\n                 self.stderr = io.TextIOWrapper(self.stderr,\n                         encoding=encoding, errors=errors)\n--&gt;      self._execute_child(args, executable, preexec_fn, close_fds,\n                             pass_fds, cwd, env,\n                             startupinfo, creationflags, shell,\n                             p2cread, p2cwrite,\n                             c2pread, c2pwrite,\n                             errread, errwrite,\n                             restore_signals,\n                             gid, gids, uid, umask,\n                             start_new_session)\n     :\n         \n          f  (, (self.stdin, self.stdout, self.stderr)):\n\nFile /opt/conda/lib/python3/subprocess.py:,  Popen._execute_child(self, args, executable, preexec_fn, close_fds, pass_fds, cwd, env, startupinfo, creationflags, shell, p2cread, p2cwrite, c2pread, c2pwrite, errread, errwrite, restore_signals, gid, gids, uid, umask, start_new_session)\n         errno_num != :\n            err_msg = os.strerror(errno_num)\n-&gt;       child_exception_type(errno_num, err_msg, err_filename)\n     child_exception_type(err_msg)\n\nPermissionError: [Errno ] Permission denied: \n</code></pre>\n<p>I tried to find the location of that file/folder, but to no avail; I was hoping I could perhaps change the permissions for that file/folder to fix the problem. Has anybody had any issues installing xFormers in an auxilliary notebook? Does anyone have any guidance on how to fix this problem?</p>",
  "messages": [
    {
      "id": 2541466,
      "postDate": "2023-11-28T14:08:49.487Z",
      "content": "<p>As per normal Kaggle practice when you need packages not available in Kaggle's regular environment, I tried to set up a notebook containing the packages I need for this competition in order to use that notebook in my submission notebook as a utility script. However, one package is giving me some grief: xFormers.</p>\n<p>As recommended in their Github repo and in line with the specs of the P100 GPU, I have installed xFormers using the following pip command: <code>pip3 install -U xformers --index-url https://download.pytorch.org/whl/cu118</code>. However, when I try to import any component from xFormers in my submission notebook, I get hit with a Errno 13 message as xFormers looks at importing Triton components (I will only show the relevant part of the error message):</p>\n<pre><code>File /kaggle/usr/lib/ubc_ocean_packages/triton/language/math.py:\n        functools\n        os\n----&gt;   .  core\n       @functools.lru_cache()\n        ():\n            torch\n\nFile /kaggle/usr/lib/ubc_ocean_packages/triton/language/core.py:\n        rvalue, rindices = reduce((, index), axis, combine_fn,\n                                  _builder=_builder, _generator=_generator)\n         rvalue, rindices\n    @jit\n-&gt;   ():\n        \n         where(x &lt; y, x, y)\n\nFile /kaggle/usr/lib/ubc_ocean_packages/triton/runtime/jit.py:,  jit(fn, version, do_not_specialize, debug, noinline, interpret)\n              JITFunction(\n                 fn,\n                 version=version,\n   (...)\n                 noinline=noinline,\n             )\n      fn   :\n--&gt;       decorator(fn)\n     :\n          decorator\n\nFile /kaggle/usr/lib/ubc_ocean_packages/triton/runtime/jit.py:,  jit.&lt;&gt;.decorator(fn)\n          GridSelector(fn)\n     :\n--&gt;       JITFunction(\n             fn,\n             version=version,\n             do_not_specialize=do_not_specialize,\n             debug=debug,\n             noinline=noinline,\n         )\n\nFile /kaggle/usr/lib/ubc_ocean_packages/triton/runtime/jit.py:,  JITFunction.__init__(self, fn, version, do_not_specialize, debug, noinline)\n     self.constexprs = [self.arg_names.index(name)  name, ty  self.__annotations__.items()    ty]\n     \n--&gt;  self.run = self._make_launcher()\n     \n     self.__doc__ = fn.__doc__\n\nFile /kaggle/usr/lib/ubc_ocean_packages/triton/runtime/jit.py:,  JITFunction._make_launcher(self)\n             args_signature = .join(name  dflt == inspect._empty    name, dflt  (self.arg_names, self.arg_defaults))\n             src = \n--&gt;          scope = {: version_key(),\n                      : get_cuda_stream,\n                      : self,\n                      : self._spec_of,\n                      : self._key_of,\n                      : self._device_of,\n                      : self._pinned_memory_of,\n                      : self.cache,\n                      : __spec__,\n                      : get_backend,\n                      : get_current_device,\n                      : set_current_device}\n             (src, scope)\n              scope[self.fn.__name__]\n\nFile /kaggle/usr/lib/ubc_ocean_packages/triton/runtime/jit.py:,  version_key()\n             contents += [hashlib.md5(f.read()).hexdigest()]\n     \n--&gt;  ptxas = path_to_ptxas()[]\n     ptxas_version = hashlib.md5(subprocess.check_output([ptxas, ])).hexdigest()\n      .join(TRITON_VERSION) +  + ptxas_version +  + .join(contents)\n\nFile /kaggle/usr/lib/ubc_ocean_packages/triton/common/backend.py:,  path_to_ptxas()\n     ptxas_bin = ptxas.split()[]\n      os.path.exists(ptxas_bin)  os.path.isfile(ptxas_bin):\n--&gt;      result = subprocess.check_output([ptxas_bin, ], stderr=subprocess.STDOUT)\n          result   :\n             version = re.search(, result.decode(), flags=re.MULTILINE)\n\nFile /opt/conda/lib/python3/subprocess.py:,  check_output(timeout, *popenargs, **kwargs)\n             empty = \n         kwargs[] = empty\n--&gt;   run(*popenargs, stdout=PIPE, timeout=timeout, check=,\n                **kwargs).stdout\n\nFile /opt/conda/lib/python3/subprocess.py:,  run(, capture_output, timeout, check, *popenargs, **kwargs)\n         kwargs[] = PIPE\n         kwargs[] = PIPE\n--&gt;   Popen(*popenargs, **kwargs)  process:\n         :\n             stdout, stderr = process.communicate(, timeout=timeout)\n\nFile /opt/conda/lib/python3/subprocess.py:,  Popen.__init__(self, args, bufsize, executable, stdin, stdout, stderr, preexec_fn, close_fds, shell, cwd, env, universal_newlines, startupinfo, creationflags, restore_signals, start_new_session, pass_fds, user, group, extra_groups, encoding, errors, text, umask, pipesize)\n              self.text_mode:\n                 self.stderr = io.TextIOWrapper(self.stderr,\n                         encoding=encoding, errors=errors)\n--&gt;      self._execute_child(args, executable, preexec_fn, close_fds,\n                             pass_fds, cwd, env,\n                             startupinfo, creationflags, shell,\n                             p2cread, p2cwrite,\n                             c2pread, c2pwrite,\n                             errread, errwrite,\n                             restore_signals,\n                             gid, gids, uid, umask,\n                             start_new_session)\n     :\n         \n          f  (, (self.stdin, self.stdout, self.stderr)):\n\nFile /opt/conda/lib/python3/subprocess.py:,  Popen._execute_child(self, args, executable, preexec_fn, close_fds, pass_fds, cwd, env, startupinfo, creationflags, shell, p2cread, p2cwrite, c2pread, c2pwrite, errread, errwrite, restore_signals, gid, gids, uid, umask, start_new_session)\n         errno_num != :\n            err_msg = os.strerror(errno_num)\n-&gt;       child_exception_type(errno_num, err_msg, err_filename)\n     child_exception_type(err_msg)\n\nPermissionError: [Errno ] Permission denied: \n</code></pre>\n<p>I tried to find the location of that file/folder, but to no avail; I was hoping I could perhaps change the permissions for that file/folder to fix the problem. Has anybody had any issues installing xFormers in an auxilliary notebook? Does anyone have any guidance on how to fix this problem?</p>",
      "rawMarkdown": "As per normal Kaggle practice when you need packages not available in Kaggle's regular environment, I tried to set up a notebook containing the packages I need for this competition in order to use that notebook in my submission notebook as a utility script. However, one package is giving me some grief: xFormers.\n\nAs recommended in their Github repo and in line with the specs of the P100 GPU, I have installed xFormers using the following pip command: `pip3 install -U xformers --index-url https://download.pytorch.org/whl/cu118 `. However, when I try to import any component from xFormers in my submission notebook, I get hit with a Errno 13 message as xFormers looks at importing Triton components (I will only show the relevant part of the error message):\n\n```python\nFile /kaggle/usr/lib/ubc_ocean_packages/triton/language/math.py:4\n      1 import functools\n      2 import os\n----> 4 from . import core\n      7 @functools.lru_cache()\n      8 def libdevice_path():\n      9     import torch\n\nFile /kaggle/usr/lib/ubc_ocean_packages/triton/language/core.py:1376\n   1370     rvalue, rindices = reduce((input, index), axis, combine_fn,\n   1371                               _builder=_builder, _generator=_generator)\n   1372     return rvalue, rindices\n   1375 @jit\n-> 1376 def minimum(x, y):\n   1377     \"\"\"\n   1378     Computes the element-wise minimum of :code:`x` and :code:`y`.\n   1379 \n   (...)\n   1383     :type other: Block\n   1384     \"\"\"\n   1385     return where(x < y, x, y)\n\nFile /kaggle/usr/lib/ubc_ocean_packages/triton/runtime/jit.py:542, in jit(fn, version, do_not_specialize, debug, noinline, interpret)\n    534         return JITFunction(\n    535             fn,\n    536             version=version,\n   (...)\n    539             noinline=noinline,\n    540         )\n    541 if fn is not None:\n--> 542     return decorator(fn)\n    544 else:\n    545     return decorator\n\nFile /kaggle/usr/lib/ubc_ocean_packages/triton/runtime/jit.py:534, in jit.<locals>.decorator(fn)\n    532     return GridSelector(fn)\n    533 else:\n--> 534     return JITFunction(\n    535         fn,\n    536         version=version,\n    537         do_not_specialize=do_not_specialize,\n    538         debug=debug,\n    539         noinline=noinline,\n    540     )\n\nFile /kaggle/usr/lib/ubc_ocean_packages/triton/runtime/jit.py:433, in JITFunction.__init__(self, fn, version, do_not_specialize, debug, noinline)\n    431 self.constexprs = [self.arg_names.index(name) for name, ty in self.__annotations__.items() if 'constexpr' in ty]\n    432 # launcher\n--> 433 self.run = self._make_launcher()\n    434 # re-use docs of wrapped function\n    435 self.__doc__ = fn.__doc__\n\nFile /kaggle/usr/lib/ubc_ocean_packages/triton/runtime/jit.py:388, in JITFunction._make_launcher(self)\n    317         args_signature = ', '.join(name if dflt == inspect._empty else f'{name} = {dflt}' for name, dflt in zip(self.arg_names, self.arg_defaults))\n    319         src = f\"\"\"\n    320 def {self.fn.__name__}({args_signature}, grid=None, num_warps=4, num_stages=3, extern_libs=None, stream=None, warmup=False, device=None, device_type=None):\n    321     from ..compiler import compile, CompiledKernel\n   (...)\n    386       return None\n    387 \"\"\"\n--> 388         scope = {\"version_key\": version_key(),\n    389                  \"get_cuda_stream\": get_cuda_stream,\n    390                  \"self\": self,\n    391                  \"_spec_of\": self._spec_of,\n    392                  \"_key_of\": self._key_of,\n    393                  \"_device_of\": self._device_of,\n    394                  \"_pinned_memory_of\": self._pinned_memory_of,\n    395                  \"cache\": self.cache,\n    396                  \"__spec__\": __spec__,\n    397                  \"get_backend\": get_backend,\n    398                  \"get_current_device\": get_current_device,\n    399                  \"set_current_device\": set_current_device}\n    400         exec(src, scope)\n    401         return scope[self.fn.__name__]\n\nFile /kaggle/usr/lib/ubc_ocean_packages/triton/runtime/jit.py:120, in version_key()\n    118         contents += [hashlib.md5(f.read()).hexdigest()]\n    119 # ptxas version\n--> 120 ptxas = path_to_ptxas()[0]\n    121 ptxas_version = hashlib.md5(subprocess.check_output([ptxas, \"--version\"])).hexdigest()\n    122 return '-'.join(TRITON_VERSION) + '-' + ptxas_version + '-' + '-'.join(contents)\n\nFile /kaggle/usr/lib/ubc_ocean_packages/triton/common/backend.py:114, in path_to_ptxas()\n    112 ptxas_bin = ptxas.split(\" \")[0]\n    113 if os.path.exists(ptxas_bin) and os.path.isfile(ptxas_bin):\n--> 114     result = subprocess.check_output([ptxas_bin, \"--version\"], stderr=subprocess.STDOUT)\n    115     if result is not None:\n    116         version = re.search(r\".*release (\\d+\\.\\d+).*\", result.decode(\"utf-8\"), flags=re.MULTILINE)\n\nFile /opt/conda/lib/python3.10/subprocess.py:421, in check_output(timeout, *popenargs, **kwargs)\n    418         empty = b''\n    419     kwargs['input'] = empty\n--> 421 return run(*popenargs, stdout=PIPE, timeout=timeout, check=True,\n    422            **kwargs).stdout\n\nFile /opt/conda/lib/python3.10/subprocess.py:503, in run(input, capture_output, timeout, check, *popenargs, **kwargs)\n    500     kwargs['stdout'] = PIPE\n    501     kwargs['stderr'] = PIPE\n--> 503 with Popen(*popenargs, **kwargs) as process:\n    504     try:\n    505         stdout, stderr = process.communicate(input, timeout=timeout)\n\nFile /opt/conda/lib/python3.10/subprocess.py:971, in Popen.__init__(self, args, bufsize, executable, stdin, stdout, stderr, preexec_fn, close_fds, shell, cwd, env, universal_newlines, startupinfo, creationflags, restore_signals, start_new_session, pass_fds, user, group, extra_groups, encoding, errors, text, umask, pipesize)\n    967         if self.text_mode:\n    968             self.stderr = io.TextIOWrapper(self.stderr,\n    969                     encoding=encoding, errors=errors)\n--> 971     self._execute_child(args, executable, preexec_fn, close_fds,\n    972                         pass_fds, cwd, env,\n    973                         startupinfo, creationflags, shell,\n    974                         p2cread, p2cwrite,\n    975                         c2pread, c2pwrite,\n    976                         errread, errwrite,\n    977                         restore_signals,\n    978                         gid, gids, uid, umask,\n    979                         start_new_session)\n    980 except:\n    981     # Cleanup if the child failed starting.\n    982     for f in filter(None, (self.stdin, self.stdout, self.stderr)):\n\nFile /opt/conda/lib/python3.10/subprocess.py:1863, in Popen._execute_child(self, args, executable, preexec_fn, close_fds, pass_fds, cwd, env, startupinfo, creationflags, shell, p2cread, p2cwrite, c2pread, c2pwrite, errread, errwrite, restore_signals, gid, gids, uid, umask, start_new_session)\n   1861     if errno_num != 0:\n   1862         err_msg = os.strerror(errno_num)\n-> 1863     raise child_exception_type(errno_num, err_msg, err_filename)\n   1864 raise child_exception_type(err_msg)\n\nPermissionError: [Errno 13] Permission denied: '/kaggle/usr/lib/ubc_ocean_packages/triton/common/../third_party/cuda/bin/ptxas'\n```\n\nI tried to find the location of that file/folder, but to no avail; I was hoping I could perhaps change the permissions for that file/folder to fix the problem. Has anybody had any issues installing xFormers in an auxilliary notebook? Does anyone have any guidance on how to fix this problem?",
      "votes": 3
    },
    {
      "id": 2547858,
      "postDate": "2023-12-04T00:29:21.837Z",
      "content": "<p>I think xformers is primarily targeted for A100 and newer GPUs, so I wouldn't be surprised if it didn't work on p100. What exactly do you want to use xformers for?</p>",
      "rawMarkdown": "I think xformers is primarily targeted for A100 and newer GPUs, so I wouldn't be surprised if it didn't work on p100. What exactly do you want to use xformers for?",
      "votes": 1,
      "replies": [
        {
          "id": 2547974,
          "postDate": "2023-12-04T04:05:10.280Z",
          "content": "<p>Without giving away too much of my solution, I have an attention mechanism in my pipeline and I need to use memory-efficient attention (MEA) to achieve it. Of all the available MEA solutions (e.g. PyTorch, Lucidrains' <code>memory_efficient_attention_pytorch</code>, Tri Dao's <code>flash-attention</code>), xFormers was the only solution flexible enough to help me achieve what I wanted to do. It's true that the most recent versions of xFormers are aimed at more recent versions of CUDA (11.8, 12.0, etc.), but the package appears to have been extensively used in a recent competition (<a href=\"url\" target=\"_blank\">https://www.kaggle.com/competitions/stable-diffusion-image-to-prompts/overview</a>) that accepted CPU/GPU submissions. Nevertheless, I couldn't find any reference as to how the participants actually installed the package to work offline. And I don't need the most performing options of xFormers (e.g. Triton, Flash-Attention), just the plain vanilla CUTLASS; in fact, I have developed my solution on a Windows machine running on an RTX 4080 using CUTLASS (Triton is not available for Windows) and it works like a charm.</p>\n<p>Since my post, I have noticed that <a href=\"https://www.kaggle.com/ksmcg90\" target=\"_blank\">@ksmcg90</a> has created a dataset containing the wheels to install the CUDA 11.8 version of xFormers. I can report that his dataset actually works (thanks KSMCG90!) and I can use it on the P100 GPU with CUDA 11.4. I don't know exactly why it works (and why my previous attempts didn't work); at one point, I thought it was failing because of the upgrade to PyTorch 2.1.0. It now looks like xFormers for CUDA 11.8 appears to work on an \"outdated\" GPU with CUDA 11.4.</p>",
          "rawMarkdown": "Without giving away too much of my solution, I have an attention mechanism in my pipeline and I need to use memory-efficient attention (MEA) to achieve it. Of all the available MEA solutions (e.g. PyTorch, Lucidrains' `memory_efficient_attention_pytorch`, Tri Dao's `flash-attention`), xFormers was the only solution flexible enough to help me achieve what I wanted to do. It's true that the most recent versions of xFormers are aimed at more recent versions of CUDA (11.8, 12.0, etc.), but the package appears to have been extensively used in a recent competition ([https://www.kaggle.com/competitions/stable-diffusion-image-to-prompts/overview](url)) that accepted CPU/GPU submissions. Nevertheless, I couldn't find any reference as to how the participants actually installed the package to work offline. And I don't need the most performing options of xFormers (e.g. Triton, Flash-Attention), just the plain vanilla CUTLASS; in fact, I have developed my solution on a Windows machine running on an RTX 4080 using CUTLASS (Triton is not available for Windows) and it works like a charm.\n\nSince my post, I have noticed that @ksmcg90 has created a dataset containing the wheels to install the CUDA 11.8 version of xFormers. I can report that his dataset actually works (thanks KSMCG90!) and I can use it on the P100 GPU with CUDA 11.4. I don't know exactly why it works (and why my previous attempts didn't work); at one point, I thought it was failing because of the upgrade to PyTorch 2.1.0. It now looks like xFormers for CUDA 11.8 appears to work on an \"outdated\" GPU with CUDA 11.4.",
          "votes": 1,
          "replies": [
            {
              "id": 2548053,
              "postDate": "2023-12-04T05:29:27.260Z",
              "content": "<p>Can you use <a href=\"https://pytorch.org/docs/stable/generated/torch.nn.functional.scaled_dot_product_attention.html\" target=\"_blank\">torch sdpa</a>? </p>",
              "rawMarkdown": "Can you use [torch sdpa](https://pytorch.org/docs/stable/generated/torch.nn.functional.scaled_dot_product_attention.html)? "
            },
            {
              "id": 2548135,
              "postDate": "2023-12-04T07:00:23.673Z",
              "content": "<p>I used it in the early stages of development, but felt that I needed a more complete and flexible solution. Hence, xFormers.</p>",
              "rawMarkdown": "I used it in the early stages of development, but felt that I needed a more complete and flexible solution. Hence, xFormers."
            },
            {
              "id": 2548148,
              "postDate": "2023-12-04T07:17:49.807Z",
              "content": "<p>Ok good luck installing it :) </p>",
              "rawMarkdown": "Ok good luck installing it :) "
            }
          ]
        }
      ]
    },
    {
      "id": 2542482,
      "postDate": "2023-11-29T09:45:04.270Z",
      "content": "<p>Have you tried copying you're packages from the dataset you've attached (which would be a read only folder) to some other folder like <code>/tmp/packages</code>, and install from there?</p>",
      "rawMarkdown": "Have you tried copying you're packages from the dataset you've attached (which would be a read only folder) to some other folder like `/tmp/packages`, and install from there?",
      "replies": [
        {
          "id": 2543352,
          "postDate": "2023-11-30T02:42:22.437Z",
          "content": "<p>I tried that, but it didn't work. I also tried to download the wheels in a dataset to perform the install when I would want to use the notebook (which would be a pain, as it takes about 7 minutes to install xFormers), but I get the following error messages:</p>\n<pre><code>ERROR: Could  find a version that satisfies the requirement xformers ( versions: .post7cu118)\nERROR: No matching distribution found  xformers\n</code></pre>\n<p>I tried different wheels for xFormers, but nothing works. This is starting to test my patience…</p>\n<p><strong>EDIT:</strong> I have also posted an issue in the xFormers Github repo. Let's see what they have to say…</p>",
          "rawMarkdown": "I tried that, but it didn't work. I also tried to download the wheels in a dataset to perform the install when I would want to use the notebook (which would be a pain, as it takes about 7 minutes to install xFormers), but I get the following error messages:\n\n```python\nERROR: Could not find a version that satisfies the requirement xformers (from versions: 0.0.22.post7cu118)\nERROR: No matching distribution found for xformers\n```\n\nI tried different wheels for xFormers, but nothing works. This is starting to test my patience...\n\n**EDIT:** I have also posted an issue in the xFormers Github repo. Let's see what they have to say..."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2547858,
      "author_name": "Nicholas Broad",
      "author_url": "",
      "post_date": "2023-12-04T00:29:21.837000",
      "content": "<p>I think xformers is primarily targeted for A100 and newer GPUs, so I wouldn't be surprised if it didn't work on p100. What exactly do you want to use xformers for?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2547974,
          "author_name": "Patrick Robitaille",
          "author_url": "",
          "post_date": "2023-12-04T04:05:10.280000",
          "content": "<p>Without giving away too much of my solution, I have an attention mechanism in my pipeline and I need to use memory-efficient attention (MEA) to achieve it. Of all the available MEA solutions (e.g. PyTorch, Lucidrains' <code>memory_efficient_attention_pytorch</code>, Tri Dao's <code>flash-attention</code>), xFormers was the only solution flexible enough to help me achieve what I wanted to do. It's true that the most recent versions of xFormers are aimed at more recent versions of CUDA (11.8, 12.0, etc.), but the package appears to have been extensively used in a recent competition (<a href=\"url\" target=\"_blank\">https://www.kaggle.com/competitions/stable-diffusion-image-to-prompts/overview</a>) that accepted CPU/GPU submissions. Nevertheless, I couldn't find any reference as to how the participants actually installed the package to work offline. And I don't need the most performing options of xFormers (e.g. Triton, Flash-Attention), just the plain vanilla CUTLASS; in fact, I have developed my solution on a Windows machine running on an RTX 4080 using CUTLASS (Triton is not available for Windows) and it works like a charm.</p>\n<p>Since my post, I have noticed that <a href=\"https://www.kaggle.com/ksmcg90\" target=\"_blank\">@ksmcg90</a> has created a dataset containing the wheels to install the CUDA 11.8 version of xFormers. I can report that his dataset actually works (thanks KSMCG90!) and I can use it on the P100 GPU with CUDA 11.4. I don't know exactly why it works (and why my previous attempts didn't work); at one point, I thought it was failing because of the upgrade to PyTorch 2.1.0. It now looks like xFormers for CUDA 11.8 appears to work on an \"outdated\" GPU with CUDA 11.4.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2548053,
              "author_name": "Nicholas Broad",
              "author_url": "",
              "post_date": "2023-12-04T05:29:27.260000",
              "content": "<p>Can you use <a href=\"https://pytorch.org/docs/stable/generated/torch.nn.functional.scaled_dot_product_attention.html\" target=\"_blank\">torch sdpa</a>? </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2548135,
              "author_name": "Patrick Robitaille",
              "author_url": "",
              "post_date": "2023-12-04T07:00:23.673000",
              "content": "<p>I used it in the early stages of development, but felt that I needed a more complete and flexible solution. Hence, xFormers.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2548148,
              "author_name": "Nicholas Broad",
              "author_url": "",
              "post_date": "2023-12-04T07:17:49.807000",
              "content": "<p>Ok good luck installing it :) </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2542482,
      "author_name": "KSMCG90",
      "author_url": "",
      "post_date": "2023-11-29T09:45:04.270000",
      "content": "<p>Have you tried copying you're packages from the dataset you've attached (which would be a read only folder) to some other folder like <code>/tmp/packages</code>, and install from there?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2543352,
          "author_name": "Patrick Robitaille",
          "author_url": "",
          "post_date": "2023-11-30T02:42:22.437000",
          "content": "<p>I tried that, but it didn't work. I also tried to download the wheels in a dataset to perform the install when I would want to use the notebook (which would be a pain, as it takes about 7 minutes to install xFormers), but I get the following error messages:</p>\n<pre><code>ERROR: Could  find a version that satisfies the requirement xformers ( versions: .post7cu118)\nERROR: No matching distribution found  xformers\n</code></pre>\n<p>I tried different wheels for xFormers, but nothing works. This is starting to test my patience…</p>\n<p><strong>EDIT:</strong> I have also posted an issue in the xFormers Github repo. Let's see what they have to say…</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2541466": "As per normal Kaggle practice when you need packages not available in Kaggle's regular environment, I tried to set up a notebook containing the packages I need for this competition in order to use that notebook in my submission notebook as a utility script. However, one package is giving me some grief: xFormers.\n\nAs recommended in their Github repo and in line with the specs of the P100 GPU, I have installed xFormers using the following pip command: `pip3 install -U xformers --index-url https://download.pytorch.org/whl/cu118 `. However, when I try to import any component from xFormers in my submission notebook, I get hit with a Errno 13 message as xFormers looks at importing Triton components (I will only show the relevant part of the error message):\n\n```python\nFile /kaggle/usr/lib/ubc_ocean_packages/triton/language/math.py:4\n      1 import functools\n      2 import os\n----> 4 from . import core\n      7 @functools.lru_cache()\n      8 def libdevice_path():\n      9     import torch\n\nFile /kaggle/usr/lib/ubc_ocean_packages/triton/language/core.py:1376\n   1370     rvalue, rindices = reduce((input, index), axis, combine_fn,\n   1371                               _builder=_builder, _generator=_generator)\n   1372     return rvalue, rindices\n   1375 @jit\n-> 1376 def minimum(x, y):\n   1377     \"\"\"\n   1378     Computes the element-wise minimum of :code:`x` and :code:`y`.\n   1379 \n   (...)\n   1383     :type other: Block\n   1384     \"\"\"\n   1385     return where(x < y, x, y)\n\nFile /kaggle/usr/lib/ubc_ocean_packages/triton/runtime/jit.py:542, in jit(fn, version, do_not_specialize, debug, noinline, interpret)\n    534         return JITFunction(\n    535             fn,\n    536             version=version,\n   (...)\n    539             noinline=noinline,\n    540         )\n    541 if fn is not None:\n--> 542     return decorator(fn)\n    544 else:\n    545     return decorator\n\nFile /kaggle/usr/lib/ubc_ocean_packages/triton/runtime/jit.py:534, in jit.<locals>.decorator(fn)\n    532     return GridSelector(fn)\n    533 else:\n--> 534     return JITFunction(\n    535         fn,\n    536         version=version,\n    537         do_not_specialize=do_not_specialize,\n    538         debug=debug,\n    539         noinline=noinline,\n    540     )\n\nFile /kaggle/usr/lib/ubc_ocean_packages/triton/runtime/jit.py:433, in JITFunction.__init__(self, fn, version, do_not_specialize, debug, noinline)\n    431 self.constexprs = [self.arg_names.index(name) for name, ty in self.__annotations__.items() if 'constexpr' in ty]\n    432 # launcher\n--> 433 self.run = self._make_launcher()\n    434 # re-use docs of wrapped function\n    435 self.__doc__ = fn.__doc__\n\nFile /kaggle/usr/lib/ubc_ocean_packages/triton/runtime/jit.py:388, in JITFunction._make_launcher(self)\n    317         args_signature = ', '.join(name if dflt == inspect._empty else f'{name} = {dflt}' for name, dflt in zip(self.arg_names, self.arg_defaults))\n    319         src = f\"\"\"\n    320 def {self.fn.__name__}({args_signature}, grid=None, num_warps=4, num_stages=3, extern_libs=None, stream=None, warmup=False, device=None, device_type=None):\n    321     from ..compiler import compile, CompiledKernel\n   (...)\n    386       return None\n    387 \"\"\"\n--> 388         scope = {\"version_key\": version_key(),\n    389                  \"get_cuda_stream\": get_cuda_stream,\n    390                  \"self\": self,\n    391                  \"_spec_of\": self._spec_of,\n    392                  \"_key_of\": self._key_of,\n    393                  \"_device_of\": self._device_of,\n    394                  \"_pinned_memory_of\": self._pinned_memory_of,\n    395                  \"cache\": self.cache,\n    396                  \"__spec__\": __spec__,\n    397                  \"get_backend\": get_backend,\n    398                  \"get_current_device\": get_current_device,\n    399                  \"set_current_device\": set_current_device}\n    400         exec(src, scope)\n    401         return scope[self.fn.__name__]\n\nFile /kaggle/usr/lib/ubc_ocean_packages/triton/runtime/jit.py:120, in version_key()\n    118         contents += [hashlib.md5(f.read()).hexdigest()]\n    119 # ptxas version\n--> 120 ptxas = path_to_ptxas()[0]\n    121 ptxas_version = hashlib.md5(subprocess.check_output([ptxas, \"--version\"])).hexdigest()\n    122 return '-'.join(TRITON_VERSION) + '-' + ptxas_version + '-' + '-'.join(contents)\n\nFile /kaggle/usr/lib/ubc_ocean_packages/triton/common/backend.py:114, in path_to_ptxas()\n    112 ptxas_bin = ptxas.split(\" \")[0]\n    113 if os.path.exists(ptxas_bin) and os.path.isfile(ptxas_bin):\n--> 114     result = subprocess.check_output([ptxas_bin, \"--version\"], stderr=subprocess.STDOUT)\n    115     if result is not None:\n    116         version = re.search(r\".*release (\\d+\\.\\d+).*\", result.decode(\"utf-8\"), flags=re.MULTILINE)\n\nFile /opt/conda/lib/python3.10/subprocess.py:421, in check_output(timeout, *popenargs, **kwargs)\n    418         empty = b''\n    419     kwargs['input'] = empty\n--> 421 return run(*popenargs, stdout=PIPE, timeout=timeout, check=True,\n    422            **kwargs).stdout\n\nFile /opt/conda/lib/python3.10/subprocess.py:503, in run(input, capture_output, timeout, check, *popenargs, **kwargs)\n    500     kwargs['stdout'] = PIPE\n    501     kwargs['stderr'] = PIPE\n--> 503 with Popen(*popenargs, **kwargs) as process:\n    504     try:\n    505         stdout, stderr = process.communicate(input, timeout=timeout)\n\nFile /opt/conda/lib/python3.10/subprocess.py:971, in Popen.__init__(self, args, bufsize, executable, stdin, stdout, stderr, preexec_fn, close_fds, shell, cwd, env, universal_newlines, startupinfo, creationflags, restore_signals, start_new_session, pass_fds, user, group, extra_groups, encoding, errors, text, umask, pipesize)\n    967         if self.text_mode:\n    968             self.stderr = io.TextIOWrapper(self.stderr,\n    969                     encoding=encoding, errors=errors)\n--> 971     self._execute_child(args, executable, preexec_fn, close_fds,\n    972                         pass_fds, cwd, env,\n    973                         startupinfo, creationflags, shell,\n    974                         p2cread, p2cwrite,\n    975                         c2pread, c2pwrite,\n    976                         errread, errwrite,\n    977                         restore_signals,\n    978                         gid, gids, uid, umask,\n    979                         start_new_session)\n    980 except:\n    981     # Cleanup if the child failed starting.\n    982     for f in filter(None, (self.stdin, self.stdout, self.stderr)):\n\nFile /opt/conda/lib/python3.10/subprocess.py:1863, in Popen._execute_child(self, args, executable, preexec_fn, close_fds, pass_fds, cwd, env, startupinfo, creationflags, shell, p2cread, p2cwrite, c2pread, c2pwrite, errread, errwrite, restore_signals, gid, gids, uid, umask, start_new_session)\n   1861     if errno_num != 0:\n   1862         err_msg = os.strerror(errno_num)\n-> 1863     raise child_exception_type(errno_num, err_msg, err_filename)\n   1864 raise child_exception_type(err_msg)\n\nPermissionError: [Errno 13] Permission denied: '/kaggle/usr/lib/ubc_ocean_packages/triton/common/../third_party/cuda/bin/ptxas'\n```\n\nI tried to find the location of that file/folder, but to no avail; I was hoping I could perhaps change the permissions for that file/folder to fix the problem. Has anybody had any issues installing xFormers in an auxilliary notebook? Does anyone have any guidance on how to fix this problem?",
    "2547858": "I think xformers is primarily targeted for A100 and newer GPUs, so I wouldn't be surprised if it didn't work on p100. What exactly do you want to use xformers for?",
    "2542482": "Have you tried copying you're packages from the dataset you've attached (which would be a read only folder) to some other folder like `/tmp/packages`, and install from there?"
  }
}