{
  "id": 388236,
  "title": "Cannot re-initialize CUDA in forked subprocess",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/388236",
  "author_name": "Gabriel Vinicius",
  "post_date": "2023-02-16T15:02:40.358000",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I am using Huging Face Acclerate for train with 2x T4 on kaggle kernel like <a href=\"https://www.kaggle.com/code/heyytanay/rsna-pytorch-multi-gpu-training-w-b-fp16\" target=\"_blank\">here</a>. Recently the following error appeared:</p>\n<p><code>RuntimeError: Cannot re-initialize CUDA in forked subprocess. To use CUDA with multiprocessing, you must use the 'spawn' start method</code></p>\n<p>Anyone haved the same error and can help me? :)</p>",
  "messages": [
    {
      "id": 2147572,
      "postDate": "2023-02-16T18:01:01.793Z",
      "content": "<p>Another solution might be <strong>not</strong> to do the following inside the dataloader:</p>\n<pre><code>img = torch.tensor(img, dtype=torch.)\ntarget = torch.tensor(target, dtype=torch.)\n</code></pre>",
      "rawMarkdown": "Another solution might be **not** to do the following inside the dataloader:\n```python\nimg = torch.tensor(img, dtype=torch.float)\ntarget = torch.tensor(target, dtype=torch.float)\n```",
      "votes": 1
    },
    {
      "id": 2147562,
      "postDate": "2023-02-16T17:51:19.927Z",
      "content": "<p>What happens if you set <code>num_workers=0</code>? And what happens if you call <code>multiprocessing.set_start_method('spawn', force=True)</code> when you have <code>num_workers</code> other than 0?</p>",
      "rawMarkdown": "What happens if you set `num_workers=0`? And what happens if you call `multiprocessing.set_start_method('spawn', force=True)` when you have `num_workers` other than 0?",
      "votes": 1,
      "replies": [
        {
          "id": 2147564,
          "postDate": "2023-02-16T17:54:41.613Z",
          "content": "<p>My quota of this week finished, waiting for request …</p>",
          "rawMarkdown": "My quota of this week finished, waiting for request ..."
        }
      ]
    },
    {
      "id": 2147356,
      "postDate": "2023-02-16T15:02:40.360Z",
      "content": "<p>I am using Huging Face Acclerate for train with 2x T4 on kaggle kernel like <a href=\"https://www.kaggle.com/code/heyytanay/rsna-pytorch-multi-gpu-training-w-b-fp16\" target=\"_blank\">here</a>. Recently the following error appeared:</p>\n<p><code>RuntimeError: Cannot re-initialize CUDA in forked subprocess. To use CUDA with multiprocessing, you must use the 'spawn' start method</code></p>\n<p>Anyone haved the same error and can help me? :)</p>",
      "rawMarkdown": "I am using Huging Face Acclerate for train with 2x T4 on kaggle kernel like [here](https://www.kaggle.com/code/heyytanay/rsna-pytorch-multi-gpu-training-w-b-fp16). Recently the following error appeared:\n\n`RuntimeError: Cannot re-initialize CUDA in forked subprocess. To use CUDA with multiprocessing, you must use the 'spawn' start method`\n\nAnyone haved the same error and can help me? :)",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 2147572,
      "author_name": "Antti Isosalo",
      "author_url": "",
      "post_date": "2023-02-16T18:01:01.793000",
      "content": "<p>Another solution might be <strong>not</strong> to do the following inside the dataloader:</p>\n<pre><code>img = torch.tensor(img, dtype=torch.)\ntarget = torch.tensor(target, dtype=torch.)\n</code></pre>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2147562,
      "author_name": "Antti Isosalo",
      "author_url": "",
      "post_date": "2023-02-16T17:51:19.927000",
      "content": "<p>What happens if you set <code>num_workers=0</code>? And what happens if you call <code>multiprocessing.set_start_method('spawn', force=True)</code> when you have <code>num_workers</code> other than 0?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2147564,
          "author_name": "Gabriel Vinicius",
          "author_url": "",
          "post_date": "2023-02-16T17:54:41.613000",
          "content": "<p>My quota of this week finished, waiting for request …</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2147572": "Another solution might be **not** to do the following inside the dataloader:\n```python\nimg = torch.tensor(img, dtype=torch.float)\ntarget = torch.tensor(target, dtype=torch.float)\n```",
    "2147562": "What happens if you set `num_workers=0`? And what happens if you call `multiprocessing.set_start_method('spawn', force=True)` when you have `num_workers` other than 0?",
    "2147356": "I am using Huging Face Acclerate for train with 2x T4 on kaggle kernel like [here](https://www.kaggle.com/code/heyytanay/rsna-pytorch-multi-gpu-training-w-b-fp16). Recently the following error appeared:\n\n`RuntimeError: Cannot re-initialize CUDA in forked subprocess. To use CUDA with multiprocessing, you must use the 'spawn' start method`\n\nAnyone haved the same error and can help me? :)"
  }
}