{
  "id": 382270,
  "title": "Notebook Out of Memory Error During Inference | Analysis [SOLUTION]",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/382270",
  "author_name": "Mark Wijkhuizen",
  "post_date": "2023-01-30T10:42:49.723000",
  "votes": 10,
  "comment_count": 9,
  "views": 0,
  "content": "<p>During the inference process I often encounter <code>Notebook Out of Memory</code> errors. The problem seems to be caused by the memory filling up when reading the dicom files.</p>\n<p>The example below simply reads all dicom files in a function and should not fill up the memory, but just 1000 dicom files make the memory grow from ~2.5GB -&gt; ~8GB. This memory can not be freed with a <code>gc.collect()</code> call. This behaviour is observed for both <code>pydicom</code> and <code>dicomdsl</code>.</p>\n<p>All dicom files are read within a function, where variables only exist within the function scope and should thus not allocate memory outside the function call. After each function call, the memory usage should be equal to the memory usage before the function call.</p>\n<p>How can the dicom files be read without making the memory consumption grow indefinitely?</p>\n<pre><code># Read Train DataFrame\ntrain = pd.read_csv('/kaggle/input/rsna-breast-cancer-detection/train.csv').head(1000)\n\n# Get File Path to DICOM files\ndef get_file_path(args):\n    patient_id, image_id = args\n    return f'/kaggle/input/rsna-breast-cancer-detection/train_images/{patient_id}/{image_id}.dcm'\n\n# Assign DICOM file path for each training sample\ntrain['file_path'] = train[['patient_id', 'image_id']].apply(get_file_path, axis=1)\n\n# Function which simply reads a DICOM file\ndef read_dicom(fp):\n    dicom = dicomsdl.open(fp)\n    image = dicom.pixelData()\n\n# For 1000 training samples, read the DICOM file\nfor fp in tqdm(train['file_path']):\n    read_dicom(fp)\n</code></pre>\n<p><strong>SOLUTION</strong></p>\n<p>The solution to this problem is rather silly, but luckily easy to implement. All that is needed is to allocate all memory, except ~1GB before the processing of the dicoms, and free this memory after the dicoms are processed. This prevents the dicom processing to magically use up all memory.</p>\n<p>Keep in mind the amount of memory you need to allocate depends on how much memory is used just before you start processing the dicoms and how much memory the dicom processing uses.</p>\n<pre><code># Define large array to allocate 10GB of memory\nhuge_10gb_array = np.arange(int(10 * 2**30 / 8), dtype=np.int64)\n\n# Process All Dicoms\nmy_dicom_processing_function()\n\n# Free Up Memory\ndel huge_10gb_array\ngc.collect()\n</code></pre>",
  "messages": [
    {
      "id": 2121476,
      "postDate": "2023-01-30T10:42:49.723Z",
      "content": "<p>During the inference process I often encounter <code>Notebook Out of Memory</code> errors. The problem seems to be caused by the memory filling up when reading the dicom files.</p>\n<p>The example below simply reads all dicom files in a function and should not fill up the memory, but just 1000 dicom files make the memory grow from ~2.5GB -&gt; ~8GB. This memory can not be freed with a <code>gc.collect()</code> call. This behaviour is observed for both <code>pydicom</code> and <code>dicomdsl</code>.</p>\n<p>All dicom files are read within a function, where variables only exist within the function scope and should thus not allocate memory outside the function call. After each function call, the memory usage should be equal to the memory usage before the function call.</p>\n<p>How can the dicom files be read without making the memory consumption grow indefinitely?</p>\n<pre><code># Read Train DataFrame\ntrain = pd.read_csv('/kaggle/input/rsna-breast-cancer-detection/train.csv').head(1000)\n\n# Get File Path to DICOM files\ndef get_file_path(args):\n    patient_id, image_id = args\n    return f'/kaggle/input/rsna-breast-cancer-detection/train_images/{patient_id}/{image_id}.dcm'\n\n# Assign DICOM file path for each training sample\ntrain['file_path'] = train[['patient_id', 'image_id']].apply(get_file_path, axis=1)\n\n# Function which simply reads a DICOM file\ndef read_dicom(fp):\n    dicom = dicomsdl.open(fp)\n    image = dicom.pixelData()\n\n# For 1000 training samples, read the DICOM file\nfor fp in tqdm(train['file_path']):\n    read_dicom(fp)\n</code></pre>\n<p><strong>SOLUTION</strong></p>\n<p>The solution to this problem is rather silly, but luckily easy to implement. All that is needed is to allocate all memory, except ~1GB before the processing of the dicoms, and free this memory after the dicoms are processed. This prevents the dicom processing to magically use up all memory.</p>\n<p>Keep in mind the amount of memory you need to allocate depends on how much memory is used just before you start processing the dicoms and how much memory the dicom processing uses.</p>\n<pre><code># Define large array to allocate 10GB of memory\nhuge_10gb_array = np.arange(int(10 * 2**30 / 8), dtype=np.int64)\n\n# Process All Dicoms\nmy_dicom_processing_function()\n\n# Free Up Memory\ndel huge_10gb_array\ngc.collect()\n</code></pre>",
      "rawMarkdown": "During the inference process I often encounter `Notebook Out of Memory` errors. The problem seems to be caused by the memory filling up when reading the dicom files.\n\nThe example below simply reads all dicom files in a function and should not fill up the memory, but just 1000 dicom files make the memory grow from ~2.5GB -> ~8GB. This memory can not be freed with a `gc.collect()` call. This behaviour is observed for both `pydicom` and `dicomdsl`.\n\nAll dicom files are read within a function, where variables only exist within the function scope and should thus not allocate memory outside the function call. After each function call, the memory usage should be equal to the memory usage before the function call.\n\nHow can the dicom files be read without making the memory consumption grow indefinitely?\n\n```\n# Read Train DataFrame\ntrain = pd.read_csv('/kaggle/input/rsna-breast-cancer-detection/train.csv').head(1000)\n\n# Get File Path to DICOM files\ndef get_file_path(args):\n    patient_id, image_id = args\n    return f'/kaggle/input/rsna-breast-cancer-detection/train_images/{patient_id}/{image_id}.dcm'\n    \n# Assign DICOM file path for each training sample\ntrain['file_path'] = train[['patient_id', 'image_id']].apply(get_file_path, axis=1)\n\n# Function which simply reads a DICOM file\ndef read_dicom(fp):\n    dicom = dicomsdl.open(fp)\n    image = dicom.pixelData()\n    \n# For 1000 training samples, read the DICOM file\nfor fp in tqdm(train['file_path']):\n    read_dicom(fp)\n```\n\n**SOLUTION**\n\nThe solution to this problem is rather silly, but luckily easy to implement. All that is needed is to allocate all memory, except ~1GB before the processing of the dicoms, and free this memory after the dicoms are processed. This prevents the dicom processing to magically use up all memory.\n\nKeep in mind the amount of memory you need to allocate depends on how much memory is used just before you start processing the dicoms and how much memory the dicom processing uses.\n\n```\n# Define large array to allocate 10GB of memory\nhuge_10gb_array = np.arange(int(10 * 2**30 / 8), dtype=np.int64)\n\n# Process All Dicoms\nmy_dicom_processing_function()\n\n# Free Up Memory\ndel huge_10gb_array\ngc.collect()\n```",
      "votes": 10
    },
    {
      "id": 2123533,
      "postDate": "2023-01-31T15:10:19.280Z",
      "content": "<p>I am not sure. But I have met that same kernel submission,  sometimes works, sometimes doesn't (oom or timeout).</p>",
      "rawMarkdown": "I am not sure. But I have met that same kernel submission,  sometimes works, sometimes doesn't (oom or timeout).\n\n",
      "votes": 2
    },
    {
      "id": 2123827,
      "postDate": "2023-01-31T17:19:46.383Z",
      "content": "<p>I created this notebook to test your code snippet and investigate the Out of Memory issue, but it runs just fine (it monitors the system memory usage, and during the loop over DCM loadings the memory usage remains on average around 1GB).<br>\n<a href=\"https://www.kaggle.com/code/rasoulmojtahedzadeh/test-bug-out-of-memory\" target=\"_blank\">https://www.kaggle.com/code/rasoulmojtahedzadeh/test-bug-out-of-memory</a></p>\n<p>I believe that the error you get has nothing to do with the code snippet you have shared. The problem shall lay in your tensorflow model loading for inference. For example, reduce the batch size for inference to 1 and try to delete any unused variable and then <code>gc.collect()</code>.</p>",
      "rawMarkdown": "I created this notebook to test your code snippet and investigate the Out of Memory issue, but it runs just fine (it monitors the system memory usage, and during the loop over DCM loadings the memory usage remains on average around 1GB).\nhttps://www.kaggle.com/code/rasoulmojtahedzadeh/test-bug-out-of-memory\n\nI believe that the error you get has nothing to do with the code snippet you have shared. The problem shall lay in your tensorflow model loading for inference. For example, reduce the batch size for inference to 1 and try to delete any unused variable and then `gc.collect()`.",
      "replies": [
        {
          "id": 2124867,
          "postDate": "2023-02-01T09:04:55.247Z",
          "content": "<p>There seems to be a mismatch between the memory usage according to <code>psutil.virtual_memory()</code> and the session memory usage, see the screenshot below. The memory usage according to <code>psutil</code> remains consistently around 1GB, however, the session memory usage grows steadily and reaches 6.3GB after reading all 1000 dicom files.</p>\n<p>The cause of these different memory usage statistics should answer why the memory usage keeps growing when reading the dicom files.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4433335%2Ffa8fb8abfefd816f3f70a3aa2be6522a%2Ftemp.png?generation=1675241703082112&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "There seems to be a mismatch between the memory usage according to `psutil.virtual_memory()` and the session memory usage, see the screenshot below. The memory usage according to `psutil` remains consistently around 1GB, however, the session memory usage grows steadily and reaches 6.3GB after reading all 1000 dicom files.\n\nThe cause of these different memory usage statistics should answer why the memory usage keeps growing when reading the dicom files.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4433335%2Ffa8fb8abfefd816f3f70a3aa2be6522a%2Ftemp.png?generation=1675241703082112&alt=media)",
          "replies": [
            {
              "id": 2124940,
              "postDate": "2023-02-01T10:06:34.393Z",
              "content": "<p>You're right! I didn't notice the session memory usage. It seems like there is some kind of memory leakage. I updated the code to test it with <code>pydicom</code> and <code>python pure binary file open</code>, in all cases we have this rapidly increasing session memory issue. Maybe some <strong>Kaggle stuff</strong> can comment on it?</p>",
              "rawMarkdown": "You're right! I didn't notice the session memory usage. It seems like there is some kind of memory leakage. I updated the code to test it with `pydicom` and `python pure binary file open`, in all cases we have this rapidly increasing session memory issue. Maybe some **Kaggle stuff** can comment on it?",
              "votes": 1
            },
            {
              "id": 2125756,
              "postDate": "2023-02-01T21:07:03.297Z",
              "content": "<p>I found a silly, almost comical, solution to the problem, as demonstrated in the following <a href=\"https://www.kaggle.com/code/markwijkhuizen/test-bug-out-of-memory/notebook?scriptVersionId=117965260\" target=\"_blank\">notebook</a>. If you allocate all available memory, except ~1GB, and free the memory afterwards, you do not allow the magical memory consumption to grow.</p>\n<p>The <code>psutil</code> memory consumption and the session memory consumption are roughly equal, the session memory usage is &lt;1GB, see screenshot below. I will validate if this method can be successfully applied to an actual inference notebook.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4433335%2Fb71f9b695a5c9ea4a2f3150a8cd28b42%2Ftemp2.png?generation=1675285581924377&amp;alt=media\" alt=\"\"></p>",
              "rawMarkdown": "I found a silly, almost comical, solution to the problem, as demonstrated in the following [notebook](https://www.kaggle.com/code/markwijkhuizen/test-bug-out-of-memory/notebook?scriptVersionId=117965260). If you allocate all available memory, except ~1GB, and free the memory afterwards, you do not allow the magical memory consumption to grow.\n\nThe `psutil` memory consumption and the session memory consumption are roughly equal, the session memory usage is <1GB, see screenshot below. I will validate if this method can be successfully applied to an actual inference notebook.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4433335%2Fb71f9b695a5c9ea4a2f3150a8cd28b42%2Ftemp2.png?generation=1675285581924377&alt=media)",
              "votes": 2
            },
            {
              "id": 2125772,
              "postDate": "2023-02-01T21:18:10.603Z",
              "content": "<p>Wow! That's really crazy but it works :)</p>",
              "rawMarkdown": "Wow! That's really crazy but it works :)"
            }
          ]
        }
      ]
    },
    {
      "id": 2122475,
      "postDate": "2023-01-31T00:50:26.990Z",
      "content": "<p>Not sure what is happening, but keep in mind <code>dicom.pixelData()</code> actually does a decompression saving the decompressed data in-place.</p>\n<p>I suggest applying bisection debugging.</p>",
      "rawMarkdown": "Not sure what is happening, but keep in mind `dicom.pixelData()` actually does a decompression saving the decompressed data in-place.\n\nI suggest applying bisection debugging."
    },
    {
      "id": 2126061,
      "postDate": "2023-02-02T04:29:55.763Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2122380,
      "postDate": "2023-01-30T22:20:01.243Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2123533,
      "author_name": "dragon zhang",
      "author_url": "",
      "post_date": "2023-01-31T15:10:19.280000",
      "content": "<p>I am not sure. But I have met that same kernel submission,  sometimes works, sometimes doesn't (oom or timeout).</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2123827,
      "author_name": "Rasoul Mojtahedzadeh",
      "author_url": "",
      "post_date": "2023-01-31T17:19:46.383000",
      "content": "<p>I created this notebook to test your code snippet and investigate the Out of Memory issue, but it runs just fine (it monitors the system memory usage, and during the loop over DCM loadings the memory usage remains on average around 1GB).<br>\n<a href=\"https://www.kaggle.com/code/rasoulmojtahedzadeh/test-bug-out-of-memory\" target=\"_blank\">https://www.kaggle.com/code/rasoulmojtahedzadeh/test-bug-out-of-memory</a></p>\n<p>I believe that the error you get has nothing to do with the code snippet you have shared. The problem shall lay in your tensorflow model loading for inference. For example, reduce the batch size for inference to 1 and try to delete any unused variable and then <code>gc.collect()</code>.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2124867,
          "author_name": "Mark Wijkhuizen",
          "author_url": "",
          "post_date": "2023-02-01T09:04:55.247000",
          "content": "<p>There seems to be a mismatch between the memory usage according to <code>psutil.virtual_memory()</code> and the session memory usage, see the screenshot below. The memory usage according to <code>psutil</code> remains consistently around 1GB, however, the session memory usage grows steadily and reaches 6.3GB after reading all 1000 dicom files.</p>\n<p>The cause of these different memory usage statistics should answer why the memory usage keeps growing when reading the dicom files.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4433335%2Ffa8fb8abfefd816f3f70a3aa2be6522a%2Ftemp.png?generation=1675241703082112&amp;alt=media\" alt=\"\"></p>",
          "votes": 0,
          "replies": [
            {
              "id": 2124940,
              "author_name": "Rasoul Mojtahedzadeh",
              "author_url": "",
              "post_date": "2023-02-01T10:06:34.393000",
              "content": "<p>You're right! I didn't notice the session memory usage. It seems like there is some kind of memory leakage. I updated the code to test it with <code>pydicom</code> and <code>python pure binary file open</code>, in all cases we have this rapidly increasing session memory issue. Maybe some <strong>Kaggle stuff</strong> can comment on it?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2125756,
              "author_name": "Mark Wijkhuizen",
              "author_url": "",
              "post_date": "2023-02-01T21:07:03.297000",
              "content": "<p>I found a silly, almost comical, solution to the problem, as demonstrated in the following <a href=\"https://www.kaggle.com/code/markwijkhuizen/test-bug-out-of-memory/notebook?scriptVersionId=117965260\" target=\"_blank\">notebook</a>. If you allocate all available memory, except ~1GB, and free the memory afterwards, you do not allow the magical memory consumption to grow.</p>\n<p>The <code>psutil</code> memory consumption and the session memory consumption are roughly equal, the session memory usage is &lt;1GB, see screenshot below. I will validate if this method can be successfully applied to an actual inference notebook.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4433335%2Fb71f9b695a5c9ea4a2f3150a8cd28b42%2Ftemp2.png?generation=1675285581924377&amp;alt=media\" alt=\"\"></p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2125772,
              "author_name": "Rasoul Mojtahedzadeh",
              "author_url": "",
              "post_date": "2023-02-01T21:18:10.603000",
              "content": "<p>Wow! That's really crazy but it works :)</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2122475,
      "author_name": "Alfredo Maussa",
      "author_url": "",
      "post_date": "2023-01-31T00:50:26.990000",
      "content": "<p>Not sure what is happening, but keep in mind <code>dicom.pixelData()</code> actually does a decompression saving the decompressed data in-place.</p>\n<p>I suggest applying bisection debugging.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2126061,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-02-02T04:29:55.763000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2122380,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-01-30T22:20:01.243000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2121476": "During the inference process I often encounter `Notebook Out of Memory` errors. The problem seems to be caused by the memory filling up when reading the dicom files.\n\nThe example below simply reads all dicom files in a function and should not fill up the memory, but just 1000 dicom files make the memory grow from ~2.5GB -> ~8GB. This memory can not be freed with a `gc.collect()` call. This behaviour is observed for both `pydicom` and `dicomdsl`.\n\nAll dicom files are read within a function, where variables only exist within the function scope and should thus not allocate memory outside the function call. After each function call, the memory usage should be equal to the memory usage before the function call.\n\nHow can the dicom files be read without making the memory consumption grow indefinitely?\n\n```\n# Read Train DataFrame\ntrain = pd.read_csv('/kaggle/input/rsna-breast-cancer-detection/train.csv').head(1000)\n\n# Get File Path to DICOM files\ndef get_file_path(args):\n    patient_id, image_id = args\n    return f'/kaggle/input/rsna-breast-cancer-detection/train_images/{patient_id}/{image_id}.dcm'\n    \n# Assign DICOM file path for each training sample\ntrain['file_path'] = train[['patient_id', 'image_id']].apply(get_file_path, axis=1)\n\n# Function which simply reads a DICOM file\ndef read_dicom(fp):\n    dicom = dicomsdl.open(fp)\n    image = dicom.pixelData()\n    \n# For 1000 training samples, read the DICOM file\nfor fp in tqdm(train['file_path']):\n    read_dicom(fp)\n```\n\n**SOLUTION**\n\nThe solution to this problem is rather silly, but luckily easy to implement. All that is needed is to allocate all memory, except ~1GB before the processing of the dicoms, and free this memory after the dicoms are processed. This prevents the dicom processing to magically use up all memory.\n\nKeep in mind the amount of memory you need to allocate depends on how much memory is used just before you start processing the dicoms and how much memory the dicom processing uses.\n\n```\n# Define large array to allocate 10GB of memory\nhuge_10gb_array = np.arange(int(10 * 2**30 / 8), dtype=np.int64)\n\n# Process All Dicoms\nmy_dicom_processing_function()\n\n# Free Up Memory\ndel huge_10gb_array\ngc.collect()\n```",
    "2123533": "I am not sure. But I have met that same kernel submission,  sometimes works, sometimes doesn't (oom or timeout).\n\n",
    "2123827": "I created this notebook to test your code snippet and investigate the Out of Memory issue, but it runs just fine (it monitors the system memory usage, and during the loop over DCM loadings the memory usage remains on average around 1GB).\nhttps://www.kaggle.com/code/rasoulmojtahedzadeh/test-bug-out-of-memory\n\nI believe that the error you get has nothing to do with the code snippet you have shared. The problem shall lay in your tensorflow model loading for inference. For example, reduce the batch size for inference to 1 and try to delete any unused variable and then `gc.collect()`.",
    "2122475": "Not sure what is happening, but keep in mind `dicom.pixelData()` actually does a decompression saving the decompressed data in-place.\n\nI suggest applying bisection debugging.",
    "2126061": "",
    "2122380": ""
  }
}