{
  "id": 350266,
  "title": "Notebook Running out of memory",
  "url": "/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/350266",
  "author_name": "cbarb15",
  "post_date": "2022-09-05T00:17:21.405000",
  "votes": -1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I have a very simple section of my notebook where I am simply trying to process data and create two lists filled with tensors for the image data and for the labels.  I have not even begun to train anything and my notebook  throws an error saying </p>\n<blockquote>\n  <p>\"your Notebook tried to allocate more memory than available. It has been restarted\"</p>\n</blockquote>\n<p>This is kind of confusing to me as I do not feel like I am writing that much data.   I have researched this error and there are some tips for reducing memory, but they seem to be with pandas data frames and I am not reading any giant data frames. Do I need to compress my images or do I need to delete things that I am not?  Thanks<br>\nBelow is my code.  </p>\n<pre><code>`!pip install -qU ../input/for-pydicom/python_gdcm-3.0.14-cp37-cp37m-manylinux_2_17_x86_64.manylinux2014_x86_64.whl \n!pip install ../input/pylibjpeg-whl/pylibjpeg-1.4.0-py3-none-any.whl --find-links frozen_packages --no-index\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nfrom pydicom import dcmread\n# from pydicom.pixel_data_handlers.util import apply_voi_lut\nimport nibabel as nib\nimport os\nos.environ['TF_CPP_MIN_LOG_LEVEL'] = '3' \nimport tensorflow as tf\nimport matplotlib.pyplot as plt\nimport tensorflow as tf\nimport re\n\nCURRENT_DIR_PATH = \"./\"\nTRAIN_IMAGE_PATH = \"../input/rsna-2022-cervical-spine-fracture-detection/train_images\"\nSEGMENTATIONS_PATH = \"../input/rsna-2022-cervical-spine-fracture-detection/segmentations\"\nTRAIN_CSV_PATH = \"../input/rsna-2022-cervical-spine-fracture-detection/train.csv\"\ntrain_data = pd.read_csv(TRAIN_CSV_PATH)\n\ndef atoi(text):\n    return int(text) if text.isdigit() else text\ndef natural_keys(text):\n    return [atoi(c) for c in re.split(r'(\\d+)', text)]\n\ndef labels_for_image(vertebrae):\n    vertebrae_for_image = list(filter(lambda vertebra: vertebra &gt; 0.0, vertebrae))\n    return vertebrae_for_image if len(vertebrae_for_image) &gt; 0 else [0]\n\ndef create_train_dataset():\n    segmentation_patient_files = os.listdir(SEGMENTATIONS_PATH)\n    train_images = []\n    train_image_labels = []\n    for segmentation_patient_file in segmentation_patient_files:\n        file_path = os.path.join(SEGMENTATIONS_PATH, segmentation_patient_file)\n        segmentation_file = nib.load(file_path).get_fdata()\n        segmentation_file_transposed = segmentation_file[:, ::-1, ::-1].transpose(2, 1, 0)\n\n        for slice_number in range(0, len(segmentation_file_transposed)):\n            dicom_slice = segmentation_file_transposed[slice_number]\n            patient_id = os.path.join(TRAIN_IMAGE_PATH, file_path.split(\"/\")[-1].replace('.nii', \"\")).split('/')[-1]\n            vertebrae = np.unique(dicom_slice)\n            labels = labels_for_image(vertebrae)\n            labels_tensor = tf.convert_to_tensor(labels)\n            train_image_labels.append(labels_tensor)\n            train_images_path = os.path.join(TRAIN_IMAGE_PATH, file_path.split(\"/\")[-1].replace('.nii', \"\"))\n\n            image_dir = os.listdir(train_images_path)\n            image_dir.sort(key=natural_keys)\n            image = os.path.join(train_images_path, image_dir[slice_number])\n            with open(image, 'rb') as dicom_file:\n                dataset = dcmread(dicom_file)\n                img = dataset.pixel_array\n                img_tensor = tf.convert_to_tensor(img)\n                train_images.append(img_tensor)\n\n\n    return train_images, train_image_labels\n\n\n\ntrain_images, train_labels = create_train_dataset()`\n</code></pre>",
  "messages": [
    {
      "id": 1926972,
      "postDate": "2022-09-05T09:41:01.270Z",
      "content": "<pre><code>train_images.append(img_tensor)\n</code></pre>\n<p>It should be OOM here. I think that it is necessary to load images only for the mini batch.</p>",
      "rawMarkdown": "```\ntrain_images.append(img_tensor)\n```\n\nIt should be OOM here. I think that it is necessary to load images only for the mini batch.",
      "votes": 1,
      "replies": [
        {
          "id": 1927479,
          "postDate": "2022-09-05T16:41:17.327Z",
          "content": "<p>Thank you so much.  I really appreciate the incite.</p>",
          "rawMarkdown": "Thank you so much.  I really appreciate the incite.\n"
        }
      ]
    },
    {
      "id": 1926556,
      "postDate": "2022-09-05T00:17:21.407Z",
      "content": "<p>I have a very simple section of my notebook where I am simply trying to process data and create two lists filled with tensors for the image data and for the labels.  I have not even begun to train anything and my notebook  throws an error saying </p>\n<blockquote>\n  <p>\"your Notebook tried to allocate more memory than available. It has been restarted\"</p>\n</blockquote>\n<p>This is kind of confusing to me as I do not feel like I am writing that much data.   I have researched this error and there are some tips for reducing memory, but they seem to be with pandas data frames and I am not reading any giant data frames. Do I need to compress my images or do I need to delete things that I am not?  Thanks<br>\nBelow is my code.  </p>\n<pre><code>`!pip install -qU ../input/for-pydicom/python_gdcm-3.0.14-cp37-cp37m-manylinux_2_17_x86_64.manylinux2014_x86_64.whl \n!pip install ../input/pylibjpeg-whl/pylibjpeg-1.4.0-py3-none-any.whl --find-links frozen_packages --no-index\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nfrom pydicom import dcmread\n# from pydicom.pixel_data_handlers.util import apply_voi_lut\nimport nibabel as nib\nimport os\nos.environ['TF_CPP_MIN_LOG_LEVEL'] = '3' \nimport tensorflow as tf\nimport matplotlib.pyplot as plt\nimport tensorflow as tf\nimport re\n\nCURRENT_DIR_PATH = \"./\"\nTRAIN_IMAGE_PATH = \"../input/rsna-2022-cervical-spine-fracture-detection/train_images\"\nSEGMENTATIONS_PATH = \"../input/rsna-2022-cervical-spine-fracture-detection/segmentations\"\nTRAIN_CSV_PATH = \"../input/rsna-2022-cervical-spine-fracture-detection/train.csv\"\ntrain_data = pd.read_csv(TRAIN_CSV_PATH)\n\ndef atoi(text):\n    return int(text) if text.isdigit() else text\ndef natural_keys(text):\n    return [atoi(c) for c in re.split(r'(\\d+)', text)]\n\ndef labels_for_image(vertebrae):\n    vertebrae_for_image = list(filter(lambda vertebra: vertebra &gt; 0.0, vertebrae))\n    return vertebrae_for_image if len(vertebrae_for_image) &gt; 0 else [0]\n\ndef create_train_dataset():\n    segmentation_patient_files = os.listdir(SEGMENTATIONS_PATH)\n    train_images = []\n    train_image_labels = []\n    for segmentation_patient_file in segmentation_patient_files:\n        file_path = os.path.join(SEGMENTATIONS_PATH, segmentation_patient_file)\n        segmentation_file = nib.load(file_path).get_fdata()\n        segmentation_file_transposed = segmentation_file[:, ::-1, ::-1].transpose(2, 1, 0)\n\n        for slice_number in range(0, len(segmentation_file_transposed)):\n            dicom_slice = segmentation_file_transposed[slice_number]\n            patient_id = os.path.join(TRAIN_IMAGE_PATH, file_path.split(\"/\")[-1].replace('.nii', \"\")).split('/')[-1]\n            vertebrae = np.unique(dicom_slice)\n            labels = labels_for_image(vertebrae)\n            labels_tensor = tf.convert_to_tensor(labels)\n            train_image_labels.append(labels_tensor)\n            train_images_path = os.path.join(TRAIN_IMAGE_PATH, file_path.split(\"/\")[-1].replace('.nii', \"\"))\n\n            image_dir = os.listdir(train_images_path)\n            image_dir.sort(key=natural_keys)\n            image = os.path.join(train_images_path, image_dir[slice_number])\n            with open(image, 'rb') as dicom_file:\n                dataset = dcmread(dicom_file)\n                img = dataset.pixel_array\n                img_tensor = tf.convert_to_tensor(img)\n                train_images.append(img_tensor)\n\n\n    return train_images, train_image_labels\n\n\n\ntrain_images, train_labels = create_train_dataset()`\n</code></pre>",
      "rawMarkdown": "I have a very simple section of my notebook where I am simply trying to process data and create two lists filled with tensors for the image data and for the labels.  I have not even begun to train anything and my notebook  throws an error saying \n>\"your Notebook tried to allocate more memory than available. It has been restarted\"\n\nThis is kind of confusing to me as I do not feel like I am writing that much data.   I have researched this error and there are some tips for reducing memory, but they seem to be with pandas data frames and I am not reading any giant data frames. Do I need to compress my images or do I need to delete things that I am not?  Thanks\nBelow is my code.  \n\n```\n`!pip install -qU ../input/for-pydicom/python_gdcm-3.0.14-cp37-cp37m-manylinux_2_17_x86_64.manylinux2014_x86_64.whl \n!pip install ../input/pylibjpeg-whl/pylibjpeg-1.4.0-py3-none-any.whl --find-links frozen_packages --no-index\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nfrom pydicom import dcmread\n# from pydicom.pixel_data_handlers.util import apply_voi_lut\nimport nibabel as nib\nimport os\nos.environ['TF_CPP_MIN_LOG_LEVEL'] = '3' \nimport tensorflow as tf\nimport matplotlib.pyplot as plt\nimport tensorflow as tf\nimport re\n\nCURRENT_DIR_PATH = \"./\"\nTRAIN_IMAGE_PATH = \"../input/rsna-2022-cervical-spine-fracture-detection/train_images\"\nSEGMENTATIONS_PATH = \"../input/rsna-2022-cervical-spine-fracture-detection/segmentations\"\nTRAIN_CSV_PATH = \"../input/rsna-2022-cervical-spine-fracture-detection/train.csv\"\ntrain_data = pd.read_csv(TRAIN_CSV_PATH)\n\ndef atoi(text):\n    return int(text) if text.isdigit() else text\ndef natural_keys(text):\n    return [atoi(c) for c in re.split(r'(\\d+)', text)]\n\ndef labels_for_image(vertebrae):\n    vertebrae_for_image = list(filter(lambda vertebra: vertebra > 0.0, vertebrae))\n    return vertebrae_for_image if len(vertebrae_for_image) > 0 else [0]\n\ndef create_train_dataset():\n    segmentation_patient_files = os.listdir(SEGMENTATIONS_PATH)\n    train_images = []\n    train_image_labels = []\n    for segmentation_patient_file in segmentation_patient_files:\n        file_path = os.path.join(SEGMENTATIONS_PATH, segmentation_patient_file)\n        segmentation_file = nib.load(file_path).get_fdata()\n        segmentation_file_transposed = segmentation_file[:, ::-1, ::-1].transpose(2, 1, 0)\n        \n        for slice_number in range(0, len(segmentation_file_transposed)):\n            dicom_slice = segmentation_file_transposed[slice_number]\n            patient_id = os.path.join(TRAIN_IMAGE_PATH, file_path.split(\"/\")[-1].replace('.nii', \"\")).split('/')[-1]\n            vertebrae = np.unique(dicom_slice)\n            labels = labels_for_image(vertebrae)\n            labels_tensor = tf.convert_to_tensor(labels)\n            train_image_labels.append(labels_tensor)\n            train_images_path = os.path.join(TRAIN_IMAGE_PATH, file_path.split(\"/\")[-1].replace('.nii', \"\"))\n            \n            image_dir = os.listdir(train_images_path)\n            image_dir.sort(key=natural_keys)\n            image = os.path.join(train_images_path, image_dir[slice_number])\n            with open(image, 'rb') as dicom_file:\n                dataset = dcmread(dicom_file)\n                img = dataset.pixel_array\n                img_tensor = tf.convert_to_tensor(img)\n                train_images.append(img_tensor)\n                \n    \n    return train_images, train_image_labels\n            \n\n\ntrain_images, train_labels = create_train_dataset()`\n```"
    }
  ],
  "comments": [
    {
      "id": 1926972,
      "author_name": "patriot",
      "author_url": "",
      "post_date": "2022-09-05T09:41:01.270000",
      "content": "<pre><code>train_images.append(img_tensor)\n</code></pre>\n<p>It should be OOM here. I think that it is necessary to load images only for the mini batch.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1927479,
          "author_name": "cbarb15",
          "author_url": "",
          "post_date": "2022-09-05T16:41:17.327000",
          "content": "<p>Thank you so much.  I really appreciate the incite.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1926972": "```\ntrain_images.append(img_tensor)\n```\n\nIt should be OOM here. I think that it is necessary to load images only for the mini batch.",
    "1926556": "I have a very simple section of my notebook where I am simply trying to process data and create two lists filled with tensors for the image data and for the labels.  I have not even begun to train anything and my notebook  throws an error saying \n>\"your Notebook tried to allocate more memory than available. It has been restarted\"\n\nThis is kind of confusing to me as I do not feel like I am writing that much data.   I have researched this error and there are some tips for reducing memory, but they seem to be with pandas data frames and I am not reading any giant data frames. Do I need to compress my images or do I need to delete things that I am not?  Thanks\nBelow is my code.  \n\n```\n`!pip install -qU ../input/for-pydicom/python_gdcm-3.0.14-cp37-cp37m-manylinux_2_17_x86_64.manylinux2014_x86_64.whl \n!pip install ../input/pylibjpeg-whl/pylibjpeg-1.4.0-py3-none-any.whl --find-links frozen_packages --no-index\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nfrom pydicom import dcmread\n# from pydicom.pixel_data_handlers.util import apply_voi_lut\nimport nibabel as nib\nimport os\nos.environ['TF_CPP_MIN_LOG_LEVEL'] = '3' \nimport tensorflow as tf\nimport matplotlib.pyplot as plt\nimport tensorflow as tf\nimport re\n\nCURRENT_DIR_PATH = \"./\"\nTRAIN_IMAGE_PATH = \"../input/rsna-2022-cervical-spine-fracture-detection/train_images\"\nSEGMENTATIONS_PATH = \"../input/rsna-2022-cervical-spine-fracture-detection/segmentations\"\nTRAIN_CSV_PATH = \"../input/rsna-2022-cervical-spine-fracture-detection/train.csv\"\ntrain_data = pd.read_csv(TRAIN_CSV_PATH)\n\ndef atoi(text):\n    return int(text) if text.isdigit() else text\ndef natural_keys(text):\n    return [atoi(c) for c in re.split(r'(\\d+)', text)]\n\ndef labels_for_image(vertebrae):\n    vertebrae_for_image = list(filter(lambda vertebra: vertebra > 0.0, vertebrae))\n    return vertebrae_for_image if len(vertebrae_for_image) > 0 else [0]\n\ndef create_train_dataset():\n    segmentation_patient_files = os.listdir(SEGMENTATIONS_PATH)\n    train_images = []\n    train_image_labels = []\n    for segmentation_patient_file in segmentation_patient_files:\n        file_path = os.path.join(SEGMENTATIONS_PATH, segmentation_patient_file)\n        segmentation_file = nib.load(file_path).get_fdata()\n        segmentation_file_transposed = segmentation_file[:, ::-1, ::-1].transpose(2, 1, 0)\n        \n        for slice_number in range(0, len(segmentation_file_transposed)):\n            dicom_slice = segmentation_file_transposed[slice_number]\n            patient_id = os.path.join(TRAIN_IMAGE_PATH, file_path.split(\"/\")[-1].replace('.nii', \"\")).split('/')[-1]\n            vertebrae = np.unique(dicom_slice)\n            labels = labels_for_image(vertebrae)\n            labels_tensor = tf.convert_to_tensor(labels)\n            train_image_labels.append(labels_tensor)\n            train_images_path = os.path.join(TRAIN_IMAGE_PATH, file_path.split(\"/\")[-1].replace('.nii', \"\"))\n            \n            image_dir = os.listdir(train_images_path)\n            image_dir.sort(key=natural_keys)\n            image = os.path.join(train_images_path, image_dir[slice_number])\n            with open(image, 'rb') as dicom_file:\n                dataset = dcmread(dicom_file)\n                img = dataset.pixel_array\n                img_tensor = tf.convert_to_tensor(img)\n                train_images.append(img_tensor)\n                \n    \n    return train_images, train_image_labels\n            \n\n\ntrain_images, train_labels = create_train_dataset()`\n```"
  }
}