{
  "id": 375119,
  "title": "Looking for a faster way to extract data into image files.",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/375119",
  "author_name": "Nizar Haytham ",
  "post_date": "2022-12-30T12:07:36.156000",
  "votes": 0,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hi,<br>\nActually I have written code to loop through folders and extract png images from the dcm files into a folder of two folders, one is named ''0'' and the other is ''1'' corresponding to the labels (I checked the labels from the csv file of course) , but I still have a problem which is that it extracts really slowly. I hope if anyone have a faster way.</p>\n<p>`i = -1<br>\nnum_fails = 0</p>\n<p>for patient in list(df_train['patient_id']) :<br>\n    list_of_files = os.listdir('/kaggle/input/rsna-breast-cancer-detection/train_images/'+str(patient))<br>\n    for file in list_of_files :<br>\n        file = file.replace('.dcm',\"\")<br>\n        label = df_train.loc[int(file) , 'cancer']<br>\n        i = i + 1 <br>\n        if i % 50 == 0 :<br>\n            print('image number {}'.format(str(i)))</p>\n<pre><code>    if label == 0 :\n        try :\n            im = pydicom.dcmread('/kaggle/input/rsna-breast-cancer-detection/train_images/'+str(patient)+'/'+str(file)+'.dcm')\n\n            im = im.pixel_array.astype(float) # Extract the pixel array from im\n            # print(im)\n            # print(np.maximum(im, 0))\n            # print(im.max())\n            rescaled_image = (np.maximum(im, 0) / im.max())*255 # Rescaling the image, put their values between 0 and 255\n            final_image = np.uint8(rescaled_image) # Convert int\n\n            final_image = Image.fromarray(final_image) # Creater image from an array\n            final_image = final_image.resize((550,550))\n            final_image.save('/kaggle/working/ReadyData/0/'+str(i)+'.png')\n        except :\n            num_fails = num_fails + 1\n            print('a class 0 fail occured , total num_fails is now : {}'.format(num_fails))\n\n\n\n\n\n    if label == 1 :\n        try :\n            im = pydicom.dcmread('/kaggle/input/rsna-breast-cancer-detection/train_images/'+str(patient)+'/'+str(file)+'.dcm')\n\n            im = im.pixel_array.astype(float) # Extract the pixel array from im\n            # print(im)\n            # print(np.maximum(im, 0))\n            # print(im.max())\n            rescaled_image = (np.maximum(im, 0) / im.max())*255 # Rescaling the image, put their values between 0 and 255\n            final_image = np.uint8(rescaled_image) # Convert int\n\n            final_image = Image.fromarray(final_image) # Creater image from an array\n            final_image = final_image.resize((550,550))\n            final_image.save('/kaggle/working/ReadyData/1/'+str(i)+'.png')\n        except :\n            num_fails = num_fails + 1\n            print('a class 1 fail occured, total num fails is now : {}'.format(num_fails))`\n</code></pre>",
  "messages": [
    {
      "id": 2080821,
      "postDate": "2022-12-30T14:16:21.563Z",
      "content": "<p>Check the code tab, folks have shared a number of different solutions.  Using a combination of dicomsdl library + joblib is fairly proven to work.  There are some experiments with dali and nvjpeg2000 library, but I don't know how proven that is.</p>",
      "rawMarkdown": "Check the code tab, folks have shared a number of different solutions.  Using a combination of dicomsdl library + joblib is fairly proven to work.  There are some experiments with dali and nvjpeg2000 library, but I don't know how proven that is.\n\n",
      "votes": 1
    },
    {
      "id": 2080708,
      "postDate": "2022-12-30T12:07:36.157Z",
      "content": "<p>Hi,<br>\nActually I have written code to loop through folders and extract png images from the dcm files into a folder of two folders, one is named ''0'' and the other is ''1'' corresponding to the labels (I checked the labels from the csv file of course) , but I still have a problem which is that it extracts really slowly. I hope if anyone have a faster way.</p>\n<p>`i = -1<br>\nnum_fails = 0</p>\n<p>for patient in list(df_train['patient_id']) :<br>\n    list_of_files = os.listdir('/kaggle/input/rsna-breast-cancer-detection/train_images/'+str(patient))<br>\n    for file in list_of_files :<br>\n        file = file.replace('.dcm',\"\")<br>\n        label = df_train.loc[int(file) , 'cancer']<br>\n        i = i + 1 <br>\n        if i % 50 == 0 :<br>\n            print('image number {}'.format(str(i)))</p>\n<pre><code>    if label == 0 :\n        try :\n            im = pydicom.dcmread('/kaggle/input/rsna-breast-cancer-detection/train_images/'+str(patient)+'/'+str(file)+'.dcm')\n\n            im = im.pixel_array.astype(float) # Extract the pixel array from im\n            # print(im)\n            # print(np.maximum(im, 0))\n            # print(im.max())\n            rescaled_image = (np.maximum(im, 0) / im.max())*255 # Rescaling the image, put their values between 0 and 255\n            final_image = np.uint8(rescaled_image) # Convert int\n\n            final_image = Image.fromarray(final_image) # Creater image from an array\n            final_image = final_image.resize((550,550))\n            final_image.save('/kaggle/working/ReadyData/0/'+str(i)+'.png')\n        except :\n            num_fails = num_fails + 1\n            print('a class 0 fail occured , total num_fails is now : {}'.format(num_fails))\n\n\n\n\n\n    if label == 1 :\n        try :\n            im = pydicom.dcmread('/kaggle/input/rsna-breast-cancer-detection/train_images/'+str(patient)+'/'+str(file)+'.dcm')\n\n            im = im.pixel_array.astype(float) # Extract the pixel array from im\n            # print(im)\n            # print(np.maximum(im, 0))\n            # print(im.max())\n            rescaled_image = (np.maximum(im, 0) / im.max())*255 # Rescaling the image, put their values between 0 and 255\n            final_image = np.uint8(rescaled_image) # Convert int\n\n            final_image = Image.fromarray(final_image) # Creater image from an array\n            final_image = final_image.resize((550,550))\n            final_image.save('/kaggle/working/ReadyData/1/'+str(i)+'.png')\n        except :\n            num_fails = num_fails + 1\n            print('a class 1 fail occured, total num fails is now : {}'.format(num_fails))`\n</code></pre>",
      "rawMarkdown": "Hi,\nActually I have written code to loop through folders and extract png images from the dcm files into a folder of two folders, one is named ''0'' and the other is ''1'' corresponding to the labels (I checked the labels from the csv file of course) , but I still have a problem which is that it extracts really slowly. I hope if anyone have a faster way.\n\n\n`i = -1\nnum_fails = 0\n\nfor patient in list(df_train['patient_id']) :\n    list_of_files = os.listdir('/kaggle/input/rsna-breast-cancer-detection/train_images/'+str(patient))\n    for file in list_of_files :\n        file = file.replace('.dcm',\"\")\n        label = df_train.loc[int(file) , 'cancer']\n        i = i + 1 \n        if i % 50 == 0 :\n            print('image number {}'.format(str(i)))\n        \n        \n        \n        if label == 0 :\n            try :\n                im = pydicom.dcmread('/kaggle/input/rsna-breast-cancer-detection/train_images/'+str(patient)+'/'+str(file)+'.dcm')\n\n                im = im.pixel_array.astype(float) # Extract the pixel array from im\n                # print(im)\n                # print(np.maximum(im, 0))\n                # print(im.max())\n                rescaled_image = (np.maximum(im, 0) / im.max())*255 # Rescaling the image, put their values between 0 and 255\n                final_image = np.uint8(rescaled_image) # Convert int\n\n                final_image = Image.fromarray(final_image) # Creater image from an array\n                final_image = final_image.resize((550,550))\n                final_image.save('/kaggle/working/ReadyData/0/'+str(i)+'.png')\n            except :\n                num_fails = num_fails + 1\n                print('a class 0 fail occured , total num_fails is now : {}'.format(num_fails))\n            \n            \n            \n            \n            \n        if label == 1 :\n            try :\n                im = pydicom.dcmread('/kaggle/input/rsna-breast-cancer-detection/train_images/'+str(patient)+'/'+str(file)+'.dcm')\n\n                im = im.pixel_array.astype(float) # Extract the pixel array from im\n                # print(im)\n                # print(np.maximum(im, 0))\n                # print(im.max())\n                rescaled_image = (np.maximum(im, 0) / im.max())*255 # Rescaling the image, put their values between 0 and 255\n                final_image = np.uint8(rescaled_image) # Convert int\n\n                final_image = Image.fromarray(final_image) # Creater image from an array\n                final_image = final_image.resize((550,550))\n                final_image.save('/kaggle/working/ReadyData/1/'+str(i)+'.png')\n            except :\n                num_fails = num_fails + 1\n                print('a class 1 fail occured, total num fails is now : {}'.format(num_fails))`"
    }
  ],
  "comments": [
    {
      "id": 2080821,
      "author_name": "@kaggleqrdl",
      "author_url": "",
      "post_date": "2022-12-30T14:16:21.563000",
      "content": "<p>Check the code tab, folks have shared a number of different solutions.  Using a combination of dicomsdl library + joblib is fairly proven to work.  There are some experiments with dali and nvjpeg2000 library, but I don't know how proven that is.</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2080821": "Check the code tab, folks have shared a number of different solutions.  Using a combination of dicomsdl library + joblib is fairly proven to work.  There are some experiments with dali and nvjpeg2000 library, but I don't know how proven that is.\n\n",
    "2080708": "Hi,\nActually I have written code to loop through folders and extract png images from the dcm files into a folder of two folders, one is named ''0'' and the other is ''1'' corresponding to the labels (I checked the labels from the csv file of course) , but I still have a problem which is that it extracts really slowly. I hope if anyone have a faster way.\n\n\n`i = -1\nnum_fails = 0\n\nfor patient in list(df_train['patient_id']) :\n    list_of_files = os.listdir('/kaggle/input/rsna-breast-cancer-detection/train_images/'+str(patient))\n    for file in list_of_files :\n        file = file.replace('.dcm',\"\")\n        label = df_train.loc[int(file) , 'cancer']\n        i = i + 1 \n        if i % 50 == 0 :\n            print('image number {}'.format(str(i)))\n        \n        \n        \n        if label == 0 :\n            try :\n                im = pydicom.dcmread('/kaggle/input/rsna-breast-cancer-detection/train_images/'+str(patient)+'/'+str(file)+'.dcm')\n\n                im = im.pixel_array.astype(float) # Extract the pixel array from im\n                # print(im)\n                # print(np.maximum(im, 0))\n                # print(im.max())\n                rescaled_image = (np.maximum(im, 0) / im.max())*255 # Rescaling the image, put their values between 0 and 255\n                final_image = np.uint8(rescaled_image) # Convert int\n\n                final_image = Image.fromarray(final_image) # Creater image from an array\n                final_image = final_image.resize((550,550))\n                final_image.save('/kaggle/working/ReadyData/0/'+str(i)+'.png')\n            except :\n                num_fails = num_fails + 1\n                print('a class 0 fail occured , total num_fails is now : {}'.format(num_fails))\n            \n            \n            \n            \n            \n        if label == 1 :\n            try :\n                im = pydicom.dcmread('/kaggle/input/rsna-breast-cancer-detection/train_images/'+str(patient)+'/'+str(file)+'.dcm')\n\n                im = im.pixel_array.astype(float) # Extract the pixel array from im\n                # print(im)\n                # print(np.maximum(im, 0))\n                # print(im.max())\n                rescaled_image = (np.maximum(im, 0) / im.max())*255 # Rescaling the image, put their values between 0 and 255\n                final_image = np.uint8(rescaled_image) # Convert int\n\n                final_image = Image.fromarray(final_image) # Creater image from an array\n                final_image = final_image.resize((550,550))\n                final_image.save('/kaggle/working/ReadyData/1/'+str(i)+'.png')\n            except :\n                num_fails = num_fails + 1\n                print('a class 1 fail occured, total num fails is now : {}'.format(num_fails))`"
  }
}