{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<p id=\"part0\"></p>\n​\n<p style=\"font-family: Arials; line-height: 2; font-size: 44px; font-weight: bold; letter-spacing: 0px; text-align: center; color: #FF8C00\">This program, first crops dicom images, and then turns them into png images</p>\n​\n<img src=\"https://th.bing.com/th/id/OIP.ZZ5i8bsrLufsWzXxAxYpDwHaD4?pid=ImgDet&rs=1\" width=\"90%\" align=\"center\" hspace=\"20%\" vspace=\"5%\"/>\n​\n<p style=\"font-family: Arials; font-size: 40px; font-style: normal; font-weight: bold; letter-spacing: -2px; color: #000000; line-height:2.0\">Content:</p>\n​\n<p style=\"font-family: Arials; font-size: 18px; font-style: normal; font-weight: bold; letter-spacing: 0px; color: #000000; line-height:.5\"><a href=\"#part1\" style=\"color:#000000\"> 1- Importing libraries</a></p>\n\n<p style=\"font-family: Arials; font-size: 18px; font-style: normal; font-weight: bold; letter-spacing: 0px; color: #000000; line-height:.5\"><a href=\"#part2\" style=\"color:#000000\"> 2-Describing dicom files!</a></p>\n\n<p style=\"text-indent: 1vw; font-family: Arials; font-size: 16px; font-style: normal; font-weight: bold; letter-spacing: 0px; color: #000000; line-height:0.5\">\n<a href=\"#part2-1\" style=\"color:#000000\">  2-1 Reading dicom files</a></p>\n\n<p style=\"text-indent: 1vw; font-family: Arials; font-size: 16px; font-style: normal; font-weight: bold; letter-spacing: 0px; color: #000000; line-height:0.5\">\n<a href=\"#part2-2\" style=\"color:#000000\">2-2 Cropping dicom files </a></p>\n\n<p style=\"text-indent: 1vw; font-family: Arials; font-size: 16px; font-style: normal; font-weight: bold; letter-spacing: 0px; color: #000000; line-height:0.5\">\n<a href=\"#part2-3\" style=\"color:#000000\">2-3 Comparing original images with cropped ones </a></p>\n\n<p style=\"text-indent: 1vw; font-family: Arials; font-size: 16px; font-style: normal; font-weight: bold; letter-spacing: 0px; color: #000000; line-height:0.5\">\n<a href=\"#part2-4\" style=\"color:#000000\">2-4 Turning dicom images to png ones </a></p>\n","metadata":{}},{"cell_type":"markdown","source":"<p id=\"part1\"></p>\n\n<p style=\"font-family: Times New Roman; font-size: 20px; font-style: bold; font-weight: bold; letter-spacing: 0px; color: #0000FF\">1- Importing Libraries<p>\n<hr style=\"height: 1px; border: 1; background-color: #0000FF\">","metadata":{}},{"cell_type":"code","source":"# For better performance in turning dicom to png files We uysed this command\n!pip install -qU python-gdcm pydicom pylibjpeg ","metadata":{"execution":{"iopub.status.busy":"2023-10-03T16:01:46.789966Z","iopub.execute_input":"2023-10-03T16:01:46.790534Z","iopub.status.idle":"2023-10-03T16:01:56.024818Z","shell.execute_reply.started":"2023-10-03T16:01:46.790401Z","shell.execute_reply":"2023-10-03T16:01:56.023621Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport pydicom, cv2,glob,os,time,gdcm\nfrom matplotlib import pyplot as plt\nfrom tqdm.notebook import tqdm\nfrom joblib import Parallel, delayed","metadata":{"execution":{"iopub.status.busy":"2023-10-03T16:01:56.027693Z","iopub.execute_input":"2023-10-03T16:01:56.028137Z","iopub.status.idle":"2023-10-03T16:01:56.246383Z","shell.execute_reply.started":"2023-10-03T16:01:56.028093Z","shell.execute_reply":"2023-10-03T16:01:56.245417Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<p id=\"part2\"></p>\n\n<p style=\"font-family: Times New Roman; font-size: 20px; font-style: bold; font-weight: bold; letter-spacing: 0px; color: #0000FF\">2- Describing dicom files!</p>\n<hr style=\"height: 1px; border: 1; background-color: #0000FF\">","metadata":{}},{"cell_type":"markdown","source":"\n\n**DICOM** is an acronym for Digital Imaging and Communications in Medicine. Files in this format are most likely saved with either a DCM or DCM30 (DICOM 3.0) file extension, but some may not have an extension at all.\n\nDICOM is both a communications protocol and a file format, which means it can store medical information, such as ultrasound and MRI images, along with a patient's information, all in one file. The format ensures that all the data stays together, as well as provides the ability to transfer said information between devices that support the DICOM format.\n\n**What package is used for this kind of file in python?**\npydicom is the library you can use in order to open dcm files and manupulate them.Actually, pydicom is a pure Python package for working with DICOM files such as medical images, reports, and radiotherapy objects. For more information , click this [link](https://pydicom.github.io/)","metadata":{}},{"cell_type":"markdown","source":"<p id=\"part2-1\"></p>\n\n<p style=\"font-family: Times New Roman; font-size: 20px; font-style: bold; font-weight: bold; letter-spacing: 0px; color: #0000FF\">2-1 Reading dicom files</p>\n<hr style=\"height: 1px; border: 1; background-color: #0000FF\">","metadata":{}},{"cell_type":"markdown","source":"# Before running the following code, please add rsna-breast-cancer dataset by clicking + on the right page","metadata":{}},{"cell_type":"code","source":"dataset_dir= '/kaggle/input/rsna-breast-cancer-detection/'\ntrain_images_dir = sorted(glob.glob(f'{dataset_dir}train_images/*/*.dcm'))\n\nnumber_of_iamges = len(train_images_dir)\nprint('Number of train images is:',number_of_iamges) # You are xepceted to see:54706","metadata":{"execution":{"iopub.status.busy":"2023-10-03T16:01:56.24788Z","iopub.execute_input":"2023-10-03T16:01:56.248264Z","iopub.status.idle":"2023-10-03T16:02:18.006093Z","shell.execute_reply.started":"2023-10-03T16:01:56.248225Z","shell.execute_reply":"2023-10-03T16:02:18.004998Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\ndef read_dicom(dicom_file):\n    '''This function takes image with dicom format and reads it, then returns image's pixel as an output'''\n    \n    dicom_reader = pydicom.dcmread(dicom_file) # it reads dciom file\n    dicom_pixel = dicom_reader.pixel_array # it collects the pixel\n\n    dicom_pixel = (dicom_pixel - dicom_pixel.min()) / (dicom_pixel.max() - dicom_pixel.min()) # Normalization process\n\n    if dicom_reader.PhotometricInterpretation == \"MONOCHROME1\":\n        dicom_pixel = 1 - dicom_pixel # a special kind of format in dicom files, you can search it on the net\n    \n   \n    return dicom_pixel","metadata":{"execution":{"iopub.status.busy":"2023-10-03T16:02:18.007429Z","iopub.execute_input":"2023-10-03T16:02:18.00841Z","iopub.status.idle":"2023-10-03T16:02:18.015015Z","shell.execute_reply.started":"2023-10-03T16:02:18.008367Z","shell.execute_reply":"2023-10-03T16:02:18.014062Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<p id=\"part2-2\"></p>\n\n<p style=\"font-family: Times New Roman; font-size: 20px; font-style: bold; font-weight: bold; letter-spacing: 0px; color: #0000FF\">2-2 Crop dicom file</p>\n<hr style=\"height: 1px; border: 1; background-color: #0000FF\">","metadata":{}},{"cell_type":"markdown","source":"Define a function to crop dicom file, sine our images have lots of unuseful pixels, using cv2.connectedComponentsWithStats\n\ncheck this [link](https://stackoverflow.com/questions/35854197/how-to-use-opencvs-connectedcomponentswithstats-in-python) for more detail","metadata":{}},{"cell_type":"code","source":"def crop_dicom(dicom_file):\n    '''This function takes dicom image and removes additional space/pixels of it based on the logic behind connectedComponentsWithStats method\n    ,then it returns a cropped dicom file'''\n    \n    result= cv2.connectedComponentsWithStats((dicom_file > 0.005).astype(np.uint8)[:, :], 8, cv2.CV_32S)\n    # result has 4 diffrenet outputs. You can see different outputs here.\n    # But for reconizing the black color(0) from others(1), we just need result[2] , called stat matrix\n    \n    stat_matrix = result[2] # include left, top, width, height, area_size columns\n    # the second row shows the boxes of pixels which are different from background.Also, we don't need are_size cloumn. So:\n    \n    second_row = stat_matrix[1:, 4].argmax() + 1\n    x1, y1, w, h = stat_matrix[second_row][:4]\n    x2 = x1 + w\n    y2 = y1 + h\n    cropped_dicom = dicom_file[y1: y2, x1: x2]\n    \n    return cropped_dicom","metadata":{"execution":{"iopub.status.busy":"2023-10-03T16:02:18.017576Z","iopub.execute_input":"2023-10-03T16:02:18.018743Z","iopub.status.idle":"2023-10-03T16:02:18.028827Z","shell.execute_reply.started":"2023-10-03T16:02:18.018705Z","shell.execute_reply":"2023-10-03T16:02:18.027968Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<p id=\"part2-3\"></p>\n\n<p style=\"font-family: Times New Roman; font-size: 20px; font-style: bold; font-weight: bold; letter-spacing: 0px; color: #0000FF\">2-3 Comparing original images with cropped ones</p>\n<hr style=\"height: 1px; border: 1; background-color: #0000FF\">","metadata":{}},{"cell_type":"code","source":"test_images = train_images_dir[:2]\n\nfor image in test_images:\n    patient_id = image.split('/')[-2] # split the patient_id from the whole address\n    image_id = image.split('/')[-1][:-4] # split the image_id from the whole address\n    \n\n    # Plot Orginial iamge\n    plt.figure(figsize=(8,8))\n    dicom_pixel = read_dicom(dicom_file=image)\n    plt.imshow(dicom_pixel,cmap='gray')\n    plt.title(f\"Original dicom image-->image_id:{image_id}\",fontsize=16)\n    plt.show()\n    \n    # Plot cropped image\n    plt.figure(figsize=(8,8))\n    cropped_iamge = crop_dicom(dicom_file=dicom_pixel)\n    plt.imshow(cropped_iamge,cmap='gray')\n    plt.title(f\"Cropped dicom image-->image_id:{image_id}\",fontsize=16)\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-03T16:02:18.030144Z","iopub.execute_input":"2023-10-03T16:02:18.031101Z","iopub.status.idle":"2023-10-03T16:02:28.851416Z","shell.execute_reply.started":"2023-10-03T16:02:18.031064Z","shell.execute_reply":"2023-10-03T16:02:28.85046Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<p id=\"part2-4\"></p>\n\n<p style=\"font-family: Times New Roman; font-size: 20px; font-style: bold; font-weight: bold; letter-spacing: 0px; color: #0000FF\">2-4 Turning dicom images  into png images</p>\n<hr style=\"height: 1px; border: 1; background-color: #0000FF\">","metadata":{}},{"cell_type":"code","source":"# Creating two folders called 'Cancer' and 'No_Cancer' for categorizing images into two groups.\n\nSAVE_FOLDER_CANCER = \"/kaggle/working/Cancer/\"\nos.makedirs(SAVE_FOLDER_CANCER, exist_ok=True)\n\nSAVE_FOLDER_NO_CANCER = \"/kaggle/working/NO_Cancer/\"\nos.makedirs(SAVE_FOLDER_NO_CANCER, exist_ok=True)","metadata":{"execution":{"iopub.status.busy":"2023-10-03T16:02:28.852783Z","iopub.execute_input":"2023-10-03T16:02:28.853656Z","iopub.status.idle":"2023-10-03T16:02:28.859575Z","shell.execute_reply.started":"2023-10-03T16:02:28.853608Z","shell.execute_reply":"2023-10-03T16:02:28.858317Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# In order to turn the dicom images into png files, firstly we need to have train_csv file to match images with the corresponding info.\ntrain_df_dir = \"/kaggle/input/rsna-breast-cancer-detection/train.csv\"\ntrain_df = pd.read_csv(train_df_dir)\ntrain_df[['patient_id','image_id','site_id','machine_id']] = train_df[['patient_id','image_id','site_id','machine_id']].astype(str)\n\n# Creating the function\ndef dcm_to_png(file,size=256,save_folder=\"/kaggle/working/\", extension ='png'):\n    '''This function takes the dicom file, reads it using read_dicom function, crops it using crop_dicom function, then resizes it to 512x512 \n    and then turns it to png format. Finally it saves images inside two created folders'''\n    \n    patient_id = file.split('/')[-2]\n    image_id = file.split('/')[-1][:-4]\n    dicom_pixel = read_dicom(dicom_file=file) \n    cropped_dicom = crop_dicom(dicom_file=dicom_pixel) \n    resized_img = cv2.resize(cropped_dicom,dsize=(size,size)) \n    \n    loc_in_train_df = np.where(train_df.image_id==image_id)\n    id_cancer = train_df.iloc[loc_in_train_df].cancer\n    if id_cancer.all():\n        cv2.imwrite(save_folder +\"Cancer/\" + f\"{patient_id}_{image_id}.{extension}\", (resized_img*255).astype(np.uint8))\n    else:\n        cv2.imwrite(save_folder +\"NO_Cancer/\" + f\"{patient_id}_{image_id}.{extension}\", (resized_img*255).astype(np.uint8))","metadata":{"execution":{"iopub.status.busy":"2023-10-03T16:02:28.861147Z","iopub.execute_input":"2023-10-03T16:02:28.86171Z","iopub.status.idle":"2023-10-03T16:02:29.052662Z","shell.execute_reply.started":"2023-10-03T16:02:28.861674Z","shell.execute_reply":"2023-10-03T16:02:29.051598Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Calling the dcm_to_png function\n# if you want to test just some files, please use train_images[:10] instead of the whole train_images_dir\n\n# Using parallel function to make process faster\nstart = time.time()\n_ = Parallel(n_jobs=4)(\n    delayed(dcm_to_png)(file, size=256)\n    for file in tqdm(train_images_dir[:30000])  # you can use train_images for all iamges\n)\nstop = time.time()\nprint(f'This process took you:{stop-start} seconds')\nprint('Now, you can have access to cropped files wtih png format')","metadata":{"execution":{"iopub.status.busy":"2023-10-03T16:02:29.054352Z","iopub.execute_input":"2023-10-03T16:02:29.055058Z","iopub.status.idle":"2023-10-03T16:05:47.890793Z","shell.execute_reply.started":"2023-10-03T16:02:29.055018Z","shell.execute_reply":"2023-10-03T16:05:47.889464Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}