{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<h1 style=\"font-size: 35px;\">📓📘📗 Reading DICOM files tutorial 📒📔📕</h1>\n\n<div style=\"color:white;\n           display:fill;\n           border-radius:5px;\n           background-color:#5642C5;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:1px\">\n<p style=\"font-size: 16px; color:white; padding: 12px;\">\n    Reading the files in this competition can be challenging if you've never had to deal with them before. At least that's how I felt when I started working on this. By now, I feel I've gained some interesting insights on how this format works and I wanted to share it with all participants!\n    </p>\n</div>\n<div style=\"color:white;\n           display:fill;\n           border-radius:5px;\n           background-color:#5642C5;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:1px\">\n    <p style=\"font-size: 16px; color:white; padding: 12px;\">\n    If you like the content remember to leave a comment or upvote. Any suggestions are welcome and I would love your feedback in order to improve the notebook!\n    </p>\n</div>","metadata":{}},{"cell_type":"markdown","source":"<h1>Setting up the notebook imports</h1>\n\n<p style=\"font-size: 16px;\">The format we are going to introduce in this notebook is tricky. Mainly because it's not designed for computer vision, but for easy sharing among medical instutitions. Be mindful of leaving the installs as they are if you fork the notebook, as they gave me a huge headache. It seems that some versions of the packages do not work correctly unless you have all the compatible versions installed, as we are doing below.</p>\n\n<p style=\"font-size: 16px;\">As it was pointed out in <a href=\"https://www.kaggle.com/code/micheldc55/how-to-read-dcm-dicom-data/comments#2127464\">this</a> thread, the notebook is required to run without an internet connection. Due to this limitation, you can also download the .whl file from the pydicom library and the .whl file from the gdcm library. If you upload them to your \"input\" folder, you can pip install the wheels by running the following command:</p>\n\n```python\n! pip install kaggle.input.pydicom_file.whl\n! pip install kaggle.input.gcdm_file.whl\n```","metadata":{}},{"cell_type":"code","source":"! pip install -U pylibjpeg pylibjpeg-openjpeg pylibjpeg-libjpeg pydicom python-gdcm\n! pip install --upgrade pydicom","metadata":{"execution":{"iopub.status.busy":"2023-02-16T19:11:22.180802Z","iopub.execute_input":"2023-02-16T19:11:22.181295Z","iopub.status.idle":"2023-02-16T19:11:48.061953Z","shell.execute_reply.started":"2023-02-16T19:11:22.181194Z","shell.execute_reply":"2023-02-16T19:11:48.06027Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<p style=\"font-size: 16px;\">Using the process above, you can easily pip install the repositories while not actually connecting to the internet through the notebook. Below you can find a list where the documentation for each library is located:</p>\n\n<ul>\n    <li>PyDICOM: <a href=\"https://pypi.org/project/pydicom/\">Link to PyDICOM library documentation</a></li>\n    <li>Python GCDM: <a href=\"https://pypi.org/project/python-gdcm/\">Link to Python GCDM library documentation</a></li>\n    <li>Pylibjpeg: <a href=\"https://pypi.org/project/pylibjpeg-libjpeg/\">Link to Python pylibjpeg library documentation</a></li>\n</ul>","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport os\nimport torch\nfrom torchvision.io import read_image\nfrom sklearn.model_selection import train_test_split\nimport pydicom\nfrom pydicom.data import get_testdata_file\nimport matplotlib.pyplot as plt","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-02-01T10:23:54.682562Z","iopub.execute_input":"2023-02-01T10:23:54.683024Z","iopub.status.idle":"2023-02-01T10:23:56.645047Z","shell.execute_reply.started":"2023-02-01T10:23:54.68298Z","shell.execute_reply":"2023-02-01T10:23:56.644197Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1>Introducing the DICOM format</h1>","metadata":{}},{"cell_type":"markdown","source":"<img src=\"https://static.opswat.com/uploads/images/Deep-CDR-blog-01.jpg\" style=\"width:800px\">","metadata":{}},{"cell_type":"markdown","source":"<p style=\"font-size: 16px;\">According to Wikipedia, the .dcm format stands for: \"Digital Imaging and Communications in Medicine (DICOM) is the standard for the communication and management of medical imaging information and related data. DICOM is most commonly used for storing and transmitting medical images enabling the integration of medical imaging devices such as scanners, servers, workstations, printers, network hardware, and picture archiving and communication systems (PACS) from multiple manufacturers. It has been widely adopted by hospitals and is making inroads into smaller applications such as dentists' and doctors' offices.\" Link to web <a href=\"https://en.wikipedia.org/wiki/DICOM\">here</a>.</p>\n\n<p style=\"font-size: 16px;\"><b>DICOM</b> is a standard format used in medical imaging to store and exchange images, such as CT scans, MRIs, and X-rays. To read DICOM images in Python, the PyDICOM library can be used. PyDICOM provides easy-to-use interfaces to read, modify, and write these files. To read a DICOM file, simply import the PyDICOM library, open the file using the dcmread() method, and access the image data using the pixel_array attribute to get the data as a numpy array.</p>\n\n<p style=\"font-size: 16px;\">Medical institutions use the DICOM format for images for several reasons:\n\n<ul style=\"font-size: 16px;\">\n    <li><b>Standardization:</b> DICOM provides a standardized format for medical images, allowing for seamless exchange of data between different systems and institutions.</li>\n    <li><b>Interoperability:</b> The use of DICOM ensures that images can be easily shared and processed by different medical devices and software applications, reducing the risk of compatibility issues.</li>\n    <li><b>Metadata:</b> DICOM files contain extensive metadata, including patient information, imaging parameters, and acquisition details, which are important for patient care and research purposes.</li>\n    <li><b>Long-term archiving:</b> DICOM provides a means for long-term archiving and retrieval of medical images, ensuring that patient data is preserved for future use.</li>\n</ul>\n\n<p style=\"font-size: 16px;\">In summary, DICOM provides a standardized, interoperable, and high-quality format for medical images that is essential for patient care, research, and archiving. This format allows for ease-of-use between medical institutions by sharing the metadata of the image together with the image in a single file.</p>","metadata":{}},{"cell_type":"markdown","source":"<h1>Reading the data:</h1>\n\n<p style=\"font-size: 16px;\">We have <b>5 files</b> in the dataset folder: \n    <ul style=\"font-size: 16px;\">\n        <li><b>train_images:</b> Folder that contains the images on which to run the training. This are not split into train/validation so this needs to be done manually. Each folder in train_images corresponds to a single subject.</li>\n        <li><b>test_images:</b> Folder that contains the images on which to run the testing process. Each folder in test_images corresponds to a single subject.</li>\n        <li><b>sample_submission.csv:</b> Shows the format of a submission, which should only contain the subject_id and the probability of having cancer.</li>\n        <li><b>train.csv:</b> The train.csv file contains the subject id, the image id (this is necessary for building the image path when reading), laterality (which side is the mammography taken in), view (The orientation of the image. The default for a screening exam is to capture two views per breast.), age (subject age), cancer (the target), etc.</li>\n        <li><b>test.csv:</b> Same as the train.csv, contains the test subject information.</li>\n    </ul>","metadata":{}},{"cell_type":"code","source":"train_df = pd.read_csv('/kaggle/input/rsna-breast-cancer-detection/train.csv')\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-01T10:24:01.617297Z","iopub.execute_input":"2023-02-01T10:24:01.617849Z","iopub.status.idle":"2023-02-01T10:24:01.762149Z","shell.execute_reply.started":"2023-02-01T10:24:01.617808Z","shell.execute_reply":"2023-02-01T10:24:01.76092Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df = pd.read_csv('/kaggle/input/rsna-breast-cancer-detection/test.csv')\ntest_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-01-30T17:43:43.376136Z","iopub.execute_input":"2023-01-30T17:43:43.376603Z","iopub.status.idle":"2023-01-30T17:43:43.396099Z","shell.execute_reply.started":"2023-01-30T17:43:43.376545Z","shell.execute_reply":"2023-01-30T17:43:43.394857Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2>Understanding the data:</h2>\n\n<p style=\"font-size: 16px;\">The data has a few columns that may interest us. The first thing we have to realize is that the <b>patient_ids are not unique</b>. If we take a look at the train_images folder below, we will see that the images are grouped into folders <b>by patient</b>. Each patient corresponds to a single patient_id.</p>\n\n<p style=\"font-size: 16px;\">If we are going to use PyTorch of HuggingFace, we are going to need to manually create a reading pipeline for accessing the data, as the Dataset implementations generally expect the data to be grouped by label (istead of by patient). This can be done moving a few things around in the base dataset object from PyTorch.</p>\n\n<p style=\"font-size: 16px;\">If you take a look at the first 10 elements of the train_images folder, you will notice that the labels correspond to different patient ids. Inside those folders we will find different image ids that correspond to each image:</p>","metadata":{}},{"cell_type":"code","source":"os.listdir('/kaggle/input/rsna-breast-cancer-detection/train_images')[:10]","metadata":{"execution":{"iopub.status.busy":"2023-01-30T22:08:16.62602Z","iopub.execute_input":"2023-01-30T22:08:16.626423Z","iopub.status.idle":"2023-01-30T22:08:16.641638Z","shell.execute_reply.started":"2023-01-30T22:08:16.626391Z","shell.execute_reply":"2023-01-30T22:08:16.640353Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"patient_id = '21867'\n\nos.listdir(f'/kaggle/input/rsna-breast-cancer-detection/train_images/{patient_id}')","metadata":{"execution":{"iopub.status.busy":"2023-01-30T22:14:48.711007Z","iopub.execute_input":"2023-01-30T22:14:48.711432Z","iopub.status.idle":"2023-01-30T22:14:48.721943Z","shell.execute_reply.started":"2023-01-30T22:14:48.711397Z","shell.execute_reply":"2023-01-30T22:14:48.720838Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2>Interacting with .dcm files:</h2>\n\n<p style=\"font-size: 16px;\">Finally, we have reached the <b>DICOM files</b>. We can discuss how to read them and build them into a dataset later, but for now we should first understand how to interact with these types of files. As mentioned above, the .dcm file type allows medical institutions to easily share information and metadata together with the image. Let's explore that by looking at one image.</p>\n\n<p style=\"font-size: 16px;\">First, let's construct the image path to read a particular one that we are interested in. We will build everything from the dataframe, so that it's easy to understand how we will build the dataset later. Changing the \"idx\" value below will select the image at index=idx from the DataFrame.</p>","metadata":{}},{"cell_type":"code","source":"idx = 5  # 0\n\nbase_img_dir = '/kaggle/input/rsna-breast-cancer-detection/train_images'\n\nimg_id = str(train_df['image_id'].iloc[idx])\nfull_img_id = img_id + '.dcm'\npat_id = str(train_df['patient_id'].iloc[idx])\n\nlabel = train_df['cancer'].iloc[idx]\n\nos.path.join(base_img_dir, pat_id, full_img_id)","metadata":{"execution":{"iopub.status.busy":"2023-01-30T22:51:02.234591Z","iopub.execute_input":"2023-01-30T22:51:02.234983Z","iopub.status.idle":"2023-01-30T22:51:02.24525Z","shell.execute_reply.started":"2023-01-30T22:51:02.234953Z","shell.execute_reply":"2023-01-30T22:51:02.243786Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<p style=\"font-size: 16px;\">If we swith the values of \"idx\" around, we will get different paths. We are saving one path example into the \"img_path\" variable for testing, and the reader is encouraged to test it with other values of idx. We are going to use the first one (idx = 0) for simplicity.</p>","metadata":{}},{"cell_type":"code","source":"img_path = os.path.join(base_img_dir, pat_id, full_img_id)","metadata":{"execution":{"iopub.status.busy":"2023-01-30T22:51:02.751488Z","iopub.execute_input":"2023-01-30T22:51:02.751891Z","iopub.status.idle":"2023-01-30T22:51:02.758828Z","shell.execute_reply.started":"2023-01-30T22:51:02.751859Z","shell.execute_reply":"2023-01-30T22:51:02.757542Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2>Reading the image from the path and converting to tensor:</h2>\n\n<p style=\"font-size: 16px;\">This is self-explanatory. We are going to read the image from the img_path using the pydicom library we installed above. Important note about the library: It seems to have some issues with the versions, so I recommend to install the latest one and upgrade the auxiliary libraries (gdcm and pylibjpeg) as shown above.</p>\n\n<p style=\"font-size: 16px;\">Let's use the pydicom library to read the data from a single image and see what this format returns. Notice below that it's easy to find what we are looking for: <b>the Pixel Data</b> attibute contains the information we need.</p>","metadata":{}},{"cell_type":"code","source":"dcm_img = pydicom.dcmread(img_path, force=True)\ndcm_img","metadata":{"execution":{"iopub.status.busy":"2023-01-30T22:51:03.443739Z","iopub.execute_input":"2023-01-30T22:51:03.444142Z","iopub.status.idle":"2023-01-30T22:51:03.541083Z","shell.execute_reply.started":"2023-01-30T22:51:03.444108Z","shell.execute_reply":"2023-01-30T22:51:03.540254Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<p style=\"font-size: 16px;\">In order to visualize the data, we are going to convert the read file into a numpy array using the \"pixel_array\" method. Then, we are going to use matplotlib's imshow function to visualize the image and the label for convenience.</p>","metadata":{}},{"cell_type":"code","source":"img_array = dcm_img.pixel_array\n\nplt.imshow(img_array)\nif label == 0:\n    category = \"doesn't have cancer\"\nelif label == 1:\n    category = \"has cancer\"\n\nplt.title(f'Patient {pat_id} {category} in image: {img_id}');","metadata":{"execution":{"iopub.status.busy":"2023-01-30T22:51:12.409448Z","iopub.execute_input":"2023-01-30T22:51:12.409888Z","iopub.status.idle":"2023-01-30T22:51:13.258417Z","shell.execute_reply.started":"2023-01-30T22:51:12.409853Z","shell.execute_reply":"2023-01-30T22:51:13.257339Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def random_9img_sample(\n    df: pd.DataFrame, \n    random_state: int = 101, \n    base_img_dir = '/kaggle/input/rsna-breast-cancer-detection/train_images'\n):\n    \"\"\"Function that shows 9 random images with labels and their img sizes\"\"\"\n    df_sample = df.sample(9, random_state=random_state).reset_index()\n    \n    fig, ax = plt.subplots(3, 3, figsize=(14, 14))\n    \n    for idx_num, row in df_sample.iterrows():\n        img_id = str(row['image_id'])\n        full_img_id = img_id + '.dcm'\n        pat_id = str(row['patient_id'])\n\n        label = row['cancer']\n\n        img_path = os.path.join(base_img_dir, pat_id, full_img_id)\n        q, r = divmod(idx_num, 3)\n        \n        img_array = pydicom.dcmread(img_path, force=True).pixel_array\n        \n        \n        ax[q][r].imshow(img_array)\n        ax[q][r].set_title(f'Image label is: {label}')","metadata":{"execution":{"iopub.status.busy":"2023-01-30T23:11:09.175272Z","iopub.execute_input":"2023-01-30T23:11:09.1757Z","iopub.status.idle":"2023-01-30T23:11:09.185139Z","shell.execute_reply.started":"2023-01-30T23:11:09.175668Z","shell.execute_reply":"2023-01-30T23:11:09.183746Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"random_9img_sample(train_df)","metadata":{"execution":{"iopub.status.busy":"2023-01-30T23:11:09.476054Z","iopub.execute_input":"2023-01-30T23:11:09.477092Z","iopub.status.idle":"2023-01-30T23:11:20.97188Z","shell.execute_reply.started":"2023-01-30T23:11:09.47705Z","shell.execute_reply":"2023-01-30T23:11:20.970636Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1>Conclusions</h1>\n\n<p style=\"font-size: 16px;\">And that's all for reading the DICOM files. There are a lot of interesting things that can be explored here, normalizing the data, rescaling, performing data augmentation, etc. Also note the imbalance of the dataset, as all images came out to be negative. If you are interested in building a PyTorch dataset using the concepts used above, leave a comment below or drop me a line! I will be sure to show how I did it if it's interesting to any reader!</p>","metadata":{}},{"cell_type":"markdown","source":"<h1 style=\"font-size: 35px;\">👀 Is this done? 👀</h1>\n\n<p style=\"font-size: 16px;\">Short answer: <b>No!</b> There's a clear problem with this dataset, the classes are imbalanced. So even if you read the images (which in this case is already complex) you still need to address this! Let's show what I mean...</p>\n\n<p style=\"font-size: 16px;\">We will import Seaborn to quickly visualize the imbalance. You can do it in matplotlib as well, I just prefer Seaborn's color palette. If we do a simple countplot on the \"cancer\" column we will get the imbalance. Also, note the below the plot I'm using a normalized value_counts method to get the percentage each of the classes represent in the entire dataset.</p>","metadata":{}},{"cell_type":"code","source":"import seaborn as sns","metadata":{"execution":{"iopub.status.busy":"2023-02-01T10:23:49.18314Z","iopub.execute_input":"2023-02-01T10:23:49.183949Z","iopub.status.idle":"2023-02-01T10:23:49.750828Z","shell.execute_reply.started":"2023-02-01T10:23:49.183514Z","shell.execute_reply":"2023-02-01T10:23:49.7499Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.countplot(data=train_df, x='cancer')","metadata":{"execution":{"iopub.status.busy":"2023-02-01T10:24:24.487267Z","iopub.execute_input":"2023-02-01T10:24:24.487666Z","iopub.status.idle":"2023-02-01T10:24:24.710538Z","shell.execute_reply.started":"2023-02-01T10:24:24.487633Z","shell.execute_reply":"2023-02-01T10:24:24.709654Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['cancer'].value_counts(normalize=True)","metadata":{"execution":{"iopub.status.busy":"2023-02-01T10:24:47.918483Z","iopub.execute_input":"2023-02-01T10:24:47.918826Z","iopub.status.idle":"2023-02-01T10:24:47.930082Z","shell.execute_reply.started":"2023-02-01T10:24:47.918796Z","shell.execute_reply":"2023-02-01T10:24:47.929071Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style=\"font-size: 35px;\">A few tips and tricks...</h1>\n\n<p style=\"font-size: 16px;\">This competition has been already active for a month, so there has been a lot of progress already! Check out this discussing brought by The Devastator: <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/374248\">Discussing approaches to the problem</a>.</p>","metadata":{}},{"cell_type":"markdown","source":"<p style=\"font-size: 16px;\">This discussion brings a few really interesting points that I will definitely using for my notebooks. I will list some of them below:</p>\n\n<ul>\n    <li>As most of the image classification problems: Try transfer learning. Fine-tuning doesn't seem to get as good results.</li>\n    <li>Within transfer learning, ResNet seems to have a really good starting configuration for transfer learning.</li>\n    <li>Optimze the images. Training the model can be resource extensive, so there are some tricks using JPEG encodings.</li>\n    <li>Use ROI (Region of Interest), since most images have a lot of empty pixels that are not very relevant.</li>\n</ul>","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}