{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# 🎗️[Custom Preprocessor] RSNA Breast Cancer\n---","metadata":{}},{"cell_type":"markdown","source":"The purpose of **preprocessing the images** of the [*RSNA Breast Cancer Detection Competition*](https://www.kaggle.com/competitions/rsna-breast-cancer-detection) is to get them into a consistent format and to remove any irrelevant or noisy data that may interfere with our model's ability to learn. By preprocessing the images, we can improve the performance and accuracy of our model, as well as make the training process more efficient.\n\nPreprocessing the images is an important step in the machine learning pipeline, as it allows us to extract valuable features from the data and feed them into our model. Without preprocessing, our model may have to work harder to find patterns in the raw data, which can lead to overfitting or poor generalization to new data.\n\nOverall, preprocessing the images in our dataset is an essential step that helps us prepare the data for training and improve the performance of our model.\n\n**Last Version (V3):** January 25th 2023\n\n[See version changelog](#Versions)","metadata":{}},{"cell_type":"markdown","source":"# Getting started\n## Import packages","metadata":{}},{"cell_type":"code","source":"# Install required libraries\n!pip install -qU python-gdcm pydicom pylibjpeg\n!pip install -qU dicomsdl\n\n# Standard library imports\nimport os\nimport time\n\n# Third-party library imports\nimport numpy as np\nimport pandas as pd\nimport cv2\nimport pydicom\nimport dicomsdl\n\n# Visualization library imports\nimport matplotlib.pyplot as plt\n\n# Progress bar library imports\nfrom tqdm.notebook import tqdm, trange\n\n# Parallel processing library imports\nfrom joblib import Parallel, delayed\n\nimport gc","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2023-01-24T09:15:34.554877Z","iopub.execute_input":"2023-01-24T09:15:34.555735Z","iopub.status.idle":"2023-01-24T09:16:04.910694Z","shell.execute_reply.started":"2023-01-24T09:15:34.555614Z","shell.execute_reply":"2023-01-24T09:16:04.9095Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Get data paths\n- Path to the CSV metadata\n- Path to the training images","metadata":{}},{"cell_type":"code","source":"csv_path = '/kaggle/input/rsna-breast-cancer-detection/train.csv'\ntrain_path = '/kaggle/input/rsna-breast-cancer-detection/train_images'\ndata = pd.read_csv(csv_path)","metadata":{"execution":{"iopub.status.busy":"2023-01-24T09:16:04.912691Z","iopub.execute_input":"2023-01-24T09:16:04.913203Z","iopub.status.idle":"2023-01-24T09:16:05.04656Z","shell.execute_reply.started":"2023-01-24T09:16:04.913166Z","shell.execute_reply":"2023-01-24T09:16:05.045454Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Utils","metadata":{}},{"cell_type":"markdown","source":"This section includes utility functions that are used in the notebook.\n- `get_paths`: prepare the inputs of the *MammographyPreprocessor* methods\n- `calculate_aspect_ratios`: preprocessing resizing shape decision support tool","metadata":{}},{"cell_type":"markdown","source":"### get_paths\nThis function generates and returns a list of n paths of the DICOM files, using data from the CSV file. If desired, the paths can be shuffled to produce a random selection, otherwise, the paths will be returned in their original order. In addition to the paths, this function also returns a cache, which is a list of dictionaries that map the patient IDs and scan IDs for each file. If n is omitted, all paths are obtained","metadata":{}},{"cell_type":"code","source":"def get_paths(n: int=len(data), shuffle: bool=False):\n    if shuffle == True:\n        df = data.sample(frac=1, random_state=0)\n    else:\n        df = data\n    paths = []\n    ids_cache = []\n    for i in range(n):\n        patient = str(df.iloc[i]['patient_id'])\n        scan = str(df.iloc[i]['image_id'])\n        paths.append(train_path + '/' + patient + '/' + scan + '.dcm')\n        ids_cache.append({'patient_id': patient, 'scan_id': scan})\n    return paths, ids_cache\n\n\n# Example\npaths, _ = get_paths(n=5, shuffle=True)\npaths","metadata":{"execution":{"iopub.status.busy":"2023-01-24T09:16:05.047845Z","iopub.execute_input":"2023-01-24T09:16:05.04824Z","iopub.status.idle":"2023-01-24T09:16:05.087662Z","shell.execute_reply.started":"2023-01-24T09:16:05.048199Z","shell.execute_reply":"2023-01-24T09:16:05.086528Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### calculate_aspect_ratios","metadata":{}},{"cell_type":"markdown","source":"This function returns a list of the aspect ratios (height to width) of the images specified by the filepaths to DICOM files. An optional preprocessor can be provided, which can crop the images before the aspect ratios are calculated. The purpose of this is to have an idea of the average aspect ratio of the cropped images over the dataset. Maybe the mammograms do not necessarily need to be square (e.g. 256x256, 512x512, 1024x1024), and a different aspect ratio may be more appropriate in order to reduce the pixel distorsion (e.g. 256x512, 512x768, 1024x1280).","metadata":{}},{"cell_type":"code","source":"def calculate_aspect_ratios(paths: list, preprocessor=None):\n    ratios = []\n    for i in trange(len(paths)):\n        if preprocessor:\n            img = preprocessor.preprocess_single_image(paths[i])\n        else:\n            scan = pydicom.dcmread(paths[i])\n            img = scan.pixel_array\n        height, width = img.shape\n        ratio = height / width\n        ratios.append(ratio)\n    return ratios\n\n\n# Example\nratios = calculate_aspect_ratios(paths)\nprint(\"Ratios:\", ratios)\nprint(\"Min:\", np.min(ratios))\nprint(\"Max:\", np.max(ratios))\nprint(\"Avg:\", np.mean(ratios))","metadata":{"execution":{"iopub.status.busy":"2023-01-24T09:16:05.089834Z","iopub.execute_input":"2023-01-24T09:16:05.090236Z","iopub.status.idle":"2023-01-24T09:16:08.69131Z","shell.execute_reply.started":"2023-01-24T09:16:05.090202Z","shell.execute_reply":"2023-01-24T09:16:08.68986Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Custom class: MammographyPreprocessor\n\n*Sources:*\n- Modular Approach: https://www.kaggle.com/code/onkur7/image-preprocessing-modular-approach/notebook\n- Windowing: https://www.kaggle.com/code/davidbroberts/mammography-apply-windowing/notebook\n\nThe MammographyPreprocessor class purpose is to preprocess the raw images of the [*RSNA Breast Cancer Detection Competition*](https://www.kaggle.com/competitions/rsna-breast-cancer-detection) using different technics.\n\n### Preprocessing steps\nThe steps for a single image are the following:\n- **Windowing**: improve the contrast of images the way they are intended to be viewed\n- **Fix the photometric interpretation**: set all image backgrounds to zero\n- **Rescale with slope and intercept**: it seems like it is not necessary but just in case\n- **Normalize between 0 and 255**: convert to 8-bits pixel values, grayscale\n- **Flip the breasts**: flip the breasts on the same side for more consistency\n- **Crop the background**: remove extra background\n- **Resize**: make all images the same size for the future model\n- **Save the image**: PNG default, JPEG optional\n\n### Create an instance\nTo create an instance of the class, you only need to specify:\n- **size** (*optional*): the dimensions of the resized images (if not specified, the images are not resized) ⚠️ **Size is given with the OpenCV convention (Width x Height)**. \n- **breast_side** (*default: 'L'*): the side where you want all the breasts to be reversed\n\n\n### Public methods\nYou can use the following public methods:\n- **get_paths()**: Return a list of n paths from the instance dataframe optional shuffling. **(V2)**\n - *n: int*\n - *shuffle: bool*\n - *return_cache: bool*\n- **read_image()**: Returns the image array of a single DICOM file.\n - *path: str*\n- **preprocess_all()**: Preprocess all the files of the list and save them by default as PNG. You can choose to do the preprocessing in parallel with the associated number of jobs.\n - *paths: list*\n - *save: bool*\n - *save_dir: bool*\n - *png: bool*\n - *parallel: bool*\n - *n_jobs: int*\n- **preprocess_single_image()**: Returns the array of a single preprocessed file. You can choose to save it if you want to.\n - *path: str*\n - *save: bool*\n - *save_dir: bool*\n - *png: bool*\n- **display()**: Display a grid of chosen rows and columns from the DICOM paths. You can choose to preprocess them before and to save the display on the working directory.\n - *paths: str*\n - *rows: int*\n - *cols: int*\n - *preprocess: bool*\n - *cmap: matplotlib colormap*\n - *cbar: matplotlib colorbar*\n - *save_fig: bool*\n - *save_name: str*","metadata":{}},{"cell_type":"code","source":"class MammographyPreprocessor():\n    \n    # Constructor\n    def __init__(self, size: tuple=None, breast_side: str='L',\n                 csv_path=None, train_path=None):\n        self.size = size\n        os.makedirs(os.getcwd(), exist_ok=True)\n        self.breast_side = breast_side\n        assert breast_side in ['L', 'R'], \"breast_side should be 'L' or 'R'\"\n        # implement the paths of the original RSNA dataset (V2)\n        self.csv_path = '/kaggle/input/rsna-breast-cancer-detection/train.csv'\n        self.train_path = '/kaggle/input/rsna-breast-cancer-detection/train_images'\n        if csv_path:\n            self.csv_path = csv_path\n        if train_path:\n            self.train_path = train_path\n        self.df = pd.read_csv(self.csv_path)\n        self.save_root = os.getcwd()\n    \n    # Get the paths from the preprocessor (V2)\n    def get_paths(self, n: int=None, shuffle: bool=False, return_cache: bool=False):\n        if n == None:\n            n = len(self.df)\n        if shuffle == True:\n            df = self.df.sample(frac=1, random_state=0).copy()\n        else:\n            df = self.df.copy()\n        paths = []\n        ids_cache = []\n        for i in range(n):\n            patient = str(df.iloc[i]['patient_id'])\n            scan = str(df.iloc[i]['image_id'])\n            paths.append(self.train_path + '/' + patient + '/' + scan + '.dcm')\n            ids_cache.append({'patient_id': patient, 'scan_id': scan})\n        if return_cache:\n            return paths, ids_cache\n        else:\n            return paths\n    \n    # Read from a path and convert to image array\n    def read_image(self, path: str):\n        scan = pydicom.dcmread(path)\n        img = scan.pixel_array\n        return img\n    \n    # Apply the preprocessing methods on one image\n    def preprocess_single_image(self, path: str, save: bool=False,\n                                save_dir: str=None, png: bool=True):\n        scan = dicomsdl.open(path)\n        img = scan.pixelData()\n        img = self._windowing(img, scan)\n        img = self._fix_photometric_interpretation(img, scan)\n        img = self._normalize_to_255(img)\n        img = self._flip_breast_side(img)\n        img = self._crop(img)\n        if self.size:\n            img = self._resize(img)\n        if save:\n            self._save_image(img, path, png, save_dir)\n            return # do not return the images to avoid memory leak\n        return img\n    \n    # Preprocess all the images from the paths\n    def preprocess_all(self, paths: list, save: bool=True,\n                       save_dir: str='train_images', png: bool=True,\n                       parallel: bool=False, n_jobs: int=4):\n        clock = time.time()\n        if parallel:\n            Parallel(n_jobs=n_jobs) \\\n            (delayed(self.preprocess_single_image) \\\n            (path, save, save_dir, png) for path in tqdm(paths, total=len(paths)))\n            print(\"Parallel preprocessing done!\")\n        else:\n            for i in trange(len(paths)):\n                self.preprocess_single_image(paths[i], save, save_dir, png)\n            print(\"Sequential preprocessing done!\")\n        print(\"Time =\", np.around(time.time() - clock, 3), 'sec')\n    \n    # Display the images from the dicom paths with optional preprocessing\n    def display(self, paths: list, rows: int, cols: int,\n                preprocess: bool=False, cmap='bone', cbar: bool=False,\n                save_fig: bool=False, save_name: str='myplot.png'):\n        assert len(paths) >= (rows * cols), \\\n        f\"Not enough paths for the display. \" \\\n        f\"Please give at least {rows * cols} paths.\"\n        plt.figure(figsize=(18, 26 * rows / cols))\n        for i in trange(rows * cols):\n            path = paths[i]\n            if preprocess:\n                img = self.preprocess_single_image(path, save=False)\n            else:\n                img = self.read_image(path)\n            plt.subplot(rows, cols, i+1)\n            plt.imshow(img, cmap=cmap)\n            if cbar:\n                plt.colorbar()\n            plt.grid(False)\n            plt.title(path.split('/')[-1][:-4])\n        plt.suptitle(\"Preprocessed images\" if preprocess \\\n                     else \"Raw images\", fontsize=25)\n        if save_fig:\n            plt.savefig(save_name, facecolor='white')\n        plt.show()\n    \n    # Adjust the contrast of an image\n    def _windowing(self, img, scan):\n        center = scan.WindowCenter\n        width = scan.WindowWidth\n        bits_stored = scan.BitsStored\n        function = scan.VOILUTFunction\n        if isinstance(center, list):\n            center = center[0]\n        if isinstance(width, list):\n            width = width[0] \n        y_range = float(2**bits_stored - 1)\n        if function == 'SIGMOID':\n            img = y_range / (1 + np.exp(-4 * (img - center) / width))\n        else: # LINEAR\n            center -= 0.5\n            width -= 1\n            below = img <= (center - width / 2)\n            above = img > (center + width / 2)\n            between = np.logical_and(~below, ~above)\n            img[below] = 0\n            img[above] = y_range\n            img[between] = ((img[between] - center) / width + 0.5) * y_range\n        return img\n    \n    # Interpret pixels in a consistant way\n    def _fix_photometric_interpretation(self, img, scan):\n        if scan.PhotometricInterpretation == 'MONOCHROME1':\n            return img.max() - img\n        elif scan.PhotometricInterpretation == 'MONOCHROME2':\n            return img - img.min()\n        else:\n            raise ValueError(\"Invalid Photometric Interpretation: {}\"\n                               .format(scan.PhotometricInterpretation))\n    \n    # Cast into 8-bits for saving\n    def _normalize_to_255(self, img):\n        if img.max() != 0:\n            img = img / img.max()\n        img *= 255\n        return img.astype(np.uint8)\n    \n    # Flip the breast horizontally on the chosen side \n    def _flip_breast_side(self, img):\n        img_breast_side = self._determine_breast_side(img)\n        if img_breast_side == self.breast_side:\n            return img\n        else:\n            return np.fliplr(img)    \n    \n    # Determine the current breast side\n    def _determine_breast_side(self, img):\n        col_sums_split = np.array_split(np.sum(img, axis=0), 2)\n        left_col_sum = np.sum(col_sums_split[0])\n        right_col_sum = np.sum(col_sums_split[1])\n        if left_col_sum > right_col_sum:\n            return 'L'\n        else:\n            return 'R'\n    \n    # Crop the useless background of the image\n    def _crop(self, img):\n        bin_img = self._binarize(img, threshold=5)\n        contour = self._extract_contour(bin_img)\n        img = self._erase_background(img, contour)\n        x1, x2 = np.min(contour[:, :, 0]), np.max(contour[:, :, 0])\n        y1, y2 = np.min(contour[:, :, 1]), np.max(contour[:, :, 1])\n        x1, x2 = int(0.99 * x1), int(1.01 * x2)\n        y1, y2 = int(0.99 * y1), int(1.01 * y2)\n        return img[y1:y2, x1:x2]\n    \n    # Binarize the image at the threshold\n    def _binarize(self, img, threshold):\n        return (img > threshold).astype(np.uint8)\n    \n    # Get contour points of the breast\n    def _extract_contour(self, bin_img):\n        contours, _ = cv2.findContours(\n            bin_img, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_NONE)\n        contour = max(contours, key=cv2.contourArea)\n        return contour\n    \n    # Set to background pixels of the image to zero\n    def _erase_background(self, img, contour):\n        mask = np.zeros(img.shape, np.uint8)\n        cv2.drawContours(mask, [contour], -1, 255, cv2.FILLED)\n        output = cv2.bitwise_and(img, mask)\n        return output\n    \n    # Resize the image to the preprocessor size\n    def _resize(self, img):\n        return cv2.resize(img, self.size)\n    \n    # Get the save path of a given dicom file\n    def _get_save_path(self, path, png, save_dir):\n        patient = path.split('/')[-2]\n        filename = path.split('/')[-1]\n        if png:\n            filename = filename.replace('dcm', 'png')\n        else:\n            filename = filename.replace('dcm', 'jpeg')\n        if save_dir:\n            save_path = os.path.join(self.save_root, save_dir, patient, filename)\n        else:\n            save_path = os.path.join(self.save_root, patient, filename)\n        return save_path\n    \n    # Save the preprocessed image\n    def _save_image(self, img, path, png, save_dir):\n        save_path = self._get_save_path(path, png, save_dir)\n        patient_folder = os.path.split(save_path)[0]\n        os.makedirs(patient_folder, exist_ok=True)\n        cv2.imwrite(save_path, img)","metadata":{"execution":{"iopub.status.busy":"2023-01-24T09:16:08.693521Z","iopub.execute_input":"2023-01-24T09:16:08.694249Z","iopub.status.idle":"2023-01-24T09:16:08.74055Z","shell.execute_reply.started":"2023-01-24T09:16:08.694198Z","shell.execute_reply":"2023-01-24T09:16:08.738843Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Tutorial\n### First, get the paths!\n","metadata":{}},{"cell_type":"code","source":"# Test with 100 files \npaths, _ = get_paths(n=100, shuffle=True)","metadata":{"execution":{"iopub.status.busy":"2023-01-24T09:16:08.742651Z","iopub.execute_input":"2023-01-24T09:16:08.743211Z","iopub.status.idle":"2023-01-24T09:16:08.817715Z","shell.execute_reply.started":"2023-01-24T09:16:08.743143Z","shell.execute_reply":"2023-01-24T09:16:08.816513Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Create a MammographyPreprocessor instance\nLet's set the final resize to 256x256 pixels. It might not be the best to train a good model, but it is fine for a baseline.","metadata":{}},{"cell_type":"code","source":"mp = MammographyPreprocessor(size=(256, 256))\nmp","metadata":{"execution":{"iopub.status.busy":"2023-01-24T09:16:08.819727Z","iopub.execute_input":"2023-01-24T09:16:08.820204Z","iopub.status.idle":"2023-01-24T09:16:08.903127Z","shell.execute_reply.started":"2023-01-24T09:16:08.820157Z","shell.execute_reply":"2023-01-24T09:16:08.902152Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Read a raw file\nThis can be useful to start a quick EDA.\n\n⚠️ **Shameful personal promotion:** You can check mine [here](https://www.kaggle.com/code/paulbacher/eda-breast-cancer-detection-competition)! 🙃","metadata":{}},{"cell_type":"code","source":"# Select the first file of the list\npath = paths[12]\n\n# Load the image array\nimg = mp.read_image(path=path)\n\n# Display what has been loaded\nplt.imshow(img, cmap='jet')\nplt.colorbar()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-01-24T09:16:08.904426Z","iopub.execute_input":"2023-01-24T09:16:08.904785Z","iopub.status.idle":"2023-01-24T09:16:09.871493Z","shell.execute_reply.started":"2023-01-24T09:16:08.904752Z","shell.execute_reply":"2023-01-24T09:16:09.87Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Preprocess a single image\nTo quickly check the result after preprocessing a given image, you can do it as follow.","metadata":{}},{"cell_type":"code","source":"# Preprocess from the same previous path\nimg = mp.preprocess_single_image(path=path, save=False)\n\n# Display the preprocessed image\nplt.imshow(img, cmap='jet')\nplt.colorbar()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-01-24T09:16:09.873446Z","iopub.execute_input":"2023-01-24T09:16:09.873991Z","iopub.status.idle":"2023-01-24T09:16:10.41039Z","shell.execute_reply.started":"2023-01-24T09:16:09.873928Z","shell.execute_reply":"2023-01-24T09:16:10.409032Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As expected, the raw image has been cropped and resized to a square of **256x256** pixels among other preprocessing steps.","metadata":{}},{"cell_type":"markdown","source":"### Preprocess all the images\nTo efficiently preprocess all of your images, you can use the specified method on your list of file paths. In this case, since there are only 100 files, the preprocessing will be completed quickly. Keep in mind that there are 54706 files to preprocess in the full training set and this will easily take several hours, depending on the way you preprocess them.\n\nThe resulting grayscale PNG images will be saved in the current working directory, the same structure as the input files were.\n\nFor the example, I am using parallel computation on the 4 cores of the CPU. This is faster but not necessarily by a factor of 4. You can also use the GPU to increase the speed of preprocessing.","metadata":{}},{"cell_type":"code","source":"mp.preprocess_all(paths, save=True, save_dir='train_images', parallel=True, n_jobs=4)","metadata":{"execution":{"iopub.status.busy":"2023-01-24T09:16:10.414767Z","iopub.execute_input":"2023-01-24T09:16:10.415623Z","iopub.status.idle":"2023-01-24T09:16:54.279851Z","shell.execute_reply.started":"2023-01-24T09:16:10.415571Z","shell.execute_reply":"2023-01-24T09:16:54.278365Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Display many examples\nThis goes along with what has been explained previously. Just use the parameters you want. 😊\n\nThe images are reloaded from the original paths and preprocessing can be done before displaying. It can take a few seconds for a large number of images.\n\nFirst I will display a grid of raw images.","metadata":{}},{"cell_type":"code","source":"# Images before preprocessing\nmp.display(paths, rows=3, cols=5, preprocess=False, cmap='jet', save_fig=False)","metadata":{"execution":{"iopub.status.busy":"2023-01-24T09:16:54.281372Z","iopub.execute_input":"2023-01-24T09:16:54.281759Z","iopub.status.idle":"2023-01-24T09:17:17.945118Z","shell.execute_reply.started":"2023-01-24T09:16:54.28172Z","shell.execute_reply":"2023-01-24T09:17:17.944164Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see, the raw images are not very consistant in there shapes, pixel values, or orientation.\n\nLet's have a look at the same images after preprocessing by just setting preprocess to *True*.","metadata":{}},{"cell_type":"code","source":"# Images after preprocessing\nmp.display(paths, rows=3, cols=5, preprocess=True, cmap='jet', save_fig=False)","metadata":{"execution":{"iopub.status.busy":"2023-01-24T09:17:17.946418Z","iopub.execute_input":"2023-01-24T09:17:17.947279Z","iopub.status.idle":"2023-01-24T09:17:32.791392Z","shell.execute_reply.started":"2023-01-24T09:17:17.94724Z","shell.execute_reply":"2023-01-24T09:17:32.790114Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It looks way better like this! Now we can expect our future model to learn better with a good preprocessing.","metadata":{}},{"cell_type":"markdown","source":"# About resizing parameter\nAt the beginning of the notebook, I wrote a function that returns a list of the image ratios. I would like to try something now.\n\nLet's calculate the average size ratio of several breasts just after preprocessing them but without applying the resize method yet.","metadata":{}},{"cell_type":"code","source":"# Get more file paths\npaths, _ = get_paths(n=300, shuffle=True)\n\n# Create a new preprocessor with no resize value\nmp = MammographyPreprocessor()\n\n# Calculate the size ratios with the preprocessor\nratios = calculate_aspect_ratios(paths, preprocessor=mp)\nprint(\"Min ratio:\", np.min(ratios))\nprint(\"Max ratio:\", np.max(ratios))\nprint(\"Med ratio:\", np.median(ratios))","metadata":{"execution":{"iopub.status.busy":"2023-01-24T09:17:32.792785Z","iopub.execute_input":"2023-01-24T09:17:32.793202Z","iopub.status.idle":"2023-01-24T09:19:29.176986Z","shell.execute_reply.started":"2023-01-24T09:17:32.793165Z","shell.execute_reply":"2023-01-24T09:19:29.174213Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Once cropped, the 300 sample images have a median aspect ratio of 2.1 and are rectangular in shape. Resizing these images is necessary for computational purposes, but it results in the loss of some information due to compression. When the cropped images are resized into squares, the information is not lost evenly in all directions. Instead, the images tend to be more compressed vertically than horizontally. We can clearly see this horizontal distortion in the previous plot.\n\nTo minimize this effect, it might be helpful to resize the images into a shape that is closer to their median aspect ratio, which appears to be around 2. Possible resize dimensions that could achieve this might include **128x256**, **256x512**, or **512x1024**.\n\nLet's have a look at it.","metadata":{}},{"cell_type":"code","source":"# Define first the new size value\nmp = MammographyPreprocessor(size=(256, 512)) # OpenCV convention (Width x Height)\nmp.display(paths, rows=3, cols=5, preprocess=True, cmap='jet', cbar=True, save_fig=False)","metadata":{"execution":{"iopub.status.busy":"2023-01-24T09:19:33.623883Z","iopub.execute_input":"2023-01-24T09:19:33.624524Z","iopub.status.idle":"2023-01-24T09:19:50.306581Z","shell.execute_reply.started":"2023-01-24T09:19:33.624468Z","shell.execute_reply":"2023-01-24T09:19:50.305306Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Here we go, pixel distortion is still present but to a lesser extent. What could completely remove this distortion would be to resize the images while keeping their aspect ratio then add a padding on one of the sides. This is a planned option for the next release.\n\nThis has been suggested to me by [Remek Kinas](https://www.kaggle.com/remekkinas).\n\n**Let me know your suggestions about this topic!**","metadata":{}},{"cell_type":"markdown","source":"# How to install the preprocessor on your notebook\nCopy-paste these two lines at the top of your notebook and you will be able to use the `MammographyPreprocessor`!","metadata":{}},{"cell_type":"code","source":"# !pip install git+https://github.com/Paul-Bacher/RSNA-Screening-Mammography-Breast-Cancer-Detection.git#subdirectory=rsna_custom\n# from rsna_custom import MammographyPreprocessor","metadata":{"execution":{"iopub.status.busy":"2023-01-20T08:44:33.622829Z","iopub.status.idle":"2023-01-20T08:44:33.623162Z","shell.execute_reply.started":"2023-01-20T08:44:33.623007Z","shell.execute_reply":"2023-01-20T08:44:33.623023Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Versions\n- **V1:**\n - The MammographyPreprocessor class can load and preprocess mammography images, including: applying windowing, rescaling and normalizing pixel values, flipping the side, cropping and resizing\n - It allows for the display and saving of raw and preprocessed images, including the option to save preprocessed images as PNG files in the original folder structure.\n - It can process a single image or multiple images from a list of file paths.\n- **V2:**\n - The `get_paths` function has been implemented inside of the preprocessor class.\n - Loading before preprocessing is now done via `dicomsdl` which seems to save about 35-40% of computation time.\n - Add a parameter for the saving destination of the image\n- **V3:**\n - Fix memory usage (cf: [comment section](https://www.kaggle.com/code/paulbacher/custom-preprocessor-rsna-breast-cancer/comments#2111242))","metadata":{}},{"cell_type":"markdown","source":"# Endnotes\n- I will keep updating my custom class with new public methods that could be useful or with new preprocessing steps that I can find.\n- Please, let me know in the comments if you find any bug. Also, I would be glad if you suggest me ways to improve my class, like saving computation time. I am also learning a lot from this work!\n- Of course, you are free to use my class and modify it for your own purpose.\n\n**Thank you for reading**, please upvote my work if you found it interesting. 😄","metadata":{}}]}