{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.10","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":39272,"databundleVersionId":4629629,"sourceType":"competition"}],"dockerImageVersionId":30474,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"### 🧬 Overview\n\nThis project is based on the RSNA Screening Mammography Breast Cancer Detection Kaggle competition, with a primary focus on data preprocessing for large-scale medical imaging data. Rather than emphasizing model complexity, this work concentrates on preparing high-quality, clinically meaningful inputs that are essential for building reliable and robust machine learning models.\n\nMedical imaging data presents unique challenges, including high resolution, class imbalance, noise, and inconsistencies across scans. Effective preprocessing is a critical step in improving downstream model performance and interpretability.\n\n***\n\n### 🎯 Project Objective\n\nThe main objectives of this project are to:\n\n- 🧹 Clean and standardize raw mammography data\n- 🖼️ Prepare images for efficient and effective model training\n- 🔍 Preserve clinically relevant features during preprocessing\n- 🧠 Establish a strong foundation for downstream modeling\n\n***\n\n### 📂 Dataset Description\n\nThe dataset consists of screening mammography exams provided by RSNA, including:\n- 🖼️ High-resolution mammogram images\n- 👩‍⚕️ Patient- and exam-level metadata\n- 🏷️ Binary cancer labels (cancer / no cancer)\n- 📅 Longitudinal screening records for selected patients\n\nThe dataset reflects real-world clinical distributions, making preprocessing both challenging and essential.\n\n***\n\n### ⚠️ Key Challenges Addressed\n\n- 🖼️ Extremely large image sizes and memory constraints\n- 🧠 Preservation of subtle diagnostic patterns\n- 🏥 Maintaining clinical validity during transformations","metadata":{}},{"cell_type":"markdown","source":"### Setup environment","metadata":{}},{"cell_type":"code","source":"!pip install -q python-gdcm pydicom pylibjpeg\n!pip install -q --upgrade scipy","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-05-10T08:17:39.164262Z","iopub.execute_input":"2023-05-10T08:17:39.164867Z","iopub.status.idle":"2023-05-10T08:18:26.020474Z","shell.execute_reply.started":"2023-05-10T08:17:39.164832Z","shell.execute_reply":"2023-05-10T08:18:26.018931Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"#### Import libraries","metadata":{}},{"cell_type":"code","source":"# utilities\nimport os\nimport glob\nimport shutil\nimport time\nimport math\n\n# data processing\nimport cv2\nimport gdcm\nimport pydicom\nimport numpy as np\nfrom skimage.exposure import rescale_intensity\nfrom pydicom.pixel_data_handlers.util import apply_voi_lut\n\n# visualisation\nimport matplotlib.pyplot as plt","metadata":{"execution":{"iopub.status.busy":"2023-05-10T08:18:26.023048Z","iopub.execute_input":"2023-05-10T08:18:26.023466Z","iopub.status.idle":"2023-05-10T08:18:26.992602Z","shell.execute_reply.started":"2023-05-10T08:18:26.023428Z","shell.execute_reply":"2023-05-10T08:18:26.991517Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"#### Configurations","metadata":{}},{"cell_type":"code","source":"IMAGE_WIDTH = 640\nIMAGE_HEIGHT = 1280\nEXTENSION = 'jpg'\n\nDATA_DIR = '/kaggle/input/rsna-breast-cancer-detection'\nSAVE_DIR = '/kaggle/working/processed'\n\nshutil.rmtree(SAVE_DIR) if os.path.exists(SAVE_DIR) else print(f'Creating directory {SAVE_DIR}')\nos.makedirs(SAVE_DIR, exist_ok=True)","metadata":{"execution":{"iopub.status.busy":"2023-05-10T08:19:54.343986Z","iopub.execute_input":"2023-05-10T08:19:54.344634Z","iopub.status.idle":"2023-05-10T08:19:54.354955Z","shell.execute_reply.started":"2023-05-10T08:19:54.344577Z","shell.execute_reply":"2023-05-10T08:19:54.353554Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"#### Utilities","metadata":{}},{"cell_type":"code","source":"# view image\ndef display_image(image):\n    plt.figure(figsize=(8, 8))\n    plt.imshow(image, cmap='bone')\n    plt.title(f'{image.shape} [{image.min():.0f}, {image.max():.0f}] {image.dtype}')\n    plt.axis('off')\n    plt.tight_layout()\n    plt.show()\n    \n# view image and the histogram\ndef display_image_histogram(image):\n    fig, ax = plt.subplots(1, 2, figsize=(12, 4))\n    \n    # plot image and histogram side by side\n    ax[0].imshow(image, cmap='bone')\n    ax[0].axis('off')\n    ax[1].hist(image.ravel(), bins=64, range=(image.min(), image.max()))\n    \n    title = f'shape: {image.shape}, pixel values: [{image.min():.0f}, {image.max():.0f}], {image.dtype}'\n    fig.suptitle(title, fontsize=10)\n    plt.tight_layout()\n    plt.show()  \n\n# view multiple images\ndef display_images(images, images_per_row=None, ratio=9, header=None):\n    num_images = len(images)\n    if not images_per_row:\n        images_per_row = num_images\n    num_rows = math.ceil(num_images / images_per_row)\n    \n    fig, axs = plt.subplots(num_rows, images_per_row, figsize=(20, ratio*num_rows))\n    axs = axs.flatten()\n    \n    for i, image in enumerate(images):\n        axs[i].imshow(image, cmap='bone')\n        axs[i].axis('off')\n        \n        if not header:\n            title = f'{image.shape} [{image.min():.0f}, {image.max():.0f}] {image.dtype}'\n        else:\n            title = f'{header[i]}\\n{image.shape} [{image.min():.0f}, {image.max():.0f}] {image.dtype}'\n        axs[i].set_title(title, fontsize=10)\n    \n    # hide any unused subplots\n    for i in range(num_images, num_rows * images_per_row):\n        fig.delaxes(axs[i])\n    \n    fig.tight_layout()\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2023-05-10T08:19:45.003298Z","iopub.execute_input":"2023-05-10T08:19:45.003717Z","iopub.status.idle":"2023-05-10T08:19:45.020059Z","shell.execute_reply.started":"2023-05-10T08:19:45.003686Z","shell.execute_reply":"2023-05-10T08:19:45.019056Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Explore DICOM data","metadata":{}},{"cell_type":"code","source":"DCM_FILES = glob.glob(f'{DATA_DIR}/train_images/10042/*.dcm')\n\nprint(f\"{len(DCM_FILES)} DICOM files found.\")\nprint(DCM_FILES[0])","metadata":{"execution":{"iopub.status.busy":"2023-05-10T08:19:57.14192Z","iopub.execute_input":"2023-05-10T08:19:57.142355Z","iopub.status.idle":"2023-05-10T08:19:57.170194Z","shell.execute_reply.started":"2023-05-10T08:19:57.142322Z","shell.execute_reply":"2023-05-10T08:19:57.16888Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"dicom = pydicom.dcmread(DCM_FILES[0])\ndicom","metadata":{"execution":{"iopub.status.busy":"2023-05-10T08:20:00.669821Z","iopub.execute_input":"2023-05-10T08:20:00.670245Z","iopub.status.idle":"2023-05-10T08:20:00.753322Z","shell.execute_reply.started":"2023-05-10T08:20:00.670206Z","shell.execute_reply":"2023-05-10T08:20:00.752078Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"display_image_histogram(dicom.pixel_array)","metadata":{"execution":{"iopub.status.busy":"2023-05-10T08:20:02.838685Z","iopub.execute_input":"2023-05-10T08:20:02.839714Z","iopub.status.idle":"2023-05-10T08:20:04.447236Z","shell.execute_reply.started":"2023-05-10T08:20:02.839657Z","shell.execute_reply":"2023-05-10T08:20:04.445055Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"dcm_imgs = []\nfor path in DCM_FILES:\n    patient_id = path.split('/')[-2]\n    \n    dicom = pydicom.dcmread(path)\n    dcm_imgs.append(dicom.pixel_array)\n    \ndisplay_images(dcm_imgs, 5, 5)","metadata":{"execution":{"iopub.status.busy":"2023-05-10T08:20:10.356684Z","iopub.execute_input":"2023-05-10T08:20:10.357162Z","iopub.status.idle":"2023-05-10T08:20:15.119756Z","shell.execute_reply.started":"2023-05-10T08:20:10.357112Z","shell.execute_reply":"2023-05-10T08:20:15.118442Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Process DICOM data","metadata":{}},{"cell_type":"code","source":"# read dicom and normalise\ndef read_dicom(path):\n    dicom = pydicom.dcmread(path)\n    dicom_img = dicom.pixel_array\n    \n    # apply VOI LUT or windowing to improve the contrast    \n    dicom_img = apply_voi_lut(dicom_img, dicom)\n\n    # fix inverted monochrome\n    if dicom.PhotometricInterpretation == \"MONOCHROME1\":\n        dicom_img = np.amax(dicom_img) - dicom_img\n    else:\n        dicom_img = dicom_img\n\n    # rescale pixel values to 0-255 range and convert to uint8\n    dicom_img = dicom_img - np.min(dicom_img)\n    dicom_img = dicom_img / np.max(dicom_img)\n    dicom_img = (dicom_img * 255).astype(np.uint8)\n    return dicom_img\n\n# crop borders by image percentage\ndef crop_borders(image, left=0.01, right=0.01, top=0.04, bottom=0.04):\n    height, width = image.shape[:2]\n    \n    # Get the start and end rows and columns\n    left_crop = int(width * left)\n    right_crop = int(width * (1 - right))\n    top_crop = int(height * top)\n    bottom_crop = int(height * (1 - bottom))\n    \n    cropped_img = image[top_crop:bottom_crop, left_crop:right_crop]\n    return cropped_img\n\n# create binary mask\ndef create_mask(image):\n    \n    # threshold and apply morphology close to remove small regions\n    thresh = cv2.threshold(image, 15, 255, cv2.THRESH_BINARY)[1]\n    kernel = cv2.getStructuringElement(cv2.MORPH_ELLIPSE, (31,31))\n    morph = cv2.morphologyEx(thresh, cv2.MORPH_CLOSE, kernel)\n    \n    # get largest contour\n    contours, _ = cv2.findContours(morph, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_NONE)\n    contour = max(contours, key=cv2.contourArea)\n\n    # create a mask from the largest contour\n    mask = np.zeros(image.shape, np.uint8)\n    cv2.drawContours(mask, [contour], -1, 255, cv2.FILLED)\n    return mask\n\n# remove artefacts from image\ndef remove_noise(image, mask):\n    \n    # expand the mask to compensate any missed area\n    kernel = cv2.getStructuringElement(cv2.MORPH_RECT, (101,101))\n    mask = cv2.morphologyEx(mask, cv2.MORPH_DILATE, kernel)\n    mask = cv2.blur(mask,(150, 150))\n    _, mask = cv2.threshold(mask, 150, 255, cv2.THRESH_BINARY)\n    \n    processed_img = cv2.bitwise_and(image, mask)\n    return processed_img\n\n# crop breast area\ndef extract_breast(image, mask):\n    \n    # get largest contour\n    contours, _ = cv2.findContours(mask, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_NONE)\n    contour = max(contours, key=cv2.contourArea)\n    x, y, w, h = cv2.boundingRect(contour)\n    \n    breast_img = image[y:y+h, x:x+w]\n    return breast_img\n\n# set the breast to same side\ndef flip_breast(image, mask):\n    height, width = image.shape[:2]\n    x_center = width // 2\n    y_center = height // 2\n\n    # Sum each column and row\n    col_sum = mask.sum(axis=0)\n    row_sum = mask.sum(axis=1)\n    left_sum = sum(col_sum[0:x_center])\n    right_sum = sum(col_sum[x_center:-1])\n\n    if left_sum > right_sum:\n        return image\n\n    flip_img = cv2.flip(image, 1)  \n    return flip_img\n\n# apply CLAHE\ndef apply_clahe(image, clip=40.0, grid=8):\n    clahe = cv2.createCLAHE(clipLimit=clip, tileGridSize=(grid, grid))\n    clahe_img = clahe.apply(image)\n    return clahe_img\n\n# resize image \ndef image_resize(image, target_width=None, target_height=None):\n    image_dimension = None\n    height, width = image.shape[:2]\n    \n    if target_width == None and target_height == None:\n        return img\n    \n    if target_width == None and target_height != None:\n        ratio = target_height / float(height)\n        image_dimension = (int(width * ratio), target_height)       \n    elif target_width != None and target_height == None:\n        ratio = target_width / float(width)\n        image_dimension = (target_width, int(height * ratio)) \n    else:\n        image_dimension = (target_width, target_height)\n\n    resize_img = cv2.resize(image, image_dimension, interpolation=cv2.INTER_AREA)\n    return resize_img\n\n# resize based on the largest dimension\ndef resize_to_ratio(image, target_size=512):\n    height, width = image.shape[:2]\n    \n    if height > width:\n        resize_img = image_resize(image, target_height = target_size)\n    elif height < width:\n        resize_img = image_resize(image, target_width = target_size)\n    else:\n        resize_img = image_resize(image, target_width = target_size, target_height = target_size)\n    return resize_img\n\n# resize and pad maintaning aspect ratio\ndef resize_and_pad(image, target_width=640, target_height=1280):\n    height, width = image.shape[:2]\n\n    # Calculate scale factors for resizing\n    width_scale = target_width / width\n    height_scale = target_height / height\n\n    # Choose the smallest scale factor to ensure the entire image fits within the target dimensions\n    scale_factor = min(width_scale, height_scale)\n\n    # Calculate the new dimensions of the resized image\n    new_width = int(width * scale_factor)\n    new_height = int(height * scale_factor)\n\n    # Resize the image using cv2.resize()\n    resized_img = cv2.resize(image, (new_width, new_height), interpolation=cv2.INTER_AREA)\n\n    # Add padding to ensure the image is the correct target size\n    left_padding = 0\n    right_padding = target_width - new_width\n    top_padding = 0\n    bottom_padding = target_height - new_height\n    \n    # Pad the image using cv2.copyMakeBorder()\n    padded_img = cv2.copyMakeBorder(resized_img, top_padding, bottom_padding, left_padding, right_padding, cv2.BORDER_CONSTANT, value=0)\n    return padded_img","metadata":{"execution":{"iopub.status.busy":"2023-05-10T08:20:19.688768Z","iopub.execute_input":"2023-05-10T08:20:19.689239Z","iopub.status.idle":"2023-05-10T08:20:19.721518Z","shell.execute_reply.started":"2023-05-10T08:20:19.689203Z","shell.execute_reply":"2023-05-10T08:20:19.719747Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def process_mammogram(path, target_width=IMAGE_WIDTH, target_height=IMAGE_HEIGHT, extension=EXTENSION, debug=True, save=False):\n    # Extract patient and image IDs from the path\n    patient_id, image_id = os.path.split(path)\n    patient_id = os.path.split(patient_id)[-1]\n    image_id = os.path.splitext(image_id)[0]\n    filename = f'{SAVE_DIR}/{patient_id}_{image_id}.{extension}'\n    \n    st = time.perf_counter()\n    \n    # read dicom and shave the borders\n    dcm_img = read_dicom(path)\n    dcm_img = crop_borders(dcm_img)\n    \n    # remove artefacts and extract breast area\n    mask = create_mask(dcm_img)\n    roi_img = remove_noise(dcm_img, mask)\n    roi_img = extract_breast(roi_img, mask)\n    roi_img = flip_breast(roi_img, mask)\n    \n    # feature enhancement using CLAHE\n    clh_img = apply_clahe(roi_img, 1.0, 32)\n    \n    # standardize dimensions and intensity range\n    processed_img = resize_and_pad(clh_img, target_width=target_width, target_height=target_height)\n    processed_img = rescale_intensity(processed_img)\n    \n    if debug:\n        print(f\"patient_id: {patient_id}, image_id: {image_id}, Processing time: {time.time() - st:.2f}s\")\n        print(f\"Filename: {filename}\")\n        display_images((dcm_img, mask, roi_img, processed_img))\n    \n    if save and extension in ['jpg', 'png']:\n        if extension == 'jpg':\n            cv2.imwrite(filename, processed_img, [int(cv2.IMWRITE_JPEG_QUALITY), 95])\n        elif extension == 'png':\n            cv2.imwrite(filename, processed_img, [cv2.IMWRITE_PNG_COMPRESSION, 9])\n            \n    return processed_img","metadata":{"execution":{"iopub.status.busy":"2023-05-10T08:36:28.603571Z","iopub.execute_input":"2023-05-10T08:36:28.604002Z","iopub.status.idle":"2023-05-10T08:36:28.617647Z","shell.execute_reply.started":"2023-05-10T08:36:28.603969Z","shell.execute_reply":"2023-05-10T08:36:28.615708Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"processed_imgs = []\n\nfor path in DCM_FILES:\n    processed_imgs.append(process_mammogram(path, debug=False, save=False))\n    \ndisplay_images(dcm_imgs, 5, 5)\ndisplay_images(processed_imgs, 5, 8)","metadata":{"execution":{"iopub.status.busy":"2023-05-10T08:37:19.282274Z","iopub.execute_input":"2023-05-10T08:37:19.282658Z","iopub.status.idle":"2023-05-10T08:37:28.420638Z","shell.execute_reply.started":"2023-05-10T08:37:19.282629Z","shell.execute_reply":"2023-05-10T08:37:28.419234Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"_ = process_mammogram(DCM_FILES[0])","metadata":{"execution":{"iopub.status.busy":"2023-05-10T08:39:53.630588Z","iopub.execute_input":"2023-05-10T08:39:53.63103Z","iopub.status.idle":"2023-05-10T08:39:56.387375Z","shell.execute_reply.started":"2023-05-10T08:39:53.630999Z","shell.execute_reply":"2023-05-10T08:39:56.383332Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Superpixels segmentation","metadata":{}},{"cell_type":"code","source":"import networkx as nx\n\nfrom skimage import graph\nfrom skimage import segmentation\nfrom skimage.segmentation import mark_boundaries\nfrom skimage.util import img_as_float\nfrom skimage.color import gray2rgb","metadata":{"execution":{"iopub.status.busy":"2023-05-10T08:52:59.421416Z","iopub.execute_input":"2023-05-10T08:52:59.421917Z","iopub.status.idle":"2023-05-10T08:53:00.443257Z","shell.execute_reply.started":"2023-05-10T08:52:59.42188Z","shell.execute_reply":"2023-05-10T08:53:00.442014Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def get_mammogram(path):\n    # read dicom and shave the borders\n    dcm_img = read_dicom(path)\n    dcm_img = crop_borders(dcm_img)\n    \n    # remove artefacts and extract breast area\n    mask = create_mask(dcm_img)\n    roi_img = remove_noise(dcm_img, mask)\n    roi_img = extract_breast(roi_img, mask)\n    \n    # feature enhancement\n    clh_img = apply_clahe(roi_img, 1.0, 4)\n    processed_img = resize_to_ratio(clh_img, target_size=512)\n    return rescale_intensity(processed_img)","metadata":{"execution":{"iopub.status.busy":"2023-05-10T08:53:03.057845Z","iopub.execute_input":"2023-05-10T08:53:03.058309Z","iopub.status.idle":"2023-05-10T08:53:03.066686Z","shell.execute_reply.started":"2023-05-10T08:53:03.058277Z","shell.execute_reply":"2023-05-10T08:53:03.065723Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"mammo_img = get_mammogram(DCM_FILES[1])\ndisplay_image_histogram(mammo_img)","metadata":{"execution":{"iopub.status.busy":"2023-05-10T08:53:04.064536Z","iopub.execute_input":"2023-05-10T08:53:04.065337Z","iopub.status.idle":"2023-05-10T08:53:05.338157Z","shell.execute_reply.started":"2023-05-10T08:53:04.0653Z","shell.execute_reply":"2023-05-10T08:53:05.336935Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Generate superpixels\nsegments_slic = segmentation.slic(mammo_img, n_segments=150, compactness=.1, sigma=0.5, channel_axis=None)\nsegments_quick = segmentation.quickshift(gray2rgb(mammo_img), sigma=1, kernel_size=5, max_dist=15, ratio=1)\nsegments_felz = segmentation.felzenszwalb(mammo_img, scale=100, sigma=0.6, min_size=100)\n\nprint(f'SLIC number of segments         : {len(np.unique(segments_slic))}')\nprint(f'Quickshift number of segments   : {len(np.unique(segments_quick))}')\nprint(f'Felzenszwalb number of segments : {len(np.unique(segments_felz))}')\n\ndisplay_images(( \n    mammo_img, \n    mark_boundaries(mammo_img, segments_slic),\n    mark_boundaries(mammo_img, segments_quick),\n    mark_boundaries(mammo_img, segments_felz)\n), header=['Mammogram', 'SLIC', 'Quickshift', 'Felzenszwalb'])","metadata":{"execution":{"iopub.status.busy":"2023-05-10T08:53:09.726024Z","iopub.execute_input":"2023-05-10T08:53:09.726502Z","iopub.status.idle":"2023-05-10T08:53:14.531314Z","shell.execute_reply.started":"2023-05-10T08:53:09.726465Z","shell.execute_reply":"2023-05-10T08:53:14.53039Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### ✅ Outcomes & Benefits\n\nThis preprocessing-focused approach:\n- 🚀 Improves data quality and training stability\n- 📉 Reduces noise and variability across exams\n- 🧠 Preserves meaningful clinical information\n- 🔧 Enables more reliable and scalable model development\n\nWell-preprocessed data lays the foundation for accurate, interpretable, and clinically relevant machine learning solutions.\n\n\n### 📎 Disclaimer\n\nThis project is intended for educational and research purposes only and does not replace professional medical diagnosis.","metadata":{}}]}