{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":39272,"databundleVersionId":4629629,"sourceType":"competition"},{"sourceId":7263201,"sourceType":"datasetVersion","datasetId":4182300}],"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<div style=\"align: center;\">\n    <br>\n    <img src=\"https://i.imgur.com/M11JoPx.png\" style=\"display:block; margin:auto; width:95%; height:250px;\">\n</div><br><br> \n\n<div style=\"letter-spacing:normal; opacity:1.;\">\n<!--   https://xkcd.com/color/rgb/   -->\n  <p style=\"text-align:center; background-color: lightsalmon; color: Jaguar; border-radius:10px; font-family:monospace; \n            line-height:1.4; font-size:32px; font-weight:bold; text-transform: uppercase; padding: 9px;\">\n            <strong>RSNA Screening Mammography Breast Cancer Detection</strong></p>  \n  \n  <p style=\"text-align:center; background-color:romance; color: Jaguar; border-radius:10px; font-family:monospace; \n            line-height:1.0; font-size:28px; font-weight:normal; text-transform: capitalize; padding: 5px;\"\n     >Deep Learning Module<br>RSNA Mammography Breast Cancer Image Classification\n      <br>Part 3: <a href='https://pypi.org/project/keras-cv-attention-models/1.3.5/'>\"ConvNextV2\"</a> Inference Notebook<br>(Convolutional Neural Network (CNN))</p>    \n</div>\n\n- https://keras.io/api/applications/","metadata":{}},{"cell_type":"markdown","source":"**Dataset Info**\n\n- **Context**\n\nAccording to the WHO, breast cancer is the most commonly occurring cancer worldwide. In 2020 alone, there were 2.3 million new breast cancer diagnoses and 685,000 deaths. Yet breast cancer mortality in high-income countries has dropped by 40% since the 1980s when health authorities implemented regular mammography screening in age groups considered at risk. Early detection and treatment are critical to reducing cancer fatalities, and your machine learning skills could help streamline the process radiologists use to evaluate screening mammograms.\n\nCurrently, early detection of breast cancer requires the expertise of highly-trained human observers, making screening mammography programs expensive to conduct. A looming shortage of radiologists in several countries will likely worsen this problem. Mammography screening also leads to a high incidence of false positive results. This can result in unnecessary anxiety, inconvenient follow-up care, extra imaging tests, and sometimes a need for tissue sampling (often a needle biopsy).\n\nThe competition host, the Radiological Society of North America (RSNA) is a non-profit organization that represents 31 radiologic subspecialties from 145 countries around the world. RSNA promotes excellence in patient care and health care delivery through education, research, and technological innovation.\n\n- **TASK**\n\nIdentify breast cancer. You'll train your model with screening mammograms obtained from regular screening.\n\nYour work improving the automation of detection in screening mammography may enable radiologists to be more accurate and efficient, improving the quality and safety of patient care. It could also help reduce costs and unnecessary medical procedures.\n\nNote: The dataset for this challenge contains radiographic breast images of female subjects.\nThe goal of this competition is to identify cases of breast cancer in mammograms from screening exams. It is important to identify cases of cancer for obvious reasons, but false positives also have downsides for patients. As millions of women get mammograms each year, a useful machine learning tool could help a great many people.\n\n<h4>Can be use 3 library for reading DICOM Images: \n    <br><br><mark>dicomsdl</mark> - <mark>pydicom</mark> - <mark>tensorflow_io</mark></h4>\n\n<h4><a href='https://www.kaggle.com/code/clkmuhammed/dicom-image-reading-benchmark-test'>\n    <mark>Click: DICOM Image Reading Benchmark Test</mark>\n    </a></h4>","metadata":{}},{"cell_type":"markdown","source":"<h4>Table of Content</h4>\n\n\n1. Reading Test Data\n    - Reading CSV Files\n    - List DICOM Images\n    \n    \n2. Data Preprocessing\n    - Convert and Save - DICOM Images to JPEG Images\n    - Apply Parallel Image Processing\n         \n    \n3. CNN MODELING - (Convnextv2 with Pretrain Model)\n    - Build Model Convnextv2 with Pretrain Model Weights\n    \n    \n4. Inference Test Data\n    - Predict Test Data\n    \n    \n5. Save Submission\n\n\n> ___Saved Model___: https://www.kaggle.com/datasets/clkmuhammed/rsna-mammography-breast-cancer-tensorflow-model<br>\n> ___TFRecord Dataset___: https://www.kaggle.com/datasets/clkmuhammed/rsna-mammography-breast-cancer-tfrecord-dataset<br>\n>> Competition Dataset: https://www.kaggle.com/competitions/rsna-breast-cancer-detection/data<br>\n>>> Thanks : https://www.kaggle.com/code/markwijkhuizen/rsna-convnextv2-inference-tensorflow<br>\n>>> Thanks : https://www.kaggle.com/code/bobdegraaf/dicomsdl-voi-lut","metadata":{}},{"cell_type":"markdown","source":"<div style=\"letter-spacing:normal; opacity:1.;\">\n  <h1 style=\"text-align:center; background-color: lightsalmon; color: Jaguar; border-radius:10px; font-family:monospace; border-radius:20px;\n            line-height:1.4; font-size:32px; font-weight:bold; text-transform: uppercase; padding: 9px;\">\n            <strong>1. Import Libraries & Ingest Data</strong></h1>   \n</div>\n\n<h4>pip freeze</h4>","metadata":{}},{"cell_type":"code","source":"# check Model available\n!ls -d -1 \"/kaggle/input/rsna-mammography-breast-cancer-tensorflow-model\"/*","metadata":{"execution":{"iopub.status.busy":"2023-12-22T20:34:44.305985Z","iopub.execute_input":"2023-12-22T20:34:44.306353Z","iopub.status.idle":"2023-12-22T20:34:45.243932Z","shell.execute_reply.started":"2023-12-22T20:34:44.306323Z","shell.execute_reply":"2023-12-22T20:34:45.242834Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%writefile requirements.txt\n# Install Libraries\n/kaggle/input/rsna-mammography-breast-cancer-tensorflow-model/dicomsdl-0.109.2-cp310-cp310-manylinux_2_12_x86_64.manylinux2010_x86_64.whl\n/kaggle/input/rsna-mammography-breast-cancer-tensorflow-model/keras_cv_attention_models-1.3.22-py3-none-any.whl\n/kaggle/input/rsna-mammography-breast-cancer-tensorflow-model/pylibjpeg-1.4.0-py3-none-any.whl\n/kaggle/input/rsna-mammography-breast-cancer-tensorflow-model/pylibjpeg_libjpeg-1.3.4-cp310-cp310-manylinux_2_17_x86_64.manylinux2014_x86_64.whl\n/kaggle/input/rsna-mammography-breast-cancer-tensorflow-model/python_gdcm-3.0.22-cp310-cp310-manylinux_2_17_x86_64.manylinux2014_x86_64.whl","metadata":{"execution":{"iopub.status.busy":"2023-12-22T20:34:45.24638Z","iopub.execute_input":"2023-12-22T20:34:45.246791Z","iopub.status.idle":"2023-12-22T20:34:45.254827Z","shell.execute_reply.started":"2023-12-22T20:34:45.246751Z","shell.execute_reply":"2023-12-22T20:34:45.253973Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os, sys, platform\n\n!{sys.executable} -m pip install -r requirements.txt --no-cache-dir -Uq --no-index --no-deps\nprint(\"Platform:\", platform.system())  # platform.platform()\nprint(\"Python  :\", platform.python_version())  # sys.version\nprint(\"Actv Env:\", os.getenv('CONDA_DEFAULT_ENV', 'Not Found Conda Env'))","metadata":{"execution":{"iopub.status.busy":"2023-12-22T20:34:45.255871Z","iopub.execute_input":"2023-12-22T20:34:45.256126Z","iopub.status.idle":"2023-12-22T20:34:49.535946Z","shell.execute_reply.started":"2023-12-22T20:34:45.256098Z","shell.execute_reply":"2023-12-22T20:34:49.534843Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Print Available Devices - 'TPU', 'GPU', 'CPU'.\n\n>For Inference Notebook:\n>>[GPU T4 Vs GPU P100 | Kaggle | GPU](https://siddhartha01writes.medium.com/gpu-t4-vs-gpu-p100-kaggle-gpu-cd852d56022c#:~:text=In%20general%2C%20the%20T4%20GPU,high%20performance%20and%20memory%20capacity.)","metadata":{}},{"cell_type":"code","source":"import os\n# Disable TensorFlow warnings, before you import tensorflow\nos.environ['TF_CPP_MIN_LOG_LEVEL'] = '3'\n\nimport tensorflow as tf\nprint(\"Tensorflow version \\t\\t:\", tf.__version__)\n\n# Determines the number of threads used by independent non-blocking operations.\n# Tensorflow set number of threads to 1 for speed up in parallell function mapping.\n# 0 means the system picks an appropriate number.\ntf.config.threading.set_inter_op_parallelism_threads(num_threads=1)\nprint(\"Number of threads \\t\\t:\", tf.config.threading.get_inter_op_parallelism_threads())\n\nprint(\"Available devices:\")\n# tf.config.list_physical_devices('GPU')\nfor i, device in enumerate(tf.config.list_logical_devices()):\n    print(\"%d) %s\" % (i, device))","metadata":{"execution":{"iopub.status.busy":"2023-12-22T20:34:49.539054Z","iopub.execute_input":"2023-12-22T20:34:49.539452Z","iopub.status.idle":"2023-12-22T20:35:03.819639Z","shell.execute_reply.started":"2023-12-22T20:34:49.539417Z","shell.execute_reply":"2023-12-22T20:35:03.818665Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"try:\n    import tensorflow as tf\n    cluster_resolver = tf.distribute.cluster_resolver.TPUClusterResolver()\n    tf.config.experimental_connect_to_cluster(cluster_resolver)\n    tf.tpu.experimental.initialize_tpu_system(cluster_resolver)\n    strategy = tf.distribute.TPUStrategy(cluster_resolver)\n    print('Running on TPU ', cluster_resolver.master(), len(tf.config.list_logical_devices('TPU')))\nexcept ValueError:\n    # If there's a GPU avaiable, to use the GPU, otherwise, using the CPU instead.\n    gpus = tf.config.list_logical_devices('GPU')\n    if len(gpus) > 1:\n        strategy = tf.distribute.MirroredStrategy([gpu.name for gpu in gpus])\n        print('Running on multiple GPUs ', gpus)\n    elif len(gpus) == 1:\n        # default strategy that works on CPU and single GPU\n        strategy = tf.distribute.get_strategy()\n        print('Running on single GPU ', gpus[0].name)\n    else:\n        # default strategy that works on CPU and single GPU\n        strategy = tf.distribute.get_strategy()\n        print('Running on CPU')\nfinally:\n    print(\"Number of accelerators: \", strategy.num_replicas_in_sync)\n\n# Enable Just-In-Time (JIT) computation graph at runtime, just before they are executed. \n# The idea is to convert the computation graph into machine code optimized for the specific hardware it will run on.\ntf.config.optimizer.set_jit(True)\ntf.config.optimizer.get_jit()","metadata":{"execution":{"iopub.status.busy":"2023-12-22T20:35:03.820802Z","iopub.execute_input":"2023-12-22T20:35:03.821334Z","iopub.status.idle":"2023-12-22T20:35:03.836991Z","shell.execute_reply.started":"2023-12-22T20:35:03.821307Z","shell.execute_reply":"2023-12-22T20:35:03.835863Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Importing Related Libraries","metadata":{}},{"cell_type":"code","source":"from kaggle_datasets import KaggleDatasets\nfrom keras_cv_attention_models import convnext\nimport kecam; print(f'keras_cv_attention_models: {kecam.__version__}')\n\nimport cv2; print('cv2', cv2.__version__), cv2.setNumThreads(1)\n# If this code prints 0, then your current installation of OpenCV does not have CUDA support.\n# print(cv2.getBuildInformation())\nprint('cv2 cuda:', cv2.cuda.getCudaEnabledDeviceCount())\n\nimport pylibjpeg; print(\"pylibjpeg\", pylibjpeg.__version__)\nimport libjpeg; print(\"libjpeg\", libjpeg.__version__)\nimport dicomsdl; print(\"dicomsdl\", dicomsdl.DICOMSDL_VERSION)\nimport pydicom; print(\"pydicom\", pydicom.__version__);\n# VOI LUT (if available by DICOM device) is used to transform raw DICOM data to \"human-friendly\" view\n# https://pydicom.github.io/pydicom/dev/reference/handlers.html#pixel-data-utilities\n# https://pydicom.github.io/pydicom/stable/old/working_with_pixel_data.html\nfrom pydicom.pixel_data_handlers.util import apply_voi_lut, apply_voi\nfrom pydicom.pixel_data_handlers.util import apply_color_lut, apply_modality_lut\nimport gdcm; print(\"gdcm\", gdcm.GDCM_VERSION);","metadata":{"execution":{"iopub.status.busy":"2023-12-22T20:35:03.840215Z","iopub.execute_input":"2023-12-22T20:35:03.840574Z","iopub.status.idle":"2023-12-22T20:35:04.682816Z","shell.execute_reply.started":"2023-12-22T20:35:03.84054Z","shell.execute_reply":"2023-12-22T20:35:04.681896Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport matplotlib as mpl\nimport matplotlib.pyplot as plt\n# !pip install scikit-plot -Uq\nimport scikitplot as skplt\n\nimport re\nimport time\nimport random\nimport datetime\nimport tempfile\nimport importlib\nfrom glob import glob\nfrom typing import cast\nfrom pathlib import Path\nfrom tqdm.auto import tqdm\nfrom multiprocessing import cpu_count\nimport joblib\nimport pickle\n\nimport gc\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2023-12-22T20:35:04.684426Z","iopub.execute_input":"2023-12-22T20:35:04.684738Z","iopub.status.idle":"2023-12-22T20:35:05.329973Z","shell.execute_reply.started":"2023-12-22T20:35:04.684712Z","shell.execute_reply":"2023-12-22T20:35:05.329071Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Parameters","metadata":{}},{"cell_type":"code","source":"MODEL_PATH = \"/kaggle/input/rsna-mammography-breast-cancer-tensorflow-model\"\nDATA_DIR = \"/kaggle/input/rsna-breast-cancer-detection\"\nprint('MODEL_PATH  :', MODEL_PATH)\nprint('DATA_DIR :', DATA_DIR)\n\n# Image Format and Config\nIMAGE_FORMAT        = 'JPG'\nIMAGE_QUALITY       = 95\nN_SAMPLES_TFRECORDS = 548\nTARGET_HEIGHT, TARGET_WIDTH, N_CHANNELS = ( 1280, 768, 1 )\nINPUT_SHAPE  = (TARGET_HEIGHT, TARGET_WIDTH, N_CHANNELS)\n\nTHRESHOLD_BEST = 0.49869878590106964","metadata":{"execution":{"iopub.status.busy":"2023-12-22T20:35:05.33118Z","iopub.execute_input":"2023-12-22T20:35:05.33151Z","iopub.status.idle":"2023-12-22T20:35:05.337623Z","shell.execute_reply.started":"2023-12-22T20:35:05.331484Z","shell.execute_reply":"2023-12-22T20:35:05.336706Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Reading Test Data","metadata":{}},{"cell_type":"markdown","source":"## Reading CSV Files","metadata":{}},{"cell_type":"code","source":"# test files include 1 patient and have 2 images both 'L/R' laterality\ntest_df              = pd.read_csv(f'{DATA_DIR}/test.csv')\n# submission files include 1 patient and have 1 average possibilities both 'L/R' laterality \nsample_submission_df = pd.read_csv(f'{DATA_DIR}/sample_submission.csv')\n\ntest_df.shape, sample_submission_df.shape","metadata":{"execution":{"iopub.status.busy":"2023-12-22T20:35:05.338504Z","iopub.execute_input":"2023-12-22T20:35:05.338791Z","iopub.status.idle":"2023-12-22T20:35:05.370689Z","shell.execute_reply.started":"2023-12-22T20:35:05.338768Z","shell.execute_reply":"2023-12-22T20:35:05.369828Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display(test_df.head(), sample_submission_df.head())","metadata":{"execution":{"iopub.status.busy":"2023-12-22T20:35:05.373439Z","iopub.execute_input":"2023-12-22T20:35:05.373725Z","iopub.status.idle":"2023-12-22T20:35:05.397118Z","shell.execute_reply.started":"2023-12-22T20:35:05.373701Z","shell.execute_reply":"2023-12-22T20:35:05.396239Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df.info(), print('\\n'), sample_submission_df.info(), ","metadata":{"execution":{"iopub.status.busy":"2023-12-22T20:35:05.398254Z","iopub.execute_input":"2023-12-22T20:35:05.398502Z","iopub.status.idle":"2023-12-22T20:35:05.425566Z","shell.execute_reply.started":"2023-12-22T20:35:05.398481Z","shell.execute_reply":"2023-12-22T20:35:05.424675Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## List DICOM Images","metadata":{}},{"cell_type":"code","source":"test_file_paths = tf.io.gfile.glob(f'{DATA_DIR}/test_images/*/*.dcm')\n\nprint(f'Test size images : {len(test_file_paths):<10}')","metadata":{"execution":{"iopub.status.busy":"2023-12-22T20:35:05.426936Z","iopub.execute_input":"2023-12-22T20:35:05.427443Z","iopub.status.idle":"2023-12-22T20:35:05.440559Z","shell.execute_reply.started":"2023-12-22T20:35:05.427409Z","shell.execute_reply":"2023-12-22T20:35:05.439672Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data Preprocessing","metadata":{}},{"cell_type":"code","source":"# Source: https://www.kaggle.com/code/bobdegraaf/dicomsdl-voi-lut\n# Source: https://www.kaggle.com/code/dangnh0611/1st-place-submission-code\ndef apply_voi_lut_dicomsdl(image_data, ds, isDicomsdl=True):\n    # Check and Load only the variables we need\n    meta_data = ds.getPixelDataInfo().keys() if isDicomsdl else ds\n        \n    if all(i in meta_data for i in ['WindowCenter', 'WindowWidth',\n                                    'BitsStored', 'RescaleSlope', 'RescaleIntercept',\n                                    'PixelRepresentation']):\n        # For sigmoid it's a list, otherwise a single value\n        center = np.array(ds.WindowCenter, dtype=np.float64).flatten()[0]\n        width  = np.array(ds.WindowWidth, dtype=np.float64).flatten()[0]\n        \n        # Set Unsigned y_min, max & range\n        y_min, y_max = 0.0, float(2**ds.BitsStored - 1)\n        # Set Slope\n        slope     = ds.RescaleSlope\n        intercept = ds.RescaleIntercept\n        if slope is not None and intercept is not None:\n            # Otherwise its the actual data range\n            y_min = y_min * float(slope) + float(intercept)\n            y_max = y_max * float(slope) + float(intercept)\n        y_range = y_max - y_min\n        \n        try:\n            # Function with default LINEAR (so for Nan, it will use LINEAR)\n            # if nan can be raise error in pydicom\n            voi_func = ds.VOILUTFunction\n        except:\n            voi_func = None\n        finally:\n            if voi_func is None: voi_func = 'LINEAR'\n\n        # Modality Lut, float64 needed (default) or just float32 ?\n        arr = image_data.astype('float64')\n        \n        # When performing a windowing operation only 'MONOCHROME1' and 'MONOCHROME2'\n        if ds.PhotometricInterpretation in ['MONOCHROME1', 'MONOCHROME2']:\n            # May be LINEAR (default), LINEAR_EXACT, SIGMOID or not present, VM 1\n            if voi_func in ['LINEAR', 'LINEAR_EXACT']:\n                # PS3.3 C.11.2.1.2.1 and C.11.2.1.3.2\n                if voi_func == 'LINEAR':\n                    if width < 1:\n                        raise ValueError(\n                            \"The (0028,1051) Window Width must be greater than or \"\n                            \"equal to 1 for a 'LINEAR' windowing operation\"\n                        )\n                    center -= 0.5\n                    width -= 1\n                elif width <= 0:\n                    raise ValueError(\n                        \"The (0028,1051) Window Width must be greater than 0 \"\n                        \"for a 'LINEAR_EXACT' windowing operation\"\n                    )\n\n                below = arr <= (center - width / 2)\n                above = arr > (center + width / 2)\n                between = np.logical_and(~below, ~above)\n\n                arr[below] = y_min\n                arr[above] = y_max\n                if between.any():\n                    arr[between] = (\n                        ((arr[between] - center) / width + 0.5) * y_range + y_min\n                    )\n            elif voi_func == 'SIGMOID':\n                # PS3.3 C.11.2.1.3.1\n                if width <= 0:\n                    raise ValueError(\n                        \"The (0028,1051) Window Width must be greater than 0 \"\n                        \"for a 'SIGMOID' windowing operation\"\n                    )\n\n                arr = y_range / (1 + np.exp(-4 * (arr - center) / width)) + y_min\n            else:\n                raise ValueError(\n                    f\"Unsupported (0028,1056) VOI LUT Function value '{voi_func}'\"\n                )\n        \n        # Normalize to have 0 as background, some images are reversed where 0 is max intensity\n        if ds.PhotometricInterpretation == 'MONOCHROME1':\n            arr = np.amax(arr) - arr\n    return arr.astype('uint16')","metadata":{"execution":{"iopub.status.busy":"2023-12-22T20:35:05.441894Z","iopub.execute_input":"2023-12-22T20:35:05.442163Z","iopub.status.idle":"2023-12-22T20:35:05.456733Z","shell.execute_reply.started":"2023-12-22T20:35:05.442138Z","shell.execute_reply":"2023-12-22T20:35:05.455769Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Crop Image drop out some black area \ndef analyze_components(img_data_voi, ds, filtering=False, threshold=cv2.THRESH_BINARY, debug=False):\n    # Normalization (Min-Max Scaling)\n    img_data_8u = cv2.normalize(img_data_voi, None, 0, 255, cv2.NORM_MINMAX, dtype=cv2.CV_8U)\n    \n    if filtering:\n        # (Optional) Gaussian filtering to reduce noise\n        # essential to strike a balance between noise reduction and preserving relevant features.\n        blur = cv2.GaussianBlur(src=img_data_8u, ksize=(5,5), sigmaX=0)\n    else:\n        blur = img_data_8u\n    \n    if threshold in range(255):\n        # Threshold the image to create a binary mask for the black background\n        # black_background_mask = cv2.adaptiveThreshold(\n        #     src=blur, maxValue=255, \n        #     adaptiveMethod=cv2.ADAPTIVE_THRESH_MEAN_C,\n        #     thresholdType=cv2.THRESH_BINARY, blockSize=11, C=2,)\n        retval, black_background_mask = cv2.threshold(\n            src=blur,       # Input image (should be grayscale)\n            thresh=25,      # Threshold value\n            maxval=255,     # Value assigned to pixels exceeding the threshold\n            type=threshold,  # Binary thresholding type\n        )\n    else:\n        # build image to binary mask 8-bit unsigned integer - for black area min as 0\n        # binary_mask_8u = cv2.convertScaleAbs(blur)\n        black_background_mask = (blur > 25).astype(np.uint8)\n    \n    # Apply connected components labeling\n    # 4-connectivity for well-defined edges, 8-connectivity for irregular and fuzzy objects in images.\n    output = cv2.connectedComponentsWithStats(image=black_background_mask, connectivity=8, ltype=cv2.CV_32S)\n    \n    # Unpack the output\n    # The first cell is the number of labels, The second cell is the label matrix\n    # The third cell is the stat matrix,      The fourth cell is the centroid matrix\n    retval, labels, stats, centroids = output\n    \n    # Find the index of the largest connected component (excluding background) (corresponds to the breast data)\n    # Skip background label (0), There is at least one connected component (excluding background)\n    if retval > 1:\n        largest_component_index = np.argmax(stats[1:, cv2.CC_STAT_AREA]) + 1\n        # Extract the region of interest for the largest connected component using slicing\n        x, y, width, height, area = stats[largest_component_index]\n        # Crop the image using TensorFlow 3D or NumPy indexing\n        # img_data_roi = tf.image.crop_to_bounding_box(np.expand_dims(img_data_32f, 2), y, x, height, width).numpy()\n        img_data_roi = img_data_8u[y:y+height, x:x+width]\n    else:\n        # Handle the case where no connected components are found\n        img_data_roi = img_data_8u\n    \n    # Show Processed Image\n    if debug:\n        plt.subplot(121)\n        plt.imshow(black_background_mask);\n        plt.title(f'image mask\\nthreshold: {threshold}\\nfiltering {filtering}')\n        \n        plt.subplot(122)\n        plt.imshow(img_data_roi);\n        plt.colorbar()\n        plt.title(f'cropped image\\nshape: {img_data_roi.shape}')\n        \n        plt.tight_layout()\n        plt.show();        \n        # Loop over first 5 detected components and print their statistics\n        for label in range(1, retval)[: 2]:\n            area     = stats[label, cv2.CC_STAT_AREA]\n            centroid = centroids[label]\n            x, y, width, height, _ = stats[label]\n            print(f\"\"\"Component {label}: \n            Area={area}, Centroid=({centroid[0]:.2f}, {centroid[1]:.2f}), \n            Bounding Box=({x}, {y}, {width}, {height})\n            \"\"\")\n    \n    return img_data_roi","metadata":{"execution":{"iopub.status.busy":"2023-12-22T20:35:05.457957Z","iopub.execute_input":"2023-12-22T20:35:05.458245Z","iopub.status.idle":"2023-12-22T20:35:05.474665Z","shell.execute_reply.started":"2023-12-22T20:35:05.458216Z","shell.execute_reply":"2023-12-22T20:35:05.47377Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# optimized: same result analyze_components, same time \ndef image_findNonZero(img_data_voi):\n    # Normalization (Min-Max Scaling)\n    img_data_8u = cv2.normalize(img_data_voi, None, 0, 255, cv2.NORM_MINMAX, dtype=cv2.CV_8U)\n    # find non-blank pixels\n    coords = cv2.findNonZero(img_data_8u)\n    # Find the bounding box that contains all non-blank pixels\n    x, y, width, height = cv2.boundingRect(coords)\n    \n    # Crop the image using NumPy indexing\n    img_data_roi = img_data_8u[y:y+height, x:x+width]\n    return img_data_roi","metadata":{"execution":{"iopub.status.busy":"2023-12-22T20:35:05.475769Z","iopub.execute_input":"2023-12-22T20:35:05.476007Z","iopub.status.idle":"2023-12-22T20:35:05.488654Z","shell.execute_reply.started":"2023-12-22T20:35:05.475986Z","shell.execute_reply":"2023-12-22T20:35:05.48786Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Convert and Save Images - DICOM to JPEG","metadata":{}},{"cell_type":"code","source":"# file path to image, label\ndef load_and_preprocess_image(file_path, scale_image=False, save_image=True, debug=False):    \n    # Load the raw data from the file, dicomsdl - more faster\n    img_dicom    = dicomsdl.open(file_path)\n    # Get image array uint16, dicomsdl\n    img_data_16u = img_dicom.pixelData(storedvalue=True)    \n    # apply_voi_lut: commonly used to adjust the contrast and enhance specific regions of interest in an image.\n    img_data_voi = apply_voi_lut_dicomsdl(img_data_16u, img_dicom, isDicomsdl=True)    \n    # Cutting out the breast data\n    # img_data_voi_roi = image_findNonZero(img_data_voi)\n    img_data_voi_roi = analyze_components(img_data_voi, img_dicom)\n    \n    # Resize the breast data\n    img_data_resized = cv2.resize(\n        src   = img_data_voi_roi,\n        dsize = INPUT_SHAPE[:2][::-1], # attention cv2 size reverse\n        interpolation = cv2.INTER_LANCZOS4,\n    )\n    # Convert image to a 3D\n    img_data_resized = np.expand_dims(img_data_resized, 2)\n    \n    # Save Only\n    if save_image:\n        patient_id, image_id = re.split('[\\./]', file_path)[-3:-1]\n        # Check whether the specified path exists or not\n        if not os.path.exists(patient_id):\n            # Create a new directory because it does not exist\n            os.makedirs(patient_id, exist_ok=True)\n            \n        if IMAGE_FORMAT == 'PNG':\n            cv2.imwrite(f'{patient_id}/{image_id}.png', img_data_resized)\n        else:\n            cv2.imwrite(f'{patient_id}/{image_id}.jpg', img_data_resized, [cv2.IMWRITE_JPEG_QUALITY, IMAGE_QUALITY])\n            \n    # Show Processed Image\n    if debug:\n        plt.imshow(img_data_resized, cmap='bone')","metadata":{"execution":{"iopub.status.busy":"2023-12-22T20:35:05.489812Z","iopub.execute_input":"2023-12-22T20:35:05.490178Z","iopub.status.idle":"2023-12-22T20:35:05.504481Z","shell.execute_reply.started":"2023-12-22T20:35:05.490147Z","shell.execute_reply":"2023-12-22T20:35:05.503596Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# samples analyze_components\n# 775 ms ± 10 ms per loop (mean ± std. dev. of 10 runs, 10 loops each)\n# %timeit -r 10 -n 10 -p 3 load_and_preprocess_image(test_file_paths[0], debug=True)","metadata":{"execution":{"iopub.status.busy":"2023-12-22T20:35:05.505514Z","iopub.execute_input":"2023-12-22T20:35:05.505775Z","iopub.status.idle":"2023-12-22T20:35:05.517827Z","shell.execute_reply.started":"2023-12-22T20:35:05.505753Z","shell.execute_reply":"2023-12-22T20:35:05.516974Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# CNN MODELING - (Convnextv2 with Pretrain Model)","metadata":{}},{"cell_type":"markdown","source":"## Load Pretrained Model","metadata":{}},{"cell_type":"code","source":"# MODEL_DIR = \"/kaggle/input/rsna-mammography-breast-cancer-tensorflow-model/breast_cancer_tfmodel_legacy.h5\"\nMODEL_DIR = \"/kaggle/input/rsna-mammography-breast-cancer-tensorflow-model/breast_cancer_tfmodel_legacy_v2.h5\"\n\nprint('MODEL_DIR:', MODEL_DIR)","metadata":{"execution":{"iopub.status.busy":"2023-12-22T20:35:27.393094Z","iopub.execute_input":"2023-12-22T20:35:27.393474Z","iopub.status.idle":"2023-12-22T20:35:27.398975Z","shell.execute_reply.started":"2023-12-22T20:35:27.393444Z","shell.execute_reply":"2023-12-22T20:35:27.397772Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tf.keras.backend.clear_session()\ngc.collect()\n\nwith strategy.scope():\n    # Load the model (architecture + weights)\n    model = tf.keras.models.load_model(MODEL_DIR, compile=False)    \n    # Load pretrained Model Weights\n    # model.load_weights('/kaggle/input/rsna-mammography-breast-cancer-tensorflow-model/checkpoint-00.weights.h5')\n    # Set the entire model to be non-trainable\n    model.trainable = False\n    # Compile the model manually\n    model.compile()    \n    model.summary()","metadata":{"execution":{"iopub.status.busy":"2023-12-22T20:35:29.392668Z","iopub.execute_input":"2023-12-22T20:35:29.393504Z","iopub.status.idle":"2023-12-22T20:35:43.799009Z","shell.execute_reply.started":"2023-12-22T20:35:29.393472Z","shell.execute_reply":"2023-12-22T20:35:43.798084Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# tf.keras.utils.plot_model(model, to_file='model_plot4a.png', show_shapes=True, show_layer_names=True)","metadata":{"execution":{"iopub.status.busy":"2023-12-22T20:36:11.631842Z","iopub.execute_input":"2023-12-22T20:36:11.632254Z","iopub.status.idle":"2023-12-22T20:36:11.636931Z","shell.execute_reply.started":"2023-12-22T20:36:11.63222Z","shell.execute_reply":"2023-12-22T20:36:11.63583Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Inference Test Data DCM2JPG","metadata":{}},{"cell_type":"markdown","source":"## Parallel Image Processing","metadata":{}},{"cell_type":"code","source":"# Preprocess a single image and saves it\ndef preprocess_and_save_image(args):\n    (np.random.rand()>0.9) and gc.collect()\n    (patient_id, laterality), group = args\n    for row_idx, row_data in group.iterrows():\n        # Define file path\n        image_path = f\"{DATA_DIR}/test_images/{row_data['patient_id']}/{row_data['image_id']}.dcm\"\n        # Call Image Save Function\n        load_and_preprocess_image(image_path)\n        \n        \ndef convert_images_dcm2jpg():   \n    # Preprocess all images in parallel using Joblib\n    jobs = [joblib.delayed(preprocess_and_save_image)(args) for args in test_df.groupby(['patient_id', 'laterality'])]\n    SUBMISSION_ROWS = joblib.Parallel(\n        n_jobs = -1,                 # -1, cpu_count()\n        verbose= 0,\n        backend= 'multiprocessing',  # 'multiprocessing', 'threading'\n        prefer = 'threads',          # processes\n    )(jobs)","metadata":{"execution":{"iopub.status.busy":"2023-12-22T20:36:15.226315Z","iopub.execute_input":"2023-12-22T20:36:15.227183Z","iopub.status.idle":"2023-12-22T20:36:15.233983Z","shell.execute_reply.started":"2023-12-22T20:36:15.227151Z","shell.execute_reply":"2023-12-22T20:36:15.233025Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 1.92 s ± 58.9 ms per loop (mean ± std. dev. of 7 runs, 2 loops each)\n# %timeit -r 7 -n 2 -p 3 convert_images_dcm2jpg()\nconvert_images_dcm2jpg()","metadata":{"execution":{"iopub.status.busy":"2023-12-22T20:36:16.156145Z","iopub.execute_input":"2023-12-22T20:36:16.156559Z","iopub.status.idle":"2023-12-22T20:36:18.383596Z","shell.execute_reply.started":"2023-12-22T20:36:16.156527Z","shell.execute_reply":"2023-12-22T20:36:18.382493Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Inference Submission","metadata":{}},{"cell_type":"code","source":"SUBMISSION_ROWS = []\n\n# Iterate over all patient_id/laterality combinations groups\nfor idx, ((patient_id, laterality), group) in enumerate(tqdm(test_df.groupby(['patient_id', 'laterality']))):\n    (np.random.rand()>0.9) and gc.collect()\n    # Cancer target is mean of predicted cancer values\n    image_datas, cancer = [], []\n    # Iterate over all scans in group\n    for row_idx, row in group.iterrows():\n        # Load Image\n        patient_id, image_id = row['patient_id'], row['image_id']\n        image_data = cv2.imread(f'{patient_id}/{image_id}.{IMAGE_FORMAT.lower()}', -1)\n        image_datas.append(image_data)\n        # Expand to Batch HxW -> 1xHxWx1\n        image_data = np.expand_dims(image_data, [0, 3])\n        # Make Prediction\n        cancer.append( model.predict_on_batch({'image': image_data}).squeeze().tolist() )\n        # Remove Image\n        os.remove(f'{patient_id}/{image_id}.{IMAGE_FORMAT.lower()}')\n        \n    # Show First Patients 2 group\n    if idx < 2:\n        fig, axes = plt.subplots(nrows=1, ncols=len(image_datas), figsize=(5,5))\n        fig.subplots_adjust(hspace=0.1, wspace=0.05)\n        for n, ax in enumerate(axes.flat):\n            ax.imshow(image_datas[n], cmap='bone') # binary, gray, bone\n            ax.set_title(f'{laterality}: {cancer[n]:.1e}')\n            ax.axis('off')            \n        plt.show()\n        fig.tight_layout()\n        \n    # Add SUBMISSION_ROWS\n    # Cancer target is mean of predicted cancer values\n    # We Save average probabilities as Prediction [0 1], based on BEST THRESHOLD\n    SUBMISSION_ROWS.append({\n        'prediction_id': f'{patient_id}_{laterality}',\n        'cancer': np.int8(np.mean(cancer) >= THRESHOLD_BEST),\n    })","metadata":{"execution":{"iopub.status.busy":"2023-12-22T20:36:18.385648Z","iopub.execute_input":"2023-12-22T20:36:18.385968Z","iopub.status.idle":"2023-12-22T20:36:33.355209Z","shell.execute_reply.started":"2023-12-22T20:36:18.38594Z","shell.execute_reply":"2023-12-22T20:36:33.354166Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Save Submission","metadata":{}},{"cell_type":"code","source":"# Create DataFrame from submission rows\nsubmission_df = pd.DataFrame(SUBMISSION_ROWS)\n\ndisplay(submission_df.head()), print('\\n'), submission_df.info(),","metadata":{"execution":{"iopub.status.busy":"2023-12-22T20:36:33.357456Z","iopub.execute_input":"2023-12-22T20:36:33.358147Z","iopub.status.idle":"2023-12-22T20:36:33.378471Z","shell.execute_reply.started":"2023-12-22T20:36:33.358111Z","shell.execute_reply":"2023-12-22T20:36:33.377187Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission_df.to_csv(\"submission.csv\", index=False)","metadata":{"execution":{"iopub.status.busy":"2023-12-22T20:36:33.380081Z","iopub.execute_input":"2023-12-22T20:36:33.380474Z","iopub.status.idle":"2023-12-22T20:36:33.387961Z","shell.execute_reply.started":"2023-12-22T20:36:33.380438Z","shell.execute_reply":"2023-12-22T20:36:33.387262Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Download Link","metadata":{}},{"cell_type":"code","source":"# from IPython.display import FileLink, FileLinks\n# local_file = FileLink(r'submission.csv', result_html_prefix=\"Click here to download: \")\n# local_file","metadata":{"execution":{"iopub.status.busy":"2023-12-17T19:04:24.217116Z","iopub.execute_input":"2023-12-17T19:04:24.218042Z","iopub.status.idle":"2023-12-17T19:04:24.221615Z","shell.execute_reply.started":"2023-12-17T19:04:24.217996Z","shell.execute_reply":"2023-12-17T19:04:24.220724Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# End of the Project","metadata":{}}]}