{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# !pip install pydicom nibabel pyarrow fastparquet\n# pyarrow fastparquet jupyter ipywidgets widgetsnbextension\n# ! pip install --upgrade pip","metadata":{"execution":{"iopub.status.busy":"2023-10-13T12:14:47.365636Z","iopub.execute_input":"2023-10-13T12:14:47.366202Z","iopub.status.idle":"2023-10-13T12:14:47.371787Z","shell.execute_reply.started":"2023-10-13T12:14:47.366161Z","shell.execute_reply":"2023-10-13T12:14:47.370842Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os\nfrom glob import glob\nimport pydicom\nimport nibabel as nib\nimport numpy as np\nimport pandas as pd\nfrom matplotlib import pyplot as plt\nfrom sklearn.model_selection import train_test_split\nimport tensorflow as tf\nimport cv2\nfrom tqdm.notebook import tqdm\nimport gc\nfrom joblib import Parallel, delayed\nfrom sklearn.model_selection import train_test_split","metadata":{"execution":{"iopub.status.busy":"2023-10-13T12:14:47.373667Z","iopub.execute_input":"2023-10-13T12:14:47.374261Z","iopub.status.idle":"2023-10-13T12:14:47.398137Z","shell.execute_reply.started":"2023-10-13T12:14:47.37423Z","shell.execute_reply":"2023-10-13T12:14:47.397021Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"main_dir = \"/kaggle/input/rsna-2023-abdominal-trauma-detection/\"\nprint(os.listdir(main_dir))","metadata":{"execution":{"iopub.status.busy":"2023-10-13T12:14:47.399513Z","iopub.execute_input":"2023-10-13T12:14:47.399977Z","iopub.status.idle":"2023-10-13T12:14:47.416805Z","shell.execute_reply.started":"2023-10-13T12:14:47.399943Z","shell.execute_reply":"2023-10-13T12:14:47.415549Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data","metadata":{}},{"cell_type":"markdown","source":"**train.csv** Target labels for the train set. Note that patients labeled healthy may still have other medical issues, such as cancer or broken bones, that don't happen to be covered by the competition labels.\n\n- **patient_id** - A unique ID code for each patient.\n- **[bowel/extravasation]_[healthy/injury]** - The two injury types with binary targets.\n- **[kidney/liver/spleen]_[healthy/low/high]** - The three injury types with three target levels.\n- **any_injury** - Whether the patient had any injury at all.","metadata":{}},{"cell_type":"code","source":"train_df = pd.read_csv(main_dir+\"train.csv\")\ntrain_df","metadata":{"execution":{"iopub.status.busy":"2023-10-13T12:14:47.418264Z","iopub.execute_input":"2023-10-13T12:14:47.419175Z","iopub.status.idle":"2023-10-13T12:14:47.448266Z","shell.execute_reply.started":"2023-10-13T12:14:47.419137Z","shell.execute_reply":"2023-10-13T12:14:47.447191Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**[train/test]_series_meta.csv** Each patient may have been scanned once or twice. Each scan contains a series of images.\n\n- **patient_id** - A unique ID code for each patient.\n- **series_id** - A unique ID code for each scan.\n- **aortic_hu** - The volume of the aorta in hounsfield units. This acts as a reliable proxy for when the scan was. For a multiphasic CT scan, the higher value indicates the late arterial phase.\nincomplete_organ - True if one or more organs wasn't fully covered by the scan. This label is only provided for the train set.","metadata":{}},{"cell_type":"code","source":"train_serises_df = pd.read_csv(main_dir+\"train_series_meta.csv\")\ntrain_serises_df","metadata":{"execution":{"iopub.status.busy":"2023-10-13T12:14:47.451992Z","iopub.execute_input":"2023-10-13T12:14:47.452686Z","iopub.status.idle":"2023-10-13T12:14:47.472912Z","shell.execute_reply.started":"2023-10-13T12:14:47.452629Z","shell.execute_reply":"2023-10-13T12:14:47.472039Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**image_level_labels.csv** Train only. Identifies specific images that contain either bowel or extravasation injuries.\n\n- **patient_id** - A unique ID code for each patient.\n- **series_id** - A unique ID code for each scan.\n- **instance_number** - The image number within the scan. The lowest instance number for many series is above zero as the original scans were cropped to the abdomen.\n- **injury_name** - The type of injury visible in the frame.","metadata":{}},{"cell_type":"code","source":"image_labels_df = pd.read_csv(main_dir+\"image_level_labels.csv\")\nimage_labels_df","metadata":{"execution":{"iopub.status.busy":"2023-10-13T12:14:47.473983Z","iopub.execute_input":"2023-10-13T12:14:47.474882Z","iopub.status.idle":"2023-10-13T12:14:47.497337Z","shell.execute_reply.started":"2023-10-13T12:14:47.474853Z","shell.execute_reply":"2023-10-13T12:14:47.496481Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**[train/test]_dicom_tags.parquet** DICOM tags from every image, extracted with Pydicom.","metadata":{}},{"cell_type":"code","source":"train_dicom = pd.read_parquet(main_dir + \"train_dicom_tags.parquet\")\ntrain_dicom.head()","metadata":{"execution":{"iopub.status.busy":"2023-10-13T12:14:47.498325Z","iopub.execute_input":"2023-10-13T12:14:47.498967Z","iopub.status.idle":"2023-10-13T12:14:53.123385Z","shell.execute_reply.started":"2023-10-13T12:14:47.498934Z","shell.execute_reply":"2023-10-13T12:14:53.122281Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**segmentations** Model generated pixel-level annotations of the relevant organs and some major bones for a subset of the scans in the training set. This data is provided in the nifti file format. The filenames are series IDs.","metadata":{}},{"cell_type":"code","source":"seg_file = main_dir+\"segmentations/12114.nii\"\n\nseg_img = nib.load(seg_file)\narray = seg_img.get_fdata()\narray.shape","metadata":{"execution":{"iopub.status.busy":"2023-10-13T12:14:53.124884Z","iopub.execute_input":"2023-10-13T12:14:53.125606Z","iopub.status.idle":"2023-10-13T12:14:53.563078Z","shell.execute_reply.started":"2023-10-13T12:14:53.125566Z","shell.execute_reply":"2023-10-13T12:14:53.561988Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"slice_img = array[:,:,array.shape[2] // 2]\n\nplt.imshow(slice_img, cmap='gray')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-13T12:14:53.564513Z","iopub.execute_input":"2023-10-13T12:14:53.564909Z","iopub.status.idle":"2023-10-13T12:14:53.840566Z","shell.execute_reply.started":"2023-10-13T12:14:53.564872Z","shell.execute_reply":"2023-10-13T12:14:53.83858Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**[train/test]_images/[patient_id]/[series_id]/[image_instance_number].dcm** The CT scan data, in DICOM format. Scans from dozens of different CT machines have been reprocessed to use the run length encoded lossless compression format but retain other differences such as the number of bits per pixel, pixel range, and pixel representation. Expect to see roughly 1,100 patients in the test set.","metadata":{}},{"cell_type":"code","source":"image_ex = main_dir+\"train_images/10494/65369/100.dcm\"","metadata":{"execution":{"iopub.status.busy":"2023-10-13T12:14:53.842469Z","iopub.execute_input":"2023-10-13T12:14:53.843427Z","iopub.status.idle":"2023-10-13T12:14:53.848602Z","shell.execute_reply.started":"2023-10-13T12:14:53.843365Z","shell.execute_reply":"2023-10-13T12:14:53.847472Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def plot_image_dcm(Image_path):\n    ds = pydicom.dcmread(Image_path)\n    plt.imshow(ds.pixel_array, cmap=plt.cm.bone)\n    plt.show()\n    \nplot_image_dcm(image_ex)","metadata":{"execution":{"iopub.status.busy":"2023-10-13T12:14:53.849995Z","iopub.execute_input":"2023-10-13T12:14:53.850517Z","iopub.status.idle":"2023-10-13T12:14:54.201435Z","shell.execute_reply.started":"2023-10-13T12:14:53.850479Z","shell.execute_reply":"2023-10-13T12:14:54.200524Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Preparing data","metadata":{}},{"cell_type":"code","source":"# Checking if patients are repeated by finding the number of unique patient IDs\n# num_rows = train_serises_df.shape[0]\n# unique_patients = train_serises_df[\"patient_id\"].nunique()\n# test_serises_df = pd.read_csv(main_dir+\"test_series_meta.csv\") \n# print(f\"{num_rows=}\")\n# print(f\"{unique_patients=}\")","metadata":{"execution":{"iopub.status.busy":"2023-10-13T12:14:54.202599Z","iopub.execute_input":"2023-10-13T12:14:54.203355Z","iopub.status.idle":"2023-10-13T12:14:54.207548Z","shell.execute_reply.started":"2023-10-13T12:14:54.203325Z","shell.execute_reply":"2023-10-13T12:14:54.206547Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# test_serises_df = pd.read_csv(main_dir+\"test_series_meta.csv\")\n# creating dicom file path column\n# def createListPaths(dcom_folders):\n#     paths = []\n#     for folder in tqdm(dcom_folders):\n#         paths+=sorted(tf.io.gfile.glob(folder+ \"*.dcm\"))\n#     return paths\n\n# train_serises_df[\"dicom_folder\"] = main_dir + \"train_images\"+ \"/\" + train_serises_df.patient_id.astype(str) + \"/\" + train_serises_df.series_id.astype(str)+ \"/\"\n# test_serises_df[\"dicom_folder\"] = main_dir + \"test_images\"+ \"/\" + test_serises_df.patient_id.astype(str) + \"/\" + test_serises_df.series_id.astype(str)+ \"/\"\n\n# train_dcm_image_paths = createListPaths(train_serises_df.dicom_folder.tolist())\n# test_dcm_image_paths = createListPaths(test_serises_df.dicom_folder.tolist())","metadata":{"execution":{"iopub.status.busy":"2023-10-13T12:14:54.208499Z","iopub.execute_input":"2023-10-13T12:14:54.209358Z","iopub.status.idle":"2023-10-13T12:14:54.22733Z","shell.execute_reply.started":"2023-10-13T12:14:54.20933Z","shell.execute_reply":"2023-10-13T12:14:54.226185Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# # create image path in a dataframe corresponding patients\n# PNG_IMAGE_DIR = \"tmp/dataset/\"\n# def createImagePathColumn(paths, subDir):\n#     image_ds = pd.DataFrame(paths, columns=[\"dicom_path\"])\n#     image_ds[\"patient_id\"] = image_ds.dicom_path.map(lambda x: x.split(\"/\")[-3]).astype(int)\n#     image_ds[\"series_id\"] = image_ds.dicom_path.map(lambda x: x.split(\"/\")[-2]).astype(int)\n#     image_ds[\"instance_number\"] = image_ds.dicom_path.map(lambda x: x.split(\"/\")[-1].replace(\".dcm\",\"\")).astype(int)\n\n#     image_ds[\"image_path\"] = f\"{PNG_IMAGE_DIR}/\"+ subDir + \"/\" + image_ds.patient_id.astype(str)\\\n#                         + \"/\" + image_ds.series_id.astype(str) + \"/\" + image_ds.instance_number.astype(str) +\".png\"\n#     return image_ds\n\n# # train_images_df = createImagePathColumn(train_dcm_image_paths, subDir=\"train_images\")\n# test_images_df = createImagePathColumn(test_dcm_image_paths, subDir=\"test_images\")","metadata":{"execution":{"iopub.status.busy":"2023-10-13T12:14:54.231768Z","iopub.execute_input":"2023-10-13T12:14:54.232193Z","iopub.status.idle":"2023-10-13T12:14:54.241107Z","shell.execute_reply.started":"2023-10-13T12:14:54.232155Z","shell.execute_reply":"2023-10-13T12:14:54.239971Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Converting dicom files to png and creating a dataset.\n## Already created the png dataset for train and test","metadata":{}},{"cell_type":"code","source":"# IMG_SIZE = (256, 256)\n# !rm -r {PNG_IMAGE_DIR}\n# os.makedirs(f\"{PNG_IMAGE_DIR}/train_images\", exist_ok=True)\n# os.makedirs(f\"{PNG_IMAGE_DIR}/test_images\", exist_ok=True)","metadata":{"execution":{"iopub.status.busy":"2023-10-13T12:14:54.242804Z","iopub.execute_input":"2023-10-13T12:14:54.243447Z","iopub.status.idle":"2023-10-13T12:14:54.25782Z","shell.execute_reply.started":"2023-10-13T12:14:54.243408Z","shell.execute_reply":"2023-10-13T12:14:54.256808Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# def standardize_pixel_array(dcm):\n#     # Correct DICOM pixel_array if PixelRepresentation == 1.\n#     pixel_array = dcm.pixel_array\n#     if dcm.PixelRepresentation == 1:\n#         bit_shift = dcm.BitsAllocated - dcm.BitsStored\n#         dtype = pixel_array.dtype \n#         new_array = (pixel_array << bit_shift).astype(dtype) >>  bit_shift\n#         pixel_array = pydicom.pixel_data_handlers.util.apply_modality_lut(new_array, dcm)\n#     return pixel_array\n\n# def resize_and_save(file_path):\n#     dicom = pydicom.dcmread(file_path)\n#     data = standardize_pixel_array(dicom)\n#     data = data - np.min(data)\n#     data = data / (np.max(data) + 1e-5)\n#     if dicom.PhotometricInterpretation == \"MONOCHROME1\":\n#         data = 1.0 - data\n#     img = (cv2.resize(data, IMG_SIZE, cv2.INTER_AREA) * 255).astype(np.float32)\n#     sub_path = file_path.split(\"/\",4)[-1].split(\".dcm\")[0] + \".png\"\n#     new_path = os.path.join(PNG_IMAGE_DIR, sub_path)\n#     os.makedirs(new_path.rsplit(\"/\",1)[0], exist_ok=True)\n#     cv2.imwrite(new_path, img)\n","metadata":{"execution":{"iopub.status.busy":"2023-10-13T12:14:54.259916Z","iopub.execute_input":"2023-10-13T12:14:54.261067Z","iopub.status.idle":"2023-10-13T12:14:54.27598Z","shell.execute_reply.started":"2023-10-13T12:14:54.261021Z","shell.execute_reply":"2023-10-13T12:14:54.274947Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# def convertDCMtoPNG(file_paths):\n#     # To better utilise the CPU performing parallel computing\n#     _ = Parallel(n_jobs=-1, backend=\"threading\")(\n#         delayed(resize_and_save)(file_path) for file_path in tqdm(file_paths, leave=True, position=0)\n#     )\n#     # delete the unused memory and automatically reclaim memory that is no longer accessible or in use by the application.\n#     # garbage collector runs automatically and periodically to clean up objects that are no longer referenced and thus are eligible for garbage collectionto occur immediately.\n#     del _\n#     gc.collect()\n    \n    \n# Split your large DataFrame into chunks (e.g., 70000 rows per chunk)\n# chunk_size = 70000\n# chunks = [train_images_df['dicom_path'][i:i + chunk_size] for i in range(0, len(train_images_df), chunk_size)]\n\n\n\n# def chunkArray(chunk, i):\n#     filename = f\"train_images_{i}.csv\"\n#     chunk['img_array'] = Parallel(n_jobs=-1, backend=\"threading\")(\n#         delayed(resize_and_save)(row) for _, row in tqdm(chunk.iterrows(), leave=True, position=0)\n#     )\n#     chunk['img_array'].astype(np.float32)\n#     chunk.to_csv(filename, index=False)\n    \n#     del chunk\n#     gc.collect()\n","metadata":{"execution":{"iopub.status.busy":"2023-10-13T12:14:54.277353Z","iopub.execute_input":"2023-10-13T12:14:54.277907Z","iopub.status.idle":"2023-10-13T12:14:54.289099Z","shell.execute_reply.started":"2023-10-13T12:14:54.277878Z","shell.execute_reply":"2023-10-13T12:14:54.288021Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# for index, chunk in enumerate(chunks):\n#     chunkArray(chunks[i], i)","metadata":{"execution":{"iopub.status.busy":"2023-10-13T12:14:54.291087Z","iopub.execute_input":"2023-10-13T12:14:54.291897Z","iopub.status.idle":"2023-10-13T12:14:54.307857Z","shell.execute_reply.started":"2023-10-13T12:14:54.29185Z","shell.execute_reply":"2023-10-13T12:14:54.306463Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# convertDCMtoPNG(test_images_df.dicom_path.tolist())","metadata":{"execution":{"iopub.status.busy":"2023-10-13T12:14:54.309598Z","iopub.execute_input":"2023-10-13T12:14:54.310076Z","iopub.status.idle":"2023-10-13T12:14:54.321878Z","shell.execute_reply.started":"2023-10-13T12:14:54.310029Z","shell.execute_reply":"2023-10-13T12:14:54.32105Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Training","metadata":{}},{"cell_type":"code","source":"train_image_dir = \"/kaggle/input/rsna-atd-256x256-png-trainimages/\"","metadata":{"execution":{"iopub.status.busy":"2023-10-13T12:14:54.323079Z","iopub.execute_input":"2023-10-13T12:14:54.32355Z","iopub.status.idle":"2023-10-13T12:14:54.338509Z","shell.execute_reply.started":"2023-10-13T12:14:54.323524Z","shell.execute_reply":"2023-10-13T12:14:54.337178Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# TRAINING ON TPU IF AVAILABLE\n# Detect hardware, return appropriate distribution strategy\ntry:\n    tpu = tf.distribute.cluster_resolver.TPUClusterResolver()  # TPU detection. No parameters necessary if TPU_NAME environment variable is set. On Kaggle this is always the case.\n    print('Running on TPU ', tpu.master())\nexcept ValueError:\n    tpu = None\n\nif tpu:\n    tf.config.experimental_connect_to_cluster(tpu)\n    tf.tpu.experimental.initialize_tpu_system(tpu)\n    strategy = tf.distribute.TPUStrategy(tpu)\nelse:\n    strategy = tf.distribute.get_strategy() # default distribution strategy in Tensorflow. Works on CPU and single GPU.\n\nprint(\"REPLICAS: \", strategy.num_replicas_in_sync)","metadata":{"execution":{"iopub.status.busy":"2023-10-13T12:14:54.340078Z","iopub.execute_input":"2023-10-13T12:14:54.340678Z","iopub.status.idle":"2023-10-13T12:14:54.355095Z","shell.execute_reply.started":"2023-10-13T12:14:54.340628Z","shell.execute_reply":"2023-10-13T12:14:54.353847Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"IMAGE_SIZE = (256, 256)\nBATCH_SIZE = 16 * strategy.num_replicas_in_sync if (strategy.num_replicas_in_sync >1) else 32 \nEPOCHS = 20\nTARGET_COLS  = [\n        \"bowel_injury\", \"extravasation_injury\",\n        \"kidney_healthy\", \"kidney_low\", \"kidney_high\",\n        \"liver_healthy\", \"liver_low\", \"liver_high\",\n        \"spleen_healthy\", \"spleen_low\", \"spleen_high\",\n    ]\nAUTOTUNE = tf.data.AUTOTUNE","metadata":{"execution":{"iopub.status.busy":"2023-10-13T12:14:54.357297Z","iopub.execute_input":"2023-10-13T12:14:54.357776Z","iopub.status.idle":"2023-10-13T12:14:54.374427Z","shell.execute_reply.started":"2023-10-13T12:14:54.357725Z","shell.execute_reply":"2023-10-13T12:14:54.373261Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"BASE_PATH = \"/kaggle/input/rsna-atd-512x512-png-v2-dataset\"\n# train\ndataframe = pd.read_csv(f\"{BASE_PATH}/train.csv\")\ndataframe[\"image_path\"] = f\"{BASE_PATH}/train_images\"\\\n                    + \"/\" + dataframe.patient_id.astype(str)\\\n                    + \"/\" + dataframe.series_id.astype(str)\\\n                    + \"/\" + dataframe.instance_number.astype(str) +\".png\"\ndataframe = dataframe.drop_duplicates()\n\ndataframe.head(2)","metadata":{"execution":{"iopub.status.busy":"2023-10-13T12:14:54.375979Z","iopub.execute_input":"2023-10-13T12:14:54.376478Z","iopub.status.idle":"2023-10-13T12:14:54.503763Z","shell.execute_reply.started":"2023-10-13T12:14:54.376441Z","shell.execute_reply":"2023-10-13T12:14:54.502417Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Function to handle the split for each group\ndef split_group(group, test_size=0.2):\n    if len(group) == 1:\n        return (group, pd.DataFrame()) if np.random.rand() < test_size else (pd.DataFrame(), group)\n    else:\n        return train_test_split(group, test_size=test_size, random_state=42)\n\n# Initialize the train and validation datasets\ntrain_data = pd.DataFrame()\nval_data = pd.DataFrame()\n\n# Iterate through the groups and split them, handling single-sample groups\nfor _, group in dataframe.groupby(TARGET_COLS):\n    train_group, val_group = split_group(group)\n    train_data = pd.concat([train_data, train_group], ignore_index=True)\n    val_data = pd.concat([val_data, val_group], ignore_index=True)","metadata":{"execution":{"iopub.status.busy":"2023-10-13T12:14:54.505406Z","iopub.execute_input":"2023-10-13T12:14:54.506057Z","iopub.status.idle":"2023-10-13T12:14:54.606382Z","shell.execute_reply.started":"2023-10-13T12:14:54.506022Z","shell.execute_reply":"2023-10-13T12:14:54.605202Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.shape, val_data.shape","metadata":{"execution":{"iopub.status.busy":"2023-10-13T12:14:54.608134Z","iopub.execute_input":"2023-10-13T12:14:54.608554Z","iopub.status.idle":"2023-10-13T12:14:54.616242Z","shell.execute_reply.started":"2023-10-13T12:14:54.608513Z","shell.execute_reply":"2023-10-13T12:14:54.614927Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def decode_image_and_label(image_path, label):\n    file_bytes = tf.io.read_file(image_path)\n    image = tf.io.decode_png(file_bytes, channels=3, dtype=tf.uint8)\n    image = tf.image.resize(image, IMAGE_SIZE, method=\"bilinear\")\n    image = tf.cast(image, tf.float32) / 255.0\n    \n    label = tf.cast(label, tf.float32)\n    #         bowel       extra       kidney      liver       spleen\n    labels = (label[0:1], label[1:2], label[2:5], label[5:8], label[8:11])\n    \n    return (image, labels)\n\n\ndef apply_augmentation(images, labels):\n    image_aug = tf.image.random_brightness(images, 0.2)\n    image_aug = tf.image.random_contrast(image_aug, 0.2, 0.5)\n    image_aug = tf.image.random_flip_left_right(image_aug)\n    image_aug = tf.image.random_flip_up_down(image_aug)\n    image_aug = tf.image.random_hue(image_aug, 0.2)\n    image_aug = tf.image.adjust_gamma(image_aug, 0.2)\n    image_aug = tf.image.random_saturation(image_aug, 0.2, 0.5)\n    return (images, labels)\n\n\ndef build_dataset(image_paths, labels):\n    ds = (\n        tf.data.Dataset.from_tensor_slices((image_paths, labels))\n        .map(decode_image_and_label, num_parallel_calls=AUTOTUNE)\n        .shuffle(BATCH_SIZE * 10)\n        .batch(BATCH_SIZE)\n        .map(apply_augmentation, num_parallel_calls=AUTOTUNE)\n        .prefetch(AUTOTUNE)\n    )\n    return ds","metadata":{"execution":{"iopub.status.busy":"2023-10-13T12:14:54.617894Z","iopub.execute_input":"2023-10-13T12:14:54.618237Z","iopub.status.idle":"2023-10-13T12:14:54.631607Z","shell.execute_reply.started":"2023-10-13T12:14:54.61821Z","shell.execute_reply":"2023-10-13T12:14:54.630311Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"paths  = train_data.image_path.tolist()\nlabels = train_data[TARGET_COLS].values\n\nds = build_dataset(image_paths=paths, labels=labels)\nimages, labels = next(iter(ds))\nimages.shape, [label.shape for label in labels]","metadata":{"execution":{"iopub.status.busy":"2023-10-13T12:14:54.633586Z","iopub.execute_input":"2023-10-13T12:14:54.634089Z","iopub.status.idle":"2023-10-13T12:14:55.899636Z","shell.execute_reply.started":"2023-10-13T12:14:54.634046Z","shell.execute_reply":"2023-10-13T12:14:55.898472Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import math\n# Define configuration parameters\nstart_lr = 0.00001\nmin_lr = 0.00001\nmax_lr = 0.00005 * strategy.num_replicas_in_sync\nrampup_epochs = 5\nsustain_epochs = 0\nexp_decay = 0.5\n\n# Define the scheduling function\ndef schedule(epoch):\n  def lr(epoch, start_lr, exp_decay):\n    return start_lr * math.exp(-exp_decay*epoch)\n  return lr(epoch, start_lr, exp_decay)\n\nlr_callback = tf.keras.callbacks.LearningRateScheduler(schedule, verbose = True)\n\nrng = [i for i in range(EPOCHS)]\ny = [schedule(x) for x in rng]\nplt.plot(rng, y)\nprint(\"Learning rate schedule: {:.3g} to {:.3g} to {:.3g}\".format(y[0], max(y), y[-1]))","metadata":{"execution":{"iopub.status.busy":"2023-10-13T12:14:55.90102Z","iopub.execute_input":"2023-10-13T12:14:55.901493Z","iopub.status.idle":"2023-10-13T12:14:56.167462Z","shell.execute_reply.started":"2023-10-13T12:14:55.901448Z","shell.execute_reply":"2023-10-13T12:14:56.166009Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\ndef build_model():\n    with strategy.scope():\n        pretrained_model = tf.keras.applications.densenet.DenseNet201(\n            input_shape=(IMAGE_SIZE[0], IMAGE_SIZE[1], 3),\n            weights='imagenet',\n            include_top=False,\n        )\n        pretrained_model.trainable = True\n        inputs = tf.keras.Input(shape=(IMAGE_SIZE[0], IMAGE_SIZE[1], 3), batch_size=BATCH_SIZE)\n        # GAP to get the activation maps\n        gap = tf.keras.layers.GlobalAveragePooling2D()\n        x = pretrained_model(inputs)\n        x = gap(x)\n\n        # Define 'necks' for each head\n        x_bowel = tf.keras.layers.Dense(128, activation='silu')(x)\n        x_extra = tf.keras.layers.Dense(128, activation='silu')(x)\n        x_liver = tf.keras.layers.Dense(128, activation='silu')(x)\n        x_kidney = tf.keras.layers.Dense(128, activation='silu')(x)\n        x_spleen = tf.keras.layers.Dense(128, activation='silu')(x)\n\n        # Define heads\n        out_bowel = tf.keras.layers.Dense(1, name='bowel', activation='sigmoid')(x_bowel) # use sigmoid to convert predictions to [0-1]\n        out_extra = tf.keras.layers.Dense(1, name='extra', activation='sigmoid')(x_extra) # use sigmoid to convert predictions to [0-1]\n        out_liver = tf.keras.layers.Dense(3, name='kidney', activation='softmax')(x_kidney) # use softmax for the liver head\n        out_kidney = tf.keras.layers.Dense(3, name='liver', activation='softmax')(x_liver) # use softmax for the kidney head\n        out_spleen = tf.keras.layers.Dense(3, name='spleen', activation='softmax')(x_spleen) # use softmax for the spleen head\n\n        # Concatenate the outputs\n        outputs = [out_bowel, out_extra, out_liver, out_kidney, out_spleen]\n\n        # Create model\n        print(\"[INFO] Building the model...\")\n        model = tf.keras.Model(inputs=inputs, outputs=outputs)\n\n        loss = {\n            \"bowel\":tf.keras.losses.BinaryCrossentropy(),\n            \"extra\":tf.keras.losses.BinaryCrossentropy(),\n            \"liver\":tf.keras.losses.CategoricalCrossentropy(),\n            \"kidney\":tf.keras.losses.CategoricalCrossentropy(),\n            \"spleen\":tf.keras.losses.CategoricalCrossentropy(),\n        }\n        metrics = {\n            \"bowel\":[\"accuracy\"],\n            \"extra\":[\"accuracy\"],\n            \"liver\":[\"accuracy\"],\n            \"kidney\":[\"accuracy\"],\n            \"spleen\":[\"accuracy\"],\n        }\n        print(\"[INFO] Compiling the model...\")\n        model.compile(\n            optimizer= tf.keras.optimizers.Adam(),\n          loss=loss,\n          metrics=metrics\n        )\n\n    return model","metadata":{"execution":{"iopub.status.busy":"2023-10-13T12:14:56.169112Z","iopub.execute_input":"2023-10-13T12:14:56.16946Z","iopub.status.idle":"2023-10-13T12:14:56.183114Z","shell.execute_reply.started":"2023-10-13T12:14:56.169412Z","shell.execute_reply":"2023-10-13T12:14:56.181763Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# get image_paths and labels\nprint(\"[INFO] Building the dataset...\")\ntrain_paths = train_data.image_path.values; train_labels = train_data[TARGET_COLS].values.astype(np.float32)\nvalid_paths = val_data.image_path.values; valid_labels = val_data[TARGET_COLS].values.astype(np.float32)\n\n# train and valid dataset\ntrain_ds = build_dataset(image_paths=train_paths, labels=train_labels)\nval_ds = build_dataset(image_paths=valid_paths, labels=valid_labels)\n\n# total_train_steps = train_ds.cardinality().numpy() * BATCH_SIZE * EPOCHS\n# warmup_steps = int(total_train_steps * 0.10)\n# decay_steps = total_train_steps - warmup_steps\n\n# print(f\"{total_train_steps=}\")\n# print(f\"{warmup_steps=}\")\n# print(f\"{decay_steps=}\")","metadata":{"execution":{"iopub.status.busy":"2023-10-13T12:14:56.184599Z","iopub.execute_input":"2023-10-13T12:14:56.185667Z","iopub.status.idle":"2023-10-13T12:14:56.442718Z","shell.execute_reply.started":"2023-10-13T12:14:56.185611Z","shell.execute_reply":"2023-10-13T12:14:56.441474Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# build the model\nprint(\"[INFO] Building the model...\")\n# model = build_model()\n\n# # train\n# print(\"[INFO] Training...\")\n# history = model.fit(\n#     train_ds,\n#     epochs=EPOCHS,\n#     validation_data=val_ds,\n#     callbacks = [lr_callback]\n# )","metadata":{"execution":{"iopub.status.busy":"2023-10-13T12:14:56.444668Z","iopub.execute_input":"2023-10-13T12:14:56.445034Z","iopub.status.idle":"2023-10-13T12:14:56.451165Z","shell.execute_reply.started":"2023-10-13T12:14:56.445007Z","shell.execute_reply":"2023-10-13T12:14:56.450141Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# # Create a 3x2 grid for the subplots\n# fig, axes = plt.subplots(5, 1, figsize=(5, 15))\n\n# # Flatten axes to iterate through them\n# axes = axes.flatten()\n\n# # Iterate through the metrics and plot them\n# for i, name in enumerate([\"bowel\", \"extra\", \"kidney\", \"liver\", \"spleen\"]):\n#     # Plot training accuracy\n#     axes[i].plot(history.history[name + '_accuracy'], label='Training ' + name)\n#     # Plot validation accuracy\n#     axes[i].plot(history.history['val_' + name + '_accuracy'], label='Validation ' + name)\n#     axes[i].set_title(name)\n#     axes[i].set_xlabel('Epoch')\n#     axes[i].set_ylabel('Accuracy')\n#     axes[i].legend()\n\n# plt.tight_layout()\n# plt.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-13T12:14:56.452856Z","iopub.execute_input":"2023-10-13T12:14:56.453242Z","iopub.status.idle":"2023-10-13T12:14:56.469214Z","shell.execute_reply.started":"2023-10-13T12:14:56.453207Z","shell.execute_reply":"2023-10-13T12:14:56.468018Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# plt.plot(history.history[\"loss\"], label=\"loss\")\n# plt.plot(history.history[\"val_loss\"], label=\"val loss\")\n# plt.legend()\n# plt.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-13T12:14:56.470873Z","iopub.execute_input":"2023-10-13T12:14:56.471207Z","iopub.status.idle":"2023-10-13T12:14:56.483198Z","shell.execute_reply.started":"2023-10-13T12:14:56.471181Z","shell.execute_reply":"2023-10-13T12:14:56.481878Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# # store best results\n# best_epoch = np.argmin(history.history['val_loss'])\n# best_loss = history.history['val_loss'][best_epoch]\n# best_acc_bowel = history.history['val_bowel_accuracy'][best_epoch]\n# best_acc_extra = history.history['val_extra_accuracy'][best_epoch]\n# best_acc_liver = history.history['val_liver_accuracy'][best_epoch]\n# best_acc_kidney = history.history['val_kidney_accuracy'][best_epoch]\n# best_acc_spleen = history.history['val_spleen_accuracy'][best_epoch]\n\n# # Find mean accuracy\n# best_acc = np.mean(\n#     [best_acc_bowel,\n#      best_acc_extra,\n#      best_acc_liver,\n#      best_acc_kidney,\n#      best_acc_spleen\n# ])\n\n# print(f'>>>> BEST Loss  : {best_loss:.3f}\\n>>>> BEST Acc   : {best_acc:.3f}\\n>>>> BEST Epoch : {best_epoch}\\n')\n# print('ORGAN Acc:')\n# print(f'  >>>> {\"Bowel\".ljust(15)} : {best_acc_bowel:.3f}')\n# print(f'  >>>> {\"Extravasation\".ljust(15)} : {best_acc_extra:.3f}')\n# print(f'  >>>> {\"Liver\".ljust(15)} : {best_acc_liver:.3f}')\n# print(f'  >>>> {\"Kidney\".ljust(15)} : {best_acc_kidney:.3f}')\n# print(f'  >>>> {\"Spleen\".ljust(15)} : {best_acc_spleen:.3f}')","metadata":{"execution":{"iopub.status.busy":"2023-10-13T12:14:56.484973Z","iopub.execute_input":"2023-10-13T12:14:56.485302Z","iopub.status.idle":"2023-10-13T12:14:56.500754Z","shell.execute_reply.started":"2023-10-13T12:14:56.485277Z","shell.execute_reply":"2023-10-13T12:14:56.499431Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Prediction","metadata":{}},{"cell_type":"code","source":"test_df = pd.read_csv(\"/kaggle/input/rsna-atd-512x512-png-v2-dataset/test.csv\")","metadata":{"execution":{"iopub.status.busy":"2023-10-13T12:14:56.502103Z","iopub.execute_input":"2023-10-13T12:14:56.502407Z","iopub.status.idle":"2023-10-13T12:14:56.523849Z","shell.execute_reply.started":"2023-10-13T12:14:56.502382Z","shell.execute_reply":"2023-10-13T12:14:56.522651Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df[\"image_path_png\"] = f\"{BASE_PATH}/test_images\"\\\n                    + \"/\" + test_df.patient_id.astype(str)\\\n                    + \"/\" + test_df.series_id.astype(str)\\\n                    + \"/\" + test_df.instance_number.astype(str) +\".png\"\ntest_df","metadata":{"execution":{"iopub.status.busy":"2023-10-13T12:14:56.525214Z","iopub.execute_input":"2023-10-13T12:14:56.525702Z","iopub.status.idle":"2023-10-13T12:14:56.545121Z","shell.execute_reply.started":"2023-10-13T12:14:56.525628Z","shell.execute_reply":"2023-10-13T12:14:56.543843Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def post_proc(pred):\n    proc_pred = np.empty((pred.shape[0], 2*2 + 3*3), dtype='float32')\n\n    # bowel, extravasation\n    proc_pred[:, 0] = 1 - pred[:, 0] # bowel-healthy\n    proc_pred[:, 1] = pred[:, 0] # bowel-injured\n    proc_pred[:, 2] = 1 - pred[:, 1] # extra-healthy\n    proc_pred[:, 3] = pred[:, 1] # extra-injured\n    \n    # liver, kidney, sneel\n    proc_pred[:, 4:7] = pred[:, 2:5]\n    proc_pred[:, 7:10] = pred[:, 5:8]\n    proc_pred[:, 10:13] = pred[:, 8:11]\n\n    return proc_pred\n\n\n\ndef decode_image(image_path):\n    file_bytes = tf.io.read_file(image_path)\n    image = tf.io.decode_png(file_bytes, channels=3, dtype=tf.uint8)\n    image = tf.image.resize(image, IMAGE_SIZE, method=\"bilinear\")\n    image = tf.cast(image, tf.float32) / 255.0\n    return image\n\ndef build_dataset(image_paths):\n    ds = (\n        tf.data.Dataset.from_tensor_slices(image_paths)\n        .map(decode_image, num_parallel_calls=AUTOTUNE)\n        .shuffle(BATCH_SIZE * 10)\n        .batch(BATCH_SIZE)\n        .prefetch(AUTOTUNE)\n    )\n    return ds","metadata":{"execution":{"iopub.status.busy":"2023-10-13T12:14:56.546377Z","iopub.execute_input":"2023-10-13T12:14:56.547107Z","iopub.status.idle":"2023-10-13T12:14:56.556512Z","shell.execute_reply.started":"2023-10-13T12:14:56.547077Z","shell.execute_reply":"2023-10-13T12:14:56.555054Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Getting unique patient IDs from test dataset\npatient_ids = test_df['patient_id'].unique()\n\n# Initializing array to store predictions\npatient_preds = np.zeros(shape=(len(patient_ids), 2*2 + 3*3), dtype='float32')\n\n# Iterating over each patient\nfor pidx, patient_id in tqdm(enumerate(patient_ids), total=len(patient_ids), desc=\"Patients \"):\n    # Query the dataframe for a particular patient\n    patient_df = test_df[test_df['patient_id']==patient_id]\n    \n    # Initializing model predictions array\n    model_preds = np.zeros(shape=(1, 11), dtype=np.float32)\n     \n    # Getting image paths for a patient\n    patient_paths = patient_df.image_path_png.tolist()\n\n\n    # Building dataset for prediction\n    dtest = build_dataset(patient_paths)\n\n    with strategy.scope():\n        # Loading a model from a fold path\n        load_model = tf.keras.models.load_model(\"/kaggle/input/rsna-saved-model-densenet/model_rsna_02.h5\", compile=False)\n\n    # Predicting with the model\n    pred = load_model.predict(dtest, verbose=1)\n    pred = np.concatenate(pred, axis=-1).astype('float32') # reducing memory footprint\n    pred = pred[:len(patient_paths), :]\n    pred = np.mean(pred.reshape(1, len(patient_paths), 11), axis=0)\n    pred = np.max(pred, axis=0) # taking max prediction of all ct scans for a patient\n\n    # Store model's prediction\n    model_preds += pred      \n            \n    # Adding processed predictions to patient_preds\n    patient_preds[pidx, :] += post_proc(model_preds)[0]\n    \n    del model_preds, dtest, patient_paths, load_model, pred; gc.collect()","metadata":{"execution":{"iopub.status.busy":"2023-10-13T12:14:56.557682Z","iopub.execute_input":"2023-10-13T12:14:56.558001Z","iopub.status.idle":"2023-10-13T12:15:36.575171Z","shell.execute_reply.started":"2023-10-13T12:14:56.557977Z","shell.execute_reply":"2023-10-13T12:15:36.574163Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create Submission\npred_df = pd.DataFrame({\"patient_id\":patient_ids,})\npred_df[['bowel_healthy','bowel_injury' , 'extravasation_healthy',\n       'extravasation_injury', 'kidney_healthy', 'kidney_low', 'kidney_high',\n       'liver_healthy', 'liver_low', 'liver_high', 'spleen_healthy',\n       'spleen_low', 'spleen_high']] = patient_preds.astype(\"float64\")\n# Align with sample submission\nsub_df = pd.read_csv(f'{BASE_PATH}/sample_submission.csv')\n\nsub_df.set_index('patient_id', inplace = True)\npred_df.set_index('patient_id', inplace = True)\n\nsub_df.update(pred_df)\nsub_df.reset_index(inplace = True)\n\nsub_df.to_csv('submission.csv', index = False)\n\n","metadata":{"execution":{"iopub.status.busy":"2023-10-13T12:56:44.73603Z","iopub.execute_input":"2023-10-13T12:56:44.736605Z","iopub.status.idle":"2023-10-13T12:56:44.769169Z","shell.execute_reply.started":"2023-10-13T12:56:44.736569Z","shell.execute_reply":"2023-10-13T12:56:44.768088Z"},"trusted":true},"execution_count":null,"outputs":[]}]}