{"metadata":{"kernelspec":{"name":"python3","display_name":"Python 3","language":"python"},"language_info":{"name":"python","version":"3.11.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":99552,"databundleVersionId":13190393,"isSourceIdPinned":false,"sourceType":"competition"},{"sourceId":1421668,"sourceType":"datasetVersion","datasetId":832340}],"dockerImageVersionId":31089,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# RSNA 2025 Intracranial Aneurysm Detection Create Dataset\n","metadata":{"execution":{"iopub.status.busy":"2024-06-12T23:49:53.776252Z","iopub.execute_input":"2024-06-12T23:49:53.777034Z","iopub.status.idle":"2024-06-12T23:49:55.048899Z","shell.execute_reply.started":"2024-06-12T23:49:53.777Z","shell.execute_reply":"2024-06-12T23:49:55.047589Z"}}},{"cell_type":"markdown","source":"\nThis notebook is to create the dataset from DICOM files to PNG files (224x224 px as default). You can adjust the size to suit your needs.\n\n**The dataset generated by this notebook can be found at : https://www.kaggle.com/datasets/dennisfong/rsna-2025-intracranial-aneurysm-png-224x224/**\n\nFiles will be created to cvt_png folder with the following sturure:\n```plaintext\ncvt_png/\n├── <location_name>/\n│   ├── <SeriesInstanceUID_1>/\n│   │   ├── 0000.png\n│   │   ├── 0001.png\n│   │   ├── 0002.png\n│   │   └── ...\n│   ├── <SeriesInstanceUID_2>/\n│   │   ├── 0000.png\n│   │   ├── 0001.png\n│   │   └── ...\n│   └── ...\n├── <location_name_2>/\n│   ├── <SeriesInstanceUID_3>/\n│   │   ├── 0000.png\n│   │   ├── 0001.png\n│   │   └── ...\n│   └── ...\n└── ...\n```\n\n* train_localizers_with_relative.csv will be created to store the original information and relative information created for the converted PNG\n* series_index_mapping.csv maps each DICOM file in the dataset to its corresponding SeriesInstanceUID, SOPInstanceUID (derived from filename), Modality, and its relative position (index) within the sorted series.","metadata":{}},{"cell_type":"markdown","source":"# Challenges in Processing DICOM Images for PNG Conversion\n**1. Handling Compressed DICOM Files**\nMany DICOM files use compression schemes that require specialized decoding libraries. Without proper decompression support, attempts to read pixel data will fail or produce invalid images. For instance, some transfer syntaxes require the use of the GDCM (Grassroots DICOM) library to decode compressed pixel data such as JPEG or JPEG2000. If GDCM is not available or not properly configured, the pipeline cannot process these compressed DICOM files.\n\n**2. Unsupported Image Channel Formats for OpenCV**\nOpenCV’s image writing function (cv2.imwrite) only supports images with 1 (grayscale), 3 (RGB), or 4 (RGBA) channels. However, DICOM pixel data may sometimes have unusual channel configurations, such as multi-frame images or color images with 2 or more than 4 channels (e.g., shape (height, width, 2) or beyond). This mismatch leads to errors or corrupt output since OpenCV cannot handle these unsupported channel formats directly.\n\n**3. Empty or Invalid DICOM Pixel Data Leading to Resize Failures**\nSome DICOM files may contain missing, corrupted, or empty pixel data. Attempting to resize such empty images results in runtime errors in OpenCV, specifically an assertion failure indicating that the target resize dimensions are empty or invalid. This commonly occurs when the pixel array is None or has zero size, causing the image processing pipeline to crash or terminate unexpectedly.\n\n**4. Incorrect Slice Ordering Due to Non-sequential Filenames**\nDICOM files are often named using unique identifiers such as SOPInstanceUID, which are not inherently ordered. Sorting files by filename—even using natural sorting—does not guarantee the correct anatomical or acquisition order of slices. This is critical in medical imaging, where proper spatial ordering is required for accurate visualization, 3D reconstruction, or volume analysis. Relying on alphabetical or natural file sorting can result in incorrect image sequences.","metadata":{}},{"cell_type":"markdown","source":"# Solutions to Address These Challenges\n**1. Integrate GDCM for Compressed DICOM Decoding**\nEnsure that the GDCM library is installed and integrated with the pydicom pixel data handlers. (Special thanks to @ronaldokun on providing GDCM offline version)\n\n\n**2. Normalize or Convert Pixel Data to Supported Formats**\nBefore passing image data to OpenCV functions, inspect the shape of the pixel array. For unsupported channel numbers:\n\n* Convert color spaces as needed (e.g., YBR_FULL to RGB).\n\n* For multi-channel or multi-frame images, extract or merge frames to produce a single 1-, 3-, or 4-channel image.\n\n* Convert to grayscale if color conversion is ambiguous or unsupported.\n\nThis preprocessing ensures the final image conforms to OpenCV requirements, avoiding runtime errors during saving or display.\n\n**3. Validate and Handle Empty or Corrupt Pixel Data Gracefully**\nAdd explicit checks after loading pixel data to verify it is non-empty and valid before any resizing or processing. If pixel data is missing or empty, skip processing the image with appropriate logging or error handling. This prevents crashes and allows the pipeline to continue processing other valid files.\n\n**4. Robust Multi-Criterion DICOM Sorting**\n\nTo ensure reliable ordering, we employ a two-level sorting strategy using metadata from each DICOM file:\n\na. **Primary Criterion: InstanceNumber**\n\nThis DICOM tag (0020,0013) is the most commonly used and directly indicates the order of acquisition or frame.\n\nIf InstanceNumber is present, it is used as the primary key for sorting.\n\nb. **Secondary Criterion: ImagePositionPatient**\n\nIf InstanceNumber is missing or unreliable, we fall back to the Z-coordinate ([2]) of ImagePositionPatient (0020,0032), which represents the spatial position of the slice along the patient’s body axis.\n\nThis is especially useful for CT/MR scans when spatial consistency is required.\n\nc. **Fallback Behavior**\n\nIf neither tag is available, the file is assigned a very large sorting key (float('inf')), effectively pushing it to the end of the list.\n\nThis ensures that invalid or unreadable files do not interfere with the valid sorting logic.","metadata":{}},{"cell_type":"markdown","source":"# Config","metadata":{}},{"cell_type":"code","source":"\nDEBUG = False  # Set to False to run on full dataset\nGLOBAL_WIDTH = 224 # the width is of the same size as height","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-01T02:11:28.425688Z","iopub.execute_input":"2025-08-01T02:11:28.425972Z","iopub.status.idle":"2025-08-01T02:11:28.434245Z","shell.execute_reply.started":"2025-08-01T02:11:28.425948Z","shell.execute_reply":"2025-08-01T02:11:28.433354Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Install and import Libraries","metadata":{}},{"cell_type":"code","source":"!pip install dicomsdl","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-01T02:11:28.435331Z","iopub.execute_input":"2025-08-01T02:11:28.435655Z","iopub.status.idle":"2025-08-01T02:11:35.135967Z","shell.execute_reply.started":"2025-08-01T02:11:28.435626Z","shell.execute_reply":"2025-08-01T02:11:35.135004Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"!pip install python-gdcm","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-01T02:11:35.138594Z","iopub.execute_input":"2025-08-01T02:11:35.138976Z","iopub.status.idle":"2025-08-01T02:11:40.0274Z","shell.execute_reply.started":"2025-08-01T02:11:35.138944Z","shell.execute_reply":"2025-08-01T02:11:40.026479Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import os\nimport re\nimport glob\nimport time\n\nimport numpy as np\nimport pandas as pd\nimport cv2\nimport pydicom\nfrom pydicom.pixel_data_handlers.util import convert_color_space\nimport dicomsdl as dicoml\nfrom matplotlib import pyplot as plt\nfrom tqdm import tqdm\nfrom joblib import Parallel, delayed\nfrom IPython.display import HTML\nfrom multiprocessing import Pool, cpu_count\nimport imageio","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-01T02:11:40.028591Z","iopub.execute_input":"2025-08-01T02:11:40.028901Z","iopub.status.idle":"2025-08-01T02:11:41.806476Z","shell.execute_reply.started":"2025-08-01T02:11:40.028868Z","shell.execute_reply":"2025-08-01T02:11:41.805649Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Some dicom files are compressed, we need to import the gdcm library to convert the files","metadata":{}},{"cell_type":"code","source":"!cp ../input/gdcm-conda-install/gdcm.tar .\n!tar -xvzf gdcm.tar\n!conda install --offline ./gdcm/gdcm-2.8.9-py37h71b2a6d_0.tar.bz2\nprint(\"done installing gdcm\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-01T02:11:41.807404Z","iopub.execute_input":"2025-08-01T02:11:41.807901Z","iopub.status.idle":"2025-08-01T02:11:42.390789Z","shell.execute_reply.started":"2025-08-01T02:11:41.80787Z","shell.execute_reply":"2025-08-01T02:11:42.389702Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Init some variables and helper functions","metadata":{}},{"cell_type":"code","source":"rd = '/kaggle/input/rsna-intracranial-aneurysm-detection'","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-01T02:11:42.39219Z","iopub.execute_input":"2025-08-01T02:11:42.392542Z","iopub.status.idle":"2025-08-01T02:11:42.39717Z","shell.execute_reply.started":"2025-08-01T02:11:42.392501Z","shell.execute_reply":"2025-08-01T02:11:42.39627Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\ndef atoi(text):\n    return int(text) if text.isdigit() else text\n\ndef natural_keys(text):\n    return [ atoi(c) for c in re.split(r'(\\d+)', text) ]\n\n\ndef get_sort_key(path):\n    \"\"\"\n    Return a sorting key for a DICOM file based on:\n    1. InstanceNumber (preferred)\n    2. ImagePositionPatient's Z-axis (fallback)\n    3. Defaults to (inf, inf) if neither is available\n    \"\"\"\n    try:\n        ds = pydicom.dcmread(path, stop_before_pixels=True)\n\n        # Try to use InstanceNumber first\n        instance_number = getattr(ds, 'InstanceNumber', None)\n\n        # Fallback: use Z-coordinate of ImagePositionPatient\n        image_position = getattr(ds, 'ImagePositionPatient', None)\n        z_position = image_position[2] if image_position and len(image_position) == 3 else None\n\n        if instance_number is not None:\n            return (int(instance_number), 0)\n        elif z_position is not None:\n            return (float('inf'), float(z_position))\n        else:\n            return (float('inf'), float('inf'))\n\n    except Exception:\n        return (float('inf'), float('inf'))\n\n\n\ndef extract_sort_key(path):\n    try:\n        ds = pydicom.dcmread(path, stop_before_pixels=True, force=True)\n        instance_number = getattr(ds, 'InstanceNumber', None)\n        position = getattr(ds, 'ImagePositionPatient', [None, None, None])\n        z = position[2] if position and len(position) == 3 else None\n\n        if instance_number is not None:\n            return (int(instance_number), 0, path)\n        elif z is not None:\n            return (float('inf'), float(z), path)\n        else:\n            return (float('inf'), float('inf'), path)\n    except:\n        return (float('inf'), float('inf'), path)\n\ndef fast_sort_dicom_paths_parallel(dcm_paths, num_workers=None):\n    if num_workers is None:\n        num_workers = max(1, cpu_count() - 1)\n\n    with Pool(processes=num_workers) as pool:\n        sort_info = pool.map(extract_sort_key, dcm_paths)\n\n    sort_info.sort()\n    return [x[2] for x in sort_info]\n\n# === Function to sort DICOM paths using metadata ===\ndef fast_sort_dicom_paths(dcm_paths):\n    sort_info = []\n    for path in dcm_paths:\n        try:\n            ds = pydicom.dcmread(path, stop_before_pixels=True, force=True)\n            instance_number = getattr(ds, 'InstanceNumber', None)\n            position = getattr(ds, 'ImagePositionPatient', [None, None, None])\n            z = position[2] if position and len(position) == 3 else None\n\n            if instance_number is not None:\n                sort_info.append((int(instance_number), 0, path))\n            elif z is not None:\n                sort_info.append((float('inf'), float(z), path))\n            else:\n                sort_info.append((float('inf'), float('inf'), path))\n        except:\n            sort_info.append((float('inf'), float('inf'), path))\n    sort_info.sort()\n    return [x[2] for x in sort_info]\n\n# === Helper function for parallel sorting ===\ndef sort_series(args):\n    series_uid, paths = args\n    sorted_paths = fast_sort_dicom_paths(paths)\n    return series_uid, sorted_paths\n\ndef apply_dicom_windowing(img, window_center, window_width):\n    \"\"\"\n    Apply DICOM windowing to the input image. Windowing is a technique used in medical imaging \n    to enhance the visibility of specific structures or tissue types by adjusting the pixel intensity range.\n    \n    The windowing process involves clipping pixel values outside a specific range (defined by \n    window center and width) and then normalizing the image to a common intensity scale (0-255) for visualization.\n    \n    Parameters:\n    ----------\n    img : np.ndarray\n        The input image to which windowing is applied. It should be a NumPy array (2D for grayscale images).\n        \n    window_center : float\n        The center value of the windowing. Typically, this represents the central intensity value for a certain tissue type.\n        \n    window_width : float\n        The width of the window. It defines the range of pixel values that will be included in the windowing process. \n        A wider window includes more intensities, while a narrower window emphasizes a smaller range.\n\n    Returns:\n    -------\n    np.ndarray\n        The windowed image, scaled to 8-bit values (0-255) and returned as a NumPy array.\n    \"\"\"\n    # Calculate the minimum and maximum pixel values for the window\n    img_min = window_center - window_width // 2\n    img_max = window_center + window_width // 2\n    \n    # Clip the pixel values to the windowed range\n    img = np.clip(img, img_min, img_max)\n    \n    # Normalize the image to 0-1 range to prepare for scaling\n    img = (img - img_min) / (img_max - img_min + 1e-7)  # Adding a small epsilon to avoid division by zero\n    \n    # Scale the image to 8-bit (0-255) and return\n    return (img * 255).astype(np.uint8)\n\ndef get_windowing_params(modality):\n    \"\"\"\n    Get the appropriate windowing parameters (center and width) based on the DICOM modality.\n    \n    Different modalities like CT, MRI, and MRA have different standard windowing values \n    to enhance specific tissue types or structures. For example, CT might have a window for bone, \n    while MRI might have a window for soft tissues.\n\n    Parameters:\n    ----------\n    modality : str\n        The imaging modality (e.g., 'CT', 'MRI', 'CTA'). This determines which set of windowing parameters to use.\n\n    Returns:\n    -------\n    Tuple[float, float]\n        A tuple containing the window center and window width for the given modality.\n    \"\"\"\n    # Dictionary mapping modalities to windowing parameters (center, width)\n    windows = {\n        'CT': (40, 80),     # Windowing parameters for CT images (e.g., bone/window for soft tissue)\n        'CTA': (50, 350),   # CTA (CT Angiography) has larger window width to visualize blood vessels\n        'MRA': (600, 1200), # MRA (Magnetic Resonance Angiography) uses a different range for blood vessels\n        'MRI': (40, 80),    # MRI typically uses windowing similar to CT for soft tissue\n        'MRI T2': (40, 80),\n        'MRI T1post': (40, 80),\n    }\n    \n    # Return the windowing parameters for the modality, defaulting to CT parameters if modality is not found\n    return windows.get(modality, (40, 80))  # Default to CT windowing if modality is unknown\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-01T02:11:42.398108Z","iopub.execute_input":"2025-08-01T02:11:42.398392Z","iopub.status.idle":"2025-08-01T02:11:42.421045Z","shell.execute_reply.started":"2025-08-01T02:11:42.398358Z","shell.execute_reply":"2025-08-01T02:11:42.419896Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Cache files sorting in each seriesfor fast retrieval ","metadata":{}},{"cell_type":"code","source":"\n# === Load train.csv ===\ndf_train = pd.read_csv('/kaggle/input/rsna-intracranial-aneurysm-detection/train.csv')\nseries_uids_full = df_train['SeriesInstanceUID'].unique()\n\nif DEBUG:\n    print(\"DEBUG MODE: Using only first 10 series_uids\")\n    series_uids = series_uids_full[:10]\nelse:\n    print(\"PRODUCTION MODE: Using all series_uids\")\n    series_uids = series_uids_full\n\n# === Build map: SeriesInstanceUID → list of DICOM paths ===\nseries_dicom_map = {\n    si: glob.glob(os.path.join(rd, 'series', si, '*.dcm'))\n    for si in series_uids\n}\n\n# === Parallel sort each series ===\nwith Pool(cpu_count()) as pool:\n    sorted_results = list(tqdm(pool.imap(sort_series, series_dicom_map.items()),\n                               total=len(series_dicom_map),\n                               desc=\"Sorting DICOM series\"))\n\n# === Generate output rows ===\nrows = []\nfor series_uid, sorted_paths in tqdm(sorted_results, desc=\"Generating CSV rows\"):\n    modality = df_train[df_train['SeriesInstanceUID'] == series_uid]['Modality'].iloc[0]  # Get modality for the current series\n    for idx, path in enumerate(sorted_paths):\n        sop_uid = os.path.splitext(os.path.basename(path))[0]  # use filename as SOPInstanceUID\n        rows.append({\n            'SeriesInstanceUID': series_uid,\n            'SOPInstanceUID': sop_uid,\n            'dicom_filename': path,\n            'relative_index': idx,\n            'Modality': modality\n        })\n\n# === Save as DataFrame ===\ndf_series_index_mapping = pd.DataFrame(rows)\ndf_series_index_mapping = df_series_index_mapping.sort_values(\n    by=['SeriesInstanceUID', 'relative_index']\n)\ndf_series_index_mapping.to_csv('series_index_mapping.csv', index=False)\nprint(\"Saved series_index_mapping.csv\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-01T02:11:42.422359Z","iopub.execute_input":"2025-08-01T02:11:42.422763Z","iopub.status.idle":"2025-08-01T02:11:52.691969Z","shell.execute_reply.started":"2025-08-01T02:11:42.422729Z","shell.execute_reply":"2025-08-01T02:11:52.690927Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Explore csv","metadata":{}},{"cell_type":"code","source":"df_localizers = pd.read_csv(f'{rd}/train_localizers.csv')\ndf_localizers.columns","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-01T02:11:52.69552Z","iopub.execute_input":"2025-08-01T02:11:52.695848Z","iopub.status.idle":"2025-08-01T02:11:52.726075Z","shell.execute_reply.started":"2025-08-01T02:11:52.695801Z","shell.execute_reply":"2025-08-01T02:11:52.725192Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df = pd.read_csv(f'{rd}/train.csv')\ndf.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-01T02:11:52.727119Z","iopub.execute_input":"2025-08-01T02:11:52.727446Z","iopub.status.idle":"2025-08-01T02:11:52.797809Z","shell.execute_reply.started":"2025-08-01T02:11:52.727418Z","shell.execute_reply":"2025-08-01T02:11:52.796909Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df.columns","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-01T02:11:52.798803Z","iopub.execute_input":"2025-08-01T02:11:52.799166Z","iopub.status.idle":"2025-08-01T02:11:52.80619Z","shell.execute_reply.started":"2025-08-01T02:11:52.799137Z","shell.execute_reply":"2025-08-01T02:11:52.805466Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df['SeriesInstanceUID'].value_counts()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-01T02:11:52.807211Z","iopub.execute_input":"2025-08-01T02:11:52.8075Z","iopub.status.idle":"2025-08-01T02:11:52.833777Z","shell.execute_reply.started":"2025-08-01T02:11:52.807471Z","shell.execute_reply":"2025-08-01T02:11:52.832804Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Export png from dcm","metadata":{}},{"cell_type":"code","source":"\n\nUNCOMPRESSED_SYNTAXES = [\n    '1.2.840.10008.1.2',       # Implicit VR Little Endian\n    '1.2.840.10008.1.2.1',     # Explicit VR Little Endian\n    '1.2.840.10008.1.2.2',     # Explicit VR Big Endian\n]\n\ndef dicom_to_png(src_path, dst_path, width=224, to_rgb=False, apply_windowing=False, modality='CT'):\n    try:\n        dicom = pydicom.dcmread(src_path, force=True)\n\n        if 'PixelData' not in dicom:\n            print(f\"[SKIP] No pixel data in: {src_path}\")\n            return\n\n        tsuid = dicom.file_meta.TransferSyntaxUID  # compression could be indicated here\n        # PixelSequence::copyDecodedFrameData - error in decoding frame data 'Unsupported marker type 0x64'\n\n        img = dicom.pixel_array\n        interp = dicom.PhotometricInterpretation\n\n        # Handle YBR color space\n        if interp == \"YBR_FULL\":\n            img = convert_color_space(img, 'YBR_FULL', 'RGB')\n\n        # Convert color to grayscale\n        if img.ndim == 3:\n            if interp in [\"RGB\", \"YBR_FULL\"]:\n                img = cv2.cvtColor(img, cv2.COLOR_RGB2GRAY)\n            elif img.shape[2] == 1:\n                img = img[:, :, 0]\n            elif img.shape[2] > 3:\n                img = img[:, :, 0]  # Fallback: take the first channel\n\n        # If windowing is requested, apply it\n        if apply_windowing:\n            window_center, window_width = get_windowing_params(modality)\n            img = apply_dicom_windowing(img, window_center, window_width)\n\n        # Normalize the image\n        img = img.astype(np.float32)\n        img_min, img_max = img.min(), img.max()\n        if img_max > img_min:\n            img = (img - img_min) / (img_max - img_min)\n        else:\n            img[:] = 0  # Flat image\n\n        # MONOCHROME1 means low values = white, invert it\n        if interp == \"MONOCHROME1\":\n            img = 1 - img\n\n        # Scale to 8-bit and resize\n        img = (img * 255).astype(np.uint8)\n\n        if img is None or img.size == 0:\n            print(f\"[SKIP] Invalid image data in: {src_path}\")\n            return\n\n        # Resize safely\n        try:\n            img = cv2.resize(img, (width, width))\n        except Exception as e:\n            print(f\"[ERROR] Resize failed: {src_path} | {e}\")\n            return\n\n        # Convert grayscale to RGB if requested\n        if to_rgb:\n            img = cv2.cvtColor(img, cv2.COLOR_GRAY2RGB)\n\n        # Ensure destination directory exists\n        os.makedirs(os.path.dirname(dst_path), exist_ok=True)\n\n        # Save the image\n        cv2.imwrite(dst_path, img)\n        # print(f\"[OK] Saved: {dst_path}\")\n\n    except Exception as e:\n        print(f\"[ERROR] Failed to process {src_path}: {e}\")\n\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-01T02:11:52.834926Z","iopub.execute_input":"2025-08-01T02:11:52.835248Z","iopub.status.idle":"2025-08-01T02:11:52.854511Z","shell.execute_reply.started":"2025-08-01T02:11:52.835214Z","shell.execute_reply":"2025-08-01T02:11:52.853484Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"seriesInst_ids = df['SeriesInstanceUID'].unique()\nseriesInst_ids[:3], len(seriesInst_ids)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-01T02:11:52.855642Z","iopub.execute_input":"2025-08-01T02:11:52.856012Z","iopub.status.idle":"2025-08-01T02:11:52.88254Z","shell.execute_reply.started":"2025-08-01T02:11:52.855979Z","shell.execute_reply":"2025-08-01T02:11:52.881373Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"modality = list(df['Modality'].unique())\nmodality","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-01T02:11:52.883556Z","iopub.execute_input":"2025-08-01T02:11:52.883947Z","iopub.status.idle":"2025-08-01T02:11:52.900136Z","shell.execute_reply.started":"2025-08-01T02:11:52.883919Z","shell.execute_reply":"2025-08-01T02:11:52.899096Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\n\n# Columns to exclude\nexclude_cols = ['SeriesInstanceUID', 'PatientAge', 'PatientSex', 'Modality', 'Aneurysm Present']\n\n# Get list of columns excluding the specified ones\nlocation = [col for col in df.columns if col not in exclude_cols]\n\nprint(location)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-01T02:11:52.901281Z","iopub.execute_input":"2025-08-01T02:11:52.901664Z","iopub.status.idle":"2025-08-01T02:11:52.92068Z","shell.execute_reply.started":"2025-08-01T02:11:52.901633Z","shell.execute_reply":"2025-08-01T02:11:52.919651Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Check the maximum number of input DICOM files per series directory to decide whether {j:03d}.png is sufficient or if {j:04d}.png is needed.","metadata":{}},{"cell_type":"code","source":"\n\"\"\"\n# Path to your input series directory\nseries_base_dir = os.path.join(rd, 'series')\n\nmax_dcm_count = 0\nseries_with_max = None\n\n# Get all series folders\nseries_dirs = glob.glob(os.path.join(series_base_dir, '*'))\n\n# Loop through each series folder and count .dcm files\nfor series_dir in series_dirs:\n    dcm_files = glob.glob(os.path.join(series_dir, '*.dcm'))\n    count = len(dcm_files)\n    if count > max_dcm_count:\n        max_dcm_count = count\n        series_with_max = series_dir\n\nprint(f\"Max number of DICOM files in a series: {max_dcm_count}\")\nprint(f\"Series folder with max DICOMs: {series_with_max}\")\n\"\"\"","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-01T02:11:52.921675Z","iopub.execute_input":"2025-08-01T02:11:52.922032Z","iopub.status.idle":"2025-08-01T02:11:52.942435Z","shell.execute_reply.started":"2025-08-01T02:11:52.922Z","shell.execute_reply":"2025-08-01T02:11:52.941627Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\n# Load precomputed mapping CSV\ndf_mapping = pd.read_csv('series_index_mapping.csv')\n\noutputList = []\n\nif DEBUG:\n    series_subset = seriesInst_ids[:10]\n    print(\"DEBUG MODE: Processing first 10 SeriesInstanceUIDs\")\nelse:\n    series_subset = seriesInst_ids\n    print(\"PRODUCTION MODE: Processing all SeriesInstanceUIDs\")\n\nfor idx, si in enumerate(tqdm(series_subset, total=len(series_subset))):\n    pdf = df[df['SeriesInstanceUID'] == si]\n    \n    for _, row in pdf.iterrows():\n        location_exist = [col for col in location if row[col] == 1]\n        \n        if not location_exist:\n            continue  # Skip if no diseases found\n\n        for loca in location_exist:\n            loca_ = loca.replace('/', '_')  # Clean the name for file system\n            \n            # === Read DICOM list from CSV ===\n            df_series = df_mapping[df_mapping['SeriesInstanceUID'] == si]\n            df_series = df_series.sort_values(by='relative_index')\n\n            if df_series.empty:\n                continue\n\n            # Create output directory\n            out_dir = f'cvt_png/{loca_}/{si}'\n            os.makedirs(out_dir, exist_ok=True)\n\n            for j, row_mapping in enumerate(df_series.itertuples(index=False)):\n                impath = row_mapping.dicom_filename\n                dst = f'{out_dir}/{j:04d}.png'\n                outputList.append({'impath': impath, 'dst': dst, 'modality': row_mapping.Modality})","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-01T02:11:52.943475Z","iopub.execute_input":"2025-08-01T02:11:52.943879Z","iopub.status.idle":"2025-08-01T02:11:53.007178Z","shell.execute_reply.started":"2025-08-01T02:11:52.943813Z","shell.execute_reply":"2025-08-01T02:11:53.006237Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print (\"Length of outputList : %d\" % len (outputList))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-01T02:11:53.008179Z","iopub.execute_input":"2025-08-01T02:11:53.008571Z","iopub.status.idle":"2025-08-01T02:11:53.013891Z","shell.execute_reply.started":"2025-08-01T02:11:53.00852Z","shell.execute_reply":"2025-08-01T02:11:53.01273Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"outputList[:2]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-01T02:11:53.014996Z","iopub.execute_input":"2025-08-01T02:11:53.015352Z","iopub.status.idle":"2025-08-01T02:11:53.035141Z","shell.execute_reply.started":"2025-08-01T02:11:53.015324Z","shell.execute_reply":"2025-08-01T02:11:53.034312Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"for img in outputList[:3]:\n    print(img)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-01T02:11:53.036031Z","iopub.execute_input":"2025-08-01T02:11:53.036361Z","iopub.status.idle":"2025-08-01T02:11:53.056919Z","shell.execute_reply.started":"2025-08-01T02:11:53.036338Z","shell.execute_reply":"2025-08-01T02:11:53.055914Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Create relative infos for the converted png","metadata":{}},{"cell_type":"code","source":"# === Load Mapping CSV ===\ndf_mapping = pd.read_csv('series_index_mapping.csv')\nmapping_dict = {\n    (row['SeriesInstanceUID'], row['SOPInstanceUID']): (row['relative_index'], row['dicom_filename'])\n    for _, row in df_mapping.iterrows()\n}\n\nif DEBUG:\n    print(\"DEBUG MODE: Using only first 10 rows of df_localizers\")\n    df_coords = df_localizers.iloc[:10].copy()\nelse:\n    print(\"PRODUCTION MODE: Using full df_localizers\")\n    df_coords = df_localizers.copy()\n\n# Output columns\nrelative_indices = []\nrelative_xs = []\nrelative_ys = []\n\n# Process each row\nfor idx, row in tqdm(df_coords.iterrows(), total=len(df_coords)):\n    key = (row['SeriesInstanceUID'], row['SOPInstanceUID'])\n\n    if key not in mapping_dict:\n        relative_indices.append(None)\n        relative_xs.append(None)\n        relative_ys.append(None)\n        continue\n\n    relative_index, dicom_path = mapping_dict[key]\n    relative_indices.append(relative_index)\n\n    try:\n        ds = pydicom.dcmread(dicom_path, stop_before_pixels=True)\n        h, w = int(ds.Rows), int(ds.Columns)\n\n        coords = row['coordinates']\n        if isinstance(coords, str):\n            coords = eval(coords)\n\n        x_rel = (coords['x'] / w) * GLOBAL_WIDTH\n        y_rel = (coords['y'] / h) * GLOBAL_WIDTH\n\n        relative_xs.append(x_rel)\n        relative_ys.append(y_rel)\n\n    except Exception as e:\n        relative_xs.append(None)\n        relative_ys.append(None)\n\n# Add columns back to DataFrame\ndf_coords['relative_index'] = relative_indices\ndf_coords['relative_x'] = relative_xs\ndf_coords['relative_y'] = relative_ys\n\n# Save\noutput_csv_path = f'train_localizers_with_relative.csv'\ndf_coords.to_csv(output_csv_path, index=False)\nprint(f\"Saved updated DataFrame with relative info to: {output_csv_path}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-01T02:11:53.057845Z","iopub.execute_input":"2025-08-01T02:11:53.058127Z","iopub.status.idle":"2025-08-01T02:11:53.248059Z","shell.execute_reply.started":"2025-08-01T02:11:53.058094Z","shell.execute_reply":"2025-08-01T02:11:53.246948Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Save the PNG from DICOM","metadata":{}},{"cell_type":"code","source":"from multiprocessing import cpu_count\nn_cores = cpu_count()\nprint(f'Number of Logical CPU cores: {n_cores}')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-01T02:11:53.248931Z","iopub.execute_input":"2025-08-01T02:11:53.249178Z","iopub.status.idle":"2025-08-01T02:11:53.254583Z","shell.execute_reply.started":"2025-08-01T02:11:53.24916Z","shell.execute_reply":"2025-08-01T02:11:53.253772Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"start_time = time.time()\n\nParallel(n_jobs=n_cores)(\n    delayed(dicom_to_png)(img['impath'], img['dst'], GLOBAL_WIDTH, apply_windowing=True, modality=img['modality'])\n    for img in outputList\n)\n\nelapsed = time.time() - start_time\n\n# Format time nicely\nhours, rem = divmod(elapsed, 3600)\nminutes, seconds = divmod(rem, 60)\nprint(f\"Total running time: {int(hours)}h {int(minutes)}m {seconds:.2f}s\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-01T02:11:53.255699Z","iopub.execute_input":"2025-08-01T02:11:53.256023Z","iopub.status.idle":"2025-08-01T02:12:07.007383Z","shell.execute_reply.started":"2025-08-01T02:11:53.255991Z","shell.execute_reply":"2025-08-01T02:12:07.006606Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# delete the offline gdcm to create the dataset for Kaggler's use\n\n# Delete folder if it exists\nfolder_path = \"/kaggle/working/gdcm\"\nif os.path.exists(folder_path) and os.path.isdir(folder_path):\n    !rm -rf /kaggle/working/gdcm\n\n# Delete file if it exists\nfile_path = \"/kaggle/working/gdcm.tar\"\nif os.path.exists(file_path) and os.path.isfile(file_path):\n    !rm -f /kaggle/working/gdcm.tar\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-01T02:12:07.008448Z","iopub.execute_input":"2025-08-01T02:12:07.008721Z","iopub.status.idle":"2025-08-01T02:12:07.26376Z","shell.execute_reply.started":"2025-08-01T02:12:07.008701Z","shell.execute_reply":"2025-08-01T02:12:07.26263Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Visualization of the Dicom and converted PNG with useful information","metadata":{}},{"cell_type":"code","source":"\ndef show_dicom_and_png_with_info(row, df_meta, rd, box_size_resized=10):\n\n    # Load the series_index_mapping.csv once\n    df_index_map = pd.read_csv(\"/kaggle/working/series_index_mapping.csv\")\n\n    # Parse identifiers\n    series_uid = row['SeriesInstanceUID']\n    sop_uid = row['SOPInstanceUID']\n    rel_index = row['relative_index']\n    rel_x = row['relative_x']\n    rel_y = row['relative_y']\n    loca = row['location'].replace('/', '_')\n\n    # Get metadata row\n    df_row = df_meta[df_meta['SeriesInstanceUID'] == series_uid].iloc[0]\n    patient_age = df_row['PatientAge']\n    patient_sex = df_row['PatientSex']\n    modality = df_row['Modality']\n\n    # Aneurysm location columns\n    aneurysm_location_cols = [\n        \"Left Infraclinoid Internal Carotid Artery\", \"Right Infraclinoid Internal Carotid Artery\",\n        \"Left Supraclinoid Internal Carotid Artery\", \"Right Supraclinoid Internal Carotid Artery\",\n        \"Left Middle Cerebral Artery\", \"Right Middle Cerebral Artery\",\n        \"Anterior Communicating Artery\", \"Left Anterior Cerebral Artery\",\n        \"Right Anterior Cerebral Artery\", \"Left Posterior Communicating Artery\",\n        \"Right Posterior Communicating Artery\", \"Basilar Tip\",\n        \"Other Posterior Circulation\"\n    ]\n    aneurysm_locations = [col for col in aneurysm_location_cols if df_row.get(col, 0) == 1]\n\n    # === Load PNG ===\n    rel_index_int = int(rel_index)\n    png_path = f'cvt_png/{loca}/{series_uid}/{rel_index_int:04d}.png'\n    img_png = cv2.imread(png_path)\n    if img_png is None:\n        print(f\"[!] PNG not found: {png_path}\")\n        return\n    img_png = cv2.cvtColor(img_png, cv2.COLOR_BGR2RGB)\n\n    # Draw box on PNG\n    x1_png = int(rel_x - box_size_resized / 2)\n    y1_png = int(rel_y - box_size_resized / 2)\n    x2_png = int(rel_x + box_size_resized / 2)\n    y2_png = int(rel_y + box_size_resized / 2)\n    cv2.rectangle(img_png, (x1_png, y1_png), (x2_png, y2_png), (255, 0, 0), 2)\n\n    # === Load DICOM using CSV ===\n    df_series = df_index_map[df_index_map['SeriesInstanceUID'] == series_uid]\n    dicom_path_row = df_series[df_series['SOPInstanceUID'] == sop_uid]\n    if dicom_path_row.empty:\n        print(f\"[!] DICOM path not found in CSV for SOPInstanceUID: {sop_uid}\")\n        return\n\n    dicom_path = dicom_path_row['dicom_filename'].values[0]\n\n    try:\n        ds = pydicom.dcmread(dicom_path)\n        img_dcm = ds.pixel_array.astype(np.float32)\n        img_dcm = (img_dcm - img_dcm.min()) / (img_dcm.max() - img_dcm.min() + 1e-6)\n        img_dcm = (img_dcm * 255).astype(np.uint8)\n        img_dcm_rgb = cv2.cvtColor(img_dcm, cv2.COLOR_GRAY2RGB)\n    except Exception as e:\n        print(f\"[!] Error reading DICOM: {dicom_path} | {e}\")\n        return\n\n    # Draw box on DICOM\n    coords = row['coordinates']\n    if isinstance(coords, str):\n        coords = eval(coords)\n    x_orig = int(coords['x'])\n    y_orig = int(coords['y'])\n\n    h_dcm, w_dcm = img_dcm.shape\n    scale_x = w_dcm / GLOBAL_WIDTH\n    scale_y = h_dcm / GLOBAL_WIDTH\n    box_size_dcm_x = int(box_size_resized * scale_x)\n    box_size_dcm_y = int(box_size_resized * scale_y)\n\n    x1_dcm = max(0, int(x_orig - box_size_dcm_x / 2))\n    y1_dcm = max(0, int(y_orig - box_size_dcm_y / 2))\n    x2_dcm = min(w_dcm - 1, int(x_orig + box_size_dcm_x / 2))\n    y2_dcm = min(h_dcm - 1, int(y_orig + box_size_dcm_y / 2))\n    cv2.rectangle(img_dcm_rgb, (x1_dcm, y1_dcm), (x2_dcm, y2_dcm), (255, 0, 0), 2)\n\n    # === Plot ===\n    fig, axes = plt.subplots(1, 2, figsize=(12, 6))\n    fig.subplots_adjust(top=0.75)\n\n    info_line = f\"Modality: {modality} | Age: {patient_age} | Sex: {patient_sex}\"\n    aneurysm_line = (\n        \"Aneurysm Present in: \" + \", \".join(aneurysm_locations)\n        if aneurysm_locations else \"No Aneurysm Present\"\n    )\n    tech_line = f\"Series UID: {series_uid}\\nLocation: {loca} | Frame Index: {rel_index}\"\n\n    fig.suptitle(\n        f\"{info_line}\\n{aneurysm_line}\\n{tech_line}\",\n        fontsize=12, y=1.1, ha='center'\n    )\n\n    axes[0].imshow(img_dcm_rgb)\n    axes[0].set_title(\"Original DICOM\")\n    axes[0].set_xlabel(\"Pixels\")\n    axes[0].set_ylabel(\"Pixels\")\n    axes[0].grid(True, linestyle='--', linewidth=0.5)\n\n    axes[1].imshow(img_png)\n    axes[1].set_title(f\"Converted & Resized PNG ({GLOBAL_WIDTH}x{GLOBAL_WIDTH})\")\n    axes[1].set_xlabel(\"Pixels\")\n    axes[1].set_ylabel(\"Pixels\")\n    axes[1].grid(True, linestyle='--', linewidth=0.5)\n\n    plt.show()\n\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-01T02:12:07.265082Z","iopub.execute_input":"2025-08-01T02:12:07.265343Z","iopub.status.idle":"2025-08-01T02:12:07.283466Z","shell.execute_reply.started":"2025-08-01T02:12:07.265316Z","shell.execute_reply":"2025-08-01T02:12:07.282345Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\nfor i in range(3):\n    show_dicom_and_png_with_info(df_coords.iloc[i], df, rd)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-01T02:12:07.287657Z","iopub.execute_input":"2025-08-01T02:12:07.288Z","iopub.status.idle":"2025-08-01T02:12:09.250678Z","shell.execute_reply.started":"2025-08-01T02:12:07.287976Z","shell.execute_reply":"2025-08-01T02:12:09.249762Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\ndef concat_dicom_series_to_image(series_uid, location=None, df_coords=None, rd='.', \n                                resize_shape=(GLOBAL_WIDTH, GLOBAL_WIDTH), box_size_resized=10,\n                                layout='vertical', show=True, save_path=None,\n                                use_png=False):\n    \"\"\"\n    Concatenate all images in a DICOM series (or converted PNGs) into a single visual strip or grid.\n\n    Args:\n        series_uid (str): SeriesInstanceUID to visualize.\n        location (str): Needed only if use_png=True. Disease location (for PNG folder path).\n        df_coords (pd.DataFrame or None): Coordinate info for box overlays.\n        rd (str): Root directory containing 'series' folder.\n        resize_shape (tuple): Target (width, height) for resizing each image.\n        box_size_resized (int): Box size (in resized image space).\n        layout (str): 'vertical', 'horizontal', or 'grid'.\n        show (bool): Whether to plot using matplotlib.\n        save_path (str or None): If set, saves the output image.\n        use_png (bool): If True, loads PNGs instead of DICOMs.\n\n    Returns:\n        None\n    \"\"\"\n\n    # Load index mapping from CSV\n    df_index = pd.read_csv('/kaggle/working/series_index_mapping.csv')\n    df_series = df_index[df_index['SeriesInstanceUID'] == series_uid]\n    dicom_paths = df_series.sort_values(by='relative_index')['dicom_filename'].tolist()\n\n    # Build coordinate map from df_coords\n    coord_map = {}\n    if df_coords is not None:\n        df_sub = df_coords[df_coords['SeriesInstanceUID'] == series_uid]\n        for _, row in df_sub.iterrows():\n            sop = row['SOPInstanceUID']\n            coords = row['coordinates']\n            if isinstance(coords, str):\n                coords = eval(coords)\n            coord_map[sop] = coords\n\n    images = []\n\n    if use_png:\n        if location is None:\n            raise ValueError(\"`location` must be provided when use_png=True\")\n\n        png_dir = os.path.join('/kaggle/working/cvt_png', location.replace('/', '_'), series_uid)\n        print(f\"Loading PNGs from: {png_dir}\")\n        png_paths = sorted(glob.glob(os.path.join(png_dir, '*.png')))\n        \n        if df_coords is not None:\n            df_sub = df_coords[df_coords['SeriesInstanceUID'] == series_uid]\n\n        for i, path in enumerate(png_paths):\n            img = cv2.imread(path)\n            if img is None:\n                continue\n            img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)\n            img_resized = cv2.resize(img, resize_shape)\n\n            if df_coords is not None:\n                match = df_sub[df_sub['relative_index'] == i]\n                if not match.empty:\n                    x = match.iloc[0]['relative_x']\n                    y = match.iloc[0]['relative_y']\n                    if pd.notna(x) and pd.notna(y):\n                        x1 = int(x - box_size_resized / 2)\n                        y1 = int(y - box_size_resized / 2)\n                        x2 = int(x + box_size_resized / 2)\n                        y2 = int(y + box_size_resized / 2)\n                        cv2.rectangle(img_resized, (x1, y1), (x2, y2), (255, 0, 0), 2)\n\n            images.append(img_resized)\n\n    else:\n        for path in dicom_paths:\n            try:\n                ds = pydicom.dcmread(path)\n                img = ds.pixel_array.astype(np.float32)\n                img = (img - img.min()) / (img.max() - img.min() + 1e-6)\n                img = (img * 255).astype(np.uint8)\n                img_rgb = cv2.cvtColor(img, cv2.COLOR_GRAY2RGB)\n                img_resized = cv2.resize(img_rgb, resize_shape)\n\n                sop = ds.SOPInstanceUID\n                if sop in coord_map:\n                    orig_h, orig_w = img.shape\n                    x_orig = coord_map[sop]['x']\n                    y_orig = coord_map[sop]['y']\n\n                    # Scale original coords to resized image\n                    x = int((x_orig / orig_w) * resize_shape[0])\n                    y = int((y_orig / orig_h) * resize_shape[1])\n\n                    x1 = int(x - box_size_resized / 2)\n                    y1 = int(y - box_size_resized / 2)\n                    x2 = int(x + box_size_resized / 2)\n                    y2 = int(y + box_size_resized / 2)\n                    cv2.rectangle(img_resized, (x1, y1), (x2, y2), (255, 0, 0), 2)\n\n                images.append(img_resized)\n            except Exception as e:\n                print(f\"Failed to load DICOM: {path}: {e}\")\n                continue\n\n    if not images:\n        print(f\"[!] No images found for series: {series_uid}\")\n        return None\n\n    # Layout image\n    if layout == 'vertical':\n        final_image = cv2.vconcat(images)\n    elif layout == 'horizontal':\n        final_image = cv2.hconcat(images)\n    elif layout == 'grid':\n        grid_size = int(np.ceil(np.sqrt(len(images))))\n        while len(images) < grid_size ** 2:\n            images.append(np.zeros_like(images[0]))\n        rows = [\n            cv2.hconcat(images[i*grid_size:(i+1)*grid_size])\n            for i in range(grid_size)\n        ]\n        final_image = cv2.vconcat(rows)\n    else:\n        raise ValueError(\"layout must be 'vertical', 'horizontal', or 'grid'\")\n\n    if show:\n        plt.figure(figsize=(12, 12))\n        plt.imshow(final_image)\n        plt.title(f\"Series UID: {series_uid} (using {'PNG' if use_png else 'DICOM'})\")\n        plt.axis('off')\n        plt.show()\n\n    if save_path:\n        cv2.imwrite(save_path, cv2.cvtColor(final_image, cv2.COLOR_RGB2BGR))\n        print(f\"Saved: {save_path}\")\n\n    return\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-01T02:12:09.25172Z","iopub.execute_input":"2025-08-01T02:12:09.252012Z","iopub.status.idle":"2025-08-01T02:12:09.27104Z","shell.execute_reply.started":"2025-08-01T02:12:09.251989Z","shell.execute_reply":"2025-08-01T02:12:09.269944Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# View entire series with annotation if coordinates available\n\nconcat_dicom_series_to_image(\n    series_uid='1.2.826.0.1.3680043.8.498.10022796280698534221758473208024838831',\n    df_coords=df_coords,\n    rd=rd,\n    show=True,\n    use_png=False,\n    layout='grid'\n)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-01T02:12:09.272196Z","iopub.execute_input":"2025-08-01T02:12:09.272534Z","iopub.status.idle":"2025-08-01T02:12:26.145079Z","shell.execute_reply.started":"2025-08-01T02:12:09.272504Z","shell.execute_reply":"2025-08-01T02:12:26.144003Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"concat_dicom_series_to_image(\n    series_uid='1.2.826.0.1.3680043.8.498.10022796280698534221758473208024838831',\n    df_coords=df_coords,\n    rd=rd,\n    show=True,\n    location='Right Middle Cerebral Artery',\n    use_png=True,\n    layout='grid'\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-01T02:12:26.146205Z","iopub.execute_input":"2025-08-01T02:12:26.146495Z","iopub.status.idle":"2025-08-01T02:12:31.217646Z","shell.execute_reply.started":"2025-08-01T02:12:26.146472Z","shell.execute_reply":"2025-08-01T02:12:31.216398Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Animation","metadata":{}},{"cell_type":"code","source":"\ndef animate_dicom_series_to_gif(series_uid, location=None, df_coords=None, rd='.',\n                                 resize_shape=(GLOBAL_WIDTH, GLOBAL_WIDTH), box_size_resized=10,\n                                 use_png=False, gif_path='aneurysm_trace.gif', duration=0.3):\n\n    frames = []\n\n    # === Step 1: Get all bounding boxes from df_coords ===\n    boxes_to_draw = []\n\n    if df_coords is not None:\n        df_sub = df_coords[df_coords['SeriesInstanceUID'] == series_uid]\n        for _, row in df_sub.iterrows():\n            if use_png:\n                rel_x = row.get('relative_x')\n                rel_y = row.get('relative_y')\n                if pd.notna(rel_x) and pd.notna(rel_y):\n                    x1 = int(rel_x - box_size_resized / 2)\n                    y1 = int(rel_y - box_size_resized / 2)\n                    x2 = int(rel_x + box_size_resized / 2)\n                    y2 = int(rel_y + box_size_resized / 2)\n                    boxes_to_draw.append((x1, y1, x2, y2))\n            else:\n                coords = row.get('coordinates')\n                if isinstance(coords, str):\n                    coords = eval(coords)\n                if isinstance(coords, dict):\n                    x = coords.get('x')\n                    y = coords.get('y')\n                    if x is not None and y is not None:\n                        boxes_to_draw.append((x, y))  # True (unscaled) position\n\n    # === Step 2: Read PNGs or DICOMs ===\n    if use_png:\n        if location is None:\n            raise ValueError(\"`location` must be provided when using PNGs.\")\n\n        png_dir = os.path.join('/kaggle/working/cvt_png', location.replace('/', '_'), series_uid)\n        png_paths = sorted(glob.glob(os.path.join(png_dir, '*.png')))\n\n        for path in png_paths:\n            img = cv2.imread(path)\n            if img is None:\n                continue\n            img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)\n\n            # Draw all bounding boxes\n            for x1, y1, x2, y2 in boxes_to_draw:\n                cv2.rectangle(img, (x1, y1), (x2, y2), (255, 0, 0), 2)\n\n            frames.append(img)\n\n    else:\n        # Load paths from CSV instead of glob\n        df_index = pd.read_csv('series_index_mapping.csv')\n        df_series = df_index[df_index['SeriesInstanceUID'] == series_uid]\n        dicom_paths = df_series.sort_values(by='relative_index')['dicom_filename'].tolist()\n\n        for path in dicom_paths:\n            try:\n                ds = pydicom.dcmread(path)\n                img = ds.pixel_array.astype(np.float32)\n                img = (img - img.min()) / (img.max() - img.min() + 1e-6)\n                img = (img * 255).astype(np.uint8)\n                img = cv2.cvtColor(img, cv2.COLOR_GRAY2RGB)\n\n                # Resize the image to target shape\n                h_orig, w_orig = img.shape[:2]\n                img = cv2.resize(img, resize_shape)\n\n                # Rescale original positions to resized image\n                scale_x = resize_shape[0] / w_orig\n                scale_y = resize_shape[1] / h_orig\n\n                for x_orig, y_orig in boxes_to_draw:\n                    x1 = int(x_orig * scale_x - box_size_resized / 2)\n                    y1 = int(y_orig * scale_y - box_size_resized / 2)\n                    x2 = int(x_orig * scale_x + box_size_resized / 2)\n                    y2 = int(y_orig * scale_y + box_size_resized / 2)\n                    cv2.rectangle(img, (x1, y1), (x2, y2), (255, 0, 0), 2)\n\n                frames.append(img)\n\n            except Exception as e:\n                print(f\"Skipped {path}: {e}\")\n                continue\n\n    # === Step 3: Save GIF ===\n    if not frames:\n        print(f\"No frames to animate for {series_uid}\")\n        return\n\n    imageio.mimsave(gif_path, frames, duration=duration)\n    print(f\"GIF saved: {gif_path}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-01T02:12:31.219249Z","iopub.execute_input":"2025-08-01T02:12:31.219673Z","iopub.status.idle":"2025-08-01T02:12:31.242939Z","shell.execute_reply.started":"2025-08-01T02:12:31.219632Z","shell.execute_reply":"2025-08-01T02:12:31.241856Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"animate_dicom_series_to_gif(\n    series_uid='1.2.826.0.1.3680043.8.498.10022796280698534221758473208024838831',\n    location='Right Middle Cerebral Artery',\n    df_coords=df_coords,\n    rd='/kaggle/input/rsna-intracranial-aneurysm-detection',\n    use_png=True,  # or False for original DICOMs\n    gif_path='aneurysm_trace.gif',\n    duration=0.4\n)\n\n#display gif\n\ngif_path = \"aneurysm_trace.gif\"\nHTML(f'<img src=\"{gif_path}\" style=\"max-width:100%;\"/>')\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-01T02:12:31.244137Z","iopub.execute_input":"2025-08-01T02:12:31.244486Z","iopub.status.idle":"2025-08-01T02:12:38.01543Z","shell.execute_reply.started":"2025-08-01T02:12:31.244458Z","shell.execute_reply":"2025-08-01T02:12:38.014184Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}