{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.12.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[]}},"nbformat_minor":4,"nbformat":4,"cells":[{"id":"700d0217-396c-4d0c-8717-52ac11850abc","cell_type":"markdown","source":"# 1/5 — DICOM Is Not PNG\n## A reproducible CT audit, HU conversion and multi-window pipeline\n\n**Series:** Anatomy-First Cervical Spine CT Triage  \n**Question:** What can silently go wrong before the first tensor reaches a model?\n\nWe will:\n\n1. discover the RSNA competition files without hard-coded notebook IDs;\n2. build a study-level manifest;\n3. audit DICOM geometry and rescale metadata;\n4. sort a CT series in physical order;\n5. convert stored pixels to Hounsfield units (HU);\n6. compare bone, soft-tissue and wide-bone windows;\n7. export a small, reusable data contract for the rest of the series.\n\nThis notebook deliberately trains no model. A broken data pipeline can produce a beautifully converged wrong answer.\n","metadata":{}},{"id":"1ffdd836-c7f9-46af-b981-92605a6b16a8","cell_type":"markdown","source":"### Data source\n\nAttach the [**RSNA 2022 Cervical Spine Fracture Detection**](https://www.kaggle.com/competitions/rsna-2022-cervical-spine-fracture-detection/data) competition data. The code searches both Kaggle and local paths. Dataset context: [RSNA challenge page](https://www.rsna.org/artificial-intelligence/ai-image-challenge/cervical-spine-fractures-ai-detection-challenge-2022) and [dataset paper](https://pmc.ncbi.nlm.nih.gov/articles/PMC10546361/).","metadata":{}},{"id":"158175bd-6017-4033-8af9-499e2d33c14b","cell_type":"code","source":"!pip install -q pylibjpeg pylibjpeg-libjpeg pylibjpeg-openjpeg python-gdcm","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-29T13:30:05.366437Z","iopub.execute_input":"2026-06-29T13:30:05.367234Z","iopub.status.idle":"2026-06-29T13:30:12.523486Z","shell.execute_reply.started":"2026-06-29T13:30:05.367202Z","shell.execute_reply":"2026-06-29T13:30:12.522439Z"}},"outputs":[],"execution_count":null},{"id":"82e10f08-cf5a-4852-bfd1-eb9de424998c","cell_type":"code","source":"from __future__ import annotations\n\nimport json\nimport os\nimport random\nfrom pathlib import Path\n\nimport matplotlib.pyplot as plt\nimport numpy as np\nimport pandas as pd\npd.set_option(\"display.max_columns\", 200)\nimport pydicom\nfrom IPython.display import display\nfrom tqdm.auto import tqdm\n\nSEED = 42\nFAST_MODE = True  # False scans all training studies before publication.\nMAX_STUDIES = 160 if FAST_MODE else None\nrandom.seed(SEED)\nnp.random.seed(SEED)\n\nKAGGLE_INPUT = Path(\"/kaggle/input\")\nWORK_ROOT = Path(\"/kaggle/working\") if Path(\"/kaggle/working\").exists() else Path(\"outputs\")\nOUT_DIR = WORK_ROOT / \"cspine_artifacts\" / \"01_data_audit\"\nOUT_DIR.mkdir(parents=True, exist_ok=True)\nprint(\"Output:\", OUT_DIR.resolve())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-29T13:30:15.377596Z","iopub.execute_input":"2026-06-29T13:30:15.377918Z","iopub.status.idle":"2026-06-29T13:30:16.87658Z","shell.execute_reply.started":"2026-06-29T13:30:15.377885Z","shell.execute_reply":"2026-06-29T13:30:16.875683Z"}},"outputs":[],"execution_count":null},{"id":"280b4746-2959-4d3a-9869-89534a384953","cell_type":"markdown","source":"## 1. Locate the competition data\n\nKaggle mount names can change when a dataset is copied or versioned. We identify the root by its required files instead of assuming one exact directory name.\n","metadata":{}},{"id":"9e78778d-2240-44f7-8a6c-3b0390f36284","cell_type":"code","source":"def find_rsna_root() -> Path:\n    fast_candidates = [\n        KAGGLE_INPUT / \"rsna-2022-cervical-spine-fracture-detection\",\n        Path(\"data/rsna-2022-cervical-spine-fracture-detection\"),\n        Path(\"data\"),\n    ]\n    \n    for root in fast_candidates:\n        if (root / \"train.csv\").exists() and (root / \"train_images\").exists():\n            return root\n            \n    if KAGGLE_INPUT.exists():\n        for p in KAGGLE_INPUT.rglob(\"train.csv\"):\n            root = p.parent\n            if (root / \"train_images\").exists():\n                return root\n                \n    raise FileNotFoundError(\n        \"RSNA data not found. Attach the competition data, then rerun this cell.\"\n    )\n\n\nDATA_ROOT = find_rsna_root()\nTRAIN_IMAGE_ROOT = DATA_ROOT / \"train_images\"\ntrain_df = pd.read_csv(DATA_ROOT / \"train.csv\")\nbbox_path = DATA_ROOT / \"train_bounding_boxes.csv\"\nbbox_df = pd.read_csv(bbox_path) if bbox_path.exists() else pd.DataFrame()\n\nprint(\"Data root:\", DATA_ROOT)\nprint(\"Training studies:\", len(train_df))\nprint(\"Bounding-box rows:\", len(bbox_df))\ndisplay(train_df.head())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-29T13:30:19.285645Z","iopub.execute_input":"2026-06-29T13:30:19.286697Z","iopub.status.idle":"2026-06-29T13:30:19.367131Z","shell.execute_reply.started":"2026-06-29T13:30:19.286664Z","shell.execute_reply":"2026-06-29T13:30:19.366175Z"}},"outputs":[],"execution_count":null},{"id":"fe7ff591-f896-478c-aab7-1a20af3dae1f","cell_type":"code","source":"bbox_df.head(2)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-29T13:22:37.794009Z","iopub.execute_input":"2026-06-29T13:22:37.794352Z","iopub.status.idle":"2026-06-29T13:22:37.805628Z","shell.execute_reply.started":"2026-06-29T13:22:37.794324Z","shell.execute_reply":"2026-06-29T13:22:37.804838Z"}},"outputs":[],"execution_count":null},{"id":"aa8da28a-8a0d-4dbe-9eaf-098354b0da7b","cell_type":"code","source":"train_df[\"patient_overall\"].value_counts()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-29T13:22:39.794018Z","iopub.execute_input":"2026-06-29T13:22:39.7949Z","iopub.status.idle":"2026-06-29T13:22:39.802502Z","shell.execute_reply.started":"2026-06-29T13:22:39.794867Z","shell.execute_reply":"2026-06-29T13:22:39.801546Z"}},"outputs":[],"execution_count":null},{"id":"9cab3988-3e44-4ec8-8760-4b87b14ae574","cell_type":"markdown","source":"## 2. Study-level label audit\n\n`patient_overall` answers whether any cervical level is fractured; `C1`–`C7` are vertebra-level targets. We keep the original study identifier intact and never create folds at slice level.\n","metadata":{}},{"id":"45e83624-cf76-4655-8ca5-3be8d0df8002","cell_type":"code","source":"# ---------------------------------------------------------\n# 定義預測目標與資料驗證 (Define Targets & Sanity Checks)\n# ---------------------------------------------------------\n\n# 定義預測目標列表，包含整體狀況 (\"patient_overall\") 與 C1 到 C7 的頸椎節段。\n# Define the target list, including overall patient status and cervical spine levels C1 to C7.\nTARGETS = [\"patient_overall\"] + [f\"C{i}\" for i in range(1, 8)]\n\n# 確保訓練集中的檢查單號 (StudyInstanceUID) 是唯一的，沒有重複值。\n# Ensure that there are no duplicate study instance UIDs in the training dataframe.\nassert train_df[\"StudyInstanceUID\"].is_unique\n\n# 確保目標欄位的值去重複後，只包含 0 或 1（確認資料是乾淨的二元分類標籤）。\n# Assert that the unique values in target columns are only 0s or 1s (confirming clean binary labels).\nassert set(np.unique(train_df[TARGETS].values)).issubset({0, 1})\n\n# ---------------------------------------------------------\n# 計算資料分佈比例 (Calculate Prevalence)\n# ---------------------------------------------------------\n\n# 計算每個目標欄位的平均值（即陽性/骨折比例），並由高到低降冪排序。\n# Calculate the positive fraction (mean) for each target and sort them in descending order.\nprevalence = train_df[TARGETS].mean().sort_values(ascending=False)\n\n# 將計算結果重新命名為 \"prevalence\"，轉換成 DataFrame 格式後在畫面上漂亮地印出。\n# Rename the series to \"prevalence\", convert it to a DataFrame, and display it formatting nicely.\ndisplay(prevalence.rename(\"prevalence\").to_frame())\n\n\n# ---------------------------------------------------------\n# 視覺化與輸出 (Visualization & Output)\n# ---------------------------------------------------------\n\n# 使用上述比例資料繪製長條圖 (bar chart)，設定圖片大小 (10x4) 與指定的藍色。\n# Plot a bar chart using the prevalence data, specifying figure size (10x4) and a specific blue color.\nax = prevalence.plot(kind=\"bar\", figsize=(10, 4), color=\"#2878B5\")\n\n# 設定圖表標題、Y 軸標籤，並將 Y 軸的顯示範圍固定在 0 到 1 之間。\n# Set the chart title, Y-axis label, and limit the Y-axis range from 0 to 1.\nax.set(title=\"Study-level target prevalence\", ylabel=\"Positive fraction\", ylim=(0, 1))\n\n# 將 X 軸的文字標籤旋轉 30 度，並設定向右對齊，避免文字互相重疊。\n# Rotate the X-axis tick labels by 30 degrees and align them to the right to prevent overlap.\nplt.xticks(rotation=30, ha=\"right\")\n\n# 自動調整圖表的邊距與佈局，確保所有標籤與標題都能完整顯示在畫布內。\n#  Automatically adjust subplot parameters to give specified padding and ensure everything fits.\nplt.tight_layout()\n\n# 將完成的圖表以 160 DPI 的解析度，存檔到指定的輸出目錄 (OUT_DIR) 中。\n#  Save the figure to the specified output directory (OUT_DIR) with a resolution of 160 DPI.\nplt.savefig(OUT_DIR / \"target_prevalence.png\", dpi=160)\n\n# 在畫面上渲染並顯示最終圖表。\n# Render and display the final plot on the screen.\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-29T13:30:22.777102Z","iopub.execute_input":"2026-06-29T13:30:22.77755Z","iopub.status.idle":"2026-06-29T13:30:23.295052Z","shell.execute_reply.started":"2026-06-29T13:30:22.777523Z","shell.execute_reply":"2026-06-29T13:30:23.294139Z"}},"outputs":[],"execution_count":null},{"id":"613ab66e-1798-4e9d-9da9-47b8dd0f7119","cell_type":"markdown","source":"## 3. Header-only DICOM audit\n\nReading every pixel array is unnecessary for a geometry audit. For each study we sample the first, middle and last filenames and read metadata only. Filename order is not treated as physical order.\n","metadata":{}},{"id":"0ee63830-4909-4e2f-b0b0-40c29bc1a304","cell_type":"code","source":"# ---------------------------------------------------------\n# 輔助清理函數：處理 DICOM 標籤的複雜格式\n# Helper cleaning function: Handle complex formats of DICOM tags\n# ---------------------------------------------------------\n\n\n# 確保數值轉換為浮點數，若發生錯誤則回傳預設值 (預設為 NaN)。\n# Ensure the value is converted to a float; return the default (NaN by default) if an error occurs.\ndef scalar(value, default=np.nan) -> float:\n    try:\n        # 如果數值是列表、元組或 DICOM 的多值型態，只提取第一個元素。\n        # If the value is a list, tuple, or DICOM MultiValue, extract only the first element.\n        if isinstance(value, (list, tuple, pydicom.multival.MultiValue)):\n            value = value[0]\n        # 將數值強制轉型為浮點數並回傳。\n        # Cast the value to float and return it.\n        return float(value)\n    except (TypeError, ValueError):\n        # 若遭遇型別或數值轉換錯誤，回傳浮點數格式的預設值。\n        # If a TypeError or ValueError is encountered, return the default value as a float.\n        return float(default)\n\n# 提取 DICOM 影像在病患體內的 Z 軸絕對物理位置。\n# Extract the absolute physical Z-axis position of the DICOM image within the patient.\ndef position_z(ds) -> float:\n    # 安全地獲取 \"ImagePositionPatient\" 標籤的值，若無則回傳 None。\n    # Safely get the value of the \"ImagePositionPatient\" tag, returning None if it doesn't exist.\n    value = getattr(ds, \"ImagePositionPatient\", None)\n    # 如果值存在且至少有三個元素 (X, Y, Z)，則回傳 Z 軸數值 (索引 2)；否則回傳 NaN。\n    # If the value exists and has at least 3 elements (X, Y, Z), return the Z-axis value (index 2); otherwise return NaN.\n    return scalar(value[2]) if value is not None and len(value) >= 3 else np.nan\n\n\n# ---------------------------------------------------------\n# 核心審查函數：提取單一病患的掃描元資料\n# Core audit function: Extract scan metadata for a single patient\n# ---------------------------------------------------------\n\n# 審查特定檢查單號 (病患) 的目錄，提取關鍵的物理與設備參數。\n# Audit a specific study UID (patient) directory to extract key physical and equipment parameters.\ndef audit_study(study_uid: str) -> dict:\n    # 定義該病患的 DICOM 影像資料夾路徑。\n    # Define the folder path for the patient's DICOM images.\n    folder = TRAIN_IMAGE_ROOT / study_uid\n\n    # 取得資料夾內所有 .dcm 檔案的路徑並進行排序。\n    # Get and sort all .dcm file paths within the folder.\n    paths = sorted(folder.glob(\"*.dcm\"))\n\n    # 初始化回傳字典，記錄檢查單號與總切片數量。\n    # Initialize the return dictionary, recording the study UID and the total number of slices.\n    row = {\"StudyInstanceUID\": study_uid, \"n_slices\": len(paths)}\n\n    # 如果找不到任何 DICOM 檔案，標記狀態為遺失並提早回傳。\n    # If no DICOM files are found, mark the status as missing and return early.\n    if not paths:\n        return {**row, \"status\": \"missing_dicom\"}\n\n    # 為了效能，只挑選 3 張切片進行檢查：第一張、中間那張、最後一張。\n    # For performance, select only 3 slices to probe: the first, the middle, and the last.\n    probe_indices = sorted(set([0, len(paths) // 2, len(paths) - 1]))\n    \n    # 讀取這 3 張切片的標頭檔 (設定 stop_before_pixels=True 以略過龐大的像素資料，大幅提升速度)。\n    # Read the headers of these 3 slices (setting stop_before_pixels=True skips bulky pixel data, massively boosting speed).\n    headers = [pydicom.dcmread(paths[i], stop_before_pixels=True) for i in probe_indices]\n    \n    # 選擇中間那張切片的標頭檔作為整份檢查的代表特徵。\n    # Select the middle slice's header as the representative feature for the entire study.\n    mid = headers[len(headers) // 2]\n    \n    # 獲取像素間距 (PixelSpacing)，若缺失則預設為 [NaN, NaN]。\n    # Get PixelSpacing; default to [NaN, NaN] if missing.\n    spacings = getattr(mid, \"PixelSpacing\", [np.nan, np.nan])\n\n    # 更新字典，填入從標頭檔提取的各項關鍵參數。\n    # Update the dictionary with various key parameters extracted from the header.\n    row.update({\n        \"status\": \"ok\",  # 狀態正常 / Status OK\n        \"rows\": int(getattr(mid, \"Rows\", 0)),  # 影像高度 (列數) / Image height (number of rows)\n        \"columns\": int(getattr(mid, \"Columns\", 0)),  # 影像寬度 (行數) / Image width (number of columns)\n        \"pixel_spacing_y\": scalar(spacings[0]),  # Y 軸像素物理尺寸 / Y-axis pixel physical size\n        \"pixel_spacing_x\": scalar(spacings[1]),  # X 軸像素物理尺寸 / X-axis pixel physical size\n        \"slice_thickness\": scalar(getattr(mid, \"SliceThickness\", np.nan)),  # 切片厚度 / Slice thickness\n        \"rescale_slope\": scalar(getattr(mid, \"RescaleSlope\", 1.0), 1.0),  # 像素值轉換斜率 (用於計算 HU) / Rescale slope (used to compute HU)\n        \"rescale_intercept\": scalar(getattr(mid, \"RescaleIntercept\", 0.0), 0.0),  # 像素值轉換截距 / Rescale intercept\n        \"manufacturer\": str(getattr(mid, \"Manufacturer\", \"unknown\")),  # 掃描儀器製造商 / Scanner manufacturer\n        \"transfer_syntax\": str(getattr(mid.file_meta, \"TransferSyntaxUID\", \"unknown\")),  # 檔案傳輸語法 / Transfer syntax\n        # 計算最大與最小 Z 軸座標的差值，得出整份檢查的 Z 軸物理跨度 (毫米)。\n        # Calculate the difference between max and min Z-coordinates to get the physical Z-span of the study (in mm).\n        \"sample_z_span_mm\": float(np.nanmax([position_z(x) for x in headers]) - np.nanmin([position_z(x) for x in headers])),\n    })\n    return row\n\n# ---------------------------------------------------------\n# 批次執行與資料合併\n# Batch execution and data merging\n# ---------------------------------------------------------\n\n# 將訓練集所有的檢查單號轉為字串格式的列表。\n# Convert all study UIDs in the training set to a list of strings.\nstudy_ids = train_df[\"StudyInstanceUID\"].astype(str).tolist()\n\n# 如果設定了最大研究數量 (MAX_STUDIES)，則進行隨機抽樣。\n# If a maximum number of studies (MAX_STUDIES) is set, perform random sampling.\nif MAX_STUDIES is not None:\n    rng = np.random.default_rng(SEED)  # 初始化隨機數生成器 / Initialize random number generator\n    # 隨機抽取指定數量的不重複檢查單號。\n    # Randomly sample the specified number of unique study UIDs without replacement.\n    study_ids = rng.choice(study_ids, size=min(MAX_STUDIES, len(study_ids)), replace=False).tolist()\n\n# 使用 tqdm 顯示進度條，對列表中的每個檢查單號執行 audit_study，並將結果轉換為 DataFrame。\n# Use tqdm for a progress bar, run audit_study on each UID in the list, and convert results to a DataFrame.\naudit_df = pd.DataFrame([audit_study(uid) for uid in tqdm(study_ids, desc=\"Header audit\")])\n\n# 以檢查單號作為鍵值，將原始訓練資料 (左表) 與剛提取的元資料 (右表) 進行左連接合併。\n# Left merge the original training data (left) with the newly extracted metadata (right) using StudyInstanceUID as the key.\nmanifest_df = train_df.merge(audit_df, on=\"StudyInstanceUID\", how=\"left\")\n\n# 在畫面上印出合併後資料表的前幾筆記錄。\n# Display the first few rows of the merged dataframe on the screen.\ndisplay(manifest_df.head())\n\n# 印出關鍵物理參數的統計描述摘要 (平均值、標準差、最大最小值等)。\n# Display the statistical description summary (mean, std, min/max, etc.) of key physical parameters.\ndisplay(manifest_df[[\"n_slices\", \"slice_thickness\", \"pixel_spacing_x\", \"pixel_spacing_y\"]].describe())\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-29T13:30:25.949864Z","iopub.execute_input":"2026-06-29T13:30:25.950686Z","iopub.status.idle":"2026-06-29T13:30:36.072Z","shell.execute_reply.started":"2026-06-29T13:30:25.950607Z","shell.execute_reply":"2026-06-29T13:30:36.071127Z"}},"outputs":[],"execution_count":null},{"id":"abe3608f-f9c8-475a-b02e-11e7c820db94","cell_type":"markdown","source":"## 4. Correct series sorting and HU conversion\n\nPreferred ordering uses `ImagePositionPatient`; `InstanceNumber` is only a fallback. HU conversion is performed per slice because slope/intercept are DICOM attributes and should not be assumed globally constant.\n","metadata":{}},{"id":"7c120a0e-a5ca-4c1b-bd26-3d4f9985c8c3","cell_type":"code","source":"# 定義 DICOM 切片的排序邏輯，確保 3D 影像是按照人體解剖結構正確排列。\n# Define the sorting logic for DICOM slices to ensure the 3D image is ordered anatomically.\ndef dicom_sort_key(ds) -> tuple:\n    # 嘗試取得切片的 Z 軸物理位置。\n    # Try to get the physical Z-axis position of the slice.\n    z = position_z(ds)\n    # 如果 Z 軸數值存在且有效，則優先以 Z 軸座標作為排序依據（回傳 0 表示高優先級）。\n    # If the Z value is finite, prioritize sorting by Z-axis coordinate (return 0 for high priority).\n    if np.isfinite(z):\n        return (0, z)\n    # 如果缺少 Z 軸資訊，則退而求其次，使用檔案內建的「實例編號 (InstanceNumber)」排序。\n    # If Z-axis info is missing, fall back to sorting by the built-in \"InstanceNumber\".\n    return (1, scalar(getattr(ds, \"InstanceNumber\", 0), 0))\n\n\n# 將原始的 DICOM 像素值轉換為具有醫學物理意義的 Hounsfield Unit (HU) 值。\n# Convert raw DICOM pixel values to medically and physically meaningful Hounsfield Units (HU).\ndef to_hu(ds) -> np.ndarray:\n    # 將像素陣列轉換為 32 位元浮點數。\n    # Convert the pixel array to 32-bit floats.\n    pixels = ds.pixel_array.astype(np.float32)\n    # 取得轉換斜率，預設為 1.0。\n    # Get the rescale slope, default to 1.0.\n    slope = scalar(getattr(ds, \"RescaleSlope\", 1.0), 1.0)\n    # 取得轉換截距，預設為 0.0。\n    # Get the rescale intercept, default to 0.0.\n    intercept = scalar(getattr(ds, \"RescaleIntercept\", 0.0), 0.0)\n    # 套用線性轉換公式：HU = 像素值 * 斜率 + 截距。\n    # Apply the linear transformation formula: HU = pixel_value * slope + intercept.\n    return pixels * slope + intercept\n\n\n# 讀取並重組單一病患完整的 3D CT 影像矩陣。\n# Load and reconstruct a complete 3D CT image volume for a single patient.\ndef load_ct_series(study_uid: str) -> tuple[np.ndarray, list]:\n    # 取得該病患資料夾內所有 .dcm 檔案的路徑並排序。\n    # Get and sort all .dcm file paths within the patient's folder.\n    paths = sorted((TRAIN_IMAGE_ROOT / study_uid).glob(\"*.dcm\"))\n    # 將所有切片檔案的標頭與像素資料讀取進記憶體。\n    # Read all slice files into memory.\n    slices = [pydicom.dcmread(path) for path in paths]\n    # 使用我們剛剛定義的排序函數，將切片依照真實物理空間排序。\n    # Sort the slices according to actual physical space using our sorting function.\n    slices.sort(key=dicom_sort_key)\n    # 將所有切片轉換為 HU 值，並堆疊成一個完整的三維 NumPy 陣列 (Volume)。\n    # Convert all slices to HU values and stack them into a complete 3D NumPy array (Volume).\n    volume = np.stack([to_hu(ds) for ds in slices]).astype(np.float32)\n    # 回傳 3D 陣列與原始的 DICOM 標頭物件列表。\n    # Return the 3D volume and the list of original DICOM header objects.\n    return volume, slices\n\n\n# 執行 CT 窗位設定 (Windowing)，將特定範圍的 HU 值正規化到 0~1 之間，藉此突顯特定組織。\n# Perform CT windowing, normalizing a specific range of HU values to 0~1 to highlight specific tissues.\ndef window_hu(image: np.ndarray, center: float, width: float) -> np.ndarray:\n    # 根據窗位中心 (center) 與窗寬 (width)，計算出數值的上下限。\n    # Calculate the lower and upper bounds based on the window center and width.\n    low, high = center - width / 2.0, center + width / 2.0\n    # 將影像數值裁切 (clip) 在上下限之內，然後縮放至 0.0 到 1.0 的範圍。\n    # Clip the image values within the bounds, then scale them to the 0.0 to 1.0 range.\n    return ((np.clip(image, low, high) - low) / (high - low)).astype(np.float32)\n\n\n# 預定義常用的醫學窗位設定字典：(Center, Width)。\n# Predefine a dictionary of common medical window settings: (Center, Width).\nWINDOWS = {\n    # 軟組織窗 (觀察肌肉、血管、器官) / Soft tissue window (for muscles, organs)\n    \"soft_tissue\": (50, 350),   \n    # 骨頭窗 (觀察頸椎骨骼結構) / Bone window (for cervical spine bone structures)\n    \"bone\": (400, 2000),        \n    # 寬骨頭窗 (用於密度極高的部位或減少金屬植入物的偽影) / Wide bone window (for high-density areas or reducing metal artifacts)\n    \"wide_bone\": (500, 3000),   \n}","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-29T13:30:38.258794Z","iopub.execute_input":"2026-06-29T13:30:38.259207Z","iopub.status.idle":"2026-06-29T13:30:38.270733Z","shell.execute_reply.started":"2026-06-29T13:30:38.259172Z","shell.execute_reply":"2026-06-29T13:30:38.269724Z"}},"outputs":[],"execution_count":null},{"id":"6362da62-d083-4064-9a9f-c74690fe53b6","cell_type":"markdown","source":"## 5. One study, three windows\n\nThese windows are task-oriented defaults, not universal clinical display standards. Their values must be stored with the model package because changing them changes the model input distribution.\n","metadata":{}},{"id":"7fcd3dc9-50e8-457e-82a3-382c569ad18f","cell_type":"code","source":"# ---------------------------------------------------------\n# 強制註冊解碼器與環境設定\n# Force register decoders and environment setup\n# ---------------------------------------------------------\n\n# 匯入處理 DICOM 醫療影像的基礎套件。\n# Import the base package for handling DICOM medical images.\nimport pydicom\n\n# 匯入 JPEG Lossless 影像解碼套件。\n# Import the JPEG Lossless image decoding package.\nimport pylibjpeg\n\n# 從 pydicom 匯入全域設定模組。\n# Import the global configuration module from pydicom.\nfrom pydicom import config\n\n# 匯入特定的像素資料處理器（用於處理壓縮的醫療影像）。\n# Import specific pixel data handlers (for handling compressed medical images).\nfrom pydicom.pixel_data_handlers import pylibjpeg_handler, gdcm_handler\n\n# 強制將剛匯入的處理器設定為優先，解決解壓縮報錯的問題。\n# Force set the imported handlers as priority to resolve decompression errors.\nconfig.image_handlers = [pylibjpeg_handler, gdcm_handler]\n\n# ---------------------------------------------------------\n# 載入病患資料與 3D 影像重構\n# Load patient data and reconstruct 3D volume\n# ---------------------------------------------------------\n\n# 從訓練標籤表中，找出第一位「整體判定為有骨折 (patient_overall=1)」的病患 ID。\n# Find the ID of the first patient marked as having a fracture overall (patient_overall=1) in the training labels.\npositive_uid = train_df.loc[train_df[\"patient_overall\"].eq(1), \"StudyInstanceUID\"].iloc[0]\n\n# 呼叫我們先前定義的函數，載入並重構該病患完整的 3D CT 影像矩陣與標頭檔。\n# Call our previously defined function to load and reconstruct the full 3D CT volume and headers for this patient.\nvolume, slices = load_ct_series(positive_uid)\n\n# 計算 Z 軸的中間切片索引（取得位於頸椎中段的切片位置）。\n# Calculate the middle slice index on the Z-axis (getting a slice positioned in the mid-cervical spine).\nz = len(volume) // 2\n\n\n# ---------------------------------------------------------\n# 視覺化：多窗位 CT 影像對比\n# Visualization: Multi-window CT image comparison\n# ---------------------------------------------------------\n\n# 建立一個 1 列 3 行的畫布，設定圖片總大小為 15x5 吋。\n# Create a figure with 1 row and 3 columns, setting the total size to 15x5 inches.\nfig, axes = plt.subplots(1, 3, figsize=(15, 5))\n\n# 走訪我們預先定義的三種窗位設定（軟組織、骨骼、寬骨骼），並將它們分別畫在三個子圖上。\n# Iterate through our predefined window settings (soft tissue, bone, wide bone) and plot them on the three subplots.\nfor ax, (name, (center, width)) in zip(axes, WINDOWS.items()):\n    \n    # 將中間那張切片套用對應的窗位轉換，並以灰階色彩對應圖 (cmap=\"gray\") 顯示，強制數值範圍為 0 到 1。\n    # Apply the window conversion to the middle slice and display it using a grayscale colormap, enforcing the value range from 0 to 1.\n    ax.imshow(window_hu(volume[z], center, width), cmap=\"gray\", vmin=0, vmax=1)\n    \n    # 設定子圖的標題，顯示窗位的名稱以及具體的中心 (WL) 與寬度 (WW) 數值。\n    # Set the subplot title, displaying the window name along with its specific center (WL) and width (WW) values.\n    ax.set_title(f\"{name}\\nWL={center}, WW={width}\")\n    \n    # 關閉子圖的座標軸與刻度線，讓畫面更乾淨。\n    # Turn off the axes and tick marks on the subplot for a cleaner look.\n    ax.axis(\"off\")\n\n# 設定整張大圖的總標題，顯示病患 ID (取前 12 碼) 以及當前顯示的切片進度。\n# Set the main title for the entire figure, displaying the patient ID (first 12 chars) and the current slice progress.\nfig.suptitle(f\"Study {positive_uid[:12]}… | slice {z}/{len(volume)-1}\")\n\n# 自動調整子圖之間的間距，避免文字與圖片互相重疊。\n# Automatically adjust the spacing between subplots to prevent text and images from overlapping.\nplt.tight_layout()\n\n# 將畫好的比較圖存檔至輸出資料夾，設定解析度為 160 DPI，並裁切掉多餘的白邊。\n# Save the plotted comparison figure to the output directory with a resolution of 160 DPI, cropping out extra white space.\nplt.savefig(OUT_DIR / \"ct_window_comparison.png\", dpi=160, bbox_inches=\"tight\")\n\n# 在畫面上渲染並顯示最終圖表。\n# Render and display the final plot on the screen.\nplt.show()\n\n\n# ---------------------------------------------------------\n# 列印物理資訊驗證\n# Print physical information for validation\n# ---------------------------------------------------------\n\n# 印出整份 3D 影像中 Hounsfield Unit (HU) 數值的最小值與最大值，確認單位轉換正確。\n# Print the minimum and maximum Hounsfield Unit (HU) values in the entire 3D volume to verify correct unit conversion.\nprint(\"HU range:\", float(volume.min()), float(volume.max()))\n\n# 印出 3D 影像矩陣的形狀，格式為 [切片數量, 高度, 寬度]。\n# Print the shape of the 3D volume array in the format [number of slices, height, width].\nprint(\"Volume shape [z, y, x]:\", volume.shape)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-29T13:34:35.501525Z","iopub.execute_input":"2026-06-29T13:34:35.502129Z","iopub.status.idle":"2026-06-29T13:34:41.503643Z","shell.execute_reply.started":"2026-06-29T13:34:35.502093Z","shell.execute_reply":"2026-06-29T13:34:41.502852Z"}},"outputs":[],"execution_count":null},{"id":"89d43cfe-a6f8-4158-9d99-65ca1878f2b2","cell_type":"markdown","source":"## 6. Quality-control flags\n\nThe thresholds below are review flags, not reasons to delete data automatically. Unusual geometry may be a real edge case that the final model should handle.\n","metadata":{}},{"id":"dbde3104-a167-4252-af9a-6c0b1c7f17b3","cell_type":"code","source":"# ---------------------------------------------------------\n# 定義品管規則 (Define Quality Control Rules)\n# ---------------------------------------------------------\n\n# 規則 1：檢查影像狀態是否為 \"ok\"\n# Rule 1: Check if the image status is \"ok\"\n# 若之前提取元資料時發現缺少 DICOM 檔案或發生錯誤，status 就會不是 \"ok\"，將其標記為 True。\nmanifest_df[\"qc_missing\"] = manifest_df[\"status\"].ne(\"ok\")\n\n# 規則 2：檢查切片數量是否過少\n# Rule 2: Check if the number of slices is too few\n# 如果切片數量 (n_slices) 為空值 (fillna(0))，或者少於 20 張，代表這份 CT 掃描可能不完整 (例如只掃了頸椎的一小部分)，將其標記為 True。\nmanifest_df[\"qc_few_slices\"] = manifest_df[\"n_slices\"].fillna(0).lt(20)\n\n# 規則 3：檢查切片是否過厚\n# Rule 3: Check if the slice thickness is too large\n# 如果切片厚度大於 5.0 毫米 (mm)，代表解析度太低，難以看清骨折細節，將其標記為 True。\nmanifest_df[\"qc_thick_slice\"] = manifest_df[\"slice_thickness\"].fillna(0).gt(5.0)\n\n# 規則 4：檢查影像長寬是否為標準的 512x512\n# Rule 4: Check if the image matrix is the standard 512x512\n# CT 影像的標準解析度通常是 512x512。如果不是這個尺寸，代表後續需要額外的縮放處理 (Resizing)，將其標記為 True。\nmanifest_df[\"qc_non_512_matrix\"] = ~(\n    manifest_df[\"rows\"].eq(512) & manifest_df[\"columns\"].eq(512)\n)\n\n# ---------------------------------------------------------\n# 統計與顯示品管結果 (Aggregate and Display QC Results)\n# ---------------------------------------------------------\n\n# 找出所有以 \"qc_\" 開頭的欄位名稱 (也就是剛剛建立的這 4 個標記欄位)。\n# Find all column names starting with \"qc_\" (the 4 flag columns just created).\nqc_cols = [c for c in manifest_df if c.startswith(\"qc_\")]\n\n# 將這些 True/False 的欄位加總 (True = 1)，算出每種異常情況各有多少位病患，並將結果輸出成漂亮的表格。\n# Sum these boolean columns (True = 1) to count the number of flagged studies for each condition, and display it as a neat table.\ndisplay(manifest_df[qc_cols].sum().rename(\"flagged_studies\").to_frame())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-29T13:36:10.544486Z","iopub.execute_input":"2026-06-29T13:36:10.544796Z","iopub.status.idle":"2026-06-29T13:36:10.561335Z","shell.execute_reply.started":"2026-06-29T13:36:10.54477Z","shell.execute_reply":"2026-06-29T13:36:10.560485Z"}},"outputs":[],"execution_count":null},{"id":"5c2536b0-870c-49f4-b963-ed61be38c760","cell_type":"markdown","source":"## 7. Export the contract\n\nThe next notebook needs labels, study IDs and the promise that split assignment occurs before slice generation. `FAST_MODE` outputs are intentionally marked so they cannot be mistaken for the full audit.\n","metadata":{}},{"id":"87a95e57-0df0-4946-a86b-7265605b9069","cell_type":"code","source":"# ---------------------------------------------------------\n# 儲存清單與建立資料契約 (Save Manifest & Create Data Contract)\n# ---------------------------------------------------------\n\n# 定義輸出的檔案路徑，分別用於儲存資料表 (CSV) 與設定檔 (JSON)。\n# Define the output file paths for saving the data table (CSV) and the configuration file (JSON).\nmanifest_path = OUT_DIR / \"study_manifest.csv\"\ncontract_path = OUT_DIR / \"data_contract.json\"\n\n# 將整理好的元資料總表 (包含物理參數與品管標記) 儲存為 CSV 檔案，不保留 index。\n# Save the consolidated metadata table (including physical parameters and QC flags) as a CSV file, without the index.\nmanifest_df.to_csv(manifest_path, index=False)\n\n# 建立「資料契約 (Data Contract)」，這是一個包含所有訓練規則與超參數的字典。\n# Create a \"Data Contract\", a dictionary containing all training rules and hyperparameters.\ncontract = {\n    # 標記此資料集的名稱。\n    # Tag the name of this dataset.\n    \"dataset\": \"RSNA 2022 Cervical Spine Fracture Detection\",\n    \n    # 指定用來作為唯一識別碼的欄位名稱 (檢查單號)。\n    # Specify the column name used as the unique identifier (study UID).\n    \"id_column\": \"StudyInstanceUID\",\n    \n    # 寫入我們先前定義的 8 個預測目標 (整體與 C1-C7)。\n    # Write in our previously defined 8 prediction targets (overall and C1-C7).\n    \"targets\": TARGETS,\n    \n    # 明確定義 3D 矩陣的軸向順序為 [深度, 高度, 寬度]，確保模型讀取方向一致。\n    # Explicitly define the 3D matrix axis order as [depth, height, width] to ensure consistent model reading direction.\n    \"volume_axis_order\": [\"z\", \"y\", \"x\"],\n    \n    # 將我們先前設定的 CT 窗位參數 (中心與寬度) 轉換為嵌套字典格式儲存。\n    # Convert our previously set CT window parameters (center and width) into a nested dictionary format for storage.\n    \"windows\": {name: {\"center\": c, \"width\": w} for name, (c, w) in WINDOWS.items()},\n    \n    # 制定嚴格的資料切割規則：必須先以病患 ID 進行 K-Fold 切割，才能進行切片。\n    # Establish a strict data splitting rule: K-Fold split MUST be done by patient ID before generating slices.\n    \"split_rule\": \"Assign StudyInstanceUID to a fold before generating slices or crops.\",\n    \n    # 記錄目前是否處於快速測試模式。\n    # Record whether it is currently in fast testing mode.\n    \"fast_mode\": FAST_MODE,\n    \n    # 統計並記錄成功完成審查的病患總數。\n    # Count and record the total number of studies that successfully completed the audit.\n    \"audited_studies\": int(manifest_df[\"status\"].notna().sum()),\n}\n\n# 將資料契約字典轉換為縮排格式的 JSON 字串，並寫入到指定的檔案路徑中。\n# Convert the data contract dictionary to an indented JSON string and write it to the specified file path.\ncontract_path.write_text(json.dumps(contract, indent=2), encoding=\"utf-8\")\n\n# 在終端機印出檔案儲存成功的確認訊息。\n# Print confirmation messages in the terminal indicating successful file saves.\nprint(\"Saved:\", manifest_path)\nprint(\"Saved:\", contract_path)\n\n# 將 JSON 字典轉換為 Pandas Series 再轉成 DataFrame，以便在畫面上以整齊的表格呈現。\n# Convert the JSON dictionary to a Pandas Series then to a DataFrame to display it as a neat table on the screen.\ndisplay(pd.Series(contract, name=\"value\").to_frame())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-06-29T13:37:21.626344Z","iopub.execute_input":"2026-06-29T13:37:21.627289Z","iopub.status.idle":"2026-06-29T13:37:21.666927Z","shell.execute_reply.started":"2026-06-29T13:37:21.627254Z","shell.execute_reply":"2026-06-29T13:37:21.666048Z"}},"outputs":[],"execution_count":null},{"id":"d4c60c57-d5b2-4f0c-a84b-47fc06aa8c28","cell_type":"markdown","source":"## Takeaways and next notebook\n\n- DICOM filenames are not a geometry model.\n- Stored pixels are not automatically HU.\n- Window parameters are part of model versioning.\n- Metadata outliers deserve review, not silent deletion.\n\n**Next:** build leakage-safe study-level folds and demonstrate why random slice splitting is misleading.\n","metadata":{}}]}